Fusion protein, method, polynucleotide, vector, methods for integrating a transgene, for modifying a cell's genome, for site-specific transposition and for generating an engineered cell, integration cassette and cell.
Patent Information
- Application Number
- BR112025021643
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Publication Date
- 2026-09-01
Smart Images

Figure 00000000_0000_ABST
Description
1 / 112 “FUSION PROTEIN, METHOD, POLYNUCLEOTIDE, VECTOR, METHODS FOR INTEGRATING A TRANSGENE, FOR MODIFYING THE GENOME OF A CELL, FOR SITE-SPECIFIC TRANSPOSITION AND FOR GENERATING AN ENGINEERED CELL, INTEGRATION CASSETTE AND CELL” Cross-reference to related orders
[001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 494,306, filed April 5, 2023, which is incorporated by reference in its entirety herein. Reference to the Electronically Submitted Sequence Listing
[002] This application contains a Sequence Listing that was submitted in XML format through the Patent Center and is incorporated by reference in its entirety herein. The said XML copy, created on March 19, 2024, is named “POTH-083_001WO_SeqList” and is 237,142 bytes in size. Field
[003] This disclosure refers generally to fusion proteins comprising transposase domains and DNA-targeting domains. Methods for using fusion proteins for site-specific transposition are also provided. Background
[004] Transposases can be used to introduce non-endogenous DNA sequences into genomic DNA and are, in many respects, advantageous compared to other gene editing methods. However, there is still an unmet need for site-specific transposases for use, for example, in gene editing.
[005] The lipoprotein (a) gene, or LPA, evolved from a duplication event of the neighboring plasminogen gene (PLG). This event of Petition 870250091181, dated 06 / 10 / 2025, page 97 / 220 2 / 112 duplication occurred during primate evolution, about 40 million years ago. Both genes contain loop structures known as kringle domains. In LPA, the kringle domains were duplicated in a segmented manner so that each copy of LPA can contain up to 50 copies of the kringle domains (Schmidt et al., J Lipid Res. Aug 2016; 57(8):1339-59). At the genomic DNA level, each kringle domain repeat spans about 5.5 kb of DNA, each consisting of two exons and two introns. The intronic portion of the repeats contains several potential target sites for site-specific transposase-mediated integration.
[006] LPA is an attractive target for site-specific transposition because it contains multiple copies of the same target site, increasing the chance of integrating a transposon into at least one. In addition, LPA is highly expressed in hepatocytes, meaning it likely has an open chromosomal landscape amenable to editing and supporting high expression of integrated transgenes. It is a non-essential gene, and knockout is associated with lower cholesterol levels. Combined, these characteristics make it a potential target site for gene therapies. Brief Description
[007] In one aspect, a fusion protein comprising a DNA-targeted domain and a transposase domain comprising the sequence set forth in SEQ ID NO: 4 is provided herein, wherein the DNA-targeted domain binds to a nucleic acid sequence encoding an LPA repeat element. In some embodiments, the DNA-targeted domain comprises one, two, or three Zinc Finger Motifs. In some embodiments, the DNA-targeted domain comprises one or more TAL domains. In some embodiments, the TAL domain comprises the sequence set forth in any of the SEQ ID NOs: 35 to 38. In some embodiments, the DNA-targeted domain Petition 870250091181, dated 06 / 10 / 2025, pp. 98 / 220 3 / 112 binds to a nucleic acid sequence that encodes a kringle domain repeat element or an intron adjacent to a sequence that encodes a kringle domain repeat element in the LPA gene.
[008] In some embodiments, the transposase domain and the DNA-targeting domain are connected by a linker. In some embodiments, the linker comprises the sequence GGGGS (SEQ ID NO: 181).
[009] In some embodiments, the domain that targets DNA is inserted into the N-terminal of the transposase domain at a position after the 82nd amino acid and before the 105th amino acid of SEQ ID NO:4. In some embodiments, the DNA-targeting domain replaces one or more amino acids in the transposase domain between, and including, the 83rd and 105th amino acids of SEQ ID NO: 4. In some embodiments, the transposase domain comprises an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102, or 1-103.
[0010] In some embodiments, the transposase domain comprises the sequence set out in any of the SEQ ID NOs: 7 to 27. In some embodiments, the transposase domain comprises (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K and D201R; or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E and R504D.
[0011] In another aspect, a polynucleotide comprising a nucleic acid sequence encoding a fusion protein described in this document is provided herein.
[0012] In another aspect, a vector comprising a polynucleotide described in this document is provided herein.
[0013] In another aspect, a document is provided Petition 870250091181, dated 06 / 10 / 2025, pp. 99 / 220 4 / 112 method for integrating a transgene into a genomic target site of a cell, the method comprising introducing into the cell a fusion protein described herein and a transposon, wherein the transposon comprises, in the order of 5' to 3': a 5'ITR, the transgene and a 3'ITR. In some embodiments, the transposon further comprises an exogenous promoter between the 5'ITR and the transgene. In some embodiments, the transgene encodes a detectable marker. In some embodiments, the detectable marker is GFP. In some embodiments, the transgene is a gene that (a) is not expressed by the cell prior to the introduction of the fusion protein and the transposon or (b) exhibits diminished, insufficient and / or altered expression by the cell prior to the introduction of the fusion protein and the transposon.
[0014] In some embodiments, the genomic target site is located in the LPA gene. In some embodiments, the genomic target site is located in a repetitive element. In some embodiments, the repetitive element is an LPA repeat element. In some embodiments, the genomic target site is located in an intron of a gene. In some embodiments, the genomic target site is located in the intron of the LPA gene. In some embodiments, the cell is in vivo.
[0015] In another aspect, a method for modifying the genome of a cell is provided herein, the method comprising: providing the cell with a fusion protein described herein, wherein the cell comprises a modified binding site comprising, in the order of 5' to 3', the sequence of a target site for the DNA-targeted domain, a first spacer, a TTAA-targeted integration site for SPB, a second spacer, and the reverse complement of the target site sequence for the DNA-targeted domain. In some embodiments, the target integration site comprises the TTAA sequence. In some embodiments, the target integration site comprises the nucleic acid sequence established in Petition 870250091181, dated 06 / 10 / 2025, pp. 100 / 220 5 / 112 any of the SEQ ID Nos: 81-88
[0016] In another aspect, an integration cassette for site-specific transposition of a nucleic acid into the genome of a cell is provided in this document, comprising a nucleic acid comprising or consisting of a central transposon ITR integration site TTAA sequence flanked by an upstream and a downstream TAL target sequence, wherein each of the upstream and downstream TAL target sequences is separated from the TTAA sequence by 12 or 13 base pairs. In some embodiments, the integration site comprises the TTAA sequence. In some embodiments, the integration site comprises the nucleic acid sequence set in any of the SEQ ID NOs: 81-88. In some embodiments, each of the upstream and downstream TAL target sequences are the same. In some embodiments, each of the upstream and downstream TAL target sequences are different.In some embodiments, each of the upstream and downstream TAL Matrix target sites targets a 7-30 bp sequence of an LPA repeat element.
[0017] In another aspect, a cell comprising an integration cassette described herein stably integrated into the cell genome is provided herein.
[0018] In another aspect, a method for site-specific transposition of a DNA molecule into the genome of a cell is provided herein, comprising the introduction into a cell comprising an integration cassette described herein: a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase; wherein the fusion protein is expressed in the cell; and a DNA molecule comprising a transposon; wherein the expressed fusion protein integrates the transposon by site-specific transposition at the site of Petition 870250091181, dated 06 / 10 / 2025, pp. 101 / 220 6 / 112 TTAA integration of the integrated cassette in a stable manner.
[0019] In another aspect, a method for generating a site-specific transposition engineered cell is provided herein, comprising introducing into a cell comprising an integration cassette described herein: a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase; wherein the fusion protein is expressed in the cell; and a DNA molecule comprising a transposon; wherein the expressed fusion protein integrates the transposon by site-specific transposition into the TTAA integration site of the stably integrated integration cassette, thereby generating the engineered cell. In some embodiments, the sequence is TTAA. In some embodiments, the integration site comprises the nucleic acid sequence set out in any of the SEQ ID Nos: 81-88 Brief description of the Figures
[0020] FIGURE 1A illustrates the introduction of DNA-binding domains into a transposase using obligate heterodimers.
[0021] FIGURE 1B illustrates the introduction of DNA-binding domains into a transposase using obligate heterodimers.
[0022] FIGURE 1C illustrates the introduction of DNA-binding domains into a transposase using obligate heterodimers.
[0023] FIGURE 1D illustrates the introduction of DNA-binding domains into a transposase using obligate heterodimers.
[0024] FIGURE 2 is a diagram showing the Split GFP Splicing SiteSpecific Reporter.
[0025] FIGURE 3 is a diagram showing the ssSPB catalytic dimer bound to an excised transposon and recognizing its genomic integration target site. Petition 870250091181, dated 06 / 10 / 2025, page 102 / 220 7 / 112 Detailed Description
[0026] Fusion proteins comprising transposase domains and DNA-targeted domains are provided in this document. In particular, the DNA-targeted domains can target the lipoprotein A (LPA) gene. Methods for manufacturing the transposase domains and fusion proteins, cells that are modified using the fusion proteins provided in this document, and treatment methods using such cells are also provided.
[0027] In some embodiments, a fusion protein comprising an SPB or PBx domain and a DNA-targeting domain is provided herein. The DNA-targeting domains are described further below. Transposase Domains
[0028] In one aspect, the present document provides fusion proteins comprising one or more transposase domains. In some embodiments, the transposase domain is a piggyBac transposase domain. In some embodiments, the piggyBac transposase domain is a hyperactive piggyBac transposase domain. In preferred embodiments, the transposase domain is a Super piggyBac™ (SPB) transposase domain. Non-limiting examples of SPB transposases are described in detail in U.S. Patent No. 6,218,182; U.S. Patent No. 6,962,810; U.S. Patent No. 8,399,643 and PCT Publication No. WO 2010 / 099296, each of which is incorporated by reference in its entirety herein for examples of transposase domains that may be used in the fusion proteins described herein.
[0029] In some embodiments, the transposase domain is a Super PiggyBac transposase domain (SPB). An SPB comprises one or more hyperactivity mutations compared to the type piggyBac transposase. Petition 870250091181, dated 06 / 10 / 2025, page 103 / 220 8 / 112 wild type. An illustrative wild-type SPB sequence comprising a nuclear localization sequence (NLS) is shown in SEQ ID NO: 1, with the NLS shown in italics, and hyperactive mutations shown in bold. The SPB transposase domain sequence numbering for the purpose of describing deletions and mutations begins at residue 12 of SEQ ID NO: 1.
[0030] MA PKKKRKVGGGGSSLDDEHILSALLQSDDELVGEDS DSEVSDHVSEDDVQSDTEEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLAS NRILTLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPL LCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAV RKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDV FTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKIL MMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDN WFTSIPLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPL TLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMC SVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNL YMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYC TYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 1).
[0031] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 1. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence set forth in SEQ ID NO: 1 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a domain Petition 870250091181, dated 06 / 10 / 2025, page 104 / 220 9 / 112 transposase comprising the amino acid sequence established in SEQ ID NO: 1.
[0032] An illustrative sequence of the wild-type SPB transposase lacking the NLS domain is set forth in SEQ ID NO: 2. The numbering of the SPB transposase domain sequence for the purpose of describing deletions and mutations begins at residue 5 of SEQ ID NO: 2. In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 2. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence set forth in SEQ ID NO: 2 with one, two, three, four, or five conservative substitutions of amino acids.In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence set forth in SEQ ID NO: 2.
[0033] The transposase domains used in the fusion proteins described herein may be isolated from or derived from an insect, vertebrate, crustacean, or urochordate, as described in more detail in PCT Publications No. WO 2019 / 173636 and No. WO 2020 / 051374. Preferably, the SPB transposase domain is isolated from or derived from the insect Trichoplusia ni (GenBank Accession No. AAA87375), Bombyx mori (GenBank Accession No. BAD11135), or Macdunnoughia crassisigna (GenBank Accession No. ABZ85926.1).
[0034] In some realizations, the transposase domain is Petition 870250091181, dated 06 / 10 / 2025, page 105 / 220 10 / 112 Integration-deficient. An integration-deficient transposase domain is a transposase that can excise its corresponding transposon but integrates the excised transposon at a lower frequency than a corresponding wild-type transposase. Examples of integration-deficient transposases are disclosed in U.S. Patent No. 6,218,185; U.S. Patent No. 6,962,810, U.S. Patent No. 8,399,643 and WO 2019 / 173636, each of which is incorporated by reference in its entirety herein for examples of transposase domains that may be used in the fusion proteins described herein. A list of integration-deficient amino acid substitutions is disclosed in U.S. Patent No. 10,041,077, which is incorporated by reference in its entirety herein for examples of mutations that may be introduced into a transposase domain described herein.
[0035] A wild-type SPB can be made integration-deficient by the introduction of mutations, for example, K93A, R372A, K375A, R376A and / or D450N (relative to SEQ ID NO: 2, with numbering starting at residue 5). The introduction of the R372A, K375A, R376A and D450N mutations is believed to render the transposase integration-deficient, but retain the excision function. An illustrative sequence of an integration-deficient transposase domain is PBx, comprising an NLS, as set forth in SEQ ID NO: 3. In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 3.In some embodiments, a fusion protein described in this document comprises a transposase domain that... Petition 870250091181, dated 06 / 10 / 2025, page 106 / 220 11 / 112 comprises the amino acid sequence set forth in SEQ ID NO: 3 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence set forth in SEQ ID NO: 3.
[0036] The sequence of a PBx integration-deficient transposition domain that does not comprise an NLS is set at SEQ ID NO: 4: GGSSLDDEHILSALLQSDDELVGEDSDSEVSDHVSEDDVQSDT EEAFIDEVHEVQPTSSGSEILDEQNVIEQPGSSLASNRILTLPQRTIRGKNKHC WSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVKWT NAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNHMSTDDLFDRSLS MVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNY TPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYL GRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLT IVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSS CDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMAL LYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPT LKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKC KKVICREHNIDMCQSCF (SEQ ID NO: 4).
[0037] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the sequence Petition 870250091181, dated 06 / 10 / 2025, page 107 / 220 12 / 112 amino acids established at SEQ ID NO: 4 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established at SEQ ID NO: 4. Transposase Domains Comprising N-Terminal Deletions
[0038] In some embodiments, transposase domains (e.g., SPB transposase domains or PBx transposase domains) comprising a deletion of a portion of the amino-terminal (also referred to as the N-terminal or N-terminal Domain, or NTD) of the transposase domain are provided herein. SPB transposase domains or PBx transposase domains comprising deletions in the N-terminal Domain were previously described in International Patent Application Publication No. PCT / US2022 / 77549, which is incorporated by reference in its entirety herein for examples of transposase domains that may be used in the fusion proteins described herein.
[0039] Illustrative sequences of an SPB transposase domain with a deletion of amino acids 1-93 from the N-terminal and of a PBx transposase domain with a deletion of amino acids 1-93 from the N-terminal are shown in SEQ ID NOs: 5 and 6, respectively: NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKL FFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDN HMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFTPVR KIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCD SGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSI PLAKNLLQEPYKLTIVGTVRSNKREIPEVLKNSRSRPVGTSMFCFDGPLTLVSY KPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLDQMCSVMT CSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLT Petition 870250091181, dated 06 / 10 / 2025, pp. 108 / 220 13 / 112 SSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPS KIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 5).
[0040] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence established in SEQ ID NO: 5. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 5 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 5.
[0041] NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYD PLCCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMT AVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLREN DVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGI KILMCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITC DNWFTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDG PLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQ MCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMR NLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRT YCTYCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 6).
[0042] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, by Petition 870250091181, dated 06 / 10 / 2025, pp. 109 / 220 14 / 112 less 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence established in SEQ ID NO: 6. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 6 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 6.
[0043] Other illustrative sequences of PBx transposase domains comprising N-terminal deletions are set out in SEQ IDs 7-27 in Table 1.
[0044] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence established in any of the SEQ ID NOs: 7 to 27. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in any of the SEQ ID NOs: 7-27 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in any of the SEQ ID NOs: 7 to 27. Petition 870250091181, dated 06 / 10 / 2025, pp. 110 / 220 15 / 112 Table 1: Illustrative sequences of NTERMINALLY deleted PBx domains Deleção Sequência PBx TLPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCR Delta 83 NIYDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNE N- DEIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRF Terminal DFLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGA HLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMIN GMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTS IPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCF DGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYN QTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIY SHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYL RDNISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANAS CKKCKKVICREHNIDMCQSCF (SEQ ID NO: 7) PBx LPQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNI Delta 84 YDPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNED N- EIYAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFD Terminal FLIRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAH LTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMING MPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIP LAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFD GPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSH NVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRD NISNILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCK KCKKVICREHNIDMCQSCF (SEQ ID NO: 8) PBx PQRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIY Petition 870250091181, dated 06 / 10 / 2025, pp. 111 / 220 16 / 112 Deleção Sequência Delta 85 DPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEI N- YAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFL Terminal IRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLT IDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMP YLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLA KNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGP LTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKG GVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNV SSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNIS NILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKC KKVICREHNIDMCQSCF (SEQ ID NO: 9) PBx QRTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIY Delta 86 DPLLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEI N- YAFFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFL Terminal IRCLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLT IDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMP YLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLA KNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGP LTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKG GVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNIS NILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKC KKVICREHNIDMCQSCF (SEQ ID NO: 10). PBx RTIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDP Delta 87 LLCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYA N- FFGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIR Terminal CLRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTID Petição 870250091181, de 06 / 10 / 2025, pág. 112 / 220 17 / 112 Deleção Sequência EQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPY LGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAK NLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPL TLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKG GVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNV SSKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNIS NILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKC KKVICREHNIDMCQSCF (SEQ ID NO: 11) PBx TIRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPL Delta 88 LCFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAF N- FGILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRC Terminal LRMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDE QLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYL GRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKN LLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLT LVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGG VDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVS SKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISN ILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCK KVICREHNIDMCQSCF (SEQ ID NO: 12) PBx IRGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLL Delta 89CFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFF N- GILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCL Terminal RMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDE QLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYL GRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKN LLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLT Petition 870250091181, dated 06 / 10 / 2025, pp. 113 / 220 18 / 112 Deleção Sequência LVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGG VDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVS SKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISN ILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCK KVICREHNIDMCQSCF (SEQ ID NO: 13) PBx RGKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLL Delta 90 CFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFF N- GILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCL Terminal RMDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDE QLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYL GRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKN LLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLT LVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGG VDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVS SKGEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISN ILPKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCK KVICREHNIDMCQSCF (SEQ ID NO: 14) PBx GKNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLC Delta 91 FKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFG N- ILVMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLR Terminal MDDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGR GTQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLL QEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLV SYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVD TLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSK GEKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNIL Petition 870250091181, dated 06 / 10 / 2025, pp. 114 / 220 19 / 112 Deleção Sequência PKEVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKV ICREHNIDMCQSCF (SEQ ID NO: 15) PBx KNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFK Delta 92 LFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGIL N- VMTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRM Terminal DDKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLL GFRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRG TQTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQ EPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVS YKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDT LNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKG EKVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPK EVPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVIC REHNIDMCQSCF (SEQ ID NO: 16) PBx NKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKL Delta 93 FFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILV N- MTAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMD Terminal DKSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLG FRGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGT QTNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQE PYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTL NQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGE KVQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKE VPGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICR EHNIDMCQSCF (SEQ ID NO: 17) Petition 870250091181, dated 06 / 10 / 2025, pp. 115 / 220 20 / 112 Deleção Sequência PBx KHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLF Delta 94 FTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVM N- TAVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDD Terminal KSIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGF RGRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQ TNGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEP YKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYK PKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLN QMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEK VQSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEV PGTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICRE HNIDMCQSCF (SEQ ID NO: 18) PBx HCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFF Delta 95 TDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMT N- AVRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDK Terminal SIRPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFR GRCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQT NGVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPY KLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKP KPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKV QSRKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVP GTSDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREH NIDMCQSCF (SEQ ID NO: 19) PBx CWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFT Delta 96 DEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTA N- VRKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSI Petition 870250091181, dated 06 / 10 / 2025, pp. 116 / 220 21 / 112 Deleção Sequência Terminal RPTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRG RCPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTN GVPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKL TIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKP AKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMC SVMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQS RKKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGT SDDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNI DMCQSCF (SEQ ID NO: 20) PBx WSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTD Delta 97 EIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAV N- RKDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIR Terminal PTLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGR CPFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNG VPLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLT IVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPA KMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCS VMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSR KKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTS DDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNID MCQSCF (SEQ ID NO: 21) PBxSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEII Delta 98 SEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVR N- KDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRP Terminal TLRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRC PFRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGV PLGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTI Petição 870250091181, de 06 / 10 / 2025, pág. 117 / 220 22 / 112 Deleção Sequência VGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPA KMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCS VMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSR KKFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTS DDSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNID MCQSCF (SEQ ID NO: 22) PBx TSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIIS Delta 99 EIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRK N- DNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPT Terminal LRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCP FRVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVP LGEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIV GTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAK MVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSV MTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRK KFMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSD DSTEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDM CQSCF (SEQ ID NO: 23) PBx SKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISE Delta IVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKD 100 N- NHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTL TerminalRENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPF RVYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPL GEYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVG TVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKM VYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVM TCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKK Petition 870250091181, dated 06 / 10 / 2025, pp. 118 / 220 23 / 112 Deleção Sequência FMRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDD STEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMC QSCF (SEQ ID NO: 24) PBx KSTRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEI Delta VKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDN 101 N- HMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLR Terminal ENDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFR VYIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLG EYYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGT VASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMV YLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMT CSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKF MRNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDS TEEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQ SCF (SEQ ID NO: 25) PBx STRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIV Delta KWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNH 102 N- MSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRE Terminal NDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRV YIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGE YYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVY LLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTC SRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFM RNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDST EEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQS CF (SEQ ID NO: 26) Petição 870250091181, de 06 / 10 / 2025, pág. 119 / 220 24 / 112 Deleção Sequência PBx TRRSRVSALNIVRSQRGPTRMCRNIYDPLLCFKLFFTDEIISEIVK Delta WTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVRKDNH 103 N- MSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRE Terminal NDVFTPVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRV YIPNKPSKYGIKILMMCDSGTKYMINGMPYLGRGTQTNGVPLGE YYVKELSKPVHGSCRNITCDNWFTSIPLAKNLLQEPYKLTIVGTV ASNAREIPEVLKNSRSRPVGTSMFCFDGPLTLVSYKPKPAKMVY LLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCSVMTC SRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFM RNLYMSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDST EEPVMKKRTYCTYCPSKIRRKANASCKKCKKVICREHNIDMCQS CF (SEQ ID NO: 27) Domínios que têm DNA como Alvo
[0045] The transposase and fusion protein domains provided herein may additionally comprise one or more DNA-targeting domains. A DNA-targeting domain may be attached to the C-terminus or N-terminus of the transposase domain or fusion protein. In some embodiments, the DNA-targeting domain is attached to the N-terminus of the transposase domain, for example, a transposase domain comprising an N-terminus deletion. Without being limited to theory, it is believed that the addition of a DNA-targeting domain to a transposase domain enhances the site-specific transposase activity by targeting the transposase fused to the DNA-targeting domain to the target site.In some embodiments, the insertion of a DNA-targeted domain enhances site-specific transposase activity by at least 2-fold, at least 3-fold, at least 4-fold, or at least 5-fold compared to the same transposase domain that does not comprise one. Petition 870250091181, dated 06 / 10 / 2025, pages 120 / 220 25 / 112 domain that targets DNA.
[0046] Any domain that has a known DNA target in the prior art may be used in the context of the transposase domains, fusion proteins, and tandem dimer transposases described herein, including, without limitation, CRISPR, Zinc Finger Motifs, TALE, and transcription factors. In some embodiments, the DNA-targeting domain comprises one, two, or three Zinc Finger Motifs. In some embodiments, the DNA-targeting domain comprises three Zinc Finger Motifs. In some embodiments, the three Zinc Finger Motifs are flanked by GGGGS ligands (SEQ ID NO: 181). In some embodiments, the three Zinc Finger Motifs flanked by GGGGS ligands (SEQ ID NO: 181) cumulatively comprise the sequence established in SEQ ID NO: 28: GGGGSERPYACPVESCDRRFSRSDELTRHIRIHTGQKPFQCRICMRNFSRSD HLTTHIRTHTGEKPFACDICGRKFARSDERKRHTKIHLRQKDGGGGS (SEQ ID NO: 28) or a sequence with at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity to it.
[0047] In one particular embodiment, a fusion protein comprising a transposase domain with an N-terminal deletion, an NLS, and three Zinc Finger Motifs is provided herein. In some embodiments, the NLS comprises or consists of the sequence set forth in SEQ ID NO: 29.
[0048] In some respects, the domain that targets DNA is a TAL array. TALEs (Transcription activator-like effectors) from Xanthomonas commonly contain a 288-amino acid N-terminal, followed by an array with a variable number of ~34 amino acid repeats, followed by a 278-amino acid C-terminal (SEQ ID NO: 30); Petition 870250091181, dated 06 / 10 / 2025, pp. 121 / 220 26 / 112 However, truncated versions have been described in the literature (e.g., see Miller et al., Nat Biotechnol 29, 143-148 (2011)). TALs fused to a FokI nuclease (called TALENs) generally contain N-terminal and C-terminal truncations. For example, the first 152 amino acids of the N-terminal are often removed (called Delta 152; SEQ ID No 31) and the C-terminal is often truncated, leaving 63 amino acids (called +63; SEQ ID NO: 32).
[0049] TALs contain arrays of 34 amino acids repeated a variable number of times. The two amino acids at positions 12 and 13 are varied and determine which nucleotide the TAL repeat will recognize. This feature allows a TAL array to be programmed to bind to a specific DNA sequence. The amino acids NG recognize T, NI recognizes A, NN recognizes G or A, HD recognizes C, NK recognizes G, NS recognizes A, C, G, or T. Other amino acids within the 34-residue repeat can also be varied. For example, position 11 is often changed to an N for repeats that recognize G. In addition, positions 4 and 32 are often varied to reduce the repeatability of the array, but not to determine binding specificity. The number of 34-amino acid repeats in an array determines the length of the DNA sequence recognized (one protein repeat binds one bp of DNA).Furthermore, the last bp is recognized by a half matrix that is 20 amino acids instead of 34.
[0050] Furthermore, the N-terminal domain of TALs (e.g., SEQ ID NO: 31) recognizes and requires a T that is located immediately 5' from the target DNA sequence. N-terminal domain mutations of TALs that no longer require a 5' T have been described in the literature (Lamb et al., Nucleic Acids Res. 2013 Nov;41(21):9779-85). For example, the NT-G mutant requires a 5'G instead of a 5'T (SEQ ID NO: 33) while the NT-βN mutant does not require any specific 5' nucleotide (SEQ ID NO: 34). These mutated sequences Petition 870250091181, dated 06 / 10 / 2025, pp. 122 / 220 27 / 112 N-terminal domains can be used to provide additional sequence options that can be targeted using TAL Arrays.
[0051] In general, each TAL array comprises nine 34-amino acid repeats followed by a 20-amino acid half-repeat. TAL arrays can be synthesized with flanking BsmBI type IIS restriction sites. In one embodiment, individual TAL modules containing 34-amino acid or 20-amino acid half-repeat repeats can be designed and synthesized flanked by BsmBI type IIS restriction sites. The complete set of TAL modules contains 4 modules capable of recognizing A, C, G, and T for each of the 10 bp positions (40 modules / 10 bp target) and one TAL half-repeat module. Illustrative TAL modules are set out in SEQ ID NOs: 35-38, where X is any amino acid: • TAL Module Version 1: LTPDQVVAIAXXXGGKQALETVQRLLPVLCQDHG (SEQ ID NO: 35) • TAL Module Version2: LTPEQVVAIAXXXGGKQALETVQRLLPVLCQAHG (SEQ ID NO: 36) • TAL Module Version 3 LTPDQVVAIAXXXGGKQALETVQRLLPVLCQAHG (SEQ ID NO: 37) • TAL Module Version4: LTPAQVVAIAXXXGGKQALETVQRLLPVLCQDHG (SEQ ID NO: 38).
[0052] An exemplary half-module TAL is set out in SEQ ID NO: 39, where X is any amino acid: LTPEQVVAIAXXXGGRPALE (SEQ ID NO: 39).
[0053] Pairs of sequences that have TAL arrays as targets in the desired gene can be designed and corresponding modules selected and grouped using Golden Gate Assembly, to assemble each TAL array in-frame. The DNA sequence encoding TAL arrays generated in this document can be further optimized in codons using Petition 870250091181, dated 06 / 10 / 2025, pp. 123 / 220 28 / 112 GeneArt algorithms (Thermo Fisher).
[0054] By designing left- and right-handed TAL arrays comprising an N-terminal domain that recognizes a C-terminal T domain and TAL being fused to an N-terminal deleted transposase sequence (i.e., TAL-ssSPB or TAL-PBx; described below), one TAL array recognizes a 5' sequence of TTAA and the other TAL array recognizes a 3' sequence of TTAA. Because the 5' sequence of TTAA is more often different from the 3' sequence of TTAA in genomic DNA targets, TAL-ssSPB will most often be used as a heterodimer consisting of two different TAL domains that recognize two different DNA sequences. Additionally, the sequence recognized by the TAL array is not directly adjacent to TTAA. Instead, it is separated from TTAA by a spacer of a specified bp length, for example, 12 bp, 13 bp, or 14 bp spacers.
[0055] A TAL array can target any DNA sequence (e.g., genomic DNA sequence) of interest. It will be evident to a person skilled in the art that any left-hand TAL array for a given target can be combined with any right-hand TAL array for the same target.
[0056] In some embodiments, a TAL array targets green fluorescent protein (GFP). TAL-piggyBac transposase fusion proteins comprising deleted N-terminal piggyBac transposase sequences and defective N-terminal piggyBac transposase that target GFP were described in Co-owned International Patent Application Publication No. PCT / 2022 / 22549.
[0057] In some embodiments, a TAL array targets an LPA gene repeat element. Illustrative sequences of TAL arrays targeting an LPA repeat element on the left are Petition 870250091181, dated 06 / 10 / 2025, pp. 124 / 220 29 / 112 established in SEQ ID NOs: 116, 118, 121, 124, 125, 127, 129, 131, 133, 135, 137, 139, and 141. Illustrative sequences of TAL arrays targeting LPA on the right are established in SEQ ID NOs: 117, 119, 120, 122, 123, 126, 128, 130, 132, 134, 136, 138, 140, and 142. In some embodiments, the TAL array targeting a left-sided LPA repeat element binds to a nucleic acid molecule comprising the sequence established in SEQ ID NOs: 89, 91, 94, 97, 98, 100, 102, 104, 106, 108, 110, 112, and 114. In some embodiments, the TAL array targeting a right-handed LPA repeat element binds to a nucleic acid molecule comprising the sequence set forth in SEQ ID NOs: 90, 92, 93, 95, 96, 99, 101, 103, 105, 107, 109, 111, 113, and 115. It will be evident to one skilled in the art that any left-handed TAL array disclosed in this document can be combined with any right-handed TAL array disclosed in this document.Illustrative genomic target sites for an LPA repeat element are established in SEQ ID Nos. 81 to 88.
[0058] The present disclosure provides fusion proteins comprising a DNA-targeted domain attached to the transposase domain in different ways. In some embodiments, the DNA-targeted domain may be fused to or attached to the N-terminus of a transposase domain comprising an N-terminal deletion. For example, the DNA-targeted domain may be inserted into a transposase domain at a suitable position in the N-terminal region of the transposase domain. In some embodiments, the DNA-targeted domain may replace one or more amino acids in the N-terminal region of the transposase domain. In some embodiments, the DNA-targeted domain is inserted into a transposase domain at a suitable position in the N-terminal region of the transposase domain without replacing an amino acid.
[0059] The domain that targets DNA can be inserted into Petition 870250091181, dated 06 / 10 / 2025, pp. 125 / 220 30 / 112 N-terminal of a transposase domain. For example, the DNA-targeting domain is inserted into the N-terminus of the transposase domain at a position after the 82nd amino acid and before the 105th amino acid of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 82nd and 83rd amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 83rd and 84th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 84th and 85th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. amino acid, respectively) or SEQ ID NO: 4.In some embodiments, the DNA-targeting domain is inserted between the 85th and 86th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 86th and 87th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 87th and 88th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 88th and 89th amino acids of SEQ ID NO: 2 or 3. (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4.In some embodiments, the DNA-targeting domain is inserted between the 89th and 90th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 90th and 91st amino acids of... Petition 870250091181, dated 06 / 10 / 2025, pp. 126 / 220 31 / 112 SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 91st and 92nd amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 92nd and 93rd amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 93rd and 94th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 94th and 95th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 95th and 96th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 96th and 97th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 97th and 98th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 98th and 99th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 99th and 100th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid). Petition 870250091181, dated 06 / 10 / 2025, pp. 127 / 220 In some embodiments, the DNA-targeting domain is inserted between the 100th and 101st amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 101st and 102nd amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 102nd and 103rd amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 100th and 101st amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain is inserted between the 101st and 102nd amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. It is inserted between the 103rd and 104th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4.In some embodiments, the DNA-targeting domain is inserted between the 104th and 105th amino acids of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain comprises the sequence of SEQ ID NO: 28 or a sequence with at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity to it. The transposase domain may additionally comprise an NLS, for example, and the NLS of SEQ ID NO: 29.
[0060] The DNA-targeting domain can replace one or more amino acids in the N-terminal region of the transposase domain. For example, the DNA-targeting domain can replace one or more amino acids in the transposase domain between, and including, the 83rd and 105th amino acids of SEQ ID NO: 2 or 3 (with numbering starting at the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain replaces the 83rd amino acid of SEQ ID NO: 2 or 3 (with Petition 870250091181, dated 06 / 10 / 2025, pp. 128 / 220 33 / 112 (numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 84th amino acid of SEQ ID NO: 2 or 3 (numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 85th amino acid of SEQ ID NO: 2 or 3 (numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 86th amino acid of SEQ ID NO: 2 or 3 (numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 87th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or SEQ ID NO: 4.In some embodiments, the DNA-targeted domain replaces the 88th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 89th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 90th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 91st amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some realizations, the domain that targets DNA replaces the 92nd amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4.In some embodiments, the DNA-targeted domain replaces the 93rd amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 3rd amino acid). Petition 870250091181, dated 06 / 10 / 2025, pp. 129 / 220 34 / 112 In some embodiments, the DNA-targeted domain replaces the 94th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 95th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 96th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 97th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. 12th amino acid, respectively) or SEQ ID NO: 4.In some embodiments, the DNA-targeting domain replaces the 98th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain replaces the 99th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain replaces the 100th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain replaces the 101st amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4.In some embodiments, the DNA-targeting domain replaces the 102nd amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeting domain replaces the 103rd amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or. Petition 870250091181, dated 06 / 10 / 2025, pp. 130 / 220 35 / 112 of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 104th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain replaces the 105th amino acid of SEQ ID NO: 2 or 3 (with numbering starting from the 5th or 12th amino acid, respectively) or of SEQ ID NO: 4. In some embodiments, the DNA-targeted domain comprises the sequence of SEQ ID NO: 28 or a sequence with at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity therewith. The transposase domain may additionally comprise an NLS, for example, an NLS with SEQ ID NO: 29.
[0061] An illustrative sequence of a fusion protein comprising a transposase domain comprising a 93-amino acid N-terminal deletion, an NLS, and three Zinc Finger Motifs flanked by GGGGS ligands (SEQ ID NO: 181) is shown in SEQ ID NO: 40, where the NLS is shown in italics, the sequence comprising the three Zinc Finger Motifs and the GGGGS ligands is underlined, and the transposase domain comprising a 93-amino acid N-terminal deletion is shown in bold: MA PKKKRKVGGGGSERPYACPVESCDRRFSRSDELTRHIRIHT GQKPFQCRICMRNFSRSDHLTTHIRTHTGEKPFACDICGRKFARSDERKRHTK IHLRQKDGGGGSNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLL CFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVR KDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFT PVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILM MCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNW FTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPTLL VSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCS Petition 870250091181, dated 06 / 10 / 2025, pp. 131 / 220 36 / 112 VMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLY MSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCT YCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 40)
[0062] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence established in SEQ ID NO: 40. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 40 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 40.An illustrative sequence of a fusion protein comprising an integration-deficient transposase domain comprising a 93-amino acid N-terminal deletion, an NLS, and three GGGGS-flanked Zinc Finger Motifs (SEQ ID NO: 181) ligands is set out in SEQ ID NO: 180, where the NLS is shown in italics, the sequence comprising the three Zinc Finger Motifs and the GGGGS ligands is underlined, and the transposase domain comprising a 93-amino acid N-terminal deletion is shown in bold. MA PKKKRKVGGGGSERPYACPVESCDRRFSRSDELTRHIRIHT GQKPFQÇRJÇMRNFSRSDHLTTHIRTHTGEKPFAÇDIÇGRKFARSDERKRHTK IHLRQKDGGGGSNKHCWSTSKSTRRSRVSALNIVRSQRGPTRMCRNIYDPLL CFKLFFTDEIISEIVKWTNAEISLKRRESMTSATFRDTNEDEIYAFFGILVMTAVR KDNHMSTDDLFDRSLSMVYVSVMSRDRFDFLIRCLRMDDKSIRPTLRENDVFT Petition 870250091181, dated 06 / 10 / 2025, pp. 132 / 220 37 / 112 PVRKIWDLFIHQCIQNYTPGAHLTIDEQLLGFRGRCPFRVYIPNKPSKYGIKILM MCDSGTKYMINGMPYLGRGTQTNGVPLGEYYVKELSKPVHGSCRNITCDNW FTSIPLAKNLLQEPYKLTIVGTVASNAREIPEVLKNSRSRPVGTSMFCFDGPLTL VSYKPKPAKMVYLLSSCDEDASINESTGKPQMVMYYNQTKGGVDTLNQMCS VMTCSRKTNRWPMALLYGMINIACINSFIIYSHNVSSKGEKVQSRKKFMRNLY MSLTSSFMRKRLEAPTLKRYLRDNISNILPKEVPGTSDDSTEEPVMKKRTYCT YCPSKIRRKANASCKKCKKVICREHNIDMCQSCF (SEQ ID NO: 180).
[0063] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to the amino acid sequence established in SEQ ID NO: 180. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 180 with one, two, three, four, or five conservative amino acid substitutions. In some embodiments, a fusion protein described herein comprises a transposase domain comprising the amino acid sequence established in SEQ ID NO: 180. Nuclear Location Signals
[0064] In some embodiments, the transposase domains and fusion proteins provided herein may comprise an in-frame nuclear localization sequence (NLS). Examples of transposases fused to a nuclear localization signal are disclosed in U.S. Patent No. 6,218,185; U.S. Patent No. 6,962,810, U.S. Patent No. 8,399,643 and WO 2019 / 173636, each of which is incorporated by reference in its entirety herein for examples of domains. Petition 870250091181, dated 06 / 10 / 2025, pp. 133 / 220 38 / 112 transposases that can be used in the fusion proteins described in this document. In some embodiments, the NLS comprises the PKKKRKV sequence (SEQ ID NO: 29). In certain aspects, the in-frame NLS is located upstream (N-terminal) of the transposase domain, which comprises an N-terminal deletion.
[0065] In general, the NLS is preferentially located at the N-terminal end of a fusion protein. In some embodiments, the NLS is fused or linked to the N-terminal of a transposase domain. In some embodiments, the NLS is fused or linked to the N-terminal of a DNA-targeted domain.
[0066] In certain aspects, the in-frame NLS is fused directly to the amino terminus of the transposase domain comprising an N-terminal deletion. In some embodiments, the NLS is attached to the N-terminus of a transposase domain comprising an N-terminal deletion by means of a ligand (e.g., a GGGGS ligand or a GGS ligand).
[0067] In some embodiments, a methionine primer is introduced before the NLS. In some embodiments, additional alanine residues are introduced before and / or after the NLS to ensure in-frame translation. As such, residue numbering in SEQ ID NOs: 1 and 3 begins at residue 12 of SEQ ID NOs: 1 and 3 for the purpose of identifying deleted and mutated residues. In SEQ ID NO: 2, which is the SPB sequence, which does not comprise an NLS, residue numbering begins at residue 5 for the purpose of identifying deleted and mutated residues. In SEQ ID NO: 4, numbering begins at residue 1 for the purpose of identifying deleted and mutated residues.
[0068] In some embodiments, a fusion protein comprises an NLS and a transposase domain that includes a 93-amino acid N-terminal deletion. In some embodiments, the fusion protein Petition 870250091181, dated 06 / 10 / 2025, pp. 134 / 220 39 / 112 comprises an amino acid sequence that is at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identical to the amino acid sequence established in SEQ ID NO: 5. In some embodiments, the fusion protein comprises the amino acid sequence established in SEQ ID NO: 5. Obligate Heterodimers and Tandem Dimers
[0069] In another aspect, tandem dimer transposases comprising two fusion proteins are provided in this document, each fusion protein comprising a transposase domain and one or both fusion proteins additionally comprising a DNA-targeting domain. In some embodiments, both fusion proteins comprise a DNA-targeting domain. In some embodiments, both fusion proteins comprise DNA-targeting domains, and the DNA-targeting domains target DNA sequences that are adjacent to the DNA sequence that is the transposase-targeted insertion site. In some embodiments, only one of the two fusion proteins in the tandem dimer transposase comprises a DNA-targeting domain. A DNA-targeting domain may be attached to the C-terminus or the N-terminus of the fusion protein.
[0070] Thus, in some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a first transposase domain and a first DNA-targeting domain; and (b) a second fusion protein comprising a first transposase domain and a second DNA-targeting domain, wherein the first DNA-targeting domain and the second DNA-targeting domain are different; wherein the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a Petition 870250091181, dated 06 / 10 / 2025, pp. 135 / 220 40 / 112 complex.
[0071] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising, in N-terminal to C-terminal order: a first NLS, a first DNA-targeting domain, and a first transposase domain comprising an N-terminal deletion; and (b) a second fusion protein comprising, in N-terminal to C-terminal order: a second NLS, a second DNA-targeting domain, and a second transposase domain comprising an N-terminal deletion; wherein the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex. In some embodiments, the first and / or second transposase domains are SPB domains. In some embodiments, the first and / or second transposase domains are PBx transposase domains.In some embodiments, the first and / or second transposase domain comprises an N-terminal deletion of 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, or 103 amino acids. In some embodiments, the first and / or second transposase domains comprise the sequence with SEQ ID NO: 5 or 6. In some embodiments, the first and / or second DNA-targeting domain comprises one, two, or three Zinc Finger Motifs. In some embodiments, the first and / or second DNA-targeting domain comprises the sequence with SEQ ID NO: 28. In some embodiments, the first and / or second DNA-targeting domain comprises TAL motifs.
[0072] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising, in N-terminal to C-terminal order: a first NLS and a first transposase domain comprising the SEQ ID sequence NO: 2, 3 or 4; and (b) a second fusion protein comprising, in N-terminal order Petition 870250091181, dated 06 / 10 / 2025, pp. 136 / 220 41 / 112 for C-terminal: a second NLS domain and a second transposase domain comprising the sequence SEQ ID NO: 2, 3, or 4; wherein the first and second transposase domains comprise a DNA-targeting domain, and wherein the transposase domain of the first fusion protein and the transposase domain of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex. In some embodiments, the first and / or second DNA-targeting domain comprises one, two, or three Zinc Finger Motifs. In some embodiments, the first and / or second DNA-targeting domain comprises the sequence SEQ ID NO: 28. In some embodiments, the first and / or second DNA-targeting domain comprises TAL motifs.In some embodiments, the first DNA-targeting domain replaces one or more amino acids between, and including, the 83rd and 105th amino acids of the first transposase domain, with numbering starting at residue 5 or 12 of SEQ ID NO: 2 or 3, respectively. In some embodiments, the first DNA-targeting domain replaces the 83rd, 84th, 85th, 86th, 87th, 88th, 89th, 90th, 91st, 92nd, 93rd, 94th, 95th, 96th, 97th, 98th, 99th, 100th, 101st, 102nd, or 103rd residue of the first transposase domain, with numbering starting at residue 5 or 12 of SEQ ID NO: 2 or 3, respectively. In some embodiments, the first DNA-targeting domain replaces one or more amino acids between, and including, the 83rd and 105th amino acids of the second transposase domain, with numbering starting at residue 5 or 12 of SEQ ID NO: 2 or 3, respectively.In some embodiments, the second domain that targets DNA replaces the 83rd, 84th, 85th, 86th, 87th, 88th, 89th, 90th, 91st, 92nd, 93rd, 94th, 95th, 96th, 97th, 98th, 99th, 100th, 101st, 102nd, or 103rd residue of the second transposase domain, with numbering starting at residue 5 or 12 of SEQ ID NO: 2 or 3, respectively.
[0073] In another aspect, the following are provided in this document Petition 870250091181, dated 06 / 10 / 2025, pp. 137 / 220 42 / 112 fusion proteins comprising a transposase domain that can form obligate heterodimers with another fusion protein comprising a transposase domain. Without being limited to theory, it is believed that two such fusion proteins assemble into a dimer structure held together by a combination of charge interactions, hydrogen bonds, pic-cation pairs, and hydrophobic interactions. Thus, each obligate heterodimer complex comprises two transposase domains. In some embodiments, two fusion proteins provided herein form a complex, said complex comprising (a) a first fusion protein comprising a transposase domain and (b) a second fusion protein comprising a transposase domain; wherein the transposase domains of the first fusion protein and the transposase domains of the second fusion protein have opposite charges that allow the two fusion proteins to form a complex.In non-limiting examples, the assembled complex may be a single dimer (2 protein molecules) or a dimer of dimers (4 protein molecules or a tetramer).
[0074] By introducing charged residues in the amino acids that contribute to dimerization with a second fusion protein, it is possible to design pairs of fusion proteins that can only associate with each other in a tandem dimer in a predetermined configuration. By introducing mutations that allow only one configuration of the tandem dimer, it becomes feasible to introduce DNA-targeted domains into the fusion proteins, thus increasing the specificity of the transposase domains. This is illustrated in FIGURES 1A and 1B for SPB and FIGURES 1C and 1D for PBx: The introduction of DNA-targeted domains into fusion proteins that can dimerize in any configuration, including homodimerization, would lead to the presence of four DNA-targeted domains in a tandem dimer transposase. However, only two DNA-targeted domains Petition 870250091181, dated 06 / 10 / 2025, pp. 138 / 220 43 / 112 would interact with DNA, leaving the other two potentially sterically hindering the transposase-DNA interaction. Any suitable DNA-targeting domain described herein or known in the art may be used in the fusion proteins described herein.
[0075] Mutations in transposase domains that confer a positive or negative charge can be determined by a person skilled in the art. In the case of a fusion protein comprising a first and a second transposase domain, the crystal structure published in Chen et al. (Nat Commun 11,3446 (2020)) can be used to identify residue pairs in the transposase domains that are in close proximity in the tandem dimer formed by two such fusion proteins. Altering the charge of these residue pairs to create a positively charged transposase domain and a negatively charged transposase domain can be performed using standard techniques such as site-directed mutagenesis.
[0076] For example, one or more of M185, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593, and / or F594 can be mutated into an SPB transposase domain (for example, the SPB established in SEQ ID NO: 1 or 2, with numbering starting at the 12th residue of SEQ ID NO: 1 and the 5th residue of SEQ ID NO: 2) to generate an SPB- or SPB+ transposase domain. Similarly, one or more of M185, R189, K190, D191, H193, M194, D198, D201, S203, L204, S205, V207, K500, R504, K575, K576, R583, N586, I587, D588, M589, C593, and / or F594 can be mutated into a PBx transposase domain (for example, the PBx transposase domain of SEQ ID NO: 3 with numbering starting at the 12th residue of SEQ ID NO: 3, or the PBx transposase domain of SEQ ID NO: 4) to generate a PBx-(minus) or PBx+(plus) transposase domain.
[0077] In some embodiments, a fusion protein described Petition 870250091181, dated 06 / 10 / 2025, pp. 139 / 220 44 / 112 in this document may comprise (i) an SPB+ transposase domain or (ii) an SPB- transposase domain.
[0078] To achieve the formation of an obligate heterodimer, mutation pairs can be introduced into fusion proteins or transposase domains to generate positively and negatively charged fusion proteins or transposase domains, which can then interact to form a heterodimer. In some embodiments, the residue pair to be mutated is one of those set out in Table 2. For example, one or more of the mutations listed in the column labeled Protein 1 can be introduced into a first SPB or PBx domain, and the corresponding mutation(s) listed in the column labeled Protein 2 can be introduced into a second SPB or PBx domain. In some embodiments, the members of a residue pair are mutated to have opposite charges. Table 2: Illustrative Residue Pairs; numbering starts at residue SEQ ID NO: 2 or at residue 12 of SEQ ID NO: 1 or 3. Protein 1 Protein 2 Protein 1 Protein 2 Protein 1 Protein 2 M185 L204 D201 R504 R583 D588 R189 R189 S203 R504 N586 D588 R189 D191 L204 R189 I587 R583 R189 M194 L204 L204 I587 I587 R189 L204 L204 S205 D588 I587 K190 K190 L204 R504 D588 D588 K190 H193 S205 L204 D588 M589 K190 M194 V207 S203 M589 M589 D191 R189 V207 L204 M589 F594 H193 K190 K500 D198 C593 M589 M194 R189 R504 D201 F594 K575 M194 K190 K575 F594 F594 K576 D198 K500 K576 F594 F594 M589
[0079] To introduce a positive charge, amino acids with uncharged side chains, such as methionine, or amino acids with negatively charged side chains, such as aspartic acid, can be replaced by positively charged amino acids, such as lysine or Petition 870250091181, dated 06 / 10 / 2025, pp. 140 / 220 45 / 112 arginine. To introduce a negative charge, amino acids with positively charged side chains, such as arginine or lysine, or amino acids with hydrophobic side chains, such as leucine, can be replaced by negatively charged amino acids, such as aspartic acid or glutamic acid.
[0080] In certain embodiments, one or more of the following mutations is / are introduced into an SPB transposase domain (e.g., the SPB established at SEQ ID NO: 1 or 2, with numbering starting at the 12th residue of SEQ ID NO: 1 and the 5th residue of SEQ ID NO: 2) of a fusion protein provided herein to generate an SPB+ fusion protein: M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, an SPB+ transposase domain comprises an M185R mutation and a D198K mutation. In some embodiments, an SPB+ transposase domain comprises an M185R mutation and a D201R mutation. In some embodiments, an SPB+ transposase domain comprises a D197K mutation and a D201R mutation. In some embodiments, an SPB+ transposase domain comprises a D198K mutation and a D201R mutation. In some embodiments, an SPB+ transposase domain comprises an M185R mutation, a D198K mutation, and a D201R mutation.
[0081] In certain embodiments, one or more of the following mutations is / are introduced into a PBx transposase domain (for example, the PBx transposase domain of SEQ ID NO: 3 with numbering starting at the 12th residue of SEQ ID NO: 3; or the PBx transposase domain of SEQ ID NO: 4) of a fusion protein provided herein to generate a PBx+ fusion protein: M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R. In some embodiments, a PBx+ transposase domain comprises an M185R mutation and a D198K mutation. In some embodiments, a PBx+ transposase domain comprises an M185R mutation. Petition 870250091181, dated 06 / 10 / 2025, pp. 141 / 220 46 / 112 and a D201R mutation. In some embodiments, a PBx+ transposase domain comprises a D197K mutation and a D201R mutation. In some embodiments, a PBx+ transposase domain comprises a D198K mutation and a D201R mutation. In some embodiments, a PBx+ transposase domain comprises an M185R mutation, a D198K mutation, and a D201R mutation.
[0082] In certain embodiments, one or more of the following mutations is / are introduced into an SPB transposase domain (e.g., the SPB established at SEQ ID NO: 1 or 2, with numbering starting at the 12th residue of SEQ ID NO: 1 and the 5th residue of SEQ ID NO: 2) of a fusion protein provided herein to generate an SPB fusion protein: L204D, L204E, K500D, K500E, R504E, and R504D. In some embodiments, an SPB transposase domain comprises an L204E mutation and a K500D mutation. In some embodiments, an SPB transposase domain comprises an L204E mutation and an R504D mutation. In some embodiments, an SPB transposase domain comprises a K500 mutation and an R504D mutation. In some embodiments, an SPB- transposase domain comprises an L204E mutation, a K500D mutation, and an R504D mutation.
[0083] In certain embodiments, one or more of the following mutations is / are introduced into a PBx transposase (e.g., the PBx transposase domain of SEQ ID NO: 3 with numbering starting at the 12th residue of SEQ ID NO: 3 or the PBx-transposase domain of SEQ ID NO: 4) of a fusion protein provided herein to generate a PBx-fusion protein: L204D, L204E, K500D, K500E, R504E, and R504D. In some embodiments, a PBx-transposase domain comprises an L204E mutation and a K500D mutation. In some embodiments, a PBx-transposase domain comprises an L204E mutation and an R504D mutation. In some embodiments, a PBx-transposase domain comprises a K500 mutation and Petition 870250091181, dated 06 / 10 / 2025, pp. 142 / 220 47 / 112 an R504D mutation. In some embodiments, a PBx transposase domain comprises an L204E mutation, a K500D mutation, and an R504D mutation.
[0084] Illustrative sequences of SPB+ transposase domains are set forth in SEQ ID NOs: 42-54. Illustrative sequences of SPB-transposase domains are set forth in SEQ ID NOs: 55-64. In some embodiments, a transposase domain provided herein comprises the amino acid sequence set forth in any of the SEQ ID NOs: 42-64. In some embodiments, a transposase domain provided herein comprises the amino acid sequence set forth in any of the SEQ ID NOs: 42-64 additionally comprising one or more conservative amino acid sequences.
[0085] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence set at any of the SEQ ID NOs: 42-54. In some embodiments, the transposase domain comprises an amino acid sequence set at any of the SEQ ID NOs: 42-54 additionally comprising one or more conservative amino acid sequences.
[0086] In some embodiments, a fusion protein described herein comprises a transposase domain comprising an amino acid sequence set at any of the SEQ ID NOs: 55-64. In some embodiments, the transposase domain comprises an amino acid sequence set at any of the SEQ ID NOs: 55-64 additionally comprising one or more conservative amino acid sequences.
[0087] In some embodiments, a complex is provided herein comprising (a) a first fusion protein comprising a transposase domain comprising the sequence of Petition 870250091181, dated 06 / 10 / 2025, pp. 143 / 220 (a) a fusion protein comprising a transposase domain comprising the amino acid sequence established in any of the SEQ ID NOs: 42-54; and (b) a second fusion protein comprising a transposase domain comprising the amino acid sequence established in any of the SEQ ID NOs: 55-64.
[0088] The SPB+, SPB-, PBx+, and PBx- fusion proteins and transposase domains may additionally comprise the N-terminal deletions of the transposase domain described herein. Thus, in some embodiments, an SPB+ fusion protein comprising a transposase domain comprising an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 100 amino acids, or about 115 amino acids is provided herein. In some embodiments, the transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 85 amino acids.In some embodiments, the transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 93 amino acids. In. Petition 870250091181, dated 06 / 10 / 2025, pp. 144 / 220 49 / 112 In some embodiments, the transposase domain comprises an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 98 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 99 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 103 amino acids.
[0089] In some embodiments provided herein is an SPB- fusion protein comprising a transposase domain with an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 82 amino acids, about 83 amino acids, about 85 amino acids, about 86 amino acids, about 88 amino acids, about 89 amino acids, about 91 amino acids, about 92 amino acids, about 94 amino acids, about 95 amino acids, about 97 amino acids, about 98 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids, or about 115 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 83 amino acids. In some realizations, the transposase domain Petition 870250091181, dated 06 / 10 / 2025, pp. 145 / 220 50 / 112 comprises an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 93 amino acids.In some embodiments, the transposase domain comprises an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 98 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 99 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 103 amino acids. Petition 870250091181, dated 06 / 10 / 2025, pp. 146 / 220 51 / 112
[0090] In some embodiments, a PBx+ fusion protein is provided herein comprising a transposase domain comprising an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 100 amino acids, or about 115 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 83 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 88 amino acids.In some embodiments, the transposase domain comprises an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 93 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 98 amino acids.In some realizations, the transposase domain. Petition 870250091181, dated 06 / 10 / 2025, pp. 147 / 220 52 / 112 comprises an N-terminal deletion of 99 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 103 amino acids.
[0091] In some embodiments, a PBx- fusion protein comprising a transposase domain comprising an N-terminal deletion of about 20 amino acids, about 40 amino acids, about 60 amino acids, about 80 amino acids, about 81 amino acids, about 82 amino acids, about 83 amino acids, about 84 amino acids, about 85 amino acids, about 86 amino acids, about 87 amino acids, about 88 amino acids, about 89 amino acids, about 90 amino acids, about 91 amino acids, about 92 amino acids, about 93 amino acids, about 94 amino acids, about 95 amino acids, about 96 amino acids, about 97 amino acids, about 98 amino acids, about 99 amino acids, about 100 amino acids, about 101 amino acids, about 102 amino acids, about 103 amino acids or about 115 amino acids is provided herein. In some embodiments, the transposase domain comprises an N-terminal deletion of 83 amino acids.In some embodiments, the transposase domain comprises an N-terminal deletion of 84 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 85 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 86 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 87 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 88 amino acids. In some embodiments, the domain... Petition 870250091181, dated 06 / 10 / 2025, pages 148 / 220 53 / 112 transposase comprises an N-terminal deletion of 89 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 90 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 91 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 92 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 93 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 94 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 95 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 96 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 97 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 98 amino acids.In some embodiments, the transposase domain comprises an N-terminal deletion of 99 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 100 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 101 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 102 amino acids. In some embodiments, the transposase domain comprises an N-terminal deletion of 103 amino acids. Integration Cassettes
[0092] Integration cassettes for site-specific transposition of a DNA molecule into the genome of a cell are also provided in this document. In some embodiments, the integration cassette comprises an integration site of the TTAA sequence. In some embodiments, the integration cassette for site-specific transposition of a nucleic acid into the genome of a cell comprises a nucleic acid that Petition 870250091181, dated 06 / 10 / 2025, pp. 149 / 220 54 / 112 comprises or consists of a central transposon ITR integration site with the CTTAAA sequence flanked by an upstream and a downstream TAL matrix target sequence, wherein each of the upstream and downstream TAL matrix target sequences is separated from the CTTAAA sequence by 12 or 13 base pairs. In some embodiments, each of the at least one upstream and downstream TAL matrix target site sequences are the same. In some embodiments, each of the at least one upstream and downstream TAL matrix target site sequences are different. In some embodiments, each of the at least one upstream and downstream TAL matrix target site sequences targets a 10 bp sequence of an LPA repeat element.
[0093] Methods are also provided for site-specific transposition of a DNA molecule into the genome of a cell comprising a stably integrated integration cassette, comprising introducing into the cell: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase; wherein the fusion protein is expressed in the cell, and b) a DNA molecule comprising a transposon; wherein the expressed fusion protein integrates the transposon by site-specific transposition into the CTTAAA sequence of the stably integrated integration cassette.
[0094] Methods are also provided for generating a site-specific transposition engineered cell, comprising: introducing into a cell comprising a stably integrated integration cassette: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase; wherein the fusion protein is expressed in the cell, and b) a DNA molecule comprising a transposon; wherein the fusion protein expressed is integrated into the Petition 870250091181, dated 06 / 10 / 2025, pages 150 / 220 55 / 112 transposon by site-specific transposition in the CTTAAA sequence of the stably integrated integration cassette, thus generating the engineered cell. Nucleic Acids
[0095] Also provided herein are polynucleotides comprising nucleic acid sequences encoding the fusion proteins described herein. In some embodiments, the polynucleotides are isolated.
[0096] The polynucleotides isolated from the disclosure can be produced using (a) recombinant methods, (b) synthetic techniques, (c) purification techniques and / or (d) combinations thereof, as well known in the art.
[0097] The methods of constructing nucleic acids encoding transposase domains comprising an N-terminal deletion described herein are well known in the art or described herein, for example, PCR-based mutagenesis.
[0098] The fusion of the present invention can be generated using any suitable method known in the art or described in this document.
[0099] The polynucleotides isolated from this disclosure, such as RNA, cDNA, genomic DNA, or any combination thereof, may be obtained from biological sources using any number of cloning methodologies known to those skilled in the art. In some respects, oligonucleotide probes that selectively hybridize, under stringent conditions, with the polynucleotides of this disclosure are used to identify the desired sequence in a cDNA or genomic DNA library.
[00100] RNA or DNA amplification methods are well known in the art and can be used in accordance with disclosure without Petition 870250091181, dated 06 / 10 / 2025, pp. 151 / 220 56 / 112 improper experimentation, based on the teaching and guidance presented in this document. Known methods of DNA or RNA amplification include, without limitation, polymerase chain reaction (PCR) and related amplification processes (see, for example, U.S. Patents Nos. 4,683,195, 4,683,202, 4,800,159, 4,965,188, by Mullis et al.; 4,795,699 and 4,921,794 by Tabor et al.; 5,142,033 by Innis; 5,122,464 by Wilson et al.; 5,091,310 by Innis; 5,066,584 by Gyllensten et al.; 4,889,818 by Gelfand et al.; 4,656,134 by Ringold) and RNA-mediated amplification that uses antisense RNA for the target sequence as a template for synthesis. of double-stranded DNA (U.S. Patent No. 5,130,238 to Malek et al., trade name NASBA), the full contents of which references are incorporated by reference herein. (See, for example, Ausubel, supra; or Sambrook, supra.)
[00101] For example, polymerase chain reaction (PCR) technology can be used to amplify disseminated polynucleotide sequences and related genes directly from genomic DNA or cDNA libraries. PCR and other in vitro amplification methods can also be useful, for example, for cloning nucleic acid sequences encoding proteins to be expressed, for producing nucleic acids to be used as probes to detect the presence of desired mRNA in samples, for nucleic acid sequencing, or for other purposes. Examples of techniques sufficient to guide those skilled in the art through in vitro amplification methods are found in Berger, supra, Sambrook, supra, and Ausubel, supra, as well as Mullis, et al., U.S. Patent No. 4,683,202 (1987); and Innis, et al., PCR Protocols: A Guide to Methods and Applications, Eds., Academic Press Inc., San Diego, Calif. (1990).Commercially available kits for genomic amplification by PCR are known in the art. See, for example, Advantage-GC Genomic PCR Kit. Petition 870250091181, dated 06 / 10 / 2025, pp. 152 / 220 57 / 112 (Clontech). Additionally, for example, the T4 32 gene protein (Boehringer Mannheim) can be used to enhance the yield of long PCR products.
[00102] Disclosure polynucleotides can also be prepared by direct chemical synthesis using known methods (see, for example, Ausubel, et al., supra). Chemical synthesis generally produces a single-stranded oligonucleotide, which can be converted into double-stranded DNA by hybridization with a complementary sequence or by polymerization with a DNA polymerase using the single strand as a template. A person skilled in the art will recognize that, although the chemical synthesis of DNA may be limited to sequences of about 100 or more bases, longer sequences can be obtained by linking shorter sequences. Expression Vectors and Host Cells
[00103] Disclosure also refers to vectors that include disclosure polynucleotides, host cells that are genetically engineered with recombinant vectors, and the production of at least one protein scaffold by recombinant techniques, as is well known in the art. See, for example, Sambrook, et al., supra; Ausubel, et al., supra, both fully incorporated herein by reference.
[00104] Polynucleotides can optionally be linked to a vector containing a selectable marker for propagation into a host. Generally, a plasmid vector is introduced into a precipitate, such as a calcium phosphate precipitate, or into a complex with a charged lipid. If the vector is a virus, it can be packaged in vitro using an appropriate packaging cell line and then transduced into host cells.
[00105] The DNA insert can be operatively ligated to an appropriate promoter. In some embodiments, the promoter is a promoter Petition 870250091181, dated 06 / 10 / 2025, pp. 153 / 220 58 / 112 EF-Iα. The expression constructs will additionally contain transcription initiation and termination sites and, in the transcribed region, a ribosome binding site for translation. The coding portion of the mature transcripts expressed by the constructs will preferentially include a translation initiation codon at the beginning (e.g., ATG) and a termination codon (e.g., UAA, UGA, or UAG) appropriately positioned at the end of the mRNA to be translated, with UAA and UAG being preferred for expression in mammalian or eukaryotic cells.
[00106] Expression vectors can include at least one selectable marker. These markers include, for example, but not limited to, ampicillin, zeocin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / Geneticin (neo gene), DHFR (which encodes dihydrofolate reductase and confers methotrexate resistance), mycophenolic acid or glutamine synthetase (GS, U.S. Patents Nos. 5,122,464; 5,770,359; 5,827,739), blasticidin (bsd gene), resistance genes for eukaryotic cell culture, as well as ampicillin, zeocin (Sh bla gene), puromycin (pac gene), hygromycin B (hygB gene), G418 / Geneticin (neo gene), kanamycin, spectinomycin, streptomycin, carbenicillin resistance genes, bleomycin, erythromycin, polymyxin B, or tetracycline for culture. in E. coli and other bacteria or prokaryotes (the above patents are incorporated by reference in their entirety herein).The appropriate culture media and conditions for the host cells described above are known in the art. Suitable vectors will be readily apparent to those skilled in the art. The introduction of a construct vector into a host cell can be effected by transfection with calcium phosphate, DEAE-dextran mediated transfection, cationic lipid-mediated transfection, electroporation, transduction, infection, or other known methods. Such methods are described in the art, as in Sambrook, supra, Chapters 1-4 and 16-18; Ausubel, supra, Chapters 1, 9, 13, 15. Petition 870250091181, dated 06 / 10 / 2025, pp. 154 / 220 59 / 112 16.
[00107] Expression vectors may include at least one selectable cell surface marker for isolating cells modified by disclosure compositions and methods. Selectable disclosure cell surface markers comprise surface proteins, glycoproteins, or protein clusters that distinguish one cell or subset of cells from another defined subset of cells. Preferably, the selectable cell surface marker distinguishes cells modified by a disclosure composition or method from cells that are not modified by a disclosure composition or method. These cell surface markers include, for example, but not limited to, designation or classification determinant cluster proteins (often abbreviated as CD), such as a truncated or full-length form of CD19, CD271, CD34, CD22, CD20, CD33, CD52, or any combination thereof.Cell surface markers additionally include the suicide gene marker RQR8 (Philip B et al. Blood. 2014 August 21; 124(8):1277-87).
[00108] Expression vectors may include at least one selectable drug resistance marker for isolation from cells modified by the disclosure compositions and methods. Selectable drug resistance markers for disclosure may include DHFR, TYMS, FRANCF, RAD51C, GCS, MDR1, ALDH1, wild-type or mutant NKX2.2, or any combination thereof.
[00109] Those skilled in the art are knowledgeable about the various expression systems available for the expression of a nucleic acid encoding a dissemination protein. Alternatively, dissemination nucleic acids can be expressed in a host cell by activation (through manipulation) in a host cell containing endogenous DNA that Petition 870250091181, dated 06 / 10 / 2025, pages 155 / 220 60 / 112 encodes a protein dissemination support. Such methods are well known in the art, for example, as described in U.S. Patents Nos. 5,580,734, 5,641,670, 5,733,746 and 5,733,761, which are incorporated herein by reference.
[00110] Illustrative cell cultures useful for the production of protein scaffolds, specified portions or variants thereof, are bacterial, yeast and mammalian cells, as known in the art. Mammalian cell systems are frequently presented in the form of cell monolayers, although mammalian cell suspensions or bioreactors may also be used. Several suitable host cell lines, capable of expressing intact glycosylated proteins, have been developed using this technique, including COS-1 (e.g., ATCC CRL 1650), COS-7 (e.g., ATCC CRL-1651), HEK293, BHK21 (e.g., ATCC CRL-10), CHO (e.g., ATCC CRL 1610), and BSC-1 (e.g., ATCC CRL-26) cell lines, Cos-7 cells, CHO cells, Hep G2 cells, P3X63Ag8653, SP2 / 0-Ag14, 293, HeLa cells, and similar cells, which are readily available, for example, from the American Type Culture Collection, Manassas, Virginia (www.atcc.org).Preferred host cells include cells of lymphoid origin, such as myeloma and lymphoma cells. Particularly preferred host cells are P3X63Ag8.653 cells (ATCC Accession Number CRL-1580) and SP2 / O-Ag14 cells (ATCC Accession Number CRL-1851). In a preferred aspect, the recombinant cell is either a P3X63Ab8.653 cell or an SP2 / O-Ag14 cell.
[00111] The expression vectors for these cells may include one or more of the following expression control sequences, such as, without limitation, an origin of replication; a promoter (e.g., late or early SV40 promoters, the CMV promoter (U.S. Patents Nos. 5,168,062; 5,385,839), an HSV tk promoter, a pgk promoter (phosphoglycerate)). Petition 870250091181, dated 06 / 10 / 2025, pp. 156 / 220 61 / 112 kinase), an EF-1 alpha promoter (U.S. Patent No. 5,266,491), at least one human promoter; an enhancer and / or information processing sites, such as ribosome binding sites, RNA splice sites, polyadenylation sites (e.g., a large SV40 poly AT Ag addition site) and transcriptional terminator sequences. Consult, for example, Ausubel et al., supra; Sambrook, et al., supra. Other cells useful for the production of nucleic acids or proteins of the present disclosure are known and / or available, for example, from the American Type Culture Collection Cell Line and Hybridoma Catalogue (www.atcc.org) or other known or commercial sources.
[00112] When eukaryotic host cells are employed, polyadenylation or transcription terminator sequences are commonly incorporated into the vector. An example of a terminator sequence is the polyadenylation sequence from the bovine growth hormone gene. In some embodiments, the polyA sequence is a polyA SV40 sequence.
[00113] Sequences for precise transcript splicing may also be included. An example of a splicing sequence is the SV40 VP1 intron (Sprague et al., J. Virol. 45:773-781 (1983)). Additionally, gene sequences to control replication in the host cell may be incorporated into the vector, as known in the art.
[00114] The plasmid constructs described herein can be used to deliver nucleic acids encoding the transposase domains or fusion proteins described herein to a cell.
[00115] The transposase domains and fusion proteins described in this document can also be delivered to a cell using mRNA constructs. Therefore, in one embodiment, it is provided in the present Petition 870250091181, dated 06 / 10 / 2025, pp. 157 / 220 62 / 112 document a sequence of mRNA encoding a transposase domain or a fusion protein described in this document. These mRNA sequences can be delivered to a cell using a nanoparticle, for example, a lipid nanoparticle. Examples of lipid nanoparticles are described, for example, in International Patent Applications No. PCT / US2021 / 055876, No. PCT / US2022 / 017570, U.S. Provisional Application No. 63 / 397,268, U.S. Provisional Application No. 63 / 301,855 and U.S. Provisional Application No. 63 / 348,614, each of which is incorporated by reference in its entirety herein for examples of lipid nanoparticles that can be used to deliver mRNA constructs encoding the fusion proteins or transposase domains described herein. An mRNA construct can also be delivered to a cell by electroporation or nucleofection. The mRNA can be encapsulated or otherwise modified. CELLS AND MODIFIED CELLS
[00116] The transposases and fusion proteins described herein can be used in conjunction with a transposon to modify cells. The transposon can be a piggyBac™ (PB) transposon. In some embodiments, when the transposon is a PB transposon, the transposase is a piggyBac™ (PB) transposase, a piggyBac-type (PBL) transposase, or a Super piggyBac™ (SPB) transposase. Non-limiting examples of PB transposons are described in detail in U.S. Patent No. 6,218,182; U.S. Patent No. 6,962,810; U.S. Patent No. 8,399,643 and PCT Publication No. WO 2010 / 099296, each of which is incorporated by reference in its entirety herein by way of example of transposons that may be used in conjunction with the transposases and fusion proteins described herein. Transposons may comprise a nucleic acid encoding a therapeutic protein or a therapeutic agent.Examples of therapeutic proteins include those disclosed in the Publications. Petition 870250091181, dated 06 / 10 / 2025, pp. 158 / 220 63 / 112 PCT No. WO 2019 / 173636 and No. WO 2020 / 051374, each of which is incorporated by reference in its entirety herein as examples of therapeutic proteins that can be encoded by a transposon used in conjunction with the transposases and fusion proteins described herein.
[00117] Therefore, modified cells comprising one or more transposons and one or more tandem dimer transposases or fusion proteins described herein are provided in this document. The cells and modified cells of the disclosure may be mammalian cells. Preferably, the cells and modified cells are human cells.
[00118] A cell modified using a site-specific transposase fusion protein described in this document may be a germline cell or a somatic cell. Disclosure cells and modified cells may be immune cells, for example, lymphoid progenitor cells, natural killer (NK) cells, T lymphocytes (T cells), memory stem cells (Tscm cells), central memory T cells (Tcm cells), stem cell-like T cells, B lymphocytes (B cells), antigen-presenting cells (APCs), cytokine-induced killer (CIK) cells, myeloid progenitor cells, neutrophils, basophils, eosinophils, monocytes, macrophages, platelets, erythrocytes, red blood cells (RBCs), megakaryocytes, or osteoclasts. The modified cell may be differentiated, undifferentiated, or immortalized. The undifferentiated modified cell may be a stem cell. The modified undifferentiated cell may be an induced pluripotent stem cell.The modified cell can be a T cell, a hematopoietic stem cell, a natural killer cell, a macrophage, a dendritic cell, a monocyte, a megakaryocyte, or an osteoclast. The modified cell can be modified while quiescent, in an activated state, at rest, or in... Petition 870250091181, dated 06 / 10 / 2025, pp. 159 / 220 64 / 112 interphase, prophase, metaphase, anaphase, or telophase. The modified cell may be fresh, cryopreserved, bulk, subpopulated, from whole blood, leukapheresis, or an immortalized cell line. A detailed description for isolating cells from a leukapheresis or blood product is provided in PCT Publications No. WO 2019 / 173636 and WO 2020 / 051374, each of which is incorporated by reference in its entirety herein.
[00119] The disclosure methods may modify and / or produce a population of modified T cells, in which at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% or any percentage among the plurality of modified T cells in the population express one or more cell surface markers of a memory stem cell (Tscm) or a Tscm-like cell; and wherein one or more cell surface markers comprise(s) CD45RA and CD62L. The cell surface markers may comprise one or more of the following: CD62L, CD45RA, CD28, CCR7, CD127, CD45RO, CD95, CD95 and IL-2Rp.Cell surface markers may include one or more of the following: CD45RA, CD95, IL-2RP, CCR7, and CD62L.
[00120] The disclosure provides methods for expressing a CAR on the surface of a cell. The method comprises (a) obtaining a cell population; (b) bringing the cell population into contact with a composition comprising a CAR or a sequence encoding the CAR, under conditions sufficient to transfer the CAR across the cell membrane of at least one cell in the cell population, thereby generating a cell population Petition 870250091181, dated 06 / 10 / 2025, pages 160 / 220 65 / 112 modified; (c) cultivate the modified cell population under conditions suitable for integration of the sequence encoding the CAR; and (d) expand and / or select at least one cell from the modified cell population that expresses the CAR on the cell surface. A more detailed description of the methods for expressing a CAR on the surface of a cell is disclosed in PCT Publications No. WO 2019 / 049816 and WO 2020 / 051374, each of which is incorporated by reference in its entirety herein.
[00121] This disclosure provides a cell or a population of cells wherein the cell comprises a composition comprising (a) an inducible transgene construct comprising a sequence encoding an inducible promoter and a sequence encoding a transgene, and (b) a receptor construct comprising a sequence encoding a constitutive promoter and a sequence encoding an exogenous receptor, such as a CAR, wherein, after integration of the construct of (a) and the construct of (b) into a genomic sequence of a cell, the exogenous receptor is expressed, and wherein the exogenous receptor, upon binding to a ligand or antigen, transduces an intracellular signal that directly or indirectly targets the inducible promoter regulating the expression of the inducible transgene (a) to modify gene expression.
[00122] The disclosure further provides a composition comprising the modified, expanded and selected cell population from the methods described herein.
[00123] Disclosure-modified cells (e.g., CAR-T cells) can be further modified to enhance their therapeutic potential. Alternatively, or additionally, the modified cells can be further modified to make them less sensitive to immunological and / or metabolic checkpoints, for example, by blocking and / or diluting specific checkpoint signals delivered to the cells (by Petition 870250091181, dated 06 / 10 / 2025, pages 161 / 220 66 / 112 example, checkpoint inhibition) naturally, within the immunosuppressive tumor microenvironment.
[00124] Disclosure-modified cells (e.g., CAR-T cells) may be further modified to silence or reduce the expression of (i) one or more genes encoding inhibitory checkpoint signaling receptors; (ii) one or more genes encoding intracellular proteins involved in checkpoint signaling; (iii) one or more genes encoding a transcription factor that hinders the effectiveness of a therapy; (iv) one or more genes encoding a cell death or apoptosis receptor; (v) one or more genes encoding a metabolic sensing protein; (vi) one or more genes encoding proteins that confer sensitivity to a cancer therapy, including a monoclonal antibody; and / or (vii) one or more genes encoding a growth advantage factor.Non-limiting examples of genes that can be modified to silence or reduce expression or to repress a gene function include, without limitation, exemplary inhibitory checkpoint signals, intracellular proteins, transcription factors, cell death or apoptosis receptors, metabolic sensing proteins, proteins that confer sensitivity to cancer therapy, and growth advantage factors, as disclosed in PCT Publication No. WO 2019 / 173636.
[00125] Modified disclosure cells (e.g., CAR-T cells) can be further modified to express a modified / chimeric checkpoint receptor. The modified / chimeric checkpoint receptor may comprise a null receptor, a decoy receptor, or a negative-dominant receptor. Examples of null, decoy, or negative-dominant intracellular receptors / proteins include, but are not limited to, downstream signaling components of an inhibitory checkpoint signal, a transcription factor, a cytokine, or a Petition 870250091181, dated 06 / 10 / 2025, pp. 162 / 220 67 / 112 cytokine receptor, a chemokine or chemokine receptor, a cell death or apoptosis receptor / ligand, a metabolic sensor molecule, a protein that confers sensitivity to cancer therapy, and an oncogene or tumor suppressor gene. Non-limiting examples of cytokines, cytokine receptors, chemokines, and chemokine receptors are disclosed in PCT Publication WO 2019 / 173636.
[00126] Genome modification may involve the introduction of a nucleic acid sequence, transgene, and / or a genome editing construct into a cell ex vivo, in vivo, in vitro, or in situ to stably integrate a nucleic acid sequence, transiently integrate a nucleic acid sequence, produce site-specific integration of a nucleic acid sequence, or produce biased integration of a nucleic acid sequence. The nucleic acid sequence may be a transgene.
[00127] Stable chromosome integration can be random integration, site-specific integration, or skewed integration. Without wishing to be bound by theory, it is believed that the addition of DNA-binding domains to the tandem dimer transposases described in this paper improves the site specificity of the transposases.
[00128] Site-specific integration can occur at a safe harbor site. Genomic safe harbor sites are capable of accommodating the integration of new genetic material in a way that ensures the newly inserted genetic elements function reliably (e.g., are expressed at a therapeutically effective expression level) and do not cause deleterious alterations to the host genome that pose a risk to the host organism. Non-limiting examples of potential genomic safe harbors include intronic sequences of the human albumin gene, adeno-associated virus site 1 (AAVS1), a site of Petition 870250091181, dated 06 / 10 / 2025, pp. 163 / 220 68 / 112 integration of naturally occurring AAV virus on chromosome 19, the chemokine receptor 5 (CC motif) gene site (CCR5) and the mouse Rosa26 human ortholog site.
[00129] Site-specific transgene integration can occur at a site that interrupts the expression of a target gene. Interruption of target gene expression can occur through site-specific integration at introns, exons, promoters, genetic elements, enhancers, suppressors, start codons, stop codons, and response elements. Non-limiting examples of target genes that are targets of site-specific integration include TRAC, TRAB, PDI, any gene encoding an immunosuppressive protein, and genes encoding proteins involved in allorejection.
[00130] Site-specific transgene integration can occur at a site that results in enhanced expression of a target gene. Enhanced expression of the target gene can occur through site-specific integration at introns, exons, promoters, genetic elements, enhancers, suppressors, start codons, stop codons, and response elements.
[00131] The site-specific transgene integration site may be an unstable chromosomal insertion. The unstable integration may be a transient nonchromosomal integration, a semi-stable nonchromosomal integration, a semi-persistent nonchromosomal insertion, or an unstable chromosomal insertion. The transient nonchromosomal insertion may be epichromosomal or cytoplasmic. In one aspect, the transient nonchromosomal insertion of a transgene does not integrate into a chromosome, and the modified genetic material is not replicated during cell division.
[00132] The site-specific transgene integration site may be a modified binding site for the DNA-targeted domain in a transposon, fusion protein, or tandem dimer domain described in Petition 870250091181, dated 06 / 10 / 2025, pp. 164 / 220 69 / 112 of this document. For example, the TTAA target DNA integration site for SPB can be modified to insert flanking DNA binding sites for the target DNA domain comprising three Zinc Finger Motifs (e.g., a target DNA domain comprising or consisting of the sequence SEQ ID NO: 28 or a sequence with at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% identity therewith). For example, a target DNA domain comprising three Zinc Finger Motifs is believed to bind to the DNA sequence GCGTGGGCG. Therefore, it is believed that the introduction of two copies of the GCGTGGGCG sequence flanking the TTAA target integration site for SPB enhances the site-specific integration of an SPB transposase domain comprising a DNA-targeted domain comprising three Zinc Finger Motifs.In some embodiments, the two copies of the GCGTGGGCG sequence are in reverse (5') and complementary (3') orientation.
[00133] In some embodiments, a polynucleotide is provided herein comprising, in the order of 5' to 3', the reverse complement of the sequence of a target site for a DNA-targeted domain, a first spacer, the TTAA-target integration site for SPB, a second spacer, and the sequence of the target site for a DNA-targeted domain. In some embodiments, the first spacer and the second spacer are the same length. In some embodiments, the first and / or the second spacer are 3 bp long. In some embodiments, the first and / or the second spacer are 4 bp long. In some embodiments, the first and / or the second spacer are 5 bp long. In some embodiments, the first and / or the second spacer are 6 bp long. In some embodiments, the first and / or the second spacer are 7 bp long. In some embodiments, the first and / or the second spacer are 7 bp long. Petition 870250091181, dated 06 / 10 / 2025, pages 165 / 220 70 / 112 spacers are 8 bp long. In some embodiments, the first and / or second spacer is 9 bp long. In some embodiments, the first and / or second spacer is 10 bp long.
[00134] The modified target site can be introduced into a cell or cell line to facilitate the genomic engineering that is targeted. For example, a cell line that has been engineered to comprise a modified target site for an SPB or a PBx provided herein can be transfected with said SPB or PBx as well as a transposon comprising donor DNA, so that the donor DNA is inserted into the modified target site. In some embodiments, the cell line is a T cell line. In some embodiments, the modified target sequence is introduced into a highly expressed genomic region. In some embodiments, the cell is an in vitro cell, for example, a cell in cell culture.
[00135] For DNA-binding domains comprising TALs, the target site is determined by the TAL sequence. A person skilled in the art will be able to modify the TAL sequences to achieve the desired target specificity.
[00136] Genome modification can be an unstable chromosomal integration of a transgene. The integrated transgene can be silenced, removed, excised, or further modified.
[00137] In some embodiments, the transposase domains, fusion proteins, and tandem dimer complexes provided herein have better transposase efficacy than their wild-type equivalents. Transposase activity can be measured by any suitable assay known in the art or described herein, for example, a Split GFP assay. For example, the transposase domains, fusion proteins, and tandem dimer complexes provided herein may have genome integration activity in the target comparable to Petition 870250091181, dated 06 / 10 / 2025, pp. 166 / 220 71 / 112 of their wild-type counterparts, but have reduced off-target genome integration activity compared to their wild-type equivalents.
[00138] In some embodiments, a transposase domain and a DNA-targeting domain provided in this document have an on-target to off-target activity ratio that is increased by at least 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500-fold, at least about 550-fold, at least about 600-fold, at least about 650-fold, at least about 700-fold, at least about 750-fold, at least about 800-fold, at least about 850-fold, at least about 900-fold, at least about 950-fold, or at least about 1000-fold compared with the unmodified SPB transposase.
[00139] In some embodiments, a transposase domain comprising a DNA-targeting domain inserted into the N-terminal region of the transposase domain provided herein has an on-target to off-target activity ratio that is increased by at least 50-fold, at least about 100-fold, at least about 150-fold, at least about 200-fold, at least about 250-fold, at least about 300-fold, at least about 350-fold, at least about 400-fold, at least about 450-fold, at least about 500-fold, at least about 550-fold, at least about 600-fold, at least about 650-fold, at least about 700-fold, at least about 750-fold, at least about 800-fold, at least about 850-fold, at least about 900-fold, at least about 950-fold, or at least about 1000 times greater compared to the wild-type transposase domain. Petition 870250091181, dated 06 / 10 / 2025, pp. 167 / 220 72 / 112
[00140] In certain embodiments, modified cells are used therapeutically in adoptive cell therapy.
[00141] Adoptive cell compositions that are universally safe for administration to any patient (not just the patient from whom they are derived) require a significant reduction or elimination of alloreactivity. To this end, the cells of the present disclosure (e.g., allogeneic cells) can be modified to disrupt the expression or function of a T-cell receptor (TCR) and / or a class of Major Histocompatibility Complex (MHC). The TCR mediates graft-versus-host (GvH) reactions, while the MHC mediates host-versus-graft (HvG) reactions. Preferably, any TCR expression and / or function is eliminated to prevent T-cell-mediated GvH that could cause the death of the individual. Thus, in a preferred aspect, the present disclosure provides a composition of pure, TCR-negative allogeneic T cells (e.g., each cell in the composition expresses itself at such a low level that it becomes undetectable or nonexistent).
[00142] The expression and / or function of MHC class I (MHC-I, specifically HLA-A, HLA-B, and HLA-C) is reduced or eliminated to prevent HvG and consequently enhance cell engraftment in an individual. Enhanced engraftment results in greater cell persistence and therefore a wider therapeutic window for the individual. Specifically, the expression and / or function of a structural element of MHC-I, Beta-2 Microglobulin (B2M), is reduced or eliminated. Non-limiting examples of guide RNAs (gRNAs) to target and delete MHC activators are disclosed in PCT Application No. PCT / US2019 / 049816.
[00143] A detailed description of unnaturally occurring chimeric stimulatory receptors, genetic modifications of endogenous sequences encoding TCR-alpha (TCR-α), TCR-beta (TCR-β) and / or Beta-2 Petition 870250091181, dated 06 / 10 / 2025, pp. 168 / 220 73 / 112 Microglobulin (β2M), and unnaturally occurring polypeptides comprising an HLA class I histocompatibility antigen, alpha E chain (HLA-E polypeptide) is disclosed in Application Publication PCT No. WO 2020 / 051374, which is incorporated by reference in its entirety herein.
[00144] Under normal conditions, full T cell activation depends on the involvement of the TCR in conjunction with a second signal mediated by one or more co-stimulatory receptors (e.g., CD28, CD2, 4-1BBL) that potentiate the immune response. However, when the TCR is not present, T cell expansion is severely reduced when stimulated with standard activation / stimulation reagents, including the anti-CD3 agonist mAb. Therefore, the present disclosure provides a non-naturally occurring chimeric stimulatory receptor (CSR) comprising: (a) an ectodomain comprising an activation component, wherein the activation component is isolated from or derived from a first protein; (b) a transmembrane domain; and (c) an endodomain comprising at least one signal transduction domain, wherein the at least one signal transduction domain is isolated from or derived from a second protein; wherein the first protein and the second protein are not identical.
[00145] The activation component may comprise a portion of one or more components of a T Cell Receptor (TCR), a component of a TCR complex, a component of a TCR co-receptor, a component of a TCR co-stimulatory protein, a component of a TCR inhibitory protein, a cytokine receptor, and a chemokine receptor to which an agonist of the activation component binds. The activation component may comprise an extracellular CD2 domain or a portion thereof to which an agonist binds.
[00146] The signal transduction domain may include a Petition 870250091181, dated 06 / 10 / 2025, pp. 169 / 220 74 / 112 or more components of a human signal transduction domain, a T cell receptor (TCR), a component of a TCR complex, a component of a TCR co-receptor, a component of a TCR co-stimulatory protein, a component of a TCR inhibitory protein, a cytokine receptor, and a chemokine receptor. The signal transduction domain may comprise a CD3 protein or a portion thereof. The CD3 protein may comprise a CD3Z protein or a portion thereof.
[00147] The endodomain may additionally comprise a cytoplasmic domain. The cytoplasmic domain may be isolated from or derived from a third protein. The first and third proteins may be identical. The ectodomain may additionally comprise a signal peptide. The signal peptide may be derived from a fourth protein. The first and fourth proteins may be identical. The transmembrane domain may be isolated from or derived from a fifth protein. The first and fifth proteins may be identical.
[00148] This disclosure also provides a non-naturally occurring chimeric stimulatory receptor (CSR) in which the ectodomain comprises a modification. The modification may comprise a mutation or truncation of the amino acid sequence of the activating component or the first protein, compared to a wild-type sequence of the activating component or the first protein. The mutation or truncation of the amino acid sequence of the activating component may comprise a mutation or truncation of an extracellular CD2 domain or a portion thereof to which an agonist binds. The mutation or truncation of the extracellular CD2 domain may reduce or eliminate binding to naturally occurring CD58.
[00149] This disclosure provides a nucleic acid sequence encoding any CSR disclosed herein. This disclosure also provides a transposon or vector that Petition 870250091181, dated 06 / 10 / 2025, pages 170 / 220 75 / 112 comprises a nucleic acid sequence that encodes any CSR disclosed in this document.
[00150] This disclosure provides a cell comprising any CSR disclosed herein. This disclosure provides a cell comprising a nucleic acid sequence encoding any CSR disclosed herein. This disclosure provides a cell comprising a vector containing a nucleic acid sequence encoding any CSR disclosed herein. This disclosure provides a cell comprising a transposon comprising a nucleic acid sequence encoding any CSR disclosed herein.
[00151] This disclosure provides a composition comprising any CSR disclosed herein. This disclosure provides a composition comprising a nucleic acid sequence encoding any CSR disclosed herein. This disclosure provides a composition comprising a vector comprising a nucleic acid sequence encoding any CSR disclosed herein. This disclosure provides a composition comprising a transposon comprising a nucleic acid sequence encoding any CSR disclosed herein. This disclosure provides a composition comprising a modified cell disclosed herein or a composition comprising a plurality of modified cells disclosed herein.
[00152] This document also provides site-specific gene integration methods. The transposase domains and fusion proteins provided herein can be used to deliver a transgene to a cell and integrate the transgene at a target site. The Petition 870250091181, dated 06 / 10 / 2025, pp. 171 / 220 76 / 112 A target site can be, for example, a genomic safe harbor, that is, a genomic site where a transgene can be integrated in a way that ensures the transgene functions predictably and does not cause alterations in the host's genomic DNA sequence. In some embodiments, the target site is a repetitive element, such as an LPA sequence. There may be one, two, or more target sites within a repetitive element. In some embodiments, the target site is located within an intron (e.g., an intron of the LPA gene).
[00153] Site-specific integration can be used in vitro or in vivo. An example of an in vivo application is gene therapy, which involves delivering a transgene to the genomic DNA of a cell. Formulations, Dosages and Methods of Administration
[00154] This disclosure provides formulations, dosages, and methods for administering the compositions and cells described herein. In one aspect, a pharmaceutical composition comprising a tandem dimer transposase or fusion protein described herein and a pharmaceutically acceptable carrier is provided herein. In another aspect, a pharmaceutical composition comprising a modified cell described herein and a pharmaceutically acceptable carrier is provided herein.
[00155] The disclosed pharmaceutical compositions and compounds may comprise at least one of any suitable excipients, such as, without limitation, diluent, binder, stabilizer, buffers, salts, lipophilic solvents, preservative, adjuvant or the like. Pharmaceutically acceptable excipients are preferred. Non-limiting examples and methods for preparing such sterile solutions are well known in the art, such as, without limitation, Gennaro, Ed., Remington's Pharmaceutical Sciences, 18th edition, Mack Publishing Co. (Easton, Pa.) 1990 and in the “Physician's Desk”. Petition 870250091181, dated 06 / 10 / 2025, pp. 172 / 220 77 / 112 Reference”, 52nd ed., Medical Economics (Montvale, NJ) 1998. Pharmaceutically acceptable vehicles may be routinely selected that are suitable for the mode of administration, solubility and / or stability of the protein carrier, fragment composition or variants, as well known in the art or as described herein.
[00156] Non-limiting examples of pharmaceutical excipients and additives suitable for use include proteins, peptides, amino acids, lipids and carbohydrates (e.g., sugars, including monosaccharides, di-, tri-, tetra- and oligosaccharides; derivatized sugars, such as alditols, aldonic acids, esterified sugars and the like; and polysaccharides or sugar polymers), which may be present alone or in combination, comprising, alone or in combination, from 1% to 99.99% by weight or volume. Non-limiting examples of protein excipients include serum albumin, such as human serum albumin (HSA), recombinant human albumin (rHA), gelatin, casein and the like. Representative amino acid / protein components, which can also act as buffers, include alanine, glycine, arginine, betaine, histidine, glutamic acid, aspartic acid, cysteine, lysine, leucine, isoleucine, valine, methionine, phenylalanine, aspartame, and the like.A preferred amino acid is glycine.
[00157] Non-limiting examples of carbohydrate excipients suitable for use include monosaccharides such as fructose, maltose, galactose, glucose, D-mannose, sorbose and the like; disaccharides such as lactose, sucrose, trehalose, cellobiose and the like; polysaccharides such as raffinose, melezitose, maltodextrins, dextrans, starches and the like; and alditols such as mannitol, xylitol, maltitol, lactitol, xylitol, sorbitol (glucitol), myo-inositol and the like. Preferably, the carbohydrate excipients are mannitol, trehalose and / or raffinose.
[00158] The compositions may also include a buffer or Petition 870250091181, dated 06 / 10 / 2025, pp. 173 / 220 78 / 112 a pH adjusting agent; commonly, the buffer is a salt prepared from an organic acid or base. Representative buffers include salts of organic acids, such as salts of citric acid, ascorbic acid, gluconic acid, carbonic acid, tartaric acid, succinic acid, acetic acid, or phthalic acid; tris, tromethamine hydrochloride, or phosphate buffers. Preferred buffers are salts of organic acids, such as citrate.
[00159] In addition, the disclosed compositions may include polymeric excipients / additives, such as polyvinylpyrrolidones, ficolls (a polymeric sugar), dextrates (e.g., cyclodextrins, such as 2-hydroxypropyl-cyclodextrin), polyethylene glycols, flavoring agents, antimicrobial agents, sweeteners, antioxidants, antistatic agents, surfactants (e.g., polysorbates, such as TWEEN 20 and TWEEN 80), lipids (e.g., phospholipids, fatty acids), steroids (e.g., cholesterol), and chelating agents (e.g., EDTA).
[00160] Many known and developed methods can be used to administer therapeutically effective amounts of the compositions or pharmaceutical compositions disclosed in this document. Non-limiting examples of administration methods include bolus, buccal, infusion, intra-articular, intrabronchial, intra-abdominal, intracapsular, intracartilaginous, intracavitary, intracerebellar, intracerebroventricular, intracolic, intracervical, intragastric, intrahepatic, intralesional, intramuscular, intramyocardial, intranasal, intraocular, intraosseous, intrapelvic, intrapericardial, intraperitoneal, intrapleural, intraprostatic, intrapulmonary, intrarectal, intrarenal, intraretinal, intraspinal, intrasynovial, intrathoracic, intrauterine, intratumoral, intravenous, intravesical, oral, parenteral, rectal, sublingual, subcutaneous, transdermal, or vaginal.In preferred embodiments, a composition comprising a modified cell described herein is administered via... Petition 870250091181, dated 06 / 10 / 2025, pp. 174 / 220 79 / 112 intravenous, for example, by intravenous infusion.
[00161] A composition disclosed herein may be prepared for parenteral use (subcutaneous, intramuscular or intravenous) or any other administration, particularly in the form of liquid solutions or suspensions. For parenteral administration, a composition disclosed herein may be formulated as a solution, suspension, emulsion, particle, powder or lyophilized powder in combination, or supplied separately, with a pharmaceutically acceptable parenteral vehicle. Formulations for parenteral administration may contain as common excipients water or sterile saline solution, polyalkylene glycols such as polyethylene glycol, vegetable oils, hydrogenated naphthalenes and the like. Aqueous or oily suspensions for injection may be prepared using an appropriate emulsifier or humidifier and a suspending agent, according to known methods.Agents for injection or infusion may be a non-toxic, non-orally administerable diluent, such as an aqueous solution, a sterile injectable solution, or a suspension in a solvent. Water, Ringer's solution, isotonic saline solution, etc., are permitted as vehicles or solvents; sterile non-volatile oil may be used as a common solvent or suspension solvent. For these purposes, any type of non-volatile oil and fatty acid may be used, including natural, synthetic, or semi-synthetic fatty oils or fatty acids; natural, synthetic, or semi-synthetic mono-, di-, or triglycerides. Parenteral administration is known in the art and includes, without limitation, conventional means of injection, a needle-free injection device under gas pressure as described in U.S. Patent No. 5,851,198, and a laser-piercing device as described in U.S. Patent No. 5,839,446.
[00162] It may be desirable to deliver the described compounds to the individual for extended periods, for example, for periods of one week to one year from a single administration. various release forms Petition 870250091181, dated 06 / 10 / 2025, pp. 175 / 220 Slow-release, depot, or implantable 80 / 112 dosage forms may be used. For example, a dosage form may contain a pharmaceutically acceptable non-toxic salt of compounds that have a low degree of solubility in body fluids, for example, (a) an acid addition salt with a polybasic acid, such as phosphoric acid, sulfuric acid, citric acid, tartaric acid, tannic acid, pamoic acid, alginic acid, polyglutamic acid, mono- or disulfonic naphthalene acids, polygalacturonic acid, and the like; (b) a salt with a polyvalent metallic cation, such as zinc, calcium, bismuth, barium, magnesium, aluminum, copper, cobalt, nickel, cadmium, and the like, or with an organic cation formed from, for example, N,N'-dibenzyl-ethylenediamine or ethylenediamine; or (c) combinations of (a) and (b), for example, a zinc tannate salt.Furthermore, the disclosed compounds, or preferably a relatively insoluble salt, such as those described above, may be formulated in a gel, for example, an aluminum monostearate gel with, for example, sesame oil, suitable for injection. Particularly preferred salts are zinc salts, zinc tannate salts, pamoate salts, and the like. Another type of slow-release depot formulation for injection would contain the compound or salt dispersed for encapsulation in a slow-degrading, non-toxic, and non-antigenic polymer, such as a polylactic acid / polyglycolic acid polymer, for example, as described in U.S. Patent No. 3,773,919. The compounds, or preferably relatively insoluble salts, such as those described above, may also be formulated in cholesterol matrix silastic pellets, particularly for use in animals.Additional slow-release, depot, or implant formulations, for example, gaseous or liquid liposomes, are known in the literature (U.S. Patent No. 5,770,222 and “Sustained and Controlled Release Drug Delivery Systems”, JR Robinson ed., Marcel Dekker, Inc., NY, 1978). Petition 870250091181, dated 06 / 10 / 2025, pages 176 / 220 81 / 112 Treatment Methods
[00163] In another aspect, methods of treating a disease or disorder in an individual are provided in this document, the method comprising administering to the individual a composition comprising the modified cells described in this document. The terms individual and patient are used interchangeably in this document. In preferred embodiments, the patient is human.
[00164] The modified cells can be allogeneic or autologous to the patient. In some preferred embodiments, the modified cell is an allogeneic cell. In some embodiments, the modified cell is an autologous T cell or a modified autologous CAR-T cell. In some preferred embodiments, the modified cell is an allogeneic T cell or a modified allogeneic CAR-T cell.
[00165] In some embodiments, the disease or disorder treated according to the methods described in this document is cancer. Non-limiting examples of cancer include leukemia, acute leukemia, acute lymphoblastic leukemia (ALL), acute lymphocytic leukemia, B-cell, T-cell or FAB ALL, acute myeloid leukemia (AML), acute myeloid leukemia, chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), hairy cell leukemia, myelodysplastic syndrome (MDS), lymphoma, Hodgkin's disease, malignant lymphoma, non-Hodgkin's lymphoma, Burkitt's lymphoma, multiple myeloma, Kaposi's sarcoma, colorectal carcinoma, pancreatic carcinoma, nasopharyngeal carcinoma, malignant histiocytosis, paraneoplastic syndrome / hypercalcemia of malignancy, solid tumors, bladder cancer, breast cancer, colorectal cancer, endometrial cancer, head cancer, neck cancer, hereditary nonpolyposis cancer, Hodgkin's lymphoma, liver cancer, lung cancer,Non-small cell lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, renal cell carcinoma, testicular cancer, adenocarcinomas, Petition 870250091181, dated 06 / 10 / 2025, pp. 177 / 220 82 / 112 sarcomas, malignant melanoma, hemangioma, metastatic disease, cancer-related bone resorption, cancer-related bone pain, and similar conditions.
[00166] In some embodiments, the disease or disorder treated according to the methods described in this document is a liver disease or disorder, a urea cycle disorder, a metabolic liver disorder, or a hemophilia disease. In some aspects, the metabolic liver disorder may be Ornithine Transcarbamylase (OTC) Deficiency. In some aspects, the metabolic liver disorder may be Methylmalonic Acidemia (MMA).
[00167] In a non-limiting example, the present disclosure provides methods for treating a hemophilic disease in an individual. In some respects, the hemophilic disease may be hemophilia A. In other respects, the hemophilic disease may be hemophilia B.
[00168] In a non-limiting example, the present disclosure provides methods for treating phenylketonuria (PKU) in an individual.
[00169] In some embodiments, the present disclosure provides methods for treating an autoimmune disease. In some embodiments, the autoimmune disease is autoimmune neutropenia, Guillain-Barré syndrome, epilepsy, autoimmune encephalitis, Isaacs syndrome, nevus syndrome, pemphigus vulgaris, pemphigus decidua, bullous pemphigoid, acquired epidermolysis bullosa, gestational pemphigoid, mucous membrane pemphigoid, antiphospholipid antibody syndrome, autoimmune anemia, myasthenia gravis, autoimmune Graves' disease, thyroid eye disease (TED), Goodpasture syndrome, multiple sclerosis, rheumatoid arthritis, lupus, idiopathic thrombocytopenic purpura (ITP), warm autoimmune hemolytic anemia (WAIHA), chronic inflammatory demyelinating polyneuropathy (CIDP), lupus nephritis, or membranous nephropathy.
[00170] The dosage of a pharmaceutical composition to be Petition 870250091181, dated 06 / 10 / 2025, pp. 178 / 220 The effect of 83 / 112 administered to an individual may vary depending on known factors, such as the pharmacodynamic characteristics of the specific agent and its mode and route of administration; the recipient's age, health, and weight; the nature and extent of symptoms; the type of concomitant treatment; the frequency of treatment; and the desired effect.
[00171] In aspects where the compositions to be administered to an individual who needs them are modified cells as disclosed in this document, between approximately 1x103 and approximately 1x104 cells; between approximately 1x104 and approximately 1x105 cells; between approximately 1x105 and approximately 1x106 cells; between approximately 1x106 and approximately 1x107 cells; between approximately 1x107 and approximately 1x108 cells; between approximately 1x108 and approximately 1x109 cells; between about 1x10⁹ and about 1x10¹⁰ cells, between about 1x10¹⁰ and about 1x10¹¹ cells, between about 1x10¹¹ and about 1x10¹² cells, between about 1x10¹² and about 1x10¹³ cells, between about 1x10¹³ and about 1x10¹⁴ cells, between about 1x10¹⁴ and about 1x10¹⁵ cells, between about 1x10¹⁵ and about 1x10¹⁶ cells, between about 1x10¹⁶ and about 1x10¹⁷ cells, between about 1x10¹⁷ and about 1x10¹⁸ cells, between about 1x10¹⁸ and about 1x10¹⁹ cells; Between approximately 1x10¹⁹ and approximately 1x10²⁰ cells can be administered.In some embodiments, the cells are administered in a dose between approximately 5x10⁶ and approximately 25x10⁶ cells.
[00172] In other embodiments, the cell dosage may depend on the person's body weight, for example, between approximately 1x10³ and approximately 1x10⁴ cells; between approximately 1x10⁴ and approximately 1x10⁵ cells; between approximately 1x10⁵ and approximately 1x10⁶ cells; between approximately 1x10⁶ and approximately 1x10⁷ cells; between approximately 1x10⁷ and approximately 1x10⁸ cells; between approximately 1x10⁸ and approximately 1x10⁹ cells; between approximately 1x10⁹ and approximately 1x10¹⁰ cells, between approximately 1x10¹⁰ and approximately 1x10¹¹ cells, between approximately 1x10¹¹ and approximately 1x10¹² cells, between approximately 1x10¹² and approximately 1x10¹³ cells, between approximately 1x10¹³ and approximately Petition 870250091181, dated 06 / 10 / 2025, pp. 179 / 220 84 / 112 of 1x1014 cells, between approximately 1x1014 and approximately 1x1015 cells, between approximately 1x1015 and approximately 1x1016 cells, between approximately 1x1016 and approximately 1x1017 cells, between approximately 1x1017 and approximately 1x1018 cells, between approximately 1x1018 and approximately 1x1019 cells; or between approximately 1x1019 and approximately 1x1020 cells may be administered per kg of the individual's body weight.
[00173] A more detailed description of the excipients, formulations, dosages and pharmaceutically acceptable methods of administration of the disclosed compositions and pharmaceutical compositions is presented in PCT Publication No. WO 2020 / 051374.
[00174] The transposase domains and fusion proteins provided in this document can be used to deliver gene therapy. Gene therapy generally involves delivering a transgene to the genomic DNA of a cell. Usually, the transgene replaces a gene that has mutated or is not adequately expressed in the cell. For example, the transgene may replace a gene that exhibits decreased, insufficient, and / or altered expression in the cell. In some embodiments, this decreased, insufficient, and / or altered expression may directly or indirectly result in a disease or disorder, such as a liver disease or disorder, a urea cycle disorder, a metabolic liver disorder, or a hemophilia disease. The fusion proteins, transposase domains, and complexes described in this document can be used to deliver a therapeutic transgene to a cell and integrate the transgene into a target site.In some embodiments, a treatment method comprises introducing into the cell a fusion protein provided in this disclosure and a transposon, wherein the transposon comprises, in the order of 5' to 3': a 5'ITR, the transgene and a 3'ITR.
[00175] In some embodiments, the therapeutic transgene is a gene that is expressed at lower levels, and the reduced expression results in a disease or disorder. In some embodiments, the therapeutic transgene is a Petition 870250091181, dated 06 / 10 / 2025, pp. 180 / 220 85 / 112 gene that is expressed in an altered pattern compared to a wild-type gene, and the altered expression results in a disease or disorder. Therefore, methods are provided in this document for treating a disease or disorder caused by or associated with altered gene expression, comprising administering a transposon described in this document and a transposase to an individual in need thereof.
[00176] The therapeutic transgene delivered to the cell by the fusion proteins, transposase domains, and complexes described in this document may encode a therapeutic polypeptide. In some embodiments, the therapeutic polypeptide is the Factor VIII polypeptide, Factor IX polypeptide, phenylalanine hydroxylase (PAH), ornithine transcarbamylase (OTC) polypeptide, or methylmalonyl-CoA mutase (MUT1) polypeptide.
[00177] In a non-limiting example, the transposase domains and fusion proteins provided herein can be used to deliver liver-targeted gene therapy. In some respects, liver-targeted gene therapy can be used to treat Ornithine Transcarbamylase (OTC) deficiency, and the therapeutic polypeptide encoded by the therapeutic transgene may comprise the ornithine transcarbamylase (OTC) polypeptide. In some respects, liver-targeted gene therapy can be used to treat methylmalonic acidemia (MMA), and at least one therapeutic protein encoded by the therapeutic transgene may comprise a methylmalonyl-CoA mutase (MUT1) polypeptide.
[00178] In some respects, a gene therapy targeting the liver can be used to treat hemophilia A, and at least one therapeutic protein encoded by the therapeutic transgene may comprise Factor VIII. In some respects, a gene therapy targeting the liver can be used to treat hemophilia B, and at least one therapeutic protein encoded by the therapeutic transgene may comprise Factor IX. Petition 870250091181, dated 06 / 10 / 2025, pages 181 / 220 86 / 112
[00179] In some respects, a gene therapy targeting the liver can be used to treat phenylketonuria (PKU) and at least one therapeutic protein encoded by the therapeutic transgene may comprise phenylalanine hydroxylase (PAH). Kits
[00180] In another aspect, a kit is provided herein comprising a cell line that has been engineered to include a modified target site for an SPB or a PBx provided herein within its genome, preferably in a highly expressed genomic region. The kit may further comprise a composition comprising one or more SPB or PBx transposase domains or fusion proteins described herein. In some embodiments, the cell line is a T cell line. Definitions
[00181] As used throughout the disclosure, the singular forms a, eeo include plural referents unless the context clearly indicates otherwise. Thus, for example, reference to “a method” includes a plurality of such methods and reference to “a dose” includes reference to one or more doses and equivalents known to those skilled in the art, and so forth.
[00182] The term about or approximately means within an acceptable range of error for the specific value, as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, for example, the limitations of the measuring system. For example, about may mean within 1 or more standard deviations. Alternatively, about may mean a range of up to 20%, or up to 10%, or up to 5%, or up to 1% of a given value. Alternatively, in particular with regard to biological systems or processes, the term may Petition 870250091181, dated 06 / 10 / 2025, pages 182 / 220 87 / 112 means within an order of magnitude, preferably within 5 times, and more preferably within 2 times, of a value. When specific values are described in the application and claims, unless otherwise indicated, the term approximately means within an acceptable error range for the specific value.
[00183] Disclosure provides compositions of isolated or substantially purified polynucleotides or proteins. An isolated or purified polynucleotide or protein, or its biologically active portion, is substantially or essentially free of components that normally accompany or interact with the polynucleotide or protein as found in its naturally occurring environment. Therefore, an isolated or purified polynucleotide or protein is substantially free of other cellular material or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized.Optimally, an isolated polynucleotide is free of sequences (optimally, protein-coding sequences) that naturally flank the polynucleotide (i.e., sequences located at the 5' and 3' ends of the polynucleotide) in the genomic DNA of the organism from which the polynucleotide is derived. For example, in various respects, the isolated polynucleotide may contain less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleotide sequence that naturally flanks the polynucleotide in the genomic DNA of the cell from which the polynucleotide is derived. A substantially cell-free protein includes protein preparations with less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of contaminating protein.When the disseminating protein or its biologically active portion is produced recombinantly, the optimized culture medium contains less than approximately 30%, 20%, 10%, 5%, or 1% (by dry weight) of chemical precursors or substances. Petition 870250091181, dated 06 / 10 / 2025, pages 183 / 220 88 / 112 non-protein chemicals of interest.
[00184] The disclosure provides fragments and variants of the disclosed DNA sequences and proteins encoded by those DNA sequences. As used throughout the disclosure, the term fragment refers to a portion of the DNA sequence or a portion of the amino acid sequence and therefore to the protein it encodes. Fragments of a DNA sequence comprising coding sequences may encode protein fragments that retain the biological activity of the native protein and therefore the DNA recognition or binding activity to a target DNA sequence as described herein. Alternatively, fragments of a DNA sequence that are useful as hybridization probes generally do not encode proteins that retain biological activity or that do not retain promoter activity.Therefore, the fragments of a DNA sequence can vary from at least about 20 nucleotides, about 50 nucleotides, about 100 nucleotides, and up to the full-length polynucleotide of the fragment.
[00185] Dissemination nucleic acids or proteins can be constructed by a modular approach, including pre-assembly of monomeric units and / or repeat units into target vectors that can be subsequently assembled into a final target vector. Dissemination polypeptides can comprise dissemination repeat monomers and can be constructed by a modular approach, pre-assembling repeat units into target vectors that can be subsequently assembled into a final target vector. Dissemination provides polypeptides produced by this method as well as nucleic acid sequences encoding these polypeptides. Dissemination provides host organisms and cells comprising nucleic acid sequences encoding polypeptides produced by this modular approach. Petition 870250091181, dated 06 / 10 / 2025, pp. 184 / 220 89 / 112
[00186] The term comprising is intended to mean that the compositions and methods include the elements mentioned, but do not exclude others. Consisting essentially of, when used to define compositions and methods, means excluding other elements of any essential importance to the combination when used for the intended purpose. Therefore, a composition consisting essentially of the elements defined herein would not exclude trace contaminants or inert carriers. Consisting of means excluding more than residual elements from other substantial ingredients and steps of the method. The aspects defined by each of these transitional terms are within the scope of this disclosure.
[00187] As used in this document, expression refers to the process by which polynucleotides are transcribed into mRNA and / or the process by which the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. If the polynucleotide is derived from genomic DNA, the expression may include mRNA splicing in a eukaryotic cell.
[00188] Gene expression refers to the conversion of the information contained in a gene into a gene product. A gene product can be the direct transcriptional product of a gene (e.g., mRNA, tRNA, rRNA, antisense RNA, ribozyme, shRNA, microRNA, structural RNA, or any other type of RNA) or a protein produced by the translation of an mRNA. Gene products also include RNAs modified by processes such as encapsulation, polyadenylation, methylation, and editing, and proteins modified by, for example, methylation, acetylation, phosphorylation, ubiquitination, ADPribosylation, myristylation, and glycosylation.
[00189] Modulation or regulation of gene expression refers to an alteration in the activity of a gene. Modulation of expression can Petition 870250091181, dated 06 / 10 / 2025, pages 185 / 220 90 / 112 include, among other things, gene activation and repression.
[00190] The term “operatively linked” or its equivalents (e.g., “operatively linked”) means that two or more molecules are positioned relative to each other in such a way that they have the capacity to interact to affect a function attributable to one or both molecules or to a combination thereof. In the context of nucleic acids, a promoter may be operatively linked to a nucleotide sequence encoding a transposable domain or fusion protein described in this document, placing the expression of the nucleotide sequence under the control of the promoter.
[00191] Non-covalently linked components and methods for producing and using non-covalently linked components are disclosed. The various components can take a variety of different forms, as described in this document. For example, non-covalently linked (i.e., operatively linked) proteins can be used to allow temporary interactions that avoid one or more problems in the art. The ability of non-covalently linked components, such as proteins, to associate and dissociate allows functional association only or primarily in circumstances where such association is required for the desired activity. The connection can last long enough to allow the desired effect.
[00192] A method for targeting proteins to a specific locus in an organism's genome is disclosed. The method may comprise the steps of providing a DNA localization component and providing an effector molecule, wherein the DNA localization component and the effector molecule have the ability to bind operatively via a non-covalent bond.
[00193] A target site or target sequence is a sequence of Petition 870250091181, dated 06 / 10 / 2025, pp. 186 / 220 91 / 112 nucleic acid that defines a portion of a nucleic acid to which a linker molecule will bind, provided there are sufficient conditions for binding.
[00194] The terms nucleic acid or oligonucleotide or polynucleotide refer to at least two nucleotides covalently linked together. The representation of a single strand also defines the sequence of the complementary strand. Therefore, a nucleic acid can also encompass the complementary strand of a single strand shown. A nucleic acid disclosure also encompasses substantially identical nucleic acids and complements of the same that retain the same structure or encode the same protein.
[00195] Disclosure nucleic acids can be single-stranded or double-stranded. Disclosure nucleic acids can contain double-stranded sequences, even when most of the molecule is single-stranded. Disclosure nucleic acids can contain single-stranded sequences, even when most of the molecule is double-stranded. Disclosure nucleic acids can include genomic DNA, cDNA, RNA, or a hybrid thereof. Disclosure nucleic acids can contain combinations of deoxyribonucleotides and ribonucleotides. Disclosure nucleic acids can contain combinations of bases, including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. Disclosure nucleic acids can be synthesized to comprise unnatural modifications of amino acids. Disclosure nucleic acids can be obtained by chemical synthesis methods or by recombinant methods.
[00196] The nucleic acids in the disclosure, whether their complete sequence or any portion thereof, may be of unnatural occurrence. The nucleic acids in the disclosure may contain one or more mutations, substitutions, deletions, or insertions that do not occur naturally, making the entire nucleic acid sequence of unnatural occurrence. Petition 870250091181, dated 06 / 10 / 2025, pages 187 / 220 92 / 112 nucleic acids in disclosure may contain one or more duplicated, inverted, or repeated sequences, the resulting sequence of which does not occur naturally, making the entire nucleic acid sequence unnatural. Nucleic acids in disclosure may contain modified, artificial, or synthetic nucleotides that do not occur naturally, making the entire nucleic acid sequence unnatural.
[00197] Given the redundancy in the genetic code, a plurality of nucleotide sequences can code for any specific protein. All such nucleotide sequences are covered in this document.
[00198] As used throughout this disclosure, the term promoter refers to a synthetic or naturally occurring molecule that has the ability to confer, activate, or enhance the expression of a nucleic acid in a cell. A promoter may comprise one or more specific transcriptional regulatory sequences to further enhance expression and / or alter the spatial and / or temporal expression thereof. A promoter may also comprise distal enhancer or repressor elements, which may be located up to several thousand base pairs from the transcription start site. A promoter may be derived from sources including viruses, bacteria, fungi, plants, insects, and animals.A promoter can regulate the expression of a genetic component constitutively or differentially with respect to the cell, tissue, or organ in which the expression occurs, or with respect to the developmental stage at which the expression occurs, or in response to external stimuli such as physiological stresses, pathogens, metal ions, or inducing agents. Representative examples of promoters include the bacteriophage T7 promoter, the bacteriophage T3 promoter, the SP6 promoter, the lac operator-promoter, the tac promoter, the SV40 late promoter, the V40 early promoter, and the RSVP promoter. (Reference 870250091181, 06 / 10 / 2025, p. 188 / 220.) 93 / 112 LTR, the CMV IE promoter, the EF-1 Alfa promoter, the CAG promoter, the early SV40 promoter or the late SV40 promoter, and the CMV IE promoter.
[00199] As used throughout the disclosure, the term vector refers to a nucleic acid sequence containing an origin of replication. A vector can be a viral vector, bacteriophage, bacterial artificial chromosome, or yeast artificial chromosome. A vector can be a DNA or RNA vector. A vector can be a self-replicating extrachromosomal vector and, preferably, is a DNA plasmid. A vector can comprise a combination of an amino acid with a DNA sequence, an RNA sequence, or both a DNA and RNA sequence.
[00200] A conservative amino acid substitution, that is, the substitution of an amino acid by a different amino acid with similar properties (e.g., hydrophilicity, degree and distribution of charged regions), is recognized in the art as commonly involving a small alteration. These small alterations can be identified, in part, by considering the hydropathic index of the amino acids, as understood in the art. Kyte et al., J. Mol. Biol. 157: 105-132 (1982). The hydropathic index of an amino acid is based on consideration of its hydrophobicity and charge. Amino acids with similar hydropathic indices can be substituted and still retain protein function. In one aspect, amino acids with hydropathic indices of ±2 are substituted. The hydrophilicity of amino acids can also be used to reveal substitutions that would result in proteins that retain biological function.Considering the hydrophilicity of amino acids in the context of a peptide allows for the calculation of the highest local average hydrophilicity of that peptide, a useful measure that has been reported to correlate well with antigenicity and immunogenicity. U.S. Patent No. 4,554,101, incorporated in full by reference herein.
[00201] The replacement of amino acids with values of Petition 870250091181, dated 06 / 10 / 2025, pages 189 / 220 94 / 112 similar hydrophilicity can result in peptides that retain biological activity, for example, immunogenicity. Substitutions can be made with amino acids with hydrophilicity values within ±2 of each other. Both the hydrophobicity index and the hydrophilicity value of amino acids are influenced by the specific side chain of that amino acid. Consistent with this observation, amino acid substitutions compatible with biological function are understood to be dependent on the relative similarity of the amino acids, and particularly of the side chains of these amino acids, as revealed by hydrophobicity, hydrophilicity, charge, size, and other properties.
[00202] As used in this document, conservative amino acid substitutions can be defined as set out in Table 3, Table 4, and Table 5 below. In some respects, fusion polypeptides and / or nucleic acids encoding such fusion polypeptides include conservative substitutions that have been introduced by modification of polynucleotides encoding dissemination polypeptides. Amino acids can be classified according to their physical properties and contribution to the secondary and tertiary structure of the protein. A conservative substitution is the replacement of one amino acid with another amino acid that has similar properties. Exemplary conservative substitutions are presented in Table 3. Table 3: Conservative Substitutions I Side chain characteristics: Aliphatic amino acid: Nonpolar; GAPILVF: Polar - uncharged; CSTMNQ: Polar - charged; DEKR: Aromatic; HFWY: Other; NQDE
[00203] Alternatively, the conservative amino acids can be grouped as described in Lehninger, (Biochemistry, Second Edition; Petition 870250091181, dated 06 / 10 / 2025, pp. 190 / 220 95 / 112 Worth Publishers, Inc. NY, NY (1975), pp. 71-77) as set forth in Table 4. Table 4: Conservative Substitutions II Side Chain Characteristics Amino Acid Nonpolar (hydrophobic) Aliphatic: ALIVP Aromatic: FWY Sulfur-containing: M Borderline: GY Polar uncharged Hydroxyl: STY Amides: NQ Sulfhydryl: C Borderline: GY Positively Charged (Basic): KRH Negatively Charged (Acidic): DE
[00204] Alternatively, exemplary conservative substitutions are set out in Table 5. Table 5: Conservative Substitutions III Original Residue Exemplary Substitution Ala (A) Val Leu Ile Met Arg (R) Lys His Asn (N) Gln Asp (D) Glu Cys (C) Ser Thr Gln (Q) Asn Glu (E) Asp Gly (G) Ala Val Leu Pro His (H) Lys Arg Ile (I) Leu Val Met Ala Phe Leu (L) Ile Val Met Ala Phe Lys (K) Arg His Met (M) Leu Ile Val Ala Phe (F) Trp Tyr Ile Pro (P) Gly Ala Val Leu Ile Ser(S) Thr Thr(T) Ser Trp (W) Tyr Phe Ile Tyr (Y) Trp Phe Thr Ser Val (V) Ile Leu Met Ala
[00205] The polypeptides and proteins of the disclosure, whether their entire sequence or any portion thereof, may be of non-occurrence. Petition 870250091181, dated 06 / 10 / 2025, pages 191 / 220 96 / 112 natural. Disclosure polypeptides and proteins may contain one or more mutations, substitutions, deletions, or insertions that do not occur naturally, making the entire amino acid sequence unnatural. Disclosure polypeptides and proteins may contain one or more duplicated, inverted, or repeated sequences, the resulting sequence of which occurs naturally, making the entire amino acid sequence unnatural. Disclosure polypeptides and proteins may contain modified, artificial, or synthetic amino acids that do not occur naturally, making the entire amino acid sequence unnatural.
[00206] As used throughout the disclosure, the identity between two sequences can be determined using the stand-alone executable program BLAST engine for comparing two sequences (bl2seq), which can be obtained from the National Center for Biotechnology Information (NCBI) ftp site, using the standardized parameters (Tatusova and Madden, FEMS Microbiol Lett., 1999, 174, 247-250; which is incorporated by reference in its entirety herein). The terms identical or identity, when used in the context of two or more nucleic acids or polypeptide sequences, refer to a specified percentage of residues that are the same in a specified region of each of the sequences. In some embodiments, sequence identity is determined along the entire length of a sequence.The percentage can be calculated by optimally aligning the two sequences, comparing them over the specified region, determining the number of positions where the identical residue occurs in both sequences to produce the number of matching positions, dividing the number of matching positions by the total number of positions in the specified region, and multiplying the result by 100 to produce the sequence identity percentage. In cases where... Petition 870250091181, dated 06 / 10 / 2025, pages 192 / 220 97 / 112 if the two sequences have different lengths or the alignment produces one or more staggered ends and the specified comparison region includes only a single sequence, the residues of the single sequence are included in the denominator but not in the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be performed manually or using a computational sequence algorithm such as BLAST or BLAST 2.0.
[00207] In certain embodiments, if a sequence has a certain sequence identity (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%) with a given SEQ ID NO, the sequence and the sequence of the SEQ ID NO have the same length. In certain embodiments, if a sequence has a certain sequence identity (e.g., 75%, 80%, 85%, 90%, 95%, 98%, or 99%) with a given SEQ ID NO, the sequence and the sequence of the SEQ ID NO differ only due to conservative amino acid substitutions.
[00208] As used throughout the disclosure, the term endogenous refers to a nucleic acid or protein sequence naturally associated with a target gene or a host cell into which it is introduced.
[00209] As used throughout the disclosure, the term exogenous refers to a nucleic acid or protein sequence not naturally associated with a target gene or a host cell into which it is introduced, including unnaturally occurring multiple copies of a naturally occurring nucleic acid, for example, a DNA sequence or a naturally occurring nucleic acid sequence located at an unnaturally occurring genomic locus.
[00210] The disclosure provides methods for introducing a polynucleotide construct comprising a DNA sequence into a Petition 870250091181, dated 06 / 10 / 2025, pp. 193 / 220 98 / 112 host cell. By introduction, it is meant presenting the polynucleotide construct to the cell in such a way that the construct has access to the interior of the host cell. Disclosure methods do not depend on a specific method for introducing a polynucleotide construct into a host cell, only that the polynucleotide construct has access to the interior of a host cell. Methods for introducing polynucleotide constructs into bacteria, plants, fungi, and animals are known in the art, including, without limitation, stable transformation methods, transient transformation methods, and virus-mediated methods. Examples
[00211] The examples in this section are provided for illustration purposes only and are not intended to limit the invention. Example 1: Construction of amino-terminal deletions of Super PiggyBac transposases
[00212] Plasmids comprising a nucleotide sequence encoding a full-length, wild-type Super PiggyBac transposase (SPB; SEQ ID NO: 2) or a nucleotide sequence encoding an integration-deficient Super PiggyBac transposase variant comprising amino acid substitutions at positions R372A, K375A, and D450N (PBx; SEQ ID NO: 3) were used as templates for PCR mutagenesis to generate N-terminal deleted transposase variants lacking the 93 N-terminal amino acids (SPBΔ1-93 and PBxΔ1-93, respectively).
[00213] In summary, forward and reverse primers were designed to amplify a portion of the SPB and PBx coding sequences corresponding to amino acids 94–594. The resulting DNA fragments encoding SPBΔ1-93 or PBxΔ1-93 were used along with a fragment of the acquired gBlock gene to construct fusion proteins. Petition 870250091181, dated 06 / 10 / 2025, pp. 194 / 220 99 / 112 DNA-binding domain transposase via a next-generation 2-fragment Gibson assembly.
[00214] Additional N-terminal deleted transposase variants, lacking the 85 N-terminal amino acids (SPBΔ1-85 and ΡΒχΔ1-85, respectively), were generated as described in this document. Example 2: Design and Construction of TAL Matrices Targeting LPA
[00215] This Example illustrates the design and construction of TAL Matrix compositions targeting the LPA gene that can be used in methods to validate the target specificity of TAL Matrices. The TAL Matrices were constructed using the design criteria as set out below.
[00216] The Lipoprotein A (LPA) gene contains up to 50 copies of a segmental duplication element, making it a potentially attractive target for optimizing the chance of a site-specific transposition event at a target sequence, thus leading to an increase in the number of transposed cells.
[00217] TAL Matrix pairs comprising an N-terminal domain that recognizes a T were designed targeting four specific 10 bp right- and left-hand pair sequences within the LPA gene repeat elements. For three of the targets, multiple TAL Matrix pairs were designed using 12 bp or 13 bp spacers.
[00218] The left and right target sequences, along with the upstream 5'T used to generate TAL arrays targeting the LPA gene, are shown in Table 6. Table 6: Illustrative TAL Matrices that have LPA as a Target pair lpa # target sequence left target sequence right 1 TGGAGACCCCA (SEQ ID NO: 89) TCTAGTAATAT (SEQ ID NO: 90) 2 TGAAGAAACAG (SEQ ID NO: 91) TTCCTACATGT (SEQ ID NO: 92) TGAAGAAACAG (SEQ ID NO: 91) TCCTACATGTC (SEQ ID NO: 93) 3 TGAAACAAAAT (SEQ ID NO: 94) TCTTACCTCTA (SEQ ID NO: 95) Petition 870250091181, dated 06 / 10 / 2025, pp. 195 / 220 100 / 112 par lpa # target sequence left target sequence right TGAAACAAAAT (SEQ ID NO: 94) TTCTTACCTCT (SEQ ID NO: 96) 4 TTAAAAAAAAT (SEQ ID NO: 97) TCCTCTCTGCA (SEQ ID NO: 99) TTTAAAAAAAA (SEQ ID NO: 98) TCCTCTCTGCA (SEQ ID NO: 99)
[00219] Individual TAL modules containing “half” repeats of 34 amino acids or 20 amino acids were synthesized flanked by BsmBI type IIS restriction sites. The complete set of modules contains 4 modules capable of recognizing A, C, G, and T for each of the 10 bp positions within a target sequence (40 modules / 10 bp target). Sequence pairs targeting TAL arrays in the LPA gene were designed, and the corresponding modules were selected and grouped using Golden Gate Assembly to be assembled into a structure and create each LPA TAL Array. All coding sequences used were codon-optimized for human expression.
[00220] The seven combinations of left and right pairs were used to design and construct the Left TAL Matrices LPA, LPAL1, LPAL2, LPAL3, LPAL4.1, and LPAL4.2 (SEQ ID Nos. 116, 118, 121, 124, and 125, respectively) and the Right TAL Matrices LPA LPAR1, LPAR2.1, LPAR2.2, LPA3.1, LPAR3.2, and LPAR4 (SEQ ID Nos. 117, 119, 120, 122, 123, and 126, respectively). Example 3: Construction and Analysis of TAL Transposase piggyBac (ss-SPB) (TAL-PBxs) Matrix Compositions Designed for Site-Specific Transposition in the LPA Gene
[00221] This example illustrates the construction of TAL Matrix-transposase Super piggyBac (TAL-ssSPB) fusion protein compositions that are useful in methods to achieve site-specific transposition at a specific target locus.
[00222] The TAL-PBx fusion constructs were prepared as follows: an expression plasmid was synthesized containing the direction Petition 870250091181, dated 06 / 10 / 2025, pp. 196 / 220 101 / 112 5' to 3': a CMV promoter, a T7 promoter, a Kozak sequence, a Flag tag 3x (SEQ ID NO: 65), an SV40 NLS (SEQ ID NO: 66), the N-terminal Delta 152 TAL domain (SEQ ID NO: 31), two BsmBI type IIS restriction enzyme sites, the C-terminal +63 TAL domain (SEQ ID NO: 32), a GGGS linker, delta 193 PBx (comprising a 93 N-terminal amino acid deletion and mutations at R372A, K375A, D450N in the Super piggyBac transposase codon sequence; SEQ ID NO: 6) and a bGH polyadenylation sequence.
[00223] Cloning a left- or right-flanked BsmBI TAL array into the BsmBI sites of the expression plasmid results in in-frame fusion of the TAL array and the PBx coding sequence via a linkage sequence, generating full-length TAL-PBx constructs. All coding sequences used were codon-optimized for human expression using GeneArt algorithms (Thermo Fisher).
[00224] The eleven TAL matrices designed and constructed in Example 2 flanked with BsmBI ends were cloned into the BsmBI restriction sites of the expression plasmid described above to generate eleven TAL-PBx constructs: LPAL1, LPAL2, LPAL3, LPAL4.1 and LPAL4.2 TAL-PBxs on the Left (SEQ ID Nos. 143, 145, 148, 151 and 152, respectively) and LPAR1, LPAR2.1, LPAR2.2, LPAR3.1, LPAR3.2 and LPAR4 TAL-PBxs on the Right (SEQ ID Nos. 144, 146, 147, 149, 150 and 153, respectively). Example 4: Demonstration of Site-Specific Transposition Using TAL Matrix Compositions - piggyBac Transposase (ss-SPB) (TAL-PBxs) and an Episomal Division Gfp Splicing Reporter System
[00225] This example illustrates exemplary compositions and methods for demonstrating site-specific transposition at specific episomal loci using TAL Matrix Fusion Proteins - SPB transposase.
[00226] An episomal splitting GFP reporter system was employed to evaluate site transposition efficiency. Petition 870250091181, dated 06 / 10 / 2025, pages 197 / 220 102 / 112 specific to the various TAL Matrix fusion proteins - SPB transposases constructed in Example 3. The reporter system consists of two plasmids. The first plasmid, the reporter, was constructed containing, from the 5' to 3' direction: an EF1a promoter (SEQ ID NO: 67), a Kozak sequence, the first portion of an open GFP reading frame (SEQ ID NO: 68), a splicing donor (SEQ ID NO: 69), and two BsaI-type restriction enzyme IIS sites. The BsaI sites allow cloning of a target TTAA sequence flanked by variable-length spacers flanked by target recognition sequences for TAL matrices.The second plasmid, the donor, was constructed containing, from the 5' to 3' direction: a TTAA sequence, the 5' minimum PiggyBac ITR of 35 bp (SEQ ID NO: 70), a splice acceptor site (SEQ ID NO: 71), the second portion of an open GFP reading frame (SEQ ID NO: 72), a synthetic polyadenylation sequence (SEQ ID NO: 73), the 3' minimum PiggyBac ITR of 63 bp (SEQ ID NO: 74), and a TTAA sequence. A schematic diagram of the split GFP reporter plasmid is shown in FIGURE 2.
[00227] Four different LPA target sequences found naturally in genomic DNA (SEQ ID Nos. 81-84) were cloned into the episomal reporter plasmid described above. Complementary oligos were synthesized containing the LPA genomic DNA sequences (SEQ ID Nos. 81-84). The complementary oligos contained 4 bp overhangs compatible with the overhangs created in the spliced GFP reporter after digestion with BsaI. The oligos were annealed and ligated to the digested vector to create a reporter compatible with each LPA TAL-PBx pair constructed in Example 3.
[00228] TAL arrays were designed and constructed to create heterodimeric pairs of TAL-ssSPBs (i.e., one TAL-PBx array on the left and one on the right). Each pair of TAL-PBx constructs was cotransfected into HEK293T cells with its corresponding reporter plasmid and plasmid Petition 870250091181, dated 06 / 10 / 2025, pages 198 / 220 103 / 112 donor. As a negative control, each pair of TAL-PBx constructs was co-transfected into HEK293T cells with a non-matching reporter plasmid (i.e., TAL-PBx pair 1 with reporter 2, TAL-PBx pair 2 with reporter 3, TAL-PBx pair 3 with reporter 4, and TAL-PBx pair 4 with reporter 1) and the donor plasmid. Transfection mixtures containing 26 ng of the TAL-ssSPB expression vector, 170 ng of the reporter plasmid, 117 ng of the donor plasmid, and 0.78 µl of the Transit-2020 transfection reagent in a total volume of 26 µl of OptiMem Serum-Free medium were prepared. 95,000 HEK293T cells were added to 250 µl of DMEM medium supplemented with 10% FBS, and the transfection mixture was placed in 48-well plates and incubated for four days at 37 °C with 5% CO2, splitting the cells 1:3 on the second day.
[00229] When reporter and donor plasmids are cotransfected into cells along with TAL-PBx, TAL-PBx catalyzes the excision of the donor plasmid transposon and its site-specific integration into the TTAA target site of the reporter plasmid. FIGURE 3 is a schematic showing the catalytic dimer ssSPB bound to an excised transposon and recognizing its genomic integration target site. After site-specific transposition, transcription, splicing, and translation, a reconstituted GFP coding sequence is produced (DNA, SEQ ID NO: 75; Amino acid; SEQ ID NO: 76) and fluorescence can be detected. The percentage of cells positive for site-specific transposition at the target for the various TAL-PBx pairs was determined by FACS analysis, and the results are shown in Table 7. Table 7 On Target Off Target Reply 1 Reply 2 Average Reply 1 Reply 2 Average LPA L1 / R1 21.9 21.1 21.5 4.3 3.9 4.1 LPA L2 / R2.1 12.3 11.6 12.0 0.0 0.1 0.1 LPA L2 / R2.2 12.9 12.5 12.7 4.8 4.5 4.7 LPA L3 / R3.1 16.9 15.7 16.3 0.0 0.0 0.0 Petition 870250091181, dated 06 / 10 / 2025, pp. 199 / 220 104 / 112 LPA L3 / R3.2 11.4 11.8 11.6 4.2 4.8 4.5 LPA L4.1 / R4 10.4 9.6 10.0 5.0 4.6 4.8
[00230] As observed in Table 7, all TAL-ssSPBs catalyzed site-specific transposition of their respective reporters on the target, but not with reporters containing a non-matching off-target. Furthermore, the greatest transposition was observed on target 1, the only target with a TTTAAA integration site. Example 5: Determination of the Optimized 5' and 3' Flanking Nucleotides Immediately Adjacent to the TTAA Integration Site
[00231] The previous example shows that the target site with the most robust integration, target 1, contains a 5'T and a 3'A immediately adjacent to the TTAA target site, generating a TTTAAA integration site. This example illustrates additional compositions and methods for preparing site-specific transposition-optimized target sites by determining the optimized 5' and 3' flanking nucleotides immediately adjacent to the TTAA integration site.
[00232] An episomal splitting GFP splice reporter, as described in Example 4, was employed to evaluate the site-specific transposition efficiency of several TAL-PBx fusion proteins that are targets of the green fluorescent protein (GFP) gene. TAL Matrix Fusion Proteins - SPB transposase GFP1 Right TAL-PBx and GFP1 Left TAL-PBx that are targets of specific 10 bp right and 10 bp left sequences in the coding region of the GFP gene were prepared as described in Examples 14 and 18 of International Patent Application Publication No. PCT / US2022 / 77549, the contents of which are incorporated by reference in their entirety.
[00233] To create a GFP1-compatible reporter plasmid TAL-PBx on the right, complementary oligos were synthesized containing the Petition 870250091181, dated 06 / 10 / 2025, pp. 200 / 220 105 / 112 target site for GFP1 TAL on the right downstream of a T followed by a 12 bp spacer followed by TTAA followed by a 12 bp spacer, followed by the reverse complement of the TAL target site followed by an A (SEQ ID No. 172). The spacer sequences were such that the nucleotide immediately 5' from TTAA is C and the nucleotide immediately 3' from TTAA is C. The complementary oligos contained 4 bp overhangs compatible with the overhangs created in the split GFP splice reporter after digestion with BsaI. The oligos were annealed and ligated to the digested vector to create a TAL-PBx GFP1-compatible reporter on the right. Similar oligos were synthesized with modified 12 bp spacer sequences to mutate the 5' and 3' flanking nucleotide immediately adjacent to the TTAA integration sequence to a T and an A, respectively, to generate a TTTAAA integration site (SEQ ID No.173), or for a C and an A, respectively, to generate a CTTAAA integration site (SEQ ID No. 174). Similar oligonucleotides were synthesized containing the target site for GFP1 TAL to the right downstream of a T followed by a 13 bp spacer followed by TTAA followed by a 13 bp spacer, followed by the reverse complement of the target site TAL followed by an A (SEQ ID No. 175). The spacer sequences were such that the nucleotide immediately 5' from TTAA is C and the nucleotide immediately 3' from TTAA is C. Similarly, similar oligos were synthesized with 13 bp spacer sequences modified to mutate the 5' and 3' flanking nucleotide immediately adjacent to the TTAA integration sequence to T and A, respectively, to generate a TTTAAA integration site (SEQ ID No. 176), or to C and A, respectively, to generate a CTTAAA integration site (SEQ ID No. 177).177), or for a T and a G, respectively, to generate a TTTAAG integration site (SEQ ID No. 178), or for a C and a G, respectively, to generate a CTTAAG integration site (SEQ ID No. 179). Petition 870250091181, dated 06 / 10 / 2025, pp. 201 / 220 106 / 112
[00234] Each reporter plasmid and donor plasmid were co-transfected into cells HEK293T cells were transfected with the GFP1 expression plasmid TAL-PBx on the Right (SEQ ID No. 77). As a negative control, the GFP1 expression plasmid TAL-PBx on the Left (SEQ ID No. 78), which does not recognize the target sequence GFP1 on the Right, was transfected in place of the GFP1 expression plasmid TAL-PBx on the Right. HEK293T cells were seeded in 24-well plates in 500 µl of DMEM medium supplemented with 10% FBS. The following day, a transfection mixture containing 50 ng of the TAL-ssSPB expression vector, 225 ng of the reporter plasmid, 225 ng of the donor plasmid, and 1 µl of JetPrime transfection reagent in a total volume of 50 µl of JetPrime buffer was prepared. The mixture was added to HEK293T cells and the cells were incubated for four days at 37 °C and 5% CO2, splitting them 1:6 on the first day. The percentage of cells positive for site-specific transposition to the target site for the various constructs was determined by FACS analysis on day 4.
[00235] When reporter and donor plasmids are cotransfected into cells along with TAL-PBx, TAL-PBx catalyzes the excision of the donor plasmid transposon and its site-specific integration into the TTAA target site of the reporter plasmid. After site-specific transposition, transcription, splicing, and translation, a reconstituted GFP coding sequence is produced (DNA SEQ ID No. 75; amino acid SEQ ID No. 76) and fluorescence can be detected. The percentage of cells positive for site-specific transposition at the target for the various spacer length constructs was determined by FACS analysis, and the results are shown in Table 8.
[00236] As shown in Table 8, GFP1 TAL-PBx on the Right catalyzed the site-specific transposition leading to the GFP signal above the Petition 870250091181, dated 06 / 10 / 2025, pp. 202 / 220 107 / 112 background levels with all target sites. Target sites TTTAAA resulted in a higher GFP signal than target sites CTTAAC and CTTAAG. Target sites CTTAAA and TTTAAG resulted in the highest GFP signal. GFP1 TAL-PBx on the Left did not result in any GFP signal above the bottom using the specific GFP1 reporters on the Right. Table 8 Percentage of GFP Cells GFP1 TAL-ssSPB Right GFP1 TAL-ssSPB Left Replicate 1 Replicate 2 Replicate 3 Average Replicate 1 Replicate 2 Replicate 3 Average 12bp (CTTA AC) 15.8 15.0 15.2 15.3 3.5 3.1 3.3 3.3 12bp (TTTA AA) 30.8 30.6 31.0 30.8 3.9 3.9 3.9 3.9 12bp (CTTA AA) 33.4 35.5 34.8 34.6 3.5 3.6 3.4 3.5 13bp (CTTA AC) 20.6 23.0 19.0 20.9 4.7 5.3 4.6 4.8 13bp (TTTA AA) 28.2 26.4 26.5 27.0 4.3 4.0 3.3 3.9 13bp (CTTA AA) 34.9 32.6 31.0 32.8 4.1 5.1 5.8 5.0 13bp (TTTA AG) 33.2 32.3 31.2 32.2 4.5 4.5 5.2 4.7 13bp (CTTA AG) 24.9 24.1 25.1 24.7 5.0 4.7 4.8 4.8 Example 6: Demonstration of Site-Specific Transposition Using TAL Matrix Compositions - Transposase piggyBac (ss-SPB) (TAL-PBxs)
[00237] Based on the results of Example 5, a second set of four different LPA target sequences found naturally in genomic DNA (SEQ ID Nos. 85-88) was cloned into the episomal reporter plasmid described in Example 4. Like the first set of targets evaluated in Example 4, each of the target sequences in the second set has sites Petition 870250091181, dated 06 / 10 / 2025, pp. 203 / 220 108 / 112 TAL connectors of 10 bp and 12 bp or 13 bp spacers on both sides of TTAA. Furthermore, each of the target sequences in the second set comprises spacer sequences such that the nucleotide immediately 5' from TTAA is T and the nucleotide immediately 3' from TTAA is A, to generate a TTTAAA integration site, or such that the nucleotide immediately 5' from TTAA is C and the nucleotide immediately 3' from TTAA is A, to generate a CTTAAA integration site. Additionally, since a thymidine is not immediately 5' from all LPA target sites, the N-terminal domain of TAL has been mutated to not require any specific nucleotide 5' from the binding site. These mutations were introduced into the wild-type TAL sequence by replacing the amino acid sequence QWS at positions 79-81 of SEQ ID NO: 31 with YH to generate the NT-βN variant (SEQ ID NO: 34).
[00238] The TAL Matrices were constructed to target these TAL binding sites using the design criteria described in this document or as set out below.
[00239] TAL Matrix pairs were designed targeting four specific sequences of 10 bp right- and left pairs within the second set of four LPA target sites. For each of the targets, multiple TAL Matrix pairs were designed using 12 bp or 13 bp spacers.
[00240] The left and right target sequences, along with the 5' nucleotide used to generate TAL arrays that target the LPA gene, are shown in Table 9. Table 9: Illustrative TAL Matrices that have LPA as a Target PAR LPA # TARGET SEQUENCE ON THE LEFT TARGET SEQUENCE ON THE RIGHT 5 GTATCCGCAGA (SEQ ID NO: 100) CGCIIIICTAC (SEQ ID NO: 101) TATCCGCAGAG (SEQ ID NO: 102) GCIII ICTACA (SEQ ID NO: 103) 6 AGTATGATAAC (SEQ ID NO: 104) TCTGCTTTCTT (SEQ ID NO: 105) GTATGATAACT (SEQ ID NO: 106) CTGCTTTCTTG (SEQ ID NO: 107) Petition 870250091181, dated 06 / 10 / 2025, pp. 204 / 220 109 / 112 LPA PAR # LEFT TARGET SEQUENCE RIGHT TARGET SEQUENCE 7 CGTTTGCTACT (SEQ ID NO: 108) CTTAATAGATT (SEQ ID NO: 109) GTTTGCTACTT (SEQ ID NO: 110) TTAATAGATTA (SEQ ID NO: 111) 8 GAAGGGAGTGA (SEQ ID NO: 112) GTACAAGTGTC (SEQ ID NO: 113) AAGGGAGTGAT (SEQ ID NO: 114) TACAAGTGTCA (SEQ ID NO: 115)
[00241] The eight combinations of left and right pairs were used to design and construct the Left TAL Matrices LPA LPAL5.1, LPAL5.2, LPAL6.1, LPAL6.2, LPAL7.1, LPAL7.2, LPAL8.1, and LPAL8.2 (SEQ ID Nos. 127, 129, 131, 133, 135, 137, 139, and 141, respectively) and the Right TAL Matrices LPA LPAR5.1, LPAR5.2, LPAR6.1, LPAR6.2, LPAR7.1, LPAR7.2, LPAR8.1, and LPAR8.2 (SEQ ID Nos. 128, 130, 132, 134, 136, 138, 140, and 142, respectively), as described in Example 2.
[00242] The TAL-PBx fusion constructs were prepared as follows: an expression plasmid was synthesized containing, from the 5' to 3' direction: a CMV promoter, a T7 promoter, a Kozak sequence, a Flag tag 3x (SEQ ID NO: 65), an SV40 NLS (SEQ ID NO: 66), the N-terminal Delta 152 TAL domain (SEQ ID NO: 31) of the NT-BN TAL variant (SEQ ID NO: 34), two BsmBI-type IIS restriction enzyme sites, the C-terminal +73 TAL domain (SEQ ID NO: 79), a GGGS linker, delta 1-85 PBx (comprising an 85-amino acid N-terminal deletion and mutations at R372A, K375A, D450N in the Super piggyBac transposase codon sequence; SEQ ID NO: 9) and a sequence of bGH polyadenylation.
[00243] Cloning a left- or right-flanked BsmBI TAL array into the BsmBI sites of the expression plasmid results in in-frame fusion of the TAL array and the PBx coding sequence via a linkage sequence, generating full-length TAL-PBx constructs. All coding sequences used were codon-optimized for human expression using GeneArt algorithms (Thermo Fisher).
[00244] The sixteen TAL matrices flanked with BsmBI ends were cloned into the BsmBI restriction sites of the plasmid. Petition 870250091181, dated 06 / 10 / 2025, pp. 205 / 220 110 / 112 of the expression described above to generate sixteen TAL-PBx constructs: LPAL5.1, LPAL5.2, LPAL6.1, LPAL6.2, LPAL7.1, LPAL7.2, LPAL8.1 and LPAL8.2 Left TAL-PBxs (SEQ ID Nos. 154, 156, 158, 160, 162, 164, 166 and 168, respectively) and LPAR5.1, LPAR5.2, LPAR6.1, LPAR6.2, LPAR7.1, LPAR7.2, LPAR8.1 and LPAR8.2 Right TAL-PBxs (SEQ ID Nos. 155, 157, 159, 161, 163, 165, 167 and 169, respectively), as described in Example 3.
[00245] In addition, the TAL LPAL1 (SEQ ID No. 116) and LPAR1 (SEQ ID No. 117) matrices, as described in Example 2, were cloned into the expression plasmid described above in this Example 6 to generate the TAL-PBx LPAL1 v2 (SEQ ID No. 170) and LPAR2 v2 (SEQ ID No. 171) constructs. A. Site-Specific Transposition of Episomal Target
[00246] The activity of the novel mutant TAL-PBx fusions was determined using their respective episomal split GFP splicing reporters. Briefly, each reporter plasmid and the donor plasmid were cotransfected into HEK293T cells with the corresponding TALPBx expression plasmid. Approximately 120,000 HEK293T cells were seeded in 24-well plates in 500 µl of DMEM medium supplemented with 10% FBS. The following day, a transfection mixture containing 50 ng of the TAL-PBx expression vector, 225 ng of the reporter plasmid, 225 ng of the donor plasmid, and 1 µl of JetPrime transfection reagent, in a total volume of 50 µl of JetPrime buffer, was prepared. This mixture was added to HEK293T cells, which were incubated for four days at 37 °C and 5% CO2, splitting the cells 1:6 on the first day. The percentage of GFP-positive cells was determined for each sample. The results are shown in Table 10. Table 10 On Target Off Target Reply 1 Reply 2 Average Reply 1 Reply 2 Average LPA L1 / R1 11.2 11.5 11.4 2.0 1.8 1.9 Petition 870250091181, dated 06 / 10 / 2025, pp. 206 / 220 111 / 112 On Target Off Target Reply 1 Reply 2 Average Reply 1 Reply 2 Average LPA L1 v2 / R1 v2 22.6 26.2 24.4 2.0 2.1 2.1 LPA L5.1 / R5.1 (13 bp spacers) 4.9 5.1 5.0 1.1 1.0 1.1 LPA L5.2 / R5.2 (12 bp spacers) 4.0 4.6 4.3 1.0 1.0 1.0 LPA L6.1 / R6.1 (13 bp spacers) 19.3 22.4 20.9 1.2 1.3 1.2 LPA L6.2 / R6.2 (12 bp spacers) 5.9 6.5 6.2 0.8 1.1 0.9 LPA L7.1 / R7.1 (13 bp spacers) 7.4 7.5 7.4 2.2 2.6 2.4 LPA L7.2 / R7.2 (12 bp spacers) 3.3 3.2 3.2 1.7 2.1 1.9 LPA L8.1 / R8.2 (13 bp spacers) 31.5 27.9 29.7 1.9 2.1 2.0 LPA L8.2 / R8.2 (12 bp spacers) 12.5 12.1 12.3 1.6 1.8 1.7
[00247] As seen in Table 10, targets that target TAL-ssSPB 1, 6, and 8 resulted in the greatest transposition. Furthermore, TALssSPBs using 13 bp spacers resulted in greater edit than those using 12 bp spacers. B. Site-Specific Transposition of Genomic Target
[00248] After confirming that the newly designed LPA TALs were functional and recognized their target sequence, the TAL-PBx constructs were used to edit the endogenous genomic targets of LPA in Huh7, an immortalized hepatocyte cell line. Briefly, 100,000 cells were plated the day before transfections in 24-well plates in RPMI medium + 10% FBS. The following day, 0.5 pg or 1 pg of mRNA encoding the TAL-ssSPB pair of target LPA 1 (SEQ ID NOs: 143 and 144) was mixed with 0.5 pg or 1 pg of Messenger Max reagent, respectively, to generate ssSPB-mRNA lipid complexes. Simultaneously, 450 pg of a transposon donor vector (SEQ ID NO: 80) were mixed with 0.5 pl of P3000 reagent and 1 pl of lipofectamine 3000 to generate DNA lipid complexes. 50 pl of ssSPB mRNA lipid complexes and 50 pl of DNA lipid complexes were also produced. Petition 870250091181, dated 06 / 10 / 2025, pages 207 / 220 112 / 112 were delivered to the cells, which were incubated at 37 °C.
[00249] To evaluate the site-specific integration of the transposon donor into LPA loci, genomic DNA was extracted from transfected cells two days after transfections and analyzed by droplet digital PCR (ddPCR) using a probe-based detection scheme. A primer that binds to the transposon was paired with a primer that binds to LPA genomic DNA near the TTAA integration site. Therefore, an amplicon should only be generated after site-specific transposition to an LPA locus. Since integration is non-directional, two assays were designed for each LPA target to detect transposon integration in both forward and reverse directions. As a negative control, genomic DNA was extracted from non-transferred cells and used as a template in the ddPCR reaction to demonstrate the specificity of the primer / probe sets. The results are shown in Table 11. Table 11 Percentage of Haploid Genomes Edited in Target LPA Direct Integration Reverse Integration Replication 1 Replication 2 Average Replication 1 Replication 2 Average Not transferred 0.00 0.00 0.00 0.00 0.00 0.00 0.5μg mRNA 0.33 0.42 0.37 0.52 0.57 0.55 1μg mRNA 0.59 0.71 0.65 1.00 1.06 1.03
[00250] As shown in Table 11, amplicons corresponding to direct and / or reverse transposon integration were detected from genomic DNA isolated from cells transfected with LPA TAL-PBx constructs along with the transposon, providing direct evidence of genomic integration at LPA loci. Petition 870250091181, dated 06 / 10 / 2025, pages 208 / 220
Claims
1 / 6 Claims 1. FUSION PROTEIN, characterized by comprising a DNA-targeting domain and a transposase domain comprising the sequence set forth in SEQ ID NO: 4, wherein the DNA-targeting domain binds to a nucleic acid sequence encoding an LPA repeat element.
2. Fusion protein, according to claim 1, characterized by the DNA-targeting domain comprising one, two, or three zinc finger motifs.
3. FUSION PROTEIN, according to claim 1, characterized by the DNA-targeting domain comprising one or more TAL domains.
4. METHOD, according to claim 3, characterized in that the TAL domain comprises the sequence set forth in any of the SEQ ID Nos: 35 to 38.
5. FUSION PROTEIN, according to any one of claims 1 to 4, characterized by the DNA-targeted domain binding to a nucleic acid sequence encoding a kringle domain repeat element or to an intron adjacent to a sequence encoding a kringle domain repeat element in the LPA gene.
6. Fusion protein, according to any one of claims 1 to 5, characterized in that the transposase domain and the DNA-targeting domain are connected by a linker.
7. Fusion protein, according to claim 6, characterized by the linker comprising the sequence GGGGS (SEQ ID NO: 181).
8. FUSION PROTEIN, according to any one of claims 1 to 7, characterized in that the DNA-targeting domain is inserted into the N-terminal of the transposase domain at a position after the 82nd amino acid and before the 105th amino acid of SEQ ID NO:
4.
9. FUSION PROTEIN, according to any one of claims 1 to 7, characterized by the DNA-targeted domain replacing one or more amino acids in the transposase domain between, and including, the 83rd amino acid and the 105th amino acid of SEQ ID NO:
4.
10. FUSION PROTEIN, according to any one of claims 1 to 9, characterized by the transposase domain comprising an N-terminal deletion of amino acids 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, 1-100, 1-101, 1-102 or 1-103.
11. Fusion protein, according to any one of claims 1 to 10, characterized by the transposase domain comprising the sequence set forth in any one of the SEQ ID Nos: 7 to 27.
12. FUSION PROTEIN, according to any one of claims 1 to 11, characterized by the transposase domain comprising (a) at least one mutation selected from the group consisting of M185R, M185K, D197K, D197R, D198K, D198R, D201K, and D201R; or (b) at least one mutation selected from the group consisting of L204D, L204E, K500D, K500E, R504E, and R504D.
13. POLYNUCLEOTIDE, characterized by comprising a nucleic acid sequence encoding the fusion protein, as defined in any one of claims 1 to 12.
14. VECTOR, characterized by comprising the polynucleotide, as defined in claim 13.
15. METHOD FOR INTEGRATING A TRANSGENE into a genomic target site of a cell, the method characterized by comprising introducing into the cell the fusion protein, as defined in any of claims 1 to 12, and a transposon, wherein the transposon comprises, in Petition 870250091181, dated 06 / 10 / 2025, pp. 210 / 220 3 / 6, in the order of 5' to 3': a 5'ITR, the transgene, and a 3'ITR.
16. METHOD, according to claim 15, characterized by the transposon further comprising an exogenous promoter between the 5' ITR and the transgene.
17. METHOD, according to any one of claims 15 to 16, characterized by the transgene encoding a detectable marker.
18. METHOD, according to claim 17, characterized in that the detectable marker is GFP.
19. METHOD, according to any one of claims 15 to 16, characterized in that the transgene is a gene that (a) is not expressed by the cell before the introduction of the fusion protein and transposon or (b) exhibits decreased, insufficient and / or altered expression by the cell before the introduction of the fusion protein and transposon.
20. METHOD, according to any one of claims 15 to 19, characterized in that the genomic target site is located in the LPA gene.
21. METHOD, according to any one of claims 15 to 19, characterized in that the genomic target site is located on a repetitive element.
22. METHOD, according to claim 21, characterized in that the repeating element is an LPA repeating element.
23. METHOD, according to any one of claims 15 to 19, characterized in that the genomic target site is located in an intron of a gene.
24. METHOD, according to claim 23, characterized in that the genomic target site is located in the intron of the LPA gene.
25. METHOD, according to any one of claims 15 to 24, characterized in that the cell is in vivo.
26. METHOD FOR MODIFYING THE GENOME OF A CELL, the method characterized by comprising: providing the cell with the fusion protein, as defined in any one of claims 1 to 12, wherein the cell comprises a modified binding site comprising, in the order of 5' to 3', the sequence of a target site for the DNA-targeted domain, a first spacer, a TTAA-targeted integration site for SPB, a second spacer, and the reverse complement of the target site sequence for the DNA-targeted domain.
27. METHOD, according to claim 26, characterized by the target integration site comprising the TTAA sequence.
28. METHOD, according to claim 26, characterized by the target integration site comprising the nucleic acid sequence set forth in any of the SEQ ID NOS: 81 to 88.
29. INTEGRATION CASSETTE for site-specific transposition of a nucleic acid in the genome of a cell, characterized by comprising a nucleic acid comprising or consisting of a central transposon integration site TTAA sequence flanked by an upstream TAL matrix target sequence and a downstream TAL matrix target sequence, wherein each of the upstream and downstream TAL matrix target sequences is separated from the TTAA sequence by 12 or 13 base pairs.
30. INTEGRATION CASSETTE, according to claim 29, characterized by the integration site comprising the TTAA sequence.
31. INTEGRATION CASSETTE, according to claim 29, characterized by the integration site comprising the nucleic acid sequence set forth in any of the SEQ ID NOS: 81 to 88.
32. INTEGRATION CASSETTE, according to any of claims 29 to 31, characterized in that each of the upstream and downstream target site sequences of Petition 870250091181, dated 06 / 10 / 2025, p. 212 / 220 5 / 6 is the same.
33. INTEGRATION CASSETTE, according to any one of claims 29 to 31, characterized in that each of the upstream and downstream TAL matrix target site sequences are different.
34. INTEGRATION CASSETTE, according to any one of claims 29 to 33, characterized by each of the upstream and downstream TAL Matrix target sites targeting a 7 to 30 bp sequence of an LPA repeat element.
35. CELL, characterized by comprising the integration cassette, as defined in any one of claims 29 to 34, stably integrated into the cell genome.
36. METHOD FOR SITE-SPECIFIC TRANSPOSITION of a DNA molecule into the genome of a cell, characterized by comprising introducing into the cell, as defined in claim 35: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase; wherein the fusion protein is expressed in the cell; and b) a DNA molecule comprising a transposon; wherein the expressed fusion protein integrates the transposon by site-specific transposition into the TTAA integration site of the stably integrated integration cassette.
37. METHOD FOR GENERATING AN ENGINEERED CELL by site-specific transposition, characterized by comprising introducing into the cell, as defined in claim 31: a) a nucleic acid encoding a fusion protein comprising a DNA-binding domain and a transposase; wherein the fusion protein is expressed in the cell; and b) a DNA molecule comprising a transposon; in Petition 870250091181, dated 06 / 10 / 2025, pp. 213 / 220 6 / 6, wherein the expressed fusion protein integrates the transposon by site-specific transposition into the TTAA integration site of the stably integrated integration cassette, thereby generating the engineered cell.
38. METHOD, according to any one of claims 36 to 37, characterized in that the integration site comprises the TTAA sequence.
39. METHOD, according to any one of claims 36 to 37, characterized by the integration site comprising the nucleic acid sequence set forth in any one of the SEQ ID Nos: 81 to 88. Petition 870250091181, dated 06 / 10 / 2025, pp. 214 / 220