Genome editing using evolved integrases
The use of mutated serine integrases addresses the inefficiencies of existing methods by integrating large DNA cargos at attachment sites, achieving high-efficiency integration of large DNA cargos at attachment sites, enabling high-efficiency integration of large DNA cargos, and large DNA integration of large DNA cargos, and large DNA cargos, and large DNA integration of large DNA integration of large DNA cargos at attachment sites, achieving high-efficiency integration of large DNA cargos up to 20 kb without double-strand breaks.
Patent Information
- Application Number
- PCT/US2025/033171
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-19
- Filing Date
- 2025-06-11
- Publication Date
- 2025-12-18
AI Technical Summary
Current integrase-mediated integration methods are inefficient for inserting large DNA cargos and lack diversity in integration sites, requiring undesired double-strand breaks that trigger DNA damage pathways.
Development of mutated serine integrases, such as Bxb1 and PhiC31, with specific amino acid mutations and nuclear localization sequences, enabling high-efficiency integration of large DNA cargos at attachment sites without causing double-strand breaks, using Cas9-directed reverse transcription to target 'safe' loci.
Achieves integration rates of up to 90% of target chromosomes with large DNA cargos ranging from 0.1 kb to 20 kb, enhancing therapeutic gene delivery and adaptability across various host cells.
Smart Images

Figure US2025033171_18122025_PF_FP_ABST
Abstract
Description
GENOME EDITING USING EVOLVED INTEGRASESFIELD
[0001] The present disclosure relates to, in part, recombinant genome editing enzymes and methods of preparing the same and uses thereof, e.g., to modify host cell genomes for protein expression.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims benefit of and priority to U.S. Non-provisional Application No. 63 / 735,975, filed December 19, 2024, and U.S. Non-provisional Application No. 63 / 658,836, filed June 11 , 2024, the entire contents of each are incorporated by reference herein.SEQUENCE LISTING
[0003] The instant application contains a sequence listing, which has been submitted in XML format via Patentcenter. The XML file is named “138774-5002_Seque nce_Listi ng . xm I, ” which was created on June 11 , 2025, and is approximately 212,992 bytes in size, the entire contents of which are incorporated herein by reference in their entirety.BACKGROUND
[0004] Serine integrases are a type of site-specific recombinase (SSR) that are capable of actively inserting large DNA cargos at defined DNA sequences. During integrase recombination between attachment (att) sites, e.g., phage attachment site, attP (donor sequence), and bacterial attachment site, attB (target sequence). The DNA is protected during cleavage, strand exchange, and re-ligation. No synthesis or degradation of DNA occurs and no sequence homology is required for the donor DNA. This process does not rely on rate-limiting host factors and does not provoke error-prone DNA repair processes caused by nuclease cleavage that can induce a p53-mediated DNA damage response and lead to unwanted selection for cells with an impaired p53 pathway. A limitation of serine integrase technology is that the att site must first be inserted at the target sequence (e.g., in a host cell genome). Fortunately, the small size of att sites has permitted their insertion using HDR into the genomes of mice and human cells to act as a insertion sites for integrase-based insertion.
[0005] By overcoming the lack of efficiency of current methods, numerous applications would be accelerated. Tailoring a therapy to correct an individual’s unique genetic abnormality or mutation is challenging, hence a full replacement of the coding sequence of a recessive gene is often needed fora single,common therapy to be effective. As treatments for increasingly complex diseases are undertaken, multi-gene insertions will be necessary requiring larger cargos beyond the limits of current virus-based approaches.
[0006] There remains a need in the art to improve the efficiency of integrase-mediated integration, including the development of improvements to the enzymatic machinery involved and the ability to diversify the repertoire of integration sites, as well as improved methods for therapeutic protein expression.SUMMARY
[0007] Accordingly, the present disclosure provides, in part, the ability to deliver large transgenes to a single genomic sequence with high efficiency using integrases (e.g., as compared to wild-type integrases). In embodiments, the integrases are mutated serine integrases (e.g., Bxb1 and PhiC31) which catalyze the integration of large DNA cargos at attachment (atf) sites. In embodiments, targeting to ati sites in the genome utilize Cas9-directed reverse transcription, where integrases target “safe” loci. In embodiments, insertion avoids double-strand breaks, which would ordinarily trigger DNA damage pathways (e.g., p53).
[0008] Described herein, in embodiments, are Bxb1 integrases comprising an amino acid sequence that is at least 90% identical to amino acids 1-480 of SEQ ID NO: 1, an amino acid sequence that is at least 90% identical to amino acids 1-488 of SEQ ID NO: 1 , or an amino acid sequence that is at least 90% identical to SEQ ID NO: 1 , where the Bxb1 integrase comprises one or more mutants at one or more amino acid positions selected from the list consisting of: 75, 76, 158, 232, 234, 236, 237, 257, 314, 316, 318, 322, 323, 325; and one or more mutants at one or more amino acid positions selected from the list consisting of: 5, 14, 20, 24, 29, 35, 40, 45, 49, 50, 51, 60, 68, 69, 70, 73, 74, 78, 84, 86, 87, 100, 105, 116, 124, 183, 197, 207, 208, 209, 229, 261 , 267, 273, 273, 287, 291 , 333, 342, 343, 347, 361 , 368, 375, 435, 449, 453, 462, 483, 494 and / or 4, 18, 34, 36, 42, 46, 61 , 62, 63, 67, 79, 85, 88, 89, 90, 92, 95, 99, 106, 110, 111 , 119, 122, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174, 175, 178, 179, 181, 187, 189, 191 , 203, 218,223, 231 , 239, 248, 251, 254, 264, 268, 272, 278, 280, 281 , 282, 283, 285, 288, 292, 295, 302, 306, 307,311, 313, 319, 321 , 328, 331 , 332, 334, 353, 355, 359, 360, 362, 369, 370, 380, 388, 397, 398, 405, 409,411, 414, 415, 416, 419, 425, 428, 434, 444, 461 , 463, 466, 468, 476, 479, 480, 484, 487, 488, 489, 496, and / or 499.
[0009] In embodiments, the Bxb1 integrase comprises an amino acid sequence that is at least 90% identical to amino acids 1-480 of SEQ ID NO: 1 ; an amino acid sequence that is at least 90% identical to amino acids 1 -488 of SEQ ID NO: 1 ; or an amino acid sequence that is at least 90% identical to SEQ ID NO: 1 and comprises one or more amino acid mutations at a position selected from 4, 5, 14, 18, 20, 24, 29, 34,35, 36, 40, 42, 45, 46, 49, 50, 51, 60, 61, 62, 63, 67, 68, 69, 70, 73, 74, 75, 76, 78, 79, 84, 85, 86, 87, 88, 89, 90, 92, 95, 99, 100, 105, 106, 110, 111, 116, 119, 122, 124, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174, 175, 178, 179, 181, 183, 187, 189, 191, 197, 203, 207, 208, 209, 218, 223, 229, 231,232, 234, 236, 237 239, 248, 251, 254, 257, 261, 264, 267, 268, 272, 273, 278, 280, 281, 282, 283, 285,287, 288, 291, 292, 295, 302, 306, 307, 311, 313, 314, 316, 318, 319, 321, 322, 323, 325, 328, 331, 332,333, 334, 342, 343, 347, 353, 355, 359, 360, 361, 362, 368, 369, 370, 375, 380, 388, 397, 398, 405, 409,411, 414, 415, 416, 419, 425, 428, 434, 435, 444, 449, 453, 461, 462, 463, 466, 468, 476, 479, 480, 483,484, 487, 488, 489, 494, 496, and 499, or a position corresponding thereto; or an amino acid sequence that is at least 90% identical to SEQ ID NO: 1 and comprises at least two amino acid substitutions at positions selected from (a) and (b), (b) and (c), or (a) and (c): (a) 75, 76, 158, 232, 234, 236, 237, 257, 314, 316, 318, 322, 323, 325, or a position corresponding thereto; and (b) 5, 14, 20, 24, 29, 35, 40, 45, 49, 50, 51, 60, 68, 69, 70, 73, 74, 78, 84, 86, 87, 100, 105, 116, 124, 183, 197, 207, 208, 209, 229, 261, 267, 273, 273, 287, 291, 333, 342, 343, 347, 361, 368, 375, 435, 449, 453, 462, 483, 494, or a position corresponding thereto; and / or (c) 4, 18, 34, 36, 42, 46, 61, 62, 63, 67, 79, 85, 88, 89, 90, 92, 95, 99, 106, 110, 111, 119, 122, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174, 175, 178, 179, 181, 187, 189, 191, 203, 218,223, 231, 239, 248, 251, 254, 264, 268, 272, 278, 280, 281, 282, 283, 285, 288, 292, 295, 302, 306, 307,311, 313, 319, 321, 328, 331, 332, 334, 353, 355, 359, 360, 362, 369, 370, 380, 388, 397, 398, 405, 409,411, 414, 415, 416, 419, 425, 428, 434, 444, 461, 463, 466, 468, 476, 479, 480, 484, 487, 488, 489, 496, and / or 499, or a position corresponding thereto.
[0010] In embodiments, the Bxb1 mutants are selected from L4I, V5I, DUN, E20K, E20Q, E24K, L29F, G34D, W35L, D36A, V40I, V40A, E42K, D45G, A49T, V50I, D51N, D51E, D51Y, N60S, L61F, R63K, F67S, E68K, E69A, E69D, Q70P, D73G, V74A, V74I, V74M, I75V, V76I, Y78H, Y78N, T84S, S86T, I87V, I87L, H89G, Q92H, H95Y, D99N, H100N, H100Y, V105A, V105I, H111P, T116P, A119S, V122M, A124S, A145T, S157G, T166I, V179A, V179I, R181K, E183L, V187I, H189N, P197T, H203Y, R207Q, R208S, G209V, E229K, A232G, A232S, A232R, A232T, A232V, A232P, A232Q, T233G, T233W, T233R, T233Y, T233D,T233G, T233N, T233H, T233S, T233Q, T233A, T233C, A234N, A234G, A234S, A234T, A234H, A234F,A234S, K236R, K236S, R237K, R237Q, R237N, R237V, R237C, M239I, A248T, N251K, T254S, D257K,A261T, A261V, V264A, E267D, R272Q, E273D, E273K, A280T, T285A, R287P, A288V, A288T, A291T,A311V, K313R, F314M, F314L, F314N, F314R, F314K, G316R, G316W, G318H, G318P, G318S, G318N, G318K, R319K, R319G, H321P, H321Q, H321 K, H321L, H321R, H321Y, H321T, H321S, P322A, P322R, P322G, P322L, R323L, R323G, R323Y, R323I, R325K, R325Q, R325Y, F331S, P332H, K333N, H334R,M342T, M342V, A343T, A347E, A347V, V353I, D355N, D359N, A360T, E361 D, R362K, V368A, A369P, A369E, A369T, A369S, V375I, V380I, T388M, A396P, A398S, R409H, A411V, A414V, A415S, R416K, A425T, E434G, T435A, R444L, A449V, T453I, T453A, L462M, T463I, V466M, G468D, L479I, Q480STOP, E483K, Q484K, R487R, del L488 (del L488 frameshift), R494S, R494Q, H496N, and / or M499T, or a position corresponding thereto, relative to SEQ ID NO: 1. In embodiments, mutations with same amino acids relate to silent mutations that do not affect amino acid profile of the integrase.
[0011] In embodiments, one or more of the following amino acids are not mutated relative to SEQ ID NO: 1 : V5, S18, E24, V40, V46, D51, A62, E69, R79, R85, R88, L90, Q92, S106, A110, A130, E133, 1137, R140, K153, G156, P160, L164, L174, V175, P178, V187, Q191 , G209, A218, R223, S231, N251 , P268, L278, A280, E281, L282, V283, R287, P292, P295, L302, V306, C307, K313, S328, K333, D359, E361 , G370, V375, R397, A405, A414, E419, A425, S428, T435, R461 , V466, G468, F476, E483, G489, and R494.
[0012] In embodiments, the Bxb1 integrase comprises one or more mutations and / or deletions, and / or one or more combinations of mutations and / or deletions as described in Table 1 and / or Table 2 relative to SEQ ID NO: 1.
[0013] In embodiments, the Bxb1 integrase comprises one or more mutations and / or deletions that alters specificity at -11, -10, and / or -9 positions of a 3 bp helix target site compared to a wild-type Bxb1 of SEQ ID NO: 1. In embodiments, the Bxb1 integrase comprises one or more mutations and / or deletions that alters specificity at the -10 position to thymine (T) compared to a wild-type Bxb1 of SEQ ID NO: 1. In embodiments, the Bxb1 integrase comprises one or more mutations and / or deletions that alters specificity at -10 and / or -9 positions of a 3 bp helix target site compared to a wild-type Bxb1 of SEQ ID NO: 1. In embodiments, the Bxb1 integrase comprises one or more mutations and / or deletions that alters specificity at -19 through -12 positions of a DNA binding site compared to a wild-type Bxb1 of SEQ ID NO: 1.
[0014] In embodiments, the Bxb1 integrase comprises an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 , and one or more amino acid mutations at positions selected from L4, E42, R57, R63, V76, H111 , V122, V187, K313, R319, A341 , D359, A369, A398, R416, A425, L479, and M499, relative to SEQ ID NO: 1.
[0015] Described herein, in embodiments, are PhiC31 integrases comprising an amino acid sequence that is at least 90% identical to amino acids 1-604 of SEQ ID NO: 2 and one or more amino acid mutations at positions selected from 1, 2, 12, 14, 18, 24, 32, 36, 41 , 43, 44, 45, 49, 51, 55, 58, 74, 77, 96, 103, 107, 117, 153, 176, 196, 197, 199, 200, 228, 229, 230, 231 , 235, 238, 240, 252, 254, 255, 260, 262, 264, 266,269, 274, 278, 302, 319, 320, 322, 331, 333, 340, 344, 345, 346, 347, 351, 353, 355, 359, 362, 364, 378,382, 393, 396, 397, 399, 406, 410, 424, 429, 431, 436, 438, 445, 448, 449, 450, 452, 457, 468, 472, 478,498, 499, 501, 505, 512, 514, 516, 517, 520, 527, 535, 536, 549, 551, 552, 553, 563, 568, 580, 585, 586,587, 590, 592, 594, 600, 603, 604, 609, 612, and / or 616, or a position corresponding thereto.
[0016] In embodiments, the PhiC31 integrase comprises one or more mutations, N-terminal additions, and / or one or more combinations of mutations and / or N-terminal additions as described in Table 1, Table 3, and / or Table 4 relative to SEQ ID NO: 2.
[0017] In embodiments, integrases herein (e.g., Bxb1 and / or PhiC31) comprise one or more nuclear localization sequences (NLSs), which improve the efficiency of the integrase relative to a cognate integrase lacking the NLS. In embodiments, the one or more NLSs is fused at the N-terminus, at the C-terminus, at both the N-terminus and C-terminus, or within the integrase sequence. In embodiments, the NLS is or comprises one or more sequence selected from: PKKKRKV (SEQ ID NO: 57), NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 58); PAAKRVKLD (SEQ ID NO: 59), RQRRNELKRSF (SEQ ID NO: 60); NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 61),RMRKFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 62), VSRKRPRP (SEQ ID NO: 63), PPKKARED (SEQ ID NO: 64), PQPKKKPL (SEQ ID NO: 65), SALIKKKKKMAP (SEQ ID NO: 66), DRLRR (SEQ ID NO: 67), PKQKKRK (SEQ ID NO: 68), RKLKKKIKKL (SEQ ID NO: 69), REKKKFLKRR (SEQ ID NO: 70), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 71), RKCLQAGMNLEARKTKK (SEQ ID NO: 72), DPKKKRKVDPKKKRKVDPKKKRKV (SEQ ID NO: 73), KRTADGSEFESPKKKRKV (SEQ ID NO: 74), KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75), GSHHHHHHGSGPKKKRKV (SEQ ID NO: 76), and GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0018] In embodiments, integrases herein (e.g., Bxb1 and / or PhiC31) comprise one or more one or more stabilization sequences, optionally located at the N-terminus, C-terminus, or both the N-terminus and C-terminus. In embodiments, the one or more stabilization sequences comprises an amino acid sequence that influences stability, solubility, folding, aggregation, degradation, purification, isolation, tracking expression and / or subcellular localization, relative to a cognate protein lacking the one or more stabilization sequences. In embodiments, the one or more stabilization sequences is or comprises about 5 amino acids to about 300 amino acids in length. In embodiments, the one or more stabilization sequences is or comprises the amino acid sequence of any one of SEQ ID NOs: 157 and 162-165, or an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes thereto.
[0019] In embodiments, integrases herein (e.g., Bxb1 or PhiC31) are fusions with one or more DNA- binding domains (e.g., to improve targeting to one or more DNA sequences). In embodiments, the one or more DNA-binding domains comprise one or more TAL domains, one or more Cas proteins, and / or one or more zinc finger protein domains. In embodiments, integrases herein (e.g., Bxb1 or PhiC31) comprise about 1 to about 30 ZFPs, e.g., to target a “left” and / or a “right” portion of an aft sequence (e.g., either endogenous genomic sequence, or an exogenous sequence knocked-in to the genome).
[0020] In embodiments, the genetic cargoes are single or multi-gene cargos. In embodiments, the integration rates are as high as 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% or more of target chromosomes in host cells. In embodiments, the hyperactive integrases (e.g., PhiC31 and Bxb1 serine integrases) insert therapeutic DNA cargoes ranging in size from about 0.1 kb to about 20 kb or more (e.g., about 0.5 kb, about 1 kb, about 2 kb, about 3 kb, about 4 kb, about 5 kb, about 6 kb, about 7 kb, about 8 kb, about 9 kb, about 10 kb, about 11 kb, about 12 kb, about 13 kb, about 14 kb, about 15 kb, about 16 kb, about 17 kb, about 18 kb, about 19 kb, and larger).
[0021] In aspects, described herein are nucleic acids encoding the integrases herein (e.g., Bxb1 and / or PhiC31). In aspects, described herein are expression vectors encoding the nucleic acids. In aspects, described herein are host cells comprising the nucleic acids and / or expression vectors.
[0022] In embodiments, the methods and compositions described accelerate therapeutic gene delivery. In embodiments, the methods described herein are adaptable to improve a variety of integrases, e.g., including any integrase that uses an att sequence, to generate novel mutants and mutant combinations.
[0023] In embodiments, integrases described herein find use various methods including (i) site-specific integration of genomic payloads into the genomes of cell lines for industrial biotechnology, the production of pharmaceuticals, and other expression of proteins; (ii) introduction of desirable traits into host cells (e.g., plant or mammalian), such as disease resistance, enhanced nutritional content, or to confer increased functionality; (iii) insertion of disease-related genes into genomes of animal models to study the mechanisms of various disorders and / or test therapeutic interventions; (iv) integration of genetic modules into synthetic genomes, enabling the creation of biological memory devices or genetic logic gates; (v) attachment of epitope tags or fusion proteins to study gene regulation and localization in situ and / or in vivo and (vi) delivering therapeutic genes into the genomes of patients' cells.
[0024] In aspects, described herein are methods of genomic insertion of a host cell comprising introducing at least one nucleic acid comprising a first nucleic acid sequence encoding the integrase (asdescribed herein), optionally introducing a second nucleic acid sequence encoding a genomic payload and one or more attB or attP site compatible with the integrases herein, and optionally introducing a third nucleic acid sequence encoding one or more elements configured for genomic integration of one or more attB or attP site compatible with the integrase.
[0025] In embodiments, the method comprises introducing the first nucleic acid sequence and the second nucleic acid sequence. In embodiments, the method comprises integrase-mediated recombination between the one or more attB or attP site and one or more naturally-occurring genomic sequences of the host cell. In embodiments, the method comprises introducing the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence. In embodiments, the method comprises integrase-mediated recombination between the one or more attB or attP sequence and one or more genomically-integrated attB or attP sequence in the genome of the host cell. In embodiments, methods herein further comprising introducing one or more additional nucleic acid sequences which encodes one or more Bxb1 that targets a “left” and a “right” flanking portion of one or more genomic site in the host cell.
[0026] In embodiments, methods herein comprise introducing one or more additional nucleic acid sequences which encodes one or more Bxb1 integrases that targets a “left” and / or a “right” flanking portion of one or more att site. In embodiments, the one or more Bxb1 integrases that targets a “left” and / or a “right” flanking portion are fusion proteins with one or more DNA-binding domains.
[0027] In embodiments, the methods further comprise using one or more elements configured for genomic integration comprises one or more of a zinc finger nuclease (ZFN), transcription-activator like effector nuclease (TALEN), endonuclease, Cas9 nickase, a reverse transcriptase (e.g., PE5 or PE7), and a Cas9-directed reverse transcription guide RNA (pegRNA), e.g., for knock-in of one or more att site.
[0028] In embodiments, described herein is an adaptation of a method of Cas9-directed reverse transcription established as a tool for precise targeting of short sequences to desired loci with advantages over HDR including high efficiency in both dividing and non-dividing cells as well as a reduction in DSBs. In embodiments, Cas9-directed reverse transcription uses a Cas9 nickase that targets the DNA using a Cas9- directed reverse transcription guide RNA (pegRNA) that encodes a reverse transcriptase template (RTT) which is “written” into the target site by a reverse transcriptase (RT). In embodiments, Cas9-directed reverse transcription is used for DNA substitutions, short insertions, and deletions. In embodiments, methods herein utilize several recent improvements to Cas9-directed reverse transcription, including pegRNA engineering,optimized Cas9-directed reverse transcription architecture, and twin Cas9-directed reverse transcription have enabled the insertion of sequences >50 bp at high efficiencies.
[0029] In embodiments, Cas9-directed reverse transcription is used to insert att sites for insertion of large DNA cargos by integrases. Although high Cas9-directed reverse transcription efficiency has been achieved, the efficiency of insertion into mammalian cells by wildtype integrases has thus far been limited. In embodiments, the method includes first delivering one or more att sites using Cas9-directed reverse transcription, subsequently followed by using mutated integrases to insert large DNA cargos at efficiencies several-fold higher than cognate wildtype integrases into a genome, e.g., a human cell genome. In embodiments, this process is compatible with a variety of host cells.DESCRIPTION OF THE DRAWINGS
[0030] Fig. 1 depicts a graphical representation of Cas9-directed reverse transcription site-specific genomic integration of a 6.6 kbp donor DNA sequence using wild-type Bxb1 compared against mutant Bxb1 with wild-type attB sequence for integration. Percent integration was measured using ddPCR.
[0031] Fig. 2 depicts a graphical representation of site-specific genomic integration of a 6.6 kbp donor DNA sequence using wild-type Bxb1 compared against mutant Bxb1 at the TRAC locus. Percent integration was measured using ddPCR.
[0032] Fig. 3 depicts a graphical representation of site-specific genomic integration of a 6.6 kbp donor DNA sequence at the TRAC locus using wild-type Bxb1 and select Bxb1 mutants compared against cognate recombinases with nuclear localization sequence (NLS). Percent integration was measured using ddPCR.
[0033] Fig. 4 depicts a graphical representation of Cas9-directed reverse transcription and site-specific genomic integration of a 6.6 kbp donor DNA sequence at the ROSA26 locus in HEK293T cells using wildtype Bxb1 and select Bxb1 mutants compared against cognate recombinases with nuclear localization sequence (NLS). NLS sequences are npNLS (SEQ ID NO: 75). Percent integration was measured using ddPCR.
[0034] Fig. 5 depicts a graphical representation of Cas9-directed reverse transcription and site-specific genomic integration of a 6.6 kbp donor DNA sequence using C-terminal NLSs with wild-type Bxb1 and select Bxb1 mutants. Percent integration was measured using ddPCR.
[0035] Fig. 6 depicts a graphical representation of Cas9-directed reverse transcription and site-specific genomic integration of a 6.6 kbp donor DNA sequence using with wild-type Bxb1 and select Bxb1 mutants compared against cognate recombinases with NLSs. Percent integration was measured using ddPCR.
[0036] Figs. 7A-7B depict a graphical representation of Cas9-directed reverse transcription and sitespecific genomic integration of a 6.6 kbp donor DNA sequence using with wild-type Bxb1 and select Bxb1 mutants compared against cognate recombinases with NLSs. Percent integration was measured using ddPCR.
[0037] Fig. 8 depicts a graphical representation of Bxb1-mediated site-specific genomic integration rates at the TRAC locus of HEK293 cells. The genomic payload is an anti-CD19 chimeric antigen receptor. Percent integration was measured by digital droplet PCR (ddPCR) with probes targeted against the TRAC locus and the CAR insertion sequence.
[0038] Fig. 9 depicts a graphical representation of Bxb1-mediated site-specific genomic integration rates at the TRAC locus of HEK293 cells. The genomic payload is an anti-CD19 chimeric antigen receptor. Percent integration was measured by digital droplet PCR (ddPCR) with probes targeted against the TRAC locus and the CAR insertion sequence.
[0039] Figs. 10A-10B depict graphical representations of Bxb1 -mediated site-specific genomic integration a 6.6 kbp donor DNA sequence at the ROSA26 locus of HEK293 cells using wild-type Bxb1 and select Bxb1 mutants compared against cognate recombinases with NLSs and / or stability tags. Percent integration was measured using ddPCR.
[0040] Figs. 11A-11B depict graphical representations of Bxb1-mediated site-specific genomic integration in K562 cells at the TRAC locus and AAVS1 locus using 3-Bxb1 systems, where two Bxb1 mutants are targeted to “left” and “right” portions of a genomic att sequence in combination with a wild-type Bxb1 or a third mutant Bxb1 (e.g., eeBxbl). Fig. 11 A depicts results from targeting genomic attB TRAC sequence (SEQ ID NO: 166) where TRAC-L (mutant Bxb1 targeting the “left” att site) had mutation of residues Y154- P159 to be YRGGLP (SEQ ID NO: 18), mutation of residues S231-R237 to be YGSALKQ (SEQ ID NO: 43), mutation of residues K313-S328 to be LARGPRKRAGYK (SEQ ID NO: 50), and a D257K mutant, relative to SEQ ID NO: 1 (bolded underlined residues represent mutants relative to a wild-type sequence); and where TRAC-R (mutant Bxb1 targeting the “right” att site) had mutations of residues Y154-P159 to be YRGGLP (SEQ ID NO: 18), mutation of residues S231-R237 to be SQWALKC (SEQ ID NO: 44), and mutation of mutation of residues K313-S328 to be RAWGKRKYAYYQ (SEQ ID NO: 51), relative to SEQ ID NO: 1 (boldedunderlined residues represent mutants relative to a wild-type sequence). Fig. 11 B depicts results from targeting genomic attB AAVS1 sequence (SEQ ID NO: 167) where AAVS1-L had mutation of residues Y154- P159 to be YRGGLP (SEQ ID NO: 18), mutation of residues S231-R237 to be YPWSLRR (SEQ ID NO: 45), and mutation of residues K314-R325 to be KAWGSRKTRLYR (SEQ ID NO: 52), relative to SEQ ID NO: 1 (bolded underlined residues represent mutants relative to a wild-type sequence); and where AAVS1-R had mutation of residues Y154-P159 to be YRGGLP (SEQ ID NO: 18), mutation of residues S231-R237 to be AGGNLKR (SEQ ID NO: 21), and mutation of residues K314-R325 to be MARGGRKSAIYY (SEQ ID NO: 53), relative to SEQ ID NO: 1 (bolded underlined residues represent mutants relative to a wild-type sequence). In both cases, significant improvement in insertion efficiency was observed (relative to wild-type Bxb1 , SEQ ID NO: 1) for mutant Bxb1 (e.g., eeBxbl), which comprises mutations V74A, E229K, and V375I relative to SEQ ID NO: 1 , and which was independently verified and performed well at ROSA26 in HEK293 cells, e.g., as described in Example 3 (Table 9 and Fig. 5). Site-specific integration was measured using ddPCR.DETAILED DESCRIPTION
[0041] The present disclosure provides, in part, the ability to deliver large transgenes to a single genomic sequence (site) with high efficiency (e.g., compared to wild-type integrase recombination efficiencies). Current genomic integration methods suffer from low insertion efficiency and many rely on undesired double-strand DNA breaks. In embodiments, the mutant serine integrases described herein (e.g., Bxb1 and PhiC31) catalyze the integration of large DNA cargos at attachment (aft) sites. In embodiments, targeting to att sites in the genome does not require knock-in of an aft sequence. In embodiments, knock-in of att sequences is achieved by using one or more zinc finger nuclease (ZFN), transcription-activator like effector nuclease (TALEN), Cas9-directed reverse transcription, endonuclease, etc., where the att and integrases target “safe” loci. In embodiments, insertion avoids double-strand breaks, which would ordinarily trigger DNA damage pathways (e.g., p53).
[0042] In embodiments, the integrases herein are serine integrases, e.g., PhiC31 and Bxb1 serine integrases. Described herein, in embodiments, are novel hyperactive mutants generated by combining synergistic mutations which function to integrate genetic cargoes. In embodiments, the integrases herein are non-naturally occurring, e.g., have one or more changes relative to a wild-type sequence.
[0043] In embodiments, the genetic cargoes are single or multi-gene cargos. In embodiments, the integration rates are as high as 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% or more of targetchromosomes in host cells. In embodiments, the hyperactive integrases (e.g., PhiC31 and Bxb1 serine integrases) insert therapeutic DNA cargoes ranging in size from about 0.1 kb to about 20 kb or more (e.g., about 0.5 kb, about 1 kb, about 2 kb, about 3 kb, about 4 kb, about 5 kb, about 6 kb, about 7 kb, about 8 kb, about 9 kb, about 10 kb, about 11 kb, about 12 kb, about 13 kb, about 14 kb, about 15 kb, about 16 kb, about 17 kb, about 18 kb, about 19 kb, about 20 kb, and larger).
[0044] In embodiments, the methods and compositions described herein accelerate therapeutic gene delivery. In embodiments, the methods described herein are adaptable to improve a variety of integrases, e.g., including any integrase that uses an aft sequence, to generate novel mutants and mutant combinations.
[0045] In embodiments, integrases described herein find use various methods including (i) site-specific integration of genomic payloads into the genomes of cell lines for industrial biotechnology, the production of pharmaceuticals, and other expression of proteins; (ii) introduction of desirable traits into host cells (e.g., plant or mammalian), such as disease resistance, enhanced nutritional content, or to confer increased functionality; (iii) insertion of disease-related genes into genomes of animal models to study the mechanisms of various disorders and / or test therapeutic interventions; (iv) integration of genetic modules into synthetic genomes, enabling the creation of biological memory devices or genetic logic gates; (v) attachment of epitope tags or fusion proteins to study gene regulation and localization in situ and / or in vivo, and (vi) delivering therapeutic genes into the genomes of patients' cells.
[0046] In embodiments, described herein is an adaptation of a method of Cas9-directed reverse transcription established as a tool for precise targeting of short sequences to desired loci with advantages over HDR including high efficiency in both dividing and non-dividing cells as well as a reduction in DSBs. In embodiments, Cas9-directed reverse transcription uses a Cas9 nickase that targets the DNA using a Cas9- directed reverse transcription guide RNA (pegRNA) that encodes a reverse transcriptase template (RTT) which is “written” into the target site by a reverse transcriptase (RT). In embodiments, Cas9-directed reverse transcription is used for DNA substitutions, short insertions, and deletions. In embodiments, methods herein utilize several recent improvements to Cas9-directed reverse transcription, including pegRNA engineering, optimized Cas9-directed reverse transcription architecture, and twin Cas9-directed reverse transcription have enabled the insertion of sequences >50 bp at high efficiencies.
[0047] In embodiments, Cas9-directed reverse transcription is used to insert aft site insertion of large DNA cargos by integrases. Although high Cas9-directed reverse transcription efficiency has been achieved, the efficiency of insertion into mammalian cells by wildtype integrases has thus far been limited. Inembodiments, the method includes first delivering one or more att sites using Cas9-directed reverse transcription, subsequently followed by using mutated integrases to insert large DNA cargos at efficiencies several-fold higher than cognate wildtype integrases into a genome, e.g., a human cell genome. In embodiments, this process is compatible with a variety of host cells.Integrases
[0048] The present disclosure provides, in embodiments, recombinant “mutant” integrases with altered functionality relative to their cognate wild-type (e.g., non-mutant variant) counterparts. As used herein, in embodiments, “wild-type” refers to a naturally-occurring nucleic acid and / or amino acid sequence, including natural polymorphisms and variants, which has not been modified. In embodiments, the altered functionality includes an improved functionality with respect to one or more functions of the integrase, including increased affinity and / or binding to nucleic acid target sites, altered recognition (e.g., different specificity at DNA target sites, including for example in Bxb1 , at attB sites -19 through -12, -11 through -9, -7, -6, and / or +4, or at attP sites -24 through -17, -11 through -9, -7, and -6) increased integration (e.g., as measured as integration at one or both loci in a chromosome, and / or calculated as a percent, ratio, or amount of host cells with integration compared to the total population of cells), increased probability of integration, decreased off-target integration, etc. In embodiments, “hyperactive,” “increased activity,” or “improved functionality,” with respect to mutant integrases refers to an increase in activity and / or functionality, relative to a cognate wild-type integrase, of about or at least about 1 %, 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% (e.g., 1-fold), 150%, 200% (e.g., 2-fold), 250%, 300% (e.g., 3-fold), 350%, 400% (e.g., 4-fold), 450%, 500% (e.g., 5-fold), 1000% (e.g., 10-fold), 1500% (e.g., 15-fold), 2000% (e.g., 20-fold) or more, in one or more respects.
[0049] In embodiments, the integrase is an enzyme which catalyzes the DNA integration using one or more attachment (att) sites. In embodiments, the integrase described herein interacts between attB and / or attP sites. In embodiments, aft sites are double-stranded DNA (dsDNA). For example with Bxb1, in embodiments, integrases herein make contacts with dsDNA at attB sites -19 through -12 (e.g., hairpin target), -11 through -9 (e.g., helix target), -7 and -6 (e.g., loop target); and / or makes contacts with dsDNA at attP sites -24 through -17 (e.g., hairpin target), -11 through -9 (e.g., helix target), -7 and -6 (e.g., loop target). In embodiments, integrases herein adopt higher order tertiary and quaternary structures (e.g., homodimers, homotetramers, etc.) which make symmetrical contacts on either side of a central dipeptide (e.g., often denoted -1 / +1), where negative (-) and positive (+) denotes a left and right half-site, respectively. Inembodiments, integrases (and mutations, additions, and / or deletions made thereto) alter binding interaction with one half-site (e.g., left or right) which symmetrically effects the binding interaction in the other half-site.
[0050] In embodiments, the integrase (e.g., also referred to herein as a recombinase, or phage integrase) is bacterial, e.g., such as Pa01 (Pseudomonas). In embodiments, the integrase is a viral integrase.
[0051] In embodiments, the integrase is a serine integrase. In embodiments, the serine integrase is Bxb1 (Mycobacterium). In embodiments, the serine integrase is PhiC31 (Streptomyces). In embodiments, the serine integrase is R4 (Streptomyces). In embodiments, the serine integrase is TP901 (Lactococcus). In embodiments, the serine integrase is yb (E. coli). In embodiments, the serine integrase is gin (phage Mu) (E. coli). In embodiments, the serine integrase is Tn3 (Klebsiella). In embodiments, compositions and methods herein comprise one or more integrase that is a large serine recombinase selected from Bxb1, PhiC31 , R4, phiBTI , MJ1 , MR11 , TP901-1 , A118, V153, phiRVI , phi370.1, TG1 , WB, BL3, SprA, phiJoe, phiK38, Int2, Int3, Int4, Int7, Int8, Int9, IntIO, Inti 1, Inti 2, Int13, L1 , peaches, Bxz2, and SV1. Persons skilled in the art, with the benefit of this disclosure in its entirety, will understand how to utilize structure-function analysis to take mutations from Bxb1 and / or PhiC31 to make equivalent mutations in other large serine recombinases and test their efficiency against wild-type enzymes e.g., using the methods described herein.
[0052] Illustrative sequences of PhiC31 and Bxb1 integrases are shown in Table 1 .Table 1 : Illustrative wild-type amino acid sequences of serine integrases, PhiC31 and Bxb1 . Residues in bold and underlined represent illustrative mutable amino acids.
[0053] In embodiments, the Bxb1 integrase includes one or more mutants at one or more amino acid positions of 4, 5, 14, 18, 20, 24, 29, 34, 35, 36, 40, 42, 45, 46, 49, 50, 51, 57, 60, 61, 62, 63, 67, 68, 69, 70, 73, 74, 75, 76, 78, 79, 84, 85, 86, 87, 88, 89, 90, 92, 95, 99, 100, 104, 105, 106, 110, 111, 116, 119, 122, 124, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174, 175, 178, 179, 181, 183, 187, 189,191, 194, 197, 203, 207, 208, 209, 218, 223, 229, 231, 232, 233, 234, 236, 237, 239, 248, 251, 254, 257,261, 264, 267, 268, 272, 273, 278, 280, 281, 282, 283, 285, 287, 288, 291, 292, 295, 302, 306, 307, 311,313, 314, 316, 318, 319, 321, 322, 323, 325, 328, 331, 332, 333, 334, 341, 342, 343, 347, 353, 355, 359,360, 361, 362, 368, 369, 370, 375, 380, 388, 396, 397, 398, 405, 409, 411, 414, 415, 416, 419, 425, 428,434, 435, 444, 449, 453, 461, 462, 463, 466, 468, 476, 479, 480, 483, 484, 487, 488, 489, 494, 496, and / or 499 relative to SEQ ID NO: 1 , or equivalent positions thereof in related integrases.
[0054] In embodiments, the Bxb1 integrase includes one or more mutants at one or more amino acid positions selected from the list consisting of 75, 76, 158, 232, 234, 236, 237, 257, 314, 316, 318, 322, 323, 325; and / or the Bxb1 integrase includes one or more mutants at one or more amino acid positions selected from the list consisting of 5, 14, 20, 24, 29, 35, 40, 45, 49, 50, 51, 60, 68, 69, 70, 73, 74, 78, 84, 86, 87, 100, 105, 116, 124, 183, 197, 207, 208, 209, 229, 261, 267, 273, 273, 287, 291, 333, 342, 343, 347, 361, 368, 375, 435, 449, 453, 462, 483, 494; and / or the Bxb1 integrase includes one or more mutants at one or more amino acid positions selected from the list consisting of 4, 18, 34, 36, 42, 46, 61, 62, 63, 67, 79, 85, 88, 89, 90, 92, 95, 99, 106, 110, 111, 119, 122, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174,175, 178, 179, 181, 187, 189, 191, 203, 218, 223, 231, 239, 248, 251, 254, 264, 268, 272, 278, 280, 281,282, 283, 285, 288, 292, 295, 302, 306, 307, 311, 313, 319, 321, 328, 331, 332, 334, 353, 355, 359, 360,362, 369, 370, 380, 388, 397, 398, 405, 409, 411, 414, 415, 416, 419, 425, 428, 434, 444, 461, 463, 466,468, 476, 479, 480, 484, 487, 488, 489, 496, and / or 499.
[0055] In embodiments, the Bxb1 integrase includes one or more mutants at one or more amino acid positions selected from the list consisting of 14, 20, 29, 35, 45, 49, 50, 60, 68, 70, 73, 74, 75, 76, 78, 84, 86, 116, 124, 157, 158, 183, 197, 207, 208, 232, 234, 236, 237, 257, 267, 273, 291, 314, 316, 318, 322, 323, 325, 342, 343, 368, 449, and 462; in combination with one or more mutants at one or more amino acid positions selected from the list consisting of 4, 5, 18, 24, 34, 36, 40, 42, 46, 51, 61, 62, 63, 67, 69, 79, 85,87, 88, 89, 90, 92, 95, 99, 100, 105, 106, 110, 111, 119, 122, 130, 133, 137, 140, 145, 153, 156, 160, 164, 166, 174, 175, 178, 179, 181, 187, 189, 191, 203, 209, 218, 223, 229, 231, 239, 248, 251, 254, 261, 264,268, 272, 278, 280, 281, 282, 283, 285, 287, 288, 292, 295, 302, 306, 307, 311, 313, 319, 321, 328, 331,332, 333, 334, 347, 353, 355, 359, 360, 361, 362, 369, 370, 375, 380, 388, 397, 398, 405, 409, 411, 414,415, 416, 419, 425, 428, 434, 435, 444, 453, 461, 463, 466, 468, 476, 479, 480, 483, 484, 487, 488, 489,494, 496, and / or 499.
[0056] In embodiments, the Bxb1 integrase includes one or more mutants at one or more amino acid positions selected from the list consisting of 5, 14, 20, 24, 29, 35, 40, 45, 49, 50, 51, 60, 68, 69, 70, 73, 74, 78, 84, 86, 87, 100, 105, 116, 124, 157, 183, 197, 207, 208, 209, 229, 261, 267, 273, 287, 291, 333, 342, 343, 347, 361 , 368, 375, 435, 449, 453, 462, 483, 494; in combination with one or more mutants at one or more amino acid positions selected from the list consisting of 4, 18, 34, 36, 42, 46, 61, 62, 63, 67, 75, 76, 79, 85, 88, 89, 90, 92, 95, 99, 106, 110, 111, 119, 122, 130, 133, 137, 140, 145, 153, 156, 158, 160, 164, 166, 174, 175, 178, 179, 181, 187, 189, 191, 203, 218, 223, 231, 232, 234, 236, 237, 239, 248, 251, 254, 257,264, 268, 272, 278, 280, 281, 282, 283, 285, 288, 292, 295, 302, 306, 307, 311, 313, 314, 316, 318, 319,321, 322, 323, 325, 328, 331, 332, 334, 353, 355, 359, 360, 362, 369, 370, 380, 388, 397, 398, 405, 409,411, 414, 415, 416, 419, 425, 428, 434, 444, 461, 463, 466, 468, 476, 479, 480, 484, 487, 488, 489, 496, and / or 499.
[0057] In embodiments, the Bxb1 integrase includes one or more mutants selected from one or more of L4I, V5V, V5I, D14N, S18S, E20K, E20Q, E24E, E24K, L29F, G34D, W35L, D36A, V40I, V40A, V40V, E42K, D45G, V46V, A49T, V50I, D51D, , D51N, D51E, D51Y, N60S, L61F, A62A, R63K, F67S, E68K, E69A, E69D, E69E, Q70P, D73G, V74A, V74I, V74M, I75V, V76I, Y78H, Y78N, R79R, T84S, R85R, S86T, I87V, I87L, R88R, H89G, L90L, Q92H, Q92Q, H95Y, D99N, H100N, H100Y, V105A, V105I, S106S, A110A, H111P, T116P, A119S, V122M, A124S, A130A, E133E, 11371, R140R, A145T, K153K, S157G, G156G, P160P, L164L, T166I, L174L, V175V, P178P, V179A, V179I, R181 K, E183L, V187I, V187V, H189N, Q191Q, P197T, H203Y, R207Q, R208S, G209G, G209V, A218A, R223R, E229K, S231S, A232G, A232S, A232R,A232T, A232V, A232P, A232Q, T233G, T233W, T233R, T233Y, T233D, T233G, T233N, T233H, T233S,T233Q, T233A, T233C, A234N, A234G, A234S, A234T, A234H, A234F, A234S, K236R, K236S, R237K,R237Q, R237N, R237V, R237C, M239I, A248T, N251K, N251N, T254S, D257K, A261T, A261V, V264A,E267D, P268P, R272Q, E273D, E273K, L278L, A280T, A280A, E281E, L282L, V283V, T285A, R287R,R287P, A288V, A288T, A291T, P292P, P295P, L302L, V306V, C307C, A311V, K313R, K313K, F314M, F314L, F314N, F314R, F314K, G316R, G316W, G318H, G318HP, G318S, G318N, G318K, R319K, R319G,H321 P, H321 Q, H321 K, H321 L, H321R, H321Y, H321T, H321S, P322A, P322R, P322G, P322L, R323L,R323G, R323Y, R323I, R325K, R325Q, R325Y, S328S, F331 S, P332H, K333K, K333N, H334R, M342T,M342V, A343T, A347E, A347V, V353I, D355N, D359N, D359D, A360T, E361 E, E361 D, R362K, V368A,A369P, A369E, A369T, A369S, G370G, V375V, V375I, V380I, T388M, R397R, A396P, A398S, A405A,R409H, A411V, A414A, A414V, A415S, R416K, E419E, A425T, A425A, S428S, E434G, T435T, T435A,R444L, A449V, T453I, T453A, R461 R, L462M, T463I, V466M, V466V, G468D, G468G, F476F, L479I, Q480STOP, E483E, E483K, Q484K, R487R, del L488 (del L488 frameshift), G489G, R494S, R494R, R494Q, H496N, and / or M499T relative to SEQ ID NO: 1, or equivalent positions thereof in related integrases. In embodiments, these mutants relate to the combinations of the lists of mutants herein. In embodiments, mutations denote Xi#Xi, where Xi is a single amino acid and # is the position number (e.g., V5V) represent positions of silent mutations, and / or or a residue that is kept unmutated in combinations with other mutants.
[0058] In embodiments, the Bxb1 integrase comprises a C-terminal deletion of one or more amino acids in the range of L488-S500 relative to SEQ ID NO: 1 (e.g., a deletion of 1-13 amino acids from the C-terminal end). In embodiments, the C-terminal deletion comprises a deletion of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, or all 13 amino acids of L488-S500 relative to SEQ ID NO: 1. In embodiments, the C-terminal deletion includes mutating a residue selected from the range of L488-S500, relative to SEQ ID NO: 1 , to a stop codon (e.g., L488STOP). In embodiments, a deletion of one or more amino acids of L488-S500 relative to SEQ ID NO: 1 results in no change in functionality of Bxb1 integrase relative to wild-type. In embodiments, a deletion of one or more amino acids of L488-S500 relative to SEQ ID NO: 1 results in improved functionality of Bxb1 integrase relative to wild-type.
[0059] In embodiments, the Bxb1 integrase comprises a C-terminal deletion of one or more amino acids in the range of Q480-S500 relative to SEQ ID NO: 1 (e.g., a deletion of 1-21 amino acids from the C-terminal end). In embodiments, the C-terminal deletion comprises a deletion of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, 13 amino acids, 14 amino acids, 15 amino acids, 16 amino acids, 17 amino acids, 18 amino acids, 19 amino acids, 20 amino acids, or all 21 amino acids of Q480-S500 relative to SEQ ID NO: 1. In embodiments, the C-terminal deletion includes mutating a residue selected from the range of Q480-S500, relative to SEQ ID NO: 1, to a stop codon (e.g., Q488STOP). In embodiments, a deletion of one or more amino acids of Q480-S500 relative to SEQ ID NO: 1 results in no change in functionality of Bxb1 integrase relative to wild-type. In embodiments, a deletion of one or more amino acidsof Q480-S500 relative to SEQ ID NO: 1 results in improved functionality of Bxb1 integrase relative to wildtype.
[0060] In embodiments, Bxb1 residues A157-P160 are mutable to alters specificity at -7 and -6 positions of the loop target DNA site. In embodiments, the mutation of Bxb1 is selected from residues A157, L158, and P160. In embodiments, mutation of Bxb1 L158 alters specificity at -7 and -6 positions of the loop target DNA site. In embodiments, the mutation is S157G in the Bxb1 at DNA loop recognition site. In embodiments, mutation at this position alters the dinucleotide preference to adenine (A) or thymine (T) at positions -7 and cytosine (C) or guanine (G) at position -6.
[0061] In embodiments, the one or more amino acid mutations comprises mutation of one or more of residues of Y154-P159; and / or wherein the one or more mutations alters specificity at -7 and / or -6 positions of a loop target DNA site compared to a wild-type Bxb1 of SEQ ID NO: 12. In embodiments, the mutation comprises S157G. In embodiments, the one or more Bxb1 integrases comprise a sequence from Y154-P159 selected from YRGSLS (SEQ ID NO: 16), WWGGTP (SEQ ID NO: 17), YRGGLP (SEQ ID NO: 18), AHPGRT (SEQ ID NO: 19), and WRGGAS (SEQ ID NO: 20), relative to SEQ ID NO: 1 or a position corresponding thereto, where the residues mutated from the wild-type sequence (e.g., SEQ ID NO: 1) are bolded and underlined.
[0062] In embodiments, there are 16 possible attB and / or attP central dinucleotide sequences which Bxb1 can target. In embodiments, the non-wild-type Bxb1 integrases herein are able to target one of the non- canonical 15 central dinucleotide sequences (e.g., central dinucleotide that is not GT, such as GG, GA, GO, AT, AA, AG, AC, TT, TA, TG, TC, CC, CA, CT, CG). In embodiments, the non-wild-type Bxb1 herein targets thymine-adenine (TA) as the core dinucleotide sequence (e.g., denoted as the -1 / +1 dinucleotide).
[0063] In embodiments, the Bxb1 S231-R237 region recognizes the 3 bp helix target site flanking the DNA cleavage site (e.g., -11 , -10, -9) in the target DNA site. In embodiments, this stretch of amino acids forms an alpha helix that docks with the target DNA site, reminiscent of the interaction between a zinc finger (ZF) and its target trinucleotide sequence. In embodiments, this site is mutated to a non-native sequence to alter target specificity, e.g., non-native genomic sites. Because this section recognizes 3 DNA bps, in embodiments, mutations are made that target all 64 possible DNA triplets at positions -11, -10, -9. In embodiments, L235 is not mutated because it is analogous to the residue often found at +4 of a zinc finger (ZF) recognition helix, .e.g., leaving residues 231-234 and 236-237 to be mutated to alter specificity at positions -11, -10, and -9 DNA target sites. In embodiments, A234 is mutated to an asparagine (e.g., A234N)to alter the specificity at the -10 position to thymine (T). In embodiments, the 3 bp helix target site flanking the DNA cleavage site (e.g., -11 , -10, -9) is altered as a function of mutation of residues S231-R237.
[0064] In embodiments, residues 231, 233, and / or 237 are mutable to alter specificity at -10 and -9 DNA target sites. In embodiments, the Bxb1 integrase comprises a sequence from S231-R237 relative to SEQ ID NO: 1 selected from AGGNLKR (SEQ ID NO: 21), LGTNLKR (SEQ ID NO: 22), SGTGLKK_(SEQ ID NO: 23), SGSALKT (SEQ ID NO: 24), AAWALRR (SEQ ID NO: 25), GGRSLKR (SEQ ID NO: 26), SGYNLRR (SEQ ID NO: 27), SGWGLKK (SEQ ID NO: 28), SGWALRQ (SEQ ID NO: 29), SARALSR (SEQ ID NO: 30), RADTLRR (SEQ ID NO: 31), YSRNLKR (SEQ ID NO: 32), SRNGLRK (SEQ ID NO: 33), RGHALKN (SEQ ID NO: 34), GGSHLKR (SEQ ID NO: 35), TTRTLKR (SEQ ID NO: 36), AVQNLKR (SEQ ID NO: 37), RAAFLKK_(SEQ ID NO: 38), RAWTLKC (SEQ ID NO: 39), RAWSLKR (SEQ ID NO: 40), HGWSLKV (SEQ ID NO: 41), HGCTLKR (SEQ ID NO: 42), YGSALKQ (SEQ ID NO: 43), SQWALKC (SEQ ID NO: 44), YPWSLRR (SEQ ID NO: 45), where the residues mutated from the wild-type sequence (e.g., SEQ ID NO: 1) are bolded and underlined. In embodiments, 1 residue, 2 residues, 3 residues, 4 residues, or 5 residues are mutated between S231-R237 (inclusive) relative to SEQ ID NO: 1.
[0065] In embodiments, mutation of Bxb1 residues K313-S328 alter specificity at DNA binding / interatction sites -19 through -12 (e.g., DNA hairpin target sequence). In embodiments, these residues form the interfacial region with the DNA major grove and / or influence DNA hairpin structure. In embodiments, Bxb1 residue P322 is maintained as a proline. In embodiments, K313, F314, A315, G316, G318, R319, H321 , P322, R323, R325, and / or S328. In embodiments, the Bxb1 has mutation of residues K313, R319, H321, and / or S328. In embodiments, residues K314-R325 are mutated to one of the following sequences: MAGGHRKQALYR (SEQ ID NO: 46), MAGGPRKKRRYR (SEQ ID NO: 47), LARGSRKLALYR (SEQ ID NO: 48), NARGNRKRGRYR (SEQ ID NO: 49), LARGPRKRAGYK (SEQ ID NO: 50), RAWGKRKYAYYQ (SEQ ID NO: 51), KAWGSRKTRLYR (SEQ ID NO: 52), MARGGRKSAIYY (SEQ ID NO: 53), MASGSRKTAIYY (SEQ ID NO: 54), LARGRRKWARYR (SEQ ID NO: 55), and LARGSRKLALYR (SEQ ID NO: 56), where the residues mutated from the wild-type sequence (e.g., SEQ ID NO: 1) are bolded and underlined. In embodiments, 1 residue, 2 residues, 3 residues, 4 residues, 5 residues, 6 residues, or 7 residues are mutated between F314-R325 (inclusive) relative to SEQ ID NO: 1.
[0066] In embodiments, the Bxb1 integrase comprises one or more mutations and / or deletions, and / or one or more combinations of mutations and / or deletions as described in Table 2. In embodiments, the Bxb1 integrase comprises one or more mutations selected from Table 2.Table 2: Illustrative Bxb1 integrase mutants and illustrative mutant combinations. “X” refers to any amino acid different from the wild-type residue. Illustrative changes are relative to SEQ ID NO: 1. bpNLS = KRTADGSEFESPKKKRKV (SEQ ID NO: 74); npNLSKRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75); HisNPL = GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0067] In embodiments, the Bxb1 integrase has at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1 , or equivalent positions thereof.
[0068] In embodiments, the Bxb1 integrase comprises at least 90% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1, or equivalent positions thereof.
[0069] In embodiments, the Bxb1 integrase comprises at least 95% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, or at least 20 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1 , or equivalent positions thereof.
[0070] In embodiments, the Bxb1 integrase comprises at least 96% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, or 20 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1 , or equivalent positions thereof.
[0071] In embodiments, the Bxb1 integrase comprises at least 97% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, or 15 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1 , or equivalent positions thereof.
[0072] In embodiments, the Bxb1 integrase comprises at least 98% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, or 10mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1 , or equivalent positions thereof.
[0073] In embodiments, the Bxb1 integrase comprises at least 99% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, or 5 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1 , or equivalent positions thereof.
[0074] In embodiments, the Bxb1 integrase comprises the amino acid sequences of any one of SEQ ID NOs: 124-153.
[0075] In embodiments, the Bxb1 integrase comprises at least 90% sequence identity to residues 1- 480 of SEQ ID NO: 1 and having one or more mutations selected from Table 2. In embodiments, the Bxb1 integrase comprises at least 92% sequence identity to residues 1-480 of SEQ ID NO: 1 and having one or more mutations selected from Table 2. In embodiments, the Bxb1 integrase comprises at least 95% sequence identity to residues 1-480 of SEQ ID NO: 1 and having one or more mutations selected from Table 2. In embodiments, the Bxb1 integrase comprises at least 96% sequence identity to residues 1-480 of SEQ ID NO: 1 and having one or more mutations selected from Table 2. In embodiments, the Bxb1 integrase comprises at least 97% sequence identity to residues 1-480 of SEQ ID NO: 1 and having one or more mutations selected from Table 2. In embodiments, the Bxb1 integrase comprises at least 98% sequence identity to residues 1-480 of SEQ ID NO: 1 and having one or more mutations selected from Table 2. In embodiments, the Bxb1 integrase comprises at least 99% sequence identity to residues 1-480 of SEQ ID NO: 1 and having one or more mutations selected from Table 2.
[0076] In embodiments, the Bxb1 integrase the an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1 and one or more amino acid mutations at positions selected from L4, E42, R57, R63, V76, H111 , V122, V187, K313, R319, A341 , D359, A369, A398, R416, A425, L479, and M499, relative to SEQ ID NO: 1.
[0077] In embodiments, the Bxb1 amino acid is at least 97% identical to the amino acid sequence of SEQ ID NO: 1 and has one or more amino acid mutations at positions selected from L4, E42, R57, R63, V76, H111 , V122, V187, K313, R319, A341, D359, A369, A398, R416, A425, L479, and M499, relative to SEQ ID NO: 1.
[0078] In embodiments, the Bxb1 amino acid is at least 98% identical to the amino acid sequence of SEQ ID NO: 1 and has one or more amino acid mutations at positions selected from L4, E42, R57, R63, V76,H111 , V122, V187, K313, R319, A341, D359, A369, A398, R416, A425, L479, and M499, relative to SEQ IDNO: 1.
[0079] In embodiments, the Bxb1 amino acid is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and has one or more amino acid mutations at positions selected from L4, E42, R57, R63, V76, H111 , V122, V187, K313, R319, A341, D359, A369, A398, R416, A425, L479, and M499, relative to SEQ ID NO: 1.
[0080] In embodiments, the Bxb1 integrase comprises one or more amino acid mutations selected from L4I, E42K, R57K, R63K, V76I, H111 P, V122M, V187, K313R, R319K, A341 D, D359A, A369P, A398S, R416K, A425T, L479I, and M499T relative to SEQ ID NO: 1. In embodiments, the Bxb1 integrase comprises one or more amino acid mutations selected from D36N, V40A, A49S, V74I, I87L or I87A or I87S, H89G, V175A, V179A, R287H, A288V, T453I, and H496N relative to SEQ ID NO: 1 (e.g., in combination with the one or more amino acid mutations selected from L4I, E42K, R57K, R63K, V76I, H111 P, V122M, V187, K313R, R319K, A341 D, D359A, A369P, A398S, R416K, A425T, L479I, and M499T relative to SEQ ID NO: 1). In embodiments, the Bxb1 integrase comprises one or more amino acid mutations selected from V40I, D45G, I75V, H95Y, A119S, A280T, A311 V, E434G, and V466M relative to SEQ ID NO: 1 (e.g., in combination with the one or more amino acid mutations selected from L4I, E42K, R57K, R63K, V76I, H111 P, V122M, V187, K313R, R319K, A341 D, D359A, A369P, A398S, R416K, A425T, L479I, and M499T relative to SEQ ID NO: 1).
[0081] In embodiments, the Bxb1 integrase comprises one or more of the following amino acids, e.g., relative to SEQ ID NO: 1 , which are not mutated: A62, H100, Q191 , P195, G209, P295, L302, C307, L387, E419, S428, E483, and R487 relative to SEQ ID NO: 1.
[0082] In embodiments, the Bxb1 integrase comprises one or more amino acid mutations selected from mutations at position V122 and / or position A369 relative to SEQ ID NO: 1 . In embodiments, the mutation at position V122 relative to SEQ ID NO: 1 is V122M. In embodiments, the mutation at position A369 relative to SEQ ID NO: 1 is A369P. In embodiments, the Bxb1 integrase comprises mutations at position V122 and position A369 relative to SEQ ID NO: 1. In embodiments, the mutations are V122M and A369P relative to SEQ ID NO: 1.
[0083] In embodiments, the Bxb1 integrase comprises one or more amino acid mutation position combinations selected from: V76 and V122; V76 and A369; I87 and V122; I87 and A369; H95 and V122; H95 and A369; V122 and E434; A369 and E434; V76, V122, and A369; I87, H95 and A369; I87, V122 andE434; 187, A369 and E434; H95, V122 and A369; H95, V122 and E434; H95, A369 and E434; V122, A369 and E434; I87, H95, V122 and E434; I87, H95, A369 and E434; and I87, V122, A369 and E434 relative to SEQ ID NO: 1.
[0084] In embodiments, the Bxb1 integrase comprises one or more amino acid mutation combinations selected from: V76I and V122M; V76I and A369P; I87L and V122M; I87L and A369P; H95Y and V122; H95Y and A369P; V122M and E434G; A369P and E434G; V76I, V122M, and A369P; I87L, H95Y and A369P; I87L, V122M and A369P; I87L, V122M and E434G; I87L, A369P and E434G; H95, V122M and A369P; H95Y, V122M and E434G; H95Y, A369P and E434G; V122, A369P and E434G; I87L, H95Y, V122M and E434G; I87L, H95Y, A369P and E434G; or I87L, V122, A369P and E434G relative to SEQ ID NO: 1.
[0085] In embodiments, the Bxb1 integrase comprises one or more amino acid mutation position combinations selected from: I87, A369, and E434, V122, A369, and E434, H95, A369, and E434, V122 and E434, V122, A369, E434, and V175, V76 and A119, A119, A369, and E434, 187, V175, and D359, 187, R57, and R63, 187 and A369, 187 and A425, R63 and V122, R63, V122, R319, and V40I, R63, A119, and M499, V122 and A369, H95 and A369, A49, I87, A396, and E434, R57, R63, I87, A396, and E434, D36, V122, A396, and E434, R57, R63 and H95, D36, I87, A396, and E434, I87, A341, A396, and E434, R57, R63, V122, A396, and E434, V122, R287, A396, and E434, 187, R287, A396, and E434, 187 and A341, 187, R287, A396, E434, and A341, V122, A396, E434, and A341, H95, and A341, A49, V122, A396, and E434, D45, A396, and E434, V122, R287, A341, A396, and E434, A49, A396, and E434, I87, A341, and R287, D36, A396, and E434, V122, D359, A369, and E434, R63 and I87, 187, D359, A369, and E434, V122, A369, A425, and E434, 187, A369, A425, and E434, V40, H95, A369, and E434, L4, H95, A369, and E434, H95, A369, E434, and V466, R63, V122, A369, and E434, V76, 187, and K313, R63, I87, and N194, 187 and D359, L4 and I87, V74 and V76, V122, A369, E434, and R319, 187, A369, E434, and R319, L4, V122, A369, and E434, I87 and A425, E42, V76, and I87, 187 and L479, H95, A369, A425, and E434, R63, 187, A369, and E434, I87 and R319, V76, A369, and E434, H95, A369, E434, and L479, V74, 175, and V76, H95, R319, A369, and E434, L4, 187, A369, and E434, V122, V175, D359, A369, and E434, R57, 187, A369, and E434, E42, H95, A369, and E434, H95, V179, A369, and E434, 175 and V76, L4 and V76, H95 and K313, V40, R63, and H89, A288, A311, A398, R416, T453, H496, and E42, and H111 and E434 relative to SEQ ID NO: 1.
[0086] In embodiments, the Bxb1 integrase comprises one or more amino acid mutation combinations selected from: I87L, A369P, and E434G, V122M, A369P, and E434G, H95Y, A369P, and E434G, I87L and E434G, V122M and E434G, V122M, A369P, E434G, and V175A, V76I and A119S, D36N and A119S,A119S, A369P, and E434G, I87L, V175A, and D359A, I87L, R57K, and R63K, I87L and D36N, I87L and A369P, I87L and A425T, R63K and V122M, R63K, V122M, R319K, and V40I, R63K, A119S, and M499T, H95Y and E434G, V122M and A369P, H95Y and V466M, H95Y and A369P, A49S, I87L, A396P, and E434G, A49S and I87L, R57K, R63K, I87L, A396P, and E434G, D36N, V122M, A396P, and E434G, R57K, R63K and H95Y, D36N, I87L, A396P, and E434G, D36N and H95Y, I87L, A341D, A396P, and E434G, A49S and H95Y, R57K, R63K, V122M, A396P, and E434G, V122M, R287H, A396P, and E434G, I87L and R287H, I87L, R287H, A396P, and E434G, I87L and A341D, I87L, R287H, A396P, E434G, and A341D, V122M, A396P, E434G, and A341D, H95Y and A341D, A49S, V122M, A396P, and E434G, D45G, A396P, and E434G, V122M, R287H, A341D, A396P, and E434G, A49S, A396P, and E434G, I87L, A341D, and R287H, H95Y and R287H, D36N, A396P, and E434G, V122M, D359A, A369P, and E434G, R63K and I87L, I87L and V466M, I87L, D359A, A369P, and E434G, V122M, A369P, A425T, and E434G, I87L, A369P, A425T, and E434G, V40I, H95Y, A369P, and E434G, L4I, H95Y, A369P, and E434G, H95Y, A369P, E434G, and V466M, R63K, V122M, A369P, and E434G, V76I, I87L, and K313R, R63K, I87V, and N194D, I87L and D359A, I87L and L4I V74I and V76I, V122M, A369P, E434G, and R319K, I87L, A369P, E434G, and R319K, L4I, V122M, A369P, and E434G, I87L and A425T, E42K, V76I, and I87L, I87L and L479I, H95Y, A369P, A425T, and E434G, V40I and I87L, R63K, I87L, A369P, and E434G, I87L and R319K, V76I, A369P, and E434G, H95Y, A369P, E434G, and L479I, V74I, I75V, and V76I, H95Y, R319K, A369P, and E434G, L4I, I87L, A369P, and E434G, V122M, V175A, D359A, A369P, and E434G, R57K, I87L, A369P, and E434G, E42K, H95Y, A369P, and E434G, H95Y, V179A, A369P, and E434G, I75V and V76I, I87L and V179A, L4I and V76I, H95Y and K313R, H95Y and A280T, V40A, R63K, and H89G, A288V, A311V, A398S, R416K, T453I, H496N, and E42K, and H111 P and E434G relative to SEQ ID NO: 1.
[0087] In embodiments, the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and has the amino acid mutations I87L, A369P, and E434G relative to SEQ ID NO: 1.
[0088] In embodiments, the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations V122M, A369P, and E434G relative to SEQ ID NO: 1.
[0089] In embodiments, the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations H95Y, A369P, and E434G relative to SEQ ID NO: 1.
[0090] In embodiments, the PhiC31 integrase includes one or more mutants at one or more amino acid positions of 1, 2, 12, 14, 18, 24, 32, 36, 41, 43, 44, 45, 49, 51, 55, 58, 74, 77, 96, 103, 107, 117, 153, 176, 196, 197, 199, 200, 228, 229, 230, 231, 235, 238, 240, 252, 254, 255, 260, 262, 264, 266, 269, 274, 278,302, 319, 320, 322, 331, 333, 340, 344, 345, 346, 347, 351, 353, 355, 359, 362, 364, 378, 382, 393, 396,397, 399, 406, 410, 424, 429, 431, 436, 438, 445, 448, 449, 450, 452, 457, 468, 472, 478, 498, 499, 501,505, 512, 514, 516, 517, 520, 527, 535, 536, 549, 551, 552, 553, 563, 568, 580, 585, 586, 587, 590, 592,594, 600, 603, 604, 609, 612, and / or 616 relative to SEQ ID NO: 2, or equivalent positions thereof in related integrases.
[0091] In embodiments, the PhiC31 integrase includes one or more mutants selected from one or more of M1E, M1V, D2V, D2M, S12N, E14G, S18N, A24A, D32A, D36A, V41I, R43R, D44A, G45G, R49R, V51M, S55S, P58P, I74I, E77E, R96R, 11031, S107S, V117V, 11531, 1153V, E176D, N196N, K197K, A199T, H200H, H228Y, L229L, P230S, F231L, S235S, A238A, H240R, D252G, D254D, A255G, G260G, T262T, G264G, K266K, S269N, P274S, M278L, T302A, L319L, R320R, V322V, E331E, A333T, A333S, A333D, A340V, G344V, G344D, R345R, G346S, R347K, L351V, L351Q, R353R, Q355Q, S359S, D362N, D362G, L364M, E378K, K382K, V393V, S396N, S396R, A397T, G399G, N406S, A410T, 14241, G429S, E431V, L436L, W438R, G445G, W448R, E449D, A450D, E452K, R457R, L468L, E472D, R478R, A498A, L499L, L501L, G505V, G505G, G505S, E512K, E514E, A516T, E517E, K520K, F527F, P535L, P535P, T536A, D549D, R551R, V552V, F553F, V563V, T568T, A580V, A585A, K586Q, P587P, D590G, D592G, D594N, D594D, T600S, T600T, V603I, V603A, V603V, A604S, P609S, V612V, and / or A616V relative to SEQ ID NO: 2, or equivalent positions thereof in related integrases.
[0092] In embodiments, the PhiC31 integrase comprises a N-terminal addition of one or more amino acids. In embodiments, the N-terminal extension comprises an addition of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, 13 amino acids, 14 amino acids, 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 35 amino acids, about 40 amino acids, about 45 amino acids, or about 50 amino acids. In embodiments, the N-terminal addition is selected from Table 3.Table 3: Illustrative N-terminal additions for PhiC31 integrases.
[0093] In embodiments, the N-terminal addition site is between M1 and D2 residues, e.g., of SEQ ID NO: 2. In embodiments, the N-terminal extension is include N-terminal relative to M1 of SEQ ID NO: 2. In embodiments, an extension of one or more amino acids at the N-terminus relative to SEQ ID NO: 2 results in no change in functionality of PhiC31 integrase relative to wild-type. In embodiments, an extension of one or more amino acids at the N-terminus relative to SEQ ID NO: 2 results in improved functionality of PhiC31 integrase relative to wild-type.
[0094] In embodiments, the PhiC31 integrase comprises one or more mutations, N-terminal additions, and / or one or more combinations of mutations and / or N-terminal additions as described in Table 3 and / or Table 4. In embodiments, the PhiC31 integrase comprises one or more mutations and / or N-terminal additions selected from Table 3 and / or Table 4.Table 4: Illustrative PhiC31 integrase mutants and illustrative mutant combinations. “X” refers to any amino acid different from the wild-type residue. Illustrative changes are relative to SEQ ID NO: 2.
[0095] In embodiments, the PhiC31 integrase has at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations, e.g., as selected from Table 1 , Table 3, and / or Table 4 relative to SEQ ID NO: 2, or equivalent positions thereof. In embodiments, mutants with respect to PhiC31 integrase include N-terminal addition.
[0096] In embodiments, the PhiC31 integrase comprises at least 90% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations, e.g., as selected from Table 1 , Table 3, and / or Table 4 relative to SEQ ID NO: 2, or equivalent positions thereof.
[0097] In embodiments, the PhiC31 integrase comprises at least 95% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, or at least 30 mutations, e.g., as selected from Table 1 , Table 3, and / or Table 4 relative to SEQ ID NO: 2, or equivalent positions thereof.
[0098] In embodiments, the PhiC31 integrase comprises at least 96% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, atleast 20 mutations, e.g., as selected from Table 1 , Table 3, and / or Table 4 relative to SEQ ID NO: 2, or equivalent positions thereof.
[0099] In embodiments, the PhiC31 integrase comprises at least 97% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, or at least 15 mutations, e.g., as selected from Table 1 , Table 3, and / or Table 4 relative to SEQ ID NO: 2, or equivalent positions thereof.
[0100] In embodiments, the PhiC31 integrase comprises at least 98% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, or 10 mutations, e.g., as selected from Table 1, Table 3, and / or Table 4 relative to SEQ ID NO: 2, or equivalent positions thereof.
[0101] In embodiments, the PhiC31 integrase comprises at least 99% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, or at least 5 mutations, e.g., as selected from Table 1, Table 3, and / or Table 4 relative to SEQ ID NO: 2, or equivalent positions thereof.
[0102] In embodiments, the PhiC31 integrase comprises at least 90% sequence identity to residues 1- 604 of SEQ ID NO: 2 and having one or more mutations selected from Table 1 and / orTable 4, and / or having an N-terminal addition of one or more amino acids (e.g., about 1 to about 50 amino acids). In embodiments, the PhiC31 integrase comprises at least 92% sequence identity to residues 1-604 of SEQ ID NO: 2 and having one or more mutations selected from Table 1 and / or Table 4, and / or having an N-terminal addition of one or more amino acids (e.g., about 1 to about 50 amino acids). In embodiments, the PhiC31 integrase comprises at least 95% sequence identity to residues 1-604 of SEQ ID NO: 2 and having one or more mutations selected from Table 1 and / or Table 4, and / or having an N-terminal addition of one or more amino acids (e.g., about 1 to about 50 amino acids). In embodiments, the PhiC31 integrase comprises at least 96% sequence identity to residues 1-604 of SEQ ID NO: 2 and having one or more mutations selected from Table 1 and / or Table 4, and / or having an N-terminal addition of one or more amino acids (e.g., about 1 to about 50 amino acids). In embodiments, the PhiC31 integrase comprises at least 97% sequence identity to residues 1-604 of SEQ ID NO: 2 and having one or more mutations selected from Table 1 and / or Table 4, and / orhaving an N-terminal addition of one or more amino acids (e.g., about 1 to about 50 amino acids). In embodiments, the PhiC31 integrase comprises at least 98% sequence identity to residues 1-604 of SEQ ID NO: 2 and having one or more mutations selected from Table 1 and / or Table 4, and / or having an N-terminal addition of one or more amino acids (e.g., about 1 to about 50 amino acids). In embodiments, the PhiC31 integrase comprises at least 99% sequence identity to residues 1-604 of SEQ ID NO: 2 and having one or more mutations selected from Table 1 and / or Table 4, and / or having an N-terminal addition of one or more amino acids (e.g., about 1 to about 50 amino acids).
[0103] In embodiments, “mutation” herein (e.g., “X” in Table 2 and Table 4) refers to a mutation that differs from the wild-type amino acid. In embodiments, the mutation (e.g., “X” in Table 2 and Table 4) refers to a mutation from an amino acid to a stop codon, which truncates the polypeptide at that position. In embodiments, the mutation (e.g., “X” in Table 2 and Table 4) refers to a deletion of an amino acid at that position. In embodiments, the mutation is a conservative substitution (e.g., mutation to an amino acid with similar physicochemical properties). For example, in embodiments, the amino acid is mutated from a first hydrophobic (or non-polar) amino acid to a second hydrophobic (or non-polar) amino acid (e.g., selected from the group consisting of glycine (G), alanine (A), valine (V), leucine (L), isoleucine (I), proline (P), phenylalanine (F), methionine (M), and tryptophan (W)); from a first polar amino acid to a second polar amino acid (e.g., selected from the group consisting of serine (S), threonine (T), cysteine (C), asparagine (N), glutamine (Q), and tyrosine (Y)); from a first basic amino acid to a second basic amino acid (e.g., selected from the group consisting of arginine (R), lysine (K), and histidine (H)); from a first acidic amino acid to a second amino acid (e.g., selected from the group consisting of aspartate (D) and glutamate (E)). In embodiments, the mutation is a non-conservative substitution, where an amino acid from a first physicochemical group (e.g., hydrophobic / non-polar, polar, basic, acidic, etc.) is mutated to an amino acid belonging to a second physicochemical group, where the groups differ in one or more physicochemical properties.
[0104] In embodiments, the mutation (e.g., “X” in Table 2 and Table 4; or the bolded underline residues in SEQ ID NOs: 1 and 2) refers to a mutation from the indicated amino acid to a deletion, a stop codon, or to an amino acid selected from alanine (A), arginine (R), asparagine (N), aspartate (D), cysteine (C), glutamate (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V).
[0105] In embodiments, integrases herein (e.g., Bxb1 and / or PhiC31) comprise one or more nuclear localization sequences (NLSs). In embodiments, "nuclear localization signal" or "NLS" refers to an amino acid sequence that promotes import of a protein into the cell nucleus, for example, by nuclear transport. Nuclear localization sequences are known in the art and would be apparent to the skilled artisan. For example, NLS sequences are described in Plank et al., International PCT application, PCT / EP2000 / 011690, filed November 23, 2000, published as WQ / 2001 / 038547 on May 31, 2001; and in Kosugi et al., J. Biol. Chem. (2009) 284: 478-485, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. Nuclear localization sequences, in embodiments, include the nuclear localization sequence of the SV40 virus large T-antigen the minimal functional unit of which is the seven amino acid sequence PKKKRKV (SEQ ID NO: 57). Other examples of nuclear localization sequences include the nucleoplasmin bipartite NLS with the sequence NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 58); the c-myc nuclear localization sequence having the amino acid sequence PAAKRVKLD (SEQ ID NO: 59) or RQRRNELKRSF (SEQ ID NO: 60); and the hRNPAI M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 61). Further examples of nuclear localization sequences are: the sequenceRMRKFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 62) of the IBB domain from importin-alpha; the sequences VSRKRPRP (SEQ ID NO: 63) and PPKKARED (SEQ ID NO: 64) of the myoma T protein; the sequence PQPKKKPL (SEQ ID NO: 65) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO: 66) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO: 67) and PKQKKRK (SEQ ID NO: 68) of the influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 69) of the Hepatitis virus delta antigen; and the sequence REKKKFLKRR (SEQ ID NO: 70) of the mouse Mx1 protein. It is also possible, in embodiments, to use bipartite nuclear localization sequences such as the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 71) of the human poly(ADP-ribose) polymerase or the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 72) of the steroid hormone receptors (human) glucocorticoid. In embodiments, a NLS of the present disclosure comprises the amino acid sequence SV40 Large T-antigen NLS comprising the amino acid sequence DPKKKRKVDPKKKRKVDPKKKRKV (SEQ ID NO: 73). In embodiments, the NLS comprises the nucleic acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 74). In embodiments, the NLS comprises the nucleic acid sequence KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75). In embodiments, the NLS comprises the nucleic acid sequence GSHHHHHHGSGPKKKRKV (SEQ ID NO: 76). In embodiments, the NLS comprises the nucleic acid sequence GSGSGSHHHHHHGSGPKKKRKV (SEQ IDNO: 77) In embodiments, a NLS of the present disclosure comprises one or more mutations (e.g., amino acid substitutions, deletions, and / or insertions) with respect to any one of above sequences.
[0106] In embodiments, the one or more NLSs further improve the efficiency to the integrase by shunting the Bxb1 and / or PhiC31 into the nucleus where it carries out its recombination. In embodiments, the Bxb1 and / or PhiC31 has a single NLS. In embodiments, the NLS is located at the N-terminus. In embodiments, the NLS is located at the C-terminus. In embodiments, a NLS is located at the N-terminus and the C-terminus. In embodiments, the NLS is located within the protein sequence.
[0107] In embodiments, integrases herein (e.g., Bxb1 and / or PhiC31) comprise one or more stabilization tags. In embodiments, “stabilization sequence,” "stabilization tag," "stability tag, or “stability sequence," refers to an amino acid sequence that influences one or more of stability, solubility, folding, aggregation, degradation, purification, isolation, expression tracking, and / or subcellular localization, and the like, e.g., by promoting improvements of protein functionality relative to a cognate protein lacking the stability tag / sequence. Stabilization sequences for proteins are known in the art and would be apparent to the skilled artisan depending upon the desired functionality. For example, stabilization sequences are described in Khurana et al., “Chapter Seventeen - Elucidating the role of an immunomodulatory protein in cancer: From protein expression to functional characterization,” Methods in Enzymology, Vol. 629 (2019): pp. 307-60, the contents of which are incorporated herein by reference for their disclosure of exemplary stabilization tag sequences.
[0108] In embodiments, the stabilization sequence is or comprises about 5 amino acids to about or at least about 300 amino acids in length (e.g., about or at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, or 300 amino acids or more in length, including lengths therebetween). In embodiments, the stabilization tag is a “C-stab” (SEQ ID NO: 157). In embodiments, the one or more stabilization sequences is or comprises the amino acid sequence of any one of SEQ ID NOs: 157 and 162- 165, or an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes thereto. In embodiments, the stabilization tag is or comprises the amino acid sequence of SEQ ID NO: 157, or an amino acid sequence having 1 , 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes thereto.
[0109] Stabilization sequences, in embodiments, include poly-amino acid sequences, such as poly-Cys, poly-Arg, poly-Asp, poly-Phe, and poly-His (e.g., hexahistidine or His6x) tags. Other exemplary stabilization tag sequences include AU1 tag, AU5 tag, Glu-Glu tag EYMPME, FLAG tag, strep-tag II, HA tag, c-Myc tag, T7 tag, VSV-G tag, KT3 epitope tag, HSV tag, Protein C tag, V5 tag, S-tag, twin-strep tag, SBP tag, CBDtag, protein eXact tag. In embodiments, the stabilization tag is a protein domain, or a portion thereof, for example of ubiquitin, human serum albumin (HSA), SmbP, SUMO, Trx, Skp, Ecotin, SNAP-tag, slyd, DsbA, GST, Halo tag, MBP, NusA, -galactosidase, Fc-domain fusion (e.g., lgG1, lgG2, lgG3, lgG4, etc.), or CaBP. In embodiments, the stabilization tag is a protein domain, or a portion thereof useful for detection, e.g., such as RFP, GFP, CFP, mCherry, mOrange, SEAP, and luciferase. In embodiments, the stabilization sequence is or comprises an amino acid sequence that is or comprises a recognition site that is cleavable by one or more of thrombin, enterokinase, TEV protease, Factor Xa, SUMO protease, or HRV 3C protease.
[0110] In embodiments, integrases herein are fusion proteins comprising one or more DNA-binding domains. In embodiments, the one or more DNA-binding domains helps to target the integrase to a genomic att site by increasing the overall level of interaction (e.g., instead of relying purely on loop, helix, hairpin interactions, such as the case with Bxb1). In embodiments, the one or more DNA-binding domain comprises one or more TAL domains, one or more Cas proteins (Cas + guide RNA), and / or one or more zinc finger protein domains.
[0111] In embodiments, the one or more DNA-binding domains comprise one or more zinc finger proteins (ZFPs). Zinc finger proteins (ZFPs), in embodiments, are protein that contain zinc finger structural motifs, which are specific DNA-binding domains. In embodiments, these structural motifs are used to target and bind to specific DNA sequences within the genome. In embodiments, the one or more ZFPs comprises a C2H2 (Cys2His2) zinc finger (e.g., illustrative PROSITE identifier PS00028), composed of a -hairpin and an a-helix stabilized by a zinc ion, and recognizes about or at least about 3-4 base pairs of a DNA sequence. Exemplary native proteins with native ZFPs include MYST family histone acetyltransferases, Myt1 myelin transcription factor, and ST18 tumor suppressor protein. In embodiments, the one or more ZFPs is or comprises a Gag knuckle structural motif (e.g., illustrative PROSITE identifier PS50158), treble clef structural motif , zinc ribbon structural motif (e.g., illustrative PROSITE identifier PS51134), Zn2Cyse structural motif (e.g., illustrative PROSITE identifier cd00067), and / or TAZ2 domain motif. In embodiments, ZFPs have modular designs to target DNA triplets which are fused to the N-terminus, C-terminus, and / or within, the Bxb1 integrase. In embodiments, other DNA-binding domain ZFPs and similar DNA-binding domains comprise one or more of B-box zinc finger, DNA-binding protein, FPG lleRS zinc finger, Kruppel associated box, RING finger domain, TAL effector, transcription activator-like effector nuclease, zinc finger inhibitor, zinc finger nuclease, and zinc finger transcription factor, and / or are useful for structure-based design based on a designed genomic sequence.
[0112] Persons skilled in the art, with the benefit of this disclosure in its entirety, will understand how to design and clone TALE repeat domains, Cas domains, and ZFP domains into integrases (e.g., Bxb1) herein, for example using standard mutagenesis techniques, as well as design tools. For instance, non-limiting examples of zinc finger design tools include ZFtools, as described in Mandell and Barbas, “Zinc Finger Tools: custom DNA-binding domains for transcription factors and nucleases,” Nucleic Acids Research, Vol. 34 (2006): pp. W516-23; ZiFiT, as described in Sander et al., “Zinc Finger Targeter (ZiFiT): an engineered zinc finger / target site design tool,” Nucleic Acids Research, Vol. 35, (2007): pp. W599-W605; ZifBASE, as described in Jayakanthan et al., “ZifBASE: a database of zinc finger proteins and associated resources,” BMC Genomics, Vol. 10, No. 421 (2009); and ZF Design, as described in Ichikawa, et al. “A universal deeplearning model for zinc finger design enables transcription factor reprogramming,” Nat Biotechnol, Vol. 41 , (2023): pp. 1117-29, the entire contents of each are incorporated herein by reference in their entireties.
[0113] In embodiments, mutant integrases herein (e.g., Bxb1 and / or PhiC31) comprise an amino acid linker sequence, e.g., between a NLS and the N-term or C-term of the integrase, or between sections of a NLS, or between a stabilization sequence and the N-term or C-term of the integrase. Amino acid linker sequences are known in the art, for example, for producing fusion proteins, etc. Exemplary amino acid linker sequences, in embodiments, include glycine and serine (e.g., GG, GS, SS, SG, GGS, SSG, SGS, GSG, etc.), as well as repeated motifs and units thereof, repeated units / motifs consisting of glycine and serine residues, aspartic acid (D) residues, and the like. Amino acid linkers, in embodiments, are about 1-30 amino acids, such as a single amino acid in length, about 2 amino acids in length, about 3 amino acids in length, about 4 amino acids in length, about 5 amino acids in length, about 6 amino acids in length, about 8 amino acids in length, about 10 amino acids in length, about 12 amino acids in length, about 15 amino acids in length, about 20 amino acids in length, about 25 amino acids in length, or about 30 amino acids in length, including lengths therebetween. In embodiments, amino acid linkers are included between NLS peptides / signals (e.g., bipartite sequences). In embodiments, an amino acid linker is between the NLS and the enzyme. In embodiments, an amino acid linker is between a stabilization sequence (stability tag) and the enzyme. In embodiments, an amino acid linker is between a stabilization sequence (stability tag) and a NLS. In embodiments, an amino acid linker is between one or more DNA-binding domains (e.g., one or more ZFPs) and / or between one or more DNA-binding domains (e.g., one or more ZFPs) and the Bxb1 sequence.Methods of Using Integrases for Genomic Insertion
[0114] Described herein are methods of using integrases to modify genomic sequences (e.g., by insertion of one or more payload sequences). In embodiments, the methods herein include generating a genomic insertion site. In embodiments, the genomic insertion site is one or more exogenous attB or attP sites inserted within a genomic sequence (e.g., a bacterial / phage genomic consensus sequence inserted within a host cell genome) to serve as a safe harbor loci to diminish or ablate off-target or uncontrolled genomic insertion.
[0115] In embodiments, methods herein do not require knock-in of an aft site into the host cell genome, e.g., target one or more naturally-occurring endogenous genomic site / sequence.
[0116] In embodiments, methods herein comprise introducing at least one nucleic acid molecule comprising a first nucleic acid sequence encoding an integrase (e.g., Bxb1 and / or PhiC31 integrase as described herein). In embodiments, methods herein introducing at least one nucleic acid molecule comprising a second nucleic acid sequence encoding a genomic payload and one or more attB or attP site compatible with the Bxb1 integrase and / or PhiC31 integrase. In embodiments, methods herein introducing at least one nucleic acid molecule comprising a third nucleic acid sequence encoding one or more elements configured for genomic integration of one or more attB or attP site compatible with the Bxb1 integrase and / or the PhiC31 integrase.
[0117] In embodiments, methods herein comprise introducing the first nucleic acid sequence and the second nucleic acid sequence (e.g., where knock-in of an att site is not needed). In embodiments, the method comprises integrase-mediated recombination between the one or more attB or attP site and one or more naturally-occurring genomic sequences of the host cell.
[0118] In embodiments, methods herein comprise introducing the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence. In embodiments, the method comprises integrase-mediated recombination between the one or more attB or attP sequence and one or more genomically-integrated attB or attP sequence in the genome of the host cell.
[0119] In embodiments, methods herein comprise introducing one or more additional nucleic acid sequences which encodes one or more Bxb1 that targets a “left” and a “right” flanking portion of one or more genomic site in the host cell and / or one or more Bxb1 that targets the central att of the donor sequence (e.g., similarly to the methods as described in Example 4). In embodiments, “left” and “right” targeting Bxb1 are term -L and -R Bxb1 , respectively. In embodiments, “left” and “right” targeting Bxb1 are mutant Bxb1 havingone or more substitutions relative to SEQ ID NO: 1. In embodiments, “left” and “right” targeting Bxb1 are fusions with one or more DNA-binding domain.
[0120] In embodiments, methods herein comprise introducing into individual cells a series of Bxb1 molecules. In embodiments, integrase-mediated recombination reactions comprise introducing multiple nucleic acid sequences, which each encode a distinct integrase (e.g., two Bxb1 enzymes, three Bxb1 enzymes, four Bxb1 enzymes, etc.). In embodiments, methods comprise 3 Bxb1 enzymes - one Bxb1 which targets a “Left” side of a genome-located site that resembles the attB site, one Bxb1 which targets a “Right” side of the genome-located site that resembles the attB site, and a third Bxb1 which binds a sequence resembling an attP on the donor DNA, in embodiments, also referred to as the “helper.”
[0121] In embodiments, methods herein comprise one or more elements configured for genomic integration comprises one or more of a zinc finger nuclease (ZFN), transcription-activator like effector nuclease (TALEN), endonuclease, Cas9 nickase, a reverse transcriptase (e.g., PE5 or PE7), and a Cas9- directed reverse transcription guide RNA (pegRNA). In embodiments, knock-in of an aft site is achieved using one or more endonuclease, zinc finger nuclease (ZFN), transcription-activator like effector nuclease (TALEN), Cas-directed system, and the like. In embodiments, the third nucleic acid sequence encodes each of the Cas9 nickase, the reverse transcriptase PE5, and the Cas9-directed reverse transcription guide RNA (pegRNA).
[0122] In embodiments, each of the first nucleic acid sequence, second nucleic acid sequence, and / or third nucleic acid sequence is encoded on a single nucleic acid molecule, on two nucleic acid molecules, or on three nucleic acid molecules.
[0123] In embodiments, this process comprises using Cas9-directed reverse transcription for sitespecific genomic integration of a attB site using a fusion protein comprising a Cas9 nickase and a reverse transcriptase PE5, as well as a Cas9-directed reverse transcription guide RNA (pegRNA). In embodiments, the pegRNA is a modified single guide RNA (sgRNA) which both directed the PE5 to the target genomic sequence and encodes the desired edit. In embodiments, both the fusion protein and / or pegRNA are encoded in a plasmid or vector. In embodiments, the pegRNA contains a 5’ spacer sequence that directs PE5 to nick the target sequence. In embodiments, the primer binding site (PBS) found on the 3’ end of the pegRNA anneals to the non-target DNA strand following the nick in the target strand. In embodiments, the pegRNA contains a reverse transcriptase template (RTT) that is reverse transcribed to generate a 3’ single-strand flap of DNA that encodes the inserted DNA, such as the aft site. During twin Cas9-directed reverse transcription,in embodiments, a second pegRNA directs the insertion of an adjacent flap that includes overlapping homologous sequence so that the two flaps form a double-stranded intermediate. Upon resolution of the intermediate, in embodiments, the new sequence becomes permanently inserted into the host genome.
[0124] For example, in embodiments, Cas9-directed reverse transcription efficiently inserts sequences <1 OObp. In embodiments, Cas9-directed reverse transcription is first used to insert the 35 bp attB site at a desired genomic locus, such as TRAC or ROSA26. In embodiments, the attB site acts as a insertion site for PhiC31 integrase which can efficiently insert multi-kb sequences. In embodiments, during a second step, the integrase recombines the attB site found in the genome with an attP site found on the 6.6 kb donor plasmid. Following recombination, in embodiments, the entire donor plasmid becomes inserted into the genome and the attB and attP sites are rearranged to form attL and attR sites flanking the insertion. Because the integrase cannot recombine the newly formed attL and attR sites, in embodiments, the donor plasmid contains an antibiotic resistance gene and DsRed florescent protein gene is permanently inserted into the genome.
[0125] Cas9-directed reverse transcription efficiency is generally higher for shorter inserts, thus in embodiments, the pegRNA is designed for inserting the minimal attB site and / or the minimal attP site.
[0126] In embodiments, the minimal attB site is about or at least about 25 nucleotides in length, about or at least about 30 nucleotides in length, about or at least about 31 nucleotides in length, about or at least about 32 nucleotides in length, about or at least about 33 nucleotides in length, about or at least about 34 nucleotides in length, about or at least about 35 nucleotides in length, about or at least about 36 nucleotides in length, about or at least about 37 nucleotides in length, about or at least about 38 nucleotides in length, about or at least about 39 nucleotides in length, about or at least about 40 nucleotides in length, or about or at least about 50 nucleotides in length.
[0127] In embodiments, the minimal attP site is about or at least about 30 nucleotides in length, about or at least about 35 nucleotides in length, about or at least about 36 nucleotides in length, about or at least about 37 nucleotides in length, about or at least about 38 nucleotides in length, about or at least about 39 nucleotides in length, about or at least about 40 nucleotides in length, about or at least about 41 nucleotides in length, about or at least about 42 nucleotides in length, about or at least about 43 nucleotides in length, about or at least about 44 nucleotides in length, about or at least about 45 nucleotides in length, about or at least about 50 nucleotides in length, or about 60 nucleotides in length. For example, in embodiments, the minimal attB site and / or the minimal attP site comprises the PhiC31 attB site (35 bp) or attP site (41 bp). In embodiments, the minimal attB site and / or the minimal attP site comprises the Bxb1 attB site or attP site.
[0128] In embodiments, the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity (e.g., at least about 75%, 80%, 90%, 95%, or 99%) to the nucleic acid sequence of: ctcgaagccgcggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccacctcacccatc (SEQ ID NO: 79), or tcggccggcttgtcgacgacggcggtctccgtcgtcaggatcatccgggc (SEQ ID NO: 78). In embodiments, the attB site for Bxb1 comprises a nucleic acid sequence having at least about 90% sequence identity to tcggccggcttgtcgacgacggcggtctccgtcgtcaggatcatccgggc (SEQ ID NO: 78). In embodiments, the attB site comprises the nucleic acid sequence of SEQ ID NO: 78. In embodiments, the one or more attB sites is or comprises a DNA sequence having the nucleic acid sequence of: ctcgaagccgcggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccacctcacccatc (SEQ ID NO: 79), or tcggccggcttgtcgacgacggcggtctccgtcgtcaggatcatccgggc (SEQ ID NO: 78).
[0129] In embodiments, the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity (e.g., at least about 75%, 80%, 90%, 95%, or 99%) to the nucleic acid sequence of: tggtgtccaggagccgaggtatcggtcctgccagggcc (SEQ ID NO: 166). In embodiments, this attB sequence is useful for targeting the TRAC locus, e.g., as a naturally-occurring genomic sequence.
[0130] In embodiments, the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity (e.g., at least about 75%, 80%, 90%, 95%, or 99%) to the nucleic acid sequence of: ctgagcgcctctcctgggcttgccaaggactcaaaccc (SEQ ID NO: 167). In embodiments, this attB sequence is useful for targeting the AAVS1 locus, e.g., as a naturally-occurring genomic sequence.
[0131] In embodiments, the one or more attP sites is or comprises a DNA sequence having about or at least about 70% sequence identity (e.g., at least about 75%, 80%, 90%, 95%, or 99%) to the nucleic acid sequence of: cgggagtagtgccccaactggggtaacctttgagttctctcagttgggggcgtagggtcg (SEQ ID NO: 81), or gtggtttgtctggtcaaccaccgcggtctcagtggtgtacggtacaaaccca (SEQ ID NO: 80). In embodiments, the attP site for Bxb1 comprises a nucleic acid sequence having at least about 90% sequence identity to gtggtttgtctggtcaaccaccgcggtctcagtggtgtacggtacaaaccca (SEQ ID NO: 80). In embodiments, the attP site comprises the nucleic acid sequence of SEQ ID NO: 80. In embodiments, the one or more attP sites is or comprises a DNA sequence having the nucleic acid sequence of: cgggagtagtgccccaactggggtaacctttgagttctctcagttgggggcgtagggtcg (SEQ ID NO: 81), or gtggtttgtctggtcaaccaccgcggtctcagtggtgtacggtacaaaccca (SEQ ID NO: 80).
[0132] In embodiments, the one or more attB site for PhiC31 comprises a nucleic acid sequence having at least about 70% sequence identity (e.g., at least about 75%, 80%, 90%, 95%, or 99%) toctcgaagccgcggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccacctcacccatc (SEQ ID NO: 79). In embodiments, the one or more attB site comprises the nucleic acid sequence of SEQ ID NO: 79. In embodiments, the one or more attP site for PhiC31 comprises a nucleic acid sequence having at least about 70% sequence identity (e.g., at least about 75%, 80%, 90%, 95%, or 99%) to cgggagtagtgccccaactggggtaacctttgagttctctcagttgggggcgtagggtcg (SEQ ID NO: 81). In embodiments, the one or more attP site comprises the nucleic acid sequence of SEQ ID NO: 81.
[0133] In embodiments, the attB and / or attP site targets to a central dinucleotide sequence comprising guanine-thymine (GT) or guanine-adenine (GA). In embodiments, there are 16 possible attB and / or attP central dinucleotide sequences which Bxb1 can target. In embodiments, the non-wild-type Bxb1 integrases herein are able to target one of the non-canonical 15 central dinucleotide sequences (e.g., central dinucleotide that is not GT, such as GG, GA, GC, AT, AA, AG, AC, TT, TA, TG, TC, CC, CA, CT, CG).
[0134] Persons skilled in the art, with the benefit of this disclosure in its entirety, will understand how to engineer an attP sequence compatible with an attB sequence (e.g., such as a sequence found in the genomic), and likewise would understand how to test attB sequences for integration (e.g., with Bxb1), using for example, standard mutagenesis, ddPCR, cell culturing, and sequencing techniques.
[0135] In embodiments, the pegRNA (epegRNA) comprises one or more nucleic acid sequences having one or more of a structured RNA motif, a cr772 guide scaffold, and / or a dominant negative mutant of human mutL homolog 1 (MLH1) involved in DNA mismatch repair. To maximize Cas9-directed reverse transcription efficiency, in embodiments, methods herein use an engineered pegRNA (epegRNA) containing a structured RNA motif to prevent degradation, an optimized cr772 guide scaffold, and expressed PE5 from the PE5max plasmid with an improved editor architecture that includes a dominant negative mutant of human mutL homolog 1 (MLH1) involved in DNA mismatch repair which has been shown to enhance Cas9-directed reverse transcription. PE5max and twin epegRNA expression plasmids for each target site were transfected into a host cell.
[0136] In embodiments, methods herein, combine the Cas9-directed reverse transcription and integrase expression plasmids into a single nucleic acid introduction event (e.g., in a single transfection). In embodiments, this strategy allows for the insertion of one or more att site using Cas9-directed reverse transcription followed by the insertion of one or more integration sequence by one or more integrases at the target sequence.
[0137] In embodiments, methods herein are compatible with a variety of methods for knock-in of att sequence within the host cell genome. In embodiments, such methods include, without limitation, introducing single strand DNA breaks (e.g., nick-mediated insertion or tandem paired nicking), double strand DNA breaks (e.g., with non-homologous end joining (NHEJ) or homology-directed repair (HDR)), transposon-based insertion, and the like. Persons skilled in the art, with the benefit of this disclosure in its entirety, will be aware of the various techniques that can be used to insert att sites within host cell genomes for use with methods of integrases herein.
[0138] In embodiments, methods herein include the use of one or more zinc-fi nger nucleases (ZFNs) to introduce the one or more att sites into the host cell genome. For example, in embodiments, the one or more nucleic acids encoding the one or more att site can be integrated into the host cell genome using ZFN mRNA and one or more ULTRAMER nucleic acid construct from INTEGRATED DNA TECHNOLOGIES. In embodiments, the ZFN generates a double-strand break that facilitates integration via homology directed repair (HDR) (e.g., double-stand DNA break). In embodiments, ZFNs include engineered ZFNs (e.g., Fokl nuclease engineered domains). In embodiments, ZFN and ZFP design and use for targeting genomic sequences and / or genomic knock-in is known in the art, for instance, as described in U.S. Patent Nos. 5,356,802, 6,265,196, 6,534,261, 6,933,113, 7,163,824, and 7,220,719, each of which is incorporated by reference herein.
[0139] In embodiments, methods herein include the use of one or more CRISPR / Cas endonucleases, transcription activator-like effector nucleases (TALENs), TALE-derived transcription factors, TALE repeat domain proteins, meganucleases, restriction enzymes, site-specific nucleases, and / or gene-editing systems to introduce the one or more att sites into the host cell genome. In embodiments, TALEN and TAL domain design and use for genomic knock-in is known in the art, for instance, as described in U.S. Patent Nos. 9,353,378, 9,453,054, 9,809,628, 8,450,471 , 8,586,363, 9,758,775, 10,400,225, 8,912,138, 9,493,750, and 8,586,526, each of which is incorporated by reference herein.
[0140] In embodiments, such systems include one or more proteins and / or nucleic acids working in concert, including without limitation, TALENs, ZFNs, RNase P RNA, RNase H, CRISPR / Cas, C2c1, C2c2, C2c3, Cas9, Cpf1 , TevCas9, Archaea Cas9, CasY.1 , CasY.2, CasY.3, CasY.4, CasY.5, CasY.6, CasX Cas omega, transposase, and / or any ortholog or homolog thereof. In embodiments, the gene editors can also include an RNA molecule (e.g., gRNA, which, as used herein, refers to guide RNA). In embodiments, the gRNA is sequence complimentary to a coding or a non-coding sequence and can be tailored to the particularsequence to be targeted (e.g., targeting to a particular genomic locus). In embodiments, the gRNA sequence can be a sense or anti-sense nucleic acid sequence. In embodiments, when a gene editor composition is administered herein, preferably without limitation, including two or more gRNAs; however, a single gRNA can also be used. In embodiments, Bxb1 mutants herein comprises one or more elements from a gene editor, e.g., to improve targeting to a genomic DNA sequence.
[0141] In embodiments, the host cell includes cells for expression of therapeutic proteins and / or for producing clonal cell populations with genetic amendments. In embodiments, the host cell is a mammalian cell, plant cell, insect cell, yeast cell, or bacterial cell. In embodiments, the mammalian host cell includes a cell line, such as without limitation, a human embryonic kidney (HEK293, HEK293T) cell line, K562 human lymphoblast cell line, U2OS human osteosarcoma cell line, primary human fibroblasts (human dermal fibroblast (HDFa) cell line), human erythroleukemia cell line (K562), Chinese hamster ovary (CHO) cell line (CHO-K1, CHO-DHB11, CHO-DXB1 , CHO-S, CHO-DG44, CHO-M), baby hamster kidney (BHK) cell line, Vero cell line (Vero, Vero 76, Vero E6), human cervical carcinoma cell line (HELA, 3T3), PERc6 cell line, CAP cell line, induced pluripotent stem cells (iPSCs), human embryonic stem cells (ESCs), mouse cell line, or monkey kidney CV1 cell line. In embodiments, methods herein include genomic manipulation of induced pluripotent stem cells (iPSCs) or human embryonic stem cells (ESCs). In embodiments, methods herein include genomic manipulation of primary cells. In embodiments, methods herein include genomic manipulation of tissue-specific cells, including in non-limiting examples, human T-cells, NK cells, HSCs, CD34+ iPSCs, human tissue cells such as heart cells, lung cells, liver cells, muscle cells, brain cells, pancreas cells, etc. Without wishing to be bound by theory, methods herein are adaptable to a wide variety of cell types and organismal systems, e.g., from human cells to plant cells, due to the high level of homology between sequences targeted herein. For example, without wishing to be bound by theory, it is anticipated that there may on the order of 1-20 nucleotide differences in target sequences between species.
[0142] In embodiments, methods herein utilize any mutant Bxb1 integrase as described herein. In embodiments, the mutant Bxb1 integrase is encoded in a nucleic acid, where the nucleic acid is introduced into the cell. In embodiments, the nucleic acid encoding the mutant Bxb1 integrase is or comprises DNA (e.g., a plasmid, vector, or DNA minicircle), where the DNA is double-stranded, single-stranded, linear, or circular, or comprises RNA (e.g., mRNA or circular RNA). In embodiments, the nucleic acid encoding the mutant Bxb1 integrase is part of a non-viral plasmid or vector, or a viral plasmid or vector.
[0143] In embodiments, the nucleic acid is or comprises a nucleic acid sequence having about or at least about 70%, about or at least about 75%, about or at least about 80%, about or at least about 85%, about or at least about 90%, about or at least about 95%, about or at least about 96%, about or at least about 97%, about or at least about 98%, or about or at least about 99% sequence identity to the nucleic acid sequence of any one of SEQ ID NOs: 89-118 (e.g., in Table 14). In embodiments, the nucleic acid is or comprises the nucleic acid sequence of any one of SEQ ID NOs: 89-118.
[0144] In embodiments, the nucleic acid encoding the mutant Bxb1 integrase (or any sequence to be inserted into the genome) belongs to a viral plasmid or vector, or comprises at least a portion of a viral backbone. In embodiments, the viral vector is an adeno-associated virus (AAV) vector, recombinant adeno- associated virus (rAAV) vector, or lentiviral-based vector used for delivery into target cells. In embodiments, the AAV vector comprises one or more serotypes selected from AAV-1 , AAV-2, AAV-3, AAV-4, AAV-5, AAV- 6, AAV-7, AAV-8, AAV-9, AAV-10, AAV-11 , AAV-12, AAV-13, AAVrh.74 and tropism modified AAV vectors, e.g., depending upon the type of cell or desired targeted delivery.
[0145] In embodiments, the nucleic acid comprises one or more elements of a promoter, enhancer, internal ribozyme entry sequence (IRES), and the like, that are useful elements for protein expression. In embodiments, the nucleic acids herein comprise one or more inducible nucleic acid elements which are useful for controlling expression of one or more operably linked nucleic acid sequences. In embodiments, the nucleic acids herein comprise one or more sequences encoding a protein associated with a selectable marker, e.g., such as an antibiotic resistance cassette which encodes an enzyme for cellular resistance to an antibiotic (e.g., ampicillin, kanamycin, neomycin, zeocin, etc.). In embodiments, the nucleic acids herein comprise one or more sequences encoding a metabolic gene for auxotrophic culturing, e.g., dihydrofolate reductase (dhfr) gene for DHFR-based expression selection system.
[0146] In embodiments, Cas9-directed reverse transcription and integrase expression plasmids are introduced into a host cell simultaneously, e.g., in a single transfection and / or transduction. In embodiments, introduction of nucleic acids into host cells is achieved using one or more of a lipid-based transfection reagent (e.g., LIPOFECTAMINE (INVITROGEN), TRANSIT-2020 (MIRUS BIO), HELA-MONSTER (MIRUS BIO), NEON TRANSFECTION SYSTEM (INVITROGEN), diethylaminoethyl (DEAE)-dextran, liposomes, cationic lipid-based reagents, etc.), electroporation, sonoporation, chemical reagent (e.g., calcium phosphate, etc.), microinjection, or via a viral vector system (e.g., AAV, lentiviral, etc.). In embodiments, integrases are delivered as DNA (e.g., plasmid or vector) or RNA (e.g., mRNA). In embodiments, integrases are deliveredas part of a viral vector system (e.g., AAV, lentiviral, etc.). In embodiments, integrases are delivered as DNA minicircle.
[0147] In embodiments, the efficiency of Cas9-directed reverse transcription insertion of att sites into host cell genomes ranges from 0.1 % to 90% or more of a population of host cells subjected to the fusion protein and pegRNA. In embodiments, the efficiency of Cas9-directed reverse transcription insertion of att sites into host cell genomes is about or at least about 0.5%, about or at least about 1 .0%, about or at least about 5.0%, about or at least about 10%, about or at least about 15%, about or at least about 20%, about or at least about 30%, about or at least about 40%, about or at least about 50%, about or at least about 60%, about or at least about 70%, or about or at least about 90%. In embodiments, efficiency of Cas9-directed reverse transcription insertion of att sites (and / or of genomic insertion of a genomic payload sequence) is measurable using amplicon sequencing with genomic primers flanking the insert followed by confirmation by a polymerase chain reaction (e.g., a droplet digital PCR (ddPCR), which a form of quantitative PCR that reports highly accurate copy number of genomic inserts over a wide range of template concentrations), and / or sequencing (next-generating sequencing (NGS), exome sequencing (WES), whole genome sequencing (WGS), or Sanger sequencing), and / or based on expression of an inserted sequence (e.g., a reporter gene). For example, in embodiments, the efficiency of Cas9-directed reverse transcription insertion of att sites into host cell genomes ranges from 0.3% to 62.1 % as measured by amplicon sequencing and ddPCR.
[0148] In embodiments, genomic insertion herein includes integration of nucleic acids ranging in size from about 0.1 kb to about 20 kb or more. In embodiments, the genomic integration includes knock-in of one or more nucleic acids of about or at least about 0.5 kb in length, about or at least about 1 kb in length, about or at least about 2 kb in length, about or at least about 3 kb in length, about or at least about 4 kb in length, about or at least about 5 kb in length, about or at least about 6 kb in length, about or at least about 7 kb in length, about or at least about 8 kb in length, about or at least about 9 kb in length, about or at least about 10 kb in length, about or at least about 11 kb in length, about or at least about 12 kb in length, about or at least about 13 kb in length, about or at least about 14 kb in length, about or at least about 15 kb in length, about or at least about 16 kb in length, about or at least about 17 kb in length, about or at least about 18 kb in length, about or at least about 19 kb in length, about or at least about 20 kb in length, about or at least about 30 kb, or about or at least about 50 kb. In embodiments, the nucleic acid is DNA (e.g., double-stranded DNA, dsDNA).
[0149] In embodiments, the epegRNA is targeted to one or more loci in the host cell genome. In embodiments, the methods herein target, without limitation, to a locus of AAVS1, Rosa26, Xq22.1 , MACO1 , GYS1, CCR5, ACTB, GBA1 , COL7A1 , FANCA, Smn1 , ALB, B2M, TRAC, Factor IX, CFTR. In embodiments, these sites are amenable to becoming safe harbor sites for site-specific targeted genomic insertion. In embodiments, the methods herein target to a chromosomal location of chromosome 1 , chromosome 2, chromosome 3, chromosome 4, chromosome 6, chromosome 7, chromosome 8, chromosome 10, chromosome 12, chromosome 13, chromosome 17, chromosome 19, chromosome 20, or chromosome 22. In embodiments, methods herein demonstrate specificity to substantially only target a single 1 chromosomal site (e.g. loci), or 2 chromosomal sites, or 3 or fewer chromosomal sites. In embodiments, methods herein demonstrate higher specificity than comparable methods using a wild-type integrase (e.g., compared to genomic integration using wild-type Bxb1 or PhiC31).
[0150] In embodiments, methods herein, in place of using a cell containing a pre-inserted att site, combine the Cas9-directed reverse transcription and integrase expression plasmids into a single nucleic acid introduction event (e.g., in a single transfection). In embodiments, this strategy allows for the insertion of one or more att site using Cas9-directed reverse transcription followed by the insertion of the donor plasmid by integrase (e.g., Bxb1 or PhiC31) at the target sequence. For example, in embodiments, co-transfection of PE5max and epegRNA plasmids insert the 41 bp PhiC31 attP site at ROSA26 along with a donor plasmid containing the attB site and a helper plasmid expressing PhiC31 integrase. In such embodiments, PhiC31 integrase insertion of the donor plasmid is detectable for all four integrase variants with a favorable efficiency of 2.3% for P3 compared to 0.7% for wild-type. In embodiments, mutant integrases (e.g., “hyperactive” integrases) outperform their cognate wild-type integrases during Cas9-directed reverse transcription- mediated gene insertion. For example, hyperactive PhiC31 variants outperform wild-type PhiC31 integrase.
[0151] In embodiments, the genetic cargoes are single or multi-gene cargos. In embodiments, the integration rates are as high as 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% or more of target chromosomes (e.g., loci) in host cells. In embodiments, the hyperactive integrases (e.g., PhiC31 and Bxb1 serine integrases) insert therapeutic DNA cargoes ranging in size from about 0.1 kb to about 20 kb or more (e.g., about 0.5 kb, about 1 kb, about 2 kb, about 3 kb, about 4 kb, about 5 kb, about 6 kb, about 7 kb, about 8 kb, about 9 kb, about 10 kb, about 11 kb, about 12 kb, about 13 kb, about 14 kb, about 15 kb, about 16 kb, about 17 kb, about 18 kb, about 19 kb, and larger). For example, in embodiments, a large, 15.7 kb therapeutic DNA cargo encoding Von Willebrand Factor (vWF) was inserted into host cell genomes with at least about 80% efficiency (e.g., as described in Example 1).
[0152] Anzalone et al. previously showed that the aft site found on the epegRNA expression plasmid may recombine with the aft site on the donor plasmid, rendering both plasmids nonfunctional, which reduces the amount of epegRNA available for Cas9-directed reverse transcription, as well as reduce the number of donor plasmids available for integrase insertion. To reduce this effect, in embodiments, the epegRNA expression plasmids are engineered to contain a partial aft site. In embodiments, during twin Cas9-directed reverse transcription, each inserted 3’ single strand flap contains only part of the aft site with an overlapping sequence to form a double-stranded intermediate. In embodiments, the remaining single-stranded sequence of the aft sites are subsequently filled in to promote insertion of the full aft site. In embodiments, this strategy is sufficient for reducing or preventing unwanted integrase recombination.
[0153] In embodiments, the genomic payload comprises one or more of a therapeutic protein. In embodiments, the therapeutic protein comprises an antibody or antibody-format protein, therapeutic enzyme, fusion protein, secretory protein, protein hormone, and / or protein toxin or antitoxin, chimeric receptor, chimeric antigen receptor.
[0154] In embodiments, the genomic payload (e.g., to be inserted) comprises a chimeric receptor (e.g, anti-CD19 chimeric antigen receptor (CAR)), a promoter or enhancer sequence, and / or a genetic cassette encoding a therapeutic protein corrective protein (e.g., a non-mutant form of a protein that is mutant or otherwise defective in the host cell, for example, for protein replacement therapy).
[0155] In embodiments, the genomic payload comprises one or more exogenous genes encoding one or more parts of a chimeric antigen receptor (CAR). In embodiments, the one or more exogenous genes encodes CAR. In embodiments, the CAR comprises a transmembrane domain of or derived from CD28, CD3 , CD4, CD8 a , ICOS, or fragment and / or combination thereof. In embodiments, the CAR comprises an intracellular domain and / or an intracellular signaling domain, optionally comprising one or more of a CD3- chain and FcRy chain. In embodiments, the CAR comprises one or more co-stimulatory molecules, optionally comprising one or more of CD28, 4-1 BB, ICOS, CD27, and 0X40. In embodiments, the CAR comprises a CD8 or I gG4-derived hinge region (e.g., linking the antigen-binding portion to the transmembrane portion). In embodiments, the CAR is a fusion protein comprising an extracellular (or outwardly facing) binding moiety (e.g., which comprises one or more scFv or antibody-format binding molecule), connected by a hinge peptide (e.g., CH2 / CH3 domains from an IgG Fc region, a peptide linker comprised predominantly of glycine and serine, a CD28 peptide, CD8a peptide, etc.) to a transmembrane domain (e.g., CD28, CD3<, CD4, CD8a, ICOS, etc.), followed by a variety of intracellular signaling domains and / or co-stimulatory domains (e.g. 4-1 BB, CD3 , CD28, 4-1 BB, I COS, CD27, 0X40, etc.). In embodiments, the CAR comprises an antigen-binding portion which includes an antibody or antibody format binding moiety, for example, comprising the antigenbinding portion of a monoclonal antibody, polyclonal antibody, antibody fragment, Fab, Fab', Fab'-SH, F(ab')2, Fv, single chain Fv (scFv), diabody, nanobody, bispecific antibody, chimeric antibody, humanized antibody, human antibody, and fusion protein comprising the antigen-binding portion of an antibody.
[0156] In embodiments, the one or more exogenous genes comprises one or more protein-coding sequence or non-protein-coding sequences that improves T cell functionality (e.g., survival, exhaustion, cytotoxicity, etc.). In embodiments, the one or more exogenous genes comprises one or more protein-coding sequence or non-protein-coding sequences that boosts the therapeutic effectiveness of primary cells (e.g., hematopoietic stem cells, liver cells, heart cells, muscle cells, neurons, etc.).
[0157] In embodiments, methods herein further comprises measuring the insertion efficiency (e.g., insertion of one or more att site or genomic payload sequence). In embodiments, insertion efficiency is measured by performing one or more of flow cytometry, confocal microscopy (confocal laser scanning microscopy, spinning-disk confocal microscopy), in situ fluorescence, and immunohistochemistry, for example, for detecting expression of one or more fluorescence proteins, e.g., expression from one or more knocked-in fluorescence marker protein, or fluorescence staining for expression of one or more knocked-in marker protein, e.g., such as a chimeric antigen receptor (CAR), or for tracking the subcellular localization of an integration. In embodiments, insertion efficiency is measured by performing one or more of SDS-PAGE, western blotting, enzyme kinetics, and / or enzyme-linked immunosorbent assay (ELISA), for example, for detecting expression of one or more knocked-in marker protein, such as a CAR, cytosolic / nuclear protein, cell surface marker, integrase, etc. In embodiments, insertion efficiency is measured by performing one or more of long-read sequencing, droplet digital PCR (ddPCR), reverse transcriptase PCR (RT-PCR), quantitative or real-time PCR (RT-PCR), amplicon sequencing, and / or Sanger sequence, for example, to measure or detected integrase-mediated recombination events at desired locations, detection of genomically- integrated att sites, and / or detection of off-target genomic destabilization.Pharmaceutical Compositions
[0158] In aspects, the disclosure provides pharmaceutical compositions comprising one or more nucleic acids encoding a mutant Bxb1 and / or one or more nucleic acids encoding an attB or attP sequence (as described herein) compatible for recombination with an attB or attP (as described herein). In embodiments, the pharmaceutical composition comprises one or more pharmaceutically acceptable excipient or carrier, thefirst nucleic acid molecule of any of the embodiments disclosed herein, and the second nucleic acid molecule of any of the embodiments disclosed herein.
[0159] In embodiments, the one or more pharmaceutically acceptable excipient or carrier comprises one or more of saline, solvents, dispersion media, coatings, antibacterial agent, antiviral agent, antifungal agents, isotonicity agent, buffer, absorption delaying agent, or chelating agent. In embodiments, the pharmaceutical composition is formulated to be compatible with its intended route of administration. In embodiments, the pharmaceutical composition of formulated for parenteral administration, e.g., intravenous, intradermal, subcutaneous, intraperitoneal, transdermal, subdermal, or transmucosal administration. In embodiments, the pharmaceutical composition is formulated for delivery and / or expression of the first nucleic acid molecule and / or the second nucleic acid molecule in a particular tissue or organ, such as the liver, muscle, skin, brain, eyes, lungs, etc.
[0160] Pharmaceutical compositions suitable for injectable use include sterile aqueous solutions (where water soluble) or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersion. For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, Cremophor EL™ (BASF, Parsippany, N.J.) or phosphate buffered saline (PBS). In embodiments, the composition is sterile and is fluidic to the extent that easy syringability exists. In embodiments, the composition is stable under the conditions of manufacture and storage. In embodiments, the composition is free of contaminating action of microorganisms such as bacteria, mycoplasma, spores, virus, and fungi. The carrier can be a solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, polyetheylene glycol, and the like), and suitable mixtures thereof. Prevention of the action of microorganisms can be achieved by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, ascorbic acid, thimerosal, and the like. In many cases, it will be preferable to include isotonic agents and / or cryoprotectants, for example, sugars, polyalcohols such as manitol, sorbitol, sodium chloride, magnesium chloride, dimethylsulfoxide (DMSO), etc., in the composition. Prolonged absorption of the injectable compositions can be brought about by including in the composition an agent which delays absorption, for example, aluminum monostearate and / or gelatin. Pharmaceutically compatible binding agents, and / or adjuvant materials can be included as part of the composition.
[0161] Sterile injectable solutions can be prepared by incorporating the active compound in the required amount in a selected solvent with one or a combination of ingredients enumerated above, as required,followed by filtered sterilization. Generally, dispersions are prepared by incorporating the active compound into a sterile vehicle, which contains a basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation are vacuum drying and freeze-drying which yields a powder of the active ingredient plus any additional desired ingredient from a previously sterile-filtered solution thereof.Methods of Treatment
[0162] In aspects, described herein are methods of treating a disease or disorder in a subject in need thereof, the method comprising administering to the subject a pharmaceutical composition or composition herein (or using a kit or components thereof, as described herein) comprising a first nucleic acid molecule of any of the embodiments disclosed herein and a second nucleic acid molecule of any of the embodiments disclosed herein.
[0163] In embodiments, the first nucleic acid molecule encodes a Bxb1 integrase variant of any of the embodiments disclosed herein. In embodiments, the second nucleic acid molecule comprises an attB or attP sequence of any of the embodiments disclosed herein, and optionally further comprising a donor DNA payload, where the attB or attP site is compatible for recombination using the Bxb1 with a cognate attB or attP site (as described in any embodiments herein). In embodiments, the pharmaceutical composition is formulated for delivery and / or expression of the first nucleic acid molecule and / or the second nucleic acid molecule in in a particular tissue or organ, such as the liver, muscle, skin, brain, eyes, lungs, etc.
[0164] In embodiments, the genomic payload comprises a gene therapy, for example, for gene replacement therapy. In embodiments, the genomic payload comprises a protein-coding sequence of one or more proteins to correct a deficiency to treat a disease or disorder in a subject. In embodiments, the disease is caused by insufficiency of a protein and the payload comprises an open reading frame capable of encoding the protein (without limitation, e.g., a blood clotting factor). In embodiments, the disease or disorder is a genetic disease or disorder, or a disease or disorder from a defective protein-coding sequence.
[0165] In embodiments, the genomic payload comprises one or more promoter sequence, enhancer sequence, internal ribosome entry site (IRES), and 3’ polyadenine (polyA) sequence, and / or a genetic cassette encoding a therapeutic protein or corrective protein. In embodiments, the genomic payload comprises one or more non-coding RNA.
[0166] In embodiments, the genomic payload encodes one or more of a therapeutic protein. In embodiments, the therapeutic protein comprises an antibody or antibody-format protein, therapeutic enzyme, fusion protein, secretory protein, protein hormone, and / or protein toxin or antitoxin, chimeric receptor, chimeric antigen receptor.
[0167] In embodiments, the genomic payload comprises one or more non-coding RNA. In embodiments, the non-coding RNA functions to assist in the expression of an exogenous sequence (e.g., a protein-coding sequence or no-protein coding sequence not typically expressed in the host cell). For example, in embodiments, the non-coding RNA is a tRNA that corrects a missense or a nonsense mutation (e.g., a suppressor tRNA), or a tRNA that assists in expression of polypeptides with rare codons for the host cell type. In embodiments, the genomic payload comprises one or more non-coding RNAs. In embodiments, the non-coding RNA is a RNA that functions to modulate one or more endogenous cell factors (e.g., to suppress or reduce the expression of an endogenous protease, reduce or increase post-translational modification, etc.). In embodiments, the non-coding RNA is a regulatory RNA. In embodiments, the noncoding RNA is an miRNA. In embodiments, the non-coding RNA is siRNA In embodiments, the non-coding RNA is shRNA. In embodiments, the non-coding RNA is aptamer. In embodiments, the non-coding RNA is riboswitch. In embodiments, the non-coding RNA is a tRNA.Nucleic Acids, Host Cells, and Kits
[0168] Described herein, in embodiments, are compositions comprising the evovled integrases, Cas9 nickase fusion proteins (for Cas9-directed reverse transcription), pegRNA, att sequences, etc., encoding in one or more nucleic acid sequences (e.g., DNA or RNA). In embodiments, the nucleic acid is DNA (e.g., a non-viral or viral plasmid, vector, or minicircle). In embodiments, the nucleic acid is RNA (e.g., mRNA, circular RNA, etc.). In embodiments, the components described herein are encoded on multiple nucleic acid; alternatively each is present on a single nucleic acid molecule. In embodiments, the nucleic acid is configured to be introduced into a host cell, e.g., using methods described herein.
[0169] Described herein, in embodiments, are host cells, e.g., as described herein which comprise one or more integrase Cas9 nickase fusion proteins, pegRNA (for Cas9-directed reverse transcription), att sequences, etc..
[0170] In embodiments, enzymes, nucleic acids, and / or host cells of the present disclosure are assembled into a kit. In embodiments, the kit comprises enzymes, nucleic acids, and / or host cells in one or more formulations for use in one or more methods as described herein.
[0171] The kit described herein may include one or more containers housing components for performing the methods described herein and optionally instructions for use. Any of the kits described herein may further comprise components needed for performing the manufacturing methods described herein. Each component of the kits, where applicable, may be provided in liquid form (e.g., in aqueous solution, a buffer, or cell media). In embodiments, some of the components are reconstitutable or otherwise processible (e.g., nucleic acids that reconstitute in nuclease-free water, or frozen cell stocks for seeding plates), for example, by the addition of a suitable solvent or other species, which may or may not be provided with the kit.
[0172] In embodiments, the kits may optionally include instructions and / or promotion for use of the components provided. As used herein, "instructions" can define a component of instruction and / or promotion, and typically involve written instructions on or associated with packaging of the disclosure. Instructions also can include any oral or electronic instructions provided in any manner such that a user will clearly recognize that the instructions are to be associated with the kit, for example, audiovisual (e.g., videotape, DVD, etc.), Internet, and / or web-based communications, etc. As used herein, "promoted" includes all methods of doing business including methods of education, engineering instruction, scientific inquiry, discovery or development, academic research, manufacturing, chemical, cosmetic, and pharmaceutical industry activity including sales, and any advertising or other promotional activity including written, oral, and electronic communication of any form, associated with the disclosure. Additionally, the kits may include other components depending on the specific application, as described herein.
[0173] The kits may have a variety of forms, such as a blister pouch, a shrinkwrapped pouch, a vacuum sealable pouch, a sealable thermoformed tray, or a similar pouch or tray form, with the accessories loosely packed within the pouch, one or more tubes, containers, a box, or a bag.
[0174] Without further elaboration, it is believed that one skilled in the art can, based on the above description, utilize the present disclosure to its fullest extent. The following specific embodiments are, therefore, to be construed as merely illustrative, and not limiting of the remainder of the disclosure in anyway whatsoever.DEFINITIONS
[0175] The following definitions are used in connection with the disclosure disclosed herein. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of skill in the art to which this disclosure belongs.
[0176] As used herein, “a,” “an,” or “the” can mean one or more than one.
[0177] Further, the term “about” when used in connection with a referenced numeric indication means the referenced numeric indication plus or minus up to 10% of that referenced numeric indication. For example, the language “about 50” covers the range of 45 to 55.
[0178] As referred to herein, all compositional percentages are by weight of the total composition, unless otherwise specified. As used herein, the word “include,” and its variants, is intended to be non-limiting, such that recitation of items in a list is not to the exclusion of other like items that may also be useful in the compositions and methods of this technology. Similarly, the terms “can” and “may” and their variants are intended to be non-limiting, such that recitation that an embodiment can or may comprise certain elements or features does not exclude other embodiments of the present technology that do not contain those elements or features.
[0179] Although the open-ended term “comprising,” as a synonym of terms such as including, containing, or having, is used herein to describe and claim the disclosure, the present disclosure, or embodiments thereof, may alternatively be described using alternative terms such as “consisting of” or “consisting essentially of.”
[0180] In embodiments, as used herein, the words “preferred” and “preferably” refer to embodiments of the technology that afford certain benefits, under certain circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful, and is not intended to exclude other embodiments from the scope of the technology.EXAMPLESExample 1: Mutant Bxb1 site-specific genomic integration compared against wild-type Bxb1
[0181] In this Example, Cas9-directed reverse transcription was used to insert an att site in host cells, as described in Hew, et al., “Directed evolution of hyperactive integrases for site specific insertion of transgenes,” Nucleic Acids Res. (2024) Vol. 52, No. 14: e64, the entire contents of which is incorporated by reference herein. A donor DNA (6.6 kbp in size) comprising a DNA sequence having a wild-type Bxb1 att was constructed. Insertion efficiency of the 6.6 kbp donor DNA was measured using droplet digital polymerase chain reaction (ddPCR) using probes that assesses the presence of the donor DNA at the locus and / or other locations in the genome. Such primers included a forward or reverse primer that is identical toor complementary with a genomic sequence at or near the target locus site and the other primer identical to or complementary with a sequence from the donor DNA. Insertion efficiency was compared between wildtype Bxb1 and select mutant Bxb1 integrases.
[0182] As shown in Fig. 4, representative efficiencies of Cas9-directed reverse transcription at the ROSA26 locus ranged from approximately 74% to 95% (with reverse transcriptase, PE5 or PE7), with integration efficiencies (Int) for each mutant outperforming the wild-type Bxb1. Cas9-directed att integration is shown in the left bar of each doublet, and Bxb1 -directed genomic integration is shown on the right bar. The inclusion of NLSs also improved both wild-type and mutant Bxb1 integration efficiency. Table 5 summarizes the efficiency statistics shown graphically in Fig. 1.Table 5: Mutant Bxb1 insertion efficiencies compared against wild-type Bxb1.
[0183] The data demonstrated, inter alia, that the mutant Bxb1 integrases exhibited an increased efficiency of site-specific insertion of the donor DNA at the target site compared to the wild type Bxb1 integrase. The data demonstrated, inter alia, that the mutant Bxb1 integrases exhibited an increased total amount of insertion of the donor DNA compared to the wild type Bxb1 integrase.Example 2: Mutant Bxb1 site-specific genomic integration at the TRAC locus
[0184] This Example demonstrates using a donor DNA containing for comparing insertion efficiency of wild-type Bxb1 integrase to mutant Bxb1 integrase derivatives using a sequence resembling the wild-type attB for integration at the TRAC locus. A donor DNA (6.6 kbp in size) comprising a DNA sequence having the wild-type Bxb1 attP was constructed. Host cells were transfected with the donor DNA and three “helper” plasmids each encoding a hyperactive integrase. The first helper plasmid encoded a Bxb1 having hyperactive mutations. The second helper plasmid encoded a Bxb1 having hyperactive mutations and additional mutations for binding the left side of the TRAC target. The third helper plasmid encoded a Bxb1 having hyperactive mutations and additional mutations for binding the right side of the TRAC target. These three helpers are co-transfected with the donor plasmid containing an att site with a matching dinucleotide core to the TRAC target.
[0185] Insertion efficiency at the TRAC locus was measured using droplet digital polymerase chain reaction (ddPCR) probes that assesses the presence of the donor DNA at the TRAC locus and / or other locations in the genome. Such primers included a forward or reverse primer that is identical to orcomplementary with a genomic sequence at or near the target TRAC locus site and the other primer identical to or complementary with a sequence from the donor DNA.
[0186] As shown in Fig. 2, several Bxb1 mutants exhibited increased site-specific insertion efficiency at the TRAC locus compared to wild-type Bxb1. Table 6 summarizes the efficiency statistics shown graphically in Fig. 2.Table 6: Bxb1 insertion efficiencies at TRAC locus.
[0187] Next, nuclear localization sequences (NLSs) were tested for the ability to further improve mutant Bxb1 efficiency. A second set of experiments were performed, as described previously, to compare the ratio of insertion relative to a cognate Bxb1 lacking the NLS.
[0188] As shown in Fig. 3, several Bxb1 mutants with a NLS exhibited increased site-specific insertion efficiency at the TRAC locus relative to the cognate mutant Bxb1 lacking the NLS. This was also observed with the wild-type Bxb1 (wt Bxb1 vs. wt Bxb1 + NLS). Table 7 summarizes the efficiency statistics shown graphically in Fig. 3.Table 7: Bxb1 insertion efficiencies at TRAC locus with NLS.
[0189] The data demonstrated, inter alia, that the mutant Bxb1 integrases exhibited an increased efficiency of site-specific insertion of the donor DNA at the target TRAC locus site(s) compared to the wild type Bxb1 integrase. The data demonstrated, inter alia, that the mutant Bxb1 integrases exhibited an increased total amount of insertion of the donor DNA compared to the wild type Bxb1 integrase. The data demonstrated, inter alia, that each of the mutant Bxb1 integrases exhibited an increased efficiency of recombination with the addition of a NLS sequence, and the same was observed for wild-type Bxb1.Example 3: Mutant Bxb1 site-specific genomic integration at ROSA26 locus
[0190] In this Example, Cas9-directed reverse transcription was used to insert an attB at the ROSA26 locus in HEK293T cells, as described in Hew, et al., ‘‘Directed evolution of hyperactive integrases for site specific insertion of transgenes,” Nucleic Acids Res. (2024) Vol. 52, No. 14: e64, the entire contents of which is incorporated by reference herein. Insertion efficiency at the ROSA26 locus of the 6.6 kbp donor DNA was measured using droplet digital polymerase chain reaction (ddPCR) using probes that assesses the presence of the donor DNA at the ROSA26 locus and / or other locations in the genome. Such primers included a forward or reverse primer that is identical to or complementary with a genomic sequence at or near the target ROSA26 locus site and the other primer identical to or complementary with a sequence from the donor DNA. Insertion efficiency was compared between wild-type Bxb1 and mutant Bxb1 , with and without nuclear localization sequences (NLSs).
[0191] As shown in Fig. 4, representative efficiencies of Cas9-directed reverse transcription at the ROSA26 locus ranged from approximately 74% to 95% (with reverse transcriptase, PE5 or PE7), with integration efficiencies (Int) for each mutant outperforming the wild-type Bxb1. Cas9-directed att integration is shown in the left bar of each doublet, and Bxb1 -directed genomic integration is shown on the right bar. The inclusion of NLSs also improved both wild-type and mutant Bxb1 integration efficiency. Table 8 summarizes the efficiency statistics shown graphically in Fig. 4.Table 8: Bxb1 insertion efficiencies at ROSA26 locus with and without NLS. att insertion is left bar of each doublet, Bxb1 -mediated integration is the right bar of each doublet.
[0192] As shown in Fig. 5, the integration efficiencies at ROSA26 using Cas9-directed reverse transcription with C-terminal NLSs were compared between select mutant Bxb1 recombinases and wild-type Bxb1. Each mutant outperformed the wild-type Bxb1. The C-terminal NLSs also improved the integration efficiency of wild-type Bxb1 and several Bxb1 mutants. Table 9 summarizes the efficiency statistics shown graphically in Fig. 5.Table 9: Bxb1 insertion efficiencies at ROSA26 locus with Cas9-directed reverse transcription with and without C-terminal NLS. Mutations are relative to SEQ ID NO: 1. npNLS = nucleoplasmin NLS KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75).
[0193] As shown in Fig. 6, the integration efficiencies at ROSA26 using Cas9-directed reverse transcription with NLSs were compared between select mutant Bxb1 recombinases and wild-type Bxb1 . Each mutant outperformed the wild-type Bxb 1 . The NLSs also improved the integration efficiency of wild-type Bxb 1 and several Bxb1 mutants. Table 10 summarizes the efficiency statistics shown graphically in Fig. 6.Table 10: Bxb1 insertion efficiencies at ROSA26 locus with Cas9-directed reverse transcription with and without NLS. Mutations are relative to SEQ ID NO: 1. bpNLS = bipartite NLS KRTADGSEFESPKKKRKV (SEQ ID NO: 74): npNLS = nucleoplasmin NLS KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75).
[0194] As shown in Fig. 7A, the integration efficiencies at ROSA26 using Cas9-directed reverse transcription with NLSs were compared between select mutant Bxb1 recombinases and wild-type Bxb1 . TheNLSs also improved the integration efficiency of wild-type Bxb1 and several Bxb1 mutants. Table 11 summarizes the efficiency statistics shown graphically in Fig. 7A.Table 11 : Bxb1 insertion efficiencies at ROSA26 locus with Cas9-directed reverse transcription with and without NLS. Mutations are relative to SEQ ID NO: 1. bpNLS = bipartite NLS KRTADGSEFESPKKKRKV (SEQ ID NO: 74): npNLS nucleoplasmin NLS KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75). HisNLS = GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0195] As shown in Fig. 7B, the integration efficiencies at ROSA26 using Cas9-directed reverse transcription with NLSs were compared between select mutant Bxb1 recombinases and wild-type Bxb1. The NLSs also improved the integration efficiency of wild-type Bxb1 and several Bxb1 mutants. Table 12 summarizes the efficiency statistics shown graphically in Fig. 7B.Table 12: Bxb1 insertion efficiencies at ROSA26 locus with Cas9-directed reverse transcription with and without C-terminal NLS. Mutations are relative to SEQ ID NO: 1. bpNLS = bipartite NLSKRTADGSEFESPKKKRKV (SEQ ID NO: 74): npNLS = nucleoplasmin NLSKRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75). HisNLS = GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0196] The data demonstrated, inter alia, that the mutant Bxb1 integrases exhibited an increased efficiency of site-specific insertion of the donor DNA at the target ROSA26 locus site(s) compared to the wild type Bxb1 integrase. The data demonstrated, inter alia, that the mutant Bxb1 integrases exhibited an increased total amount of insertion of the donor DNA compared to the wild type Bxb1 integrase. The data demonstrated, inter alia, that Bxb1 integrases (mutant or wild-type) exhibited an increased efficiency of recombination with the addition of a NLS sequence.Example 4: Mutant Bxb1 site-specific genomic integration into of Chimeric Antigen Receptor (CAR) at TRAC locus
[0197] In this Example, insertion an exogenous attB at the TRAC locus in HEK293T cells was not needed. Mutant Bxb1 were used to target a naturally-occurring genomic site at the TRAC locus. The data reflected in Figs. 2-3 and 8-9 show data for Bxb1-mediated integration at this naturally-occurring site. Figs.1 and 4-7 show data for Bxb1-mediated integration at an exogenous, inserted wild-type att site, e.g., using Cas9-directed reverse transcription-mediated insertion. The data herein show, inter alia, that mutant Bxb1 are useful to target native att sequences, as well as modified sequences.
[0198] HEK293T cells were plated in 24-well plates 1 day prior to transfection. The TRANSIT 2020 transfection reagent was used to transfect 600 ng of DNA. The cells were pelleted 3 days later and lysed for digital droplet PCR (ddPCR) analysis. ddPCR forward primer and probe were designed to bind in the genome at the TRAC locus and the ddPCR reverse primer extended out from the donor DNA, e.g., as shown in Table 13.Table 13. Exemplary primer and probe sequences for TRAC locus insertion.
[0199] For each transfection, four different plasmids were used: (1) a helper plasmid, (2) a donor plasmid, (3) a TRAC-L plasmid, and (4) a TRAC-R plasmid, which collectively encode 3 Bxb1 enzymes - one Bxb1 targets the Left side of the genome-located site that resembles the attB site, one Bxb1 targets the Right side of the genome-located site that resembles the attB site, and a third Bxb1 binds the attP on the donor DNA, termed the “helper.” 150 ng of donor plasmid containing the wildtype Bxb1 attP site and the donor DNA payload to be integrated into the host cell genome was mixed with an expression plasmid encoding a mutant Bxb1 integrase. The donor plasmid encoded an Ef1 a promoter driving an anti-CD19 Chimeric Antigen Receptor (CAR) and a red fluorescent protein (DsRed). Three integrase plasmids were included in each transfection: 1) 360 ng of a ‘helper’ plasmid that encoded a Bxb1 having one or more hyperactive mutations but contained wildtype regions of the protein that bind the attP site, 2) 45 ng of a TRAC-L integrase expressing plasmid, and 3) 45 ng of a TRAC-R integrase expression plasmid. The helper plasmid is what binds the donor plasmid in contact with the wildtype attP sequence. These plasmids each encoded both hyperactive Bxb1 mutants that generally improve efficiency as well as mutations that aid in the binding to either the left (TRAC-L) or right (TRAC-R) of the endogenous TRAC sequence that resembles an attB site that is found in the human genome. Upon binding of the helper plasmid to the attP on the donor plasmid, the TRAC-L binding to the left side of the target sequence and the TRAC-R to the right side of the target sequence, the multi-subunit complex recombines the donor DNA and inserts it into the TRAC locus. ddPCR was used to amplify the junction where the donor is inserted.
[0200] The frequency of insertions were calculated as a ratio to the frequency of a housekeeping gene called UBE2D2 which is found at approximately 3 copies per cell in HEK293T cells. Nuclear localization signals (bpNLS, npNLS, HisNLS) were fused to the C-terminus of the integrase to promote import into the nucleus and increase targeting efficiency to TRAC. Additionally, stability tags (e.g., such as C-Stab) were fused to the C-terminus to stabilize the protein and improve targeting efficiency.Table 14: Experimental data for Fig. 8.Table 15: Experimental data for Fig. 9.
[0201] As shown in Figs. 8 and 9, with data statistics in Tables 14 and 15, respectively, each of the mutants tested (e.g., nuclear localization sequence (NLS), stabilization tag, and / or mutant) outperformed wild-type Bxb1 recombination efficiency. The presence of a NLS on its own (e.g., bpNLS, npNLS, or HisNLS) improved Bxb1 recombination efficiency relative to wild-type Bxb1. Each of the mutants tested (e.g., R63K, V76I, I87L, A119S, V122M, A369P, and E434G), including combinations therebetween improved the recombination efficiency relative to wild-type Bxb1. The presence of the stabilization tag (e.g., C-stab) improved Bxb1 recombination efficiency relative to wild-type Bxb1. The presence of the stability tag improved Bxb1 recombination efficiency in mutant Bxb1 relative to NLS alone.
[0202] Additional testing with stability tags was performed using Cas9 reverse transcriptase insertion of wild-type aft at the ROSA26 locus in HEK293 cells, e.g., as described in the Examples herein. Table 17 provides DNA and amino acid sequences of stability tags (SEQ ID NOs: 122 and 158-191 for nucleic acid sequences, SEQ ID NOs: 157 and 162-165 for amino acid sequences). Table 16 provides data statistics for the mutants tested. As shown in Figs. 10A-10B, each of the Bxb1 mutants tested (e.g., nuclear localization sequence (NLS), stabilization tag, and / or mutant) outperformed wild-type Bxb1 recombination efficiency.Table 15: Experimental data for Fig. 10A. N-term = stability tag and NLS attached at the N-terminus of the integrase: C-term = stability tag and NLS attached at the C-terminus of the integrase: no = npNLS, KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75): bp = bpNLS, KRTADGSEFESPKKKRKV (SEQ ID NO: 74): C-stab - stability tag attached at the C-terminus of the integrase without a NLS: N-stab = stability tag attached at the N-terminus of the integrase without a NLS.Table 16: Experimental data for 10B. N-term - stability tag and NLS attached at the N-terminus of theintegrase; C-term = stability tag and NLS attached at the C-terminus of the integrase; np = npNLS, KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75); bp = bpNLS, KRTADGSEFESPKKKRKV (SEQ ID NO: 74); C-stab = stability tag attached at the C-terminus of the integrase without a NLS; N-stab = stability tag attached at the N-terminus of the integrase without a NLS.
[0203] In embodiments, Table 17 below provides exemplary, non-limiting, DNA expression sequences and amino acid sequence of mutant Bxb1 mutants, locus-specific Bxb1 mutants, nuclear localization sequences, and stabilization tags. Persons skilled in the art, with the benefit of this disclosure in its entirety, will understand that changes to the mutations, nuclear localization sequences, and stabilization sequences, as well as combinations between each e.g., among elements as described herein) are expected to function similarly e.g., with improvements in functionality relative to wild-type Bxb1), expected to function at a variety of genomic sites, and expected to function across different cell types. In embodiments, DNA sequences are adjustable, e.g., based on codon optimization for the desired host cell.
[0204] Collectively, these data demonstrated, inter alia, that non-naturally occurring Bxb1 mutations described herein (and mutation combinations thereof), including the presence of nuclear localization sequences and / or stability tags, are useful to improve Bxb1 genomic recombination efficiency of large insertions (>6 kbp) in comparison to wild-type Bxb1. The data demonstrated successful Bxb1 -mediated genomic insertion, both using inserting at exogenous att sites inserted using Cas9 RT-mediated insertion, and insertion by targeting endogenous sequences resembling att sites. Successful integrase-mediated insertion was shown at two different cellular locis. Both wild-type Bxb1 and mutant Bxb1 are amenable to targeting att sites that differ from wild-type sequences.Table 17: Exemplary Sequences for Bxb1 mutants, NLSs, and stabilization tag. SEQ ID NO: 88 = nucleic acid sequences encoding wt Bxb1, SEQ ID NO: 1 = amino acid sequences encoding wt Bxb1, SEQ ID NOs: 89-118 = nucleic acid sequences encoding Bxb1 mutants, SEQ ID NOs: 124-153 = amino acid seguences encoding Bxb1 mutants, SEQ ID NOs: 119-121 = nucleic acid sequences encoding NLSs, SEQ ID NOs: 154- 156 = amino acid sequences encoding NLSs, SEQ ID NOs: 122 and 158-161 = nucleic acid seguences encoding stability tags, SEQ ID NOs: 157 and 162-165 = amino acid seguences encoding stability tags.DEFINITIONS
[0205] The following definitions are used in connection with the disclosure disclosed herein. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of skill in the art to which this disclosure belongs.
[0206] As used herein, “a,” “an,” or “the” can mean one or more than one.
[0207] Further, the term “about” when used in connection with a referenced numeric indication means the referenced numeric indication plus or minus up to and including 10% of that referenced numeric indication. For example, the language “about 50” covers the range of 45 to 55.
[0208] In embodiments, the terms “at least one” and “one or more” are used interchangeably.
[0209] As used herein, the word “include,” and its variants, is intended to be non-limiting, such that recitation of items in a list is not to the exclusion of other like items that may also be useful in the enzymes, nucleic acids, compositions, and methods of the technologies herein. Similarly, the terms “can” and “may” and their variants are intended to be non-limiting, such that recitation that an embodiment can or may comprise certain elements or features does not exclude other embodiments of the present technology that do not contain those elements or features.
[0210] Although the open-ended term “comprising,” as a synonym of terms such as including, containing, or having, is used herein to describe and claim the disclosure, the present disclosure, or embodiments thereof, may alternatively be described using alternative terms such as “consisting of” or “consisting essentially of.”
[0211] In embodiments, reference to amino acid sequences or stretches of amino acids are inclusive, meaning the range includes the amino acids identified on the ends of the range. For example, “one or more of residues A157-P160, ” includes A157 and P160.
[0212] In embodiments, as used herein, the words “preferred” and “preferably” refer to embodiments of the technology that afford certain benefits, under certain circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful, and is not intended to exclude other embodiments from the scope of the technology.
[0213] In embodiments, the disclosure is directed to the following embodiments:
[0214] Embodiment 1. A Bxb1 integrase comprising: (i) an amino acid sequence that is at least 90% identical to amino acids 1-480 of SEQ ID NO: 1, (ii) an amino acid sequence that is at least 90% identical to amino acids 1-488 of SEQ ID NO: 1, (iii) or an amino acid sequence that is at least 90% identical to SEQ ID NO: 1, and comprises one or more amino acid mutations at a position selected from 4, 5, 14, 18, 20, 24, 29, 34, 35, 36, 40, 42, 45, 46, 49, 50, 51, 60, 61, 62, 63, 67, 68, 69, 70, 73, 74, 75, 76, 78, 79, 84, 85, 86, 87, 88, 89, 90, 92, 95, 99, 100, 105, 106, 110, 111, 116, 119, 122, 124, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174, 175, 178, 179, 181, 183, 187, 189, 191, 197, 203, 207, 208, 209, 218, 223, 229, 231,232, 234, 236, 237 239, 248, 251, 254, 257, 261, 264, 267, 268, 272, 273, 278, 280, 281, 282, 283, 285,287, 288, 291, 292, 295, 302, 306, 307, 311, 313, 314, 316, 318, 319, 321, 322, 323, 325, 328, 331, 332,333, 334, 342, 343, 347, 353, 355, 359, 360, 361, 362, 368, 369, 370, 375, 380, 388, 397, 398, 405, 409,411, 414, 415, 416, 419, 425, 428, 434, 435, 444, 449, 453, 461, 462, 463, 466, 468, 476, 479, 480, 483,484, 487, 488, 489, 494, 496, and 499, or a position corresponding thereto; or wherein (i), (ii), or (iii) further comprises at least two amino acid substitutions at positions selected from (a) and (b), (b) and (c), (a) and (c):(a) 75, 76, 158, 232, 234, 236, 237, 257, 314, 316, 318, 322, 323, 325, or a position corresponding thereto,(b) 5, 14, 20, 24, 29, 35, 40, 45, 49, 50, 51, 60, 68, 69, 70, 73, 74, 78, 84, 86, 87, 100, 105, 116, 124, 183,197, 207, 208, 209, 229, 261, 267, 273, 273, 287, 291, 333, 342, 343, 347, 361, 368, 375, 435, 449, 453, 462, 483, 494, or a position corresponding thereto, and / or (c) 4, 18, 34, 36, 42, 46, 61, 62, 63, 67, 79, 85, 88, 89, 90, 92, 95, 99, 106, 110, 111, 119, 122, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174, 175, 178, 179, 181, 187, 189, 191, 203, 218, 223, 231, 239, 248, 251, 254, 264, 268, 272, 278, 280, 281,282, 283, 285, 288, 292, 295, 302, 306, 307, 311, 313, 319, 321, 328, 331, 332, 334, 353, 355, 359, 360,362, 369, 370, 380, 388, 397, 398, 405, 409, 411, 414, 415, 416, 419, 425, 428, 434, 444, 461, 463, 466,468, 476, 479, 480, 484, 487, 488, 489, 496, and / or 499, or a position corresponding thereto.
[0215] Embodiment 2. The Bxb1 integrase of embodiment 1, wherein the one or more amino acid mutations are selected from L4I, V5I, D14N, E20K, E20Q, E24K, L29F, G34D, W35L, D36A, V40I, V40A, E42K, D45G, A49T, V50I, D51N, D51E, D51Y, N60S, L61F, R63K, F67S, E68K, E69A, E69D, Q70P, D73G, V74A, V74I, V74M, I75V, V76I, Y78H, Y78N, T84S, S86T, I87V, I87L, H89G, Q92H, H95Y, D99N, H100N, H100Y, V105A, V105I, H111P, T116P, A119S, V122M, A124S, A145T, S157G, T166I, V179A, V179I, R181 K, E183L, V187I, H189N, P197T, H203Y, R207Q, R208S, G209V, E229K, A232G, A232S, A232R,A232T, A232V, A232P, A232Q, T233G, T233W, T233R, T233Y, T233D, T233G, T233N, T233H, T233S,T233Q, T233A, T233C, A234N, A234G, A234S, A234T, A234H, A234F, A234S, K236R, K236S, R237K,R237Q, R237N, R237V, R237C, M239I, A248T, N251 K, T254S, D257K, A261T, A261V, V264A, E267D,R272Q, E273D, E273K, A280T, T285A, R287P, A288V, A288T, A291T, A311V, K313R, F314M, F314L, F314N, F314R, F314K, G316R, G316W, G318H, G318P, G318S, G318N, G318K, R319K, R319G, H321P, H321Q, H321 K, H321L, H321R, H321Y, H321T, H321S, P322A, P322R, P322G, P322L, R323L, R323G, R323Y, R323I, R325K, R325Q, R325Y, F331S, P332H, K333N, H334R, M342T, M342V, A343T, A347E, A347V, V353I, D355N, D359N, A360T, E361D, R362K, V368A, A369P, A369E, A369T, A369S, V375I, V380I, T388M, A396P, A398S, R409H, A411V, A414V, A415S, R416K, A425T, E434G, T435A, R444L, A449V, T453I, T453A, L462M, T463I, V466M, G468D, L479I, Q480STQP, E483K, Q484K, R487R, del L488 (del L488 frameshift), R494S, R494Q, H496N, and / or M499T, or a position corresponding thereto, relative to SEQ ID NO: 1.
[0216] Embodiments. The Bxb1 integrase of embodiment 1 or 2, the one or more of the following amino acids are not mutated relative to SEQ ID NO: 1: V5, S18, E24, V40, V46, D51, A62, E69, R79, R85, R88, L90, Q92, S106, A110, A130, E133, 1137, R140, K153, G156, P160, L164, L174, V175, P178, V187, Q191 , G209, A218, R223, S231, N251, P268, L278, A280, E281, L282, V283, R287, P292, P295, L302, V306, C307, K313, S328, K333, D359, E361, G370, V375, R397, A405, A414, E419, A425, S428, T435, R461, V466, G468, F476, E483, G489, and R494.
[0217] Embodiment 4. The Bxb1 integrase of any one of claims 1-3, wherein the Bxb1 integrase comprises a C-terminal deletion of one or more amino acids in the range of L488-S500 relative to SEQ ID NO: 1.
[0218] Embodiment 5. The Bxb1 integrase of embodiment 4, wherein the C-terminal deletion comprises a deletion of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, or all 13 amino acids of L488-S500 relative to SEQ ID NO: 1
[0219] Embodiment 6. The Bxb1 integrase of embodiment 4 or 5, wherein L488 is mutated to a stop codon (L488STOP).
[0220] Embodiment 7. The Bxb1 integrase of any one of any one of the preceding embodiments, wherein the Bxb1 integrase comprises a C-terminal deletion of one or more amino acids in the range of Q480- S500 relative to SEQ ID NO: 1.
[0221] Embodiments. The Bxb1 integrase of embodiment 7, wherein the C-terminal deletion comprises a deletion of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, 13 amino acids, 14 amino acids, 15 amino acids, 16 amino acids, 17 amino acids, 18 amino acids, 19 amino acids, 20 amino acids, or all 21 amino acids of Q480-S500 relative to SEQ ID NO: 1
[0222] Embodiment 9. The Bxb1 integrase of any one of embodiments claims 1-3, wherein Q480 is mutated to a stop codon (Q480STOP).
[0223] Embodiment 10. The Bxb1 integrase of any one of the preceding embodiments, wherein the one or more amino acid mutations comprises mutation of one or more of residues A157-P160; and / or wherein the mutation alters specificity at -7 and / or -6 positions of a loop target DNA site compared to a wild-type Bxb 1 of SEQ ID NO: 1.
[0224] Embodiment 11 . The Bxb1 integrase of any one of the preceding embodiments, wherein the one or more amino acid mutations comprises mutation of one or more of residues Y154-P159; and / or wherein the one or more mutations alters specificity at -7 and / or -6 positions of a loop target DNA site compared to a wild-type Bxb1 of SEQ ID NO: 1.
[0225] Embodiment 12. The Bxb1 integrase of embodiment 10 or 11 , wherein the mutation comprises S157G.
[0226] Embodiment 13. The Bxb1 integrase of embodiment 11 , wherein the one or more Bxb1 integrases comprise a sequence of residues Y154-P159 selected from YRGSLS (SEQ ID NO: 16), WWGGTP (SEQ ID NO: 17), YRGGLP (SEQ ID NO: 18), AHPGRT (SEQ ID NO: 19), and WRGGAS (SEQ ID NO: 20) relative to SEQ ID NO: 1 , or a position corresponding thereto.
[0227] Embodiment 14. The Bxb1 integrase of any one of embodiments 10-13, wherein the one or more amino acid mutations at this position alters the dinucleotide preference of adenine (A) or thymine (T) at position -7 and / or cytosine (C) or guanine (G) at position -6.
[0228] Embodiment 15. The Bxb 1 integrase of any one of the preceding embodiments, wherein the one or more amino acid mutations comprises mutation of one or more of residues S231-R237; and / or wherein the mutation alters specificity at -11 , -10, and / or -9 positions of a 3 bp helix target site compared to a wildtype Bxb1 of SEQ ID NO: 1.
[0229] Embodiment 16. The Bxb1 integrase of any one of the preceding embodiments, wherein L235 is not mutated.
[0230] Embodiment 17. The Bxb1 integrase of embodiment 15 or 16, wherein the one or more amino acid mutations comprises mutation of one or more residues of 231-234 and / or 236-237
[0231] Embodiment 18. The Bxb1 integrase of any one of embodiments 15-17, wherein A234 is mutated to an asparagine (A234N); and / or wherein this mutation alters specificity at the -10 position to thymine (T) compared to a wild-type Bxb1 of SEQ ID NO: 1 .
[0232] Embodiment 19. The Bxb1 integrase of any one of embodiments 15-18, wherein the one or more amino acid mutations comprises mutation of one or more of residues 231 , 233, and 237; and / or wherein the mutation alters specificity at -10 and / or -9 positions of a 3 bp helix target site compared to a wild-type Bxb1 of SEQ ID NO: 1.
[0233] Embodiment 20. The Bxb1 integrase of any one of embodiments 15-19, wherein the Bxb1 integrase comprises 1 , 2, 3, 4, or 5 amino acids mutated between residues S231-R237 (inclusive) relative to SEQ ID NO: 1.
[0234] Embodiment 21. The Bxb1 integrase of any one of embodiments 15-20, wherein the Bxb1 integrase comprises a sequence from residues S231-R237 selected from AGGNLKR (SEQ ID NO: 21), LGTNLKR (SEQ ID NO: 22), SGTGLKK (SEQ ID NO: 23), SGSALKT (SEQ ID NO: 24), AAWALRR (SEQ ID NO: 25), GGRSLKR (SEQ ID NO: 26), SGYNLRR (SEQ ID NO: 27), SGWGLKK (SEQ ID NO: 28), SGWALRQ (SEQ ID NO: 29), SARALSR (SEQ ID NO: 30), RADTLRR (SEQ ID NO: 31), YSRNLKR (SEQ ID NO: 32), SRNGLRK (SEQ ID NO: 33), RGHALKN (SEQ ID NO: 34), GGSHLKR (SEQ ID NO: 35), TTRTLKR (SEQ ID NO: 36), AVQNLKR (SEQ ID NO: 37), RAAFLKK (SEQ ID NO: 38), RAWTLKC (SEQ ID NO: 39), RAWSLKR (SEQ ID NO: 40), HGWSLKV (SEQ ID NO: 41), HGCTLKR (SEQ ID NO: 42), YGSALKQ (SEQ ID NO: 43), SQWALKC (SEQ ID NO: 44), YPWSLRR (SEQ ID NO: 45), relative to SEQ ID NO: 1 , or a position corresponding thereto.
[0235] Embodiment 22. The Bxb1 integrase of any one of the preceding embodiments, wherein the Bxb1 integrases comprises a mutation of D257K relative to SEQ ID NO: 1 .
[0236] Embodiment 23. The Bxb 1 integrase of any one of the preceding embodiments, wherein the one or more amino acid mutations comprises mutation of one or more of residues K313-S328; and / or wherein the mutation alters specificity at alter specificity at -19 through -12 positions of a DNA binding site compared to a wild-type Bxb1 of SEQ ID NO: 1.
[0237] Embodiment 24. The Bxb1 integrase of any one of the preceding embodiments, wherein the Bxb1 residue P322 is maintained as a proline.
[0238] Embodiment 25. The Bxb1 integrase of embodiment 23 or 24, wherein the one or more amino acid mutations comprises mutation of one or more residues of K313, F314, A315, G316, G318, R319, H321 , P322, R323, R325, and S328.
[0239] Embodiment 26. The Bxb1 integrase of embodiment 25, wherein the one or more amino acid mutations comprises mutation of one or more residues of K313, R319, H321 , and / or S328.
[0240] Embodiment 27. The Bxb1 integrase of any one of the preceding embodiment, wherein the Bxb1 integrase comprises 1 , 2, 3, 4, 5, 6, or 7 amino acids mutated between residues F314-R325 (inclusive) relative to SEQ ID NO: 1.
[0241] Embodiment 28. The Bxb1 integrase of embodiment 27, wherein the Bxb1 integrase comprises a sequence from residues K314-R325 selected from MAGGHRKQALYR (SEQ ID NO: 46), MAGGPRKKRRYR (SEQ ID NO: 47), LARGSRKLALYR (SEQ ID NO: 48), NARGNRKRGRYR (SEQ ID NO: 49), LARGPRKRAGYK (SEQ ID NO: 50), RAWGKRKYAYYQ (SEQ ID NO: 51), KAWGSRKTRLYR (SEQ ID NO: 52), MARGGRKSAIYY (SEQ ID NO: 53), MASGSRKTAIYY (SEQ ID NO: 54), LARGRRKWARYR (SEQ ID NO: 55), and LARGSRKLALYR (SEQ ID NO: 56), realtive to SEQ ID NO: 1.
[0242] Embodiment 29. The Bxb1 integrase of any one of the preceding embodiments, wherein the Bxb1 integrase comprises one or more mutations and / or deletions, and / or one or more combinations of mutations and / or deletions as described in Table 1 and / or Table 2 relative to SEQ ID NO: 1.
[0243] Embodiment 30. The Bxb1 integrase of any one of the preceding embodiments, wherein the Bxb1 integrase comprises at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1.
[0244] Embodiment 31. The Bxb1 integrase of any one of claims 1-30, wherein the Bxb1 integrase comprises at least 90% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations.
[0245] Embodiment 32. The Bxb1 integrase of any one of embodiments 1-30, wherein the Bxb1 integrase comprises at least 95% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, or at least 20 mutations.
[0246] Embodiment 33. The Bxb1 integrase of any one of embodiments 1-30, wherein the Bxb1 integrase comprises at least 96% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, or 20 mutations.
[0247] Embodiment 34. The Bxb1 integrase of any one of embodiments 1-30, wherein the Bxb1 integrase comprises at least 97% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, or 15 mutations.
[0248] Embodiment 35. The Bxb1 integrase of any one of embodiments 1-30, wherein the Bxb1 integrase comprises at least 98% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, or 10 mutations.
[0249] Embodiment 36. The Bxb1 integrase of any one of embodiments 1-30, wherein the Bxb1 integrase comprises at least 99% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, or 5 mutations.
[0250] Embodiment 37. The Bxb1 integrase of embodiment 1 , wherein the Bxb1 integrase comprises: an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO: 1, and further comprising one or more amino acid mutations at positions selected from L4, E42, R57, R63, V76, H111 , V122, V187, K313, R319, A341, D359, A369, A398, R416, A425, L479, and M499, relative to SEQ ID NO: 1.
[0251] Embodiment 38. The Bxb1 integrase of embodiment 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of SEQ ID NO: 1 .
[0252] Embodiment 39. The Bxb 1 integrase of embodiment 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of SEQ ID NO: 1
[0253] Embodiment 40. The Bxb1 integrase of embodiment 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 .
[0254] Embodiment 41. The Bxb1 integrase of any one of embodiments 37-40, wherein the Bxb1 integrase comprises one or more amino acid mutations are selected from L4I, E42K, R57K, R63K, V76I, H111 P, V122M, V187, K313R, R319K, A341 D, D359A, A369P, A398S, R416K, A425T, L479I, and M499T relative to SEQ ID NO: 1.
[0255] Embodiment 42. The Bxb1 integrase of any one of embodiments 37-41 , wherein the Bxb1 integrase further comprises one or more amino acid mutations selected from D36N, V40A, A49S, V74I, I87L or I87A or I87S, H89G, V175A, V179A, R287H, A288V, T453I, and H496N relative to SEQ ID NO: 1.
[0256] Embodiment 43. The Bxb1 integrase of any one of embodiments 37-42, wherein the Bxb1 integrase further comprises one or more amino acid mutations selected from V40I, D45G, I75V, H95Y, A119S, A280T, A311 V, E434G, and V466M relative to SEQ ID NO: 1
[0257] Embodiment 44. The Bxb1 integrase of any one of embodiments 37-43, wherein one or more of the following amino acids relative to SEQ ID NO: 1 are not mutated: A62, H100, Q191 , P195, G209, P295, L302, C307, L387, E419, S428, E483, and R487 relative to SEQ ID NO: 1.
[0258] Embodiment 45. The Bxb1 integrase of any one of embodiments 37-44, wherein the one or more amino acid mutations are selected from mutations at position V122 and / or position A369 relative to SEQ ID NO: 1.
[0259] Embodiment 46. The Bxb1 integrase of embodiment 45, wherein the mutation at position V122 relative to SEQ ID NO: 1 is V122M.
[0260] Embodiment 47. The Bxb 1 integrase of embodiment 45, wherein the mutation at position A369 relative to SEQ ID NO: 1 is A369P.
[0261] Embodiment 48. The Bxb1 integrase of any one of embodiments 37-47, wherein the one or more amino acid mutations relative to SEQ ID NO: 1 are selected from a combination of positions selected from: V76 and V122; V76 and A369; I87 and V122; I87 and A369; H95 and V122; H95 and A369; V122 and E434; A369 and E434; V76, V122, and A369; I87, H95 and A369; I87, V122 and A369; I87, V122 and E434; I87, A369 and E434; H95, V122 and A369; H95, V122 and E434; H95, A369 and E434; V122, A369 and E434; I87, H95, V122 and E434; I87, H95, A369 and E434; and I87, V122, A369 and E434.
[0262] Embodiment 49. The Bxb1 integrase of embodiment 48, wherein the one or more mutations relative to SEQ ID NO: 1 are a combination selected from: V76I and V122M; V76I and A369P; I87L and V122M; I87L and A369P; H95Y and V122; H95Y and A369P; V122M and E434G; A369P and E434G; V76I, V122M, and A369P; I87L, H95Y and A369P; I87L, V122M and A369P; I87L, V122M and E434G; I87L, A369P and E434G; H95, V122M and A369P; H95Y, V122M and E434G; H95Y, A369P and E434G; V122, A369P and E434G; I87L, H95Y, V122M and E434G; I87L, H95Y, A369P and E434G; and I87L, V122, A369P and E434G.
[0263] Embodiment 50. The Bxb1 integrase of any one of embodiments 37-47, wherein the one or more amino acid mutations relative to SEQ ID NO: 1 are a combination of positions selected from: I87, A369, and E434; V122, A369, and E434; H95, A369, and E434; V122 and E434; V122, A369, E434, and V175; V76 and A119; A119, A369, and E434; I87, V175, and D359; I87, R57, and R63; I87 and A369; I87 and A425; R63 and V122; R63, V122, R319, and V40I; R63, A119, and M499; V122 and A369; H95 and A369; A49, I87, A396, and E434; R57, R63, 187, A396, and E434; D36, V122, A396, and E434; R57, R63 and H95; D36, I87, A396, and E434; 187, A341 , A396, and E434; R57, R63, V122, A396, and E434; V122, R287, A396, and E434; I87, R287, A396, and E434; I87 and A341 ; I87, R287, A396, E434, and A341 ; V122, A396, E434, and A341 ; H95, and A341 ; A49, V122, A396, and E434; D45, A396, and E434; V122, R287, A341 , A396, and E434; A49, A396, and E434; I87, A341 , and R287; D36, A396, and E434; V122, D359, A369, and E434; R63 and I87; I87, D359, A369, and E434; V122, A369, A425, and E434; I87, A369, A425, and E434; V40, H95, A369, and E434; L4, H95, A369, and E434; H95, A369, E434, and V466; R63, V122, A369, and E434; V76, I87, and K313; R63, 187, and N194; I87 and D359; L4 and I87; V74 and V76; V122, A369, E434, and R319; I87, A369, E434, and R319; L4, V122, A369, and E434; I87 and A425; E42, V76, and I87; I87 and L479; H95, A369, A425, and E434; R63, 187, A369, and E434; I87 and R319; V76, A369, and E434; H95, A369,E434, and L479; V74, I75, and V76; H95, R319, A369, and E434; L4, I87, A369, and E434; V122, V175, D359, A369, and E434; R57, 187, A369, and E434; E42, H95, A369, and E434; H95, V179, A369, and E434; I75 and V76; L4 and V76; H95 and K313; V40, R63, and H89; A288, A311, A398, R416, T453, H496, and E42; and H111 and E434.
[0264] Embodiment 51. The Bxb1 integrase of embodiment 50, wherein the one or more mutations relative to SEQ ID NO: 1 are a combination selected from: I87L, A369P, and E434G; V122M, A369P, and E434G; H95Y, A369P, and E434G; I87L and E434G; V122M and E434G; V122M, A369P, E434G, and V175A; V76I and A119S; D36N and A119S; A119S, A369P, and E434G; I87L, V175A, and D359A; I87L, R57K, and R63K; I87L and D36N; I87L and A369P; I87L and A425T; R63K and V122M; R63K, V122M, R319K, and V40I; R63K, A119S, and M499T; H95Y and E434G; V122M and A369P; H95Y and V466M; H95Y and A369P; A49S, I87L, A396P, and E434G; A49S and I87L; R57K, R63K, I87L, A396P, and E434G; D36N, V122M, A396P, and E434G; R57K, R63K and H95Y; D36N, I87L, A396P, and E434G; D36N and H95Y; I87L, A341 D, A396P, and E434G; A49S and H95Y; R57K, R63K, V122M, A396P, and E434G; V122M, R287H, A396P, and E434G; I87L and R287H; I87L, R287H, A396P, and E434G; I87L and A341D; I87L, R287H, A396P, E434G, and A341 D; V122M, A396P, E434G, and A341 D; H95Y and A341 D; A49S, V122M, A396P, and E434G; D45G, A396P, and E434G; V122M, R287H, A341D, A396P, and E434G; A49S, A396P, and E434G; I87L, A341D, and R287H; H95Y and R287H; D36N, A396P, and E434G; V122M, D359A, A369P, and E434G; R63K and I87L; I87L and V466M; I87L, D359A, A369P, and E434G; V122M, A369P, A425T, and E434G; I87L, A369P, A425T, and E434G; V40I, H95Y, A369P, and E434G; L4I, H95Y, A369P, and E434G; H95Y, A369P, E434G, and V466M; R63K, V122M, A369P, and E434G; V76I, I87L, and K313R; R63K, I87V, and N194D; I87L and D359A; I87L and L4I; V74I and V76I; V122M, A369P, E434G, and R319K; I87L, A369P, E434G, and R319K; L4I, V122M, A369P, and E434G; I87L and A425T; E42K, V76I, and I87L; I87L and L479I; H95Y, A369P, A425T, and E434G; V40I and I87L; R63K, I87L, A369P, and E434G; I87L and R319K; V76I, A369P, and E434G; H95Y, A369P, E434G, and L479I; V74I, I75V, and V76I; H95Y, R319K, A369P, and E434G; L4I, I87L, A369P, and E434G; V122M, V175A, D359A, A369P, and E434G; R57K, I87L, A369P, and E434G; E42K, H95Y, A369P, and E434G; H95Y, V179A, A369P, and E434G; I75V and V76I; I87L and V179A; L4I and V76I; H95Y and K313R; H95Y and A280T; V40A, R63K, and H89G; A288V, A311V, A398S, R416K, T453I, H496N, and E42K; and H111P and E434G.
[0265] Embodiment 52. The Bxb1 integrase of embodiment 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations I87L, A369P, and E434G relative to SEQ ID NO: 1.
[0266] Embodiment 53. The Bxb 1 integrase of embodiment 37, wherein the Bxb 1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations V122M, A369P, and E434G relative to SEQ ID NO: 1.
[0267] Embodiment 54. The Bxb 1 integrase of embodiment 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations H95Y, A369P, and E434G relative to SEQ ID NO: 1; or wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations V74A, E229K, and V375I relative to SEQ ID NO: 1.
[0268] Embodiment 55. The Bxb1 integrase of any one of the preceding embodiments, wherein the Bxb1 integrase comprises one or more nuclear localization sequences.
[0269] Embodiment 56. The Bxb1 integrase of embodiment 55, wherein the one or more nuclear localization sequences is fused at the N-terminus, at the C-terminus, at both the N-terminus and C-terminus, or within the integrase sequence.
[0270] Embodiment 57. The Bxb1 integrase of embodiment 56 or 57, wherein the one or more nuclear localization sequences comprises a sequence selected from: PKKKRKV (SEQ ID NO: 57), NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 58); PAAKRVKLD (SEQ ID NO: 59), RQRRNELKRSF (SEQ ID NO: 60); NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 61), RMRKFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 62), VSRKRPRP (SEQ ID NO: 63), PPKKARED (SEQ ID NO: 64), PQPKKKPL (SEQ ID NO: 65), SALIKKKKKMAP (SEQ ID NO: 66), DRLRR (SEQ ID NO: 67), PKQKKRK (SEQ ID NO: 68), RKLKKKIKKL (SEQ ID NO: 69), REKKKFLKRR (SEQ ID NO: 70), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 71), RKCLQAGMNLEARKTKK (SEQ ID NO: 72), DPKKKRKVDPKKKRKVDPKKKRKV (SEQ ID NO: 73), KRTADGSEFESPKKKRKV (SEQ ID NO: 74), KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75), GSHHHHHHGSGPKKKRKV (SEQ ID NO: 76), and GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0271] Embodiment 58. The Bxb1 integrase of embodiments 55-57, wherein the one or more nuclear localization sequences is or comprises KRTADGSEFESPKKKRKV (SEQ ID NO: 74), KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75), or GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0272] Embodiment 59. The Bxb1 integrase of any one of the preceding embodiments, wherein the Bxb1 integrase comprises one or more stabilization sequences, optionally located at the N-terminus, C- terminus, or both the N-terminus and C-terminus.
[0273] Embodiment 60. The Bxb1 integrase of embodiment 59, wherein to the one or more stabilization sequences comprises an amino acid sequence that influences stability, solubility, folding, aggregation, degradation, purification, isolation, tracking expression and / or subcellular localization, relative to a cognate protein lacking the one or more stabilization sequences.
[0274] Embodiment 61. The Bxb1 integrase of claim 59 or 60, wherein the one or more stabilization sequences is or comprises about 5 amino acids to about 300 amino acids in length
[0275] Embodiment 62. The Bxb1 integrase of any one of claims 59-61 , wherein the one or more stabilization sequences is or comprises the amino acid sequence of any one of SEQ ID NOs: 157 and 162- 165, or an amino acid sequence having 1 , 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes thereto.
[0276] Embodiment 63. The Bxb1 integrase of any one of embodiments 59-61 , wherein the one or more stabilization sequences is or comprises a poly-amino acid sequence, optionally selected from poly-Cys, polyArg, poly-Asp, poly-Phe, and poly-His; wherein the one or more stabilization sequences is or comprises one or more of a AU1 tag, AU5 tag, Glu-Glu tag EYMPME, FLAG tag, strep-tag II, HA tag, c-Myc tag, T7 tag, VSV-G tag, KT3 epitope tag, HSV tag, Protein C tag, V5 tag, S-tag, twin-strep tag, SBP tag, CBD tag, and a protein eXact tag; or wherein the one or more stabilization sequences is or comprises a protein domain, or a portion thereof, optionally selected from ubiquitin, human serum albumin (HSA), SmbP, SUMO, Trx, Skp, Ecotin, SNAP-tag, slyd, DsbA, GST, Halo tag, MBP, NusA, p-galactosidase, Fc-domain fusion, CaBP, RFP, GFP, CFP, mCherry, mOrange, SEAP, and luciferase.
[0277] Embodiment 64. The Bxb1 integrase of any one of the preceding embodiments, further comprising fusion to one or more DNA-binding domains, and / or wherein the one or more DNA-binding domains comprises one or more TAL domains, one or more Cas proteins, and / or one or more zinc finger proteins (ZFPs).
[0278] Embodiment 65. The Bxb1 integrase of embodiment 64, wherein the Bxb1 is a fusion with about 1 to about 30 ZFPs, or about 1 to 5 ZFPs, about 5 to 10 ZFPs, about 10 to 15 ZFPs, about 15 to 20 ZFPs, about 20 to 25 ZFPs, or about 25 to 30 ZFPs.
[0279] Embodiment 66. A PhiC31 integrase comprising: an amino acid sequence that is at least 90% identical to amino acids 1-604 of SEQ ID NO: 2, and one or more amino acid mutations at positions selected from 1, 2, 12, 14, 18, 24, 32, 36, 41, 43, 44, 45, 49, 51, 55, 58, 74, 77, 96, 103, 107, 117, 153, 176, 196, 197, 199, 200, 228, 229, 230, 231, 235, 238, 240, 252, 254, 255, 260, 262, 264, 266, 269, 274, 278, 302, 319,320, 322, 331, 333, 340, 344, 345, 346, 347, 351, 353, 355, 359, 362, 364, 378, 382, 393, 396, 397, 399,406, 410, 424, 429, 431, 436, 438, 445, 448, 449, 450, 452, 457, 468, 472, 478, 498, 499, 501, 505, 512,514, 516, 517, 520, 527, 535, 536, 549, 551, 552, 553, 563, 568, 580, 585, 586, 587, 590, 592, 594, 600,603, 604, 609, 612, and / or 616, or a position corresponding thereto.
[0280] Embodiment 67. The P hi C31 integrase of embodiment 66, wherein the one or more amino acid mutations comprise M1E, M1V, D2V, D2M, S12N, E14G, S18N, A24A, D32A, D36A, V41I, R43R, D44A, G45G, R49R, V51M, S55S, P58P, I74I, E77E, R96R, 11031, S107S, V117V, 11531, 1153V, E176D, N196N, K197K, A199T, H200H, H228Y, L229L, P230S, F231L, S235S, A238A, H240R, D252G, D254D, A255G, G260G, T262T, G264G, K266K, S269N, P274S, M278L, T302A, L319L, R320R, V322V, E331E, A333T,A333S, A333D, A340V, G344V, G344D, R345R, G346S, R347K, L351V, L351Q, R353R, Q355Q, S359S,D362N, D362G, L364M, E378K, K382K, V393V, S396N, S396R, A397T, G399G, N406S, A410T, 14241, G429S, E431V, L436L, W438R, G445G, W448R, E449D, A450D, E452K, R457R, L468L, E472D, R478R,A498A, L499L, L501L, G505V, G505G, G505S, E512K, E514E, A516T, E517E, K520K, F527F, P535L,P535P, T536A, D549D, R551R, V552V, F553F, V563V, T568T, A580V, A585A, K586Q, P587P, D590G, D592G, D594N, D594D, T600S, T600T, V603I, V603A, V603V, A604S, P609S, V612V, and / or A616V, or a position corresponding thereto, relative to SEQ ID NO: 2.
[0281] Embodiment 68. The PhiC31 integrase of embodiment 66 or 67, wherein the PhiC31 integrase comprises a N-terminal addition of one or more amino acids.
[0282] Embodiment 69. The PhiC31 integrase of embodiment 68, wherein the N-terminal addition comprises an addition of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, 13 amino acids, 14 amino acids, 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 35 amino acids, about 40 amino acids, about 45 amino acids, or about 50 amino acids relative to SEQ ID NO: 2.
[0283] Embodiment 70. The PhiC31 integrase of embodiment 68 or 69, wherein the N-terminal addition comprises a sequence selected from: MEQGWSG (SEQ ID NO: 3), MIQGWSG (SEQ ID NO: 4),MIQEWSG (SEQ ID NO: 5), MTMIYPFPAQLTLTKGNKSWSSLVTAASVLEFATMIQGVAG (SEQ ID NO: 6),MTMITPSAQLTLTKSNKSWNSLVTAASVLEFATMIQGVAG (SEQ ID NO: 7),MTMITPSAQLTLTKGNKSWSSLVTAASVLEFATMIQGVAG (SEQ ID NO: 8),MTMITPSAQLTLTKSNKSWSSLVTAASVLEFATMIQGVTG (SEQ ID NO: 9),MTMITPSAQLTLTKDNKSWSSLVTAASVLEFATMIQGVAG (SEQ ID NO: 10),MTMITPSAQLTLTKSNKSWSSLVTAASVLEFATMIQGVAG (SEQ ID NO: 11),MTMITPSAQLTLTKGNKSWSSLVTAASVLEFATVIQGVAG (SEQ ID NO: 12),MTMITPSAQLTLTKSNKSWSSLVTAASVLEFATVIQGVTG (SEQ ID NO: 13),MTMITPSAQLTLTKDNKSWSSLVTAASVLEFATVIQGVAG (SEQ ID NO: 14), andMTMITPSAQLTLTKSNKSWSSLVTAASVLEFATVIQGVAG (SEQ ID NO: 15).
[0284] Embodiment 71 . The PhiC31 integrase of any one of embodiments 66-70, wherein the PhiC31 integrase comprises one or more mutations, N-terminal additions, and / or one or more combinations of mutations and / or N-terminal additions as described in Table 1 , Table 3, and / or Table 4 relative to SEQ ID NO: 2.
[0285] Embodiment 72. The PhiC31 integrase of any one of embodiments 66-71 , wherein the PhiC31 integrase has at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations, e.g., as selected from Table 1 , Table 3, and / or Table 4 relative to SEQ ID NO: 2.
[0286] Embodiment 73. The PhiC31 integrase of any one of embodiments 66-72, wherein the PhiC31 integrase at least 90% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations.
[0287] Embodiment 74. The PhiC31 integrase of any one of embodiments 66-72, wherein the PhiC31 integrase comprises at least 95% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, or at least 30 mutations
[0288] Embodiment 75. The PhiC31 integrase of any one of embodiments 66-72, wherein the PhiC31 integrase comprises at least 96% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations.
[0289] Embodiment 76. The PhiC31 integrase of any one of embodiments 66-72, the PhiC31 integrase comprises at least 97% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, or at least 15 mutations.
[0290] Embodiment 77. The PhiC31 integrase of any one of embodiments 66-72, wherein the PhiC31 integrase comprises at least 98% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, or 10 mutations
[0291] Embodiment 78. The PhiC31 integrase of any one of embodiments 66-72, the PhiC31 integrase comprises at least 99% sequence identity to SEQ ID NO: 2 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, or at least 5 mutations.
[0292] Embodiment 79. The PhiC31 integrase of any one of embodiments 66-78, wherein the Bxb1 integrase comprises one or more nuclear localization sequences.
[0293] Embodiment 80. The PhiC31 integrase of embodiment 79, wherein the one or more nuclear localization sequences is fused at the N-terminus, at the C-terminus, at both the N-terminus and C-terminus, or within the integrase sequence.
[0294] Embodiment 81. The PhiC31 integrase of claim 70 or 71, wherein the one or more nuclear localization sequences comprises a sequence selected from: PKKKRKV (SEQ ID NO: 57),NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 58); PAAKRVKLD (SEQ ID NO: 59), RQRRNELKRSF (SEQ ID NO: 60); NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 61), RMRKFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 62), VSRKRPRP (SEQ ID NO: 63), PPKKARED (SEQ ID NO: 64), PQPKKKPL (SEQ ID NO: 65), SALIKKKKKMAP (SEQ ID NO: 66), DRLRR (SEQ ID NO: 67), PKQKKRK (SEQ ID NO: 68), RKLKKKIKKL (SEQ ID NO: 69), REKKKFLKRR (SEQ ID NO: 70), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 71), RKCLQAGMNLEARKTKK (SEQ ID NO: 72), DPKKKRKVDPKKKRKVDPKKKRKV (SEQ ID NO: 73), KRTADGSEFESPKKKRKV (SEQ ID NO: 74), KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75), GSHHHHHHGSGPKKKRKV (SEQ ID NO: 76), and GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0295] Embodiment 82. The PhiC31 integrase of embodiment 81, wherein the one or more nuclear localization sequences is or comprises KRTADGSEFESPKKKRKV (SEQ ID NO: 74), KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75), or GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
[0296] Embodiment 83. The PhiC31 integrase of any one of embodiments 66-82, wherein the Bxb1 integrase comprises one or more stabilization sequences, optionally located at the N-terminus, C-terminus, or both the N-terminus and C-terminus.
[0297] Embodiment 84. The PhiC31 integrase of embodiment 83, wherein to the one or more stabilization sequences comprises an amino acid sequence that influences stability, solubility, folding, aggregation, degradation, purification, isolation, tracking expression and / or subcellular localization, relative to a cognate protein lacking the one or more stabilization sequences.
[0298] Embodiment 85. The PhiC31 integrase of embodiment 83 or 84, wherein the one or more stabilization sequences is or comprises about 5 amino acids to about 300 amino acids in length
[0299] Embodiment 86. The PhiC31 integrase of any one of embodiments 83-85, wherein the one or more stabilization sequences is or comprises the amino acid sequence of SEQ ID NO: 157, or an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes thereto.
[0300] Embodiment 87. The PhiC31 integrase of any one of embodiments 83-85, wherein the one or more stabilization sequences is or comprises a poly-amino acid sequence, optionally selected from poly-Cys, poly-Arg, poly-Asp, poly-Phe, and poly-His; and / or wherein the one or more stabilization sequences is or comprises one or more of a AU1 tag, AU5 tag, Glu-Glu tag EYMPME, FLAG tag, strep-tag II, HA tag, c-Myctag, T7 tag, VSV-G tag, KT3 epitope tag, HSV tag, Protein C tag, V5 tag, S-tag, twin-strep tag, SBP tag, CBD tag, and a protein eXact tag; and / or wherein the one or more stabilization sequences is or comprises a protein domain, or a portion thereof, optionally selected from ubiquitin, human serum albumin (HSA), SmbP, SUMO, Trx, Skp, Ecotin, SNAP-tag, slyd, DsbA, GST, Halo tag, MBP, NusA, -galactosidase, Fc-domain fusion, CaBP, RFP, GFP, CFP, mCherry, mOrange, SEAP, and luciferase.
[0301] Embodiment 88. The PhiC31 integrase of any one of embodiments 66-87, further comprising fusion to one or more DNA-binding domains, and / or wherein the one or more DNA-binding domains comprises one or more zinc finger proteins (ZFPs).
[0302] Embodiment 89. The PhiC31 integrase of embodiment 88, wherein the Bxb1 is a fusion with about 1 to about 30 ZFPs, or about 1 to 5 ZFPs, about 5 to 10 ZFPs, about 10 to 15 ZFPs, about 15 to 20 ZFPs, about 20 to 25 ZFPs, or about 25 to 30 ZFPs.
[0303] Embodiment 90. A nucleic acid encoding the BxB1 integrase of any one of claims 1-65 and / or the P hiC31 integrase of any one of claims 66-89.
[0304] Embodiment 91 . An expression vector encoding the nucleic acid of claim 90.
[0305] Embodiment 92. A host cell comprising the nucleic acid or vector of claim 90 or 91
[0306] Embodiment 93. A method of genomic insertion of a genomic payload comprising: providing a host cell; introducing at least one nucleic acid comprising: a first nucleic acid sequence encoding the Bxb1 integrase of any one of embodiments 1-65 and / or the PhiC31 integrase of any one of embodiments 66-89; and, optionally a second nucleic acid sequence encoding a genomic payload and one or more attB or attP site compatible with the Bxb1 integrase and / or PhiC31 integrase; and, optionally, a third nucleic acid sequence encoding one or more elements configured for genomic integration of one or more attB or attP site compatible with the Bxb1 integrase and / or the PhiC31 integrase.
[0307] Embodiment 94. The method of embodiment 93, wherein the method comprises introducing the first nucleic acid sequence and the second nucleic acid sequence.
[0308] Embodiment 95. The method of embodiment 94, wherein the method comprises integrase- mediated recombination between the one or more attB or attP site and one or more naturally-occurring genomic sequences of the host cell.
[0309] Embodiment 96. The method of embodiment 93, wherein the method comprises introducing the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence.
[0310] Embodiment 97. The method of embodiment 96, wherein the method comprises recombination between the one or more attB or attP sequence and one or more genomically-integrated attB or attP sequence in the genome of the host cell.
[0311] Embodiment 98. The method of any one of embodiments 93-97, further comprising introducing one or more additional nucleic acid sequences which encodes one or more mutant Bxb1 integrases that targets a “left” and / or a “right” flanking portion of one or more att site (e.g., having one or more mutations that alters DNA-binding relative to wild-type Bxb1), and / or wherein the one or more mutant Bxb1 integrases that targets a “left” and / or a “right” flanking portion are fusion proteins comprising a wild-type or mutant Bxb1 and one or more DNA-binding domains; and / or wherein the one or more DNA-binding domains comprise one or more TAL domains, one or more Cas proteins, and / or one or more zinc finger proteins (ZFPs); and / or further comprising integrase-mediated recombination using a mutant Bxb1 , a “left” Bxb1 fusion protein having one or more ZFPs targeting a “left” portion of an att sequence, and a “right” Bxb1 fusion protein having one or more ZFPs targeting a “right” portion of an att sequence.
[0312] Embodiment 99. The method of any one of embodiments 93-98, wherein the one or more elements configured for genomic integration comprises one or more of a zinc finger nuclease (ZFN), transcription-activator like effector nuclease (TALEN), endonuclease, Cas9 nickase, a reverse transcriptase PE5, and a Cas9-directed reverse transcription guide RNA (pegRNA)
[0313] Embodiment 100. The method of embodiment 99, wherein the third nucleic acid sequence encodes each of the Cas9 nickase, the reverse transcriptase PE5, and the Cas9-directed reverse transcription guide RNA (pegRNA).
[0314] Embodiment 101. The method of any one of embodiments 93-100, wherein each of the first nucleic acid sequence, second nucleic acid sequence, and / or third nucleic acid sequence is encoded on a single nucleic acid molecule, on two nucleic acid molecules, or on three nucleic acid molecules.
[0315] Embodiment 102. The method of any one of embodiments 93-101 , wherein the attB site is about or at least about 25 nucleotides in length, about or at least about 30 nucleotides in length, about or at least about 31 nucleotides in length, about or at least about 32 nucleotides in length, about or at least about 33 nucleotides in length, about or at least about 34 nucleotides in length, about or at least about 35 nucleotides in length, about or at least about 36 nucleotides in length, about or at least about 37 nucleotides in length, about or at least about 38 nucleotides in length, about or at least about 39 nucleotides in length, about or at least about 40 nucleotides in length, or about or at least about 50 nucleotides in length.
[0316] Embodiment 103. The method of any one of embodiments 93-102, wherein the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: ctcgaagccgcggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccacctcacccatc (SEQ ID NO: 79), or tcggccggcttgtcgacgacggcggtctccgtcgtcaggatcatccgggc (SEQ ID NO: 78); or wherein the one or more attB sites is or comprises a DNA sequence having the nucleic acid sequence of: ctcgaagccgcggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccacctcacccatc (SEQ ID NO: 79), or tcggccggcttgtcgacgacggcggtctccgtcgtcaggatcatccgggc (SEQ ID NO: 78).
[0317] Embodiment 104. The method of any one of embodiments 93-103, wherein the one or more atP sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: tggtgtccaggagccgaggtatcggtcctgccagggcc (SEQ ID NO: 166), or wherein the one or more attB sites is or comprises a DNA sequence having the nucleic acid sequence of: tggtgtccaggagccgaggtatcggtcctgccagggcc (SEQ ID NO: 166).
[0318] Embodiment 105. The method of any one of embodiments 93-104, wherein the one or more atP sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: tggtgtccaggagccgaggtatcggtcctgccagggcc (SEQ ID NO: 166), or wherein the one or more attB sites is or comprises a DNA sequence having the nucleic acid sequence of: tggtgtccaggagccgaggtatcggtcctgccagggcc (SEQ ID NO: 166).
[0319] Embodiment 106. The method of any one of embodiments 93-105, wherein the one or more atP sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: cgggagtagtgccccaactggggtaacctttgagttctctcagttgggggcgtagggtcg (SEQ ID NO: 81), or gtggtttgtctggtcaaccaccgcggtctcagtggtgtacggtacaaaccca (SEQ ID NO: 80), or wherein the one or more atP sites is or comprises a DNA sequence having the nucleic acid sequence of: cgggagtagtgccccaactggggtaacctttgagttctctcagttgggggcgtagggtcg (SEQ ID NO: 81), or gtggtttgtctggtcaaccaccgcggtctcagtggtgtacggtacaaaccca (SEQ ID NO: 80).
[0320] Embodiment 107. The method of any one of embodiments 93-106, wherein the one or more atB and / or atP sites targets to a central dinucleotide sequence comprising GT, GA, GG, GC, AT, AA, AG, AC, TT, TA, TG, TC, CO, CA, CT, CG.
[0321] Embodiment 108. The method of embodiment 107, wherein the one or more atB and / or atP sites targets to a central dinucleotide sequence comprising guanine-thymine (GT) or guanine-adenine (GA).
[0322] Embodiment 109. The method of any one of embodiments 93-108, wherein the pegRNA (epegRNA) comprises one or more nucleic acid sequences having one or more of a structured RNA motif, a cr772 guide scaffold, and / or a dominant negative mutant of human mutL homolog 1 (MLH1) involved in DNA mismatch repair.
[0323] Embodiment 110. The method of any one of embodiments 93-109, wherein the host cell comprises a mammalian cell, plant cell, insect cell, yeast cell, or bacterial cell.
[0324] Embodiment 112. The method of embodiment 110, wherein the mammalian comprises one or more of a human embryonic kidney (HEK293, HEK293T) cell line, K562 human lymphoblast cell line, U2OS human osteosarcoma cell line, primary human fibroblasts (human dermal fibroblast (HDFa) cell line), human erythroleukemia cell line (K562), Chinese hamster ovary (CHO) cell line (CHO-K1 , CHO-DHB11 , CHO-DXB1 , CHO-S, CHO-DG44, CHO-M), baby hamster kidney (BHK) cell line, Vero cell line (Vero, Vero 76, Vero E6), human cervical carcinoma cell line (HELA, 3T3), PERc6 cell line, CAP cell line, or monkey kidney CV1 cell line.
[0325] Embodiment 113. The method of any one of embodiments 93-112, wherein the efficiency of insertion of one or more attB and / or attP ranges from 0.1 % to 90%.
[0326] Embodiment 114. The method of any one of embodiments 93-113, wherein the at least one nucleic acid are introduced into the host cell by one or more of a lipid-based transfection reagent, diethylaminoethyl (DEAE)-dextran, liposomes, cationic lipid-based reagents, electroporation, sonoporation, chemical reagent (calcium phosphate), microinjection, or via a viral vector system (AAV, lentiviral).
[0327] Embodiment 115. The method of any one of embodiments 93-114, wherein the method inserts the genomic payload into the host cell and the genomic payload comprises a double-stranded DNA molecule ranging in size from about 0.1 kb to about 20 kb or more.
[0328] Embodiment 116. The method of embodiments 115, wherein the genomic payload is about or at least about 0.5 kb in length, about or at least about 1 kb in length, about or at least about 2 kb in length, about or at least about 3 kb in length, about or at least about 4 kb in length, about or at least about 5 kb in length, about or at least about 6 kb in length, about or at least about 7 kb in length, about or at least about 8 kb in length, about or at least about 9 kb in length, about or at least about 10 kb in length, about or at least about 11 kb in length, about or at least about 12 kb in length, about or at least about 13 kb in length, about or at least about 14 kb in length, about or at least about 15 kb in length, about or at least about 16 kb in length,about or at least about 17 kb in length, about or at least about 18 kb in length, about or at least about 19 kb in length, about or at least about 20 kb in length, about or at least about 30 kb, or about or at least about 50 kb.
[0329] Embodiment 117. The method of any one of embodiments 93-116, wherein the methods target to a locus of AAVS1 , Rosa26, Xq22.1 , MACO1, GYS1 , CCR5, ACTB, GBA1 , COL7A1 , FANCA, Smn1 , ALB, B2M, TRAC, Factor IX, CFTR.
[0330] Embodiment 118. The method of any one of embodiments 93-117, wherein the genomic payload comprises one or more of a chimeric receptor, chimeric antigen receptor, promoter sequence, enhancer sequence, and / or a genetic cassette encoding a therapeutic protein or corrective protein.
[0331] Embodiment 119. The method of any one of embodiments 93-118, wherein the integration rates are about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of target loci in a population of the host cell.EQUIVALENTS
[0332] While the disclosure has been described in connection with specific embodiments thereof, it will be understood that it is capable of further modifications and this application is intended to cover any variations, uses, or adaptations of the disclosure following, in general, the principles of the disclosure and including such departures from the present disclosure as come within known or customary practice within the art to which the disclosure pertains and as may be applied to the essential features hereinbefore set forth and as follows in the scope of the appended claims.
[0333] Those skilled in the art will recognize, or be able to ascertain, using no more than routine experimentation, numerous equivalents to the specific embodiments described specifically herein. Such equivalents are intended to be encompassed in the scope of the following claims.INCORPORATION BY REFERENCE
[0334] All patents and publications referenced herein are hereby incorporated by reference in their entireties.
[0335] The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure.
[0336] As used herein, all headings are simply for organization and are not intended to limit the disclosure in any manner. The content of any individual section may be equally applicable to all sections.
Claims
CLAIMSWhat is claimed is:
1. A Bxb1 integrase comprising:(i) an amino acid sequence that is at least 90% identical to amino acids 1-480 of SEQ ID NO: 1;(ii) an amino acid sequence that is at least 90% identical to amino acids 1-488 of SEQ ID NO: 1 ; or(iii) an amino acid sequence that is at least 90% identical to SEQ ID NO: 1; and wherein (i), (ii), or (iii) further comprises at least two amino acid substitutions at positions selected from (a) and (b), (b) and (c), (a) and (c), wherein:(a) is 75, 76, 158, 232, 234, 236, 237, 257, 314, 316, 318, 322, 323, 325, or a position corresponding thereto;(b) is 5, 14, 20, 24, 29, 35, 40, 45, 49, 50, 51, 60, 68, 69, 70, 73, 74, 78, 84, 86, 87, 100,105, 116, 124, 183, 197, 207, 208, 209, 229, 261, 267, 273, 273, 287, 291, 333, 342, 343, 347, 361, 368, 375, 435, 449, 453, 462, 483, 494, or a position corresponding thereto; and(c) is 4, 18, 34, 36, 42, 46, 61 , 62, 63, 67, 79, 85, 88, 89, 90, 92, 95, 99, 106, 110, 111 , 119,122, 130, 133, 137, 140, 145, 153, 156, 157, 158, 160, 164, 166, 174, 175, 178,179, 181, 187, 189, 191, 203, 218, 223, 231, 239, 248, 251, 254, 264, 268, 272,278, 280, 281, 282, 283, 285, 288, 292, 295, 302, 306, 307, 311, 313, 319, 321,328, 331, 332, 334, 353, 355, 359, 360, 362, 369, 370, 380, 388, 397, 398, 405,409, 411, 414, 415, 416, 419, 425, 428, 434, 444, 461, 463, 466, 468, 476, 479,480, 484, 487, 488, 489, 496, and / or 499, or a position corresponding thereto.
2. The Bxb1 integrase of claim 1 , wherein the one or more amino acid mutations are selected from L4I, V5I, DUN, E20K, E20Q, E24K, L29F, G34D, W35L, D36A, V40I, V40A, E42K, D45G, A49T, V50I, D51N, D51E, D51Y, N60S, L61F, R63K, F67S, E68K, E69A, E69D, Q70P, D73G, V74A, V74I, V74M, I75V, V76I, Y78H, Y78N, T84S, S86T, I87V, I87L, H89G, Q92H, H95Y, D99N, H100N, H100Y, V105A, V105I, H111P, T116P, A119S, V122M, A124S, A145T, S157G, T166I, V179A, V179I, R181K, E183L, V187I, H189N, P197T, H203Y, R207Q, R208S, G209V, E229K, A232G, A232S, A232R, A232T, A232V, A232P, A232Q, T233G, T233W, T233R, T233Y, T233D, T233G, T233N, T233H, T233S, T233Q, T233A, T233C, A234N, A234G, A234S, A234T, A234H, A234F, A234S, K236R, K236S, R237K, R237Q, R237N, R237V, R237C, M239I, A248T, N251K, T254S,D257K, A261T, A261V, V264A, E267D, R272Q, E273D, E273K, A280T, T285A, R287P, A288V, A288T, A291T, A311V, K313R, F314M, F314L, F314N, F314R, F314K, G316R, G316W, G318H, G318P, G318S, G318N, G318K, R319K, R319G, H321 P, H321 Q, H321 K, H321 L, H321 R, H321Y,H321T, H321 S, P322A, P322R, P322G, P322L, R323L, R323G, R323Y, R323I, R325K, R325Q,R325Y, F331 S, P332H, K333N, H334R, M342T, M342V, A343T, A347E, A347V, V353I, D355N,D359N, A360T, E361 D, R362K, V368A, A369P, A369E, A369T, A369S, V375I, V380I, T388M,A396P, A398S, R409H, A411V, A414V, A415S, R416K, A425T, E434G, T435A, R444L, A449V, T453I, T453A, L462M, T463I, V466M, G468D, L479I, Q480STOP, E483K, Q484K, R487R, del L488 (del L488 frameshift), R494S, R494Q, H496N, and / or M499T, or a position corresponding thereto, relative to S EQ ID NO: 1.
3. The Bxbl integrase of claim 1 or 2, wherein one or more of the following amino acids are not mutated relative to SEQ ID NO: 1 : V5, S18, E24, V40, V46, D51 , A62, E69, R79, R85, R88, L90, Q92, S106, A110, A130, E133, 1137, R140, K153, G156, P160, L164, L174, V175, P178, V187, Q191, G209,A218, R223, S231, N251 , P268, L278, A280, E281 , L282, V283, R287, P292, P295, L302, V306,C307, K313, S328, K333, D359, E361, G370, V375, R397, A405, A414, E419, A425, S428, T435,R461 , V466, G468, F476, E483, G489, and R494.
4. The Bxb1 integrase of any one of claims 1-3, wherein the Bxb1 integrase comprises a C-terminal deletion of one or more amino acids in the range of L488-S500 relative to SEQ ID NO: 1 .
5. The Bxb1 integrase of claim 4, wherein the C-terminal deletion comprises a deletion of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, or all 13 amino acids of L488- S500 relative to SEQ ID NO: 1.
6. The Bxb1 integrase of claim 4 or 5, wherein L488 is mutated to a stop codon (L488STOP).
7. The Bxb1 integrase of any one of the preceding claims, wherein the Bxb 1 integrase comprises a C- terminal deletion of one or more amino acids in the range of Q480-S500 relative to SEQ ID NO: 1 .
8. The Bxb1 integrase of claim 7, wherein the C-terminal deletion comprises a deletion of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 6 amino acids, 7 amino acids, 8 amino acids, 9 amino acids, 10 amino acids, 11 amino acids, 12 amino acids, 13 amino acids, 14 aminoacids, 15 amino acids, 16 amino acids, 17 amino acids, 18 amino acids, 19 amino acids, 20 amino acids, or all 21 amino acids of Q480-S500 relative to SEQ ID NO: 1 .
9. The Bxb1 integrase of any one of claims 1 -3, wherein Q480 is mutated to a stop codon (Q480STOP).
10. The Bxb1 integrase of any one of the preceding claims, wherein the one or more amino acid mutations comprises mutation of one or more of residues A157-P160; and / or wherein the mutation alters specificity at -7 and / or -6 positions of a loop target DNA site compared to a wild-type Bxb1 of SEQ ID NO: 1.
11. The Bxb1 integrase of any one of the preceding claims, wherein the one or more amino acid mutations comprises mutation of one or more of residues Y154-P159; and / or wherein the one or more mutations alters specificity at -7 and / or -6 positions of a loop target DNA site compared to a wild-type Bxb1 of SEQ ID NO: 1.
12. The Bxb1 integrase of claim 10 or 11 , wherein the mutation comprises S157G.
13. The Bxb1 integrase of claim 11, wherein the one or more Bxb1 integrases comprise a sequence from residues Y154-P159 selected from YRGSLS (SEQ ID NO: 16), WWGGTP (SEQ ID NO: 17), YRGGLP (SEQ ID NO: 18), AHPGRT (SEQ ID NO: 19), and WRGGAS (SEQ ID NO: 20) relative to SEQ ID NO: 1 , or a position corresponding thereto.
14. The Bxb1 integrase of any one of claims 10-13, wherein the one or more amino acid mutations at this position alters the dinucleotide preference of adenine (A) or thymine (T) at position -7 and / or cytosine (C) or guanine (G) at position -6.
15. The Bxb1 integrase of any one of the preceding claims, wherein the one or more amino acid mutations comprises mutation of one or more of residues S231-R237; and / or wherein the mutation alters specificity at -11, -10, and / or -9 positions of a 3 bp helix target site compared to a wild-type Bxb1 of SEQ ID NO: 1.
16. The Bxb1 integrase of any one of the preceding claims, wherein L235 is not mutated.
17. The Bxb1 integrase of claim 15 or 16, wherein the one or more amino acid mutations comprises mutation of one or more residues of 231-234 and / or 236-237.
18. The Bxb1 integrase of any one of claims 15-17, wherein A234 is mutated to an asparagine (A234N); and / or wherein this mutation alters specificity at the -10 position to thymine (T) compared to a wildtype Bxb1 of SEQ ID NO: 1.
19. The Bxb1 integrase of any one of claims 15-18, wherein the one or more amino acid mutations comprises mutation of one or more of residues 231, 233, and 237; and / or wherein the mutation alters specificity at -10 and / or -9 positions of a 3 bp helix target site compared to a wild-type Bxb1 of SEQ ID NO: 1.
20. The Bxb1 integrase of any one of claims 15-19, wherein the Bxb1 integrase comprises 1 , 2, 3, 4, or 5 amino acids mutated between residues S231-R237 (inclusive) relative to SEQ ID NO: 1.
21. The Bxb1 integrase of any one of claims 15-20, wherein the Bxb1 integrase comprises a sequence from residues S231-R237 selected from AGGNLKR (SEQ ID NO: 21), LGTNLKR (SEQ ID NO: 22), SGTGLKK (SEQ ID NO: 23), SGSALKT (SEQ ID NO: 24), AAWALRR (SEQ ID NO: 25), GGRSLKR (SEQ ID NO: 26), SGYNLRR (SEQ ID NO: 27), SGWGLKK (SEQ ID NO: 28), SGWALRQ (SEQ ID NO: 29), SARALSR (SEQ ID NO: 30), RADTLRR (SEQ ID NO: 31), YSRNLKR (SEQ ID NO: 32), SRNGLRK (SEQ ID NO: 33), RGHALKN (SEQ ID NO: 34), GGSHLKR (SEQ ID NO: 35), TTRTLKR (SEQ ID NO: 36), AVQNLKR (SEQ ID NO: 37), RAAFLKK (SEQ ID NO: 38), RAWTLKC (SEQ ID NO: 39), RAWSLKR (SEQ ID NO: 40), HGWSLKV (SEQ ID NO: 41), HGCTLKR (SEQ ID NO: 42), YGSALKQ (SEQ ID NO: 43), SQWALKC (SEQ ID NO: 44), YPWSLRR (SEQ ID NO: 45), relative to SEQ ID NO: 1 , or a position corresponding thereto.
22. The Bxb1 integrase of any one of the preceding claims, wherein the Bxb1 integrase comprises a mutation of D257K relative to SEQ ID NO: 1 .
23. The Bxb1 integrase of any one of the preceding claims, wherein the one or more amino acid mutations comprises mutation of one or more of residues K313-S328; and / or wherein the mutation alters specificity at alter specificity at -19 through -12 positions of a DNA binding site compared to a wild-type Bxb1 of SEQ ID NO: 1.
24. The Bxbl integrase of any one of the preceding claims, wherein the Bxb1 residue P322 is maintained as a proline.
25. The Bxb1 integrase of claim 23 or 24, wherein the one or more amino acid mutations comprises mutation of one or more of residues K313, F314, A315, G316, G318, R319, H321, P322, R323, R325, and S328.
26. The Bxb1 integrase of claim 25, wherein the one or more amino acid mutations comprises mutation of one or more of residues K313, R319, H321 , and / or S328.
27. The Bxb1 integrase of any one of the preceding claims, wherein the Bxb1 integrase comprises 1 , 2, 3, 4, 5, 6, or 7 amino acids mutated between residues F314-R325 (inclusive) relative to SEQ ID NO: 1.
28. The Bxb1 integrase of claim 27, wherein the Bxb1 integrase comprises a sequence from residues K314-R325 selected from MAGGHRKQALYR (SEQ ID NO: 46), MAGGPRKKRRYR (SEQ ID NO: 47), LARGSRKLALYR (SEQ ID NO: 48), NARGNRKRGRYR (SEQ ID NO: 49), LARGPRKRAGYK (SEQ ID NO: 50), RAWGKRKYAYYQ (SEQ ID NO: 51), KAWGSRKTRLYR (SEQ ID NO: 52), MARGGRKSAIYY (SEQ ID NO: 53), MASGSRKTAIYY (SEQ ID NO: 54), LARGRRKWARYR (SEQ ID NO: 55), and LARGSRKLALYR (SEQ ID NO: 56), realtive to SEQ ID NO: 1.
29. The Bxb1 integrase of any one of the preceding claims, wherein the Bxb1 integrase comprises one or more mutations and / or deletions, and / or one or more combinations of mutations and / or deletions as described in Table 1 and / or Table 2 relative to SEQ ID NO: 1.
30. The Bxb1 integrase of any one of the preceding claims, wherein the Bxb1 integrase comprises at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations, e.g., as selected from Table 1 and / or Table 2 relative to SEQ ID NO: 1 .
31. The Bxb1 integrase of any one of claims 1-29, wherein the Bxb1 integrase comprises at least 90% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations,at least 17 mutations, at least 18 mutations, at least 19 mutations, at least 20 mutations, at least 30 mutations, at least 40 mutations, or at least 50 mutations.
32. The Bxb1 integrase of any one of claims 1-29, wherein the Bxb1 integrase comprises at least 95% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, or at least 20 mutations.
33. The Bxb1 integrase of any one of claims 1-29, wherein the Bxb1 integrase comprises at least 96% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, at least 15 mutations, at least 16 mutations, at least 17 mutations, at least 18 mutations, at least 19 mutations, or 20 mutations.
34. The Bxb1 integrase of any one of claims 1-29, wherein the Bxb1 integrase comprises at least 97% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, at least 10 mutations, at least 11 mutations, at least 12 mutations, at least 13 mutations, at least 14 mutations, or 15 mutations.
35. The Bxb1 integrase of any one of claims 1-29, wherein the Bxb1 integrase comprises at least 98% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, at least 5 mutations, at least 6 mutations, at least 7 mutations, at least 8 mutations, at least 9 mutations, or 10 mutations.
36. The Bxb1 integrase of any one of claims 1-29, wherein the Bxb1 integrase comprises at least 99% sequence identity to SEQ ID NO: 1 and having at least 1 mutation, at least 2 mutations, at least 3 mutations, at least 4 mutations, or 5 mutations.
37. The Bxb1 integrase of claim 1, wherein the Bxb1 integrase comprises: an amino acid sequence that is at least 95% identical to the amino acid sequence of SEQ ID NO:1 ; andfurther comprises one or more amino acid mutations at positions selected from L4, E42, R57, R63, V76, H111 , V122, V187, K313, R319, A341 , D359, A369, A398, R416, A425, L479, and M499, relative to SEQ ID NO: 1 .
38. The Bxb1 integrase of claim 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 97% identical to the amino acid sequence of SEQ ID NO: 1 .
39. The Bxb1 integrase of claim 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 98% identical to the amino acid sequence of SEQ ID NO: 1 .
40. The Bxb1 integrase of claim 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 .
41. The Bxb1 integrase of any one of claims 37-40, wherein the Bxb1 integrase comprises one or more amino acid mutations are selected from L4I, E42K, R57K, R63K, V76I, H111 P, V122M, V187, K313R, R319K, A341 D, D359A, A369P, A398S, R416K, A425T, L479I, and M499T relative to SEQ ID NO: 1.
42. The Bxb1 integrase of any one of claims 37-41, wherein the Bxb1 integrase further comprises one or more amino acid mutations selected from D36N, V40A, A49S, V74I, I87L or I87A or I87S, H89G, V175A, V179A, R287H, A288V, T453I, and H496N relative to SEQ ID NO: 1.
43. The Bxb1 integrase of any one of claims 37-42, wherein the Bxb1 integrase further comprises one or more amino acid mutations selected from V40I, D45G, I75V, H95Y, A119S, A280T, A311V, E434G, and V466M relative to SEQ ID NO: 1 .
44. The Bxb1 integrase of any one of claims 37-43, wherein one or more of the following amino acids relative to SEQ ID NO: 1 are not mutated: A62, H100, Q191, P195, G209, P295, L302, C307, L387, E419, S428, E483, and R487 relative to SEQ ID NO: 1.
45. The Bxb1 integrase of any one of claims 37-44, wherein the one or more amino acid mutations are selected from mutations at position V122 and / or position A369 relative to SEQ ID NO: 1 .
46. The Bxb1 integrase of claim 45, wherein the mutation at position V122 relative to SEQ ID NO: 1 is V122M.
47. The Bxb1 integrase of claim 45, wherein the mutation at position A369 relative to SEQ ID NO: 1 is A369P.
48. The Bxb1 integrase of any one of claims 37-47, wherein the one or more amino acid mutations relative to SEQ ID NO: 1 are selected from a combination of positions selected from:V76 and V122;V76 and A369;I87 and V122;I87 and A369;H95 and V122;H95 and A369;V122 and E434;A369 and E434;V76, V122, and A369;I87, H95 and A369;I87, V122 and A369;I87, V122 and E434;I87, A369 and E434;H95, V122 and A369;H95, V122 and E434;H95, A369 and E434;V122, A369 and E434;I87, H95, V122 and E434;I87, H95, A369 and E434; andI87, V122, A369 and E434.
49. The Bxb1 integrase of claim 48, wherein the one or more mutations relative to SEQ ID NO: 1 are a combination selected from:V76I and V122M;V76I and A369P;I87L and V122M;I87L and A369P;H95Y and V122;H95Y and A369P;V122M and E434G;A369P and E434G;V76I, V122M, and A369P;I87L, H95Y and A369P;I87L, V122M and A369P;I87L, V122M and E434G;I87L, A369P and E434G;H95, V122M and A369P;H95Y, V122M and E434G;H95Y, A369P and E434G;V122, A369P and E434G;I87L, H95Y, V122M and E434G;I87L, H95Y, A369P and E434G; andI87L, V122, A369P and E434G.
50. The Bxb1 integrase of any one of claims 37-47, wherein the one or more amino acid mutations relative to SEQ ID NO: 1 are a combination of positions selected from: I87, A369, and E434, V122, A369, and E434, H95, A369, and E434, V122 and E434,V122, A369, E434, and V175,V76 and A119,A119, A369, and E434,I87, V175, and D359,I87, R57, and R63,I87 and A369,I87 and A425,R63 and V122,R63, V122, R319, and V40I,R63, A119, and M499,V122 and A369,H95 and A369,A49, 187, A396, and E434,R57, R63, 187, A396, and E434,D36, V122, A396, and E434,R57, R63 and H95,D36, 187, A396, and E434,I87, A341. A396, and E434,R57, R63, V122, A396, and E434,V122, R287, A396, and E434,I87, R287, A396, and E434,I87 and A341,I87, R287, A396, E434, and A341,V122, A396, E434, and A341,H95, and A341,A49, V122, A396, and E434,D45, A396, and E434,V122, R287, A341, A396, and E434,A49, A396, and E434,I87, A341, and R287,D36, A396, and E434,V122, D359, A369, and E434,R63 and I87,I87, D359, A369, and E434,V122, A369, A425, and E434,I87, A369, A425, and E434,V40, H95, A369, and E434,L4, H95, A369, and E434,H95, A369, E434, and V466,R63, V122, A369, and E434,V76, 187, and K313,R63, 187, and N194,I87 and D359,L4 and I87,V74 and V76,V122, A369, E434, and R319,I87, A369, E434, and R319,L4, V122, A369, and E434,I87 and A425,E42, V76, and 187,I87 and L479,H95, A369, A425, and E434,R63, 187, A369, and E434,I87 and R319,V76, A369, and E434,H95, A369, E434, and L479,V74, 175, and V76,H95, R319, A369, and E434,L4, 187, A369, and E434,V122, V175, D359, A369, and E434,R57, 187, A369, and E434,E42, H95, A369, and E434,H95, V179, A369, and E434,I75 and V76,L4 and V76,H95 and K313,V40, R63, and H89,A288, A311, A398, R416, T453, H496, and E42, andH111 and E434.
51. The Bxb1 integrase of claim 50, wherein the one or more mutations relative to SEQ ID NO: 1 are a combination selected from:I87L, A369P, and E434G,V122M, A369P, and E434G,H95Y, A369P, and E434G,I87L and E434G,V122M and E434G,V122M, A369P, E434G, and V175A,V76I and A119S,D36N and A119S,A119S, A369P, and E434G,I87L, V175A, and D359A,I87L, R57K, and R63K,I87L and D36N,I87L and A369P,I87L and A425T,R63K and V122M,R63K, V122M, R319K, and V40I,R63K, A119S, and M499T,H95Y and E434G,V122M and A369P,H95Y and V466M,H95Y and A369P,A49S, I87L, A396P, and E434G,A49S and I87L,R57K, R63K, I87L, A396P, and E434G,D36N, V122M, A396P, and E434G,R57K, R63K and H95Y,D36N, I87L, A396P, and E434G,D36N and H95Y,I87L, A341D, A396P, and E434G,A49S and H95Y,R57K, R63K, V122M, A396P, and E434G,V122M, R287H, A396P, and E434G,I87L and R287H,I87L, R287H, A396P, and E434G,I87L and A341D,I87L, R287H, A396P, E434G, and A341D,V122M, A396P, E434G, and A341D,H95Y and A341D,A49S, V122M, A396P, and E434G,D45G, A396P, and E434G,V122M, R287H, A341D, A396P, and E434G,A49S, A396P, and E434G,I87L, A341D, and R287H,H95Y and R287H,D36N, A396P, and E434G,V122M, D359A, A369P, and E434G,R63K and I87L,I87L and V466M,I87L, D359A, A369P, and E434G,V122M, A369P, A425T, and E434G,I87L, A369P, A425T, and E434G,V40I, H95Y, A369P, and E434G,L4I, H95Y, A369P, and E434G,H95Y, A369P, E434G, and V466M,R63K, V122M, A369P, and E434G,V76I, I87L, and K313R,R63K, I87V, and N194D,I87L and D359A,I87L and L4IV74I and V76I,V122M, A369P, E434G, and R319K,I87L, A369P, E434G, and R319K,L4I, V122M, A369P, and E434G,I87L and A425T,E42K, V76I, and I87L,I87L and L479I,H95Y, A369P, A425T, and E434G,V40I and I87L,R63K, I87L, A369P, and E434G,I87L and R319K,V76I, A369P, and E434G,H95Y, A369P, E434G, and L479I,V74I, I75V, and V76I,H95Y, R319K, A369P, and E434G,L4I, I87L, A369P, and E434G,V122M, V175A, D359A, A369P, and E434G,R57K, I87L, A369P, and E434G,E42K, H95Y, A369P, and E434G,H95Y, V179A, A369P, and E434G,I75V and V76I,I87L and V179A,L4I and V76I,H95Y and K313R,H95Y and A280T,V40A, R63K, and H89G,A288V, A311V, A398S, R416K, T453I, H496N, and E42K, andH111 P and E434G.
52. The Bxb1 integrase of claim 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations I87L, A369P, and E434G relative to SEQ ID NO: 1.
53. The Bxb1 integrase of claim 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations V122M, A369P, and E434G relative to SEQ ID NO: 1.
54. The Bxb1 integrase of claim 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations H95Y, A369P, and E434G relative to SEQ ID NO: 1.
55. The Bxb1 integrase of claim 37, wherein the Bxb1 integrase comprises an amino acid sequence that is at least 99% identical to the amino acid sequence of SEQ ID NO: 1 and having the amino acid mutations V74A, E229K, and V375I relative to SEQ ID NO: 1.
56. The Bxb1 integrase of any one of the preceding claims, wherein the Bxb1 integrase comprises one or more nuclear localization sequences.
57. The Bxb1 integrase of claim 56, wherein the one or more nuclear localization sequences is fused at the N-terminus, at the C-terminus, at both the N-terminus and C-terminus, or within the integrase sequence.
58. The Bxb1 integrase of claim 56 or 57, wherein the one or more nuclear localization sequences comprises a sequence selected from: PKKKRKV (SEQ ID NO: 57), NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 58); PAAKRVKLD (SEQ ID NO: 59), RQRRNELKRSF (SEQ ID NO: 60); NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 61), RMRKFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 62), VSRKRPRP (SEQ ID NO: 63), PPKKARED (SEQ ID NO: 64), PQPKKKPL (SEQ ID NO: 65), SALIKKKKKMAP (SEQ ID NO: 66), DRLRR (SEQ ID NO: 67), PKQKKRK (SEQ ID NO: 68), RKLKKKIKKL (SEQ ID NO: 69), REKKKFLKRR (SEQ ID NO: 70), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 71), RKCLQAGMNLEARKTKK (SEQ ID NO: 72), DPKKKRKVDPKKKRKVDPKKKRKV (SEQ ID NO: 73), KRTADGSEFESPKKKRKV (SEQ ID NO: 74),KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75), GSHHHHHHGSGPKKKRKV (SEQ ID NO: 76), and GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
59. The Bxb1 integrase of claim 58, wherein the one or more nuclear localization sequences is or comprises KRTADGSEFESPKKKRKV (SEQ ID NO: 74),KRPAATKKAGQAKKKKGGGGSGGGGSGSKRPAATKKAGQAKKKK (SEQ ID NO: 75), or GSGSGSHHHHHHGSGPKKKRKV (SEQ ID NO: 77).
60. The Bxb1 integrase of any one of the preceding claims, wherein the Bxb1 integrase comprises one or more stabilization sequences, optionally located at the N-terminus, C-terminus, or both the N- terminus and C-terminus.
61. The Bxb1 integrase of claim 60, wherein the one or more stabilization sequences comprises an amino acid sequence that influences stability, solubility, folding, aggregation, degradation, purification, isolation, expression tracking, and / or subcellular localization, relative to a cognate protein lacking the one or more stabilization sequences.
62. The Bxb1 integrase of claim 60 or 61 , wherein the one or more stabilization sequences is or comprises about 5 amino acids to about 300 amino acids in length.
63. The Bxb1 integrase of any one of claims 60-62, wherein the one or more stabilization sequences is or comprises the amino acid sequence of any one of SEQ ID NOs: 157 and 162-165, or an amino acid sequence having 1 , 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes thereto.
64. The Bxb1 integrase of any one of claims 60-63, wherein the one or more stabilization sequences is or comprises a poly-amino acid sequence, optionally selected from poly-Cys, poly-Arg, poly-Asp, poly-Phe, and poly-His; and / or wherein the one or more stabilization sequences is or comprises one or more of a AU1 tag, AU5 tag, Glu-Glu tag EYMPME, FLAG tag, strep-tag II, HA tag, c-Myc tag, T7 tag, VSV-G tag, KT3 epitope tag, HSV tag, Protein C tag, V5 tag, S-tag, twin-strep tag, SBP tag, CBD tag, and a protein eXact tag; and / or wherein the one or more stabilization sequences is or comprises a protein domain, or a portion thereof, optionally selected from ubiquitin, human serum albumin (HSA), SmbP, SUMO, Trx, Skp, Ecotin, SNAP-tag, slyd, DsbA, GST, Halo tag, MBP, NusA, P-galactosidase, Fc-domain fusion, CaBP, RFP, GFP, CFP, mCherry, mOrange, SEAP, and luciferase.
65. The Bxb1 integrase of any one of the preceding claims, further comprising fusion to one or more DNA-binding domains.
66. The Bxb1 integrase of claim 65, wherein the one or more DNA-binding domains comprises one or more TAL domains, one or more Cas proteins, and / or one or more zinc finger proteins (ZFPs).
67. The Bxbl integrase of claim 66, wherein the Bxb1 is a fusion with about 1 to about 30 ZFPs, or about 1 to 5 ZFPs, about 5 to 10 ZFPs, about 10 to 15 ZFPs, about 15 to 20 ZFPs, about 20 to 25 ZFPs, or about 25 to 30 ZFPs.
68. A nucleic acid encoding one or more Bxb1 integrase of any one of claims 1-67.
69. An expression vector encoding the nucleic acid of claim 68.
70. A host cell comprising the nucleic acid of claim 68 or vector of claim 69.
71. A method of genomic insertion of a genomic payload comprising: introducing into one or more host cells at least one nucleic acid molecule comprising: a first nucleic acid sequence encoding one or more Bxb1 integrases of any one of claims 1- 67; and, optionally, a second nucleic acid sequence encoding the genomic payload and one or more attB or attP site compatible with the one or more Bxb1 integrases; and, optionally, a third nucleic acid sequence encoding one or more elements configured for genomic integration of one or more attB or attP sites compatible with the one or more Bxb1 integrases.
72. The method of claim 71 , comprising introducing the first nucleic acid sequence and the second nucleic acid sequence.
73. The method of claim 72, further comprising integrase-mediated recombination between the one or more attB or attP site and one or more naturally-occurring genomic sequences of the host cell.
74. The method of claim 71 , comprising introducing the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence.
75. The method of claim 74, further comprising integrase-mediated recombination between the one or more attB or attP sequence and one or more genomically-integrated attB or attP sequence in the genome of the host cell.
76. The method of any one of claims 71-75, further comprising introducing one or more additional nucleic acid sequences which encodes one or more Bxb1 integrases that targets a “left” and / or a “right” flanking portion of one or more aft site, optionally wherein the one or more Bxb1 integrases that targets a “left” and / or a “right” flanking portion are fusion proteins comprising a wild-type or mutant Bxb1 and one or more DNA-binding domains.
77. The method of claim 76, wherein the one or more DNA-binding domains comprise one or more zinc finger proteins (ZFPs).
78. The method of claim 77, further comprising integrase-mediated recombination using a mutant Bx b 1 , a “left” Bxb1 fusion protein having one or more ZFPs targeting a “left” portion of an att sequence, and a “right” Bxb1 fusion protein having one or more ZFPs targeting a “right” portion of an att sequence.
79. The method of any one of claims 71-78, wherein the one or more elements configured for genomic integration comprises one or more of a zinc finger nuclease (ZFN), transcription-activator like effector nuclease (TALEN), endonuclease, Cas9 nickase, a reverse transcriptase, and a Cas9-directed reverse transcription guide RNA (pegRNA).
80. The method of claim 79, wherein the third nucleic acid sequence encodes each of the Cas9 nickase, the reverse transcriptase, and the Cas9-directed reverse transcription guide RNA (pegRNA).
81. The method of any one of claims 71-80, wherein each of the first nucleic acid sequence, second nucleic acid sequence, and / or third nucleic acid sequence is encoded on a single nucleic acid molecule, on two nucleic acid molecules, or on three nucleic acid molecules.
82. The method of any one of claim 71 -81 , wherein the attB site is about or at least about 25 nucleotides in length, about or at least about 30 nucleotides in length, about or at least about 31 nucleotides in length, about or at least about 32 nucleotides in length, about or at least about 33 nucleotides in length, about or at least about 34 nucleotides in length, about or at least about 35 nucleotides in length, about or at least about 36 nucleotides in length, about or at least about 37 nucleotides in length, about or at least about 38 nucleotides in length, about or at least about 39 nucleotides in length, about or at least about 40 nucleotides in length, or about or at least about 50 nucleotides in length.
83. The method of any one of claims 71-82, wherein the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: ctcgaagccgcggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccacctcacccatc (SEQ ID NO: 79), or tcggccggcttgtcgacgacggcggtctccgtcgtcaggatcatccgggc (SEQ ID NO: 78); or wherein the one or more attB sites is or comprises a DNA sequence having the nucleic acid sequence of: ctcgaagccgcggtgcgggtgccagggcgtgcccttgggctccccgggcgcgtactccacctcacccatc (SEQ ID NO: 79), or tcggccggcttgtcgacgacggcggtctccgtcgtcaggatcatccgggc (SEQ ID NO: 78); or wherein the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: tggtgtccaggagccgaggtatcggtcctgccagggcc (SEQ ID NO: 166), orwherein the one or more attB sites is or comprises a DNA sequence having the nucleic acid sequence of: tggtgtccaggagccgaggtatcggtcctgccagggcc (SEQ ID NO: 166); or wherein the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: ctgagcgcctctcctgggcttgccaaggactcaaaccc (SEQ ID NO: 167), or wherein the one or more attB sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: ctgagcgcctctcctgggcttgccaaggactcaaaccc (SEQ ID NO: 167).
84. The method of any one of claim 71-83, wherein the one or more attP sites is about or at least about30 nucleotides in length, about or at least about 35 nucleotides in length, about or at least about 36 nucleotides in length, about or at least about 37 nucleotides in length, about or at least about 38 nucleotides in length, about or at least about 39 nucleotides in length, about or at least about 40 nucleotides in length, about or at least about 41 nucleotides in length, about or at least about 42 nucleotides in length, about or at least about 43 nucleotides in length, about or at least about 44 nucleotides in length, about or at least about 45 nucleotides in length, about or at least about 50 nucleotides in length, or about 60 nucleotides in length.
85. The method of claim 71-84, wherein the one or more attP sites is or comprises a DNA sequence having about or at least about 70% sequence identity to the nucleic acid sequence of: cgggagtagtgccccaactggggtaacctttgagttctctcagttgggggcgtagggtcg (SEQ ID NO: 81), or gtggtttgtctggtcaaccaccgcggtctcagtggtgtacggtacaaaccca (SEQ ID NO: 80), or wherein the one or more attP sites is or comprises a DNA sequence having the nucleic acid sequence of: cgggagtagtgccccaactggggtaacctttgagttctctcagttgggggcgtagggtcg (SEQ ID NO: 81), or gtggtttgtctggtcaaccaccgcggtctcagtggtgtacggtacaaaccca (SEQ ID NO: 80).
86. The method of any one of claims 71-85, wherein the one or more attB and / or attP sites targets to a central dinucleotide sequence comprising GT, GA, GG, GC, AT, AA, AG, AC, TT, TA, TG, TO, CO, CA, CT, CG.
87. The method of claim 86, wherein the one or more attB and / or attP sites targets to a central dinucleotide sequence comprising guanine-thymine (GT) or guanine-adenine (GA).
88. The method of any one of claims 71-87, wherein the one or more elements comprises one or more nucleic acid sequence encoding one or more pegRNA (epegRNA) comprising one or more nucleicacid sequences having one or more of a structured RNA motif, a cr772 guide scaffold, and / or a dominant negative mutant of human mutL homolog 1 (MLH1) involved in DNA mismatch repair.
89. The method of any one of claims 71-88, wherein the host cell comprises a mammalian cell, plant cell, insect cell, yeast cell, or bacterial cell.
90. The method of claim 89, wherein the mammalian comprises one or more of a human embryonic kidney (HEK293, HEK293T) cell line, K562 human lymphoblast cell line, U2OS human osteosarcoma cell line, primary human fibroblasts (human dermal fibroblast (HDFa) cell line), human erythroleukemia cell line (K562), Chinese hamster ovary (CHO) cell line (CHO-K1 , CHO-DHB11, CHO-DXB1 , CHO-S, CHO-DG44, CHO-M), baby hamster kidney (BHK) cell line, Vero cell line (Vero, Vero 76, Vero E6), human cervical carcinoma cell line (HELA, 3T3), PERc6 cell line, CAP cell line, or monkey kidney CV1 cell line.
91. The method of any one of claims 71 -90, wherein the efficiency of insertion of one or more attB and / or attP ranges from 0.1 % to 90%.
92. The method of any one of claims 71 -91 , wherein the at least one nucleic acid are introduced into the host cell by one or more of a lipid-based transfection reagent, diethylaminoethyl (DEAE)-dextran, liposomes, cationic lipid-based reagents, electroporation, sonoporation, chemical reagent (e.g., calcium phosphate), microinjection, or via a viral vector system (e.g., AAV, lentiviral).
93. The method of any one of claims 71-92, wherein the method inserts the genomic payload into the host cell and the genomic payload comprises a double-stranded DNA molecule ranging in size from about 0.1 kb to about 20 kb or more.
94. The method of claim 93, wherein the genomic payload is about or at least about 0.5 kb in length, about or at least about 1 kb in length, about or at least about 2 kb in length, about or at least about 3 kb in length, about or at least about 4 kb in length, about or at least about 5 kb in length, about or at least about 6 kb in length, about or at least about 7 kb in length, about or at least about 8 kb in length, about or at least about 9 kb in length, about or at least about 10 kb in length, about or at least about 11 kb in length, about or at least about 12 kb in length, about or at least about 13 kb in length, about or at least about 14 kb in length, about or at least about 15 kb in length, about or at least about 16 kb in length, about or at least about 17 kb in length, about or at least about 18 kb in length, aboutor at least about 19 kb in length, about or at least about 20 kb in length, about or at least about 30 kb, or about or at least about 50 kb.
95. The method of any one of claims 71-94, wherein the methods target to one or more locus of AAVS1 , Rosa26, TRAC, Xq22.1, MACO1, GYS1 , CCR5, ACTB, GBA1 , COL7A1 , FANCA, Smn1 , ALB, B2M, Factor IX, and CFTR.
96. The method of any one of claims 71-95, wherein the genomic payload comprises one or more of a chimeric receptor, chimeric antigen receptor, promoter sequence, enhancer sequence, and / or a genetic cassette encoding a therapeutic protein or corrective protein.
97. The method of any one of claims 71-96, wherein the integration rates are about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% of target loci in a population of the host cell.
Citation Information
Patent Citations
Site-specific recombination systems for use in eukaryotic cells
US20060046294A1
Site-specific serine recombinases and methods of their use
US20060172377A1