Engineered modular recombinase compositions, methods, and systems for site-specific DNA recombination
Patent Information
- Application Number
- PCT/US2025/023078
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-08
- Filing Date
- 2025-04-03
- Publication Date
- 2025-11-13
AI Technical Summary
The utility of large serine recombinases in genetic engineering and gene therapy is limited by their rarity of binding sites in the human genome and the inability to efficiently engineer them to target desired DNA sequences.
Development of variant recombinases with specific amino acid mutations in DNA binding regions to alter target site specificity, allowing recognition of a variety of DNA attachment sites, including pseudo-sites, attB-like, and attP-like target sites, and genomic sites, through modular combinations of RD and ZD domains.
Enables efficient site-specific integration and recombination of DNA sequences in human cells, expanding the repertoire of recognizable sites and enhancing the versatility of recombinases for therapeutic applications.
Smart Images

Figure US2025023078_13112025_PF_FP_ABST
Abstract
Description
Engineered Modular Recombinase Compositions, Methods, and Systems for Site-Specific DNA RecombinationCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U. S. Provisional Patent Application No. 63 / 574,237, filed on April 3, 2024 and entitled “Engineered Recombinase Compositions, Methods, and Systems for Site-Specific DNA Recombination,” U.S. Provisional Patent Application No. 63 / 643,897, filed on May 7, 2024 and entitled “Engineered Recombinase Compositions, Methods, and Systems for Site-Specific DNA Recombination,” and U.S. Provisional Patent Application No. 63 / 644,432, filed on May 8, 2024 and entitled “Engineered Recombinase Compositions, Methods, and Systems for Site-Specific DNA Recombination,” the entire contents of each are incorporated by reference herein.FIELD OF THE INVENTION
[0002] The present invention relates to the field of biotechnology, and more specifically to the field of genomic modification. Disclosed herein are altered recombinases, including compositions thereof, expression vectors, and methods of use thereof, for the generation of transgenic cells, tissues, plants, and animals. The compositions, vectors, and methods of the present invention are also useful in gene therapy techniques.BACKGROUND
[0003] Recombination typically includes binding of a recombinase to a recognition sequence and performing concerted cutting and ligation, resulting in strand exchanges between two recombining recognition sites.
[0004] Large serine recombinases (LSRs)
[0005] Large serine recombinases (LSRs) are a family of enzymes encoded by temperate phages. These enzymes can precisely cut and recombine DNA in a site-specific manner, moving DNA elements into and out of bacterial chromosomes between short DNA attachment sites in the phage (attP) and in the host bacteria (attB) (Duyne and Rutherford, Crit Rev Biochem Mol Biol. (2013) 48(5):476-91; Smith, Microbiol Spectr. (2015) 3(4): doilO.l 128 / microbiolspec). The highly directional and controlled process of DNA recombination mediated by LSRs render them a potentialtool in genetic engineering and gene therapy where site-specific DNA integration, excision, inversion, or cassette exchange in a genome is desired.
[0006] The protein-DNA interaction rules of LSR enzymes for their DNA recombination sites are poorly understood, but the interaction is seen as unlikely to be modular like zinc fingers or TALEs (Fanton et al. doi 10.1101 / 2024.11.01.621560vl). This has caused those in the field to use an inefficient process of LSR engineering to recognize new DNA sequences consisting of i) creating a large collection of recombinase variants with mutations spread throughout the entire coding sequence of the LSR and ii) identifying members from this collection, through multiple rounds of experimental selection, that can recombine a desired DNA sequence (Fanton et al. doi 10.1101 / 2024.11.01.621560vl WO / 2025 / 046147). Using this process, quantifiable levels of integration at endogenous DNA sites in mammalian cells using variant recombinases has only been achieved at pseudo-sites where the corresponding wildtype LSR has detectable activity. An LSR engineering process is likely to be substantially more efficient if the mutations were focused only at the positions in the LSR that directly interact with the target DNA sequence. Additional gains in efficiency could be achieved if pre-existing LSR domain variants could be combined in a modular fashion to create a full-length LSR capable of binding a desired DNA sequence.
[0007] Another way to change the DNA site preference of an LSR may be to fuse it with a DNA targeting domain such as dCas9 or an array of zinc fingers (Fanton et al. www.biorxiv.org / content / 10.1101 / 2024.11.01.621560vl). Fanton et al. fused both engineered and wildtype LSRs with dCas9 and observed improved activity of their LSR-dCas9 fusions with a wide range of orientations between the LSR recombination site and the guide RNA binding site (from less than 20 bp from the center of the LSR recombination site to more than 80 bp from the center of the LSR recombination site) and guide RNAs that bind either strand of DNA. This study found that the size of dCas9 and the required additional guide RNA component are not ideal for therapeutics applications and smaller DNA binding domains that don’t require a guide RNA such as zinc fingers would be preferable.
[0008] Large serine recombinases bind attB and attP sites and comprise a catalytic N-Terminal domain (NTD) and a C-Terminal domain (CTD) involved in sequence-specific DNA recognition. The CTD is further divided into an RD domain and a domain ZD containing a coiled-coil motif (FIG. 3). The NTD contains the catalytic domain (CD).
[0009] Each left DNA halfsite comprises a portion recognized by the LSR ZD domain, and a portion recognized by the RD domain. Each Right DNA halfsite comprises a portion recognized by the ZD domain, and a portion recognized by the LSR RD domain. The loop and helix region are both within the RD domain. The hairpin is within the ZD domain.
[0010] Recombinases can generate a range of dimers that recognize and specifically bind their DNA recombination sites, such as attB and attP attachment sites. When the attB and attP sites are present in the same cell, a pair of recombinase dimers that bind the attB and attP sites form a tetramer, bringing together the DNA segments containing the attB and attP sites and initiate DNA recombination. When the DNA segments are on the same DNA molecule (e.g., chromosome), this recombination event can lead to DNA inversion or excision. When the DNA segments are on different DNA molecules (e.g., a chromosome and a synthetic donor molecule), this recombination event may lead to DNA integration. Two recombination events on two different DNA molecules may lead to cassette exchange where a portion of each DNA molecule is exchanged with the other DNA molecule. The disclosed recombinase variants contain mutations that enable them to catalyze the recombination between non-wildtypeattB and attP sequences, greatly expanding the repertoire of recognizable endogenous sites and increasing the versatility of the enzymes. The disclosed recombinase variants can produce recombination events of interest, such as integration, inversion, excision, translocations, and Recombinase Mediated Cassette Exchange (RMCE). Integration requires the presence of donor DNA with an attachment site that is compatible to the attachment site of the target DNA. Excision requires two complementary attachment sites similarly orientated on the same DNA molecule. Inversion requires two complementary attachment sites oppositely orientated on the same DNA molecule. Translocations require two complementary attachment sites similarly orientated on two separate linear DNA molecules, such as chromosomes. RMCE requires donor and target DNA molecules each contain two complementary attachment sites that are not crosscompatible. For example, a donor DNA molecule containing gene X flanked by two different attB sites and a target DNA molecule containing gene Y flanked by two different attP sites where the compatible attachment sites upstream of gene X and Y are complementary and the compatible attachment sites downstream of genes X and Y are complementary, but the upstream and downstream sites are not cross-compatible. Cross-compatibility can be avoided, for example, by using different center dinucleotides in the upstream and downstream attachment sites. This system allows genes X and Y to be exchanged in the presence of the recombinase variants of the present disclosure. Only in the presence of the appropriate recombination directionality factor can serine recombinases bind to attR and attL sites and mediate the reverse recombination event.
[0011] Bxbl recombinase
[0012] Bxbl recombinase, also known as Bxbl integrase, is an LSR encoded by phage Bxbl that facilitates integration of the phage DNA into the Mycobacterium smegmatis genome (Russell et al., Biotechniques (2018) 40(4): doi.org / 10.2144 / 000112150). Each recombinase monomer contains an N-terminal catalytic domain similar to that of smaller resolvases / invertases and a larger C-terminal domain responsible for coordination of activities unique to LSRs. This C-terminal domain is further divided into a recombinase domain (RD) and a Zinc ribbon domain (ZD) that contains a coiled-coil (CC) motif (Rutherford et al., Nucleic Acids Res. (2013) 41 (17):8341-56). In order for recombination to occur, a pair of Bxbl monomers (a Bxbl dimer) must bind the double-stranded DNA attB attachment site and a second pair of Bxbl monomers (a second Bxbl dimer) must bind the doublestranded DNA attP attachment site. The ZD and RD domains of each Bxbl monomer recognize distinct portions of the attB and attP attachment sites. The subsequent interaction of the two dimers leads to the formation of a Bxbl recombinase tetramer and brings together the attP and attB sites on the two DNA molecules. Serine residues at the tetramer’s active sites form covalent bonds at those center dinucleotide base pairs and this process cleaves the double-stranded DNA and leaves each end of the double-stranded DNA covalently attached to one of the four copies of the Bxbl enzyme in the tetramer until the recombination mechanism is complete. This cleavage makes the center dinucleotides into 3’ overhangs. The resulting protein / DNA complex is rotated, and the previously separate DNA molecules or segments are ligated if the center dinucleotides at the attB and attP sequences are identical. This process produces attachment R (attR) and attachment L (attL) sites that are no longer substrates for the recombinase without the presence of additional co-factors. A key aspect of this process is that the DNA ends are covalently attached to a Bxbl active site serine residue while the reaction is occurring and thus do not trigger / require mammalian DNA repair machinery.
[0013] Despite the precision of gene editing afforded by large serine recombinases, their utility as a genetic engineering and therapeutic tool has been limited by the rarity of their binding sites in the human genome, including in therapeutically relevant genes, and the inability to efficiently engineer them to target desired DNA sequences in the human genome. Thus, there remains a need for new recombinases that can recognize a variety of DNA attachment sites and there remains a need for new methods to efficiently create such new recombinases that can target desired human DNA sequences and thus realize the therapeutic potential of this category of enzymes.SUMMARY
[0014] Disclosed is the generation, identification, isolation, cloning, expression, and methods of use of altered / variant recombinases. In some embodiments, disclosed is a method of site-specifically integrating a polynucleotide sequence of interest in a genome of a target cell using an altered / variant recombinase.
[0015] Variant Recombinases
[0016] Disclosed are variant recombinases. Disclosed is a variant recombinase, wherein the variant recombinase comprises, consists of, or consists essentially of, one or more amino acid mutations in at least two DNA binding (or DNA interacting) regions that control DNA target site specificity for the relevant portions of the target site (recombination site) of a wildtype recombinase, and wherein the variant recombinase has altered DNA target site specificity relative to the cognate wildtype recombinase binding region. In some embodiments, the variant recombinase comprises, or consists of, or consists essentially of at least two mutations relative to the wild-type recombinase in the binding regions. In some embodiments, the variant recombinase, comprises, consists of, or consists essentially of at least two mutations relative to the wild-type recombinase in each one of the binding regions. In some embodiments, the binding regions of the variant recombinase are small enough to allow randomization and re-selection of up to 7 residues within each region to completely alter the DNA target site specificity corresponding to the region. In some embodiments, the binding regions of the variant recombinase comprise a loop region, a helix region, and a hairpin region, and wherein the loop region comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its or their wildtype counterpart, the helix region, comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its or their wildtype counterpart and / or the hairpin region comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to its or their wildtype counterpart. In some embodiments, the hairpin region of the variant recombinase is within a zinc ribbon domain (ZD) region and the ZD region comprises, or consists of, or consists essentially of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to its wildtype counterpart. In some embodiments, the loop region of the variant recombinase is within a recombinase region (RD) and wherein the loop region comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart. In some embodiments, the helix region of the variant recombinase is within a recombinase region (RD) and wherein the helix region comprises 1, 2, 3, 4, 5, 6, or 7 amino acids changes relative to its wildtype counterpart. In some embodiments, the binding regions of the variant recombinase comprise one or more of a variant loop, helix, or hairpin amino acid sequence asshown in any of Figures 5B (helix), 5C(1) (helix), 5C(2) (helix), 5D (helix), 5E (helix), 6A (hairpin), 6B (hairpin), 6C (hairpin), 9C (helix and hairpin), 10C(l) (helix), 10C(2) (helix), 10D (helix), 10E(l) (helix), 10E(2) (helix), 11C (loop), HE (loop), 12D(1) (helix), 12D(2) (helix), 12E(1) (hairpin), 12E(2) (hairpin), 12F(1) (loop), 12F(2) (loop)13B (loop, helix, hairpin), 13D (loop, helix, hairpin), 13E(1) (loop, helix, hairpin), 14B (loop, helix, hairpin), 14C(1) (loop, helix, hairpin), 16 (loop, helix, hairpin), and / or 21 (loop, helix, hairpin). In some embodiments, the binding regions of the variant recombinase recognize one or more of the DNA target sites (including pseudo-sites, attB-like, attP- like target sites, genomic sites, or half-sites) as shown in Figures 5A, 9A, 9B, 10A, 10B(l), 10B(2), 11 A, 11D, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22. In some embodiments, the variant recombinase comprises an asparagine amino acid residue (N) at the 4thposition of the helix region, a tryptophan (W) amino acid residue at the 3rdposition of the helix region, two lysine residues (K) at the 6thand 7thposition of the helix region, or any one or combinations thereof. In some embodiments, the variant recombinase is Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901.
[0017] In some embodiments, the variant recombinase comprises: 1) a helix amino acid sequence as shown in Figs. 5B (helix), 5C(1) (helix), 5C(2) (helix), 5D (helix), 5E (helix), 9C (helix), 10C(l) (helix), 10C(2) (helix), 10D (helix), 10E(l) (helix), 10E(2) (helix), 12D(1) (helix), 12D(2) (helix), 13B (helix), 13D (helix), 13E(1) (helix), 14B (helix), 14C(1) (helix), 16 (helix), or 21 (helix); and / or 2) a loop amino acid sequence as shown in 11C (loop), 1 IE (loop), 12F(1) (loop), 12F(2) (loop), 13B (loop), 13D (loop), 13E(1) (loop), 14B (loop), 14C(1) (loop), 16 (loop), or 21 (loop); and / or 3) a hairpin amino acid sequence as shown in Figs. 6A (hairpin), 6B (hairpin), 6C (hairpin), 9C (hairpin), 12E(1) (hairpin), 12E(2) (hairpin), 13B (hairpin), 13D (hairpin), 13E(1) (hairpin), 14B (hairpin), 14C(1) (hairpin), 16 (hairpin), or 21 (hairpin). In some embodiments, a variant recombinase comprises a helix, hairpin, and / or loop sequence, wherein the helix, hairpin, and / or loop sequence binds to a DNA target sequence as shown for a helix, hairpin and / or loop sequence in Figures 5B, 5C(1), 5C(2), 5D, 5E, 6A, 6B, 6C, 9C, 10C(l), 10C(2), 10D, 10E(l), 10E(2), 11C, HE, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13B, 13D, 13E(1), 14B, 14C(1), 16, and / or 21..
[0018] In some embodiments, the variant recombinase comprises an amino acid sequence YRGGLP in its loop region, an amino acid sequence AGGNLKR, YPWSLRR, SQWALKC, SGWALKC, YGSALKA, or YGSALKQ in its helix region, an amino acid sequenceRAWGKRKYGYYQ, RAWGKRKYAYYI, RAWGKRKYAYYQ, RAWGKRKYAYYL,LARGGRKRAGYK, LARGVRKRAGYK, or LARGPRKRAGYK in its hairpin region,
[0019] In some embodiments, a variant recombinase as described herein comprises loop, helix, and / or hairpin sequences that bind to or recognize a target sequence as described herein (including pseudo-sites, attB-like, attP-like target sites, genomic sites, or halfsite), for example, as shown in Figures 5A, 9A, 9B, 10A, 10B(l), 10B(2), 11 A, 11D, HE, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22. In some embodiments, a variant recombinase comprises a helix sequence as indicated used to target or bind a DNA sequence as indicated: amino acid sequence AAWALRR or ASHALKR binds DNA bases CAC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence AVQNLKR binds DNA bases TTC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence GGRSLKR binds DNA bases AGC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence GGSHLKR binds DNA bases GCC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence LGTNLKR binds DNA bases ATC at positions -11 to -9; helix amino acid sequence RAAFLKK binds DNA bases ACA at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence RADTLRR binds DNA bases CGC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence RAWTLKC binds DNA bases CTG at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence RGHALKN binds DNA bases ACT at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence RGSSLKV binds DNA bases ATT at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SARALSR binds DNA bases GAC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SGSALKT binds DNA bases AAT at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SGWALRQ binds DNA bases CAT at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SGWGLKK binds DNA bases CAA at positions -11 to - 9 of the attB or attP DNA target; helix amino acid sequence SGYNLRR binds DNA bases CTC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SRNGLRK binds DNA bases GAA at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence TTRTLKR binds DNA bases GGC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence YSRNLKR binds DNA bases GTC at positions -11 to -9 of the attB or attP DNA target; or helix amino acid sequence SGTGLKK binds DNA bases AAA, CAA, or GAA at positions -11 to -9 of the attB or attP DNA target.
[0020] In one embodiment, the variant recombinase comprises an amino acid sequence YRGGLP in the loop region, an amino acid sequence YGSALKQ in the helix region, and an amino acid sequence LARGPRKRAGYK in the hairpin region.
[0021] In one embodiment, the variant recombinase comprises an amino acid sequence YRGGLP in the loop region, an amino acid sequence YGSALKQ in the helix region, and an amino acid sequence LARGGRKRAGYK in the hairpin region.
[0022] In one embodiment, the variant recombinase comprises an amino acid sequence YRGGLP in the loop region, an amino acid sequence YGSALKQ in the helix region, and an amino acid sequence LARGVRKRAGYK in the hairpin region.
[0023] In one embodiment, the variant recombinase comprises an amino acid sequence YRGGLP in the loop region, an amino acid sequence SQWALKC in the helix region, and an amino acid sequence RAWGKRKYAYYQ in the hairpin region.
[0024] In one embodiment, the variant recombinase comprises an amino acid sequence YRGGLP in the loop region, an amino acid sequence SQWALKC in the helix region, and an amino acid sequence RAWGKRKYAYYI in the hairpin region.
[0025] In one embodiment, the variant recombinase comprises an amino acid sequence YRGGLP in the loop region, an amino acid sequence SQWALKC in the helix region, and an amino acid sequence RAWGKRKYAYYL in the hairpin region.
[0026] In one embodiment, the variant recombinase comprises an amino acid sequence YRGGLP in the loop region, an amino acid sequence SGWALKC in the helix region, and an amino acid sequence RAWGKRKYAYYQ in the hairpin region.
[0027] In one embodiment, a pair of variant recombinases recognizes an attB site with the DNA sequence TGGTGTCCAGGAGCCGAGGTATCGGTCCTGCCAGGGCC (Seq ID No. 18) in the human TRAC locus with genomic coordinates hg38:chrl4:22549434-22549471, wherein the first variant recombinase recognizes the left halfsite of Seq ID No. 18 and the second variant recombinase recognizes the right half site of Seq ID NO. 18.
[0028] In one embodiment, a pair of variant recombinases recognizes an attB site with the DNA sequence CTGAGCGCCTCTCCTGGGCTTGCCAAGGACTCAAACCC (Seq ID No. 17) in the human AAVS1 locus with genomic coordinates hg38:chrl9:55115013-55115050, wherein the firstvariant recombinase recognizes the left halfsite of Seq ID No. 17 and the second variant recombinase recognizes the right half site of Seq ID NO. 17..
[0029] In one embodiment, a variant recombinase recognizes both halfsites of an attB site with the DNA sequence TGGTGTCCAGGAGCCGAGGTATCGGTCCTGCCAGGGCC (Seq ID No. 18) in the human TRAC locus with genomic coordinates hg38:chrl4:22549434-22549471.
[0030] In one embodiment, a variant recombinase recognizes both half sites of an attB site with the DNA sequence CTGAGCGCCTCTCCTGGGCTTGCCAAGGACTCAAACCC (Seq ID No. 17) in the human AAVS1 locus with genomic coordinates hg38:chrl9:55115013-55115050.
[0031] In some embodiments, the variant recombinase binds to a DNA target site sequence (including pseudo-sites, attB-like, attP-like target sites, genomic sites, or halfsite) as shown in Figures 5A, 9A, 9B, 10A, 10B(l), 10B(2), 11 A, HD, HE, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22. In some embodiments, the variant recombinase binds to a target site comprising, or consisting of, or consisting essentially of: TGGTGTCCAGGAGCCGAGGTATCGGTCCTGCCAGGGCC (Seq ID No. 18). In some embodiments, the variant recombinase binds to a target site comprising, or consisting of, or consisting essentially of: CTGAGCGCCTCTCCTGGGCTTGCCAAGGACTCAAACCC (Seq ID No. 17).
[0032] In some embodiments, the variant recombinase targets AAVS1 or TRAC. In some embodiments, a method of targeting AAVS1 or TRAC is disclosed, comprising recognizing TGGTGTCCAGGAGCCGAGGTATCGGTCCTGCCAGGGCC (Seq ID No. 18) or CTGAGCGCCTCTCCTGGGCTTGCCAAGGACTCAAACCC (Seq ID No. 17) with a modified recombines as described herein.
[0033] In some embodiments, the variant recombinase further comprises a zinc finger array fused to the C-terminus of the variant recombinase using a polypeptide linker, wherein the zinc finger array recognizes a DNA sequence, and wherein the 3 ’ edge of the zinc finger array DNA target site is separated by 4, 5, 6, 7, 8, 9, or 10 basepairs from the edge of the recombination site for the variant recombinase, site for the variant recombinase.
[0034] Also disclosed are nucleic acid molecules encoding the variant recombinases disclosed herein. Also disclosed are vectors comprising, or consisting of, or consisting essentially of, the variant recombinases disclosed herein. In come embodiments, a recombinant virus comprising anucleic acid construct encoding the variant recombinases disclosed herein in provided, optionally wherein the recombinant virus is a recombinant AAV. In some embodiments, host cell comprising the nucleic acid construct encoding the variant recombinases disclosed herein in provided
[0035] Bxbl Variant Recombinases
[0036] In some embodiments, the variant recombinase is a Bxbl variant recombinase. In some embodiments, the Bxbl variant recombinase comprises an asparagine amino acid residue (N) at position 234 of the integrase, a tryptophan (W) amino acid at position 233 of the integrase, two lysines (K) at positions 236 and 237 of the integrase, or any one or combinations thereof. In some embodiments, the Bxbl variant recombinase comprises an amino acid sequence YRGGLP in the loop region instead of an amino acid sequence YRGSLP from wildtype Bxbl, an amino acid sequence AGGNLKR, YPWSLRR, SQWALKC, SGWALKC, YGSALKA, or YGSALKQ in the helix region instead of an amino acid sequence SATALKR from wildtype Bxbl, an amino acid sequence RAWGKRKYGYYQ, RAWGKRKYAYYI, RAWGKRKYAYYQ, RAWGKRKYAYYL, LARGGRKRAGYK, LARGVRKRAGYK, or LARGPRKRAGYK in the hairpin region instead of an amino acid sequence FAGGGRKHPRYR from wildtype Bxbl, a lysine (K) at position 257 of the integrase, or any one or combination thereof.
[0037] In some embodiments, a variant recombinase as described herein comprises loop, helix, and / or hairpin sequences that bind to or recognize a target sequence as described herein (including pseudo-sites, attB-like, attP-like target sites, genomic sites, or halfsite), for example, as shown in Figures 5A, 9A, 9B, 10A, 10B(l), 10B(2), 11 A, HD, HE, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22. In some embodiments, a variant recombinase comprises a helix sequence as indicated used to target or bind a DNA sequence as indicated: amino acid sequence AAWALRR or ASHALKR binds DNA bases CAC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence AVQNLKR binds DNA bases TTC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence GGRSLKR binds DNA bases AGC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence GGSHLKR binds DNA bases GCC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence LGTNLKR binds DNA bases ATC at positions -11 to -9; helix amino acid sequence RAAFLKK binds DNA bases ACA at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence RADTLRR binds DNA bases CGC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence RAWTLKC binds DNA bases CTG at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence RGHALKN binds DNA bases ACT at positionsATT at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SARALSR binds DNA bases GAC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SGSALKT binds DNA bases AAT at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SGWALRQ binds DNA bases CAT at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SGWGLKK binds DNA bases CAA at positions -11 to - 9 of the attB or attP DNA target; helix amino acid sequence SGYNLRR binds DNA bases CTC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence SRNGLRK binds DNA bases GAA at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence TTRTLKR binds DNA bases GGC at positions -11 to -9 of the attB or attP DNA target; helix amino acid sequence YSRNLKR binds DNA bases GTC at positions -11 to -9 of the attB or attP DNA target; or helix amino acid sequence SGTGLKK binds DNA bases AAA, CAA, or GAA at positions -11 to -9 of the attB or attP DNA target.
[0038] In one embodiment the variant recombinase comprises an amino acid sequence YRGGLP in the loop region instead of an amino acid sequence YRGSLP from wildtype Bxbl, an amino acid sequence YGSALKQ in the helix region instead of an amino acid sequence SATALKR from wildtype Bxbl, a lysine (K) at position 257 of the LSR, and an amino acid sequence LARGPRKRAGYK in the hairpin region instead of an amino acid sequence FAGGGRKHPRYR from wildtype Bxbl .
[0039] In one embodiment the variant recombinase comprises an amino acid sequence YRGGLP in the loop region instead of an amino acid sequence YRGSLP from wildtype Bxbl, an amino acid sequence SQWALKC in the helix region instead of an amino acid sequence SATALKR from wildtype Bxbl, and an amino acid sequence RAWGKRKYAYYQ in the hairpin region instead of an amino acid sequence FAGGGRKHPRYR from wildtype Bxbl.
[0040] In one embodiment, a pair of variant recombinases collectively recognize an attB site with the DNA sequence TGGTGTCCAGGAGCCGAGGTATCGGTCCTGCCAGGGCC (Seq ID No. 18) in the human TRAC locus with genomic coordinates hg38:chrl4:22549434-22549471.
[0041] In one embodiment, a pair of variant recombinases collectively recognize an attB site with the DNA sequence CTGAGCGCCTCTCCTGGGCTTGCCAAGGACTCAAACCC (Seq ID No. 17) in the human AAVS1 locus with genomic coordinates hg38:chrl9:55115013-55115050.
[0042] In other preferred embodiments the helix and / or loop region is engineered to recognize positions -11 to -1 or positions +1 to +11 of a genomic attB target and separately the hairpin is engineered to recognize positions -19 to -12 or positions +12 to +19 of the genomic attB and then the CD and RD domain comprising the variant helix and loop sequences is combined with the ZD domain comprising the variant hairpin to create a variant recombinase with variant hairpin and variant helix and / or variant loop. The exact boundary between the ZD and RD domains is not critical as long as it is understood that the hairpin is in the ZD domains and both the loop and helix are in the RD domains. The engineered helix and / or loop within the RD domain may be engineered to bind the relevant portion of a desired genomic attB recombination site or it may be chosen from a pre-existing “archive” of variants characterized well enough to predict that the pre-existing variant will be able to recognize the relevant portion of the desired genomic attB recombination site. The engineered hairpin within the ZD domain may be engineered to bind the relevant portion of a desired genomic attB recombination site or it may be chosen from a pre-existing “archive” of variants characterized well enough to predict that the pre-existing variant will be able to recognize the relevant portion of the desired genomic attB recombination site.
[0043] In some embodiments, The variant recombinase comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to wildtype Bxbl at positions 308, 309, 310, 311, 312, 313, 314, 316, 318, 321, 322, 323, and 325 in the hairpin region or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to wildtype Bxbl at positions 154, 155, 156, 157, 158, and 159 in the loop region, wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to wildtype Bxbl at positions 231, 232, 233, 234, 236, and 237 of the helix region, or any one or combinations thereof. In some embodiments, the Bxbl variant recombinase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to wildtype Bxbl at positions 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, and 325 in the hairpin region.
[0044] Methods of Making Variant Recombinases
[0045] Disclosed is a method for generating a large serine recombinase variant having an altered DNA target specificity, comprising: 1) identifying positions of amino acids that control DNA target specificity in a starting recombinase; 2) randomizing the positions of amino acids that control DNA target specificity to generate recombinase variants; and 3) selecting a recombinase variant having the altered DNA specificity.
[0046] Disclosed are methods for generating a large serine recombinase variant that can recognize a desired DNA sequence, comprising: 1) Dividing the desired genomic DNA target site into 4 separate portions that correspond to the region recognized by i) the left halfsite ZD domain, ii) the left halfsite RD domain, iii) the right halfsite RD domain and iv) the right halfsite ZD domain 2) for each portion of the desired target site recognized by a ZD domain, generating a library of recombinase variants with at least two amino acid changes within the hairpin region of the recombinase and selecting this library for members that can recombine a target site comprising the relevant portion of the desired genomic target site 3) for each portion of the desired target site recognized by an RD domain, generating a library of recombinase variants with at least two amino acid changes with the helix and / or loop regions, and selecting members of this library that can recombine target sites comprising the relevant portion of the genomic target site 4) combining selected recombinase variants selected against two regions of the same halfsite and then testing such composite variants with amino acid changes in both RD and ZD domains for recombination activity with a target site comprising the relevant desired halfsite 5) delivering a mixture of a recombinase variant active on the desired left halfsite, a recombinase variant active on the right halfsite, and an appropriate donor construct into eukaryotic cells and then measuring the amount of targeted integration of the donor construct into the desired genomic location.
[0047] Modular Combinatorial Generation of Novel Variant Recombinases
[0048] The modified RD and ZD domains can be combined together in any modular fashion. For example, disclosed are variant recombinases comprising RD and / or ZD domains from different recombinases. In some embodiments, the RD and / or ZD domains are from a different wildtype recombinase. In some embodiments, the RD and / or ZD domains are from a different modified recombinase. In some embodiments, the RD and / or ZD domains comprise one or more amino acid changes relative to the wildtype counterpart domains.
[0049] Methods of Using Variant Recombinases
[0050] In some embodiments, a method of integrating a synthetic DNA donor construct into a non-coding region, a safe harbor locus, or into or near a gene of interest is provided, comprising expressing a variant recombinase described herein in a cell. In some embodiments, the method comprises expressing a first and a second variant recombinase in a cell, wherein the first and the second variant recombinase is each according to a recombinase as described herein, wherein the first variant recombinase recognizes a left halfsite of a genomic attB DNA site and the second variantrecombinase recognizes a right halfsite of the same genomic attB DNA site, alternatively wherein the first variant recombinase recognizes a left halfsite of a genomic attP DNA site and the second variant recombinase recognizes a right halfsite of the same genomic attP DNA site. In some embodiments, the first and the second variant recombinases are cotransfected with a donor construct with an attP site, wherein the variant recombinase that recognizes the left halfsite of a genomic attB site recognizes both halfsites of the attP site on the donor construct, alternatively wherein the variant recombinase that recognizes the left halfsite of a genomic attP site recognizes both halfsites of the attB site on the donor construct. In some embodiments, the first and the second variant recombinases are cotransfected with a donor construct with an attP site, wherein the variant recombinase that recognizes the right halfsite of a genomic attB site recognize both halfsites of the attP site on the donor construct, alternatively wherein the variant recombinase that recognizes the right halfsite of a genomic attP site recognize both halfsites of the attB site on the donor construct. In some embodiments, the first and the second variant recombinases are cotransfected with a donor construct with an attP site, wherein the variant recombinase that recognizes the left halfsite of a genomic attB site recognize a first halfsites of the attP site on the donor construct and the variant recombinase that recognizes the right halfsite of a genomic attB site recognize the second halfsite of the attP site on the donor construct, alternatively wherein the first and the second variant recombinases are cotransfected with a donor construct with an attB site, wherein the variant recombinase that recognizes the left halfsite of a genomic attP site recognize a first halfsites of the attB site on the donor construct and the variant recombinase that recognizes the right halfsite of a genomic attP site recognize the second halfsite of the attB site on the donor construct.
[0051] In some embodiments, a method of expressing or repressing a therapeutically or industrially relevant gene of interest in a cell is provided, comprising introducing to the cell a nucleic acid as described herein, wherein the nucleic acid encodes a variant recombinase as described herein.
[0052] In some embodiments, a method of treating a disease in a patient is provided, comprising administering to the patient a variant recombinase as described herein.
[0053] In some embodiments, use of a variant recombinase, a nucleic acid construct, or a recombinant virus as described herein is provided, for the manufacture of a medicament in the method as described herein.
[0054] Directed Evolution System
[0055] Disclosed is a method for generating a recombinase variant that can recognize a desired DNA sequence comprising: 1) generating a library of variant recombinases from a starting recombinase, wherein the variant recombinases include on or more modifications of amino acids that control DNA target specificity for the starting recombinase 2) screening the variant recombinase library for altered DNA specificity as compared to the starting recombinase; and 3) selecting a variant recombinase from 2.
[0056] Unintegrated Plasmid Removal
[0057] Disclosed is a modified plasmid donor comprising at least two Dpnl sites, wherein the Dpnl sites are placed within close proximity on both sides of a DNA target site, preferably at least one Dpnl site within 100 bp 5’ of the recombinase target site and at least one Dpnl site within 100 bp 3’ of the recombinase target site. In some embodiments, the DNA target site of the modified plasmid donor is for a recombinase or nuclease.
[0058] Disclosed is a method for unintegrated plasmid removal, comprising: 1) administering a modified plasmid donor construct comprising at least two Dpnl sites; 2) allowing the administered modified plasmid donor construct comprising at least two Dpnl sites to combine with a target gene in a cell; 3) exposing the cell comprising the recombinase variant to restriction enzyme Dpnl thereby digesting any unincorporated plasmid. This method of unintegrated plasmid removal increases integrated plasmid signal in a ligation-mediated PCR-based assay. As such, this method allows for potential integration site identification and genome-wide mapping.
[0059] Disclosed is a method for genome-wide mapping of plasmid integration sites in a genome, the method comprising: 1) generating a modified plasmid donor as described herein; 2) combining the modified plasmid donor with the genome in the presence of a reagent that facilitates site-specific integration into the genome to create a mixture; 3) extracting and purifying DNA from the mixture; 4) exposing the purified DNA to a Dpnl enzyme to generate a DpnI-digested sample; 4) sequencing the DpnI-digested sample; and 5) using the sequencing results from step 4) to identify the plasmid integration sites. In some embodiments, the reagent that facilitates site-specific integration is a recombinase or nuclease.
[0060] Antibiotic Resistant Plasmids or Viral Vectors
[0061] Disclosed is a split antibiotic resistance gene comprising a recombinase target sequence, preferably anattP and an attB sequence. In some embodiments, the split antibiotic resistance geneencodes an aaCCl protein. In some embodiments, the target sequence of the split antibiotic resistance gene is inserted into a surface-exposed loop of the protein. In some embodiments, target sequence of the split antibiotic resistance gene is inserted between the following residues: 83-84.
[0062] Disclosed is a method for selecting or identifying a recombinase having activity against a target sequence, comprising: 1) generating a split antibiotic resistance gene as described herein; 2) introducing the gene into bacterial cells; 3) introducing into the bacterial cells a recombinase from a pool of active and inactive recombinases, wherein each bacterial cell preferentially expresses one or more types of recombinase; 4) administering an antibiotic to the bacterial cells; and 5) identifying the recombinase in surviving bacterial cells.BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Fig. 1A shows integrase-mediated integration of a circular double-stranded DNA donor construct into the genome of a eukaryotic cell.
[0064] Fig. IB is a diagram of the synaptic complex with an integrase dimer bound to the attP site and second integrase dimer bound to the attB site.
[0065] Fig. 1C shows the wildtype attB and attP sequences recognized by Bxbl
[0066] Fig. 2 shows a natural Bxbl protein sequence and annotated wildtype Bxbl attB target sequence.
[0067] Fig. 3 is a schematic showing the strategy for engineering Bxbl variants against a desired target site.
[0068] Fig. 4 is a schematic showing an overview of the directed evolution system used herein.
[0069] Fig. 5A is a comparison of wildtype Bxbl attB sequence and site 1-41 pseudo attB sequence in chromosome 3 of the reference human genome (hg38).
[0070] Fig. 5B shows sequence information for Bxbl directed evolution helix selections.
[0071] Fig. 5C(1) shows plots of directed evolution of Bxbl helix using a wildtype attB target site.
[0072] Fig. 5C(2) shows plots of directed evolution of Bxbl helix using a portion of site 1-41 pseudo attB target sequence.
[0073] Fig. 5C(3) shows plots of a control directed evolution of Bxbl helix without the antibiotic that selects for cells containing recombined target plasmids.
[0074] Fig. 5D shows activity of Bxbl variants with selected helices in human cells.
[0075] Fig. 5E shows comparison of molecular specificity of wild-type Bxbl vs. helix variants.
[0076] Fig. 6A shows hairpin sequence from wildtype Bxbl and the randomization scheme used for the hairpin region.
[0077] Fig. 6B shows a sequence logo of hairpin sequences selected to recognize a portion of site 1-41 pseudo attB target site.
[0078] Fig. 6C shows the DNA target used in the hairpin selection and amino acid sequence of selected hairpin variant ZD32.
[0079] Fig. 6D shows the activity of the selected ZD32 hairpin variant and activity of Bxbl variant with a combination of a hairpin and helix variants.
[0080] Fig. 7A is a schematic of the assay used to characterize genome-wide specificity of Bxbl variants.
[0081] Fig. 7B(1) shows genome-wide specificity results for Bxbl with wild-type helix and hairpin.
[0082] Fig. 7B(2) shows genome-wide specificity of Bxbl with altered helix and hairpin intended to recognize the right halfsite of site 1-41 on human chromosome 3.
[0083] Fig. 7B(3) shows a comparison of the modified assay shown in Figure 7A with and without the Dpnl treatment.
[0084] Fig. 8 shows a recombined gentamicin resistance gene.
[0085] Fig. 9A shows additional Bxbl pseudo attB sites within the reference human genome.
[0086] Fig. 9B shows the attB and attP sites used in the directed evolution hairpin selections to recognize the sites within the reference human genome.
[0087] Fig. 9C shows the integration activity in human cells at the indicated full target site of the indicated Bxbl variant with helices and hairpins selected to recognize the indicated halfsite.
[0088] Fig. 9D shows the integration activity in human cells at the indicated full target site of the indicated mixture of Bxbl variants.
[0089] Fig. 10A shows schematic of target sites used for selecting Bxbl helices against 64 different 3 bp DNA targets.
[0090] Fig. 10B shows full attB and attP target sites used for systematic helix selections.
[0091] Fig. 10C(l) shows amino acid motifs enriched when selecting Bxbl helices against the indicated 3 bp DNA targets.
[0092] Fig. 10C(2) shows amino acid motifs enriched when selecting Bxbl helices against the indicated 3 bp DNA targets.
[0093] Fig. 10D shows molecular specificity of the indicated Bxbl helix variants.
[0094] Fig. 10E(l) shows selected helices that recognize DNA with indicated sequence at -11 to -9.
[0095] Fig. 10E(2) shows selected helices that recognize DNA with indicated sequence at -11 to -9.
[0096] Fig. 11 A shows schematic of target sites used for directed evolution of Bxbl loop.
[0097] Fig. 1 IB shows the randomization scheme used to build the Bxbl variant library used in the loop selections.
[0098] Fig. 11C shows examples of selected Bxbl loop sequences.
[0099] Fig. 1 ID shows target sites used for systematic directed evolution of Bxbl loop.
[0100] Fig. 1 IE shows molecular specificity data for the indicated selected loop sequences.
[0101] Fig 12A shows how the plasmid screening system is used as part of the step-wise process to generate Bxbl variants that can recognize a desired genomic target sequence.
[0102] Fig. 12B shows an overview of the DNA sequences used in the plasmid screening system in human K562 cells.
[0103] Fig. 12C shows an example of how different portions of an endogenous DNA target site and the wildtype Bxbl attB site are used to test activity of a given Bxbl variant against halfsites and quarter sites of a given full genomic target site (attB site).
[0104] Fig. 12D(1) shows helices with improved activity vs. wild-type Bxbl for the indicated target site in human cells.
[0105] Fig. 12D(2) shows helices with improved activity vs. wild-type Bxbl for the indicated target site in human cells.
[0106] Fig. 12E(1) shows Bxbl hairpin variants that show improved activity vs. wild-type Bxbl for the indicated target site in human cells.
[0107] Fig. 12E(2) shows Bxbl hairpin variants that show improved activity vs. wild-type Bxbl for the indicated target site in human cells.
[0108] Fig. 12F(1) shows Bxbl loop variants that show improved activity vs. wild-type Bxbl for the indicated target site in human cells.
[0109] Fig. 12F(2) shows Bxbl loop variants that show improved activity vs. wild-type Bxbl for the indicated target site in human cells.
[0110] Fig. 13A compares the natural Bxbl attB target site to the endogenous AAVS1 5032 attB site.
[0111] Fig. 13B shows loops, helices, and hairpins with improved recognition of the indicated AAVS1 5032 quarter sites vs. the loop, helix, or hairpin from wild-type Bxbl.
[0112] Fig. 13C shows target sites used for AAVS1 hairpin selections.
[0113] Fig. 13D shows activity relative of wild-type Bxbl of Bxbl variants with the indicated loop, helix, and hairpin sequences when tested against the indicated half-site of AAVS1 5032.
[0114] Fig. 13E(1) shows Loop, helix, and hairpin sequences for the Bxbl variants that were combined and tested in human K562 cells to yield the data shown in Fig. 13E(2).
[0115] Fig. 13E(2) shows the results of a PCR-based assay to characterize targeted integration.
[0116] Fig. 13F shows other sites in AAVS1 that can be recognized with this method.
[0117] Fig. 14A compares the natural Bxbl attB target site to the Bxbl pseudo site in the human TRAC locus.
[0118] Fig. 14B shows integration into the human TRAC locus of the indicated Bxbl variants.
[0119] Fig. 14C(1) shows additional Bxbl variants intended to recognize the left or right halfsite of the attB site in the human TRAC locus.
[0120] Fig. 14C(2) shows integration activity at the endogenous human TRAC locus using bxbl variants described in Figure 14C(1).
[0121] Fig 15A shows how additional attB and attP sites can be placed in linear DNA donor so that an integrase can cause circularization of the liner donor construct.
[0122] Fig. 15B Shows the results of tested different donor constructs in human cells with the wildtype Bxbl attB target pre-integrated into the human cells to generate a “landing pad” cell line.
[0123] Fig. 15C(1) strategy to convert Bxbl target site in single-stranded DNA donor construct into double-stranded DNA by annealing an oligo.
[0124] Fig. 15C(2) targeting integration of ssAAV + oligo donor results.
[0125] Fig. 16 shows integration data at the endogenous AAVS1 5032 locus for donor constructs with variant Bxbl target sites tested in human K562 cells.
[0126] Fig. 17A shows annotated amino acid sequence for Theia integrase.
[0127] Fig. 17B shows annotated amino acid sequence for Veracruz integrase.
[0128] Fig. 17C shows annotated amino acid sequence for Kp03 integrase.
[0129] Fig. 17D shows annotated amino acid sequence for PaOl integrase.
[0130] Fig. 17E shows annotated amino acid sequence for Nm60 integrase.
[0131] Fig. 17F shows annotated amino acid sequence for Si74 integrase.
[0132] Fig. 17G shows annotated amino acid sequence for Bcylnt integrase.
[0133] Fig. 17H shows annotated amino acid sequence for Bcelnt integrase.
[0134] Fig. 171 shows annotated amino acid sequence for Ssclnt integrase.
[0135] Fig. 17J shows annotated amino acid sequence for Ssalnt integrase.
[0136] Fig. 17K shows a structural alignment of Bxbl, PaOl, Kp03, Nm60, Si74, Bcylnt, andSsclnt
[0137] Fig. 18 shows a schematic for single stranded circular DNA as donor for a serine integrase.
[0138] Fig. 19A shows the genome-wide integration data for s5-6L and s5-6R
[0139] Fig. 19B shows the genome-wide integration data for s5-l IL and s5- 11R
[0140] Fig. 20 shows a diagram of a pair of Bxbl variant-fusions bound to the TRAC attB site
[0141] Fig. 21 shows the peptide linker sequences tested and an example integrase-ZFP fusion
[0142] Fig. 22 shows the alignment of the zinc finger binding sites tested
[0143] Fig. 23 A shows the results of testing zinc fingers fused to TRAC Bxbl variants
[0144] Fig. 23B shows key results from figure 23 A as a table
[0145] Fig. 24 shows the results of testing combinations of two zinc finger fusions
[0146] Fig. 25 shows the results of testing Bxbl variants fused to zinc fingers in T cellsDETAILED DESCRIPTION OF THE INVENTION
[0147] Throughout this application, various publications, patents, and published patent applications are referred to by an identifying citation. The disclosures of these publications, patents, and published patent specifications referenced in this application are hereby incorporated by reference into the present disclosure to more fully describe the state of the art to which this invention pertains.
[0148] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of molecular biology, microbiology, cell biology and recombinant DNA, which are within the skill of the art. See, e g., Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2ndedition (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, (F. M. Ausubel et al. eds., 1987); the series METHODS IN ENZYMOLOGY (Academic Press, Inc ); PCR 2: A PRACTICAL APPROACH (M I McPherson, B. D. Hames and G. R. Taylor eds., 1995) and ANIMAL CELL CULTURE (R. I. Freshney. Ed., 1987).
[0149] Definitions
[0150] It must also be noted that, as used in the specification and the appended claims, the singular forms “a,” “an” and “the” include plural referents such as “at least one” or “one or more” unless the context clearly dictates otherwise. By “comprising” or “containing” or “including” or “has,” or “having,” it is meant that at least the named compound, element, particle, or method step is present in the composition or article or method, but does not exclude the presence of other compounds, materials, particles, method steps, even if the other such compounds, material, particles, method steps have the same function as what is named. By “consisting of’ it is meant that at the named compound, element, particle, or method step is present in the composition or article or method, and does not include the presence of other compounds, materials, particles, method steps. While some embodiments comprise / include the disclosed features and may therefore include additional features not specifically described, other embodiments may be essentially free of or completely free of non-disclosed elements - that is, non-disclosed elements may optionally be essentially omitted or completely omitted.
[0151] In this disclosure, relative terms, such as “about,” “substantially,” or “approximately” are used to indicate a possible variation of ±10% in the stated value.
[0152] In describing example embodiments, terminology will be resorted to for the sake of clarity. It is intended that each term contemplates its broadest meaning as understood by those skilled in the art and includes all technical equivalents that operate in a similar manner to accomplish a similar purpose. It is also to be understood that the mention of one or more steps of a method does not preclude the presence of additional method steps or intervening method steps between those steps expressly identified. Steps of a method may be performed in a different order than those described herein without departing from the scope of the disclosed technology. Similarly, it is also to beunderstood that the mention of one or more components in a device or system does not preclude the presence of additional components or intervening components between those components expressly identified.
[0153] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the systems, devices and methods for performing multianalyte detection in a biological sample belongs.
[0154] attB target site means a DNA sequence about 38 bp in length that is recognized by an LSR when the RD and ZD domains recognize DNA subsites adjacent to each other and the flexible linker between the RD and ZD domains is not in the extended conformation. An attB target site can be recognized by a wildtype recombinase or a variant recombinase. One LSR monomer binds to the 5’ most 19 bp of an attB target and the second LSR monomer binds to the 3’ most 19 bp of an attB target. The center 2 bp of an attB must match the Core Dinucleotide (CDN) of an attP sequence for recombination to occur. attB sequences for two copies of the same LSR form a pseudo inverted repeat where the 5’ 19 bp are similar to the reverse-complement of the 3’ 19 bp. Fig. IB shows a diagram of a Bxbl dimer bound to the wildtype Bxbl attP target site and a second Bxbl dimer bound to the wildtype Bxbl attP target site.
[0155] attP target site means a DNA sequence about 48 bp in length that is recognized by an LSR when the RD and ZD domains recognize DNA subsites separated by 5 base pairs and the linker between the RD and ZD domains is in the extended conformation. There are 5 base pairs between the portion of attP sites recognizes by the hairpin region and the portion of the attP site recognized by the helix while the region recognized by the hairpin and the region recognized by the helix are adjacent to each other in attB sites. This spacing difference is a key differentiator between attP and attB sites. Fig. IB shows a diagram of a Bxbl dimer bound to the wildtype attP target site and a second Bxbl dimer bound to the wildtype attP target site. AttB and attP target sites for the same wildtype LSR often have at least some sequence homology between positions -12 to +12 of the attB site and positions -12 to +12 of the attP sites; this corresponds to the portion of the target site recognized by the RD domain. AttB and attP target sites for the same wildtype LSR may have at least some sequence homology between positions -19 to -13 of the attB site and positions -24 to -18 of the attP site and have at least some sequence homolgoy between positions +13 to +19 of the attB site and positions +18 to +24 of the attP sites; this corresponds to the portion of the target site recognized by the ZD domain.
[0156] CDN (also referred to as core dinucleotide, and center 2 bp) is the center two base pairs (center dinucleotide) of each attB and attP target sites. The CDN of an attB target and attP target site must either match each other or the CDN of an attB target must match the reverse complement of the CDN in the attP target site for recombination to occur between an attB and attP site. These two situations will result in the attP and attB sites being recombined in different orientations. The positions of the CDN are shown in Fig IB and in Fig. 1C.
[0157] Hairpin means the region of an LSR that is likely to form a beta hairpin and primarily interacts with bases at or about positions -19 to -12 of attB halfsites and primarily interacts with bases at or about positions -24 to -17 of attP halfsites. For the Bxbl, the hairpin region involved in sequence specific DNA interactions spans amino acid residues 308 to 325 with key specificitydetermining residues at positions 314, 316, 318, 321, 322, 323, and 325. For Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, and Ssalnt the hairpin region is residues 307-326, 320- 343, 306-323, 286-302, 322-339, 317-343, 289-313, 333-350, 296-315, and 283-311 respectively. The key DNA contacting residues for Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, and Ssalnt the hairpin region are (313, 312, 315,317, 322, 324, 326), (326,328,330, 332, 339, 341, 343), (311, 312, 314, 315, 318, 321, 323), (290, 292, 294, 297, 299, 300, 302), (328, 330, 333, 334, 335, 337, 339), (323, 325, 328, 326, 339, 341, 343), (295, 297, 298, 299, 307, 311, 313), (337, 339, 341, 345, 346, 348, 350), (302, 304, 310,311, 313, 315), and (289, 291, 293, 295, 305,307, 309) respectively.
[0158] Halfsite means the portion of an attB or attP site bound by a single LSR monomer.
[0159] Helix means the region of an LSR RD is the portion of the putative alpha helix that primarily interacts with bases at or about positions 11 to 9 in each attP and attB halfsite. For Bxbl, the helix region is from residue 231 to 237 with key DNA-interacting residues at positions 231, 232, 233, 234, 236, and 237. For Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, and Ssalnt the helix regions are 229-235, 241-247, 241-247, 221-227, 230-236, 234-240, 218-224, 241- 247, 223-228, and 218-224 respectively. The fifth residue in the helix region is likely to face away from the DNA and thus isn’t a key DNA-interacting residue.
[0160] Loop means the region of an LSR is likely to form a peptide loop and primarily interacts with bases at or about positions 6 and 7 in each attP and attB halfsite. For Bxbl the loop region is from residues 154 to 159. For Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, andS saint the helix regions are 151-157, 155-160, 155-160, 167-172, 167-172, 156-161, 162-167, 179- 183, 161-166, and 164-169 respectively.
[0161] Left halfsite means the 5’ most 19 bp of an attB site or 5’ most 24 bp of an attP site.
[0162] A “pseudo-site” is a DNA sequence recognized by a wildtype recombinase enzyme even though that the recognition site differs in one or more base pairs from the wild-type recombinase recognition sequence and / or is present as an endogenous sequence in a genome that differs from the genome where the wild-type recognition sequence for the recombinase resides. “Pseudo attP site” or “pseudo attB site” refer to pseudo sites that are similar to wild-type phage or bacterial attachment site sequences, respectively, for phage integrase enzymes. In the presence of the corresponding wildtype recombinase and a complimentary target site (e.g. a separate attP site if the pseudo-site is an attB site) recombination at the pseudo-site can be detected under certain experimental conditions.“Pseudo att site” is a more general term that can refer to either a pseudo attP site or a pseudo attB site.
[0163] The terms “Variant Recombinase” or “Recombinase Variant(S)” and the like refers to recombinases with one or more amino acid mutations that alter the DNA recognition specificity of the variants relative to their wild-type counterparts.
[0164] Right halfsite means the reverse-complement of the 5’ most 1 bp of an attB site or the reverse-complement of the 5’ most 24 bp of an attP site.
[0165] “Staffer Sequence” means a DNA sequence of sufficient length to form a loop (at least 300 bp) that splits and inactivates an antibiotic resistance gene. The length rather than the sequence of such a stuffer sequence is the critical aspect.
[0166] “Wildtype Bxbl attP”: SEQ ID. NO: 2 is an attP site recognized by wildtype Bxbl recombinase, where the center dinucleotides are italicized and in boldface, the single-underlined regions are recognized by the ZD domains, and the double-underlined regions are recognized by the RD domains.
[0167] “Wildtype Bxbl attB”: SEQ ID. NO: 3 is an attB site recognized by wildtype Bxbl recombinase, where the center dinucleotides are italicized and in boldface, the single-underlined regions are recognized by the ZD domains, and the double-underlined regions are recognized by the RD domains.
[0168] Other features, objectives, and advantages of the invention are apparent in the detailed description that follows. It should be understood, however, that the detailed description, while indicating embodiments and aspects of the invention, is given by way of illustration only, not limitation. Various changes and modification within the scope of the invention will become apparent to those skilled in the art from the detailed description. In some embodiments, the integrase is selected by its ability to reassemble a non-functional target gene.
[0169] Variant Recombinases
[0170] Disclosed are variant recombinases. In some embodiments, the variant recombinase comprises one or more amino acid mutations in at least two DNA binding (or DNA interacting) regions that control DNA recognition specificity for the relevant portions of the target site of a wildtype recombinase. In some embodiments, the variant recombinase has altered DNA recognition specificity relative to the cognate wildtype recombinase binding region. In some embodiments, the relevant portions of the target site are determined in terms of positions in the target site, wherein Loop specifies -7,-6 and +6, +7; helix specifies -11,-10, -9 and +9.+10.+11; hairpin specifies -19 to - 12 and +12 to +19.
[0171] In some embodiments, the binding regions of the variant recombinase comprise a loop region, a helix region, and a hairpin region. In some embodiments, the variant recombinase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes in any one of the loop region, the helix region, and / or the hairpin region. In some embodiments, the variant recombinase is Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes in any one of the loop region, the helix region, and / or the hairpin region. In some embodiments, the variant recombinase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes in any one of the loop region, the helix region, and / or the hairpin region. In some embodiments, the variant recombinase is Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901 and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes in any one of the loop region as defined in the definition section for the applicable recombinase, the helix region as defined in the definition section for the applicable recombinase, and / or the hairpin region as defined in the definition section for the applicable recombinase. In some embodiments, the variant recombinase comprises 1, 2, 3, 4,5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes in at least two binding regions wherein the binding regions are selected from the group comprising loop region, helix region, and hairpin region. In some embodiments, binding regions for wild-type or non-wild-type recombinases may be modified in the generation of variant recombinases. In some embodiments, at least two of the loop, helix, hairpin regions of at least one monomer are modified. In some embodiments, each of the at least two of the loop, helix, hairpin regions of at least one monomer that are modified comprise at least one modification.
[0172] In some embodiments, the recombinase variant may comprise mutations within the loop domain, helix domain, hairpin domain, and combinations thereof. In some embodiments, the recombinase variant may further comprise mutations within the NTD domain or CC coiled motif amino acid sequences. In some embodiments, the recombinase variant may further comprise mutations outside the loop, helix, and hairpin domains. Disclosed are any recombinases with one or more mutations in the “loop”, one or more mutations in the “helix”, one or more mutations in the “hairpin” submotif and combinations thereof.
[0173] In some embodiments, only loop, and / or helix, and / or hairpin structures are modified in the generation of variant recombinases. In each instance variant recombinases have an altered binding specificity compared to their wild-type counterpart.
[0174] In some embodiments the donor construct is designed to interact with wild-type Bxbl and only differs from a natural attP or attB sequence in the center 2 bp (CDN). In some embodiments the donor construct is designed to interact with Bxbl variants such as the Bxbl variants intended to recognize the left and right halfsites of the intended genomic target. In some embodiments both halfsites of the donor construct are designed to interact with the same variant Bxbl variant and in some embodiments each halfsite of the donor is designed to interact with a different Bxbl variant. In other embodiments the donor construct is designed to interact with integrases other than Bxbl.
[0175] Disclosed are large serine recombinase variants, and systems comprising, or consisting of, or consisting essentially of, said variants for altering the genomic DNA sequence of the host cell. These recombinase variants have altered DNA recognition preferences compared to their wildtype counterparts and can be used to recognize endogenous DNA sequences within the genome of organisms of interest, including humans. These enzymes can integrate donor DNA into the genome or excise or invert a desired genomic sequence. Thus, they can be used to integrate therapeutic, industrial or agricultural genes into the genome or to supplement, repair, or remove pathogenic genesfrom the genome, to achieve therapeutically, industrial, or agriculturally beneficial effects. Similarly, the disclosed large serine recombinase variants can be used for engineering of any eukaryotic or prokaryotic genome to achieve beneficial effects, incl synthetic biology, cell line engineering, or microbial strain engineering.
[0176] In some embodiments, the variant recombinase exhibits increased recombination as compared to non-variant recombinases.
[0177] In some embodiments, the variant recombinase disclosed herein is a modified wild-type recombinase. In some embodiments, the variant recombinase disclosed herein is derived from a previously modified recombinase, i.e., not a wild-type recombinase. In some embodiments, the variant recombinase disclosed herein is a Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901 recombinase as discussed in more detail below. In some embodiments, the variant recombinase is computationally generated using rules learned from naturally occuring recombinases. In some embodiments, the variant recombinase preferentially binds the intended target site.
[0178] In some embodiments, identification of target recombination sequences can be accomplished, for example, by using sequence alignment and analysis, where the query sequence is the recombination site of interest (for example, attP and / or attB). In some embodiments, the genome of a target cell may be searched for sequences having sequence identity to the selected recombination site for a given recombinase, for example, the wildtype attP and / or wildtype attB of Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901.
[0179] In some embodiments, after recombination the post-recombination sites are no longer able to act as substrate for the variant recombinase. This results in stable recombination with little or no recombinase mediated excision of the integrated sequence.
[0180] In some embodiments, the variant recombinase recognizes endogenous DNA sequences in a genome of interest. The mutations present in a variant recombinase may comprise amino acid substitutions, deletions, insertions, and / or other rearrangements in the amino acid sequence of the recombinase, and / or any combination of such mutations, either singly or in groups. The variant recombinase may exhibit enhanced DNA recognition specificity, reduced DNA recognition specificity, altered DNA recognition specificity (i.e., binds to a different target DNA site) ascompared to the same wild-type enzyme and / or to a different non-variant recombinase. In some embodiments, the variant recombinase has greater or lesser catalytic activity toward a particular DNA sequence, including a wild-type or non-wild-type recombinase recognition site. In some embodiments, the variant recombinase has enhanced or reduced specificity to the attB and / or attP site compared to the wildtype attB and / or wildtype attP site.
[0181] In some embodiments, the recombinase variant mediates no measurable integration above background levels at an endogenous target site in a eukaryotic genome, relative to the wild-type recombinase. In some embodiments, the recombinase variant increases transgene insertion at least 1-, 2-, 3-, 4-, 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, 15-, 16-, 17-, 18-, 19-, 20-, 50-, 60-. 70-, 100-, 200- fold, relative to the wild-type recombinase.
[0182] In some embodiments, the recombinase variant can recognitize a desired DNA sequence. In some embodiments, the desired target DNA sequence is different than the wildtype DNA binding sequence for the wildtype recombinase. In some embodiments, the desired target DNA sequence is the same as the wildtype DNA binding sequence for the wildtype recombinase. Such variant binding specificity permits the recombinase to react with a given target DNA sequence differently than would the native enzyme, while a variant level of activity permits the recombinase to carry out the reaction at greater or lesser efficiency. Stated another way, provided is a non-naturally occurring variant of a recombinase, comprising, or consisting of, or consisting essentially of, one or more amino acid mutations (e.g., substitutions, deletions, or insertions) that produce altered DNA target site specificity relative to the wildtype recombinase counterpart. In some embodiments, the variant recombinase comprises, or consists of, or consists essentially of, one or more amino acid mutations (e.g., substitutions, deletions, or insertions) within a loop structure, helix structure, hairpin structure and / or combinations thereof.
[0183] In some embodiments, the variant recombinase mediates increased transgene insertion at an endogenous target site, relative to the same wild-type recombinase and / or a different non-variant recombinase. In some embodiments, the variant recombinases recognize non-endogenous sequences in a cell, i.e., recognizes an artificial sequence that was previously integrated into the host cells’ genome. In some embodiments the variant recombinase mediates detectable levels of integration at an endogenous target site where the wild-type recombinase has no or minimal detectable integration activity.
[0184] The recombination sites used as substrates in the methods disclosed herein include, but are not limited to, wildtype attB, wildtype attP, pseudo-attB, pseudo-attP, non-wildtype attB, and non-wildtype attP, and combinations thereof. The two att sites that are recombined must consist of a combination of an attB site and attP site. In addition, the CDN of the attB and attP sites must match or the reverse complement of the CDN in the attB site must match the CDN in the attP site. If a donor construct is being integrated into genomic sequence then the orientation of donor integration can be controlled by choosing to have the CDNs match or choosing to have one CDN match the reverse complement of the other CDN. However, if the CDN sequence is palindromic then the CDN will match its own reverse-complement and both integration orientations will be achieved. Typically, at least one of the recombination sites provide a substrate for the variant recombinase.Recombination target sites may be identified, using methods described herein, in the genome of essentially any target cell, including, but not limited to, prokaryotes, eukaryotes, E. coli or other bacteria, yeast or other fungi, plant protoplast or other plant cells, human cells, rodent cells, as well as K562 cells or any other mammalian or more generally animal cell line.
[0185] In some embodiments, the variant recombinases are capable of binding DNA target site sequences other than the wildtype recombinase binding sequence. In some embodiments, the variant DNA target site sequences of interest can be a safe harbor locus, such as AAVS1, and / or TRAC. In some embodiments, recombinase target site sequence in AAVS1 is Seq ID No. 17. The recombinase target sequence in TRAC is Seq ID No. 18. In some embodiments the variant DNA target site sequence of interest in within the human ribosomal DNA repeat. The sequence of the human ribosomal repeat has genbank accession KY962518.1.
[0186] In some embodiments, the disclosed recombinase variants are more versatile than their wild-type counterparts in that the variants can bind a wider repertoire of DNA recognition sites. For example, they can bind target sequences (e.g., attB, attP that may be more efficiently bound, less efficiently bound or have the same binding efficiency by one or mixtures of the variants herein than by the wildtype recombinase. A wildtype recombinase may be engineered as described above by, e.g., adopting the amino acid modifications, such that the engineered protein can now bind to other endogenous genomic target sites in human cells.
[0187] Disclosed is knowing which positions in the recombinase to vary to recognize desired DNA sequences and the strategy for combining the engineered portions into composite variants that can be delivered as a pair to recognize the left and right halfsites of an endogenous target site. As discussed above, the selection of the mutation in the recombinase is based on the target site. Once askilled artisan knows the DNA binding regions in a recombinase (which can includeloop, helix, and hairpin;), then the recombinase of interest can be engineered using a variety of methods known to someone skilled in the art to recognize the desired DNA sequence into which one may wish to insert the genetic cargo .
[0188] In some embodiments, the donor molecule is single stranded DNA. In some embodiments, the donor molecule is double stranded DNA. In some embodiments, the target recognition site on the donor molecule is single stranded and is made at least partially doublestranded - sufficient to have the attachment site accessible for a (variant) integrase. To achieve this, different types of oligos can be used, or have the single-stranded DNA form a double-stranded region with itself by adding a self-complementary loop - or loops.
[0189] It is understood that parts of recombinase other than loop, helix, and hairpin could be altered in any known ways without affecting the binding or recombination of the variant recombinase.
[0190] Disclosed is a non-naturally occurring recombinase variant that has at least two mutations relative to wild-type integrase with a first mutation in the helix region and a second mutation in the loop region, and / or hairpin region and wherein each variant has at least 80-90% identity to the wildtype integrase in the residues which are outside the mutated loop, helix region, and hairpin regions. If the non-naturally occurring recombinase with altered helix, loop, and / or hairpin regions comprises an RD domain similar to a first naturally occurring recombinase and a ZD domain similar to a second naturally occurring recombinase then the RD domain will have at least one mutation in the loop or helix regions relative to the cognate naturally occurring RD domain and have at least 80-90% identity to the cognate naturally occurring RD domain outside of the helix and loop regions and the ZD domain will have at least one mutation in the hairpin region relative to the cognate naturally occurring ZD domain and at least 80-90% identity to the cognate naturally occurring ZD domain outside of the hairpin region. It is expected that such integrases will show detectable activity at an endogenous target site in a prokaryotic or eukaryotic genome, and the wild-type integrase will not show detectable activity at the same target site. It is understood that parts of the integrase other than loop, helix, and hairpin could be altered in any known ways without affecting the binding or recombination of the variant recombinase. In some embodiments, the mutation is a single amino acid mutation. In some embodiments, the mutation is at least one amino acid mutation.
[0191] In some embodiments, potential attB target and potential attP target sites in the genome can be identified by looking for DNA sequences that have a G at both position -4 on the top strand and position +4’ on the bottom strand In some embodiments, the potential attB target site or potential attP target site is divided, and the relevant recombinase domain and cognate quarter site behave in a modular fashion allowing each quarter site to be combined in order to recognize an entire half site.
[0192] In some embodiments, the recombinase variant recognizes both halfsites of the genonic recombination site and is co-delivered with a donor construct in order to integrate the donor construct into the desired genomic locus. A further aspect of this embodiment is that the recombinase variant also recognizes both halfsites of the genomic recombination site on the donor construct.
[0193] In some embodiments the genomic target site is an attB site and there is an attP site in the donor. In other embodiments the genomic target iste is an attP site and there is an attB site in the donor.
[0194] In some embodiments, the recombinase variant that recognizes the left half site of the att site in the genome is co-delivered with a recombinase variant that recognizes the right half site of the att site in the genome and a donor construct in order to integrate the donor construct into the desired genomic locus. The att site in the donor construct may be recognized by the recombinase variant that recognizes the left half site of the att site in the genome, the att site in the donor construct may be recognized by the recombinase variant that recognizes the right half site of the att site in the genome, or the att site in the donor construct may be recognized by a combination of the recombinase variant that recognizes the left half site of the att site in the genome and the recombinase variant that recognizes the right half site of the att site in the genome.
[0195] In some embodiments, the recombinase variant that recognizes the left half site of the att site in the genome is co-delivered with i) a recombinase variant that recognizes the right half site of the att site in the genome, ii) a donor construct in order to integrate the donor construct into the desired genomic locus, and iii) a wildtype recombinase that recognizes both half sites of the att site in the donor. In some embodiments, the recombinase variant that recognizes the left half site of the att site in the genome is co-delivered with i) a recombinase variant that recognizes the right half site of the att site in the genome, ii) a donor construct in order to integrate the donor construct into the desired genomic locus, and iii) a variant recombinase that recognizes both half sites of the att site in the donor.
[0196] In some embodiments, the recombinase variant that recognizes the left half site of the att site in the genome is co-delivered with i) a recombinase variant that recognizes the right half site of the att site in the genome, ii) a donor construct in order to integrate the donor construct into the desired genomic locus, and iii) a variant recombinase that recognizes the left hafl site of the att site in the donor, and iii) a variant recombinase that recognizes the right half site of the att site in the donor.
[0197] In preferred embodiments the attB recombination site is endogenous genomic sequence and donor contains an attP recombination site. This attP recombination site in the synthetic donor construct may be generated by inserting 5 basepairs between positions -12 and -11 of the genomic attB recombination site and inserting 5 basepairs between positions +12 and +13 of the genomic attB recombination site. These 5 basepairs may have the DNA sequence 5’-TGGTC-3’ or the DNA sequence 5’-CCGTA-3’. This attP recombination site in the synthetic donor construct may also be generated by inserting 5 basepairs between positions -12 and -11 of the genomic attB left half site site and inserting 5 basepairs between positions +12 and +13 of the reverse complement of the genomic attB halfsite. In further embodiments, this attP recombination site in the synthetic donor construct may be generated by inserting 5 basepairs between positions -12 and -11 of the reverse complement of genomic attB right half site site and inserting 5 basepairs between positions +12 and +13 of the genomic attB right halfsite.
[0198] In some embodiments, recombinase variants are generated using a donor that has the center 2 bp matching the center 2 bp of the intended target and is otherwise identical the wild type attP sequence. In some embodiments, the system co-delivers wild-type recombinase and variant recombinase. In some embodiments, the wild-type recombinase is not delivered with the variant recombinase. In some embodiments, the center 2 bp of the donor sequence match the center 2 bp of the target.
[0199] In some embodiments both attB and attP sites are in the genome. In such cases recombination will either excise the sequence in between the genomic attB site and the genomic attP site or invert the sequence in between the genomic attB site and the genomic attP site depending on the whether i) the CDN in the attB site matches the CDN in the attP site or ii) the CDN in the attB site matches the reverse-complement of the CDN in the attP site. If the CDN is palindromic then some recombination events will yield exicisions while other recombination events will yield inversions.
[0200] In some embodiments, the loop has a single mutation (e.g., substitutions, deletions, or insertions) and the helix and hairpin have no mutations. In some embodiments, the loop has more than one mutation (e g., substitutions, deletions, or insertions) and the helix and hairpin have no mutations. In some embodiments, the loop has a single mutation, and the helix has a single mutation, and the hairpin has no mutations. In some embodiments, the loop has more than one mutation and the helix has a single mutation, and the hairpin has no mutations. In some embodiments, the loop has more than one mutation and the helix has more than one mutation and the hairpin has no mutations. In some embodiments, the loop has more than one mutation, the helix has more than one mutation and the hairpin has more than one mutation.
[0201] In some embodiments, the loop has a single mutation, and the hairpin has a single mutation, and the helix has no mutations. In some embodiments, the loop has more than one mutation and the hairpin has a single mutation, and the helix has no mutations. In some embodiments, the loop has more than one mutation and the hairpin has more than one mutation and the helix has no mutations.
[0202] In some embodiments, the loop has a single mutation, and the hairpin has a single mutation, and the helix has a single mutation. In some embodiments, the loop has more than one mutation and the hairpin has a single mutation, and the helix has a single mutation. In some embodiments, the loop has more than one mutation and the hairpin has more than one mutation and the helix has a single mutation. In some embodiments, the loop has more than one mutation and the helix has more than one mutation and the hairpin has a single mutation. In some embodiments, the loop has more than one mutation and the helix has more than one mutation and the hairpin has more than one mutation.
[0203] In some embodiments, the loop has one mutation, and the helix has more than one mutation and the hairpin has more than one mutation. In some embodiments, the loop has more than one mutation and the helix has one mutation, and the hairpin has more than one mutation. In some embodiments, the loop has more than one mutation and the helix has more than one mutation and the hairpin has one mutation.
[0204] In some embodiments, the provided recombinase variants may comprise linkers, His tags, and / or nuclear localization sequences at the N- or C-terminal ends.
[0205] In some embodiments, the recombinase excises a gene or portion of a gene such that in the target cell / plasmid the gene of interest is reassembled. In some embodiments, the recombinaseexcises a gene or portion of a gene such that a resistance gene such as gentamicin is reassembled in a cell / plasmid.
[0206] In some embodiments, the activity of the provided recombinases may be improved by adding mutations outside of the loop, helix, and hairpin regions that have been identified by making wildtype Bxbl more active for wildtype Bxbl attB and attP target sites.
[0207] In some embodiments, the activity of the provided recombinases may be improved by adding mutations outside of the loop, helix, and hairpin regions that have been identified by making variant recombinases more active for non-wildtype attB and / or non-wildtype attP target sites.
[0208] In some embodiments, the activity and / or specificity of the provided recombinases may be improved by fusing a programmable DNA-binding domain such as a zinc finger (ZF), TALE, or dCas9 to the C-terminus of the recombinase.
[0209] In preferred embodiments, the activity and / or specificity of the provided recombinases may be improved by fusing an engineered zinc finger array to the C-terminus of the provided recombinase using a peptide linker of between 15 and 40 residues such that the ‘3 edge of the zinc finger binding site is separated by 4, 5, 6, 7, 8, 9, or 10 basepairs from the edge of att site for the provided recombinase. Such zinc finger fusions can be made to one or multiple recombinase monomers and the zinc finger target sites can be present in the endogenous human genome and / or the synthetic donor construct.
[0210] The present gene editing systems comprising, or consisting of, or consisting essentially of, the recombinase variants are advantageous over other gene editing systems in several important ways. First, the present systems avoid undesired changes to the genome. The most widely used gene editing systems are CRISPR / Cas (clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated protein (Cas)) systems, zinc-finger nuclease (ZFN) systems, and transcription activator-like effector nuclease (TALEN) systems. Those systems depend upon the activity of nucleases and DNA repair mechanisms, such as homologous recombination and non- homologous end joining, which are error prone and may introduce insertions or deletions (indels) and / or translocations. By having both the cutting and ligating functions and remaining covalently attached to the DNA ends during the recombination reaction, the present recombinases cleave DNA and ligate its breaks at a highly precise, site-specific manner, avoiding the introduction of harmful indels into the genome. Second, there is no inherent size limit to the DNA that can be integrated into a host genome using the present editing systems. Third, the recombination generated by the presentsystems is not easily reversible without an accessory protein called recombination directionality factor (RDF); thus, the gene edits are stable and heritable. In sum, the present recombinase variants can be used to stably and precisely remove DNA from, integrate large synthetic and / or exogenous donor DNA into, or invert a segment of, a host genome at sites that are specifically recognized by the present recombinase variants.
[0211] Bxbl variants
[0212] Target sites of variant recombinases
[0213] In some embodiments, the binding regions of the variant recombinases recognize one or more of the DNA target site sequence as shown in Figures 5A, 9A, 9B, 10A, 10B(l), 10B(2), 11 A, 11D, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22. In some embodiments, the variant recombinase recognizes AAVS1 or TRAC. In some embodiments, target site used for AAVS1 hairpin selection is the CTGAGCGC ZD submotif at positions -19 to -12 of the AAVS1 5032 left halfsite. In some embodiments, target site used for AAVS1 hairpin selection is the GGGTTTGA ZD submotif if positions -19 to -12 of the AAVS1 5032 right halfsite.
[0214] A preferred location within the human genome to integrate synthetic donor constructs is within a ribosomal DNA repeat region because there are 400-600 copies of this sequence in each cell (www.ncbi.nlm nih.gov / pmc / articles / PMC7403359 / incorporated by reference in its entirety). If the synthetic donor construct expresses a CAR construct, then even modest levels of integration at any one of the 400-600 ribosomal DNA repeats would result in a large fraction of cells with at least one integrated CAR construct. This would have utility in applications such as CAR T therapy. Genomewide integration experiments indicate that Bxbl variants can integrate synthetic DNA donor constructs where the center 2 bp of the donor matches the center 2 bp of the target site at the following sequences within the human ribosomal DNA repeat: TTGATCCTGCCAGTAGCATATGCTTGTCTCAAAGATTA, TGACCTTCATTTGTGGAATCCTCAGTCATCGACACACA, GCACAGTACCCAACCGGTATTGCAGTGGTGAGAAGCTA, TGTGGCGCAACACTTGGACAGGCAGTTGCTAAAGCTCT, and GACTCTCGCTCTGCTGCCCAGGCTGGAGTGCAGCGGCG. The most preferred of these sequences are TTGATCCTGCCAGTAGCATATGCTTGTCTCAAAGATTA and TGTGGCGCAACACTTGGACAGGCAGTTGCTAAAGCTCT since these have the highestintegration counts and integration at these sequences was observed with multiple Bxbl variants. The methods described herein to recognize desired DNA target sites with modular integrases could be used to generate integrases that can recognize one or more of the aforementioned endogenous human DNA target sequences. Thus, in some embodiments, a recombinase variant described herein integrates as described at the aforementioned sequences or target sites. In a preferred embodiment, a Bxbl variant described herein integrates as described at the aforementioned sequences or target sites
[0215] Bxbl is a Known Large Serine Recombinase
[0216] As set forth above, each RD and ZD domains can be identified and engineered as described in the context of general recombinases (and specifically large serine recombinases) and as exemplified here by Bxbl. The loop, helix, hairpin regions are defined for Bxbl, and equivalent regions are known or could be determined in other recombinases. As such, other known ortholog integrases can be altered in a similar way to reprogram specificity as outlined above.
[0217] Bxbl is a known large serine recombinase. The NTD domain may approximately correspond to amino acids 1-145 (numbering according to SEQ ID NO: 1) in wildtype Bxbl recombinase or the functionally analogous sequence in Bxbl recombinase variants. The loop domain may approximately correspond to amino acids 154-159 (numbering according to SEQ ID NO: 1) in wildtype Bxbl recombinase or the functionally analogous sequence in Bxbl recombinase variants, or the functionally analogous sequence(s) in orthologous serine recombinase(s) and their variants. The helix domain may approximately correspond to amino acids 231-237 (numbering according to SEQ ID NO: 1) in wildtype Bxbl recombinase or the functionally analogous sequence in Bxbl recombinase variants, or the functionally analogous sequence(s) in orthologous serine recombinase(s) and their variants. The RD domain may approximately correspond to amino acids 140-287 (numbering according to SEQ ID NO: 1) in wildtype Bxbl recombinase or the functionally analogous sequence in Bxbl recombinase variants, or the functionally analogous sequence(s) in orthologous serine recombinase(s) and their variants. The hairpin domain may approximately correspond to amino acids 308-325 (numbering according to SEQ ID NO: 1) in wildtype Bxbl recombinase or the functionally analogous sequence in Bxbl recombinase variants, or the functionally analogous sequence(s) in orthologous serine recombinase(s) and their variants. In some embodiments, the recombinase variant may further comprise mutations within the NTD domain or CC coiled motif amino acid sequences.
[0218] The ZD domain may approximately correspond to amino acids 302-500 (numbering according to SEQ ID NO: 1) in wildtype Bxbl recombinase or the functionally analogous sequences in Bxbl recombinase variants, or the functionally analogous sequence(s) in orthologous serine recombinase(s) and their variants. The coiled-coil (CC) motif may approximately correspond to that found within the ZD domain in wildtype Bxbl recombinase or the functionally analogous sequences in Bxbl recombinase variants, or the functionally analogous sequence(s) in orthologous serine recombinase(s) and their variants.
[0219] Fig. 1A. A synthetic plasmid DNA donor construct is shown along with the endogenous target site in a eukaryotic cell. The light gray and dark gray boxes indicate attB and attP sites, respectively. Pl and P2 refer to primer binding locations used to measure integration frequency in the NGS assay. Shown at bottom are the products resulting from successful integration events. More specifically, shown is a schematic illustrating the chromosomal recombination assay in human cells. Bxbl facilitates the targeted integration (TI) of a donor plasmid into a chromosomal Bxbl attB pseudo-site (an endogenous sequence where the wildtype LSR has measurable activity).
[0220] Fig. IB: A diagram showing the synaptic complex prior to integration. The two natural target sites for wild-type Bxbl, wildtype Bxbl attP (SEQ ID. NO: 2) and wildtype Bxbl attB (SEQ ID. NO: 3) are shown. Bxbl binds each of these sites as a dimer with one monomer of Bxbl bound to the left halfsite and one monomer bound to the right halfsite. The different domains of each Bxbl monomer in each Bxbl dimer are shown as in the linker between the RD and ZD domains. NTD indicates the N-terminal domain that contains the catalytic domain (CD) and is not altered in any of the disclosures in this application. RD indicates the recombinase domain, and this domain contains both the loop region and helix region involved in DNA recognition and these regions are altered in some of our disclosed Bxbl variants. ZD indicates the zinc ribbon domain, and this domain contains the hairpin region. The hairpin region is altered in some of our disclosed Bxbl variants. Only a combination of a first Bxb 1 dimer bound to an attP site and a second Bxb 1 dimer bound to an attB site will trigger recombination between these two sites. The center 2 bp of the attP and attB sites (CDN, underlined) must match each other for recombination to occur. Recombination results in the left halfsite of attB being combined with the right halfsite of attP and the left halfsite of attP being combined with the right halfsite of attB. If the attP site is in a circular DNA construct, then the recombination between attP and attB will result in the integration of the entire circular donor construct in the DNA bearing the attB site. The attP site naturally found in the Bxbl phage and has 5 additional bp between the portions that are recognized by the recombinase domain (RD) and the zincribbon domain (ZD) in both half-sites while the attB site is in the bacterial genome where the phage sequence integrates and has the portion recognized by the RD and ZD domains adjacent to each other. A peptide linker between the RD and ZD domains can adopt an extended conformation to accommodate the RD and ZD domains binding an attP site with 5 bp separating the DNA sequence recognized by the RD domain and the DNA sequence recognized by the ZD domain. The portions of the site that are recognized by each of these domains are boxed in Fig 1C. If viable attP and attB sites are present in a cell that expresses Bxbl and the CDNs match each other than the Bxbl enzyme exchanges DNA strands between the attP and attB sites and integrates the donor into the target site. This process converts the attP and attB sites into two new sites with different sequences (attL and attR) that are no longer substrates for Bxbl and thus the integration is irreversible. Figla shows a diagram of how this recombination between attP and attB can integrate the circular donor construct into the desired portion of the genome of interest.
[0221] Fig. 1C. Annotated versions of the two natural target sites for wildtype Bxbl, wildtype Bxbl attP (SEQ ID. NO: 2) and wildtype Bxbl attB (SEQ ID. NO: 3). The left and right halfsites of attP and attB and the CDN are indicated. Regions of each halfsite recognized by the RD and ZD domains are boxed.
[0222] Fig. 2 shows Bxbl protein sequence and annotated wildtype Bxbl attB (SEQ ID. NO: 3) target sequence: The top panel is a more detailed description of the attB target site that shows a numbering scheme for each base and shows the portion of each half site that is recognized by the loop, helix and hairpin regions of Bxbl . Note that bases in the bottom strand of DNA are denoted by a number followed by an apostrophe. The bottom panel shows amino acid sequence of wild-type Bxbl enzyme with boxes around the loop, helix, and hairpin regions that are key aspects disclosed herein. The active site serine that is not modified and residue 257 where the D257K mutation seems to increase activity for many different target sites are also boxed and labeled.
[0223] Disclosed are variants of SEQ ID NO: 1 that have altered DNA recognition sequences. In one aspect, the disclosure provides a non-naturally occurring variant of Bxbl recombinase with altered DNA target site specificity relative to wildtype Bxbl recombinase (e.g., SEQ ID NO: 1), comprising, or consisting of, or consisting essentially of, a first mutation (e.g., substitutions, deletions, or insertions) within a loop, helix and / or hairpin structure and at least a second mutation (e.g., substitutions, deletions, or insertions) within a loop, helix and / or hairpin structure wherein the first mutation and second mutation are in different loop, helix and / or hairpin structures.
[0224] In some embodiments, the Bxbl variant recombinase comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes in any one of the loop region, the helix region, and / or the hairpin region. In some embodiments, the Bxbl variant recombinase comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to wildtype Bxbl at positions 308, 309, 310, 311, 312, 313, 314, 316, 318, 321, 322, 323, and 325 in the hairpin region or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to wildtype Bxbl at positions 154, 155, 156, 157, 158, and 159 in the loop region or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to wildtype Bxbl at positions 231, 232, 233, 234, 236, and 237 of the helix region or one or more combination thereof. In some embodiments, the Bxbl variant recombinase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to wildtype Bxbl at positions 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, and 325 in the hairpin region.
[0225] In some embodiments, the one or more amino acid mutations occur at one or more of positions 154-159 [loop], 231-237 [helix region], and / or 308 to 325 [hairpin region] (numbering according to SEQ ID NO: 1). In some embodiments, the one or more amino acid mutations occur at one or more of positions 154, 155, 156, 157, 158, or 159 (numbering according to SEQ ID NO: 1). In some embodiments, the one or more amino acid mutations occur at one or more of positions 154, 155, 156, 157, 158, or 159 and one or more of positions 231, 232, 233, 234, 235, 236, 237 and / or 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, and 325 (numbering according to SEQ ID NO: 1). In some embodiments, the one or more amino acid mutations occur at one or more of positions 231, 232, 233, 234, 235, 236, 237 and / or 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, and 325 (numbering according to SEQ ID NO: 1).
[0226] In some embodiments, an asparagine amino acid residue (N) at position 234 specifies a thymine DNA base (T) at position +10 or -10 of the relevant halfsite. In some embodiments, a tryptophan amino acid residue at position 233 specifies a cysteine base (C) at position +11 or -11 of the relevant target halfsite. In some embodiments, two lysine residues (K) at positions 236 and 237 specifies an adenine DNA base (A) at position +9 or -9 of the relevant target halfsite. In some embodiments, an asparagine amino acid residue (N) at position 234 specifies a thymine DNA base (T) at position +10 or -10 of the relevant halfsite, a tryptophan amino acid residue at position 233 specifies a cytosine base (C) at position +11 or -11 of the relevant target halfsite, two lysine residues (K) at positions 236 and 237 specifies an adenine DNA base (A) at position +9 or -9 of the relevanttarget halfsite and / or combinations thereof. In some embodiments, the variant recombinase comprises an asparagine amino acid residue (N) at position 234 and a thymine DNA base (T) at position +10 of the attB halfsite or -10 of the attP halfsite.
[0227] Fig. 3 is a schematic showing the strategy for engineering Bxbl variants against a desired target site. Potential attB sites in the genome can be identified by looking for DNA sequences that have a G at both position -4 on the top strand and position +4’ on the bottom strand. From here on, all positions that start with ‘+’ are assumed to be on the bottom strand of the attB target site in the genome of interest. If the potential attB target site is divided as indicated, then the relevant Bxbl domain and cognate quarter site will behave in a modular fashion and will allow the results of parallel directed evolution selections to target each quarter site to be combined in order to recognize an entire half site. The boundaries of the quarter sites being a key aspect disclosed herein. The Bxbl variant that recognizes the left half site can then be co-delivered with a Bxbl variant that recognizes the right half site and a donor construct in order to integrate the donor construct into the desired genomic locus. To simplify the process, initial tests will be performed with a donor that has the center 2 bp matching the center 2 bp of the intended target but will otherwise be identical the wildtype attP sequence and will require co-delivery of wild-type Bxbl to bind such a donor. After a successful pair of reagents is identified the Bxbl binding site on the donor will be modified to interact with one or both of the variants to obviate the need for co-delivery of wild-type Bxb 1.
[0228] Stated another way, Fig. 3 shows rapid parallel testing in human cells of variant recombinases attachment site. The example describes this process in the context of Bxbl attB sites. A skilled artisan will recognize this can be done with other attachment sites. Hence, disclosed here are compositions, expression vectors, and methods of use thereof, for the generation of transgenic cells, tissues, plants, and animals wherein the recombinase is modified to bind a different target site than the corresponding wildtype recombinase. In some embodiments, the recombinase variant that recognizes the left half site is modified in its ZD domain, its RD domain or combinations thereof. In some embodiments, the recombinase variant that recognizes the right half site is modified in the ZD domain, the RD domain, or combinations thereof. In preferred embodiments both the recombinase variant that recognizes the left halfsite and the recombinase variant that recognizes the right halfsite are each modified in both their RD and ZD domains. It being understood that once it is recognized by a skilled artisan that the recombinase can be modified to bind to a new / different target site, a skilled artisan will be able to practice our invention and make such modifications to the recombinase to direct the recombinase to bind a different / new target site of interest.
[0229] Disclosed is a pair of non-naturally occurring Bxbl recombinase variants that each have at least two mutations relative to wild-type Bxbl the first mutation in the region of residues 231-237 [helix region], and the second mutation in the region of residues 145-154 [loop], and / or 308 to 325 [hairpin region] and wherein each variant has at least 80-90% identity to wild-type Bxbl in the region of residues 1-144, 155-230, the region of residues 238-307, and the region of residues 326- 500, and that shows detectable binding activity at an endogenous target site in a eukaryotic genome, and where wild-type Bxbl does not show detectable binding activity at the same target site. It is understood that parts of Bxbl other than loop, helix, and hairpin could be altered in any known ways without affecting the binding or recombination of the variant Bxbl recombinase.
[0230] In some embodiments, the Bxbl variant recombinase comprises an amino acid sequence YRGGLP in the loop region instead of an amino acid sequence YRGSLP from wildtype Bxbl. In some embodiments, the Bxbl variant recombinase comprises an amino acid sequence AGGNLKR or YPWSLRR in the helix region instead of an amino acid sequence SATALKR from wildtype Bxbl. In some embodiments, the Bxbl variant recombinase comprises an amino acid sequence KAWGSRKTRLYR, MASGSRKTAIYY, or MARGGRKSAIYY in the hairpin region instead of an amino acid sequence FAGGGRKHPRYR from wildtype Bxbl. In some embodiments, the Bxbl variant recombinase comprises an amino acid sequence YRGGLP in the loop region instead of an amino acid sequence YRGSLP from wildtype Bxbl, an amino acid sequence AGGNLKR or YPWSLRR in the helix region instead of an amino acid sequence SATALKR from wildtype Bxbl, and / or an amino acid sequence KAWGSRKTRLYR, MASGSRKTAIYY, or MARGGRKSAIYY in the hairpin region instead of an amino acid sequence FAGGGRKHPRYR from wildtype Bxbl.
[0231] In some embodiments, the variant recombinase comprises, consists of or consists essentially of AGGNLKR instead of the SATALKR helix sequence from wildtype Bxbl and ZD 32. ZD32 refers to a Bxbl hairpin variant with a hairpin sequence of LARGRRKWARYR (SEQ ID No. 20). ZD32 is defined in Fig. 6D . In some embodiments, the recombinase variant utilizes a mixture of 16 donors bearing all possible sequences for the CDN of the attP or attB site in donor. In some embodiments, the donor CDN matches the center 2 bp of site 1-41 (numbering according to SEQ ID NO: 3).
[0232] In some embodiments, the cells are K562 cells. In some embodiments, the cells are primary human T cells. In some embodiments, the cells are human tissue (lung, brain, kidney, liver, etc.) in human patients, or cells of plant crop species (com, wheat, soybean, canola, tomato, etc.). In some embodiments, administering comprises one or more recombinase variants. In someembodiments, the one or more recombinase variants are used to multiplex and are recognizing different attB / attP pairs with different CDNs.
[0233] In some embodiments, the variant Bxbl recombinase comprises an asparagine amino acid residue (N) at position 234 and a Thymine (T) base at position +10 or -10 of attB and / or position +10 and -10 of attP, a tryptophan amino acid residue (W) at position 233 and a cysteine base (C) at position +11 or -11 of the attB halfsite and / or +11 or -11 of the attP halfsite, , two lysine residues (K) at positions 236 and 237, an adenine base (A) at position +9 or -9 of the attB halfsite and / or +9 or -9 of the attP halfsite, and combinations thereof. In some embodiments, the engineered Bxbl variant shows a strong preference for the intended site. In some embodiments, the variant Bxbl recombinase comprises a lysine residues (K) at position 236 with or without an, an adenine base (A) at position +9 or -9 of the attB halfsite and / or +9 or -9 of the attP halfsite.
[0234] In some embodiments, the variant recombinase comprises, consists of or consists essentially of AGGNLKR (helix REGO) and ZD 32. Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, one or more modification at positions 231- 237. Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, one or more modification at positions 154-159. Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, one or more modification at positions 308-314 and specifically at 314-325.
[0235] Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, one or more modification at positions 231-237 and 154-159. Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, one or more modification at positions 231-237 and 308-325. disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, one or more modification at positions 154- 159 and 308-325. Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, one or more modification at positions 231-237, 154-159 and 308-325. Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, the permutations of the mutations above.
[0236] Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, modified non-wildtype attB sequence(s). Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, modified attB sequence(s) whereinthe attB sequence(s) have been modified at position(s) (-19, -18, -17,-16, -15, -14, -13, -12, -11, -10, -9, -7 -6) and / or (19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 7, 6).
[0237] Disclosed is a Bxbl serine recombinase variant comprising, or consisting of, or consisting essentially of, modified attB sequence(s) (also referred to as non-wildtype attB sequences) wherein the attB sequence(s) have been modified at position(s) (23, -22, -21, -20, -19, -18, -17, -11, -10, -9, - 7, -6). In some embodiments, the mutation(s) are contingent upon the identity of the engineered recombinase for which such attB or attP sequences are suitable substrates. In some embodiments, the mutation(s) are contingent upon the identity of the engineered recombinase for which native attB or attP sequences are suitable substrates.
[0238] Disclosed is a variant Bxbl protein sequence wherein the residue 257 is modified, such as D257K.
[0239] Other Recombinases
[0240] In some embodiments, the variant recombinase is Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, or Ssalnt and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 307-326, 306-323, 286-302, 322-339, 317-343, 320-343, 289-313, 333-350, 296-315, or 283-311 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 151-157, 155-160, 167-172, 156-161, 162-167, 178-183, 161-166, 165-169, 169-174 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 229-235, 241-247, 221-227, 229-235, 230-236, 234-240, 218-224, 241-247, 222-228, or 218-224 of the helix region and / or one or more combination thereof, e.g., as shown in Figs. 17A-17J.
[0241] In some embodiments, the variant recombinase is Theia and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 307-326 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 151-157 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 229-235 of the helix region, and / or one or more combination thereof. In some embodiments, the variant recombinase is Theia and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 or 19 amino acid changes relative to the wildtype recombinase at positions 307- 326 in the hairpin region. Fig. 17A shows amino acid sequence of Theia with annotation showingthe predicted loop (positions 151-157), helix (positions 229-235), and hairpin (positions 307-326) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Theia indicated. In some embodiments, the variant Theia recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Theia in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 313, 312, 315, 317, 322, 324, and 326.
[0242] In some embodiments, the variant recombinase is Veracruz and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 320-343 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, or 5 amino acid changes relative to the wildtype recombinase at positions 155-160 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 241-247 of the helix region, and / or one or more combination thereof. In some embodiments, the variant recombinase is Veracruz and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12,13, 14, 15, 16, 17, 18, 19, 20, 21, 22 or 23 amino acid changes relative to the wildtype recombinase at positions 307-326 in the hairpin region. Fig. 17B shows amino acid sequence of Veracruz with annotation showing the predicted loop (positions 155-160), helix (positions 241-247), and hairpin (positions 320-343) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Veracruz indicated. In some embodiments, the variant Veracruz recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Veracruz in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 326, 328, 330, 332, 339, 341, and 343.
[0243] In some embodiments, the variant recombinase is Kp03 and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 306-323 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, or 5 amino acid changes relative to the wildtype recombinase at positions 169-174 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 229-235 of the helix region, and / or one or more combination thereof. In some embodiments, the variant recombinase is Kp03 and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13,14, 15, 16, or 17 amino acid changes relative to the wildtype recombinase at positions 306-323 in the hairpin region. Fig. 17C shows amino acid sequence of Kp03 with annotation showing the predicted loop (positions 155-160), helix (positions 229-235), and hairpin (positions 306-323) regionsavailable for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Kp03 indicated. In some embodiments, the variant Kp03 recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Kp03 in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 311, 312, 314, 315, 318, 321, and 323.
[0244] In some embodiments, the variant recombinase is PaOl and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 286-302 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, or 5 amino acid changes relative to the wildtype recombinase at positions 167-172 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 221-227 of the helix region, and / or one or more combination thereof. In some embodiments, the variant recombinase is PaOl and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 amino acid changes relative to the wildtype recombinase at positions 286-302 in the hairpin region. Fig. 17D shows amino acid sequence of PaOl with annotation showing the predicted loop (positions 167-172), helix (positions 221-227), and hairpin (positions 286-302) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of PaOl indicated. In some embodiments, the variant PaOl recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for PaOl in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 290, 292, 294, 297, 299, 300, and 302.
[0245] In some embodiments, the variant recombinase is Nm60 and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 322-339 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 167-172 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 230-236 of the helix region, and / or one or more combination thereof. In some embodiments, the variant recombinase is Nm60 and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or 17 amino acid changes relative to the wildtype recombinase at positions 322-339 in the hairpin region. See Fig. 17E. Fig. 17E shows amino acid sequence of Nm60 with annotation showing the predicted loop (positions 167-172), helix (positions 230-236), and hairpin (positions 322-339) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Nm60 indicated. In some embodiments, thevariant Nm60 recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Nm60 in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 328, 330, 333, 334, 335, 337, and 339.
[0246] In some embodiments, the variant recombinase is Si74 and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 317-343 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 156-161 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 234-240 of the helix region, and / or one or more combination thereof. In some embodiments, the variant recombinase is Si74 and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 amino acid changes relative to the wildtype recombinase at positions 317-343 in the hairpin region. Fig. 17F shows amino acid sequence of Si74 with annotation showing the predicted loop (positions 156-161), helix (positions 234-240), and hairpin (positions 317-343) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Si74 indicated. In some embodiments, the variant Si74 recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Si74 in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 323, 325, 328, 326, 339, 341, and 343.
[0247] In some embodiments, the variant recombinase is Bcylnt and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 289-313 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 162-167 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 218-224 of the helix region, and / or one or more combination thereof. In some embodiments, the variant recombinase is Bcylnt and comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or 24 amino acid changes relative to the wildtype recombinase at positions 289-313 in the hairpin region. Fig. 17G shows amino acid sequence of Bcylnt with annotation showing the predicted loop (positions 162-167), helix (positions 218-224), and hairpin (positions 289-313) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Bcylnt indicated. In some embodiments, the variant Bcylnt recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Bcylnt in thehairpin region include 1, 2, 3, 4, 5, or all positions selected from position 295, 297, 298, 299, 307, 311, and 313.
[0248] In some embodiments, the variant recombinase is Bcelnt and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 333-350 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 178-183 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 241-247 of the helix region, and / or one or more combination thereof. See Fig. 17H. Fig. 17H shows amino acid sequence of Bcelnt with annotation showing the predicted loop (positions 178-183), helix (positions 241-247), and hairpin (positions 333-350) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Bcelnt indicated. In some embodiments, the variant Bcelnt recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Bcelnt in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 337, 339, 341, 345, 346, 348, and 350.
[0249] In some embodiments, the variant recombinase is Ssclnt, and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 296-315 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 161-166 in the loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 222-228 of the helix region, and / or one or more combination thereof. See Fig. 171. Fig. 171 shows amino acid sequence of Ssclnt with annotation showing the predicted loop (positions 161- 166), helix (positions 222-228), and hairpin (positions 296-315) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Ssclnt indicated. In some embodiments, the variant Ssclnt recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Ssclnt in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 302, 304, 310,311, 313, and 315.
[0250] In some embodiments, the variant recombinase is Ssalnt and comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions 283-311 in the hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions 164-169 in the loop region and / or wherein the variantrecombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions 218-224 of the helix region, and / or one or more combination thereof. See Fig. 17J. Fig. 17J shows amino acid sequence of Ssalnt with annotation showing the predicted loop (positions 165- 169), helix (positions 218-224), and hairpin (positions 283-311) regions available for modification as described herein, and the target site with portions recognized by the loop, helix, and hairpin regions of Ssalnt indicated. In some embodiments, the variant Ssalnt recombinase comprises conserved bases at -4 and +4’ that will not vary. In some embodiments, the DNA contacting residues to engineer or randomize for Ssalnt in the hairpin region include 1, 2, 3, 4, 5, or all positions selected from position 289, 291, 293, 295, 307,309, and 311.
[0251] A person skilled in the art may readily perform a structural alignment of any large serine recombinase with Bxbl to identify the amino acid positions of the large serine recombinase that corresponds to the loop, helix, and / or hairpin regions of Bxbl in order to engineer the large serine recombinase to bind a desired DNA target site similar to engineered Bxbl as described herein. The first step of this structural alignment is to use a structure prediction tool such as RoseTTAfold or Alphafold2 to create a predicted 3-dimensional protein structure from the amino acid sequence of the large serine recombinase. Once predicted 3-dimensional structures of the large serine recombinase and Bxbl are created, then these structures can be structurally aligned using a structural alignment tool such as foldseek (search.foldseek.com / search) or software packages containing this functionality such as PyMol (www.pymol.org / ) . The amino acid positions of the large serine recombinase that structurally align with the Bxbl loop, helix, and hairpin can then be modified in the manner as disclosed herein to engineer Bxbl. The quality of a structural alignment is typically measured by the RMSD bewteen corresponding atoms in the alignment (e.g. alpha carbon atoms). RMSD stands for Root Mean Square Deviation and is a measure of the average distance between the atoms of two superimposed molecular structures and a lower value indicates a better alignment. For example, an alignment between two structures with an RMSD of less than 2.0, 2.5, or 3.0 Angstroms is generally considered a meaningful alignment between the two molecular structures. In some cases, the structural alignment is performed separately on the RD domain and the ZD domain of Bxbl to structurally align to the predicted structure of the large serine recombinase. Accordingly, in some embodiments, the variant recombinase is Dn29, PhiRvl, Al 18, TP901, or another protein sequence classified as an LSR or computationally designed to mimic an LSR, wherein the variant recombinase or LSR comprises hairpin, loop, and / or helix region(s) that structurally align with the hairpin, loop, and / or helix region(s) of wildtype Bxbl, wherein the variant recombinase or LSR comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to the wildtype recombinase at positions that structurallyalign with the Bxbl hairpin region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to the wildtype recombinase at positions that structurally align with the Bxbl loop region and / or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to the wildtype recombinase at positions that structurally align with the Bxbl helix region, and / or one or more combination thereof.
[0252] Methods for Generating Variant Recombinases
[0253] Disclosed are methods for the engineering of recombinases. In some embodiments, the structure of a recombinase is predicted using Al tools such as RoseTTAfold or Alphafold2. In some embodiments, the predicted Al structure is then broken up into three separate domains, termed the catalytic domain (CD), the recombinase domain (RD) and Zinc ribbon domain (ZD). In some embodiments, the predicted Al structure is then broken up into loop, helix, and hairpin domains. Each domain can be distinguished by those skilled in the art due to their structural similarity to other known recombinase structures. Generally, the first 150 amino acids of a recombinase comprise its catalytic domain (CD), such a domain is characterized by the presence of a long alpha helix at its C- terminus termed herein as the “RD boundary helix.” Generally, amino acids 280 -300 of a recombinase comprise its recombinase domain. The recombinase domain is characterized by the presence of an alpha helix at its C-terminus and is predicted by Alphafold2 to be separated from the ZD domain by a linker without defined secondary structure. The positioning of the RD and ZD portions of the LSR can then be roughly mapped to an attP DNA sequence by alignment of each portion of the structure to a known LSR structure, particularly that of the Listeria innocua integrase as described in (Li, H., Sharp, R., Rutherford, K., Gupta, K., Van Duyne, G.D. (2018) J Mol Biol 430: 4401-4418). This can be accomplished using molecular visualization software such as Pymol, for example, using the cealign command. Alignment of the recombinase and Zinc ribbon domains can thus reveal that two different structural submotifs interact with DNA at the major groove. While specific interactions cannot be deduced from the sidechain orientations given by an Alphafold2 or RoseTTAfold model, a putative interaction can be clearly deduced between DNA and an alpha helix located approximately 20-30 residues upstream from the RD boundary helix, as used herein this submotif is referred to as the “Helix.”
[0254] In some recombinases, the portion of the helix which may contact the DNA comprises residues 231-237 for Bxbl or residues 229-235, 241-247, 241-247, 221-227, 230-236, 234-240, 218- 224, 241-247, 223-228, and 218-224 for Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, and Ssalnt, respectively. A second structurally conserved motif is a beta hairpin locatedbetween the two inner cysteine residues of a highly conserved tetracysteine motif, as used herein this submotif is referred to as the “Hairpin.”
[0255] Lastly, the first 20 residues comprising the RD portion of LSRs contains no defined secondary structure and appears to contact DNA at the minor groove, as used herein this submotif is referred to as the “Loop.”
[0256] Once each of these submotifs is defined, residues from each submotif can be modified to alter the sequence specificity for their respective attB and attP sites in a modular fashion. In other words, permutations of mutant “loop”, “Helix”, “Hairpin” submotifs and combinations thereof may be assembled to target attB and attP sites that diverge from the natural attB and attP sequences. In some embodiments, permutations of mutant “loop”, “Helix” or “Hairpin” submotifs target attB and attP sites that diverge from the natural attB and attP sequences.
[0257] In some embodiments, the “loop”, “Helix” and “Hairpin” regions are predicted by using sequence alignments manually improved with secondary structure predictions.
[0258] In some embodiments, disclosed are methods for generating recombinase variants wherein the variant recombinase exhibits altered DNA target site specificity compared to a different non-variant / wild-type counterpart or the starting recombinase. In some embodiments, the method comprises: a) generating a library of variant recombinases from a starting recombinase, wherein the variant recombinases include on or more modifications of amino acids that control DNA target site specificity for the starting recombinase; b) screening said variant recombinase library for altered DNA specificity as compared to a different non-variant recombinase and / or the starting recombinase; and c) selecting a variant recombinase from b.
[0259] In some embodiments, disclosed is a method for generating a recombinase variant having an altered DNA target site specificity, comprising: 1) identifying positions of amino acids that control DNA target site specificity in a starting recombinase; 2) randomizing the positions of amino acids that control DNA target site specificity to generate recombinase variants; and 3) selecting a recombinase variant having the altered DNA specificity. In some embodiments, the method further comprises determining the sequence of the recombinase variant. Mapping the interactions between specific parts of the DNA target site and the helix, hairpin, and loop is a key attribute disclosed here.
[0260] Disclosed is a method for generating a large serine recombinase variant that can target a desired DNA sequence, comprising: 1) Dividing the desired genomic DNA target site into 1-4separate portions that correspond to the region recognized by i) the left halfsite ZD domain, ii) the left halfsite RD domain, iii) the right halfsite RD domain and iv) the right halfsite ZD domain; 2) for each portion of the desired target site recognized by a ZD domain, generating a library of recombinase variants with at least two amino acid changes within the hairpin region of the recombinase and selecting this library for members that can recombine a target site comprising the relevant portion of the desired genomic target site.
[0261] Disclosed is a method for generating a large serine recombinase variant that can target a desired DNA sequence, comprising: 1) Dividing the desired genomic DNA target site into 1-4 separate portions that correspond to the region recognized by i) the left halfsite ZD domain, ii) the left halfsite RD domain, iii) the right halfsite RD domain and iv) the right halfsite ZD domain; 2) for each portion of the desired target site recognized by an RD domain, generating a library of recombinase variants with at least two amino acid changes with the helix region and / or loop region, and selecting members of this library that can recombine target sites comprising the relevant portion of the genomic target site.
[0262] In some embodiments, the method further comprises combining selected recombinase variants selected against two regions of the same halfsite (i.e., modifications in the left helix-loop, helix-hairpin, hairpin-loop or helix, hairpin and loop or a modification in the right helix-loop, helixhairpin, hairpin-loop or helix, hairpin and loop) and then testing such composite variants with a target site comprising the relevant desired halfsite.
[0263] In some embodiments, the method further comprises combining selected recombinase variants selected against two regions of different halfsites (a modification in the left or right helix and modification in the left or right loop; modification in the left or right helix and modification in the left or right hairpin; modification in the left or right hairpin and modification in the left or right loop helix; modification in the left or right hairpin, modification in the left or right loop and modification in the left or right helix) and then testing such composite variants with a target site comprising the relevant desired halfsite.
[0264] In some embodiments, the method further comprises combining selected recombinase variants selected against two regions of the same halfsite (e.g., modifications in the left helix-loop, helix-hairpin, hairpin-loop or helix, hairpin and loop; or modifications in the right helix-loop, helixhairpin, hairpin-loop or helix, hairpin and loop) and then testing such composite variants with aminoacid changes in both RD and ZD domains (left and right) for recombination activity with a target site comprising the relevant desired halfsite.
[0265] In some embodiments, a method for a recombinase variant comprises 1) identifying the positions of amino acids that are randomized in the starting variant library 2) identifying the positions of the DNA target site recognized by the loop, helix, hairpin and / or combinations thereof that will be replaced by the corresponding region of the desired DNA target sequence, and 3) applying a stepwise strategy for combining and screening large serine recombinase variants identified from the directed evolution selections in order to obtain a pair of large serine recombinases that can dimerize to recognize the desired DNA sequence. In some embodiments, directed evolution selections using variant libraries with randomized residues within the helix, loop, hairpin and / or combinations thereof can also be performed systematically against large numbers of different DNA target sites in order create an archive of pre-existing and pre-characterized large serine recombinase variants that can be combined using the same stepwise strategy in order to generate large serine recombinase variants that have altered DNA target site specificity. In some embodiments, if the amino acid changes corresponding to each variant and the DNA target site specificity of each variant are known, this stepwise strategy to combine the variants can be utilized. In some embodiments screening variants of loop, helix, or hairpin regions against the appropriate portions of the desired genomic target site is performed computationally using computer software such as Rosetta.
[0266] In some embodiments, the method further comprises delivering a mixture of a recombinase variant active on the desired left halfsite, a recombinase variant active on the right halfsite, and an appropriate donor construct into eukaryotic cells and then measuring the amount of recognized integration of the donor construct into the desired genomic location.
[0267] This method for generating large serine recombinase variants can be used for any recombinase and specifically for large serine recombinases such as Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, A118, or TP901 as exemplified below by Bxbl.
[0268] The recombination sites used as substrates in the methods disclosed herein include, but are not limited to, wild-type attB, wild-type attP, pseudo-attB, pseudo-attP, non-wildtype attB, nonwildtype attP, and combinations thereof. Recombination target sites may be identified, using methods described herein, in the genome of essentially any target cell, including, but not limited to, prokaryotes, eukaryotes, E. coli or other bacteria, yeast or other fungi, plant protoplast or other plantcells, human cells, rodent cells, as well as K562 cells or any other mammalian or, more generally, animal cell line.
[0269] As discussed above, variant recombinases identified by the methods disclosed herein may have increased, decreased, or similar recombination efficiencies relative to the wild-type recombinase. Disclosed are nucleic acid sequences encoding the polypeptide sequences of the variant recombinases.
[0270] In some embodiments, the recombinase variant comprises a codon optimized ORF cloned into the plasmid. In some embodiments, the plasmid containing the ORF comprises a pBR322 origin of replication. In some embodiments, the plasmid containing the ORF contains a promoter. In some embodiments, the ORF comprises a L-rhamnose inducible pRhaBAD promoter. In some embodiments, the “loop” submotif is modified to diversify residues 154, 155, 156, 157, 158 and 159. In some embodiments, the “loop” submotif is modified to diversify residues 154, 155, 156, 157, 158 and 159 using forward primer (SEQ ID No. 4) and reverse primer (SEQ ID No. 5). In some embodiments, the “loop” submotif is modified using primers that target residues 154-159. In some embodiments, the “alpha-helix” submotif is modified using primers that target residues 231-234, and residues 236-237. In some embodiments, the “Beta-hairpin” submotif is modified with primers that target residues 314, 316, 318, 321, 323, and 325. In some embodiments, the “Beta-hairpin” submotif is modified with primers that target residues 314, 316, 318, 321, 323, and 325 using an NNK randomization scheme and / or residue 322 is randomized using an SSK randomization scheme.
[0271] In some embodiments, the Helix region of the recombinase variant is modified to recognize the designed non-wildtype attB (SEQ ID No. 8) and / or non-wildtype attP (SEQ ID No. 9) sequences with the desired DNA bases replacing the variable region designated with one more “N” characters. In some embodiments, in each attB or attP site, the right trinucleotide is the reverse complement of the left trinucleotide.
[0272] In some embodiments, the “alpha-helix” submotif is modified with primers designed to randomize residues 231,232,233,234,236 and 237, using an NNK scheme. In some embodiments, the “alpha-helix” submotif is modified with primers designed to randomize residues 231,232,233,234,236 and 237, using an NNK scheme using forward primer (SEQ ID No. 10) and reverse primer (SEQ ID No. 11).
[0273] In some embodiments, the Hairpin submotif is modified to recognize designed attB (SEQ ID No. 12) and / or attP (SEQ ID No. 13) sequences.
[0274] In some embodiments, the “Hairpin” submotif is modified to randomize residues 314, 316, 318, 321, 323 and 325 using an NNK scheme and / or residue 322 was randomized using an SSK scheme. In some embodiments, the “Hairpin” submotif is modified to randomize residues 314, 316, 318, 321, 323 and 325 using an NNK scheme and / or residue 322 was randomized using an SSK scheme using forward primer (SEQ ID No. 14) and reverse primer (SEQ ID No. 15).
[0275] In some embodiments, the variant recombinase mediates efficient integration in the host genome e.g., human cell environment) at any site. In some embodiments, the variant recombinase mediates efficient integration in the host genome (e.g., human cell environment) at non-wildtype attB and non-wildtpye attP sites.
[0276] Disclosed is a method of converting attP and attB sequences in a cell into attL and attR sequences encoding an in-frame peptide insertion.
[0277] In some embodiments, single stranded circular DNA is used as the donor for a variant integrase. In these embodiments a host mammalian cell is transfected with circular ssDNA and inside the host, the ssDNA is converted to dsDNA in the nucleus.
[0278] In some embodiments, the variant recombinase reassembles a target gene in the host In some embodiments, the host is transfected with a first target plasmid comprising a non-functional gene and a stuffer sequence; the host is then transfected with a second target plasmid designed to excise the stuffer sequence and reassemble the non-functional gene to produce a functional gene. In some embodiments, the stuffer sequence is excised from a plasmid. In some embodiments, a circular DNA construct that cannot replicate and does not express an antibiotic resistance gene is inserted into the plasmid.
[0279] In some embodiments, the variant recombinase reassembles a target gentamicin resistance gene in the host. In some embodiments, the host is transfected with a first target plasmid comprising a non-functional gentamicin resistance gene and a stuffer sequence; the host is then transfected with a second target plasmid designed to excise the stuffer sequence and reassemble the non-functional gentamicin resistance gene to produce a functional gentamicin resistance gene.
[0280] In some embodiments, disclosed is a plasmid comprising a stuffer sequence that was flanked with modified attB sequence at its 5’ end, and with a modified attP sequence at its 3’ end. In some embodiments, disclosed is a plasmid comprising, or consisting of a variant recombinase and an attL sequence encoding an in-frame peptide insertion, giving rise to an active aacCl.
[0281] In some embodiments, disclosed is a method of directed evolution comprising delivering two separate plasmids to the same bacterial cell where the first plasmid expresses a Bxbl variant being tested and the second target plasmid contains an attB and attP version of a desired target site; recombining the Bxbl variant and the target plasmid such that the target plasmid is able to express a functional antibiotic resistance gene; contacting the bacterial cell with an antibiotic. In some embodiments, the attP target site has the same bases as the attB target site at positions +6 to +7, -6 to -7, +11 to +9, and -11 to -9. In some embodiments, the sequences of the hairpin target sites are also the same as in the attB target but have 5 bp inserted between the hairpin and helix targets. In some embodiments, recombination excises a stuffer sequence and leaves an inert circular DNA construct that cannot replicate and does not express an antibiotic resistance gene.
[0282] Disclosed are methods for engineering variant recombinases against desired sites comprising: 1) identify potential sites within desired region with G at -4 and at +4; 2) testing variant recombinases in the hairpin region against portions of target site from -12 to -19 and +12’ to +19’; and 3) testing variant recombinases in helix and / or loop domain against portions of target sites from - 11 to -6 and from +11’ to +6’.
[0283] In some embodiments, promising variants for different quarter sites are combined and tested in a stepwise fashion. In some embodiments, promising variants for different RD and ZD domains sites are combined and tested in a stepwise fashion. See Fig. 3. In some embodiments, promising variants for different loop, hairpin and / or helix regions are combined and tested in a stepwise fashion.
[0284] Disclosed is a variant recombinase comprising loop, hairpin and / or helix regions domains from different recombinases. In some embodiments, the loop, hairpin and / or helix regions are from a different wildtype recombinase. In some embodiments, the loop, hairpin and / or helix regions are from a different modified recombinase. In some embodiments, the loop, hairpin and / or helix regions comprise one or more amino acid changes relative to the wildtype counterpart domains. In some embodiments, the loop, hairpin and / or helix regions are from Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901.
[0285] Disclosed are proteins that are engineered using these methods.
[0286] Directed Evolution System
[0287] Directed evolution is the laboratory process by which biological entities with desired traits are created through iterative rounds of genetic diversification and library screening or selection. Directed evolution has become one of the most useful and widespread tools in basic and applied biology. Directed evolution consists of subjecting a gene to iterative rounds of mutagenesis (creating a library of variants), selection (expressing those variants and isolating members with the desired function) and amplification (generating a template for the next round). In one embodiment, disclosed is a method for identifying a variant recombinase using directed evolution. Directed evolution performed in bacteria is generally limited by the transformation efficiency of those bacteria and is typically used with randomized “libraries” of up to 10A9 or even 10A10 members. This roughly corresponds to libraries with 7 amino acids positions completely randomized.
[0288] In some embodiments, disclosed is a method, composition, and system for controlling large serine recombinase DNA target site specificity. The regions for controlling large serine recombinase DNA target site specificity are referred to as loop, helix, and hairpin. The hairpin region is within the zinc ribbon domain (“ZD”), while the loop and helix regions are both within the recombinase domain (“RD”). These regions are small enough to allow standard directed evolution techniques to be employed to completely randomized and re-select enough residues within each region to completely alter the DNA target site specificity corresponding to the region.
[0289] Fig. 4 is an overview of the directed evolution system used to find novel recombinase variants. In some embodiments, two different plasmids are delivered to the same bacterial cell where the first plasmid expresses a recombinase variant being tested, and the second target plasmid contains an attB and attP version of a desired target site. In some embodiments, the bacteria directed evolution system comprises a synthesized target site where key parts of the halfsites are inverted repeats of each other so that only a single recombinase variant is used. If the recombinase variant is able to interact with the attB and attP sequence in the target plasmid within the same bacterial cell, then the recombinase variant will recombine the target plasmid and the recombined target plasmid is able to express a functional antibiotic resistance gene and the cell expressing the recombinase variant can survive selection with the relevant antibiotic. In some embodiments, the plasmids encoding the recombinase variants in cells that survive the antibiotic selection are then sequenced and analyzed. In some embodiments, the attP target site has the same bases as the attB target site at positions +6 to +7, -6 to -7, +11 to +9, and -11 to -9. In some embodiments, the sequences of the hairpin target sites are also the same as in the attB target but have 5 bp inserted between the hairpin and helix recognizes as indicated in Fig. 1C. In some embodiments, successful recombination excises a stuffer sequence andleaves an inert circular DNA construct (not pictured in the Fig.) that cannot replicate and does not express an antibiotic resistance gene relevant to the selection pressure being applied.
[0290] Disclosed is a method for recombining a DNA sequence in a target plasmid, a population of cells is typically provided wherein cells of the population comprise a variant plasmid and a target plasmid (e.g., a plasmid containing an attB and attP version of a desired target site, whereby the attB and attP versions of a desired target site flank a DNA sequence that disrupts the open reading frame of a selectable marker). In some embodiments, the variant plasmid recombines the target plasmid, and the recombined target plasmid is able to express a selectable marker. Specifically, in some embodiments, the variant recombinase will recombine the target plasmid and the recombined target plasmid is now able to express a functional antibiotic resistance gene and the recombined target plasmid can survive selection with the relevant antibiotic.
[0291] Disclosed is a bacterial cell comprising a first plasmid and a second plasmid wherein the first plasmid and second plasmid are different. In some embodiments, the first plasmid expresses a recombinase variant being tested. In some embodiments, the second target plasmid contains an attB and attP version of a desired target site. In some embodiments, the attP target site has the same bases as the attB target site at positions +6 to +7, -6 to -7, +11 to +9, and -11 to -9. In some embodiments, the sequences of the hairpin target sites are also the same as in the attB target but have 5 bp inserted between the hairpin and helix target subsites as indicated in Fig. 1. In some embodiments, successful recombination excises the “stuffer” sequence and leaves an inert circular DNA construct (not pictured in the Fig.) that cannot replicate and does not express an antibiotic resistance gene.
[0292] Disclosed is a method for directed evolution comprising: 1) delivering two different plasmids to the same bacterial cell where the first plasmid expresses a recombinase variant being tested, and the second target plasmid contains an attB and attP version of a desired target site; 2) recombining the recombinase variant with the attB and attP sequence in the target plasmid; 3) expressing by the recombined target plasmid a functional antibiotic resistance gene. In some embodiments, the method further comprises, exposing the bacterial cell to an antibiotic. In some embodiments, the method further comprises, sequencing the plasmids encoding the recombinase variants in cells that survive the antibiotic selection. In some embodiments, the method further comprises, excising a stuffer sequence and leaving an inert circular DNA construct that cannot replicate and does not express an antibiotic resistance gene.
[0293] In some embodiments, the plasmids encoding the Bxbl variants in cells that survive the antibiotic selection comprise an attP target site that has the same bases as the attB target site at positions +6 to +7, -6 to -7, +11 to +9, and -11 to -9. In some embodiments, the plasmids encoding the Bxbl variants in cells that survive the antibiotic selection comprise sequences of the hairpin target sites that are the same as in the attB target but have 5 bp inserted between the hairpin and helix target subsites as indicated in Fig. 3. In some embodiments, the plasmids encoding the Bxbl variants in cells that survive the antibiotic selection comprise an attP target site that has the same bases as the attB target site at positions +6 to +7, -6 to -7, +11 to +9, and -11 to -9 and the hairpin target sites that are the same as in the attB target, but have 5 bp inserted between the hairpin and helix target subsites as indicated in Fig. 3.
[0294] Directed evolution selections using variant libraries with randomized residues within the helix, loop, or hairpin can also be performed systematically against large numbers of different DNA target sites in order create an archive of pre-existing and pre-characterized large serine recombinase variants that can be combined using the same stepwise strategy in order to generate large serine recombinase variants that have altered DNA target site specificity. As long as the amino changes corresponding to each variant and the DNA target site specificity of each variant are known, this stepwise strategy to combine the variants can be utilized.
[0295] As discussed, the method Applicant used successfully to reprogram Bxbl will work for any other recombinase by modifying the binding (or interacting) regions that control DNA target site specificity for the relevant portions of the wildtype recombinase. In some embodiments, the binding (or interacting) regions will be completely randomized. In some embodiments, residues with specific binding (or interacting) capability should be randomized. In some embodiments, the specific binding (or interacting) regions should be changed to match the desired endogenous target site. In some embodiments, if there is more than one binding (or interacting) region all binding (or interacting) regions are randomized. In some embodiments, the binding (or interacting) regions each comprise at least one modification. Recombinase variant sequences are then selected. In some embodiments, the recombinase variant sequences that match enriched 4-residue motifs will be screened in the plasmid-based screening system shown in Figure 12 A. In some embodiments, target sites will be designed according to the scheme shown in Figure 12 A, except WT_ZD and WT_RD sequences will use the relevant portions of the natural target site for the relevant integrase instead of the relevant portions of the natural Bxbl target site. In some embodiments, desired endogenous sites are recognized with these new recombinases using the same overall strategy diagramed in Figure 3,except that the boundaries of the “quarter sites” may differ slightly for some of these alternative recombinases.
[0296] Modular Combinatorial Generation of Variant Recombinases
[0297] In some embodiments a variant recombinase as described herein comprises different DNA binding modules. Disclosed is a variant recombinase comprising RD and / or ZD domains from different recombinases. In some embodiments, the RD and / or ZD domains of the variant recombinase are from a different wildtype recombinase. In some embodiments, the RD and / or ZD domains of the variant recombinase are from a different modified recombinase. In some embodiments, the RD and / or ZD domains of the variant recombinase comprise one or more amino acid changes relative to the wildtype counterpart domains. In some embodiments, the RD and / or ZD domains are from Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901.
[0298] In some embodiments, the DNA binding modules are comprised of combinations from the left half side and right half site binding domains. In some embodiments, the DNA binding modules are comprised of combinations from one or more quarter site of the left half side and from one or more quarter site of the right half site. In some embodiments, the DNA binding modules are comprised of helix, loop, and / or hairpin regions as described herein. In some embodiments, the DNA binding modules are comprised of helix, loop, and / or hairpin regions as described herein from combinations of the left and right half site. In some embodiments, the DNA binding modules are comprised of RD and ZD domains. In some embodiments, the DNA binding modules contain 1, 2, 3, 4, 5, 6, or 7 mutations as described herein. In some embodiments, the DNA binding modules are from Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901.
[0299] As can be appreciated, modified binding regions from recombinases can be combined in any combination. In some embodiments, modified binding regions from the left half site and right half site binding domains are obtained from the same wildtype and / or non- wildtype recombinases. In some embodiments, modified binding regions from one or more quarter site of the left half site and from one or more quarter site of the right half site are obtained from the same wildtype and / or nonwildtype recombinases. In some embodiments, modified binding regions from the RD and ZDdomains are obtained from the same wildtype and / or non-wildtype recombinases. In some embodiments, modified binding regions from the helix, loop, and / or hairpin regions are obtained from the same wildtype and / or non-wildtype recombinases.
[0300] In some embodiments, modified binding regions from the left half side and right half site binding domains are obtained from different wildtype and / or non-wildtype recombinases. In some embodiments, modified binding regions from one or more quarter site of the left half side and from one or more quarter site of the right half site are obtained from different wildtype and / or nonwildtype recombinases. In some embodiments the RD and ZD domains are obtained from different wildtype and / or non-wildtype recombinases. In some embodiments, modified binding regions from the helix, loop, and / or hairpin regions are obtained from different wildtype and / or non-wildtype recombinases.
[0301] In some embodiments the left half site and right half site binding domains are obtained from the same or different wildtype and / or non- wildtype recombinases. In some embodiments, binding regions from one or more quarter site of the left half site and from one or more quarter site of the right half site are obtained from the same or different wildtype and / or non-wildtype recombinases. In some embodiments the left and right RD domains are obtained from the same or different wildtype and / or non-wildtype recombinases. In some embodiments the left and right ZD domains are obtained from the same or different wildtype and / or non-wildtype recombinases. In some embodiments the left and right helix, loop, and / or hairpin regions are obtained from the same or different wildtype and / or non-wildtype recombinases.
[0302] In some embodiments, the RD and ZD domains are structurally independent from each other so that RD and ZD domains from different wildtype recombinases can be combined and retain their original DNA target site specificity and original function present in the wildtype recombinase. In some embodiments, the helix, loop, and / or hairpin regions are structurally independent from each other so that helix, loop, and / or hairpin regions from different wildtype recombinases can be combined and retain their original DNA target site specificity and original function present in the wildtype recombinase. In some embodiments, left half site and right half site binding domains or one or more quarter site of the left half site and from one or more quarter site of the right half site binding domains are structurally independent from each other so that left half side and right half site binding domains or one or more quarter site of the left half side and from one or more quarter site of the right half site binding domains from different wildtype recombinases can be combined and retain their original DNA target site specificity and original function present in the wildtype recombinase.
[0303] In some embodiments, RD domain variants are combined with ZD domain variants to yield a single full-length recombinase polypeptide. In some embodiments, a full-length recombinase variant that recognizes the left halfsite is mixed with a separate / different full-length recombinase variant that recognizes the right halfsite. In some embodiments, helix, loop, and / or hairpin domain variants are combined with helix, loop, and / or hairpin domain variants to yield a single full-length recombinase polypeptide.
[0304] In some embodiments, a variant recombinase comprises RD and / or ZD domains each from the same recombinase, or each from a different recombinase. In some embodiments, a variant recombinase comprises right helix, right loop, and / or right hairpin and / or left helix, left loop, and / or left hairpin domains each from the same recombinase, or each from a different recombinase.
[0305] Each of the left halfsite ZD domain, left halfsite RD domain, right halfsite RD domain and right halfsite ZD domain can come from a different recombinase, the same recombinase or combinations thereof. Each of the left helix, left loop, left hairpin, right helix, right loop, right hairpin domain can come from a different recombinase, the same recombinase or combinations thereof.
[0306] In some embodiments, the RD (left / right) and / or ZD (left / right) domains of a variant recombinase are each from a wildtype recombinase, each from a non-wildtype recombinase or a combination thereof. In some embodiments, the loop (left / right) and / or helix (left / right) and / or hairpin (left / right) domains of a variant recombinase are each from a wildtype recombinase, each from a non-wildtype recombinase or a combination thereof.
[0307] In some embodiments, one or more RD and / or one or more ZD domains of a variant recombinase are from a wildtype recombinase, from a non-wildtype recombinase, or a combination thereof. In some embodiments, one or more loop and / or one or more hairpin and / or one or more helix domains of a variant recombinase are from a wildtype recombinase, from a non-wildtype recombinase, or a combination thereof. In some embodiments, the left halfsite RD domain is from a wild type and the right halfsite RD domain from a non-wildtype. In some embodiments, the left halfsite ZD domain is from a wild type and the right ZD domain halfsite from a non-wildtype. In some embodiments, each of the modified left halfsite ZD domain, left halfsite RD domain, right halfsite RD domain and / or right halfsite ZD domain are from the same modified recombinase, or from different modified recombinases. In some embodiments, the left halfsite ZD domain, left halfsite RD domain, right halfsite RD domain and / or right halfsite ZD domain comprise one or moreamino acid changes relative to the wildtype counterpart domain. In some embodiments, the system comprises a mixture of different recombinase variants.
[0308] In some embodiments, the left loop region is from a wild type and the right loop region is from a non-wildtype. In some embodiments, the left helix region is from a wild type and the right helix region is from a non-wildtype. In some embodiments, the left hairpin region is from a wild type and the right hairpin region is from a non-wildtype. In some embodiments, each of the modified left helix, left loop, left hairpin, right helix, right loop, and / or right hairpin regions are from the same modified recombinase, or from different modified recombinases. In some embodiments, the left helix, left loop, left hairpin, right helix, right loop, and / or right hairpin comprise one or more amino acid changes relative to the wildtype counterpart domain.
[0309] In some embodiments, a first halfsite ZD domain is from a wild-type recombinase and a second halfsite ZD domain is not from a wild-type recombinase. In some embodiments, the first halfsite ZD domain is the left halfsite ZD domain. In some embodiments, the second halfsite ZD domain is the right halfsite ZD domain. In some embodiments, a first halfsite ZD domain is from a first recombinase and a second halfsite ZD domain is from a second recombinase wherein the first and second recombinase are different. In some embodiments, the first recombinase is a wild-type or variant Bxbl recombinase. In some embodiments, the second recombinase is a wild-type or variant Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, Ssalnt, or Dn29, PhiRvl, Al 18, or TP901 recombinase.
[0310] In some embodiments, a first halfsite RD domain is from a wild-type recombinase and a second halfsite RD domain is not from a wild-type recombinase. In some embodiments, the first halfsite RD domain is the left halfsite RD domain. In some embodiments, the second halfsite RD domain is the right halfsite RD domain. In some embodiments, a first halfsite RD domain is from a first recombinase and a second halfsite RD domain is from a second recombinase wherein the first and second recombinase are different. In some embodiments, the first recombinase is a wild-type or variant Bxbl recombinase. In some embodiments, the second recombinase is a wild-type or variant Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901, or Ssalnt recombinase.
[0311] In some embodiments, a first halfsite loop, helix and / or hairpin region is from a wild-type recombinase and a second halfsite loop, helix and / or hairpin region is not from a wild-type recombinase. In some embodiments, a first halfsite loop, helix and / or hairpin region is from a firstrecombinase and a second halfsite loop, helix and / or hairpin region is from a second recombinase wherein the first and second recombinase are different. In some embodiments, the first recombinase is a wild-type or variant Bxbl recombinase. In some embodiments, the second recombinase is a wild-type or variant Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901or Ssalnt recombinase.
[0312] In some embodiments, a first halfsite ZD domain is from a wild-type recombinase and a second halfsite RD domain is not from a wild-type recombinase. In some embodiments, the first halfsite ZD domain is the left halfsite ZD domain. In some embodiments, the second halfsite RD domain is the right halfsite RD domain. In some embodiments, a first halfsite ZD domain is from a first recombinase and a second halfsite RD domain is from a second recombinase wherein the first and second recombinase are different or are the same. In some embodiments, the first recombinase is Bxbl. In some embodiments, the second recombinase is PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901, or Ssalnt . In some embodiments, the first recombinase is a wild-type or variant Bxbl recombinase. In some embodiments, the second recombinase is a wildtype or variant PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901, or Ssalnt recombinase. In some embodiments, the first recombinase is Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901, or Ssalnt. In some embodiments, the first recombinase is a wild-type or variant Bxbl recombinase. In some embodiments, the second recombinase is a wild-type or variant Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901, or Ssalnt recombinase.
[0313] In some embodiments, the variant recombinase comprises a hairpin region within a zinc ribbon domain (ZD) region. In some embodiments, the ZD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart. In some embodiments, the ZD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart in the binding region.
[0314] In some embodiments, the ZD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart such that hairpin (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart. In other embodiments, the hairpin (leftand / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart in the binding region. In other embodiments, the hairpin (left and / or right) comprises 1, 2, 3, 4, 5, 6,7 ,8 ,9 ,10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to its wildtype counterpart in the binding region. In some embodiments, the ZD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 308-325, 307-326, 320-343, 306-323, 286-302, 322-339, 317-343, 289-313, 333-350, 296-315, or 283-311 in the hairpin region. In some embodiments, the ZD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions (313, 312, 315,317, 322, 324, 326), (326,328,330, 332, 339, 341, 343), (311, 312,314, 315, 318, 321, 323), (290, 292, 294, 297, 299, 300, 302), (328, 330, 333, 334, 335, 337, 339), (323, 325, 328, 326, 339, 341, 343), (295, 297, 298, 299, 307, 311, 313), (337, 339, 341, 345, 346, 348, 350), (302, 304, 310,311, 313, 315), or (289, 291, 293, 295, 305,307, 309) in the hairpin region. In some embodiments, the ZD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 308, 309, 310, 311, 312, 313, 314, 316, 318, 321, 322, 323, or 325 in the hairpin region. In some embodiments, the ZD variant has 90-99% identity compared to its wildtype counterpart. In some embodiments, the variant hairpin domain has 90-99% identity compared to its wildtype counterpart. In some embodiments, the ZD variant has 90-99% identity compared to its wildtype counterpart and yet still possess broader or narrower DNA recognition specificity (binding specificity) compared to the wild-type enzyme and / or greater or lesser catalytic activity toward a particular DNA sequence, including a wild-type or non-wild-type recombinase recognition site. In some embodiments, the hairpin domain has 90-99% identity compared to its wildtype counterpart and yet still possess broader or narrower DNA recognition specificity (binding specificity) compared to the wild-type enzyme and / or greater or lesser catalytic activity toward a particular DNA sequence, including a wild-type or non-wild-type recombinase recognition site.
[0315] In some embodiments, the variant recombinase comprises a helix region within a recombinase region (RD). In some embodiments, the variant recombinase comprises a loop region within the RD. In some embodiments, the variant recombinase comprises a helix and loop region within the RD. In some embodiments, the RD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart. In some embodiments, the RD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart in the binding region. In other embodiments, the loop (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart. In other embodiments, the loop (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart in the binding region. In yet otherembodiments the helix (left and / or right) comprises 1, 2, 3, 4, 5, 6, or 7 amino acids changes relative to its wildtype counterpart. In other embodiments, the helix (left and / or right) comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart in the binding region. In some embodiments, the RD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart such that the loop (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart and / or the helix (left and / or right) comprises 1, 2, 3, 4, 5, 6, or 7 amino acids changes relative to its wildtype counterpart. In some embodiments, the helix comprises an asparagine amino acid residue (N) at the 4thposition of the helix region. In some embodiments, the helix comprises a tryptophan (W) amino acid residue at the 3rd position. In some embodiments, the helix comprises an asparagine amino acid residue (N) at the 4thposition and a tryptophan amino (W) acid residue at the 3rd position of the helix region. In some embodiments, the helix comprises an asparagine amino acid residue (N) at position 234. In some embodiments, the helix comprises two lysine residues (K) at the 6thand 7thposition of the helix. In some embodiments, the helix comprises a lysine residue (K) at the 6thor 7thposition of the helix. In some embodiments, the helix comprises an asparagine amino acid residue (N) at the 4thposition, a tryptophan amino (W) acid residue at the 3rd position of the helix region, two lysine residues (K) at the 6thand 7thposition of helix and combinations thereof.
[0316] In some embodiments, the RD variant has 90-99% identity compared to its wildtype counterpart. In some embodiments, the variant helix domain has 90-99% identity compared to its wildtype counterpart. In some embodiments, the variant loop domain has 90-99% identity compared to its wildtype counterpart. In some embodiments, the variant helix domain and variant loop domain have 90-99% identity compared to its wildtype counterpart. In some embodiments, the variant helix domain variant loop domain and variant hairpin domain have 90-99% identity compared to its wildtype counterpart. In some embodiments, the RD variant has 90-99% identity compared to its wildtype counterpart and yet still possess broader or narrower DNA recognition specificity (binding specificity) compared to the wild-type enzyme and / or greater or lesser catalytic activity toward a particular DNA sequence, including a wild-type or non-wild-type recombinase recognition site. In some embodiments, the helix and / or loop domain(s) has 90-99% identity compared to its wildtype counterpart and yet still possess broader or narrower DNA recognition specificity (binding specificity) compared to the wild-type enzyme and / or greater or lesser catalytic activity toward a particular DNA sequence, including a wild-type or non-wild-type recombinase recognition site.
[0317] In some embodiments, the loop (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart. In some embodiments, the loop (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart at positions 151- 157, 155-160, 155-160, 167-172, 167-172, 156-161, 162-167, 179-183, 161-166, 164-169, or 154- 159 in the loop region. In some embodiments, the loop (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart at positions 154, 155, 156, 157, 158, and 159 in the loop region.
[0318] In some embodiments, the helix (left and / or right) comprises 1, 2, 3, 4, 5, 6, or 7 amino acids changes relative to its wildtype counterpart. In some embodiments, the helix (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to its wildtype counterpart at positions 229-235, 241-247, 241-247, 221-227, 230-236, 234-240, 218-224, 241-247, 223-228, 218-224, or 231- 237 of the helix region. In some embodiments, the helix (left and / or right) comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to its wildtype counterpart at positions 231, 232, 233, 234, 236, and 237 of the helix region.
[0319] In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart in the helix region and 1, 2, 3, 4, 5, or 6 amino acid changes in the loop region relative to its wildtype counterpart. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart in the helix region and 1, 2, 3, 4, 5, 6,7 ,8 ,9 ,10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes in the hairpin region relative to its wildtype counterpart. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6,7 ,8 ,9 ,10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to its wildtype counterpart in the hairpin region and 1, 2, 3, 4, 5, or 6 amino acid changes in the loop region relative to its wildtype counterpart. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart in the helix region, 1, 2, 3, 4, 5, 6,7 ,8 ,9 ,10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes in the hairpin region relative to its wildtype counterpart, and 1, 2, 3, 4, 5, or 6 amino acid changes in the loop region relative to its wildtype counterpart.
[0320] In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 229-235, 241-247, 241-247, 221-227, 230-236, 234-240, 218-224, 241-247, 223-228, 218-224, 231-237 or (231, 232, 233, 234, 236, and 237) of the helix region. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions (308, 309, 310, 311, 312, 313,314, 316, 318, 321, 322, 323, 325), (313, 312, 315,317, 322, 324, 326), (326,328,330, 332, 339, 341, 343), (311, 312,314, 315, 318, 321, 323), (290, 292, 294, 297, 299, 300, 302), (328, 330, 333, 334, 335, 337, 339), (323, 325, 328, 326, 339, 341, 343), (295, 297, 298, 299, 307, 311, 313), (337, 339, 341, 345, 346, 348, 350), (302, 304, 310,311, 313, 315), (289, 291, 293, 295, 305,307, 309), or (308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325) in the hairpin region. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 151-157, 155-160, 155-160, 167-172, 167- 172, 156-161, 162-167, 179-183, 161-166, 164-169 or 154-159 in the loop region. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 231, 232, 233, 234, 236, and 237 of the helix region and 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 308, 309, 310, 311, 312, 313, 314, 316, 318, 321, 322, 323, and 325 in the hairpin region. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 231, 232, 233, 234, 236, and 237 of the helix region and 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 154, 155, 156, 157, 158, and 159 in the loop region. In some embodiments, the recombinase variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 308, 309, 310, 311, 312, 313, 314, 316, 318, 321, 322, 323, and 325 in the hairpin region and 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 154, 155, 156, 157, 158, and 159 in the loop region. In some embodiments, the ZD variant comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 231, 232, 233, 234, 236, and 237 of the helix region, 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 308, 309, 310, 311, 312, 313, 314, 316, 318, 321, 322, 323, and 325 in the hairpin region and 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its wildtype counterpart at positions 154, 155, 156, 157, 158, and 159 in the loop region.
[0321] In some embodiments, the variant recombinase disclosed herein comprises at least two mutations relative to the wild-type recombinase in the binding regions. In some embodiments, the variant recombinase comprises 1, 2, 3, 4, 5, 6, or 7 mutations relative to the wild-type recombinase in the binding regions. In some embodiments, the variant recombinase disclosed herein comprises binding regions in a loop region, a helix region, and a hairpin region. In some embodiments, the variant recombinase disclosed herein comprises at least two mutations relative to the wild-type recombinase in one of the three binding regions In some embodiments, the variant recombinase disclosed herein comprises at least two mutations relative to the wild-type recombinase in two of thethree binding regions. In some embodiments, the variant recombinase disclosed herein comprises at least two mutations relative to the wild-type recombinase in each one of the binding regions. In some embodiments, the variant recombinase comprises 1, 2, 3, 4, 5, 6, or 7 mutations relative to the wildtype recombinase in each one of the binding regions. In some embodiments, the variant recombinase disclosed herein comprises at least two mutations relative to the wild-type recombinase in one or more binding region.
[0322] In some embodiments, the variant recombinase disclosed herein comprises binding regions small enough to allow standard directed evolution techniques to be employed to completely randomize and re-select enough residues within each region to completely alter the DNA target site specificity corresponding to the region. In some embodiments, the variant recombinase disclosed herein comprises binding regions small enough to allow standard directed evolution techniques to be employed to completely randomize and re-select enough residues within each region to partially alter the DNA target site specificity corresponding to the region.
[0323] In some embodiments, disclosed is a variant recombinase wherein the binding regions comprise one or more of a variant loop, helix, or hairpin amino acid sequence as shown in any of Figures 5B, 5C(1), 5C(2), 5D, 5E, 6A, 6B, 6C, 9C, 10C(l), 10C(2), 10D, 10E(l), 10E(2), 11C, HE, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13B, 13D, 13E(1), 14B, 14C(1), 16, and / or 21..
[0324] In some embodiments, the variant recombinase disclosed herein comprising binding regions capable of recognizing one or more of the DNA target site sequence as shown in Figures 5A, 9A, 9B, 10A, 10B(l), 10B(2), HA, 11D, HE, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22.
[0325] In some embodiments, a variant recombinase with a modification in the helix is grafted into another known integrase / recombinase. In some embodiments, a variant recombinase with a modification in the loop is grafted into another known integrase / recombinase. In some embodiments, a variant recombinase with a modification in the hairpin is grafted into another known integrase / recombinase. In some embodiments, a variant recombinase with a modification in the helix, loop, hairpin and / or combinations thereof are grafted into another known integrase / recomb inase .
[0326] In some embodiments, a method of making variant recombinase is disclosed comprising: 1) obtaining a recombinase variant library; 2) generating a plasmid target library; 3) screening selected variants against quarter sites wherein the plasmid library includes all four quarter sites foreach target; 4) screening recombinase variants to target half sites; and 5) combining recombinase variants to recognize endogenous target sites. See Fig. 12A.
[0327] An improved genome-wide method for detecting donor construct integration into a genome
[0328] An improved genome-wide specificity assay is provided herein. This method involves a substantial improvement to AMP-seq (anchored multiplex PCR followed by sequencing) and the improvement involves enzymatically removing unintegrated plasmid DNA from the sample that tends to cause background signal in order to enable the assay to be run on cells that have only been in culture for 7 days vs. the usual requirement for cells to be in culture for 21 to 28 days (the longer culture times are intended to minimize the amount of unintegrated donor remaining in the sample). Enzymatically removing unintegrated plasmid also causes the assay to use fewer DNA sequence reads to obtain the same amount of usable data. There are two aspects to the invention: i) a modified plasmid donor construct with Dpnl sites placed within close proximity on both sides of the recombinase target site (preferably at least one Dpnl site within 100 bp 5’ of the recombinase target site and at least one Dpnl site within 100 bp 3’ of the recombinase target she) and ii) an additional step added to the standard AMP-seq protocol where the purified DNA extracted from the sample is enzymatically digested with the Dpnl enzyme.
[0329] In an embodiment, this method is shown in Fig. 7A. Persons skilled in the art will understand that this method may be modified and used to test the specificity of any other recombinase, nuclease, or other approach involving integration of a construct into a genome, wherein the construct is initially provided as a donor plasmid.
[0330] Disclosed is a modified plasmid donor comprising at least two Dpnl sites, wherein the Dpnl sites are placed within close proximity on both sides of a DNA target site, preferably at least one Dpnl site within 100 bp 5’ of the recombinase target site and at least one Dpnl site within 100 bp 3’ of the recombinase target site. In some embodiments, the DNA target site is for a recombinase or nuclease.
[0331] Disclosed is a method for genome-wide mapping of plasmid integration sites in a genome, the method comprising: 1) generating a modified plasmid donor; 2) combining the modified plasmid donor with the genome in the presence of a reagent that facilitates site-specific integration into the genome to create a mixture; 3) extracting and purifying DNA from the mixture; 4) exposing the purified DNA to a Dpnl enzyme to generate a DpnI-digested sample; 4) sequencing the Dpnl-digested sample; and 5) using the sequencing results from step 4) to identify the plasmid integration sites. In some embodiments, the reagent that facilitates site-specific integration is a recombinase or nuclease.
[0332] When a circular plasmid is integrated, unintegrated plasmid creates a great deal of background signal unless the cells are expanded for at least 3 weeks in order for the cells to replicate many times and thus increase the ratio of integrated donor vs. unintegrated donor that does not replicate when the cells divide. To speed up this assay, the donor construct was modified, and the experimental procedure modified to use restriction enzyme Dpnl to digest any unincorporated plasmid due to its methylation marks caused when the plasmid is replicated in bacterial cells.
[0333] Disclosed is method of applying Dpnl for unintegrated plasmid removal (“Dpnl method”). In some embodiments, applying Dpnl for unintegrated plasmid removal increases integrated plasmid signal in a ligation-mediated PCR-based assay for integration detection. Indeed, the addition of Dpnl sites improves any method that results in insertion of bacterial -based circular DNA constructs into a mammalian genome and attempts to identify these insertion events by amplification of sequence in the construct. In this way, a Dpnl integration assay helps to identify potential integration sites.
[0334] Disclosed is a modified plasmid donor construct comprising at least one Dpnl site. Disclosed is a modified plasmid donor construct comprising at least one Dpnl site and a recombinase. Disclosed is a modified plasmid donor construct comprising at least one Dpnl site and a variant recombinase, the variant recombinase having altered DNA specificity.
[0335] In some embodiments, disclosed is a method for unintegrated plasmid removal, comprising: 1) administering a modified plasmid donor construct comprising at least one Dpnl site to a cell; 2) exposing the cell comprising the modified plasmid donor construct comprising at least one Dpnl site to a Dpnl restriction enzyme thereby digesting any unincorporated plasmid. For clarity, the Dpnl method does not need to be performed with a variant recombinase but can be. Indeed, the Dpnl method could also be useful to identify sites within an organism of interest for a naturally occurring recombinase.
[0336] In some embodiments, the addition of Dpnl is used to screen variant recombinase (including Bxbl, Theia, Veracruz, Kp03, PaOl, Nm60, Si74, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901, and Ssalnt) integration. Disclosed is a method for unintegrated plasmid removal, comprising: 1) administering a modified plasmid donor construct comprising at least one Dpnl siteand a recombinase; 2) allowing the administered modified plasmid donor construct comprising at least one Dpnl site and a recombinase to combine with a target gene in a cell; 3) exposing the cell comprising the recombinase variant to restriction enzyme Dpnl thereby digesting any unincorporated plasmid.
[0337] In some embodiments, disclosed is a method for delivering a recombinase variant, comprising: 1) administering a modified plasmid donor construct comprising at least one Dpnl site and a recombinase variant having altered DNA specificity to a composition comprising at least one cell; 2) integrating the modified plasmid donor construct into at least one cell; 3) exposing the composition to a restriction enzyme thereby digesting any unincorporated plasmid. In some embodiments, the restriction enzyme is Dpnl.
[0338] In some embodiments, the plasmid donor integrated into the cell genome loses methylation pattern with replication. In some embodiments, unintegrated plasmid donor retains methylation. In some embodiments, the method further comprises PCR after Dpnl digestion. In some embodiments, the PCR is ligation-mediated PCR.
[0339] In some embodiments, the at least one Dpnl site is within close proximity on both sides of the recombinase target site. In some embodiments, the at least one Dpnl sites are within 100 bp 5’ of the recombinase target site, are within 100 bp 3’ of the recombinase target site or both.
[0340] The same Dpnl method can be used to enrich a cell population for cells comprising a plasmid, the plasmid comprising at least one Dpnl site by contacting the cell population with a Dpnl digestion enzyme.
[0341] The same Dpnl method can be used to digest any unincorporated plasmid in a cell population by contacting the cell population with a Dpnl digestion enzyme.
[0342] The same Dpnl method can be used to increase the ratio of integrated donor plasmid vs. unintegrated donor plasmid.
[0343] The same Dpnl method can be used to test the specificity of a recombinase’s integration into a host genome. If the integrated plasmid signal in a ligation-mediated PCR-based assay is high, it can be said that the recombinase is specific to the target site.
[0344] In the same way, the specificity assay can be used to identify potential integration sites.If recombination events occur, integrated plasmid signal in a ligation-mediated PCR-based assay increases.
[0345] In some embodiments, integration assays can be performed after about 6 days, 7 days, or 1 week after integration, compared to 3-4 weeks without utilizing this method.
[0346] In some embodiments, this Dpnl method to remove unintegrated plasmids is used to identify the location of integrated donor sequence within a mammalian genome as discussed further below. The Dpnl aspect is an addition to an existing experimental method to determine integration locations. The original method is similar to this publication except integration of circular plasmid DNA is mapped to the genome instead of oligos: www.nature.com / articles / nbt.3117 incorporated by reference in its entirety.
[0347] In some embodiments, identifying potential integration sites comprises removing both attP sequence 5’ and 3’ of the CDN. In some embodiments, identifying potential integration sites comprises removing all sequence up to the start of the CDN (up to and including the 5’ attP sequence). In some embodiments, identifying potential integration sites comprises filtering the CDN in the read to match the CDN used in the donor for each experiment. In some embodiments, identifying potential integration sites comprises aligning sequences to the hg38 genome. In some embodiments, identifying potential integration sites comprises removing alignments with a MAPQ less than 23. In some embodiments, identifying potential integration sites comprises summing all reads per reaction per alignment position in the genome and per alignment orientation (top and bottom strands). In some embodiments, identifying potential integration sites comprises combining positions into a single potential integration location per reaction and alignment orientation if they fell within a 50 bp window of one another, and summing all reads per this grouping. In some embodiments, identifying potential integration sites comprises identifying the coordinate with the most reads. In some embodiments, identifying potential integration sites comprises combining potential integration locations across top and bottom strand alignments and between both “plus” and “minus” reactions per original transfected sample. In some embodiments, identifying potential integration sites comprises merging common potential integration locations into one potential integration location if within 50 bp of one another. In some embodiments, this combining encompasses alignments separated by a dinucleotide as a result of sequencing upstream and downstream of integration in both reactions. In some embodiments, identifying potential integration sites comprises merging common potential integration locations into one potential integrationlocation if there are a minimum of 2 instances across orientations and / or reactions. In some embodiments, identifying potential integration sites comprises inspecting the final list of potential integration loci for expected integration genotypes (2 merged locations, in opposite alignment orientation, separated by a dinucleotide that corresponds to the donor plasmid dinucleotide used in the assay).
[0348] An Antibiotic Resistant Selection System
[0349] Previous recombinase selection systems have used either strategies whereby an antibiotic expression cassette must be inverted to allow transcription of the antibiotic resistance gene, or strategies whereby a Stuffer Sequence must be removed by an active recombinase in order to restore the open reading frame of the antibiotic resistance gene. A successful system was used in the past by Gaj and colleagues (www.pnas.org / doi / full / 10.1073 / pnas.1014214108 incorporated by reference in its entirety) where the beta-lactamase gene, which provides bacteria with resistance against bet- lactam antibiotics such as ampicillin or carbenicillin, was “split” into two halves, with a stuffer sequence encoding a green fluorescent protein gene. In their system, when a recombinase was active against a new target sequence, the open reading frame of the beta lactamase gene was restored, leading to enrichment of active recombinases from a pool of active and inactive recombinases. However, Applicant found that such a system still yielded bacterial cells that survived in the presence of standard concentrations of carbenicillin or ampicillin, even in the absence of active recombinase genes. This may be due to the fact that the position at which the beta lactamase gene was split makes it possible for each half of the split enzyme to re-assemble inside cells and give rise to a resistant phenotype.
[0350] Antibiotic Recombinase Selection System Compositions
[0351] Disclosed is a split antibiotic resistance gene comprising a recombinase target sequence, preferably an attP, and an attB sequence. In some embodiments, the split antibiotic resistance gene encodes an aaCCl protein. In some embodiments, the target sequence of the split antibiotic resistance gene is inserted into a surface-exposed loop of the protein. In some embodiments, the target sequence of the split antibiotic resistance gene is inserted between the following residues: 83- 84.
[0352] Disclosed is an antibiotic resistance gene, split by a stuffer sequence. Disclosed is a plasmid comprising an antibiotic resistance gene, split by a stuffer sequence. Disclosed is a cellcomprising a plasmid, the plasmid comprising an antibiotic resistance gene, split by a stuffer sequence.
[0353] In some embodiments, the antibiotic resistant system comprises recombinase variants disclosed herein. For example, the recombinase variants can have a modified helix, a modified hairpin, a modified loop or combinations thereof. In some embodiments, the recombinase variant recognizes the left halfsite, recognizes the right halfsite or both.
[0354] In some embodiments, the recombination cassette is a spectinomycin resistance cassette. In some embodiments, the recombination cassette comprises a promoter upstream of an open reading frame encoding a gentamicin resistance gene (aacCl) disrupted by a “stuffer sequence”. In some embodiments, the stuffer sequence is flanked with modified attB sequence at its 5’ end, with a modified attP sequence at its 3’ end, or both. In some embodiments, the attB and / or attP sequences are modified such that recombination by an active recombinase would leave behind an attL sequence encoding an in-frame peptide insertion, giving rise to an active aacCl .
[0355] In some embodiments, for a repeat library selection the recombinase variant loop submotif recognizes the designed attB is SEQ ID. NO: 6, and / or recognizes the designed attP SEQ ID NO: 7.
[0356] Disclosed herein is a recombinase selection system where the split antibiotic selection gene is completely inactive in the presence of a stuffer sequence. In some embodiments, the antibiotic selection gene encodes the aaCCl protein. Disclosed is a DNA sequence encoding a recombinase attL sequence capable of being inserted into the gentamicin resistance gene, while maintaining its enzymatic activity. In some embodiments, insertion points are on surface-exposed loops located away from the active site or cofactor binding sites of the enzyme. In some embodiments, the insertion points are attL insertions between the following residues: 34-35; 52-53; 62-63; 74-75; 84-85; 138-139.
[0357] Disclosed is a recombination cassette, containing an attB sequence, an attP sequence and a stuffer sequence between the attB and attP sequence. In some embodiments, the attB and attP target site of interest is placed between the sequence that encodes for residues 83 and 84 of the expressed gentamicin resistance gene.
[0358] Antibiotic Resistant Selection Method
[0359] Disclosed is a method to provide a cell with antibiotic resistance, comprising: 1) generating a recombination cassette comprising a split antibiotic resistance gene, wherein the split antibiotic resistance gene is inactive in the presence of a staffer sequence; 2) transforming the recombination cassette into a bacterial cell comprising an active recombinase where successful recombination creates a functional antibiotic resistance gene; and 3) contacting the transformed bacterial cell with an antibiotic such as gentamicin, ampicillin, carbenicillin, or similar antibiotic. In some embodiments, the antibiotic is gentamicin. In some embodiments, the antibiotic resistance gene is aaCCl. In some embodiments, the antibiotic resistance gene produces the aaCCl protein.
[0360] Disclosed is a method for selecting or identifying a recombinase having activity against a target sequence, comprising: 1) generating a split antibiotic resistance gene described herein; 2) introducing the gene into bacterial cells; 3) introducing into the bacterial cells a recombinase from a pool of active and inactive recombinases, wherein each bacterial cell preferentially expresses a one or more types of recombinase; 4) administering an antibiotic to the bacterial cells; and 5) identifying the recombinase in surviving bacterial cells.
[0361] Method Of Site-Specifically Integrating a Polynucleotide Sequence of Interest in A Genome of a Cell
[0362] Disclosed is a method of site-specifically integrating a polynucleotide sequence of interest into a genome of a cell. The method comprises introducing (i) one or more variant recombinases into the cell capable of interacting with a target recombination site within the genome of the cell (ii) delivering a synthetic polynucleotide sequence of interest. This results in recombination between the synthetic polynucleotide sequence of interest and the target recombination site with the genome of the cell. The result of the recombination is site-specific integration of the polynucleotide sequence of interest in the genome of the cell. In some embodiments, the endogenous sites of interest are not wild-type attachment sites of the recombinase.
[0363] The variant recombinase may be introduced into the target cell, for example, as a polypeptide, or a nucleic acid (such as RNA or DNA) encoding the variant recombinase. The variant recombinase may comprise other useful components, such as a bacterial origin of replication and / or a selectable marker. The synthetic polynucleotide sequence of interest may be delivered to the cell as circular single or double-stranded DNA such as a standard plasmid, mini-circle, or nanoplasmid. The synthetic polynucleotide sequence of interest can also be delivered by a viral vector such as AAV or lentivirus. In some embodiments the cell itself converts the synthetic polynucleotide sequence ofinterest into the required double- stranded circular format. In other embodiments one or more of the variant recombinases are responsible for converting the synthetic polynucleotide sequence of interest into the required format. In yet other embodiments, a combination of a synthetic single-stranded oligonucleotide and one or more variant recombinases converts the synthetic polynucleotide sequence of interest into the required circular double-stranded format. In even further embodiments two pairs of variant recombinases exchange a portion of the synthetic polynucleotide sequence of interest with a portion of the genome of the cell in a process known as recombinase-mediated cassette exchange (RMCE).
[0364] In some embodiments, the donor molecules (plasmids or circular DNA) contemplated herein may contain additional nucleic acid fragments such as control sequences, marker sequences, selection sequences and the like.
[0365] Generally, cells are maintained under conditions that allow recombination to occur between the variant plasmid comprising a coding sequence of interest and target plasmid without the coding sequence of interest. The population of transformed cells can then be screened (or a genetic selection is applied) to identify cells containing the coding sequence of interest that excised the stuffer sequence from the target plasmid, allowing for the expression of a selectable marker as a gene product. Such a product may include, but is not limited to, a product identifiable by screening or selection, such as an RNA product or, a polypeptide product such as and antibiotic resistance gene such as gentamicin.
[0366] Donor Molecules and Methods
[0367] A donor molecule may comprise a nucleic acid coding sequence of interest that may encode a number of different products (e.g., a functional RNA and / or a polypeptide, see below). The product produced from the coding sequence of interest may be used in screening and / or selection studies, or in therapeutically, or agriculturally beneficial methods.
[0368] The nucleic acid construct (e.g., a donor molecule) useful in this embodiment may additionally be comprise one or more nucleic acid fragments of interest. Nucleic acid fragments of interest are therapeutic genes and / or control regions. The choice of nucleic acid sequence will depend on the nature of the disorder to be treated. For example, a nucleic acid construct intended to treat hemophilia B, which is caused by a deficiency of coagulation factor IX, may comprise a nucleic acid fragment encoding functional factor IX. A nucleic acid construct intended to treat obstructive peripheral artery disease may comprise nucleic acid fragments encoding proteins that stimulate thegrowth of new blood vessels, such as, for example, vascular endothelial growth factor, platelet- derived growth factor, and the like. Other applications include integrating CAR constructs into human T cells to enable CAR T cell therapies for cancer, Cystic Fibrosis, Duchenne muscular dystrophy, spinal muscular atrophy, Rhett syndrome, and autosomal dominant kidney disease. Those of skill in the art would readily recognize which nucleic acid fragments of interest would be useful in the treatment of a particular disorder.
[0369] A donor molecule may be introduced at, within, or near a DNA target site via a variant recombinase as described herein. In some embodiments, the DNA target site is specific to a particular allele of a gene. In some embodiments, the DNA target site is in a cell, such as a eukaryotic cell, e.g., a mammalian cell (e.g., a human cell). In some embodiment, the DNA target site is in a plant cell.
[0370] In some embodiments, the DNA target site is in a cell, such as a eukaryotic cell, e.g., a mammalian cell (e.g., a human cell).
[0371] Disclosed are methods for targeted insertion of a polynucleotide (or nucleic acid sequence(s)) of interest into a genome by, for example, (i) providing a variant recombinase, wherein the variant recombinase is capable of facilitating recombination between a first recombination site and a second recombination site, (ii) providing a donor molecule having a first recombination sequence and a polynucleotide of interest, (iii) introducing the variant recombinase and the donor molecule into a cell which contains in its nucleic acid the second recombination site, wherein said introducing is done under conditions that allow the variant recombinase to facilitate a recombination event between the first and second recombination sites.
[0372] In some embodiments, at least one recombination site for a selected variant recombinase is identified in a target cell of interest.
[0373] Introducing Recombinases into Cells
[0374] Also provided herein are nucleic acid molecules encoding the recombinase variants herein and expression vectors comprising, or consisting of, or consisting essentially of, the coding sequences. The vector may be, e.g., a plasmid or a viral vector (e.g., an adeno-associated viral vector, an adenoviral vector, or a lentiviral vector).
[0375] Disclosed are methods of introducing a site-specific, variant recombinase into a cell whose genome is to be modified. Methods of introducing functional proteins into cells are well known in the art. Introduction of purified variant recombinase protein ensures a transient presence of the protein and its function, which is often a preferred embodiment.
[0376] Alternatively, a gene encoding the variant recombinase can be included in an expression vector used to transform the cell. It is generally preferred that the variant recombinase be present for only such time as is necessary for insertion of the nucleic acid fragments into the genome being modified. Thus, the lack of permanence associated with most expression vectors is not expected to be detrimental.
[0377] Disclosed is a system for editing DNA in a cell, comprising, or consisting of, or consisting essentially of, the recombinase variant, the nucleic acid molecule, or a vector herein. In some embodiments, the system further comprises donor DNA, such as a circularized DNA or a linear DNA. The donor DNA may be delivered through, e g., a plasmid or viral vector.
[0378] The editing by the system herein may comprise integration of DNA into the genome of the cell, excision or inversion of DNA in the genome of the cell, or a chromosomal translocation in the genome of the cell.
[0379] In another aspect, the present disclosure provides a method of editing the genome of a cell, the method comprising, or consisting of, or consisting essentially of, providing to the cell the gene editing system herein. In some embodiments, more than one genomic region is edited. The editing method may result in excising DNA from the genome, inverting DNA in the genome, chromosomal translocation, recombinase-mediated cassette exchange (RMCE), and / or integrating donor DNA into the genome.
[0380] Cells
[0381] The variant recombinases disclosed herein can be delivered into a variety of host cells. Host cells are known in the art and include, but are not limited to, eukaryotic cells, plant cells, human embryonic stem cells, human adult tissue stem cells, pluripotent stem cells, induced pluripotent stem cells, reprogrammed stem cells, organoid stem cells, bone marrow stem cells, primary fibroblast, hepatocyte and myoblast cells, mammalian cells including human, monkey, mouse, rat, rabbit, and hamster. The variant recombinases disclosed herein can be delivered into any type of mammaliancell, i.e., fibroblast, hepatocyte, tumor cell, etc. The requirement for the cell used is it is capable of recombination. As used herein, a cell can be a eukaryotic or bacterial cell.
[0382] Provided herein are also cells comprising, or consisting of, or consisting essentially of, the present system, cells edited by the present methods, or descendent cells thereof. The cells may be eukaryotic cells (e.g., mammalian such as human cells).
[0383] Cells suitable for modification employing the methods disclosed herein include both prokaryotic cells and eukaryotic cells, provided that the cell's genome contains a recombination sequence recognizable by a variant recombinase as disclosed herein. Prokaryotic cells are cells that lack a defined nucleus. Examples of suitable prokaryotic cells include bacterial cells, mycoplasmal cells and archaebacterial cells.
[0384] Disclosed are isolated genetically engineered cells. Suitable cells may be prokaryotic or eukaryotic. The genetically engineered cells may be unicellular organisms or may be derived from multicellular organisms. In some embodiments, the isolated cells are outside a living body, whether plant or animal, and in an artificial environment. The use of the term isolated does not imply that the genetically engineered cells are the only cells present.
[0385] In one embodiment, the genetically engineered cells contain a variant recombinase construct disclosed herein. Thus, the genetically engineered cells possess a modified genome.
[0386] The genetically engineered cells disclosed herein are additionally useful as tools to screen for substances capable of modulating the activity of a protein encoded by a nucleic acid fragment of interest. Thus, an embodiment comprises methods of screening comprising, or consisting of, or consisting essentially of, contacting genetically engineered cells with a test substance and monitoring the cells for a change in cell phenotype, cell proliferation, cell differentiation, enzymatic activity of the protein or the interaction between the protein and a natural binding partner of the protein when compared to test cells not contacted with the test substance.
[0387] Cells modified by the methods disclosed herein can be maintained under conditions that, for example, (i) keep them alive but do not promote growth, (ii) promote growth of the cells, and / or (iii) cause the cells to differentiate or dedifferentiate. Cell culture conditions are typically permissive for the action of the recombinase in the cells, although regulation of the activity of the recombinase may also be modulated by culture conditions (e.g., raising or lowering the temperature at which thecells are cultured). For a given cell, cell-type, tissue, or organism, culture conditions are known in the art.
[0388] Delivery of Variant Recombinase Systems
[0389] The recombinase variants disclosed herein may be introduced to target cells as a protein, through a variety of methods (e.g., electroporation, lipid nanoparticles, cationic or anionic liposomes, or a nuclear localization signal (e.g., in combination with liposomes)). In some embodiments, a provided recombinase variant is introduced to target cells through a nucleic acid molecule encoding it, for example, a DNA plasmid or mRNA. The nucleic acid molecule may be in a nucleic acid expression vector, which may include expression control sequences such as promoters, enhancers, transcription signal sequences, and transcription termination sequences that allow expression of the coding sequence for the provided recombinase variants. “Delivery of a system” as described herein may refer to either delivery of a system comprising, or consisting of, or consisting essentially of, a recombinase variant as described herein or delivery of nucleic acid molecules encoding said system of a provided recombinase variant or vectors or expression constructs comprising, or consisting of, or consisting essentially of, the nucleic acid molecules.
[0390] In some embodiments, the promoter on the vector for directing a provided recombinase variant’s expression is a constitutively active promoter or an inducible promoter. Suitable promoters include, without limitation, a Rous sarcoma virus (RSV) long terminal repeat (LTR) promoter (optionally with an RSV enhancer), a cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), a CMV immediate early promoter, a simian virus 40 (SV40) promoter, a dihydrofolate reductase (DHFR) promoter, a P-actin promoter, a phosphoglycerate kinase (PGK) promoter, an EFla promoter, a Moloney murine leukemia virus (MoMLV) LTR, a creatine kinase-based (CK6) promoter, a transthyretin promoter (TTR), a thymidine kinase (TK) promoter, a tetracycline responsive promoter (TRE), a hepatitis B Virus (HBV) promoter, a human al -antitrypsin (hAAT) promoter, chimeric liver-specific promoters (LSPs), an E2 factor (E2F) promoter, the human telomerase reverse transcriptase (hTERT) promoter, a CMV enhancer / chicken -actin / rabbit P- globin promoter (CAG promoter; Niwa et al., Gene (1991) 108(2): 193-9), and an RU-486-responsive promoter. In addition, the promoter may include one or more self-regulating elements whereby a provided recombinase variant can bind to and repress its own expression level to a preset threshold.
[0391] Any method of introducing the nucleotide sequence into a cell may be employed, including but not limited to, electroporation, calcium phosphate precipitation, microinjection,cationic or anionic liposomes, liposomes in combination with a nuclear localization signal, naturally occurring liposomes (e.g., exosomes), or viral transduction. In certain embodiments, the nucleotide sequence is in the form of mRNA and is delivered to a cell via electroporation.
[0392] For in vivo delivery of an expression vector, viral transduction may be used. A variety of viral vectors known in the art may be adapted by one of skill in the art for use in the present disclosure, for example, vaccinia vectors, adenoviral vectors, lentiviral vectors, poxyviral vectors, adeno-associated viral (AAV) vectors, retroviral vectors, and hybrid viral vectors. In some embodiments, the viral vector used herein is a recombinant AAV (rAAV) vector. Any suitable AAV serotype may be used. For example, the AAV may be AAV1, AAV2, AAV3, AAV3b, AAV4, AAV5, AAV6, AAV7, AAV8, AAV8.2, AAV9, AAV.PHP.B, AAV.PHP.eB, or AAVrhlO, or of a novel serotype or a pseudotype such as AAV2 / 8, AAV2 / 5, AAV2 / 6, AAV2 / 9, or AAV2 / 6 / 9. In some embodiments, the expression vector is an AAV viral vector and is introduced to the target human cell by a recombinant AAV virion whose genome comprises the construct, including having the AAV Inverted Terminal Repeat (ITR) sequences on both ends to allow the production of the AAV virion in a production system such as an insect cell / baculovirus production system or a mammalian cell production system. The AAV may be engineered such that its capsid proteins have reduced immunogenicity or enhanced transduction ability in humans. Viral vectors described herein may be produced using methods known in the art. Any suitable permissive or packaging cell type may be employed to produce the viral particles. For example, mammalian (e.g., 293) or insect (e.g., sf9) cells may be used as the packaging cell line. In some embodiments, the target is a self- complementary AAV. In some embodiments, the target is any other linear double-stranded DNA molecule.
[0393] Any type of cell may be targeted for the gene editing methods described herein. For example, the cells may be eukaryotic or prokaryotic. In some embodiments, the cells are mammalian (e.g., human) cells or plant cells. Human cells may include, for example, T cells, Natural Killer (NK) cells, NK T cells, alpha-beta T cells, gamma-delta T-cells, cytotoxic T lymphocytes (CTL), regulatory T cells, B cells, human embryonic stem cells, tumor-infiltrating lymphocytes (TIL) or a pluripotent stem cell from which lymphoid cells may be differentiated (e.g., an induced pluripotent stem cell (iPSC)). In some embodiments, the systems can be used to modify pluripotent stem cells prior to their differentiation into multiple cell types. For example, a lymphoid cell precursor may be modified prior to differentiation into lymphoid cell types such as regulatory T cells, effector T cells, natural killer cells, etc. Systems of the present disclosure comprising, or consisting of, or consistingessentially of, more than one Bxbl recombinase variant or orthologous recombinase variant, in particular, can be used to prepare cells with multiple integrated, excised, or inverted genes at once, including pluripotent cells. In some embodiments, systems containing more than one of the provided recombinase variants may be used to prepare, e.g., allogeneic T cells.
[0394] For agricultural applications, any method for introduction of proteins or nucleic acid molecules to a plant cell is also contemplated, such as Agrobacterium tumefaciens-mediated T-DNA delivery.
[0395] Gene Therapy and Disorders
[0396] In another aspect, the present disclosure provides compositions and methods of treating a disease in a subject in need thereof, comprising, or consisting of, or consisting essentially of, administering the edited cells to the subject. Also included are cells for use in treating a disease in a subject in need thereof, and use of the cells in the manufacture of a medicament for treating a disease in a subject in need thereof. The administered cells are generally eukaryotic cells, such as mammalian cells, and typically are human cells. The administered cells can include additional nucleic acids, e.g., genes for introduction are those to improve the efficacy of therapy, such as by promoting viability and / or function of administered cells; genes to provide a genetic marker for selection and / or evaluation of the administered cells, such as to assess in vivo survival or localization; genes to improve safety, for example, by making the cell susceptible to negative selection in vivo. The administered cell populations can be allogeneic or autologous to the subject receiving the cell populations.
[0397] A further embodiment comprises a method of treating a disorder in a subject in need of such treatment. In one embodiment of the method, at least one cell or cell type (or tissue, etc.) of the subject has a target recombination sequence for a variant recombinase disclosed herein. This cell(s) is transformed with a nucleic acid construct (a “targeting construct”) comprising, or consisting of, or consisting essentially of, a variant recombinase and one or more polynucleotides of interest (typically a therapeutic gene). The variant recombinase specifically recognizes the recombination sequences of the host cell under conditions such that the nucleic acid sequence of interest is inserted into the genome via a recombination event. Subjects treatable using the methods disclosed herein include both humans and non-human animals.
[0398] A variety of disorders may be treated by employing the methods disclosed herein including monogenic disorders, infectious diseases, acquired disorders, cancer, and the like.Exemplary monogenic disorders include ADA deficiency, cystic fibrosis, familial- hypercholesterolemia, hemophilia, chronic ganulomatous disease, Duchenne muscular dystrophy, Fanconi anemia, sickle-cell anemia, Gaucher's disease, Hunter syndrome, X-linked SCID, and the like.
[0399] Infectious diseases treatable by employing the methods disclosed herein include infection with various types of viruses including human T-cell lymphotropic virus, influenza virus, papilloma virus, hepatitis virus, herpes virus, Epstein-Bar virus, immunodeficiency viruses (HIV, and the like), cytomegalovirus, and the like. Also included are infections with other pathogenic organisms such as Mycobacterium Tuberculosis, Mycoplasma pneumoniae, and the like or parasites such as Plasmadium falciparum, and the like.
[0400] The present disclosure provides methods of integrating, excising, and inverting a gene or sequence of DNA in cellular DNA, comprising, or consisting of, or consisting essentially of, delivering a Bxbl recombinase variant or orthologous recombinase variant system described herein to a cell (e.g., from a patient). The cell may be within a patient (in vivo treatment), or a method as described herein may be performed on a cell removed from a patient and then the edited cell delivered to the patient (ex vivo treatment). In some embodiments, the cells are further manipulated ex vivo prior to use as a treatment. The term “treating” encompasses alleviation of symptoms, prevention of onset of symptoms, slowing of disease progression, improvement of quality of life, and increased survival. In some embodiments, a patient treated by the methods described herein is a mammal, e.g., a human.
[0401] In some embodiments, the methods of the present disclosure are used to insert or excise a gene or regulatory sequence associated with a disease towards restoring normal gene expression or activity. In some embodiments, the methods of the present disclosure may recognize a particular allele of a gene, e.g., a wild-type or mutated allele. In certain embodiments, the allele may be associated with cancer.
[0402] In some embodiments, the patient has cancer. In certain embodiments, the cell from the patient is further modified before or after gene editing to provide resistance to a chemotherapeutic agent. The patient may then be treated with the chemotherapeutic agent, which in some embodiments may result in greater survival of edited over unedited cells.
[0403] In some embodiments, the patient has an autoimmune disorder.
[0404] In some embodiments, the patient has an autosomal dominant disease, such as autosomal dominant polycystic kidney disease.
[0405] In some embodiments, the patient has a neuro-developmental disorder.
[0406] In some embodiments, the patient has a mitochondrial disorder.
[0407] In some embodiments, the patient has sickle cell disease, hemophilia (e.g., hemophilia A,B, or C), cystic fibrosis, phenylketonuria, Tay-Sachs, prion disease, color blindness, a lysosomal storage disease (e.g., Fabry disease), Friedreich’s ataxia, prostate cancer, beta thalassemia, Huntington’s disease, renal transplant, inflammatory bowel disease, multiple sclerosis, amyotrophic lateral sclerosis, or frontotemporal dementia.
[0408] The present disclosure further provides a pharmaceutical composition comprising, or consisting of, or consisting essentially of, elements of the gene editing system described herein, such as a Bxbl recombinase variant or orthologous recombinase variant, or nucleotide sequences encoding said elements (e.g., in viral or non-viral vectors as described herein). The pharmaceutical composition may further comprise a pharmaceutically acceptable carrier such as water, saline (e.g., phosphate-buffered saline), dextrose, glycerol, sucrose, lactose, gelatin, dextran, albumin, or pectin. In addition, the composition may contain auxiliary substances, such as, wetting or emulsifying agents, pH-buffering agents, stabilizing agents, or other reagents that enhance the effectiveness of the pharmaceutical composition. The pharmaceutical composition may contain delivery vehicles such as liposomes, nanocapsules, microparticles, microspheres, lipid particles, and vesicles.
[0409] In some embodiments, the provided recombinase variants described herein may be used in a method of treatment described herein, may be for use in a treatment described herein, or may be used in the manufacture of a medicament for a treatment described herein.
[0410] The nucleic acid construct can be administered to the subject being treated using a variety of methods. Administration can take place in vivo or ex vivo. By “in vivo,” it is meant in the living body of an animal. By “ex vivo” it is meant that cells or organs are modified outside of the body, such cells or organs are typically returned to a living body.
[0411] In general, at least 1-10% of the cells targeted for genomic modification should be modified in the treatment of a disorder. Thus, the method and route of administration will optimally be chosen to modify at least 0.1-1% of the target cells per administration. In this way, the number ofadministrations can be held to a minimum in order to increase the efficiency and convenience of the treatment.
[0412] Depending on the specific conditions being treated, such agents may be formulated and administered systemically or locally.
[0413] In some embodiments, the variant recombinase comprises a modular combination of submotifs. In some embodiments, the variant recombinase comprises a first recombinase variant that recognizes both half-sites of a single full genomic site. In some embodiments, variant recombinase comprises a first recombinase variant that recognizes site s5-6L. In some embodiments, the variant recombinase comprises a first recombinase variant that recognizes site 5- 6R. In some embodiments, the variant recombinase comprises a first recombinase variant that recognizes site s5-6L and a second recombinase variant that recognizes site 5-6R.
[0414] The systems, methods and compositions disclosed herein will now be described in greater detail by reference to the following non-limiting Examples.EXAMPLES
[0415] Example 1
[0416] Directed evolution of Bxbl helix to recognize an endogenous sequence in the human genome.
[0417] Fig. 5A compares natural Bxbl attB sequence and site 1-41 pseudo attB sequence in chromosome 3 of the reference human genome (hg38). Bxbl variants that recognize the right halfsite of this sequence are generated and the portions of the right half-site of site 1-41 that can be recognized by the helix, hairpin, and loop regions of Bxbl are boxed.
[0418] Fig. 5B provides sequence information for Bxbl directed evolution helix selections. The selections use a library of plasmids encoding Bxbl variants that randomized the DNA encoding residues 231, 232, 233, 234, 236, and 237 of the helix region. The helix sequence of wild-type Bxbl (residues 231-237), the randomization scheme used in this example, three sequences selected in this example, and the four-residue motif shared by these three selected sequences is shown. “X” represents any of the 20 naturally occurring amino acids (A, C, D, E, F, G, H, J, K, L, M, N, P, Q, R, S, T, V, W, or Y). Note that the L at position 235 was not varied as this residue is expected to face away from the DNA bases when Bxbl is bound to its target sequence. L is not considered to be partof the selected 4 residue motif because it is not varied- only residues that can be enriched by the selection are counted in the 4 residue motifs.
[0419] Fig. 5C1 is a plot of enriched 4 residue sequence motifs from a directed evolution selections using a library of plasmids encoding Bxbl variants described in Figure 5B. that is selected against the wild-type Bxbl attB sequence with GACGAC at positions -11 to -6 and at positions +11 to +6. Many of the enriched motifs such as SXTALXR, SXXAXKR, SXTXXKR, and XXTAXKR resemble the natural sequence at positions 231 through 237 (SATALKR).
[0420] Fig. 5C2 represents a selection using the same Bxbl library with randomized helix region as in Fig. 5C1 except that the target plasmid used in this selection has the GTCTTC sequence at positions +11 to +6 from the right half site of site 1-41 at positions -11 to -6 and at positions +11’ to +6’. Note that none of the enriched motifs in the site 1-41 helix selection resemble the wild-type sequence SATALKR.
[0421] Fig. 5C3 shows a control selection using the same Bxbl randomized helix library, but where no antibiotic was used thus no selective pressure is applied to the cells. There are no enriched motifs in this control selection.
[0422] Fig. 5D shows a plot of integration activity for Bxbl variants when tested in human K562 cells. Integration was detected by PCR amplification using primers that bound at sites Pl and P2 of the endogenous target site shown in Fig 1A followed by DNA sequencing using an Illumina instrument. The selection was performed against the right half site of the site 1-41 target site, so the indicated Bxbl variant is combined with wild-type Bxbl in order to recognize the left half site and the attP sequence in the donor construct. The helix sequence of individual variants that demonstrate at least 10-fold higher integration than wild-type Bxbl alone are shown on the right side of the plot.
[0423] Fig. 5E demonstrates DNA target specificity for wild-type Bxbl and 5 variants selected in example 1 that showed good integration activity at site 1-41 in human K562 cells. Plasmids encoding the indicated Bxbl variant were combined with a mixture of target plasmid bearing all possible DNA sequences at positions -11 to -6. Each target plasmid was generated from discreet oligo sequences such that the same sequence is present at both positions -11 to -6 and at positions +11’ to +6’.Recombined plasmids were selectively amplified by PCR and then sequenced. The proportion of each base at each position of the recombined sequences is plotted. The information content for the distribution of bases at each position informs the height of the stack of letters in this type of plot. The proportion of each base is reflected in the proportion of the total stack represented by each letter. Aposition that contained 100% of a single base would have an information content of 2.0 bits since there are four possible DNA bases in the unselected initial library of target sites. Note the strong shift in preference to T at position +10 for helices with a motif of XXXNXXX (motifs where an asparagine was selected at position 234). Also note the increased information content for bases at position +10 for samples with helices WSSSLKR and WSSNLKR vs. wild-type Bxbl (helix sequence SATALKR).
[0424] Example 2
[0425] Directed evolution of Bxbl hairpin and combining with helix variant to recognize an endogenous sequence in the human genome.
[0426] Figs. 6A-D describe selecting Bxbl hairpins and combining with selected helix.
[0427] Fig. 6A shows directed evolution of the Bxbl hairpin to recognize positions -19 to -12 of the site 1-41 right halfsite Fig. 6A shows residues at positions 314, 316, 318, 321, 323, and 325 were completely randomized in the Bxbl hairpin library. Residue 322 was allowed to be either an A, G, P, or R. The sequence of wild-type Bxbl from residue 313 to residue 326 is shown in the top line. Randomization scheme for libraries is shown on lower lines where X represents any of the 20 naturally occurring amino acids (A, C, D, E, F, G, H, J, K, L, M, N, P, Q, R, S, T, V, W, or Y).
[0428] Fig. 6B is a plot of Bxbl hairpin sequences after one round of selection against a target plasmid bearing the relevant portion of the site 1-41 right half-site sequence (positions -12 to -19) at positions -19 to -12 and at positions +19 to +12 of the target plasmid. Residues at positions that were not randomized (315, 317, 319, 320, 324) are shown in gray. In this case, the maximum information content for an amino acid is 4.32 bits since there are 20 possible naturally occurring amino acids.The Y-axis of the plot shows the bits of information present in the distribution of amino acid residues present at each position in the selected Bxbl variants. An amino acid residue that occurs in all sequences would have approximately 4.3 bits of information while a position with an equal mixture of all 20 amino acid residues will have 0 bits of information. The height of the stack of letters at each position indicates the information content at that position while the relative sizes of the letters indicate the proportion of each amino acid observed at that position.
[0429] Fig. 6C shows the DNA target used in positions -19 to -12 and +19 to +12 of the attB sequence used in the hairpin selection is shown in the top line and the sequence of the most active hairpin sequence selected in this experiment (referred to as ZD32) is shown on the lower line.
[0430] Fig. 6D shows the activity of the selected ZD32 hairpin variant when combined with the selected RD75 helix yields a greater boost in activity vs. wild-type Bxbl than either the ZD32 variant of the RD75 variant alone.
[0431] Example 3
[0432] Genome-wide integration specificity assay shows dramatic shift in specificity for engineered Bxbl variant and a strong preference for the intended site.
[0433] Fig. 7A is a schematic of rapid genome-wide integration characterization. A rapid genome-wide specificity assay can be used to identify the location of integrated donor sequence within a mammalian genome.
[0434] Fig. 7B(1) shows genome-wide integrations identified after one week using the donor and procedure shown in Fig. 7A using a Bxbl construct with wild-type helix and hairpin sequences and further containing the D257K (aspartic acid at position 257 is changed to lysine) mutation that enhances integration activity.
[0435] Fig. 7B(2)shows the genome-wide integration profile of the combination of AGGNLKR (helix RD30) and ZD 32. This result indicates greater genome-wide specificity than observed in Fig. 7B(l).The intended target, site 1-41, is the top integration location for this variant and the genomewide specificity is substantially changed vs. the construct tested in Fig. 7B(1) with the helix and hairpin sequence from wild-type Bxbl. This experiment utilized a mixture of 16 donors bearing all possible CDN sequences. If this experiment had been performed with only a single donor with a CDN that matched the CDN of site 1-41 then there would have been fewer integrations at locations other than site 1-41.
[0436] Fig. 7B(3) shows a comparison of improved genome-wide specificity assay with and without Dpnl digest.
[0437] Transfection
[0438] K562 cells were transfected using conditions known to those of skill in the art. (e.g., electroporated into K562 cells using the SF cell line 96-well Nucleofector kit (Lonza, V4SC-2960) following the manufacturer’s instructions. See www.nature.com / articles / s41587-019-0186-z incorporated by reference in its entirety). Transfections involving integrations with multiple donors with different core dinucleotides were performed separately per donor plasmid and then cells werepooled and expanded for 1 week before being spun down for genomic DNA extraction. Without the Dpnl site addition to donor plasmids, the cells were grown out for 3-4 weeks prior to DNA extraction and the Dpnl digestion step below was not followed.
[0439] Genomic DNA isolation
[0440] Genomic DNA was extracted from K562 cells using Qiagen DNeasy Blood & Tissue kits following the Purification of Total DNA from Animal Blood or Cells Spin-Column Protocol for cultured cells. The optional RNase A incubation was followed for all samples. DNA was eluted in 60 pL Elution Buffer and quantified using the Qubit fluorometer and the Qubit dsDNA BR Assay Kit following the recommended protocol.
[0441] Adapter annealing
[0442] 10 uM adapters were prepared by annealing the MiSeq Common oligo to each GuideSequence _i5 oligo in a 96-well plate format to make a barcoded Y adapter plate. Annealing was performed with IX oligo annealing buffer (10 mM Tris HCL pH 7.5, 50 mM NaCl, and 0.1 mM EDTA) by following the below thermocycling method:
[0443] Chart 1 :95 °C 2 min4 °C holdAdapters were stored at -20 °C, and before use were thawed on ice.
[0444] Shearing
[0445] 400 ng (133,000 haploid human genomes) genomic DNA was brought up to 50-60 uL using IX IDTE pH 7.5 (IDT #11-05-01-05) in each tube of a Covaris 8 microTUBE-130 AFA Fiber H Slit Strip V2 (Covaris #520239). Samples were sonicated on a Covaris ME220 using the settings shown below on a ME220 Rack 8 AFA- TUBE TPX Strip (Covaris PN 500609) using the waveguide (Covaris PN 500526).
[0446] Chart 2:
[0447] Sheared DNA was purified using 1 volume Ampure XP beads (Beckman Coulter #A63880). After beads were added, the solution was mixed and incubated for 5 minutes at room temperature. The mixture was then incubated on a magnet for 5 minutes before the supernatant was removed. 150 uL freshly made 70% ethanol was then used to wash the beads twice, allowing the solution and beads to sit for 30 seconds each time After the second wash, the beads were dried for 6 minutes before adding 15 uL IDTE pH 7.5 and mixing off the magnet. After 2 minutes the mixture was placed on a magnet and incubated for another 2 minutes. 14.5 uL of the supernatant was collected for the next step.
[0448] Dpnl treatment
[0449] The reaction was brought up to 50 uL with the addition of CutSmart (final concentration IX) and 1 uL Dpnl (NEB #R0176S) and incubated for 1 hour at 37 °C. DNA was purified using 0.8x Ampure XP beads using the same bead clean-up protocol as before. The rest of this protocol is similar to the published GUIDE-seq protocol known to those of skill in the art. The Dpnl treatment itself is not unique. However, the application of using Dpnl for unintegrated plasmid removal to increase integrated plasmid signal in a ligation-mediated PCR-based assay for integration detection is unique. The addition of Dpnl sites and the experimental step would not just benefit Bxb 1-based integration screening but any method that results in insertion of bacterial -based circular DNA constructs into a mammalian genome and attempts to identify these insertion events by amplification off sequence in the construct.
[0450] End Repair / A-tailing / ligation
[0451] Next, the End repair and A-tailing mixture below was added to each 14.5 uL DNA mixture, while the reaction was kept on ice.
[0452] Chart 3:Components Volume (ul)DNA from previous step 14.510 mM dNTP mix (Invitrogen, 18427013) 0.510X T4 DNA Ligase Buffer (Enzymatics, B6030) 2.5End-repair mix (Enzymatics, Y9140-LC-L) 210X Platinum Taq Polymerase PCR Rxn Buffer (-Mg2 free) (Invitrogen, 10966034)Taq DNA Polymerase Recombinant (5u / ul) (Invitrogen, 10342020) 0.5 dsH2O 0.5Total vol. 22.5
[0453] The solution was mixed and incubated on a thermocycler with the following program inChart 4:12 °C 15 min37 °C 15 min72 °C 15 min4 °C hold
[0454] Next, 2 uL T4 DNA Ligase (Enzymatics, L6030-LC-L) and 1 uL of one of the 10 uM annealed Y adapters (chosen from a 96-well plate of GS_i5 adapters) was added per end-repaired and A-tailed reaction. Each reaction was mixed and incubated on a thermocycler with the following program in Chart 5:16 °C 30 min22 °C 30 min4 °C hold
[0455] DNA was purified using 0.9 volumes Ampure XP beads using the same bead clean-up protocol but using 23 uL IDTE pH 7.5 to resuspend the DNA-bead mixture for elution and collecting 22 uL supernatant after incubation on the magnet.
[0456] PCR 1
[0457] Next, ligated and sheared DNA fragments were amplified using a primer specific to all adapters (P5_l) and a primer specific to the sequence of interest (Guide Sequence Pl + / -):
[0458] To each tube on ice, the following reagents were added: Chart 6:Reagent Vol. (ul)DNA from previous step 77 0 10X Platinum Taq Polymerase PCR Rxn Buffer (-Mg2 free) (Invitrogen,10966034) 3.050 mM MgC12 (Invitrogen, 10966034) 1.210 mM dNTP mix (Invitrogen, 18427013) 0.610 pM P5_l primer 0.510 pM Guide Sequence P1+ / - 1.00.5 M TMAC (Sigma Aldrich T3411) 1.5Platinum Taq DNA polymerase (5 U / pL) (Invitrogen, 10966034) 0.3Total 30.1
[0459] Each reaction was mixed and incubated with the thermocycler conditions in Chart 7:95 °C 5 min14 repeats after first (Total 15 cycles)9 repeats after first (Total 10 cycles)72 °C 5 min4 °C hold
[0460] Amplified DNA was purified using the bead clean-up protocol but with 1.2 volumes of Ampure XP beads and using 21 uL IDTE pH 7.5 to resuspend the DNA-bead mixture for elution and collecting 20.4 uL supernatant after incubation on the magnet.
[0461] PCR 2
[0462] The second round of PCR amplification was performed to add plate barcodes (Guide Sequence _i7 sequences). To each tube on ice, the following reagents were added as shown in Chart 8:Reagent Vol. (ul)10 pM Plate adapter (GS_i7) 1.5DNA from previous step 20.410X Platinum Taq Polymerase PCR Rxn Buffer (-Mg2 free) (Invitrogen,10966034) 3.050 mM MgC12 (Invitrogen, 10966034) 1.210 mM dNTP mix (Invitrogen, 18427013 0.610 pM P5_2 primer 0.510 pM GUIDE SEQUENCE P2+ / - 1.00.5 M TMAC (Sigma Aldrich T3411) 1.5Platinum Taq DNA polymerase (5 U / pL) (Invitrogen, 10966034) 0.3Total 30.0
[0463] Each reaction was mixed and incubated with the thermocycler conditions in Chart 9:95 °C 5 min14 repeats after first (Total 15 cycles)72 °C 5 min4 °C hold
[0464] Afterwards, 25 uL of each barcoded reaction was combined into one pool and 0.7 volumes of Ampure XP beads were added. Samples were mixed 10 times and incubated for 5 minutes at room temperature. They were then added to the magnet for 5 minutes then the supernatant was discarded. 2 lx volumes of freshly made 70% ethanol were added, incubating for 30 seconds each time before discarding the supernatant. After the last wash, the beads were air-dried for 6-8 minutes. 75 uL IDTE pH 7.5 was added, and the tubes were removed from the magnet and mixed, then incubated for 2 minutes. The reaction was separated on the magnet for 2-4 minutes and then the supernatant was collected into a new tube for NGS library sample submission. Pooled eluate was quantified using the Qubit and the Qubit dsDNA HS kit following the recommended protocol.
[0465] NGS
[0466] Final products were sequenced on a MiSeq or NextSeq 2000 with paired-end 150 bp reads with the cycle settings: 148-10-22-148 for MiSeq or 151-10-22-151 for NextSeq 2000. Samples were sequenced to obtain 3,000,000-fold coverage per sample. For MiSeq reactions, 3 uL of 100 uM custom Indexl primer was added to MiSeq Reagent cartridge position 13 and 3 uL of 100 uM custom Read2 primer was added to MiSeq Reagent cartridge position 14 in Chart 10:Primer name Sequence (5' —> 3')Custom Indexl ATCACCGACTGCCCATAGAGAGGACTCCAGT primer CACCustom Read2 GTGACTGGAGTCCTCTCTATGGGCAGTCGGTG primer AT
[0467] For NextSeq 2000 reactions, 1.98 ul of 100 uM custom Read2 primer was added to 600 ulIllumina HP21 primer mix for a 0.3 uM final concentration and 3.98 ul of 100 uM custom Indexlprimer was added to 600 ul BPM primer mix for a 0.6 uM final concentration. Then 550 pl of each custom primer mix was added to custom 1 well or custom 2 well on the reagent cartridge separately.
[0468] Analysis
[0469] NGS reads were demultiplexed, adapter trimmed, and filtered for a minimum quality threshold of 14 over all bases. Samples then underwent separate analyses for plasmid integration site detection.
[0470] NGS samples were processed to remove remaining contaminant unintegrated plasmid reads due to incomplete Dpnl digestion or fragment removal, aligned to the hg38 genome, and potential integration sites were summarized. First, reads that contained both AttP sequence 5’ and 3’ of the dinucleotide were removed from analysis, corresponding to unintegrated donor plasmid reads. Then all sequence up to the start of the dinucleotide (up to and including the 5’ AttP sequence) was removed, leaving the remaining sequence to align to the hg38 genome using Bowtie2. Alignments with a MAPQ less than 20 were removed from the analysis. Next, all reads per reaction were summed per alignment position in the genome and per alignment orientation (top and bottom strands). Then, positions were combined into a single potential integration location per reaction and alignment orientation if they fell within a 50 bp window of one another, all reads per this grouping were summed and the coordinate with the most reads was kept per group. Lastly, potential integration locations across top and bottom strand alignments and between both “plus” and “minus” reactions per original transfected sample were combined. Common potential integration locations were merged into one potential integration location if within 50 bp of one another (this would encompass alignments separated by a dinucleotide as a result of sequencing upstream and downstream of integration in both reactions) and a minimum of 2 instances across orientations and reactions were required. The final list of potential integration loci was inspected for expected integration genotypes (2 merged locations, in opposite alignment orientation, separated by a dinucleotide that corresponds to the donor plasmid dinucleotide used in the assay).
[0471] Example 4
[0472] Gentamicin version of directed evolution system.
[0473] When generating an antibiotic resistant plasmid aaCCl protein was chosen, which acts on gentamicin and deactivates its cell-killing ability, since it is relatively small, and does not act on other antibiotics that are commonly used to work with bacteria. Since no reports of a split gentamicinresistance gene exist in the literature, Applicant decided to test empirically where the products of a recombination reaction i.e., a DNA sequence encoding a Bxbl attL sequence could be inserted into the gentamicin resistance gene, while maintaining its enzymatic activity. To this end Applicant employed the crystal structure of the product of the aaCCl gene from Serratia marcescens, PDB ID 1BO4, and Applicant chose, as insertion points, surface-exposed loops which were located away from the active site or cofactor binding sites of the enzyme.
[0474] Applicant tested attL insertions between the following residues: 34-35; 52-53; 62-63; 74.75; 84-85; 138-139.
[0475] When tested for their ability to generate colonies in an agar plate containing gentamicin, only the aaCCl variant with an insertion between residues 84-85 gave colonies. Furthermore, when Applicant inserted a recombination cassette, containing an attB sequence, a stuffer sequence and an attP sequence at this position of the gene, cells containing active Bxbl were able to survive in the presence of gentamicin, whereas cells which contained an inactive Bxbl variant (missense mutation), did not show any growth in the presence of gentamicin. This established the split antibiotic resistant system as a suitable reporter system for the selection of active serine integrases.
[0476] Fig. 8 shows the result of successful recombination of the gentamicin directed evolution system. The directed evolution system used in example 1 was based on a published system (www.pnas.org / doi / full / 10.1073 / pnas.1014214108 incorporated by reference in its entirety) that creates a functional beta-lactamase gene that allows cells to survive in ampicillin, carbenicillin, or similar antibiotic(doi.org / 10.1093 / nar / gkql25 incorporated by reference in its entirety). This yielded substantial amounts of false positive sequences that degraded the signal so a second version was created where successful recombination creates a functional gentamicin resistance gene, and the cell can then survive in the presence of gentamicin. The attB and attP target site of interest is placed between the sequence that encodes for residues 83 and 84 of the expressed gentamicin resistance gene. This Fig. shows the result of successful recombination. The sequence of the unrecombined initial sequence is this sequence with the “attL” sequence replaced with the attB, stuffer sequence, and attP sequence diagramed in Fig. 4. The attB and attP sequence are varied to match the relevant portions of the target DNA sequence the selection is being performed to recognize.
[0477] Fig. 9A shows four additional sites in the human genome where wild-type Bxbl has detectable integration of a donor sequence where the CDN of the donor match the CDN of the target (underlined). The DNA sequence and genomic coordinates of these sites are shown.
[0478] Fig. 9B shows the target sites used in the modified directed evolution system utilizing the antibiotic gentamicin used to select hairpins for the relevant portions of these target sites.
[0479] Fig. 9C shows the integration activity in human K562 cells using the indicated Bxbl variants.
[0480] Fig, 9D shows the integration activity in human K562 cells when cells are exposed to the appropriate DNA donor construct and the indicate mixture of Bxbl variants (along with wild-type Bxbl to interact with the donor construct). Mixtures of a Bxbl variant intended to interact with the left halfsite of a genomic target and a Bxbl variant intended to interact with a right halfsite of the same genomic target show substantial improvements in integration activity vs. only using a Bxbl variant intended to interact with a single halfsite as shown in figure 9C.
[0481] As will be understood by a skilled artisan, an integrase can be selected which will excise a gene or portion of a gene such that in the target cell / plasmid the gene of interest is reassembled. For example, it is possible to reassemble a resistance gene such as gentamicin in a cell / plasmid.
[0482] Preparation of selection cassette plasmids
[0483] Selection plasmid pRep2 was based on a pl 5 A origin of replication and contained a Spectinomycin resistance cassette. The selection cassette contained a Tet promoter upstream of an open reading frame encoding a gentamicin resistance gene (aacCl) disrupted by a “stuffer sequence”. The stuffer sequence was flanked with modified attB sequence at its 5’ end, and with a modified attP sequence at its 3’ end. The attB and attP sequences were modified such that recombination by an active integrase would leave behind an attL sequence encoding an in-frame peptide insertion, giving rise to an active aacCl . Each selection cassette was generated by the cloning of two ~1050bp DNA fragments into a linearized pRep2 backbone plasmid using the NEB HiFi assembly master mix, according to the manufacturer’s instructions. Sequence-confirmed plasmids were either prepared by Mini- or Midi- Kits (Qiagen) according to the manufacturer’s instructions.
[0484] Selection of loop submotifs to recognize the designed attB [GGCTTGTCGACGTTGGCGGTCTCCAACGTCAGGATCAT] (SEQ ID NO: 6) and attP [GGTTTGTCTGGTCAACCTTCGCGGTCTCAAAGGTGTACGGTACAAACC] (SEQ ID. NO: 7) sequences, where the underlined residues are all sixteen different dinucleotide combinations, and in each attB or attP site, the right dinucleotide is the reverse complement of the left dinucleotide.
[0485] Preparation of Integrase mutant libraries
[0486] As a starting point for engineering of Bxbl variants in E. coli, a codon optimized ORF was cloned into plasmid pRex2, which contains a pBR322 origin of replication as well as the L- rhamnose inducible pRhaBAD promoter. Libraries were prepared using inverse PCR and primers encoding degenerate bases using an NNK degeneracy scheme.
[0487] Sequence of pRex2 plasmid with wild type Bxbl
[0488] GGGAGACGACAACGGTTTCCCTCTAGAAATAATTTTGTTTAACTATAAGAAG GAGATATACATATGCGTGCGCTTGTAGTGATCCGCCTGTCACGCGTTACCGATGCTACT ACATCTCCGGAACGTCAACTTGAGAGTTGCCAACAGTTATGTGCCCAGCGTGGGTGGGA TGTGGTTGGAGTCGCCGAAGATTTGGATGTAAGCGGTGCGGTCGACCCGTTCGATCGCAAGCGTCGCCCTAATCTGGCACGTTGGCTGGCTTTCGAAGAGCAACCTTTTGATGTTATCG TCGCTTATCGTGTCGATCGTCTTACACGTAGCATTCGCCATCTTCAACAACTGGTACATT GGGCGGAGGATCATAAAAAGTTGGTCGTGTCGGCGACGGAGGCTCACTTTGACACAAC GACGCCATTTGCTGCGGTTGTTATCGCTCTGATGGGTACCGTAGCTCAGATGGAGCTGG AAGCCATTAAAGAGCGCAATCGTAGTGCAGCTCATTTTAACATCCGTGCGGGCAAGTAT CGCGGGAGCCTTCCACCATGGGGGTACCTTCCTACTCGTGTTGACGGGGAATGGCGCTT AGTTCCTGACCCGGTACAACGCGAACGTATCCTTGAAGTCTACCATCGCGTTGTGGACA ACCACGAACCTTTGCATTTAGTTGCGCACGACTTGAATCGTCGCGGGGTCCTTTCCCCCA AGGATTACTTTGCACAGCTTCAAGGCCGCGAACCTCAGGGGCGCGAATGGTCCGCCACC GCCTTAAAGCGTAGTATGATTTCCGAAGCAATGTTGGGATATGCGACACTTAATGGGAA GACGGTCCGCGACGACGATGGCGCTCCATTAGTCCGTGCCGAGCCAATTTTGACCCGTG AACAACTTGAAGCATTGCGCGCTGAATTGGTTAAGACTAGCCGTGCAAAGCCGGCGGTG TCGACCCCTTCACTGCTTCTTCGCGTCTTGTTCTGCGCTGTGTGTGGGGAACCGGCTTAT AAATTTGCGGGCGGCGGACGCAAGCACCCGCGTTATCGCTGCCGCTCCATGGGATTTCC CAAGCATTGCGGTAATGGCACAGTTGCCATGGCCGAGTGGGACGCTTTCTGCGAAGAAC AAGTCCTGGATTTGCTTGGGGACGCTGAGCGTTTAGAGAAGGTGTGGGTAGCGGGGTCG GATAGTGCTGTTGAGTTGGCTGAGGTCAACGCCGAGCTGGTCGACCTGACATCTTTGAT TGGCTCGCCCGCTTATCGTGCAGGGTCACCACAACGTGAAGCTCTTGACGCTCGTATCG CTGCTCTGGCGGCTCGCCAGGAAGAACTTGAGGGTCTGGAAGCCCGCCCCAGTGGCTGG GAATGGCGCGAAACAGGTCAACGTTTTGGAGACTGGTGGCGCGAGCAGGACACCGCGG CAAAAAATACTTGGTTGCGCTCTATGAATGTCCGCCTGACATTCGACGTACGTGGAGGT TTAACGCGCACAATCGACTTTGGCGATTTGCAAGAATACGAACAGCACTTGCGTCTGGGGTCAGTCGTTGAACGCTTGCACACTGGTATGAGCGGATCGGGGTCTGGCAGCCACCACCATCATCATCACTAATGAGCGGTCTTCAATAAGGATCCGGGCCTGTAACAGAGCATTAGCGCAAGGTGATTTTTGTCTTCTTGCGCTAATTTTTTGCGTCAGACCCCGTAGAAAAGATCAAAGGATCTTCTTGAGATCCTTTTTTTCTGCGCGTAATCTGCTGCTTGCAAACAAAAAAACCACCGCTACCAGCGGTGGTTTGTTTGCCGGATCAAGAGCTACCAACTCTTTTTCCGAAGGTAACTGGCTTCAGCAGAGCGCAGATACCAAATACTGTCCTTCTAGTGTAGCCGTAGTTAGGCCACCACTTCAAGAACTCTGTAGCACCGCCTACATACCTCGCTCTGCTAATCCTGTTACCAGTGGCTGCTGCCAGTGGCGATAAGTCGTGTCTTACCGGGTTGGACTCAAGACGATAGTTACCGGATAAGGCGCAGCGGTCGGGCTGAACGGGGGGTTCGTGCACACAGCCCAGCTTGGAGCGAACGACCTACACCGAACTGAGATACCTACAGCGTGAGCTATGAGAAAGCGCCACGCTTCCCGAAGGGAGAAAGGCGGACAGGTATCCGGTAAGCGGCAGGGTCGGAACAGGAGAGCGCACGAGGGAGCTTCCAGGGGGAAACGCCTGGTATCTTTATAGTCCTGTCGGGTTTCGCCACCTCTGACTTGAGCGTCGATTTTTGTGATGCTCGTCAGGGGGGCGGAGCCTATGGAAAAACGCCAGCAACGCGGCCTTTTTACGGTTCCTGGCCTTTTGCTGGCCTTTTGCTCACATGTTCTTTCCTGCGTTATCCCCTGATTCTGTGGATAACCGTATTACCGCCTTTGAGTGAGCTGATACCGCTCGCCGCAGCCGAACGACCGAGCGCAGCGAGTCAGTGAGCGAGGAAGCGGAAGAGCGCCTGATGCGGTATTTTCTCCTTACGCATCTGTGCGGTATTTCACACCGCATATGGTGCACTCTCAGTACAATCTGCTCTGATGCCGCATAGTTAAGCCAGTATACACTCCGCTATCGCTACGTGACTGGGTCATGGCTGCGCCCCGACACCCGCCAACACCCGCTGACGCGCCCTGACGGGCTTGTCTGCTCCCGGCATCCGCTTACAGACAAGCTGTGACCGTCTCCGGGAGCTGCATGTGTCAGAGGTTTTCACCGTCATCACCGAAACGCGCGAGGCAGCTGCGGTAAAGCTCATCAGCGTGGTCGTGAAGCGATTCACAGATGTCTGCCTGTTCATCCGCGTCCAGCTCGTTGAGTTTCTCCAGAAGCGTTAATGTCTGGCTTCTGATAAAGCGGGCCATGTTAAGGGCGGTTTTTTCCTGTTTGGTCACTGATGCCTCCGTGTAAGGGGGATTTCTGTTCATGGGGGTAATGATACCGATGAAACGAGAGAGGATGCTCACGATACGGGTTACTGATGATGAACATGCCCGGTTGAAATTCTGCCTCGTGATACGCCTATTTTTATAGGTTAATGTCATGATAATAATGGTTTCTTAGACGTCAGGTGGCACTCGAGTTGATCGGGCACGTAAGAGGTTCCAACTTTCACCATAATGAAATAAGATCACTACCGGGCGTATTTTTTGAGTTATCGAGATTTTCAGGAGCTAAGGAAGCTAAAATGGAGAAAAAAATCACTGGATATACCACCGTTGATATATCCCAATGGCATCGTAAAGAACATTTTGAGGCATTTCAGTCAGTTGCTCAATGTACCTATAACCAGACCGTTCAGCTGGATATTACGGCCTTTTTAAAGACCGTAAAGAAAAATAAGCACAAGTTTTATCCGGCCTTTATTCACATTCTTGCCCGCCTGATGAATGCTCATCCGGAATTTCGTATGGCAATGAAAGACGGTGAGCTGGTGATATGGGATAGTGTTCACCCTTGTTACACCGTTTTCCATGAGCAAACTGAAACGTTTTCATCGCTCTGGAGTGAATACCACGACGATTTCCGGCAGTTTCTACACATATATTCGCAAGATGTGGCGTGTTACGGT GAAAACCTGGCCTATTTCCCTAAAGGGTTTATTGAGAATATGTTTTTCGTCTCAGCCAAT CCCTGGGTGAGTTTCACCAGTTTTGATTTAAACGTGGCCAATATGGACAACTTCTTCGCC CCCGTTTTCACCATGGGCAAATATTATACGCAAGGCGACAAGGTGCTGATGCCGCTGGC GATTCAGGTTCATCATGCCGTTTGTGATGGCTTCCATGTCGGCAGAATGCTTAATGAATT ACAACAGTACTGCGATGAGTGGCAGGGCGGGGCGTAATTTTTTTAAGGCAGTTATTGGT GCCCTTAAACGCCTGGTGCTACGCCTGAATAAGTGATAATAAGCGGATGAATGGCAGA AATTCGAAAGCAAATTCGACCCGGTCGTCGGTTCAGGGCAGGGTCGTTAAATAGCCGCT TATGTCTATTGCTGGTTTACCGGTTTCACCACAATTCAGCAAATTGTGAACATCATCACG TTCATCTTTCCCTGGTTGCCAATGGCCCATTTTCCTGTCAGTAACGAGAAGGTCGCGAAT TCAGGCGCTTTTTAGACTGGTCGTA
[0489] In order to diversify the “loop” submotif, primers were designed to diversify residues 154, 155, 156, 157, 158 and 159.
[0490] Forward primer (SEQ ID No. 4)
[0491] GAGTAGGGTCTCGGCAAGNNKNNKNNKNNKNNKNNKCCATGGGGGTACCTT CCT
[0492] Reverse primer (SEQ ID No. 5)
[0493] GAGTAGGGTCTCCTTGCCCGCACGGATGTT
[0494] In order to diversify the “loop” submotif, primers were designed to target residues 154- 159. In order to diversify the “alpha-helix” submotif, primers were designed to target residues 231- 234, and residues 236-237. In order to diversify the “Beta-hairpin” submotif, primers were designed to target residues 314, 316, 318, 321, 323, and 325 using an NNK randomization scheme, while residue 322 was randomized using an SSK randomization scheme. Briefly iPCR reactions were set up in 50uL volumes, using lx KOD ONE Master mix, 0.5uM forward primer, 0.5uM reverse primer, and lOng of template plasmid DNA. Following purification of the PCR products via silica column (Qiagen PCR cleanup kit), and elution in 50uL of buffer EB, compatible overhangs for ligation were generated by digestion of PCR products using 5 units each of B sal (NEB) and Dpnl (NEB) in 60uL of lx Cutsmart buffer. The resulting digest products were then ligated at lOng / uL in lx T4 Ligase buffer, at 8U / uL T4 Ligase, for 1 hour at room temperature. Following ligation, DNA products were purified using the Qiagen PCR cleanup kit and eluted in 50uL of buffer EB. The resulting ligatedproducts were used to transform 50uL of electrocompetent cells (NEB lOBeta or ThermoFisher OneShot Top 10) using a BMX electroporator and a 96-well cuvette (BTX catalog number 45-0450- M) using the manufacturer’s protocol and allowed to recover for 1 hour at 37C in 1ml total volume of SOC media. The resulting transformations yielded libraries in the order of 4e8 to 5e9 CFUs. After recovery, cells were transferred to IL of LB media containing 34ug / mL chloramphenicol and grown overnight with shaking at 37C. The resulting culture was harvested, and plasmids were purified using a Qiagen Plasmid Plus Giga Kit.
[0495] Selection of Helix submotifs to recognize the designed attB[GGCTTGTCNNNGACGGCGGTCTCCGTCGACAGGATCAT] (SEQ ID No. 8) and attP sequences [GGTTTGTCTGGTCNNNCACCGCGGTCTCAGTGNNNTACGGTACAAACC] (SEQ ID No. 9), where the underlined residues are all 64 different trinucleotide combinations, and in each attB or attP site, the right trinucleotide is the reverse complement of the left trinucleotide.
[0496] In order to diversify the “alpha-helix” submotif, primers were designed to randomize residues 231, 232, 233, 234, 236 and 237, using an NNK scheme. The libraries were generated as in the loop selection example.
[0497] Forward primer gagtagGGTCTCAATGGNNKNNKNNKNNKTTANNKNNKAGTATGATTTCCGAAGCAATGT TGGGA (SEQ ID No. 10)
[0498] Reverse primer GAGTAGGGTCTCCCATTCGCGCCCCTGAGGTTCGCGGCCTTGAAGCTGTGCAAA (SEQ ID No. 11)
[0499] The selections were carried out as detailed in the loop selection example.
[0500] Example: Selection of Hairpin submotifs to recognize the designed attB [AGGATCGGGGGTTTGAGACGACCGCGGACTCAGTGGTCTCAAACCC] (SEQ ID No. 12) and attP[GGGTTTGATGGTCGACGACCGCGGACTCAGTGGTCTACGGTCAAACCCAGGCAG]sequen ces (SEQ ID No. 13).
[0501] In order to diversify the “Hairpin” submotif, primers were designed to randomize residues 314, 316, 318, 321, 323 and 325 using an NNK scheme, while residue 322 was randomized using an SSK scheme The libraries were generated as in the loop selection example.
[0502] Forward primer gagtagggtctcCGCAAGNNKSSKNNKTATNNKTGCCGCTCCATGGGATTT (SEQ ID No. 14)
[0503] Reverse primer gagtagggtctcCTTGCGMNNGCCMNNCGCMNNTTTATAAGCCGGTTCCCCACACA (SEQ ID No. 15)
[0504] The selections were carried out as detailed in the loop selection example.
[0505] Example 5
[0506] Systematic selections of Bxbl helix against 64 different 3 bp DNA target sites.
[0507] Fig. 10A shows systematic selection of the Bxbl randomized helix library against all possible target sequences at positions -11, -10, and -9. These target sites were generated from synthetic oligos such that the same 3 bp sequence is at positions -11, -10, and -9 and is at positions +11’, +10’, and +9’. The randomized helix library described in example 1 was selected against all possible 64 sequences at these three positions.
[0508] Fig. 10B(l) and Fig. 10B(2) shows attB and attP sequences used for each of the 64 selections described in Fig. 10A.
[0509] Fig. 10C(l) and 10C(2) shows the top 10 enriched 4-residue motifs selected with each target sequence DNA triplet (-11 to -9). The DNA sequence used in the selection is shown in the column labeled “triplet”. Only patterns whose enrichment passed a statistical significance test are shown. Some triplet sequences such as TTT did not yield statistically enriched motifs in this example. In some embodiments, the Bxbl variant recombinase comprises a sequence selected from the sequences found in Fig. 10C(l) and 10C(2).
[0510] Fig. 10D: DNA target site specificity plots for the indicated helix variant. The experiment is similar to the experiment shown in Fig. 5E, except that only the DNA bases at positions -11, -10, and -9 and +11’, +10‘, and +9’ at the target site were varied. This yields a profile containing lessexperimental noise than when more positions of the DNA target site are randomized. Data for helices that yield good DNA target site specificity for a variety of different triplet sequences are shown.
[0511] Fig. 10E(l) and 10E(2) shows helices that can recognize the indicated triplet based on the DNA target site specificity experiment. Up to 10 helix sequences are shown for each triplet. Some helix sequences are shown corresponding to multiple triplets if the molecular specificity indicates they can interact with more than one triplet.
[0512] Example 6
[0513] Systematic selections of Bxbl loops against 16 different 2 bp targets.
[0514] Fig. 11 A shows an overview of target sites used for systematic directed evolution selections of the Bxbl loop against 16 different target sites that vary at positions -7 to -6 and +7’ to +6’.
[0515] Fig. 1 IB shows the sequence of the loop region of wild-type Bxbl and the randomization scheme used in the directed evolution selections. For loop selections, residues 154-159 of Bxbl are completely randomized.
[0516] Fig. 11C shows example loop variants selected against loop targets.
[0517] Fig. 1 ID shows attB and attP target sites used in the systematic Bxbl loop selections inExample 6.
[0518] Fig. 1 IE shows loop variants that were active in a DNA target site specificity assay using a mixture of the target sites shown in Fig. 1 ID. The loop variant is indicated in the left column of the table. Data for the wild-type Bxbl loop sequence YRGSLP is in the top row. The next 5 rows are loop sequences that are tolerant of different sequences at positions -7 and -6 of the target site. The loop sequences on the sixth row and below show some specificity for a variety of different 2 bp sequences at positions -7 and -6. The values in the table are the fraction of recombined reads corresponding to the 2bp sequence at -7 and -6 indicated above the top row of data.
[0519] Example 7
[0520] Measuring recombination of Bxbl variants with a mammalian plasmid assay.
[0521] Fig. 12A shows an overview of the plasmid-based screening system used for screening experiments against synthetic target sites prior to the final test against the endogenous target site in human cells.
[0522] Fig. 12B is a general description of the target sites and priming sequence used for the plasmid-based screening assay shown in Fig. 12A. Plasmids bearing many such target sites are generated from arrays of synthetic oligos in a pool. This pool of plasmid sequences is co-delivered with either mRNA or plasmid expressing the Bxbl variant being tested to human cells. Successfully recombined sequences can then be amplified and sequenced to determine the relative preference of that Bxbl variant for all target sequences in the pool of plasmid targets.
[0523] Fig. 12C is an example of sequences in this pool that correspond to portions of a single full endogenous recombination target site within the human AAVS1 locus. Six different target sites are derived from one endogenous target site- either an inverted repeat of the full left halfsite, an inverted repeat of the full right halfsite, or an inverted repeat of one “quarter site” of the endogenous target combined with a portion of the attB target site for wild-type Bxbl. WT_ZD indicates the site has a portion of the left halfsite of the natural Bxbl attB sequence that contains the portion that is most important for recognition of the wild-type hairpin sequence as well as some additional sequence context (-24 to -13). WT_RD indicates the sequence contains the portion of the left halfsite of the wild-type attB sequence that is critical for recognition by the wild-type Bxbl loop and helix (-10 to - 2).
[0524] Figs. 12D(1) and 12D(2) shows Bxbl helix variants that show improved activity vs. wildtype Bxbl in human cells. Bxbl variants were tested using the plasmid-based assay described in Figs. 12A, 12B, and 12C in human K562 cells. Figs. 12D(1) and 12D(2) show examples of Bxbl variant helices that show improved activity at the indicated target site vs. the wild-type Bxbl construct with a SATALKR helix sequence. For each panel, the first column shows the hg38 coordinates for the full endogenous site the indicated test sequence was derived from, the second column shows where the test sequence was derived from the left (L) or right (R) halfsite of the full genome target site, the third column shows the test sequence used to generate the attB target used in the assay, the fourth column show the sequence of the bases at -11 to -9 of the test sequence that can be specified by a Bxbl helix, the fifth column shows the amino acid sequence of the helix that was tested and the last column shows the activity when divided by the activity for the wild-type Bxbl construct when tested at this same test sequence. A “pseudo count” was added to the data for both the variant and wild-type Bxbl to avoid dividing by zero. Note that while the third column shows theentire halfsite, only the quartersite corresponding to portion recognized by the RD domain was used to obtain the measurements in the right-most column. In other words, the actual sites used to obtain the data resemble the two sequences in Fig. 12C with names ending “WT_ZD” and contained the DNA sequence -24 to -13 from the natural attB target site for Bxbl and contain the sequence from - 12 to -2 of the indicated half-site.
[0525] Figs. 12E(1) and 12E(2) show Bxbl hairpin variants that show improved activity vs. wild-type Bxbl in human cells. Data is presented in the same manner as in Fig. 12D (1-2) except that the target DNA sequence in the fourth column is positions -19 to -12 which corresponds to the region that can be recognized by Bxbl hairpin variants, and the fifth column shows the amino acid sequence of the hairpin region of the Bxbl variant. In other words, the sequences used to obtain the measurements in the right column resemble the two sequences in Fig. 12C with names ending “WT RD” and contain the DNA sequence -11 to -2 from the natural attB target site for Bxbl and contain the sequence from -19 to -12 of the indicated half-site.
[0526] Fig. 12F(1) and 12F(2) shows Bxbl loop variants that show improved activity vs. wildtype Bxbl in human cells. Data is presented in the same manner as in Fig. 12D (1-2) except the target DNA sequence shown in the fourth column is positions -7 and -6 that correspond to the portion of the halfsite that can be recognized by Bxbl loop variants, and the amino acid sequence shown in the fifth column is the sequence of the loop region of the Bxbl variant being tested. The target sites are similar to the “WT_ZD” target sites described in Fig. 12C and used in experimental results described in Fig. 12D(1) and Fig. 12D(2).
[0527] Example 8
[0528] recognizing the human AAVS1 locus with a pair of engineered Bxbl variants.
[0529] Fig. 13A is a comparison of endogenous target site in the human AAVS1 locus that can be recognized with Bxbl variants with the wild-type AAVS1 target. The key elements of both targets are annotated as in Fig. 1C.
[0530] Fig. 13B shows results of screening a collection of previously selected loop, helix, and hairpin variants against quarter site targets corresponding to the left and right halfsites of the target site shown in Fig. 13 A.
[0531] Fig. 13C shows the target sites used to select Bxbl hairpins for the left and right halfsites of the AAVS1 target site shown in Fig. 13 A.
[0532] Fig. 13D is screening data for the indicated combination of loop, helix, and hairpin against the left and right halfsites of the AAVS1 target shown in Fig. 13A. Wild-type Bxbl has no detectable activity on these target sites. If the variant also had no detectable activity, then the comparison shown in the column labeled “activity vs. wt” is 1.0 due to the “pseudo count” applied to raw data for both variant and wild-type Bxbl. Variants yielding improved activity vs. wild-type Bxbl for the left halfsite were then combined with variants yielding improved activity vs. wild-type Bxbl for the right halfsite and tested against the endogenous locus in human K562 cells.
[0533] Fig. 13E(1) shows loop, helix, and hairpin sequences for the Bxbl variants that were combined and tested in human K562 cells to yield the data shown Fig. 13E(2).
[0534] Fig. 13E(2) shows the results of a PCR-based assay designed to detect a junction between the human genome ‘3 of the AAVS1 5032 target site and the integrated donor sequence. The clear bands in the samples treated with a mixture of AAVSl_5032L20 (L20), AAVSl_5032R33 (R33), and wt Bxbl that binds the attP site on the donor construct indicates targeted integration. The lack of detectable bands in the control samples treated with donor and GFP indicates no detectable targeted integration in these samples. The same samples with the indicated treated were assayed with three different pairs of PCR primers that anneal at different parts of the junction between integrated donor and relevant portion of the genome. Adjacent bands with the same indicated sample used different annealing temperatures in the PCR amplification step of the assay. All detectable bands are consistent with the expected amplicon size for the indicated primer pair amplifying the junction sequence.
[0535] Fig. 13F shows additional target sites within the human AAVS1 locus that can be recognized with the invention.
[0536] Example 9
[0537] recognizing the human TRAC locus with a pair of engineered Bxbl variants.
[0538] Fig. 14A shows a comparison of an endogenous target site in the human TRAC locus that can be recognizeed with Bxbl variants and the wild-type natural Bxbl attB target. The key elements of both targets are annotated as in Fig. 1C.
[0539] Fig. 14B shows loop, helix, and hairpin sequences for the Bxbl variants that yielded the highest activity at the endogenous human sequence in the TRAC locus shown in Fig. 14A. The table at the bottom of the figure shows targeted integration levels of a donor construct at the TRAC locus in human K562 cells. The identity of the most active variant pair is shown in the table at the top of the figure.
[0540] Figs. 14C(1) and 14C(2) shows performance of additional Bxbl variants at the same sequence in the TRAC locus shown in Fig. 14B. Fig. 14C(1) shows the helix, loop, and hairpin sequences for these additional variants for each variant and Fig. 14C(2) shows the %TI for the indicated combinations of left and right variant. Wild-type Bxbl was also present in the experiment to bind the donor sequence.
[0541] Example 10
[0542] Donor delivery methods compatible with primary human cells.
[0543] Donor delivery using circular single- stranded DNA. Fig. 18 is a schematic of donor delivery using circular single-stranded DNA.
[0544] Donor delivery using circular double-stranded DNA. Previous examples in this application demonstrate the use of circular double-stranded plasmid DNA as donor molecule. Other circular double-stranded DNA molecules can be used that are better tolerated by primary human cells, including “minicircle” DNA or Nanoplasmids.
[0545] AAV-mediated donor delivery
[0546] Fig. 15A is a schematic of the generation of a circular donor molecule from linear self- complementary AAV. This strategy can be applied to all linear double-stranded DNA molecules and AAV serves as an example.
[0547] Self-complementary AAV constructs were cloned containing genetic cargo flanked by various combinations of variant recombinase attachment sites. The Bxbl attP-GT attachment site was used to integrate the donor into a corresponding attB-GT landing pad cell line, where the attB- GT sequence was installed within the human AAVS1 gene in K562 cells. Bxbl attB-GA and attP- GA attachment sites were used to circularize the linear self-complementary AAV donor, and where applicable flank the attP-GT site. Intramolecular recombination (here “circularization”) is more efficient than intermolecular recombination (here “targeted integration” into the chromosomallanding pad). Consequently, the linear self-complementary AAV donor is likely to first circularize and as such function as a circular donor for subsequent Bxbl-mediated targeted integration into the attB-GT landing pad
[0548] A PCR-based Next Generation Sequencing assay was used to quantify targeted integration into the landing pad, as well as circularization of the donor molecule. There was no observable integration of the donor AAV construct in the absence of an attP-GT sequence. The presence of an attP-GT sequence alone, without any additional attB / attP-GA for donor circularization, sequences resulted in measurable integration. Adding attB / attP-GA sequences and therefore donor circularization capabilities resulted in a substantial increase of targeted integration.
[0549] Material and Methods
[0550] An attB-GT landing pad cell line was established in human K562 cells. Here, attB-GT stands for a Bxbl attB attachment site with GT dinucleotide. Zinc Finger Nuclease were used to introduce a DNA double-strand break within the human AAVS1 gene and integrated a DNA ultramer (see SEQ ID No. 16 below) consisting of the attB-GT sequence flanked by 40-bp homology arms on both sites. A clonal cell line with a single integration event (one allele, one copy) of the attB-GT sequence was selected.
[0551] The attB-GT landing pad cell line was then transfected with a wild-type Bxbl expression construct (Plasmid DNA) and co-transduced with different donor molecules.
[0552] Sequences
[0553] Ultramer (homologies underlines; attB sequence in bold with GT underlined) AGGAGACTAGGAAGGAGGAGGCCTAAGGATGGGGCTTTTCGGCCGGCTTGTCGACGA CGGCGGTCTCCGTCGTCAGGATCATCCGGCAGATAAAAGTACCCAGAACCAGAGCCA CATTAACCGGCC (SEQ ID No. 16)
[0554] Data
[0555] Generation of a circular donor molecule from linear self-complementary AAV. Fig. 15B summarizes the data obtained for scAAV donor results and demonstrates the donor sites and percentage of target integration. Interestingly, the percentage circular donor is higher without binding site / attB-GT.
[0556] Generation of a partially double-stranded ssAAV donor. Fig. 15C(1) shows the oligo binding ssDNA AAV; attP-GT underlined.
[0557] In some scenarios, single- stranded DNA donors may be used (e.g., single-stranded AAV).Targeted integration efficiency levels can be increased by making the donor partially doublestranded. This was achieved by co-transfecting a complementary DNA oligonucleotide resulting in a double- stranded attP-GT sequence for Bxbl.
[0558] Material & Methods
[0559] - The same process was used here as in the above re “Generation of a circular donor molecule from linear self-complementary AAV” and as shown in Fig. 15.
[0560] Data
[0561] Fig. 15C(1) shows the forward and reverse primers used.
[0562] Fig. 15C(2) shows targeted integration efficiency levels are increased by making the donor partially double-stranded.
[0563] Generation of a circular donor from an episomal or chromosomal double-stranded DNA molecule
[0564] Description of the underlying principles
[0565] - The basic idea of this hypothetical example follows the principles of Example 10:Generation of a circular donor molecule from linear self-complementary AAV. Here, recombination occurs between the attB / P-GA sites on a scAAV donor as donor circularization. Stated another way, a circular donor from the scAAV construct is excised. The same could be done from any other double- stranded episomal or chromosomal sequence.
[0566] - Therapeutic application: a lentivirus stably integrates into a genome and variant integrases are then used to excise a circular donor from the integrated lenti.
[0567] Agricultural application: a T-DNA stably integrates into a genome and variant integrases are then used to excise a circular donor from the integrated T-DNA. A similar approach was used with nucleases. Fauser et al., In planta gene targeting, PNAS 109 (19) 7535-7540 (2012)
[0568] In this example attB / P-GA is used for circularization but any other CDN can be used to achieve this. To avoid cross-compatibility with the attachment sites used for e.g. integration, one may choose a different CDN for circularization.
[0569] Example 11
[0570] Donor constructs that work with Bxbl variants
[0571] Fig. 16 shows integration data at the endogenous AAVS1 5032 locus for donor constructs with variant Bxbl target sites tested in human K562 cells. Integration data at endogenous human target sites in human K562 cells shown in Figures 5D, 6D, 9C, 9D, 13E(2), and 14B all used donor constructs with attP sequences recognized by wt Bxbl so wt Bxbl was included in all these experiments. To avoid having to co-deliver wt Bxbl along with the mixture of Bxbl variants required to recognize the left and right halfsites of the endogenous target, donor constructs that can be recognized by one or both Bxbl variants is required. The top panel of this figure shows integration activity for the indicated mixture of Bxbl variants and the indicate donor construct. A “- “ indicates no detectable integration and one or more “+” symbols indicate detectable targeted integration with more “+” symbols indicating higher levels of targeted integration. The attP sequence in each donor is shown to the right of the variant donor’s name. The portions of the attP sequence recognized by the hairpins and helices of the Bxbl variants are underlined. Only the top strand of each donor sequence is shown so the reverse-complement of the underlined portions of the right half- site will describe the helix and hairpin targets in the same way as described in the invention description. Donors LL1 to LL4 are intended to be recognized by two copies of the Bxbl left variant so the underlined portions of the right half are the reverse complement of the underlined portions of the left half of the donor and match the helix and hairpin targets shown in the left side of the AAVS1 target site shown in figure 13A. Similarly, donors RR1, RR2, and RR3 are intended to be bound by two copies of the Bxbl variant that recognizes the right halfsite of the endogenous target in the human AAVS1 locus shown in figure 13 A. Donors LR1 to LR9 are intended to be bound to one copy of the Bxbl variant that recognizes the left side of the target in Figure 13A and one copy of the Bxbl variant that recognizes the right halfsite of the target AAVS1 target shown in figure 13 A. The reverse complement of the underlined GAC on the right side of donors LR1 to LR9 and RR1 to RR3 is GTC and is intended to be recognized by the AGGNLKR or TGGNLRR helices in Bxbl variants AAVS1_5O32R21 (R21), AAVS1_5O32R28 (R28), AAVSl_5032R33 (R33), and AAVSl_5032R55 (R55). The shorthand names in parenthesis are used to indicate these variants in the top panel. AI l lsimilar shorthand notation is used for the variants intended to recognize the left halfsite of the AAVS1 target site: AAVSl_5032L20 (L20), AAVS1_5O32L27 (L27), AAVSl_5032L30 (L30).
[0572] Example 12
[0573] Retargeting other serine integrases
[0574] The method Applicant used successfully to reprogram Bxb 1 will work for any other serine integrase that is structurally similar enough to Bxbl to interact with the DNA target site in the same manner. Further, the method Applicant used successfully to reprogram Bxbl will work for any other recombinase where the binding (or interacting) regions that control DNA target specificity are known or discovered.
[0575] Applicant will perform similar directed evolution selections with the serine integrase sequences shown in figures 17A through 17J. The helix, loop, and hairpin regions for each serine integrase are indicated on the appropriate figure. The 6-residue loop region will be completely randomized with trimer 20 DNA libraries and in a separate experiment completely randomized with NNK randomization, the seven-residue helix region will be randomized completely using a trimer 20 DNA library and in a separate experiment 6 residues randomized with an NNK scheme where the residue at the 5th position of the helix region stay constant. The residues of some of the hairpin region disclosed herein will be randomized. The relevant target site for each region is also shown on each figure indicating which portions of the target site will be changed to match the desired endogenous target site for directed evolution experiments which recombinase variant libraries randomized each of the three regions. Selected recombinase variant sequences that match enriched 4- residue motifs will be screened in the plasmid-based screening system shown in Figure 12 A. Target sites will be designed according to the scheme shown in Figure 12 A, except WT_ZD and WT_RD sequences will use the relevant portions of the natural target site for the relevant integrase instead of the relevant portions of the natural Bxb 1 target site. Applicant will target desired endogenous sites with these new integrases using the same overall strategy diagramed in Figure 3, except that the boundaries of the “quarter sites” may differ slightly for some of these alternative integrases.Examples of other recombinases are shown in Fig. 17A-J.
[0576] Fig. 17A shows amino acid sequence of the integrase named Theia with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Theia indicated. The central dinucleotide (CDN) is also indicated as well as likely conserved bases at -4 and +4’ that will not vary.
[0577] Fig. 17B shows amino acid sequence of the “Veracruz” integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Veracruz indicated. The central dinucleotide (CDN) is also indicated as well as likely conserved bases at -4 and +4’ that will not vary.
[0578] Fig. 17C shows amino acid sequence of the Kp03 integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Kp03 indicated.
[0579] Fig. 17D shows amino acid sequence of the PaOl integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of PaOl indicated.
[0580] Fig. 17E shows amino acid sequence of the Nm60 integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Nm60 indicated.
[0581] Fig. 17F shows amino acid sequence of the Si74 integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Si74 indicated.
[0582] Fig. 17G shows amino acid sequence of the Bcylnt integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Bcylnt indicated.
[0583] Fig. 17H shows amino acid sequence of the Bcelnt integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Bcelnt indicated.
[0584] Fig. 171 shows amino acid sequence of the Ssclnt integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Ssclnt indicated.
[0585] Fig. 17J shows amino acid sequence of the Ssalnt integrase with annotation showing the likely loop, helix, and hairpin regions and target site with portions recognized by the loop, helix, and hairpin regions of Ssalnt indicated.
[0586] Fig. 17K shows a structural alignment of predicted structures of Bxbl, PaOl, Kp03, Nm60, Si74, Bcylnt, and Ssclnt. These structures were predicted from the amino acid sequences of the relevant protein domains using RoseTTAfold and then the predicted structures were aligned with each other and docked with DNA using PyMol. The structures all align well with each other and anyone skilled in the art can use this procedure to identify the loop, helix, and hairpin regions of any LSR whose structure can be accurately predicted using existing or future Al-assisted structure prediction tools. The circles and labels indiate the loop, helix, and hairpin regions of the aligned predicted structures.
[0587] Example 13
[0588] Donor delivery using circular single- stranded DNA
[0589] A double stranded circular DNA produced from a plasmid, minicircle DNA, Nanoplasmid, or synthetic double-stranded circular DNA, can be transformed to a single-stranded circular DNA through the use of a specific endonuclease which nicks only one DNA strand, either the plus strand or the minus strand. The nicked strand is removed through the action of an exonuclease such as T7 exonuclease, or Exo III. The circular single stranded DNA product can then be purified using standard molecular biology methods. This process is diagramed in Fig. 18.
[0590] SEQ ID No. 19 circular plasmid DNA is an example for a DNA molecule that can be used as the double stranded circular DNA starting point
[0591] Approximately 600ug of such DNA will be digested using 10 units of NtBspql in 2000uL of buffer 3.1 (NEB) at 50C for 16 hours. The resulting digested DNA is purified by phenol: chloroform extraction and isopropanol precipitation. After the resulting pellet is redissolved, the nicked plasmid DNA is treated with 1000 units of Exonuclease III (NEB) in 2000uL of Buffer 1. The resulting single-stranded circular DNA is purified by phenol chloroform extraction and ethanol precipitation before being redissolved in nuclease free water.
[0592] Several other methods can be used to produce single-stranded circular DNA by those skilled in the art, including the production of phagemid DNA, or other examples as follows.
[0593] Circular single-stranded DNA can be produced from a linear single-stranded product. Singular single-stranded DNA for instance, can be produced from a linear double-stranded DNA product, if for instance, such a linear double-stranded DNA molecule is a PCR product where one ofthe primers is biotinylated, and after PCR, the products are captured on a solid support (such as a paramagnetic bead) where streptavidin has been immobilized. After the PCR product is captured on a solid support, the non-biotinylated strand can be separated from the complex by incubation with mild sodium hydroxide. The resulting single-stranded DNA can be purified using standard molecular biology techniques, then made circular through splint ligation.
[0594] Circular single-stranded DNA can be produced from a linear single-stranded product. Singular single-stranded DNA for instance, can be produced from a linear double-stranded DNA product, if for instance, such a linear double-stranded DNA molecule is a PCR product where one of the primers is phosphorylated. After PCR, the phosphorylated strand can be selectively removed by the action of a 5 ’phosphate-dependent exonuclease such as lambda exonuclease. The resulting single-stranded DNA can be purified using standard molecular biology techniques, then made circular through splint ligation.
[0595] The circular single-stranded DNA donor can then be delivered to cells using methods described in this application, or other methods known to someone skilled in the art. Cells may require additional time to turn single-stranded DNA into double- stranded DNA and the genome editing outcome can be assayed after 3, 4, 5, 6, and / or 7 days.
[0596] Circular single-stranded synthetic DNA donor molecules can be made partially doublestranded by using a complementary oligonucleotide similar as described in Example 10 in the context of a single- stranded AAV donor.
[0597] Example 14
[0598] Genome-wide integration analysis of Bxbl variants engineered to recognize both half-sites of an endogenous target site show a dramatic shift in specificity towards the intended genomic target
[0599] When the engineered integrases from Figure 9D are assayed for genome-wide integration activity using the method described in Example 3 and the appropriate donor construct we observed a dramatic preference for the intended genomic target site in both cases. Both the s5-6 and s5- 11 DNA target sites have central dinucleotide sequence of CA so the donor construct used in both experiments contained a CA central dinucleotide. The results of the assay performed on 5-6L and 5- 6R are shown in Figure 19A. The intended target site at s5-6 had 573 sequence reads (averaged over multiple replicates) while the next highest locus had 25 sequence reads on average demonstrating a23 -fold preference for the intended target site versus the locus with the second highest number of sequence reads in the human genome. The results of the assay performed on s5- 1 IL and s5- 11R are shown in Figure 19B. The intended target site at s5-l 1 has 818 sequence reads (averaged over multiple replicates) while the next highest locus has 17.25 sequence reads (averaged over multiple replicates) demonstrating a 47-fold preference for the intended locus versus the locus with the second highest number of sequence reads in the human genome.
[0600] Example 15
[0601] Fusion of engineered zinc finger constructs to engineered Bxbl variants can improve integration activity at the TRAC locus in human K562 cells.
[0602] To improve the activity of our Bxbl variants that recognize the human TRAC locus, we designed a panel of zinc finger constructs to recognize different sites adjacent to the endogenous target site for our TRAC Bxbl variants. Molecular modeling indicated that such zinc fingers would need to be fused to the C-terminus of the engineered Bxbl variant in order to function properly. A diagram showing a pair of Bxbl variant-ZFP fusions bound to the human TRAC locus and forming a tetramer with a pair of Bxbl variants bound to the donor is shown in Figure 20. The linker sequences tested as well as the complete amino acid sequence of an example Bxbl variant-ZFP fusion are shown in Figure 21 Amino acid sequences and lengths for LSR-ZFP linkers LI, L2, L3, and L4 are shown at the top of the page while the complete amino acid sequence of the TRAC LI 12 LSR variant (with loop, helix, and hairpin sequences shown in bold) fused to ZFP cflOa (underlined) using linker L2 (in italics) is shown at the bottom of the page. Multiple zinc finger constructs with target sites of different distances from the TRAC Bxbl variant’s attB target site were tested in order to determine if there were strict requirements for the relative positions of the zinc finger target sites and the Bxbl variant target sites. Figure 22 shows the target sites for all ZFPs tested aligned with the human TRAC locus each with four different linker sequences. The number of basepairs between the 3 ’ edge of each ZFP target site and the nearest edge of the TRAC Bxb 1 variant target site are indicated. Figure 23 shows the results of testing different combinations of zinc fingers and linkers fused to either TRAC Ll 12 or TRAC R73 and combined with wildtype Bxbl and the indicated TRAC Bxbl variant without a zinc finger fusion are shown in Figure 32. The results of testing combinations of two different TRAC Bxbl variant-ZFP fusions are shown in Figure 24.
[0603] The linker and Zinc Finger Protein (ZFP) were designed for C-terminal fusion with the Bxbl protein. The linkers were designed in varying lengths from 15 to 40 amino acids. ZFPsequences were generated with internal customized ZFP design pipeline recognizing the regions within 5Obp of the selected TRAC attB target site but don’t overlap the attB sequence). An NLS was added to the C-terminal of ZFPs. The linker and ZFP parts were ordered as eblocks (IDT) and assembled with Bxbl through NEB HiFi assembly (NEB). All plasmids were sequence verified. The Bxbl-ZFP fusion plasmids were transfected into K562 cells using Amaxa HT Nucleofector system (AAU-1001, Lonza). Basically, le5 K562 cells (ATCC) were premixed with SF Cell Line Nucleofector solution with supplement (V5SC-2010, Lonza), 67 ng Bxbl wild-type plasmid, 67 ng each of left and right Bxbl variant-ZFP fusion plasmids, and 1400 ng TRAC attP donor plasmids, and electroporated using the Amaxa HT Nucleofector with program code FF / 120 / DA. Electroporated K562 cells were placed in a 37°C incubator for 3 days. The cells were harvested and lysed with Quick Extract DNA extraction solution (QE09050, Lucigen). To assess the targeted integration (TI) levels at the endogenous TRAC locus, cell lysate was subjected to amplicon sequencing using Illumina NGS technology. Basically, 2ul of the cell lysate was used as template for amplicon amplification with Phusion Hot Start II High-Fidelity PCR Master Mix (F565L, ThermoFisher) following the manufacturer’s instructions using primer oNJS-323 and oNJS-324. Then 2ul of products from the first round of PCR amplification were used in a second round of PCR amplification using primers designed to introduce a sample specific identifier sequence (“barcode”) with i5 and i7 adapter primers. The barcoded amplicons were pooled and sequenced using a Miseq or NextSeq2000 (Illumina, 2x150 PE). The TI results were analyzed through customized bioinformatic pipeline. The sequences of each zinc finger array are as follows:
[0604] >cfla
[0605] RPFQCRICMRNFSTSSNRKTHIRTHTGEKPFACDICGRKFARSDALARHTKIHTGS QKPFQCRICMRKFAQWGTRYRHTKIHTGEKPFQCRICMRNFSQSANRTTHIRTHTGEKPFAC DICGRKFAQRTPRAKHTKIHLRQKDGSGSGSHHHHHHGSGPKKKRKV[1][2] >cf3b[3] RPFQCRICMRNFSQSAHRKNHIRTHTGEKPFACDICGRKFAHRSNLNKHTKIHTGSQKPFQCRICMRNFSQSGSLTR HIRTHTGEKPFACDICGRKFAHRWHLQTHTKIHTGSQKPFQCRICM RNFSDRSNRTTHIRTHTGEKPFACDICGRKF AQNATRINHTKIHLRQKDGSGSGSHHHHHHGSGPKKKRKV[4][5] >cf5a[6] RPFQCRICMRNFSRPYTLRLHIRTHTGEKPFACDICGRKFAQRTPRAKHTKIHTGSQKPFQCRICMRKFAWRSCRSA HTKIHTGEKPFQCRICMRNFSRSDNLSEHIRTHTGEKPFACDICGRKFAQNANRKTHTKIHLRQKDGSGSGSHHHH HHGSGPKKKRKV[7][8] >cf7a[9] RPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFAHRNTLLGHTKIHTGSQKPFQCRICMRNFSTSSNRKT HIRTHTGEKPFACDICGRKFARSDALARHTKIHTGEKPFQCRICMRKFAQWGTRYRHTKIHLRQKDGSGSGSHHHH HHGSGPKKKRKV
[0010]
[0011] >cf9a
[0012] RPFQCRICMRNFSQSGALARHIRTHTGEKPFACDICGRKFAVAEYRYKHTKIHTGSQKPFQCRICMRKFATSSNRKTH TKIHTGEKPFQCRICMRNFSQSGSLTRHIRTHTGEKPFACDICGRKFAHRWHLQTHTKIHLRQKDGSGSGSHHHHH HGSGPKKKRKV
[0013]
[0014] >cfl0a
[0015] RPFQCRICMRNFAQSGNRTTHTKIHTGEKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFAHRNTLLGH TKIHTGSQKPFQCRICMRNFSTSSNRKTHIRTHTGEKPFACDICGRKFARSDALARHTKIHLRQKDGSGSGSHHHHH HGSGPKKKRKV
[0016]
[0017] >cfl5a
[0018] RPFQCRICMRNFSQNAHRKTHIRTHTGEKPFACDICGRKFATKQNRTTHTKIHTGSQKPFQCRICMRNFSQSGALA RHIRTHTGEKPFACDICGRKFAVAEYRYKHTKIHTGEKPFQCRICMRKFATSSNRKTHTKIHLRQKDGSGSGSHHHHH HGSGPKKKRKV
[0019]
[0020] >cfl6a
[0021] RPFQCRICMRNFSQSANRTKHIRTHTGEKPFACDICGRKFAQRTPRAKHTKIHTGSQKPFQCRICMRKFAQSGNRTT HTKIHTGEKPFQCRICMRNFSRSDNLARHIRTHTGEKPFACDICGRKFAHRNTLLGHTKIHLRQKDGSGSGSHHHHH HGSGPKKKRKV
[0022]
[0023] >cr85a
[0024] RPFQCRICMRNFSMVCCRTLHIRTHTGEKPFACDICGRKFARSANLTRHTKIHTGSQKPFQCRICMRNFSRSDHLSQ.HIRTHTGEKPFACDICGRKFAASSTRTKHTKIHLRQKDGSGSGSHHHHHHGSGPKKKRKV
[0025]
[0026] >cr90a
[0027] RPFQCRICMRNFSRSANLARHIRTHTGEKPFACDICGRKFAQSGHLSRHTKIHTGEKPFQCRICMRKFARLDNRTAH TKIHTHPRAPIPKPFQCRICMRNFSRSDVLSTHIRTHTGEKPFACDICGRKFADTRNLRAHTKIHLRQKDGSGSGSHH HHHHGSGPKKKRKV
[0028]
[0029] >cr91g
[0030] RPFQCRICMRNFSMVCCRTLHIRTHTGEKPFACDICGRKFARSANLTRHTKIHTGSQKPFQCRICMRNFSRSDHLSQ HIRTHTGEKPFACDICGRKFAASSTRTKHTKIHTGSQKPFQCRICMRNFSTGQTLRGHIRTHTGEKPFACDICGRKFA QNATRTKHTKIHLRQKDGSGSGSHHHHHHGSGPKKKRK...
Claims
CLAIMS1. A variant recombinase, wherein the variant recombinase comprises, consists of, or consists essentially of, one or more amino acid mutations in at least two binding (or interacting) regions that control DNA target site specificity for the relevant portions of the target site of a wildtype recombinase, wherein the wildtype recombinase has a cognate wildtype recombinase binding region, and wherein the variant recombinase has altered DNA target site specificity relative to the cognate wildtype recombinase binding region.
2. The variant recombinase according to claim 1, comprising, or consisting of, or consisting essentially of at least two mutations relative to the wild-type recombinase in the binding regions.
3. The variant recombinase according to claim 1, comprising, or consisting of, or consisting essentially of at least two mutations relative to the wild-type recombinase in each one of the binding regions.
4. The variant recombinase according to any one of claims 1-3, wherein the binding regions are small enough to allow randomization and re-selection of up to 7 residues within each region to completely alter the DNA target site specificity corresponding to the region.
5. The variant recombinase according to any one of claims 1-4., wherein the binding regions comprise a loop region, a helix region, and a hairpin region, and wherein the loop region comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its or their wildtype counterpart, the helix region, comprises 1, 2, 3, 4, 5, 6, or 7 amino acid changes relative to its or their wildtype counterpart and / or the hairpin region comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to its or their wildtype counterpart.
6. The variant recombinase according to claim 5, wherein the hairpin region is within a zinc ribbon domain (ZD) region and the ZD region comprises, or consists of, or consists essentially of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to its wildtype counterpart.
7. The variant recombinase according to claim 5, wherein the loop region is within a recombinase region (RD) and wherein the loop region comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to its wildtype counterpart.
8. The variant recombinase according to claim 5, wherein the wherein the helix region is within a recombinase region (RD) and wherein the helix region comprises 1, 2, 3, 4, 5, 6, or 7 amino acids changes relative to its wildtype counterpart.
9. The variant recombinase according to any one of claims 1-8, wherein the binding regions comprise one or more of a loop, helix, or hairpin amino acid sequence as shown in any of Figures 5B, 5C(1), 5C(2), 5D, 5E, 6A, 6B, 6C, 9C, 10C(l), 10C(2), 10D, 10E(l), 10E(2), 11C, HE, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13B, 13D, 13E(1), 14B, 14C(1), 16, and / or 21..
10. The variant recombinase according to any one of claims 1-9, wherein the binding regions recognize one or more of the DNA target site sequence as shown in Figures 5 A, 9A, 9B, 10A, 10B(l), 10B(2), HA, 11D, HE, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22.
11. The variant recombinase according to any one of claims 1-10, wherein the variant recombinase comprises an asparagine amino acid residue (N) at the 4thposition of the helix region, a tryptophan (W) amino acid residue at the 3rdposition of the helix region, two lysine residues (K) at the 6thand 7thposition of the helix region, or any one or combinations thereof.
12. The variant recombinase according to any one of claims 1-10, wherein the variant recombinase comprises an asparagine amino acid residue (N) at position 234 of the integrase, a tryptophan (W) amino acid at position 233 of the integrase, two lysines (K) at positions 236 and 237 of the integrase, or any one or combinations thereof.
13. The variant recombinase according to any one of claims 1-10, wherein the variant recombinase comprises an amino acid sequence YRGGLP in the loop region instead of an amino acid sequence YRGSLP from wildtype Bxbl, an amino acid sequence AGGNLKR or YPWSLRR in the helix region instead of an amino acid sequence SATALKR from wildtype Bxbl, an amino acid sequence KAWGSRKTRLYR, MASGSRKTAIYY, or MARGGRKSAIYY in the hairpin region instead of an amino acid sequence FAGGGRKHPRYR from wildtype Bxbl, amino acid sequence AAWALRR or ASHALKR that recognize CAC at positions -11 to -9, amino acid sequence AVQNLKR that recognize TTC at positions -11 to -9, amino acid sequence GGRSLKR that recognize AGC at positions -11 to -9, amino acid sequence GGSHLKR that recognize GCC at positions -11 to -9, amino acid sequence LGTNLKR that recognize ATC at positions -11 to -9, amino acid sequence RAAFLKK that recognize ACA at positions -11 to -9, amino acid sequence RADTLRR that recognize CGC at positions -11 to -9, amino acid sequence RAWTLKC that recognize CTG atpositions -11 to -9, amino acid sequence RGHALKN that recognize ACT at positions -11 to -9, amino acid sequence RGSSLKV that recognize ATT at positions -11 to -9, amino acid sequence SARALSR that recognize GAC at positions -11 to -9, amino acid sequence SGSALKT that recognize AAT at positions -11 to -9, amino acid sequence SGWALRQ that recognize CAT at positions -11 to -9, amino acid sequence SGWGLKK that recognize CAA at positions -11 to -9, amino acid sequence SGYNLRR that recognize CTC at positions -11 to -9, amino acid sequence SRNGLRK that recognize GAA at positions -11 to -9, amino acid sequence TTRTLKR that recognize GGC at positions -11 to -9, amino acid sequence YSRNLKR that recognize GTC at positions -11 to -9, or any one or combinations thereof.
14. The variant recombinase according to any one of claims 1-13, wherein the recombinase is Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Ssalnt, Dn29, PhiRvl, Al 18, or TP901.
15. The variant recombinase according to any one of claims 1-14, wherein the variant recombinase comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 amino acid changes relative to wildtype Bxbl at positions 308, 309, 310, 311, 312, 313, 314, 316, 318, 321, 322, 323, and 325 in the hairpin region or wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acid changes relative to wildtype Bxbl at positions 154, 155, 156, 157, 158, and 159 in the loop region, wherein the variant recombinase comprises 1, 2, 3, 4, 5, or 6 amino acids changes relative to wildtype Bxbl at positions 231, 232, 233, 234, 236, and 237 of the helix region, or any one or combinations thereof.
16. The variant recombinase according to any one of claims 1-15, wherein the recombinase recognizes AAVS1 or TRAC, optionally wherein the recombinase recognizes a sequence comprising Seq ID No. 17 or Seq ID No. 18.
17. A variant recombinase comprising binding (or interacting) regions from different recombinases.
18. The variant recombinase of claim 17, wherein the binding regions comprise RD and / or ZD domains.
19. The variant recombinase of claim 18, wherein the RD and / or ZD domains are from a different wildtype recombinase.
20. The variant recombinase of any one of claims 18-19, wherein the RD and / or ZD domains arefrom a different modified recombinase.
21. The variant recombinase of any one of claims 18-20, wherein the RD and / or ZD domains comprise one or more amino acid changes relative to the wildtype counterpart domains.
22. The variant recombinase of any one of claims 17-21, wherein the binding regions comprise one or more of a loop, helix, or hairpin amino acid sequence as shown in any of Figures 5B, 5C(1), 5C(2), 5D, 5E, 6A, 6B, 6C, 9C, 10C(l), 10C(2), 10D, 10E(l), 10E(2), 11C, HE, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13B, 13D, 13E(1), 14B, 14C(1), 16, and / or 21..
23. The variant recombinase of any one of claims 17-21, wherein the binding regions recognizeone or more of the DNA target sites as shown in Figures 5A, 9A, 9B, 10A, 10B(l), 10B(2), 11 A, 1 ID, HE, 12C, 12D(1), 12D(2), 12E(1), 12E(2), 12F(1), 12F(2), 13A, 13B, 13C, 13D, 13F, 14A, 16, and / or 22.
24. The variant recombinase of any one of claims 17-23, wherein the binding regions (e g , RD and / or ZD domains) are from Bxbl, PhiC31, LI integrase, Theia integrase, Veracruz Integrase, Kp03 integrase, PaOl integrase, Nm60 integrase, Si74 integrase, Bcylnt, Bcelnt, Ssclnt, Dn29, PhiRvl, Al 18, TP901, or Ssalnt.
25. The variant recombinase according to any one of claims 17-24, wherein the recombinase recognizes AAVS1 or TRAC, optionally wherein the recombinase recognizes a sequence comprising Seq ID No. 17 or Seq ID No. 18.
26. The variant recombinase according to any one of claims 1-25, wherein the variant recombinase further comprises a zinc finger array fused to the C-terminus of the variant recombinase using a polypeptide linker, wherein the zinc finger array recognizes a DNA sequence, and wherein the 3’ edge of the zinc finger array DNA target site is separated by 4, 5, 6, 7, 8, 9, or 10 basepairs from the edge of the recombination site for the variant recombinase, site for the variant recombinase.
27. A nucleic acid molecule encoding the variant recombinase of any one of claims 1-26.
28. A vector encoding, or comprising, or consisting of, or consisting essentially of, the variant recombinase of any one of claims 1-26.
29. A recombinant virus comprising the nucleic acid construct of claim 27, optionally wherein the recombinant virus is a recombinant AAV.
30. A host cell comprising the nucleic acid construct of claim 27.
31. A method of integrating a synthetic DNA donor construct into a non-coding region, a safe harbor locus, or into or near a gene of interest, comprising expressing a variant recombinase according to any one of claims 1-26 in a cell.
32. A method of integrating a synthetic DNA donor construct into a non-coding region, a safe harbor locus, or into or near a gene of interest, comprising expressing a first and a second variant recombinase in a cell, wherein the first and the second variant recombinase is each according to any one of claims 1-26, wherein the first variant recombinase recognizes a left halfsite of a genomic attB DNA site and the second variant recombinase recognizes a right halfsite of the same genomic attB DNA site, alternatively wherein the first variant recombinase recognizes a left halfsite of a genomic attP DNA site and the second variant recombinase recognizes a right halfsite of the same genomic attP DNA site.
33. The method of claim 32, wherein the first and the second variant recombinases are cotransfected with a donor construct with an attP site, wherein the variant recombinase that recognizes the left halfsite of a genomic attB site recognizes both halfsites of the attP site on the donor construct, alternatively wherein the variant recombinase that recognizes the left halfsite of a genomic attP site recognizes both halfsites of the attB site on the donor construct.
34. The method of claim 32, wherein the first and the second variant recombinases are cotransfected with a donor construct with an attP site, wherein the variant recombinase that recognizes the right halfsite of a genomic attB site recognize both halfsites of the attP site on the donor construct, alternatively wherein the variant recombinase that recognizes the right halfsite of a genomic attP site recognize both halfsites of the attB site on the donor construct.
35. The method of claim 32, wherein the first and the second variant recombinases are cotransfected with a donor construct with an attP site, wherein the variant recombinase that recognizes the left halfsite of a genomic attB site recognize a first halfsites of the attP site on the donor construct and the variant recombinase that recognizes the right halfsite of a genomic attB site recognize the second halfsite of the attP site on the donor construct, alternatively wherein the first and the second variant recombinases are cotransfected with a donor construct with an attB site, wherein the variant recombinase that recognizes the left halfsite of a genomic attP site recognize a first halfsites of the attB site on the donor construct and the variant recombinase that recognizes the right halfsite of a genomic attP site recognize the second halfsite of the attB site on the donor construct.
36. A method of expressing or repressing a therapeutically or industrially relevant gene of interest in a cell, comprising introducing to the cell a nucleic acid according to claim 27.
37. A method of treating a disease in a patient, comprising administering to the patient a variant recombinase according to any one of claims 1-26.
38. Use of a variant recombinase of any one of claims 1-26, a nucleic acid construct of claim 27, or a recombinant virus of claim 29 for the manufacture of a medicament in the method of any one of claims 36-37.
39. A method for generating a large serine recombinase variant having an altered DNA target site specificity, comprising:1) identifying positions of amino acids that control DNA target specificity in a starting recombinase;2) randomizing the positions of amino acids that control DNA target specificity to generate recombinase variants; and3) selecting a recombinase variant having the altered DNA specificity.
40. A method for generating a large serine recombinase variant that can recognize a desired DNA sequence, comprising:1) Dividing the desired genomic DNA target site into 4 separate portions that correspond to the region recognized by i) the left halfsite ZD domain, ii) the left halfsite RD domain, iii) the right halfsite RD domain and iv) the right halfsite ZD domain 2) for each portion of the desired target site recognized by a ZD domain, generating a library of recombinase variants with at least two amino acid changes within the hairpin region of the recombinase and selecting this library for members that can recombine a target site comprising the relevant portion of the desired genomic target site 3) for each portion of the desired target site recognized by an RD domain, generating a library of recombinase variants with at least two amino acid changes with the helix and / or loop regions, and selecting members of this library that can recombine target sites comprising the relevant portion of the genomic target site 4) combining selected recombinase variants selected against two regions of the same halfsite and then testing such composite variants with amino acid changes in both RD and ZD domains for recombination activity with a target site comprising the relevant desired halfsite 5) delivering a mixture of a recombinase variant active on the desired left halfsite, a recombinase variant active on the right halfsite, and an appropriate donor construct into eukaryotic cells and then measuring the amount of targeted integration of the donor construct into the desired genomic location.
41. A split antibiotic resistance gene comprising a recombinase target sequence, preferably an attL, attP, and / or an attB sequence.
42. The split antibiotic resistance gene of claim 41, wherein the gene encodes an aaCCl protein.
43. The split antibiotic resistance gene of any one of claims 41-42, wherein the target sequence is inserted into a surface-exposed loop of the protein.
44. The split antibiotic resistance gene of any one of claims 41-43, wherein the target sequence is inserted between the following residues: loop 83-84.
45. A method for selecting or identifying a recombinase having activity against a target sequence, comprising: 1) generating a split antibiotic resistance gene according to any one of claims 41-44; 2) introducing the gene into bacterial cells; 3) introducing into the bacterial cells a recombinase from a pool of active and inactive recombinases, wherein each bacterial cell preferentially expresses a one or more types of recombinase, 4) administering an antibiotic to the bacterial cells; and 5) identifying the recombinase in surviving bacterial cells.
Citation Information
Patent Citations
Serine recombinase systems for site-specific gene editing
WO2023147507A1