Serine recombinases for gene editing

A gene editing system using serine recombinases with high sequence identity and specific binding sites, delivered via vectors, addresses the challenge of integrating nucleic acid sequences into eukaryotic genomes, achieving precise and efficient integration.

JP2025539240APending Publication Date: 2025-12-04METAGENOMI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025525684
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-27
Filing Date
2023-11-06
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Current gene editing technologies lack efficient and precise methods for integrating nucleic acid sequences into eukaryotic genomes using serine recombinases, particularly in terms of sequence identity and delivery mechanisms.

Method used

A gene editing system utilizing serine recombinases with at least 80% sequence identity to specific SEQ IDs, combined with donor polynucleotides and binding site sequences, delivered via vectors such as plasmids or viruses, enables targeted integration and recombination in eukaryotic genomes.

Benefits of technology

The system achieves precise and efficient integration of nucleic acid sequences into eukaryotic genomes, facilitating the delivery of therapeutic agents, reporters, and markers, with high transduction efficiency and sequence specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539240000001_ABST
    Figure 2025539240000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a gene editing system comprising a serine recombinase and a method for using such a serine recombinase for the integration of nucleic acid sequences. More specifically, the present disclosure relates to a sequence-defined serine recombinase having a binding site such as a bacterial genome recombination sequence (attB). Methods for the recombinant production of the recombinase are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 382,690, filed November 7, 2022, and U.S. Provisional Patent Application No. 63 / 510,567, filed June 27, 2023, each of which is incorporated by reference in its entirety. Summary of the Invention

[0002] The present disclosure is based, in part, on the development of serine recombinases for use in gene editing systems to integrate nucleic acid sequences.

[0003] The present disclosure describes a gene editing system comprising: a) a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214, or a nucleic acid encoding the serine recombinase; and b) a donor polynucleotide and a nucleic acid comprising a first binding site sequence. In some embodiments, the first binding site sequence is 5' to the donor polynucleotide. In some embodiments, the nucleic acid encoding the serine recombinase further comprises a second binding site sequence. In some embodiments, the second binding site sequence is 5' to the serine recombinase. In some embodiments, the first binding site sequence and the second binding site sequence are capable of recombination. In some embodiments, the first binding site sequence is a bacterial genomic recombination sequence (attB). In some embodiments, the first binding site sequence is a phage genomic recombination sequence (attP). In some embodiments, the second binding site sequence is a bacterial genomic recombination sequence (attB). In some embodiments, the second binding site sequence is a phage genomic recombination sequence (attP). In some embodiments, the attB sequence comprises from about 20 nucleotides to about 500 nucleotides. In some embodiments, the attP sequence comprises from about 20 nucleotides to about 500 nucleotides. In some embodiments, the attB sequence is set forth in SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7400, 7410, 7411, 7412, 7413, 7414, 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402.In some embodiments, the attB sequence comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attP sequence comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399, and 7404. In some embodiments, the attP sequence comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the nucleic acid comprising the donor polynucleotide and the first binding sequence is delivered by a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid. In some embodiments, the nucleic acid encoding the serine recombinase is delivered by a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus.In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10 , AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV In some embodiments, the herpesvirus is AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, or AAV-HSC16, or a derivative thereof. In some embodiments, the herpesvirus is HSV-1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8. In some embodiments, the donor polynucleotide has a size of at least about 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, or greater than 120 kb. In some embodiments, the donor polynucleotide encodes a therapeutic agent, a reporter, or a marker. In some embodiments, the reporter comprises a fluorescent protein. In some embodiments, the fluorescent protein is GFP, EBFP, EBFP2, azurite, mKalamal, ECFP, cerulean, CyPet, YFP, citrine, Venus, YPet, RFP, CFP, or a derivative thereof.In some embodiments, the reporter is acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), beta-galactosidase (LacZ), beta-glucuronidase (GUS), chloramphenicol acetyltransferase (CAT), horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, or a derivative thereof. In some embodiments, the marker is an antibiotic resistance marker. In some embodiments, the antibiotic resistance marker is kanamycin, spectinomycin, streptomycin, ampicillin, carbenicillin, bleomycin, erythromycin, polymyxin B, tetracycline, chloramphenicol, neomycin, zeocin, or a derivative thereof. In some embodiments, the marker is a cell surface marker.

[0004] A eukaryotic genome comprising a donor polynucleotide sequence and an attL sequence on the 5' side of the donor polynucleotide sequence, wherein the attL sequence is selected from the group consisting of SEQ ID NOs: 17 to 18, 7145, 7147, 7150, GCATCCCC, TATTCGAT, GGGCAACC, GGGCACCC, CAAGTTC, ACCGCC, CATATGT, 7219, 7224, 7231, 7232, 7237, 7242, ATGGTGGGC, 7252, GCCATTTC, TCAGCT

[0013] The present disclosure provides a eukaryotic genome comprising a sequence selected from the group consisting of: CCA, 7269, 7270, 7275, 7276, 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC. In some embodiments, the eukaryotic genome further comprises an attR sequence 3' to the donor polynucleotide sequence.

[0005] A eukaryotic genome comprising a donor polynucleotide sequence and an attL sequence on the 3' side of the donor polynucleotide sequence, wherein the attL sequence is selected from the group consisting of SEQ ID NOs: 17 to 18, 7145, 7147, 7150, GCATCCCC, TATTCGAT, GGGCAACC, GGGCACCC, CAAGTTC, ACCGCC, CATATGT, 7219, 7224, 7231, 7232, 7237, 7242, ATGGTGGGC, 7252, GCCATTTC, TCAGCT

[0013] The present disclosure describes a eukaryotic genome comprising a sequence selected from the group consisting of: CCA, 7269, 7270, 7275, 7276, 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC. In some embodiments, the eukaryotic genome further comprises an attR sequence 3' to the donor polynucleotide sequence.

[0006] A donor polynucleotide sequence and an attL sequence on the 5' or 3' side of the donor polynucleotide sequence, which are SEQ ID NOs: 17 to 18, 7145, 7147, 7150, GCATCCCC, TATTCGAT, GGGCAACC, GGGCACCC, CAAGTTC, ACCGCC, CATATGT, 7219, 7224, 7231, 7232, 7237, 7242, ATGGTGGGC, 7252, GCCATTTC, TCAGCTCCA, 7260, 7261, 7262, 7263, 7264, 7265, 7266, 7267, 7268, 7269, 7270, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 73 an attL sequence comprising a sequence selected from the group consisting of 269, 7270, 7275, 7276, 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC; and attR sequences on the 5' or 3' side of the polynucleotide sequence, and and an attR sequence comprising a sequence selected from the group consisting of 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC. In some embodiments, the attL sequence and the attR sequence are identical. In some embodiments, the attL sequence is a recombination sequence of a first binding site sequence and a second binding site sequence. In some embodiments, the attR sequence is a recombination sequence of a first binding site sequence and a second binding site sequence. In some embodiments, the first binding site sequence is a bacterial genome recombination sequence (attB).In some embodiments, the first binding site sequence is a phage genomic recombination sequence (attP). In some embodiments, the second binding site sequence is a bacterial genomic recombination sequence (attB). In some embodiments, the second binding site sequence is a phage genomic recombination sequence (attP). In some embodiments, attB comprises from about 20 nucleotides to about 500 nucleotides. In some embodiments, attP comprises from about 20 nucleotides to about 500 nucleotides. In some embodiments, attB is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7292, 7293, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 741 77, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402. In some embodiments, attB comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, attP is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 740 79, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399, and 7404.In some embodiments, attP comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, attL comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 3, 4, 7, 8, 11, 12, 15, 16, 7153, 7157, 7161, 7165, 7169, 7173, 7177, and 7181. In some embodiments, attR comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 3, 4, 7, 8, 11, 12, 15, 16, 7154, 7158, 7162, 7166, 7170, 7174, 7178, and 7182.

[0007] The present disclosure describes a mammalian cell comprising a eukaryotic genome described herein. In some embodiments, the mammalian cell is a human cell. In some embodiments, the mammalian cell further comprises a serine recombinase. In some embodiments, the serine recombinase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214. In some embodiments, the serine recombinase comprises a sequence having at least about 80% sequence identity to SEQ ID NO: 21. In some embodiments, the serine recombinase comprises a sequence having at least about 80% sequence identity to SEQ ID NO: 22. In some embodiments, the serine recombinase comprises a sequence having at least about 80% sequence identity to SEQ ID NO: 23. In some embodiments, the serine recombinase comprises a sequence having at least about 80% sequence identity to SEQ ID NO: 24. In some embodiments, the serine recombinase comprises a transduction efficiency of at least about 5%. In some embodiments, the serine recombinase comprises a transduction efficiency of at least about 25%. In some embodiments, the serine recombinase comprises a transduction efficiency of at least about 50%. In some embodiments, the serine recombinase can target a gene comprising a catalase domain or a synthase domain. In some embodiments, the catalase is manganese catalase. In some embodiments, the synthase is queosine synthase. In some embodiments, the serine recombinase can target a gene comprising a DUF4244 Pfam domain.

[0008] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142.

[0009] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:21.

[0010] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:22.

[0011] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:23.

[0012] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:24.

[0013] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:1848.

[0014] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:7111.

[0015] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:7115.

[0016] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:7131.

[0017] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:7136.

[0018] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:7139.

[0019] In this disclosure, we describe a eukaryotic cell that contains a serine recombinase that contains at least about 80% sequence identity to SEQ ID NO:7140.

[0020] In some embodiments, the eukaryotic cell is a mammalian cell, hi some embodiments, the eukaryotic cell is a human cell.

[0021] The present disclosure describes a vector comprising: a) a nucleic acid encoding a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOS: 21-7060 and 7105-7142; and b) one or more regulatory elements. In some embodiments, the one or more regulatory elements comprise a promoter, an enhancer, an intron, a microRNA, a linker, a splicing element, or a polyA signal. In some embodiments, the promoter is selected from a constitutive promoter, an inducible promoter, a minipromoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, polH, EM7, OpIE1, and derivatives thereof.

[0022] The present disclosure describes a vector comprising a nucleic acid encoding a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142, wherein the vector is selected from the group consisting of a plasmid, a nanoplasmid, a phagemid, a phage derivative, a bacmid, a bacterial artificial chromosome (BAC), a minicircle, a doggybone, a yeast artificial chromosome (YAC), and a cosmid.

[0023] The present disclosure describes a method for gene editing, comprising: a) providing or specifying a first binding site sequence to or in a host genome; b) providing a host cell with a nucleic acid comprising a donor polynucleotide and a second binding site sequence; and c) contacting the host cell with a serine recombinase having at least about 80% sequence identity to any of SEQ ID NOS: 21-7060 and 7105-7142, or a nucleic acid encoding the serine recombinase, wherein the first binding site sequence and the second binding site sequence are capable of recombination. In some embodiments, the first binding site sequence is endogenous to the host genome. In some embodiments, the first binding site sequence is provided using viral delivery. In some embodiments, the first binding site sequence is provided using a transposase. In some embodiments, the first binding site sequence is provided using a nuclease. In some embodiments, the nuclease is a double-stranded nuclease. In some embodiments, the nuclease is a Type II CRISPR endonuclease. In some embodiments, the nuclease is a Type V CRISPR endonuclease. In some embodiments, the nuclease is Cas9. In some embodiments, the first binding site sequence is provided using a reverse transcriptase. In some embodiments, the second binding site sequence is 5' to the donor polynucleotide. In some embodiments, the first binding site sequence is a bacterial genomic recombination sequence (attB). In some embodiments, the first binding site sequence is a phage genomic recombination sequence (attP). In some embodiments, the second binding site sequence is a bacterial genomic recombination sequence (attB). In some embodiments, the second binding site sequence is a phage genomic recombination sequence (attP). In some embodiments, the attB sequence comprises from about 20 nucleotides to about 500 nucleotides. In some embodiments, the attP sequence comprises from about 20 nucleotides to about 500 nucleotides.In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402. In some embodiments, the attB sequence comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402. In some embodiments, the attP sequence comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the nucleic acid comprising the donor polynucleotide and the second binding site sequence is delivered by a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid.In some embodiments, the nucleic acid encoding the serine recombinase is delivered by a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10 , AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV In some embodiments, the herpesvirus is AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, or AAV-HSC16, or a derivative thereof. In some embodiments, the herpesvirus is HSV-1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8. In some embodiments, the donor polynucleotide has a size of at least about 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, or greater than 120 kb.In some embodiments, the donor polynucleotide encodes a therapeutic agent, a reporter, or a marker. In some embodiments, the reporter comprises a fluorescent protein. In some embodiments, the fluorescent protein is GFP, EBFP, EBFP2, azurite, mKalamal, ECFP, cerulean, CyPet, YFP, citrine, venus, YPet, RFP, CFP, or a derivative thereof. In some embodiments, the reporter is acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), beta-galactosidase (LacZ), beta-glucuronidase (GUS), chloramphenicol acetyltransferase (CAT), horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, or a derivative thereof. In some embodiments, the marker is an antibiotic resistance marker. In some embodiments, the antibiotic resistance marker is kanamycin, spectinomycin, streptomycin, ampicillin, carbenicillin, bleomycin, erythromycin, polymyxin B, tetracycline, chloramphenicol, neomycin, zeocin, or a derivative thereof, hi some embodiments, the marker is a cell surface marker. [Brief explanation of the drawings]

[0024] The features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings.

[0025] [Figure 1] Figure 1 shows a multiple sequence alignment of the MG178 family large serine recombinase (LSR) candidate versus the Bxb1 LSR reference sequence. The resolving domain, recombinase domain, and Zn finger domain are shown as boxes, and the catalytic residues required for activity are highlighted as bars below each residue. [Figure 2A]Figures 2A and 2B show the phylogenetic protein tree of the LSRs of this disclosure. The tree was inferred from a global multiple sequence alignment of LSR sequences clustered at 90% amino acid identity (AAI). Selected MG178 family candidates are highlighted with large dots and color-coded by the bacterial host they target (Figure 2A) or the host gene into which they are inserted (Figure 2B). [Figure 2B] Figures 2A and 2B show the phylogenetic protein tree of the LSRs of this disclosure. The tree was inferred from a global multiple sequence alignment of LSR sequences clustered at 90% amino acid identity (AAI). Selected MG178 family candidates are highlighted with large dots and color-coded by the bacterial host they target (Figure 2A) or the host gene into which they are inserted (Figure 2B). [Figure 3] Figure 3 shows an analysis of exemplary LSR integration sites identified from alignments of genomic fragments with and without prophage. The top panel shows a multiple sequence alignment of the genomic fragment with the integrated prophage (top) and its non-integrated host (bottom). Genes are predicted as arrows, and functional domains supporting functional annotation are represented by black bars below the gene. The prophage integrates into a gene predicted by CheckV (top) and with a quienosin biosynthesis protein annotation (bottom). The bottom panel shows a graph demonstrating that validation of the prophage boundaries allows the determination of a common core motif shared with the non-integrated host. The LSR gene is located on one of the prophage edges (black box). [Figure 4A] 4A-4C show schematic diagrams of exemplary in vitro screening procedures for serine recombinase recombination activity: Figure 4A shows a schematic diagram of in vitro recombinase expression from linear or circular dsDNA constructs. [Figure 4B]Figures 4A-4C show schematics of an exemplary in vitro screening procedure for serine recombinase recombination activity. Figure 4B shows a schematic of a recombination reaction using integrase, which is added to the recombination reaction along with attP and attB dsDNA fragments specific for the serine recombinase. [Figure 4C] Figures 4A-4C show a schematic diagram of an exemplary in vitro screening procedure for serine recombinase recombination activity. Figure 4C shows a schematic diagram of PCR analysis by agarose gel electrophoresis of recombinant DNA amplified with attL- and attR-specific primers. [Figure 5A]Figures 5A-5B show the results of in vitro recombinase assays for LSRs MG178-4, MG178-9, MG178-10, and MG178-11. Arrows indicate positive recombination event products. Figure 5A shows the results of in vitro recombinase assays using AttL-specific primers used to amplify potential recombination events. Lane 1 shows the negative control of MG178-4, which contains the MG178-4 attB and MG178-4 attP dsDNA fragments. Lane 2 shows the experimental conditions for MG178-4, which contains the MG178-4 attB and MG178-4 attP dsDNA fragments and expresses the MG178-4 recombinase. Lane 3 shows the negative control of MG178-9, which contains the MG178-9 attB and MG178-9 attP dsDNA fragments. Lane 4 shows the experimental condition of MG178-9, which contained the MG178-9 attB and MG178-9 attP dsDNA fragments and expressed the MG178-9 recombinase. Lane 5 shows the negative control, MG178-10, which contained the MG178-10 attB and MG178-10 attP dsDNA fragments. Lane 6 shows the experimental condition of MG178-10, which contained the MG178-10 attB and MG178-10 attP dsDNA fragments and expressed the MG178-10 recombinase. Lane 7 shows the negative control, MG178-11, which contained the MG178-11 attB and MG178-11 attP dsDNA fragments. Lane 8 shows the experimental conditions for MG178-11, which contained the MG178-11attB and MG178-11attP dsDNA fragments and expressed the MG178-11 recombinase. [Figure 5B]Figures 5A-5B show the results of in vitro recombinase assays for LSRs MG178-4, MG178-9, MG178-10, and MG178-11. Arrows indicate positive recombination event products. Figure 5B shows the results of in vitro recombinase assays using AttR-specific primers used to amplify potential recombination events. Lane 1 shows the negative control of MG178-4, which contains the MG178-4 attB and MG178-4 attP dsDNA fragments. Lane 2 shows the experimental conditions for MG178-4, which contains the MG178-4 attB and MG178-4 attP dsDNA fragments and expresses the MG178-4 recombinase. Lane 3 shows the negative control of MG178-9, which contains the MG178-9 attB and MG178-9 attP dsDNA fragments. Lane 4 shows the experimental condition of MG178-9, which contained the MG178-9 attB and MG178-9 attP dsDNA fragments and expressed the MG178-9 recombinase. Lane 5 shows the negative control, MG178-10, which contained the MG178-10 attB and MG178-10 attP dsDNA fragments. Lane 6 shows the experimental condition of MG178-10, which contained the MG178-10 attB and MG178-10 attP dsDNA fragments and expressed the MG178-10 recombinase. Lane 7 shows the negative control, MG178-11, which contained the MG178-11 attB and MG178-11 attP dsDNA fragments. Lane 8 shows the experimental conditions for MG178-11, which contained the MG178-11attB and MG178-11attP dsDNA fragments and expressed the MG178-11 recombinase. [Figure 6]Figure 6 shows a schematic of the experimentally verified MG178-10 attL sequence aligned with the bioinformatically identified MG178-10 attL. Black bars indicate 100% identity with the reference sequence (Found attL). The bottom panel shows a magnified sequence view of the alignment of the reconstructed attP, reconstructed attB, and experimentally determined attL sites to the bioinformatically identified attL site for MG178-10. The gray highlighted sequences reflect the identity of the reconstructed attP and attB sites, while lighter bases indicate mismatched alignments from the reference sequence (bioinformatically identified attL). Boxed sequences highlight the conservation of the common core across the found attL, attP, attB, and sequenced attL. Figure 6 discloses SEQ ID NOs: 7435-7437, and 7435, respectively, in order of appearance. [Figure 7] Figure 7 shows a schematic of the experimentally verified MG178-10 attR sequence aligned with the bioinformatically identified MG178-10 attR. Black bars indicate 100% identity to the reference sequence (Found attR). The bottom panel shows a magnified sequence view of the alignment of the reconstructed attP, reconstructed attB, and experimentally determined attR sites to the bioinformatically identified attR site for MG178-10. The gray-highlighted sequence reflects the identity of the reconstructed attP and attB sites, while lighter-colored bases indicate mismatched alignments from the reference sequence (bioinformatically identified attR). Boxed sequences highlight the conservation of the common core across the found attR, attP, attB, and sequenced attR. Figure 7 discloses SEQ ID NOs: 7438-7440 and 7438, respectively, in order of appearance. [Figure 8] Figure 8 shows a multiple sequence alignment of the MG178 LSR candidate versus the Bxb1 LSR reference sequence. The resolving domain, recombinase domain, and zinc finger domain are shown as boxes, and the catalytic residues required for activity are highlighted as bars below each residue. [Figure 9A] Figures 9A-9C show pairwise alignments of the 3' and 5' regions flanking the proviruses of MG178-7202 (Figure 9A), MG178-1859 (Figure 9B), and MG178-7193 (Figure 9C). Annotated are the proviral boundaries and common core. Proviral boundaries were predicted and determined by aligning provirus-containing contigs with contigs lacking provirus. The common core was identified by finding conserved regions within the alignment. When the alignment did not show conservation (Figure 9C), repeats were identified within and outside the proviral boundaries, and the alignment was manually refined. In order of appearance, Figure 9A discloses SEQ ID NOs: 7441-7443, Figure 9B discloses SEQ ID NOs: 7444-7446, and Figure 9C discloses SEQ ID NOs: 7447-7449, respectively. [Figure 9B] Figures 9A-9C show pairwise alignments of the 3' and 5' regions flanking the proviruses of MG178-7202 (Figure 9A), MG178-1859 (Figure 9B), and MG178-7193 (Figure 9C). Annotated are the proviral boundaries and common core. Proviral boundaries were predicted and determined by aligning provirus-containing contigs with contigs lacking provirus. The common core was identified by finding conserved regions within the alignment. When the alignment did not show conservation (Figure 9C), repeats were identified within and outside the proviral boundaries, and the alignment was manually refined. In order of appearance, Figure 9A discloses SEQ ID NOs: 7441-7443, Figure 9B discloses SEQ ID NOs: 7444-7446, and Figure 9C discloses SEQ ID NOs: 7447-7449, respectively. [Figure 9C]Figures 9A-9C show pairwise alignments of the 3' and 5' regions flanking the proviruses of MG178-7202 (Figure 9A), MG178-1859 (Figure 9B), and MG178-7193 (Figure 9C). Annotated are the proviral boundaries and common core. Proviral boundaries were predicted and determined by aligning provirus-containing contigs with contigs lacking provirus. The common core was identified by finding conserved regions within the alignment. When the alignment did not show conservation (Figure 9C), repeats were identified within and outside the proviral boundaries, and the alignment was manually refined. In order of appearance, Figure 9A discloses SEQ ID NOs: 7441-7443, Figure 9B discloses SEQ ID NOs: 7444-7446, and Figure 9C discloses SEQ ID NOs: 7447-7449, respectively. [Figure 10A] Figures 10A-10B show LSR-mediated binding site recombination events and cellular plasmid recombination activity. Figure 10A shows a schematic diagram illustrating LSR-mediated binding site recombination events. [Figure 10B] Figures 10A-10B show LSR-mediated binding site recombination events and cellular plasmid recombination activity. Figure 10B shows a bar graph depicting recombination activity. Active LSRs with greater than 5% recombination are plotted relative to BxB1 as a reference. Each bar represents an experimental condition using recombinase, AttB, and AttP plasmids transfected into HEK293T cells. Plasmid recombination was quantified by flow cytometry after 48 hours, and percent recombination was calculated based on cells expressing both eGFP (recombinase protein) and mCherry (recombination events). Error bars are included for candidates with replicates. [Figure 11]Figure 11 shows the results of in vitro recombinase assays for the LSR strains MG178-7202, MG178-1859, MG178-7193, and MG178-7177. Lane 1 shows the ladder. Lane 2 shows the experimental conditions for MG178-7202, containing the MG178-7202 attB and MG178-7202 attP dsDNA fragments, with the addition of expressed MG178-7202 recombinase. Lane 3 shows the experimental conditions for MG178-1859, containing the MG178-1859 attB and MG178-1859 attP dsDNA fragments, with the addition of expressed MG178-1859 recombinase. Lane 4 shows the experimental condition of MG178-7193 containing the MG178-7193attB and MG178-7193attP dsDNA fragments with the addition of the expressed MG178-7193 recombinase. Lane 5 shows the experimental condition of MG178-7177 containing the MG178-7177attB and MG178-7177attP dsDNA fragments with the addition of the expressed MG178-7177 recombinase. Lane 6 shows the negative control ladder. Lane 7 shows the negative control of MG178-7202 containing the MG178-7202 attB and MG178-7202 attP dsDNA fragments but without the enzyme. Lane 8 shows the negative control MG178-1859 containing the MG178-1859 attB and MG178-1859 attP dsDNA fragments but no enzyme. Lane 9 shows the negative control MG178-7193 containing the MG178-7193 attB and MG178-7193 attP dsDNA fragments but no enzyme. Lane 10 shows the negative control MG178-7177 containing the MG178-7177 attB and MG178-7177 attP dsDNA fragments but no enzyme. [Figure 12] Figure 12 shows a bar graph depicting active candidates in human cells. Percent recombination was determined as the percentage of mCherry-positive cells (recombinant) divided by the total number of eGFP-positive cells (integrase transfection and expression). [Figure 13A]Figures 13A-B show the discovery of plasmid dosage for optimal plasmid transfection concentration in human cells. Figure 13A shows a bar graph showing percent recombination. [Figure 13B] Figures 13A-13B show the plasmid dosage finding for optimal plasmid transfection concentrations in human cells. Figure 13B shows a table outlining the conditions tested. Optimal performance for MG178-7202 was found with equal weights of 250 ng of integrase, attB, and attP plasmids per transfection, respectively. [Figure 14] Figure 14 shows a bar graph demonstrating the minimization of the binding site of MG178-7202(-47) in human cells. The AttB site was tested from 108 nt to 28 nt, and the AttP from 68 nt to 48 nt. While measurable recombination can be measured down to a 32 nt AttB, the optimal conditions were determined to be a 48 nt AttB and a 58 nt AtP. [Figure 15] Figure 15 shows a bar graph illustrating the minimization of the binding site of MG178-7193(-36) in human cells. The AttB site was tested from 72 to 52 nt, and AttP from 72 to 52 nt. The optimal condition was determined to be 52 nt AttB and 72 nt AttP. [Figure 16A] Figures 16A-C show the results of purification and activity analysis of MG178-7202. Protein expression induction and purification were monitored via SDS-PAGE (Figure 16A). The predicted protein MW was approximately -76 kDa. The concentrated fusion protein was run on an S200i 10 300 SEC column (Figure 16B). Eluted fractions were visualized via SDS-PAGE (Figure 16C), and fractions containing purified protein were collected and concentrated (Figure 16B). [Figure 16B]Figures 16A-C show the results of purification and activity analysis of MG178-7202. Protein expression induction and purification were monitored via SDS-PAGE (Figure 16A). The predicted protein MW was approximately -76 kDa. The concentrated fusion protein was run on an S200i 10 300 SEC column (Figure 16B). Eluted fractions were visualized via SDS-PAGE (Figure 16C), and fractions containing purified protein were collected and concentrated (Figure 16B). [Figure 16C] Figures 16A-C show the results of purification and activity analysis of MG178-7202. Protein expression induction and purification were monitored via SDS-PAGE (Figure 16A). The predicted protein MW was approximately -76 kDa. The concentrated fusion protein was run on an S200i 10 300 SEC column (Figure 16B). Eluted fractions were visualized via SDS-PAGE (Figure 16C), and fractions containing purified protein were collected and concentrated (Figure 16B). [Figure 17]Figure 17 shows the results of in vitro recombinase assays of LSR MG178-7202 and MG178-1859. The expected bands for MG178-7202 and MG178-1859 are 1027 bp and 1167 bp, respectively. Lane 1 shows a ladder of in vitro expressed proteins. Lane 2 shows a negative control of MG178-7202 containing the MG178-7202 attB and MG178-7202 attP dsDNA fragments. Lane 3 shows the experimental conditions for MG178-7202 containing the MG178-7202 attB and MG178-7202 attP dsDNA fragments and expressing the MG178-7202 recombinase. Lane 4 shows the negative control, MG178-1859, containing the MG178-1859 attB and MG178-1859 attP dsDNA fragments. Lane 5 shows the experimental conditions for MG178-1859, containing the MG178-1859 attB and MG178-1859 attP dsDNA fragments and expressing the MG178-1859 recombinase. Lane 6 shows a purified protein ladder. Lane 7 shows the negative control, MG178-7202, containing the MG178-7202 attB and MG178-7202 attP dsDNA fragments. Lane 8 shows the experimental conditions for MG178-7202, containing the MG178-7202 attB and MG178-7202 attP dsDNA fragments and purified MG178-7202 recombinase. Lane 9 shows the negative control of MG178-1859 containing the MG178-1859 attB and MG178-1859 attP dsDNA fragments, and lane 10 shows the experimental conditions of MG178-1859 containing the MG178-1859 attB and MG178-1859 attP dsDNA fragments and purified MG178-1859 recombinase.

[0026] Brief description of the sequence listing The Sequence Listing submitted with this application provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.

[0027] SEQ ID NOs: 1-16, 7151-7210, 7215-7218, 7220-7223, 7225-7230, 7233-7236, 7238-7241, 7243-7246, 7248-7251, 7253-7256, 7258-7261, 7263-7268, 7271-7274, 7277-7280, 7282-7285, 7287-7290, 7292-7295, 7297-7300, 7302-7305, 7307-7310, 7312-7315, 73 17-7320, 7322-7325, 7327-7330, 7332-7335, 7337-7340, 7342-7345, 7347-7350, 7352-7355, 7357-7360, 7362-7365, 7367-7370, 7372-7375, 7377-7380, 7382-7385, 7387-7390, 7392-7395, 7397-7400, and 7402-7405 indicate the nucleotide sequences of the MG178 recombinase binding sites.

[0028] SEQ ID NOs: 17-18, 7145, 7147, 7150, 7219, 7224, 7231, 7232, 7237, 7242, 7252, 7269, 7270, 7275, 7276, 7281, 7301, 7306, 7311, 7316, 7326, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, and 7401 show the nucleotide sequences of the MG178 conserved core.

[0029] SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214 set forth the amino acid sequences of MG178 family large serine recombinases suitable for use in gene editing as described herein.

[0030] SEQ ID NOs: 7412-7415 and 7418 show the amino acid sequences of the MG178 recombinase protein tags.

[0031] SEQ ID NOs: 7407 to 7411 and 7416 to 7417 show the nucleotide sequences of the primers. DETAILED DESCRIPTION OF THE INVENTION

[0032] Site-specific gene editing systems are powerful tools for site-specific genome engineering in cells. Most current gene editing systems rely on DNA double-strand breaks (DSBs) to direct cellular DNA repair pathways, such as homologous recombination (HR). However, these gene editing systems are often correlated with high indel rates, low insertion efficiency, high off-target activity, and limited cargo size.

[0033] Furthermore, repairing or inserting longer DNA fragments remains challenging, and safe and efficient methods for targeted integration of large templates into genomes, such as in gene therapy or engineered cell therapy, are lacking. To date, lentiviruses or adeno-associated viruses (AAVs) have been used in combination with CRISPR nucleases to insert large DNA fragments, such as entire genes. However, lentivirus-mediated integration lacks targeting capabilities because integration often occurs randomly in open chromatin. AAV-mediated delivery has limited cargo capacity and is not available for all cell types. A safe and efficient targeted genome editing system that enables large template integration is needed.

[0034] The present disclosure is based, in part, on the development of a gene editing system comprising a large serine recombinase (LSR) or serine recombinase for targetably and programmably integrating large fragments of DNA into the genome of a eukaryotic organism. In some embodiments, the serine recombinase described in this disclosure can integrate multi-kilobase DNA sequences.

[0035] definition Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter belongs. It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of any claimed subject matter. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0036] The practice of some of the methods disclosed in this disclosure employs, unless otherwise indicated, techniques in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M.A.usubel, et al. eds.); the series Methods in Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J.MacPherson, B.D.Hames, and G.R.Taylor eds. (1995)); Harlow and Lane, eds. (1988); Antibodies, A Laboratory Manual; and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0037] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof, are used in either the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."

[0038] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0039] As used herein, the term "nucleotide" refers to a base-sugar-phosphate combination. Contemplated nucleotides include naturally occurring and synthetic nucleotides. Nucleotides are monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide encompasses dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of ddNTP include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP.Nucleotide can be unlabeled, or can be detectably labeled, for example, by using an optically detectable moiety (e.g., fluorophore) or a moiety containing quantum dot.Detectable labels include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [TAMRA]dUTP, and [dROX]ddTTP available from Perkin Elmer (Foster City, CA); FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham (Arlington Heights, IL); and FluoroLink Cy5-dUTP available from Boehringer Ingelheim. Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2′-dATP available from Mannheim (Indianapolis, IN), and Molecular Examples of chromosomal labeled nucleotides available from Probes (Eugene, Oregon) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0040] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably to refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in either single-, double-, or multiple-stranded form. Contemplated polynucleotides include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, multiple loci (single loci) defined from binding analyses, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. When referring to T, in polynucleotides, T refers to U (uracil) in RNA and T (thymine) in DNA. Polynucleotides can be exogenous or endogenous to a cell and / or can exist in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure are imparted before or after assembly of the polymer. Non-limiting examples of modifications include 5-bromouracil, peptide nucleic acids, heterologous nucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. The sequence of nucleotides can be interrupted by non-nucleotide components.

[0041] The terms "transfection" or "transfected" generally refer to the introduction of nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a gene sequence encoding an entire protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0042] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein to refer to a polymer of at least two amino acid residues joined by a peptide bond. The term does not denote a specific length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified by, for example, disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and any other manipulation, such as conjugation with a labeling component. As used in this disclosure, the terms "amino acid" and "amino acids" refer to natural and unnatural amino acids, including, but not limited to, modified amino acids. Modified amino acids include amino acids that have been chemically modified to include a group or chemical moiety that does not occur naturally on the amino acid. The term "amino acid" includes both D- and L-amino acids.

[0043] As used in this disclosure, "non-naturally occurring" refers to a nucleic acid or polypeptide sequence that does not occur in nature. Non-naturally occurring refers to a nucleic acid or polypeptide sequence that does not occur in nature, including modifications such as mutations, insertions, or deletions. The term non-naturally occurring encompasses fusion nucleic acids or polypeptides in which the non-naturally occurring sequence encodes or exhibits an activity (e.g., an enzymatic activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) of the nucleic acid or polypeptide sequence to which it is fused. Non-naturally occurring nucleic acid or polypeptide sequences include those that are linked by genetic engineering to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid or polypeptide sequence that encodes the chimeric nucleic acid or polypeptide.

[0044] As used in this disclosure, the term "promoter" refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and may be located adjacent to or overlapping the nucleotide or region of nucleotide at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often referred to as transcription factors, that facilitate the binding of RNA polymerase to DNA, resulting in gene transcription. Basal promoters in eukaryotes typically, but not necessarily, contain a TATA-box and / or a CAAT box.

[0045] The term "expression," as used in this disclosure, refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. Collectively, the transcript and the encoded polypeptide may be referred to as a "gene product." If the polynucleotide is derived from genomic DNA, the term expression includes splicing of mRNA in a eukaryotic cell.

[0046] As used in this disclosure, "operably linked," "operable linkage," "operatively linked," or their grammatical equivalents refer to the arrangement of genetic elements, e.g., promoters, enhancers, polyadenylation sequences, etc., such that operation (e.g., movement or activation) of a first genetic element has some effect on a second genetic element. The effect on the second genetic element can be, but need not be, of the same type as operation of the first genetic element. For example, two genetic elements are operably linked if movement of the first element causes activation of the second element. A regulatory element, which may include, for example, a promoter sequence and / or an enhancer sequence, is operably linked to a coding region if the regulatory element helps to initiate transcription of the coding sequence. There can be intervening residues between the regulatory element and the coding region so long as this functional relationship is maintained.

[0047] As used in this disclosure, "vector" refers to a polymer or association of polymers that contains or is associated with a polynucleotide and mediates delivery of the polynucleotide to a cell. Examples of vectors include nucleic acid-based vectors (e.g., plasmids and viral vectors) and liposomes. Exemplary nucleic acid-based vectors generally include genetic elements, e.g., regulatory elements, operably linked to a gene to promote expression of the gene in the target.

[0048] As used in this disclosure, "expression cassette" and "nucleic acid cassette" are used interchangeably to refer to a component of a vector that contains a combination of nucleic acid sequences or elements (e.g., therapeutic gene, promoter, and terminator) that are expressed together or operably linked for expression. The term encompasses expression cassettes that contain a combination of one or more genes with regulatory elements operably linked for expression.

[0049] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes the ability to affect expression in a manner attributable to the full-length sequence.

[0050] The terms "engineered," "synthetic," and "artificial" are used interchangeably herein to refer to entities modified by human intervention. For example, the terms refer to polynucleotides or polypeptides that do not occur in nature. Engineered peptides may, but need not, have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains. Non-limiting examples include: nucleic acids that are modified by changing their sequence to one that does not occur in nature; nucleic acids that are modified by ligating to a nucleic acid with which they are not naturally associated such that the ligated product possesses a function not present in the original nucleic acid; engineered nucleic acids that are synthesized in vitro using a sequence that does not occur in nature; proteins that are modified by changing their amino acid sequence to a sequence that does not occur in nature; and engineered proteins that acquire new functions or properties. An "engineered" system contains at least one engineered component.

[0051] As used in this disclosure, "guide nucleic acid" or "guide polynucleotide" refers to a nucleic acid that can hybridize to a target nucleic acid, thereby directing an associated nuclease to the target nucleic acid. A guide nucleic acid can be, but is not limited to, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. A guide nucleic acid can include crRNA or tracrRNA, or a combination of both. The term guide nucleic acid encompasses engineered guide nucleic acids and programmable guide nucleic acids that specifically bind to a target nucleic acid. A portion of a target nucleic acid can be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to a guide nucleic acid and hybridizes with the guide nucleic acid is the complementary strand. A strand of a double-stranded target polynucleotide that is complementary to the complementary strand and therefore not complementary to the guide nucleic acid is referred to as the non-complementary strand. A guide nucleic acid having a polynucleotide strand is a "single guide nucleic acid." A guide nucleic acid having two polynucleotide strands is a "dual guide nucleic acid." Unless otherwise specified, the term "guide nucleic acid" is inclusive and refers to both single and double guide nucleic acids. A guide nucleic acid may include a segment referred to as a "nucleic acid targeting segment" or "nucleic acid targeting sequence" or "spacer." A nucleic acid targeting segment may include a subsegment referred to as a "protein binding segment" or "protein binding sequence" or "Cas protein binding segment."

[0052] The term "tracrRNA" or "tracr sequence" refers to a transactivating CRISPR RNA. The tracrRNA interacts with the CRISPR (cr) RNA to form a guide nucleic acid (e.g., guide RNA or gRNA) that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid.

[0053] As used in this disclosure, the term "RuvC_III domain" refers to the third, non-contiguous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three non-contiguous segments, RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).

[0054] As used in this disclosure, the term "HNH domain" refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF01844 for domain HNH).

[0055] As used in this disclosure, the term "transposon" refers to mobile elements that move in and out of the genome and carry with them "cargo DNA." These transposons can differ by the type of nucleic acid they transfer, the type of repeats at the ends of the transposon, the type of cargo they transfer, or the mode of transfer (i.e., self-repair or host-repair).

[0056] As used in this disclosure, the term "transposase" or "transposases" refers to an enzyme that binds to the ends of a transposon and catalyzes its movement to another part of the genome. Types of movement include cut-and-paste and copy-and-transpose mechanisms.

[0057] As used herein, the term "Tn7" or "Tn7-like transposase" refers to a family of transposases that includes three major components: a heteromeric transposase (TnsA and / or TnsB) along with a regulatory protein (TnsC). In addition to the TnsABC transposition proteins, Tn7 elements can encode dedicated target site selection proteins, TnsD and TnsE. In conjunction with TnsABC, the sequence-specific DNA-binding protein TnsD directs transposition into a conserved site termed the "Tn7 binding site," i.e., attTn7. TnsD is a member of a large family of proteins that also includes TniQ. TniQ has been shown to target transposition into degradation sites on plasmids.

[0058] As used in this disclosure, the terms "gene editing" and "genome editing" can be used interchangeably. Gene editing or genome editing refers to changing the nucleic acid sequence of a gene or genome. Genome editing can include, for example, insertions, deletions, and mutations. Genome editing can be performed by a gene editing system, for example, a nuclease, a reverse transcriptase, a recombinase, or a base editor.

[0059] As used in this disclosure, the term "recombinase" refers to an enzyme that mediates the recombination of DNA fragments located between recombinase recognition sequences, resulting in excision, insertion, inversion, exchange, or transposition) of the DNA fragments located between the recombinase recognition sequences.

[0060] As used herein, the terms "recombining" or "recombination" in the context of nucleic acid modification (e.g., genomic modification) refer to a process in which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can result in, among other things, the insertion, inversion, excision, or transposition of a nucleic acid sequence within or between one or more nucleic acid molecules.

[0061] As used in this disclosure, the term "complex" refers to the joining of at least two components. Each of the two components may retain the properties / activities it had prior to forming the complex or may acquire properties as a result of forming the complex. Joining may be by, but is not limited to, covalent bonding, non-covalent bonding (i.e., hydrogen bonding, ionic interactions, van der Waals interactions, and hydrophobic bonding), use of a linker, fusion, or any other suitable method. Contemplated components of the complex include polynucleotides, polypeptides, or combinations thereof. For example, the complex may include an endonuclease and a guide polynucleotide.

[0062] The term "contig" refers to a set of DNA segments or sequences that overlap in a way that provides a contiguous representation of a genomic region.

[0063] The term "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a specified percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and presence of 11, gap cost at extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of word length (W) of 2, expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE using default parameters; MAFFT using parameters of retries of 2 and maximum repeats of 1,000; Novafold using default parameters; and HMMER hmmalign using default parameters.

[0064] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences aligned for maximum amino acid residue or nucleotide correspondence, e.g., as determined by the alignment producing the highest or "optimized" percent identity score.

[0065] The present disclosure also encompasses variants of any of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions may be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length for each other. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by identifying amino acid residues (e.g., non-conserved residues) that vary between species without altering the basic function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of the large serine recombinase protein sequences described in this disclosure (e.g., MG178 family serine recombinases, or any other family of large serine recombinases described in this disclosure). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with non-abolishing substitutions of one or more key active site residues.

[0066] The present disclosure also includes variants (e.g., reduced activity variants) of any of the enzymes described herein that have substitutions of one or more catalytic residues to reduce or eliminate the activity of the enzyme. In some embodiments, reduced activity variants of the proteins described herein include at least one, at least two, or all three catalytic residues.

[0067] Conservative substitution tables providing functionally similar amino acids are available in various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993))). The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G); 2) aspartic acid (D), glutamic acid (E); 3) asparagine (N), glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), valine (V); 6) phenylalanine (F), tyrosine (Y), tryptophan (W); 7) serine (S), threonine (T); and 8) Cysteine ​​(C), methionine (M).

[0068] Serine recombinase gene editing system Current gene editing systems lack the ability to integrate multi-kilobase nucleic acid sequences. Some of these gene editing systems rely primarily on nuclease-specific DNA double-strand breaks (DSBs) to induce cellular DNA repair pathways such as homologous recombination (HR). Despite significant advances in optimizing HR for specific situations, these approaches generally suffer from low insertion efficiency, high indel rates, and cargo size limitations, with limited success seen for cargoes larger than 1 kilobase (kb).

[0069] Large serine recombinases (LSRs) can integrate large fragments of DNA into eukaryotic genomes in a nonrandom, site-specific manner. Viral LSRs range in length from 400 to 700 amino acids and facilitate the integration of the phage genome into the bacterial host genome once the virus enters its lysogenic life cycle. The mechanism of phage integration involves the LSR recognizing specific binding sites, attB sites, in the host genome and phage binding sites, attP sites, on the phage genome. Viral genome integration occurs via recombination at these binding sites; this process results in the generation of two new binding sites, attL and attR, flanking the prophage.

[0070] By recognizing these binding sites (i.e., recognition sequences found on DNA donor and acceptor molecules), recombinases can catalyze targeted cleavage, strand exchange, and DNA rejoining. This mechanism allows for site-specific DNA insertion without the need for cellular cofactors and without generating exposed double-strand breaks. However, some LSRs have limited DNA integration efficiency. Therefore, improved LSRs are needed.

[0071] The serine recombinases described in this disclosure provide for genome engineering due to their ability to integrate desired cargoes into specific target sites.

[0072] Serine recombinase Described herein is a gene editing system comprising a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142, or a nucleic acid encoding the serine recombinase. Further described herein are nucleic acids, vectors, and cells comprising the serine recombinase described herein. Further described herein are means for integrating nucleic acid sequences into a genome.

[0073] Serine recombinases are enzymes that catalyze site-specific recombination events by promoting DNA strand exchange between two DNA segments containing cognate recombinase recognition sites. Examples of serine recombinases include the small serine recombinases gamma-delta resolvus (derived from the Tn1000 transposon) and Tn3 resolvus (derived from the Tn3 transposon), and the large serine recombinases (LSRs) φC31-integrase (derived from the φC31 phage), Bxb1-integrase (derived from mycobacteriophage), and R4 integrase. Serine recombinases are characterized by a conserved catalytic serine amino acid residue that attacks DNA phosphodiesters and covalently binds to the DNA strand terminus during catalysis. Serine recombinases recognize cognate binding site sequences, designated attB on the acceptor DNA strand (e.g., bacterial genome) and attP on the donor DNA strand (e.g., phage genome). After the recombination event, the attB and attP sites recombine to form attL and attR sites flanking the newly integrated sequence. The attB and attP sites are typically up to approximately 50 bases in length. During the recombination event, serine recombinase forms a tetrameric complex, with each protein dimer binding to the attB or attP binding site. The serine recombinase cleaves each strand, creating a double-stranded break and leaving a 2-bp overhang, followed by strand exchange and strand ligation. Typically, serine recombinase does not require other enzymes to perform the reaction.

[0074] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs:21-7060 and 7105-7142, or a variant thereof. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to any one of SEQ ID NOs:21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to any one of SEQ ID NOs:21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to any one of SEQ ID NOs:21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to any one of SEQ ID NOs:21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to any one of SEQ ID NOs:21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to any one of SEQ ID NOs:21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to any one of SEQ ID NOs:21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 21-7060 and 7105-7142.In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 21-7060 and 7105-7142. In some embodiments, the serine recombinase comprises a sequence having 100% identity to any one of SEQ ID NOs: 21-7060 and 7105-7142.

[0075] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:21. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:21.

[0076] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:22. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:22.

[0077] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:23. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:23.

[0078] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:24. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:24.

[0079] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:7140. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:7140.

[0080] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:7131. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:7131.

[0081] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:7115. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:7115.

[0082] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NO: 7139. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO: 7139. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO: 7139. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO: 7139. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:7139. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:7139. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:7139. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:7139. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:7139. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:7139. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:7139. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:7139.

[0083] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO: 1848. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO: 1848.

[0084] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:7111. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:7111.

[0085] In some embodiments, the serine recombinase comprises a sequence having at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 70% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 75% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 80% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 85% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 90% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 95% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 96% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 97% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 98% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having at least about 99% identity to SEQ ID NO:7136. In some embodiments, the serine recombinase comprises a sequence having 100% sequence identity to SEQ ID NO:7136.

[0086] Further described herein is a eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142. Further described herein is a eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO: 21. Further described herein is a eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO: 22. Further described herein is a eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO: 23. Further described herein is a eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO: 24.

[0087] In some embodiments, the eukaryotic cell is a mammalian cell, hi some embodiments, the eukaryotic cell is a human cell.

[0088] In some embodiments, the serine recombinases described herein have improved transduction efficiency. In some embodiments, the serine recombinases described herein comprise a transduction efficiency of at least about 5%. In some embodiments, the serine recombinases described herein comprise a transduction efficiency of at least about 25%. In some embodiments, the serine recombinases described herein comprise a transduction efficiency of at least about 50%. In some embodiments, the serine recombinases described herein comprise a transduction efficiency of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or greater than 95%. In some embodiments, the serine recombinase described herein has improved transduction efficiency compared to a serine recombinase selected from the group consisting of β-six, CinH, ParA γδ, Bxb1, φC31, TP901, TG1, φBT1, R4, φRV1, φFC1, MR11, A118, U153, and gp29.

[0089] In some embodiments, the serine recombinase is a viral, prokaryotic, or eukaryotic serine recombinase. In some embodiments, the serine recombinase can target a gene containing a catalase domain or a synthase domain. In some embodiments, the catalase is manganese catalase. In some embodiments, the synthase is queosine synthase. In some embodiments, the serine recombinase can target a gene containing a DUF4244 Pfam domain.

[0090] In some embodiments, the serine recombinase described herein comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the serine recombinase. In some embodiments, the NLS comprises any of the sequences in Table 1 below, or a combination thereof. [Table 1]

[0091] In some embodiments, the serine recombinase comprises a tag. In some embodiments, the tag is an affinity tag. Exemplary affinity tags include, but are not limited to, a His tag, a Flag tag, a Myc tag, a MBP tag, and a GST tag.

[0092] In some embodiments, the serine recombinase comprises a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, a TEV site, a C3 site, a factor Xa site, and an enterokinase site. Recombination site

[0093] This disclosure describes a gene editing system comprising a serine recombinase having at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142, or a nucleic acid encoding the serine recombinase, and a donor polynucleotide and a nucleic acid comprising a first binding site sequence.

[0094] In some embodiments, the first binding site sequence is 5' to the donor polynucleotide.

[0095] In some embodiments, the nucleic acid encoding the serine recombinase further comprises a second binding site sequence. In some embodiments, the second binding site sequence is 5' to the serine recombinase. In some embodiments, the nucleic acid encoding the serine recombinase comprises one or more binding site sequences. In some embodiments, the nucleic acid encoding the serine recombinase comprises 1, 2, 3, 4, 5, or more than 5 binding site sequences.

[0096] In some embodiments, the nucleic acid comprising the donor polynucleotide comprises one or more binding site sequences, hi some embodiments, the nucleic acid comprising the donor polynucleotide comprises 1, 2, 3, 4, 5, or more than 5 binding site sequences.

[0097] In some embodiments, the first binding site sequence and said second binding site sequence can be recombined.

[0098] In some embodiments, the first binding site sequence is a bacterial genome recombination sequence (attB). In some embodiments, the attB sequence comprises from about 20 nucleotides to about 500 nucleotides. In some embodiments, the attB sequence comprises from about 20 to about 450, from about 20 to about 400, from about 20 to about 350, from about 20 to about 300, from about 20 to about 250, from about 20 to about 200, from about 20 to about 250, from about 20 to about 100, from about 20 to about 50, from about 50 to about 450, from about 50 to about 400, or from about 50 to about 350. about 50 to about 300, about 50 to about 250, about 50 to about 200, about 50 to about 150, about 50 to about 100, about 100 to about 450, about 100 to about 400, about 100 to about 350, about 100 to about 300, about 100 to about 250, about 100 to about 200, or about 100 to about 150 nucleotides.

[0099] In some embodiments, the first binding site sequence is a phage genome recombination sequence (attP). In some embodiments, the attP sequence comprises about 20 to about 450, about 20 to about 400, about 20 to about 350, about 20 to about 300, about 20 to about 250, about 20 to about 200, about 20 to about 250, about 20 to about 100, about 20 to about 50, about 50 to about 450, about 50 to about 400, or about 50 to about 350 about 50 to about 300, about 50 to about 250, about 50 to about 200, about 50 to about 150, about 50 to about 100, about 100 to about 450, about 100 to about 400, about 100 to about 350, about 100 to about 300, about 100 to about 250, about 100 to about 200, or about 100 to about 150 nucleotides.

[0100] In some embodiments, the second binding site sequence is a bacterial genome recombination sequence (attB). In some embodiments, the attB sequence comprises from about 20 nucleotides to about 500 nucleotides. In some embodiments, the attB sequence comprises from about 20 to about 450, from about 20 to about 400, from about 20 to about 350, from about 20 to about 300, from about 20 to about 250, from about 20 to about 200, from about 20 to about 250, from about 20 to about 100, from about 20 to about 50, from about 50 to about 450, from about 50 to about 400, or from about 50 to about 350. about 50 to about 300, about 50 to about 250, about 50 to about 200, about 50 to about 150, about 50 to about 100, about 100 to about 450, about 100 to about 400, about 100 to about 350, about 100 to about 300, about 100 to about 250, about 100 to about 200, or about 100 to about 150 nucleotides.

[0101] In some embodiments, the second binding site sequence is a phage genome recombination sequence (attP). In some embodiments, the attP sequence comprises about 20 to about 450, about 20 to about 400, about 20 to about 350, about 20 to about 300, about 20 to about 250, about 20 to about 200, about 20 to about 250, about 20 to about 100, about 20 to about 50, about 50 to about 450, about 50 to about 400, or about 50 to about 350 about 50 to about 300, about 50 to about 250, about 50 to about 200, about 50 to about 150, about 50 to about 100, about 100 to about 450, about 100 to about 400, about 100 to about 350, about 100 to about 300, about 100 to about 250, about 100 to about 200, or about 100 to about 150 nucleotides.

[0102] In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7277, 7282, 7287, 7288-7290, 7291-7292, 7293, 7294, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7330, 7331, 7332, 7333, 7334, 7335, 7336, 7337, 7338, 7340, 7341, 7342, 7343, 7344, 734 72, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402. In some embodiments, the attB sequence comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7277, 7282, 7287, 7397, and 7402. In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397 and 7402.In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397 and 7402. In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397 and 7402.In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397 and 7402. In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397 and 7402.In some embodiments, the attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7410, 7411, 7412, 7413, 7 277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397 and 7402. In some embodiments, the attB sequence comprises any one of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7397, and 7402.

[0103] In some embodiments, the attB sequence comprises a sequence having at least 70% (e.g., 75%, 80%, 90%, 95%, 97%, 98%, or 99%) sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13. In some embodiments, the attB sequence comprises any one of SEQ ID NOs: 1, 5, 9, and 13.

[0104] In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7279, 7284, 7289, 7290, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7410, 7411, 7412, 7413, 7414, 7415, 7416, 7417, 7 7399, 7404, 7410, 7420, 7430, 7449, 7450, 7460, 7470, 7471, 7480, 7481, 7492, 7493, 7494, 7495, 7496, 7497, 7498, 7499, 7504, 7510, 7511, 7512, 7513, 7514, 7515, 7516, 7517, 7518, 7519, 7520, 7521, 7522, 7523, 7524, 7525, 7526, 7527, 7528, 7529, 7530, 7531, 7532, 7533, 7534, 7540, 7541, 7542, 7543, 7544, 7545, 7546, 7547, 7548, 7549, 7550, 7551, 7552, 7553, 75544, 7555, 7556, 7557, 7558, 7559, 7560, 7561, 7562, 7563, 75644, 7565, 7566, 7567, 7568, 7569, 7570, 7571, 7572, 7573, 75744, 7575, 7576, 7577, 757 In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7278, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7410, 7 279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399 and 7404.In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7278, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7410, 7 279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399 and 7404. In some embodiments, the attP sequence comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7399, and 7404. In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7278, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7410, 7 279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399 and 7404.In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7278, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7410, 7 279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399 and 7404. In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7278, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7410, 7 279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399 and 7404.In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7278, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7410, 7 279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399 and 7404. In some embodiments, the attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183-7187, 7201-7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7261, 7262, 7263, 7264, 7265, 7266, 7267, 7268, 7269, 7270, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7 6, 7267, 7273, 7279, 7284, 7289, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399 and 7404.

[0105] In some embodiments, the attP sequence comprises a sequence having at least 70% (e.g., 75%, 80%, 90%, 95%, 97%, 98%, or 99%) sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the attP sequence comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the ORF comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the attP sequence comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the attP sequence comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the attP sequence comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the attP sequence comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the attP sequence comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14. In some embodiments, the attP sequence comprises any one of SEQ ID NOs: 2, 6, 10, and 14.

[0106] In some embodiments, the nucleic acid comprising the donor polynucleotide and the first binding site sequence is delivered by a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid.

[0107] In the present disclosure, there is provided a eukaryotic genome comprising a donor polynucleotide sequence and an attL sequence 5' to the donor polynucleotide sequence, wherein the attL sequence is selected from the group consisting of SEQ ID NOs: 17 to 18, 7145, 7147, 7150, GCATCCCC, TATTCGAT, GGGCAACC, GGGCACCC, CAAGTTC, ACCGCC, CATATGT, 7219, 7224, 7231, 7232, 7237, 7242, ATGGTGGGC, 7252, GCCATTTC,

[0010] Described herein are eukaryotic genomes comprising a sequence selected from the group consisting of TCAGCTCCA, 7269, 7270, 7275, 7276, 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC. In some embodiments, the eukaryotic genome further comprises an attR sequence 3' to the donor polynucleotide sequence.

[0108] In the present disclosure, there is provided a eukaryotic genome comprising a donor polynucleotide sequence and an attL sequence 3' to the donor polynucleotide sequence, wherein the attL sequence is selected from the group consisting of SEQ ID NOs: 17 to 18, 7145, 7147, 7150, GCATCCCC, TATTCGAT, GGGCAACC, GGGCACCC, CAAGTTC, ACCGCC, CATATGT, 7219, 7224, 7231, 7232, 7237, 7242, ATGGTGGGC, 7252, GCCATTTC,

[0010] Described herein are eukaryotic genomes comprising a sequence selected from the group consisting of TCAGCTCCA, 7269, 7270, 7275, 7276, 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC. In some embodiments, the eukaryotic genome further comprises an attR sequence 3' to the donor polynucleotide sequence.

[0109] In the present disclosure, donor polynucleotide sequences and attL sequences 5' or 3' to the donor polynucleotide sequence, which are SEQ ID NOs: 17 to 18, 7145, 7147, 7150, GCATCCCC, TATTCGAT, GGGCAACC, GGGCACCC, CAAGTTC, ACCGCC, CATATGT, 7219, 7224, 7231, 7232, 7237, 7242, ATGGTGGGC, 7252, GCCATTTC, T a sequence selected from the group consisting of CAGCTCCA, 7269, 7270, 7275, 7276, 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC A ttL sequence and an attR sequence on the 5' or 3' side of the donor polynucleotide sequence, which are SEQ ID NOs: 17 to 18, 7145, 7147, 7150, GCATCCCC, TATTCGAT, GGGCAACC, GGGCACCC, CAAGTTC, ACCGCC, CATATGT, 7219, 7224, 7231, 7232, 7237, 7242, ATGGTGGGC, 7252, GCCATTTC, TCAGCTCCA, 7269, 7270 , 7275, 7276, 7281, GGGTC, TTCATGAG, ATGGTGGGC, 7301, 7306, 7311, 7316, GGGATCCC, 7326, GCCGA, 7336, 7341, 7346, 7351, 7356, 7361, 7366, 7371, 7376, 7381, 7386, 7391, AGGCGG, 7401, and GGATGC.

[0110] The conserved core of exemplary recombination sites is shown in Table 2. [Table 2]

[0111] In some embodiments, the attL sequence and the attR sequence are identical.

[0112] In some embodiments, the attL sequence is a recombination of the first and second binding site sequences. In some embodiments, the attR sequence is a recombination of the first and second binding site sequences.

[0113] Donor polynucleotide The serine recombinases described herein can provide for the integration of large polynucleotides (e.g., donor polynucleotides). In some embodiments, the donor polynucleotide has a size of at least about 1 kilobase (kb), 2 kb, 50 kb, or greater than 50 kb. In some embodiments, the donor polynucleotide has a size of at least about 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 50 kb, 100 kb, 200 kb, 300 kb, 400 kb, or 500 kb. In some embodiments, the donor polynucleotide has a size of about 200 base pairs (bp) to about 500 kb, 200 bp to about 250 kb, or 200 bp to about 100 kb. In some embodiments, the donor polynucleotide has a size of about 1 kb to about 10 kb, about 1 to about 7.5 kb, about 1 to about 5 kb, about 1 to about 3 kb, about 2 to about 10 kb, about 2 to about 7.5 kb, about 2 to about 5 kb, about 2 to about 3 kb, about 3 to about 10 kb, about 3 to about 7.5 kb, or about 3 to about 5 kb. In some embodiments, the donor polynucleotide has a size of about 10 kb to about 500 kb, 10 kb to about 400 kb, 10 kb to about 300 kb, 10 kb to about 200 kb, 10 kb to about 100 kb, about 10 kb to about 75 kb, about 10 kb to about 50 kb, about 10 kb to about 30 kb, about 20 kb to about 100 kb, about 20 to about 75 kb, about 20 kb to about 50 kb, about 20 kb to about 30 kb, about 30 kb to about 100 kb, about 30 kb to about 75 kb, or about 30 kb to about 50 kb. In some embodiments, the donor polynucleotide has a size of about 10 to about 500, 20 to about 400, 10 to about 300, 10 to about 200, or 10 to about 100. In some embodiments, the donor polynucleotide is circular. In some embodiments, the donor polynucleotide is linear.

[0114] In some embodiments, the donor polynucleotide encodes a therapeutic agent, a reporter, or a marker.

[0115] In some embodiments, the reporter comprises a fluorescent protein, ie, GFP, EBFP, EBFP2, azurite, mKalamal, ECFP, cerulean, CyPet, YFP, citrine, Venus, YPet, RFP, CFP, or a derivative thereof.

[0116] In some embodiments, the reporter is acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), beta-galactosidase (LacZ), beta-glucuronidase (GUS), chloramphenicol acetyltransferase (CAT), horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, or a derivative thereof.

[0117] In some embodiments, the marker is an antibiotic resistance marker, hi some embodiments, the antibiotic resistance marker is kanamycin, spectinomycin, streptomycin, ampicillin, carbenicillin, bleomycin, erythromycin, polymyxin B, tetracycline, chloramphenicol, neomycin, zeocin, or a derivative thereof.

[0118] In some embodiments, the marker is a cell surface marker. In some embodiments, the cell surface marker is a membrane protein, a sugar moiety, or a small molecule (e.g., biotin) displayed on the cell surface. In some embodiments, the cell surface marker is CD3, B2M, CD4, CD8, CD28, an HLA protein, an MHC complex, streptavidin, or avidin. In some embodiments, the cell surface marker is an antibody, such as, for example, an IgG, or an antibody fragment, such as, for example, an scFv or Fc. In some embodiments, the cell surface marker can be bound by a specific antibody. In some embodiments, cells are analyzed for expression of the cell surface marker by flow cytometry.

[0119] Delivery and Vectors This disclosure, in some embodiments, describes nucleic acid sequences encoding the serine recombinases or serine recombinase gene editing systems disclosed in this disclosure.

[0120] In some embodiments, the nucleic acid encoding the serine recombinase or the serine recombinase gene editing system is DNA, e.g., linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid is RNA, e.g., mRNA.

[0121] The present disclosure describes a vector comprising: a) a nucleic acid encoding a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142; and b) one or more regulatory elements. The present disclosure also describes a vector comprising a nucleic acid encoding a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142, wherein the vector is selected from the group consisting of a plasmid, a nanoplasmid, a phagemid, a phage derivative, a bacmid, a bacterial artificial chromosome (BAC), a minicircle, a doggybone, a yeast artificial chromosome (YAC), and a cosmid.

[0122] In some embodiments, the nucleic acid encoding the serine recombinase or the serine recombinase gene editing system is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., a circular DNA molecule that can replicate autonomously inside a cell), a cosmid (e.g., a pWE or sCos vector), an artificial chromosome, a human artificial chromosome (HAC), a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), a P1-derived artificial chromosome (PAC), a phagemid, a phage derivative, a bacmid, or a virus. In some embodiments, the nucleic acid-based vector is selected from the group consisting of pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, pRI 101-AN. DNA, pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.

[0123] In some embodiments, the one or more regulatory elements include a promoter, an enhancer, an intron, a microRNA, a linker, a splicing element, or a polyA signal. In some embodiments, the promoter is selected from the group consisting of a constitutive promoter, an inducible promoter, a minipromoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, polH, EM7, OpIE1, and derivatives thereof. In some embodiments, the promoter is a U6 promoter. In some embodiments, the promoter is a CAG promoter.

[0124] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the virus is a retrovirus.

[0125] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh 10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, In some embodiments, the herpesvirus is AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or a derivative thereof. In some embodiments, the herpesvirus is HSV-1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0126] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof.In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.

[0127] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.

[0128] In some embodiments, the nucleic acid encoding the serine recombinase or the serine recombinase gene editing system is delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is associated with a lipid. The lipid-associated nucleic acid is, in some embodiments, encapsulated within the aqueous interior of the liposome, interspersed within the lipid bilayer of the liposome, attached to the liposome via a linking molecule associated with both the liposome and the nucleic acid, entrapped in the liposome, complexed with the liposome, dispersed in a solution containing a lipid, mixed with the lipid, combined with the lipid, contained as a suspension in the lipid, contained in or complexed with the micelle, or otherwise associated with the lipid. In some embodiments, the nucleic acid is contained in a lipid nanoparticle (LNP).

[0129] In some embodiments, the nucleic acid encoding the serine recombinase or the serine recombinase gene editing system is introduced into the cell by any suitable method, either stably or transiently. In some embodiments, the serine recombinase or serine recombinase gene editing system is transfected into the cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct encoding the serine recombinase or serine recombinase gene editing system. For example, the cell is transduced (e.g., with a virus encoding the serine recombinase or serine recombinase gene editing system), or transduced with a nucleic acid encoding the serine recombinase or serine recombinase gene editing system (e.g., with a plasmid encoding the serine recombinase or serine recombinase gene editing system), or transduced with a translated serine recombinase or serine recombinase gene editing system. In some embodiments, the transduction is stable or transient transduction. In some embodiments, a plasmid expressing a serine recombinase or a serine recombinase gene editing system is introduced into cells via electroporation, transient (e.g., lipofection) and stable genome integration (e.g., piggybac), and viral transduction (e.g., lentivirus or AAV), or other methods known to those of skill in the art. In some embodiments, the gene editing system is introduced into cells as one or more polypeptides. In some embodiments, delivery is achieved through the use of RNP complexes. Methods for delivering polypeptides and / or RNPs into cells, for example, by electroporation or by cell squeezing, are known in the art.

[0130] Exemplary methods for nucleic acid delivery include lipofection, nucleofection, electroporation, stable genome integration (e.g., piggyback), microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycations or lipid-nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™, Lipofectin™, and SF Cell Line 4D-Nucleofector X Kit™ (Lonza)). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include the lipids of WO91 / 17424 and WO91 / 16024. In some embodiments, delivery is to a cell (e.g., in vitro or ex vivo administration) or a target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in a liposome or nanoparticle that specifically targets the host cell.

[0131] Other methods for delivery of nucleic acids into cells are known to those of skill in the art, see, e.g., US2003 / 0087817.

[0132] In some embodiments, delivering a serine recombinase or serine recombinase gene editing system to a target nucleic acid site comprises delivering a nucleic acid comprising an open reading frame encoding the serine recombinase or serine recombinase gene editing system. In some embodiments, the nucleic acid comprises a promoter. In some embodiments, the open reading frame encoding the serine recombinase or serine recombinase gene editing system is operably linked to the promoter. In some embodiments, the promoter is a ribonucleic acid (RNA) pol III promoter.

[0133] In some embodiments, delivering a serine recombinase or serine recombinase gene editing system to the target nucleic acid site comprises delivering a capped mRNA containing an open reading frame encoding the serine recombinase or serine recombinase gene editing system. In some embodiments, delivering a serine recombinase or serine recombinase gene editing system to the target nucleic acid site comprises delivering a translated polypeptide. In some embodiments, delivering a serine recombinase or serine recombinase gene editing system to the target nucleic acid site comprises delivering deoxyribonucleic acid (DNA) encoding the serine recombinase or serine recombinase gene editing system operably linked to a ribonucleic acid (RNA) pol III promoter.

[0134] lipid nanoparticles The present disclosure, in certain embodiments, discloses lipid nanoparticles comprising the serine recombinase or serine recombinase gene editing system of the present disclosure for delivery into cells.

[0135] In some embodiments, the lipid nanoparticles comprise a serine recombinase or a serine recombinase gene editing system, or a nucleic acid encoding a serine recombinase or a serine recombinase gene editing system. In some embodiments, the lipid nanoparticles comprise one or more components of the serine recombinase gene editing system. In some embodiments, the lipid nanoparticles comprise the serine recombinase or a nucleic acid comprising the serine recombinase. In some embodiments, the lipid nanoparticles comprise an engineered donor polynucleotide.

[0136] In some embodiments, the lipid nanoparticle is tethered to said serine recombinase gene editing system.

[0137] The lipid nanoparticles described herein may be four-component lipid nanoparticles. Such nanoparticles may be configured for delivery of RNA or other nucleic acids (e.g., synthetic RNA, mRNA, or in vitro synthesized mRNA) and may generally be formulated as described in WO2012 / 135805(A2). Such nanoparticles may generally comprise (a) a cationic lipid, (b) a neutral lipid (e.g., DSPC or DOPE), (c) a sterol (e.g., cholesterol or a cholesterol analog), or (d) a PEG-modified lipid (e.g., PEG-DMG).

[0138] The cationic lipid referred to in this disclosure as "C12-200" is disclosed in Love et al., Proc Natl Acad Sci USA. 2010 107:1864-1869 and Liu and Huang, Molecular Therapy. 2010 669-670. Cationic lipid formulations may contain particles containing any of three or more components in addition to polynucleotides, primary constructs, or RNA (e.g., mRNA). By way of example, a formulation containing a particular cationic lipid, including but not limited to 98N12-5, may contain 42% lipidoid, 48% cholesterol, and 10% PEG (alkyl chain length of C14 or greater). As another example, a formulation with a particular lipidoid, including but not limited to C12-200, may contain 50% cationic lipid, 10% disteroylphosphatidylcholine, 38.5% cholesterol, and 1.5% PEG-DMG.

[0139] In some embodiments, the cationic lipid nanoparticles comprise a cationic lipid, a PEG-modified lipid, a sterol, and a non-cationic lipid. In some embodiments, the cationic lipid nanoparticles have a molar ratio of about 20-60% cationic lipid: about 5-25% non-cationic lipid: about 25-55% sterol, and about 0.5-15% PEG-modified lipid. In some embodiments, the cationic lipid nanoparticles comprise a molar ratio of about 50% cationic lipid, about 1.5% PEG-modified lipid, about 38.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid nanoparticles comprise a molar ratio of about 55% cationic lipid, about 2.5% PEG-modified lipid, about 32.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid is an ionic cationic lipid, the non-cationic lipid is a neutral lipid, and the sterol is cholesterol. In some embodiments, the cationic lipid nanoparticles have a molar ratio of cationic lipid:cholesterol:PEG2000-DMG:DSPC or DMG:DOPE of 50:38.5:10:1.5. In some embodiments, the lipid nanoparticles described herein may comprise cholesterol, 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,1'-((2-(4-(2-((2-(bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-1-yl)ethyl)azanediyl)bis(dodecan-2-ol) (C12-200), and DMG-PEG-2000 in a molar ratio of 47.5:16:35:1.5.

[0140] Gene editing methods The present disclosure describes, in some embodiments, a method of gene editing comprising: a) providing or identifying a first binding site sequence in a host genome; b) providing a host cell with a nucleic acid comprising a donor polynucleotide and a second binding site sequence; and c) contacting the host cell with a serine recombinase having 80% sequence identity to any one of SEQ ID NOs: 21-7060 and 7105-7142, or a nucleic acid encoding a serine recombinase, wherein the first binding site sequence and the second binding site sequence are capable of recombining.

[0141] In some embodiments, the first binding site sequence is endogenous to the host genome.

[0142] In some embodiments, the first binding site sequence is provided using viral delivery, in some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh 10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, In some embodiments, the herpesvirus is AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or a derivative thereof. In some embodiments, the herpesvirus is HSV-1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0143] In some embodiments, the first binding site sequence is provided using a transposase. In some embodiments, the transposase is a transposase (Tnp), Tn5, Sleeping Beauty transposase, or Tn7 transposon. In some embodiments, the gene editing system comprises an enzyme having transposase activity. Additional enzymes having transposase activity include, but are not limited to, retrons and IS200 / IS605 transposons.

[0144] In some embodiments, the first binding site sequence is provided using a nuclease. In some embodiments, the nuclease is a double-stranded nuclease.

[0145] In some embodiments, the nuclease is a type II CRISPR endonuclease. In some embodiments, the nuclease is Cas9. Type II CRISPR systems are considered the simplest in terms of components. In type II CRISPR systems, processing of the CRISPR array into mature crRNA does not require the presence of a special endonuclease subunit, but rather a small trans-encoded crRNA (tracrRNA) with a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure consisting of a RuvC-like endonuclease domain that adopts an RNase H fold with an unrelated HNH nuclease domain inserted into the RuvC-like nuclease domain fold. The RuvC-like domain is responsible for cleaving the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleaving the displaced DNA strand.Exemplary CRISPR Cas9 proteins include those from Streptococcus pyogenes (UniProtKB-Q99ZW2 (CAS9 STRP1)), Streptococcus thermophilus (UniProtKB-G3ECR1 (CAS9 STRTR)), Staphylococcus aureus (UniProtKB-J7RUA5 (CAS9 STAAU), Campylobacter jejuni (UniProtKB-Q0P897 (CAS9 CAMJE)), Campylobacter lari (UniProtKB-A0A0A8HTA3 (A0A0A8HTA3 CAMLA), and Helicobacter canadensis (UniProtKB-C5ZYI3 (C5ZYI3 9HELI)), Francisella Examples of Cas9 nucleases include, but are not limited to, Cas9 from Bacillus tularensis subsp. Novicida (UniProtKB-A0Q5Y3(CAS9_FRATN)). Additional type II nucleases are described in International Patent Applications WO 2021 / 226363, WO 2022 / 159758, and WO 2022 / 056324.

[0146] In some embodiments, the nuclease is a CRISPR nuclease. In some embodiments, the CRISPR nuclease is a class 2 type II SpCas9 or a class 2 type VA Cas12a (formerly Cpf1). In some embodiments, the type VA nuclease has a guide RNA of 42-44 nucleotides, compared to approximately 100 nt for SpCas9. In some embodiments, the type VA nuclease produces staggered cleavage sites. In some embodiments, the type VA nuclease produces staggered cleavage sites, facilitating directed repair pathways such as microhomology-dependent targeted integration (MITI).

[0147] In some embodiments, the nuclease is a type V CRISPR endonuclease. Type V CRISPR systems are characterized by a nuclease effector (e.g., Cas12) structure similar to that of type V effectors, including a RuvC-like domain. Like type II CRISPR systems, most (if not all) type V CRISPR systems use tracrRNA to process pre-crRNA into mature crRNA. However, unlike type II systems, which require RNAse III to cleave pre-crRNA into multiple crRNAs, type V systems can cleave pre-crRNA using the effector nuclease itself. Like type II CRISPR systems, type V CRISPR systems are known as DNA nucleases. Unlike type II CRISPR systems, some type V enzymes (e.g., Cas12a) appear to have robust single-strand nonspecific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.

[0148] The most commonly used Type VA enzymes require a 5' protospacer adjacent motif (PAM) adjacent to the selected target site: 5'-TTTV-3' for Lachnospiraceae bacterium ND2006 LbCas12a and Acidaminococcus species AsCas12a; and 5'-TTV-3' for Francisella novicida FnCas12a. In some embodiments, the PAM sequence is YTV, YYN, or TTN. Additional Type II nucleases are described in International Patent Application Publication No. WO 2021 / 226363.

[0149] In some embodiments, the first binding site sequence is provided using a reverse transcriptase. Reverse transcription is the translation of an RNA template into complementary DNA. Reverse transcription is performed by an enzyme called reverse transcriptase (RT), which has RNA-dependent DNA polymerase activity to generate a complementary DNA (cDNA) strand from an RNA template. Some RT enzymes also have DNA-dependent DNA polymerase activity to generate double-stranded dsDNA. Reverse transcriptases can be of viral origin (e.g., HIV, hepatitis B, Moloney murine leukemia virus (MMLV), or avian myeloblastosis virus (AMV)) or bacterial origin (e.g., group II intron, retron / retron-like RT, diversity-generating retroelement (DGR), Abi-like RT, CRISPR-associated RT, and group II-like RT (G2L)). Reverse transcriptases of eukaryotic origin include telomerase reverse transcriptase, which maintains the telomeres of eukaryotic chromosomes. Reverse transcription allows site-specific insertions, deletions, and mutations to be introduced into the cDNA by encoding them on the RNA template.

[0150] In some embodiments, the reverse transcriptase is a viral, prokaryotic, or eukaryotic reverse transcriptase. In some embodiments, the reverse transcriptase is an MG151, MG153, or MG160 family reverse transcriptase. In some embodiments, the reverse transcriptase is an MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, or MG176 family reverse transcriptase. In some embodiments, the reverse transcriptase is any one of the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, or MG176 family reverse transcriptases or retrotransposases. Includes sequences with at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of the MG140, MG146, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, or MG176 family reverse transcriptases or retrotransposases, or a variant thereof. In some embodiments, the reverse transcriptase is smaller than 300 amino acids. In some embodiments, the reverse transcriptase is smaller than 250 amino acids.

[0151] In some embodiments, the method is used to introduce a modification into the genome of a cell. In some embodiments, the modification is an insertion, deletion, or mutation. In some embodiments, the method is used to introduce site-specific insertions, deletions, and / or mutations (e.g., insertions and mutations) into the genome of a cell. In some embodiments, the method is used in combination with a nucleic acid template to facilitate site-specific insertion into the genome of a cell. In some embodiments, the cell is a human cell. In some embodiments, the genome of the cell or a vector contained in the cell is modified. In some embodiments, the genome of the cell is modified ex vivo. In some embodiments, the genome of the cell is modified in vivo.

[0152] In some embodiments, the methods described herein further include detecting the genomic modification. In some embodiments, after the cell genome is modified, the cells are cultured for a certain period of time. In some embodiments, DNA or RNA is extracted and sequenced, and the modified sequence region is mapped and compared to the unmodified sequence. In some embodiments, the cells are stained with an antibody to the protein product translated from the modified nucleic acid, and the resulting stained protein or polypeptide in the cells is analyzed, for example, by flow cytometry.

[0153] cell In certain embodiments, the present disclosure describes a cell comprising the system described herein. In some embodiments, the cell (e.g., a mammalian cell) comprises a eukaryotic genome described in this disclosure. In some embodiments, the cell is a human cell.

[0154] In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungal cell), a mammalian cell (Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embryonic kidney (HEK), mouse myeloma (NS0), or human retinal cell), an immortalized cell (e.g., a HeLa cell, a COS cell, a HEK-293T cell, an MDCK cell, a 3T3 cell, a PC12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, an N2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, an S2 cell, or a Heliothis virescens cell), a yeast cell (e.g., a Saccharomyces In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.

[0155] In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.

[0156] In some embodiments, the cell is a liver cell.

[0157] kit In some embodiments, the present disclosure provides kits that include one or more nucleic acid constructs encoding various components of a serine recombinase described in the present disclosure, including, for example, nucleotide sequences encoding components of a serine recombinase capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequences include a heterologous promoter that drives expression of a serine recombinase described in the present disclosure.

[0158] In some embodiments, the serine recombinases disclosed in the present disclosure are incorporated into pharmaceutical, diagnostic, or research kits to facilitate their use in therapeutic, diagnostic, or research applications. The kits may include one or more containers housing any of the vectors disclosed in the present disclosure and instructions for use.

[0159] Kits can be designed to facilitate researchers' use of the methods described herein and can take many forms. Each of the components of the kit can be provided in liquid form (e.g., in solution) or solid form (e.g., dry powder), if applicable. In some embodiments, the compositions are configurable or otherwise processable (e.g., to an active form), for example, by the addition of a suitable solvent or other species (e.g., water or cell culture medium), which may or may not be provided with the kit. As used herein, "instructions" defines instructional and / or promotional components and may typically involve written instructions on or associated with the packaging of the present disclosure. Instructions may also include any oral or electronic instructions provided in any manner that clearly identifies the instructions to the user as being associated with the kit, such as audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communications, etc. The written instructions, in some embodiments, are in a form prescribed by a government agency that regulates the manufacture, use, or sale of pharmaceutical or biological products, and the instructions may also reflect approval by the agency of manufacture, use, or sale for animal administration. [Example]

[0160] The following examples are given for the purpose of illustrating various embodiments of the present disclosure and are not intended to limit the present disclosure in any way. The examples, together with the methods described in the present disclosure, represent currently preferred embodiments, are exemplary, and are not intended as limitations on the scope of the present disclosure. Modifications therein and other uses encompassed within the spirit of the present disclosure as defined by the scope of the claims will occur to those skilled in the art.

[0161] [Example 1] Bioinformatic identification of large serine recombinases This example describes the identification of proteins with large serine recombinase function through a bioinformatics approach.

[0162] Putative large serine recombinases (LSRs) were identified in an extensive database of viral, prokaryotic, and eukaryotic proteins. The search yielded 163,797 non-partial homologs with a score >50. LSRs were further filtered by requiring contigs with 1 kbp flanks on either side of the LSR and duplicates with an average amino acid identity (AAI) of 90%. After duplicate removal, 8,364 LSRs were globally aligned and a phylogenetic tree was constructed. Closely related contigs lacking an LSR were identified by searching for contigs containing two genes flanking the approximate proviral boundary. To ensure that contigs were from closely related lineages, local alignments were performed, requiring the two genes to share an AAI of ≥99%. Precise proviral boundaries were identified by locally aligning contigs containing and lacking the LSR at the nucleotide level. Once the integration boundary was delineated, the common core of binding sites as well as attL and attR sites flanking the prophage were identified by searching for imperfect repeats near the boundary.

[0163] LSR candidates were identified based on the presence of resolving domains, recombinase domains, and Zn finger domains, as well as catalytic residues required for activity (Figure 1). Selected LSR candidates belonging to the MG178 family share 26.8% AAI among themselves and <37% AAI with known Bxb1 LSR references (Figure 1). Phylogenetic analysis of the LSR candidates showed that these enzymes are encoded in highly diverse genomes, and many gene boundaries were predicted (Figure 2A and Figure 2B). LSR-integrated prophages appeared to be inserted into genes containing the Mn_catalase domain, queosin_synthesis domain, and DUF4244 Pfam domain (Figure 2A) and were shown to infect several Phila hosts, including Actinobacteria, Firmicutes, and Proteobacteria (Figure 2B). The prophage genome mobilized by the LSR reached approximately 94 kb in length. Prophage boundaries were identified by aligning contigs containing LSRs with highly similar contig sequences lacking LSRs, which likely represent hosts without integration events (Fig. 3). With integration boundaries delineated, a common core of binding sites was identified by searching for repeats near the boundaries (Fig. 3).

[0164] [Example 2] In vitro assay of serine recombinase activity In vitro recombination reaction

[0165] To test the functionality of the serine recombinase, attL, attR, and attP and attB sites from the consensus core sequence from the native integrated prophage genome context (SEQ ID NOS: 1-18, GCATCCCC, and TATTCGAT) were bioinformatically determined and tested in in vitro recombination reactions. The attB and attP sites were synthesized in approximately 300 bp long gene fragments with unique primer binding sites at each binding site end (Figure 4C). The serine recombinase was expressed in vitro, while a negative control included a template-free in vitro expression reaction (null) (Figure 4A). A negative recombination reaction control was set up in a 10 µL reaction using 50 ng of attB, 50 ng of attP, recombination buffer (20 mM HEPES pH 7.5, 50 µg / mL bovine serum albumin (BSA), 2 mM TCEP, 5 mM MgCl, 100 mM KCl, 5 mM spermidine, 0.2 mM ZnCl, and 5% glycerol), and 1 µL of a null reaction (no recombinase template). The experimental condition included 50 ng of attB, 50 ng of attP, and 1 µL of in vitro-expressed recombinase (Figure 4B). The recombination reaction was incubated at 30 °C for 1 h and diluted 1:10 in water. The fragments were then purified by PCR using attL- (attB5 and attP3) or attR- (attB3 and attP5) specific primer sets (Figure 4C) and analyzed on a 2% agarose gel to measure the amplification and size of the resulting products. Product-forming reactions were Sanger sequenced and aligned with the predicted attL and attR sequences determined bioinformatically.

[0166] LSR candidates were expressed in vitro and added to the reaction buffer containing the putative attB and attP dsDNA fragments (Figures 4A-C). Four LSRs (MG178-4 (SEQ ID NO: 21), MG178-9 (SEQ ID NO: 22), MG178-10 (SEQ ID NO: 23), and MG178-11 (SEQ ID NO: 24)) were active based on the formation of both attL and attR recombination products (Figure 5). PCR amplifications were then Sanger sequenced to confirm the predicted attB- and attP-forming attL and attR sequence crossover events. The results show that Sanger sequencing confirmed the predicted recombination events of active recombinase-containing reactions for both attL and attR, as well as the conservation of the common core in both the reactants and the recombination products (Figures 6 and 7).

[0167] [Example 3] Predictive Experiments - Intracellular Plasmid Recombination Recombinases are tested for their activity in human cells by synthesizing an attP fragment into a donor plasmid (pDonor) with an attP site upstream of a promoterless mCherry-encoding ORF. An attB fragment is synthesized into a pTarget plasmid encoding a pCMV promoter upstream of an attB site, without a downstream coding ORF. When cotransfected with an active recombinase, the pCMV promoter in pTarget recombines with pDonor mCherry, and the junction of the pCMV promoter to mCherry drives transcription and translation of the mCherry-encoding region. Recombinase efficiency is compared to a negative control of a cell population transfected with both pDonor and pTarget without the recombinase plasmid.

[0168] [Example 4] Predictive experiments - Landing pad activity in mammalian cells To introduce exogenous donor DNA into the human genome using large serine recombinases, landing pad, attP, or attB sequence sites are either (1) found endogenous to the human genome sequence, or (2) introduced using viral delivery or via a transposable element, (3) integrated into the genome using HDR coupled with a nuclease, or (4) reverse transcribed into the genome using a targeted reverse transcriptase.

[0169] After introducing or identifying a landing pad (either attP or attB) site into the genome, LSR activity into the genome is measured by using a DNA donor containing (1) a promoter-driven fluorescent protein construct, or (2) a promoterless fluorescent coding construct with the cognate attachment (attB / attP) site, and / or (3) an antibiotic resistance marker, or (4) a screenable cell surface marker. The donor is introduced into cells as a plasmid, minicircle, bacterial artificial chromosome, nanoplasmid, or linear dsDNA construct and integrated into the landing pad.

[0170] In addition to introducing the donor into cells, the LSR is transfected into cells using either (1) a plasmid encoding transcription and translation of the LSR, (2) mRNA encoding LSR translation, or (3) purified protein. Landing pad efficiency is measured by flow analysis in the case of fluorescent protein and / or cell surface marker donors, or by colony formation under selective conditions and subsequent PCR analysis of exogenous / endogenous DNA junction formation.

[0171] [Example 5] In silico identification of large serine recombinases in the MG178 family In silico identification of LSRs and their putative binding sites Putative large serine recombinases (LSRs) were identified using the following modifications: LSR domain-specific (PF00239 and PF07508) hmm searches yielded 987,835 non-partial homologs with scores >50 and lengths >450 aa. LSRs with at least 1 kbp flanks on both sides were de-replicated at 99% AAI, yielding 146,897 non-redundant homologs. LSR attL and attR sites were identified.

[0172] result LSR candidates were identified based on the presence of resolving domains, recombinase domains, and Zn-finger domains, as well as catalytic residues required for activity (Figure 8). Selected LSR candidates belonging to the MG178 family share a 16.9% AAI among themselves and a <18% AAI with known BxB1 LSRs. The LSRs identified in this study were integrated into genes belonging to the radical SAM superfamily, including glycosyl hydrolase family 18, the helix-turn-helix domain of the transportase family ISL3, the peptidase family M3, transcriptional regulators, outer membrane protein beta-barrel domains, type II / IV secretion system proteins, the acetyltransferase (GNAT) family, the MFS_1-like family, magnesium chelatases, and manganese containing catalytic Pfam domains, and transfected into unannotated genes, including messenger RNAs, T-box leader RNAs, and intergenic regions. Viruses encoding the identified LSRs infect diverse host sequences, including Actinobacteria, Proteobacteria, Bacteroidetes, Firmicutes, Lentisphaerota, Fusobacteria, Candidatus aminicenantes, and unknown phyla. The proviral genomes recruited by LSRs reached lengths of approximately 62 kbp. Proviral boundaries were identified by aligning LSR-containing contigs with highly similar sequences lacking LSRs, likely representing hosts without integration events. Once integration boundaries were delineated, a common core of LSR binding sites was identified by searching for direct repeats near the boundaries. Complete and incomplete repeats representing the common core were identified by finding conserved regions in local alignments of proviral boundaries (Figure 9A and Figure 9B). If the alignment did not show conservation, repeats were visually identified and the alignment was manually refined (Figure 9C).

[0173] [Example 6] Intracellular Plasmid Recombination Cellular Plasmid Recombination Reaction 150,000 HEK293T cells seeded for 24 hours were transfected with 1 μg of integrase, 0.5 μg of an attP-containing plasmid, and 0.5 μg of an attB-containing plasmid using LT1 transfection reagent. Transfected cells were incubated at 37°C for 48 hours, then harvested using 0.25% trypsin reagent, washed with 1x PBS, and stained with Fixable near-IR Live / Dead reagent. Treated cells were then analyzed by flow cytometry using negative controls, gating on cells expressing neither eGFP (integrase) nor mCherry (recombinant), and eGFP only (integrase expression only). Cell analysis was performed by calculating the percentage of mCherry-positive (recombinant) cells over the total number of eGFP-positive (integrase transfected and expressed) cells.

[0174] result

[0175] Selected LSR recombinases were tested for their activity in human cells by synthesizing the recombinase and an attP fragment into a donor plasmid (pDonor) containing an attP site upstream of a promoterless mCherry-encoding ORF. The attB fragment was synthesized into a pTarget plasmid encoding a pCMV promoter upstream of an attB site without a downstream coding ORF (Figure 10A). When cotransfected with the active recombinase, the pTarget pCMV promoter recombines with pDonor mCherry, and the junction of the pCMV promoter to mCherry drives transcription and translation of the mCherry-encoding region. Recombinase efficiency was compared to a negative control of a cell population transfected with both the recombinase plasmid and pDonor without the pTarget plasmid. Of the recombinases tested, MG178-7202 (SEQ ID NO: 7140) was active only against the single predicted binding site 1 (GGGCACCC) in 50% of all transfected cells, MG178-7193 (SEQ ID NO: 7131) was active in up to 45%, MG178-1859 (SEQ ID NO: 1848) and MG178-7177 (SEQ ID NO: 7115) were active in up to 30%, MG178-7201 (SEQ ID NO: 7139) recombined in up to 20%, and MG178-7173 (SEQ ID NO: 7111) and MG178-7198 (SEQ ID NO: 7136) recombined in less than 20% of all transfected cells (Figure 10B).

[0176] [Example 7] In vitro recombination of the LSR system In vitro testing of recombination To test the functionality of the serine recombinase, attP and attB sites were predicted from the attL, attR, and consensus core sequences from the natural integrated prophage genome context. The attB and attP sites were synthesized in approximately 300 bp long gene fragments with unique primer binding sites at the ends of each binding site (Figure 4C). The serine recombinase was expressed in vitro, while negative controls included in vitro expression reactions without template (null) (Figures 4A-4C). A negative recombination reaction control was set up in a 10 µL reaction using 100 ng of attB, 100 ng of attP, recombination buffer (20 mM HEPES pH 7.5, 50 µg / ml bovine serum albumin (BSA), 2 mM TCEP, 5 mM MgCl, 100 mM KCl, 5 mM spermidine, 0.2 mM ZnCl, and 5% glycerol), and 1 µL of a used null reaction (no recombinase template). The experimental conditions included 100 ng of attB, 100 ng of attP, and 1 µL of in vitro-expressed recombinase. The recombination reaction was incubated at 30 °C for 1 h and diluted 1:10 with water. PCR reactions were then performed using a recombinase-specific primer set (sequence numbers 7407-7411) and run on a 2% agarose gel to measure the amplification and size of the resulting products.

[0177] result

[0178] LSR candidates were expressed in vitro and added to the reaction buffer containing the attB and attP dsDNA fragments determined by cell recombination. Four LSRs (MG178-7202, SEQ ID NO: 7096, MG178-7193, SEQ ID NO: 7087, MG178-1859, SEQ ID NO: 1848, and MG178-7177, SEQ ID NO: 7071) were active based on strong PCR-amplified recombination products, which were not observed under negative control conditions that did not contain the recombinase enzyme (Figure 11).

[0179] [Example 8] Intracellular plasmid recombination by active MG178 candidates Cellular Plasmid Recombination Reaction Using LT1 transfection reagent, 150,000 HEK293T cells seeded for 24 hours were transfected with 1 μg of integrase, 0.5 μg of an attP-containing plasmid, and 0.5 μg of an attB-containing plasmid. Transfected cells were incubated at 37°C for 48 hours, then harvested using 0.25% trypsin, washed with 1x PBS, and stained with Fixable near-IR Live / Dead reagent. Treated cells were then analyzed by flow cytometry using negative controls, gating on cells expressing neither eGFP (integrase) nor mCherry (recombinant), and eGFP only (integrase expression only). Cell analysis was performed by calculating the percentage of mCherry-positive (recombinant) cells over the total number of eGFP-positive (integrase transfected and expressed) cells.

[0180] result

[0181] Selected LSR recombinases were tested for their activity in human cells by synthesizing the recombinase and attP fragment into a donor plasmid (pDonor) containing an attP site upstream of a promoterless mCherry-encoding ORF. The attB fragment was synthesized into a pTarget plasmid encoding a pCMV promoter upstream of an attB site without a downstream coding ORF (Figures 10A and 10B). When cotransfected with the active recombinase, the pTarget pCMV promoter recombines with pDonor mCherry, and the junction of the pCMV promoter to mCherry drives transcription and translation of the mCherry-encoding region. Recombinase efficiency was compared to a negative control of a cell population transfected with both the recombinase plasmid and pDonor without the pTarget plasmid. Of the recombinases tested, MG178-7178 (SEQ ID NO: 7072), MG178-7199 (SEQ ID NO: 7093), and MG178-7170 (SEQ ID NO: 7064) recombined in less than 5% of total transfected cells, while MG178-7201 (SEQ ID NO: 7095) promoted recombination in greater than 15% (Figure 12).

[0182] [Example 9] Human cell recombination as a result of plasmid dosage The level of intracellular recombinase activity is easily affected by the amount of recombinase, target, and donor plasmids introduced into the cells. To test the most effective ratio of plasmid administration, we modified the recombinase plasmid individually or with the target and donor plasmids.

[0183] Cellular Plasmid Recombination Reaction

[0184] Using LT1 transfection reagent, 150,000 HEK293T cells seeded for 24 hours were transfected with various levels (0.1–1 μg) of integrase, 0.1–0.5 μg of attP-containing plasmid, and 0.1–0.5 μg of attB-containing plasmid. Transfected cells were incubated at 37°C for 48 hours, then harvested using 0.25% trypsin reagent, washed with 1x PBS, and stained with Fixable near-IR Live / Dead reagent. Treated cells were then analyzed by flow cytometry using negative controls, gating on cells expressing neither eGFP (integrase) nor mCherry (recombinant), and eGFP only (integrase expression only). Cell analysis was performed by calculating the percentage of cells positive for mCherry (recombinant) over the total number of cells positive for eGFP (integrase transfection and expression).

[0185] result

[0186] The candidate MG178-7202 (SEQ ID NO: 7096) LSR recombinase was tested for its activity in human cells by administering various levels of integrase, donor, and target plasmids to measure the increase in recombination efficiency. MG178-7202 was found to be most active at 250 ng per transfection, with an amount of integrase plasmid equal to the target and donor plasmids. This represents a 30% increase in recombinase activity conferred by the concentration of plasmid in the cells (Figures 13A and 13B).

[0187] [Example 10] Minimization of cell-binding sites Construction of minimal binding sites Minimizing the binding site is important for understanding the limits of recombinase activity in eukaryotic cells. A smaller binding site footprint allows efficient integration of attB or attP sites into any target locus in the human genome, either via a dsDNA donor or via RNA-templated addition to the genome for binding site integration. To identify minimal sequences, a series of AttP and attB variant sites were synthesized with the promoterless mCherry described above for attP, and with a markerless promoter for attB. The size reduction of both attB and attP was benchmarked against a 300-nt active binding site (SEQ ID NO: 7100). For MG178-7202, the attB sequences tested corresponded to sizes of 108, 88, 68, 58, 48, 46, 44, 42, 40, 38, 36, 32, and 28 nt (SEQ ID NOs: 7188-7200), and the attP sequences were tested at 108, 88, 68, 58, and 48 nt (SEQ ID NOs: 7183-7187). For MG178-7193, the attB sites were tested at 112, 92, 72, 62, and 52 nt (SEQ ID NOs: 7206-7210), and the attP sites were tested at sizes of 112, 92, 72, 62, and 52 nt (SEQ ID NOs: 7201-7205).

[0188] Cellular plasmid synthesis and recombination reactions

[0189] Using LT1 transfection reagent, 150,000 HEK293T cells seeded for 24 hours were transfected with 250 ng of integrase, 250 ng of attP-containing plasmid, and 250 ng of attB-containing plasmid. Transfected cells were incubated at 37°C for 48 hours, then harvested using 0.25% trypsin reagent, washed with 1x PBS, and stained with Fixable near-IR Live / Dead reagent. Treated cells were then analyzed by flow cytometry using negative controls, gating on cells expressing neither eGFP (integrase) nor mCherry (recombinant), and eGFP only (integrase expression only). Cell analysis was performed by calculating the percentage of cells positive for mCherry (recombinant) over the total number of cells positive for eGFP (integrase transfection and expression).

[0190] result

[0191] A range of minimal binding sites was tested for the two LSR recombinases. All combinations were benchmarked against the original 300 bp binding site for comparison of recombination efficiency. For MG178-7202, the most efficient recombination occurred between 58 nt for attP and 48 nt for attB (Figure 14). Recombinase activity was further detected at 48 nt attP and decreased to 32 nt for attB. The MG178-7193 binding site was shown to be most active at the 52 / 72 attB / attP size, although recombinase activity was measured down to 52 / 62 attB / attP (Figure 15).

[0192] [Example 11] MG178s purification and in vitro activity assay Isolating pure and functional proteins is essential for extensive in vitro analysis of biochemical properties and mechanistic studies. MG178 candidates were expressed and purified to obtain sufficient quantity and quality of protein for such characterization. MG178-1859 (SEQ ID NO: 1848) was expressed as an N-terminal samo fusion protein in a carbenicillin-resistant pMGF expression vector, while MG178-7202 (SEQ ID NO: 7096) was expressed as an N-terminal samo fusion protein in a kanamycin-resistant pET28 expression vector. All constructs were expressed in E. coli.

[0193] Protein expression

[0194] Protein expression plasmids were transformed into competent cells and cultured overnight at 37 °C in 50 mL of 2xYT medium (1.6% tryptone, 1% yeast extract, 0.5% NaCl) with 100 μg / mL carbenicillin or 50 μg / mL kanamycin, depending on the expression vector. The next day, 7 mL from each overnight culture was used to inoculate 1000 mL of TB medium (1.2% tryptone, 2.4% yeast extract, 0.4% glycerol, 17 mM potassium phosphate monobasic, 72 mM potassium phosphate dibasic) containing 100 μg / L carbenicillin or 50 μg / mL kanamycin at 37 °C. Cultures were grown at 37 °C with shaking. At OD600 ~0.8-1.2, the cultures were cooled on ice and then induced with 0.3 mM IPTG and 0.2% w / v L-(+)-arabinose and further incubated at 16°C with shaking for approximately 18 h. The cultures were then harvested by centrifugation at 6,000 × g for 10 min, and the pellets were resuspended in Nickel-A buffer (50 mM HEPES, 500 mM NaCl, 10 mM MgCl2, 1 mM EDTA, 20 mM imidazole, 5% glycerol, pH 7.5) containing protease inhibitors (EDTA-free) plus 2 mg / mL lysozyme (lysozyme from chicken egg white, Research Product International L38100) and stored at -80°C. Culture samples were taken pre- and post-induction, and cells were pelleted via centrifugation (15,000 xg, 1.5 min) and resuspended in 100 μL of 2x Laemmli buffer per 1 OD cells.

[0195] Protein purification MG178-7202 (SEQ ID NO: 7096) is shown here as an example of the protein purification process. The expressed protein has the following sequence structure: 6xHis-(GS)1-Sumo-GSGSGGSGS-PSP-SV40 NLS-HA-MG178. The cell pellet was thawed and the volume was supplemented to 120 mL with nickel buffer (P1P1P1, CI-00234) containing 0.5% s-octylglucoside. The sample was sonicated in an ice-water bath at 75% amplitude using a 5-second on / 15-second off cycle for a total processing time of 3 minutes. The lysate was clarified by centrifugation at 30,000 × g for 15 min, and the supernatant batch was bound to 5 mL of Ni-NTA resin (≥ 15 min. The sample was loaded onto a gravity column, washed with 10 CV of Nickel_A buffer, washed again with 10 CV of Nickel_A2 buffer (Nickel_A buffer + 100 mM imidazole), then washed with 2 CV of Nickel_B buffer (Nickel_A buffer + 300 mM imidazole) and 2 CV of Nickel_B2 buffer (Nickel_A buffer + 500 mM imidazole). The purified protein was eluted with 1000 kDa MWCO. Fractions collected in Nickel_B and Nickel_B2 buffers were pooled and then concentrated using a 50 kDa MWCO concentrator. Samples were collected throughout the purification process and run on SDS-PAGE protein gels, which were imaged on a ChemiDoc in the unstained channel after 5 minutes of UV activation. These gels were used to track the progress of the purification throughout the protocol (Figure 16A). The MG178-7202 sample was then filtered through a 0.22 μm cellulose acetate membrane before being loaded onto an S200i 10 / 300GL column and treated with SEC buffer (50 mM HEPES, 250 mM NaCl, 10 mM MgCl2, 1 mM EDTA, 5% glycerol, 0.5 mM TCEP, pH 7.5) to further isolate the purified protein (Figures 16B and 16C).

[0196] In vitro testing of recombination To test the functionality of the serine recombinase, we predicted attL, attR, and attP and attB sites from the consensus core sequence from the native integrated prophage genome context. The attB and attP sites were synthesized in a gene fragment approximately 300 bp long, using unique primer binding sites at the end of each binding site. The serine recombinase was expressed in vitro, while a negative control included a template-free in vitro expression reaction (null). A negative recombination reaction control was set up in a 10 μL reaction using 100 ng of attB, 100 ng of attP, recombination buffer (20 mM HEPES pH 7.5, 50 μg / ml bovine serum albumin (BSA), 2 mM TCEP, 5 mM MgCl, 100 mM KCl, 5 mM spermidine, 0.2 mM ZnCl, and 5% glycerol), and 1 μL of the spent null reaction (no recombinase template). Experimental conditions included 100 ng of attB, 100 ng of attP, and 1 μL of in vitro-expressed recombinase. Recombination reactions were incubated at 30°C for 1 hour and diluted 1:10 with water. PCR reactions were then performed with specific primer sets (sequence numbers 7416 and 7417) and run on a 2% agarose gel to measure the amplification and size of the resulting products. Product-forming reactions were Sanger sequenced and aligned with predicted attL and attR sequences to determine bioinformatics.

[0197] result LSR candidates were expressed in vitro and added to the reaction buffer containing the attB and attP dsDNA fragments determined by cell recombination. Two LSR candidates (MG178-7202 (SEQ ID NO: 7096) and MG178-1859 (SEQ ID NO: 1848)) were active based on strong PCR-amplified recombination products that were not observed under negative control conditions that did not contain the recombinase enzyme and were more specific compared to the in vitro-expressed controls (Figure 17). The results support previous observations of active protein expression from cell-free extracts for in vitro recombination activity (Example 7). [Prior art documents]

Non-licensed literature

[0198] [Non-licensed document 1] Anzalone AV, Gao XD, Podracky CJ, Nelson AT, Koblan LW, Raguram A, Levy JM, Mercer JAM, Liu DR. Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing. Nat Biotechnol.2022,40(5):731-740.doi:10.1038 / s41587-021-01133-w. Epub 2021 Dec 9. PMID: 34887556; PMCID: PMC9117393.

[0199] [Non-licensed document 2] Durrant MG, Fanton A, Tycko J, Hinks M, Chandrasekaran SS, Perry NT, Schaepe J, Du PP, Lotfy P, Bassik MC, Bintu L, Bhatt AS, Hsu PD. Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome. Nat Biotechnol. 2022 Oct 10.doi:10.1038 / s41587-022-01494-w. PMID: 36217031

[0200] [Non-licensed document 3] Smith MCM. Phage-encoded Serine Integrases and Other Large Serine Recombinases. Microbiol Spectr. 2015 Aug;3(4).doi:10.1128 / microbiolspec.MDNA3-0059-2014. PMID: 26350324

[0201] [Non-patent document 4] Robert C. Edgar, Search and clustering order of magnitude faster than BLAST, Bioinformatics, Volume 26, Issue 19, 1 October 2010, Pages 2460-2461, doi.org / 10.1093 / bioinformatics / btq461

[0202] [Non-Patent Document 5] Nayfach, S., Camargo, AP, Schulz, F. et al. CheckV assesses the quality and completeness of metagenome-assembled viral genomes. Nat Biotechnol 39, 578-585 (2021). doi.org / 10.1038 / s41587-020-00774-7

[0203] [Non-patent document 6] Price MN, Dehal PS, Arkin AP (2010) FastTree 2 - Approximately Maximum-Likelihood Trees for Large Alignments. PLoS ONE 5(3): e9490. doi.org / 10.1371 / journal.pone.0009490

[0204] [Non-Patent Document 7] Katoh K,Standley DM. MAFFT multiple sequence alignment software version 7:improvements in performance and usability. Mol Biol Evol.2013;30(4):772-780.doi:10.1093 / molbev / mst010

[0205] [Non-patent document 8] Steinegger, M., Soding, J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol 35, 1026-1028 (2017). doi.org / 10.1038 / nbt.3988.

[0206] equivalent The present disclosure may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the foregoing embodiments are to be considered in all respects as illustrative and not limiting of the disclosure set forth herein. The scope of the present disclosure is therefore indicated by the appended claims, rather than the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein. [Table 3-1] [Table 3-2] [Table 3-3] [Table 3-4] Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 3-10 Table 3-11 Table 3-12 Table 3-13 Table 3-14 Table 3-15 Table 3-16 Table 3-17 Table 3-18 Table 3-19 Table 3-20 Table 3-21 Table 3-22 Table 3-23 Table 3-24 Table 3-25 Table 3-26 Table 3-27 Table 3-28 Table 3-29 Table 3-30 Table 3-31 Table 3-32 Table 3-33 Table 3-34 Table 3-35 Table 3-36 Table 3-37 Table 3-38 Table 3-39 Table 3-40 Table 3-41 Table 3-42 Table 3-43 Table 3-44 Table 3-45 Table 3-46 Table 3-47 Table 3-48 Table 3-49 Table 3-50 Table 3-51 Table 3-52 Table 3-53 Table 3-54 Table 3-55 Table 3-56 Table 3-57 Table 3-58 Table 3-59 Table 3-60 Table 3-61 Table 3-62 Table 3-63 Table 3-64 Table 3-65 Table 3-66 Table 3-67 Table 3-68 Table 3-69 Table 3-70 Table 3-71 Table 3-72 Table 3-73 Table 3-74

Claims

1. A gene editing system, a) a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214, or a nucleic acid encoding said serine recombinase; b) a donor polynucleotide and a nucleic acid comprising a first binding site sequence.

2. 2. The gene editing system of Claim 1, wherein the first binding site sequence is 5' to the donor polynucleotide.

3. The gene editing system of any one of claims 1 to 2, wherein the nucleic acid encoding the serine recombinase further comprises a second binding site sequence.

4. The gene editing system of Claim 3, wherein the second binding site sequence is 5' to the serine recombinase.

5. The gene editing system of any one of claims 3 to 4, wherein the first binding site sequence and the second binding site sequence are recombinable.

6. The gene editing system of any one of claims 1 to 5, wherein the first binding site sequence is a bacterial genome recombination sequence (attB).

7. The gene editing system of any one of claims 1 to 5, wherein the first binding site sequence is a phage genome recombination sequence (attP).

8. The gene editing system of any one of claims 3 to 7, wherein the second binding site sequence is a bacterial genome recombination sequence (attB).

9. The gene editing system of any one of claims 3 to 7, wherein the second binding site sequence is a phage genome recombination sequence (attP).

10. 10. The gene editing system of any one of claims 6 to 9, wherein the attB sequence comprises from about 20 to about 500 nucleotides.

11. The gene editing system of any one of claims 7 to 10, wherein the attP sequence comprises from about 20 to about 500 nucleotides.

12. The attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7277, 7282, 7287, 7288, 7290, 7291, 7292, 7293, 7294, 7295, 7296, 7297, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7400, 7410, 7411, 7412, 7413, 7414, 7415, 7416, 7417, 7418, 7419, 7420, 7421, 7422, 7423, 7424, 7425, 7426, 7427, 7428, 7429, 7430, 7431, 7432, 7433, 7434, 7435, 7436, 7437, 7438, 7439, 732, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402.

13. The gene editing system of any one of claims 6 to 12, wherein the attB has at least about 80% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13.

14. The attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183 to 7187, 7201 to 7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7279, 7284, and 7289. 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399, and 7404. The gene editing system of any one of claims 7 to 13, comprising at least about 80% sequence identity with any one of the following:

15. The gene editing system of any one of claims 7 to 14, wherein the attP comprises at least about 80% sequence identity to any one of SEQ ID NOs: 2, 6, 10 and 14.

16. 16. The gene editing system of any one of claims 1 to 15, wherein the nucleic acid comprising the donor polynucleotide and the first attachment sequence is delivered using a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid.

17. 17. The gene editing system of any one of claims 1 to 16, wherein the nucleic acid encoding the serine recombinase is delivered using a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid.

18. 18. The gene editing system of any one of claims 16 to 17, wherein the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus.

19. The AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AA V11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10, AAV-rh20, A AV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK03 , AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, or AAV-HSC16, or a derivative thereof.

20. 19. The gene editing system of Claim 18, wherein the herpesvirus is HSV-1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7 or HHV-8.

21. 21. The gene editing system of any one of claims 1-20, wherein the donor polynucleotide comprises a size of at least about 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, or greater than 120 kb.

22. 22. The gene editing system of any one of claims 1-21, wherein the donor polynucleotide encodes a therapeutic agent, a reporter, or a marker.

23. 23. The gene editing system of Claim 22, wherein the reporter comprises a fluorescent protein.

24. 24. The gene editing system of Claim 23, wherein the fluorescent protein is GFP, EBFP, EBFP2, azurite, mKalamal, ECFP, Cerulean, CyPet, YFP, Citrine, Venus, YPet, RFP, CFP, or a derivative thereof.

25. 23. The gene editing system of Claim 22, wherein the reporter is acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), beta-galactosidase (LacZ), beta-glucuronidase (GUS), chloramphenicol acetyltransferase (CAT), horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, or a derivative thereof.

26. The gene editing system of any one of claims 22 to 25, wherein the marker is an antibiotic resistance marker.

27. 27. The gene editing system of Claim 26, wherein the antibiotic resistance marker is kanamycin, spectinomycin, streptomycin, ampicillin, carbenicillin, bleomycin, erythromycin, polymyxin B, tetracycline, chloramphenicol, neomycin, zeocin, or a derivative thereof.

28. The gene editing system of any one of claims 22 to 27, wherein the marker is a cell surface marker.

29. A genome of a eukaryotic organism comprising a donor polynucleotide sequence and an attL sequence 5' to the donor polynucleotide sequence, wherein the attL sequence is selected from the group consisting of SEQ ID NOs: 3, 4, 7, 8, 11, 12, 15, 16, 7153, 7157, 7161, 7165, 7169, 7173, 7177, 7181, 7216, 7221, 7227, 7234, 7239, 7244, 7249, 7254, 7260, 7261, 7262, 7263, 7264, 7265, 7266, 7267, 7268, 7269, 7270, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405 7333, 7338, 7343, 7348, 7353, 7358, 7363, 7368, 7373, 7378, 7383, 7388, 7393, 7398, and 7403.

30. 30. The eukaryotic genome of claim 29, further comprising an attR sequence 3' to the donor polynucleotide sequence.

31. A genome of a eukaryotic organism comprising a donor polynucleotide sequence and an attL sequence 3' to the donor polynucleotide sequence, wherein the attL sequence is selected from the group consisting of SEQ ID NOs: 3, 4, 7, 8, 11, 12, 15, 16, 7153, 7157, 7161, 7165, 7169, 7173, 7177, 7181, 7216, 7221, 7227, 7234, 7239, 7244, 7249, 7254, 7260, 7261, 7262, 7263, 7264, 7265, 7266, 7267, 7268, 7269, 7270, 7271, 7272, 7273, 7274, 7275, 7276, 7277, 7278, 7279, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7400, 7401, 7402, 7403, 7404, 7405 7333, 7338, 7343, 7348, 7353, 7358, 7363, 7368, 7373, 7378, 7383, 7388, 7393, 7398, and 7403.

32. 32. The eukaryotic genome of claim 31, further comprising an attR sequence 3' to the donor polynucleotide sequence.

33. A eukaryotic genome comprising: a donor polynucleotide sequence; The attL sequence on the 5' or 3' side of the donor polynucleotide sequence is SEQ ID NO: 3, 4, 7, 8, 11, 12, 15, 16, 7153, 7157, 7161, 7165, 7169, 7173, 7177, 7181, 7216, 7221, 7227, 7234, 7239, 7244, 7249, 7254, 7259, 7265, 7272, 7273, 7274, 7275, 7276, 7277, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7292, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7400, 7401, 7402, 7403, 7404, 7405, 7406, 7407, 7408, 7409, 7500, 7510, 7511, 7512, 7513, 7514, 7515, 7516, 7517, 7518, 7519, 7600, 7610, 7611, 7612, 7613, 7614, 7615, 7616, 7617 7383, 7388, 7393, 7398, 7403, and an attL sequence comprising a sequence selected from the group consisting of: 7383, 7388, 7393, 7398, 7403, 7408, 7413, 7418, 7423, 7428, 7433, 7438, 7443, 7453, 7458, 7463, 7468, 7473, 7478, 7483, 7488, 7493, 7498, 7503, 7513, 7518, 7523, 7528, 7533, 7548, 7553, 7558, 7563, 7568, 7573, 7578, 7583, 7588, 7593, 7598, and 7603; The attR sequence on the 5' or 3' side of the donor polynucleotide sequence is SEQ ID NO: 3, 4, 7, 8, 11, 12, 15, 16, 7154, 7158, 7162, 7166, 7170, 7174, 7178, 7182, 7218, 7223, 7230, 7236, 7241, 7246, 7251, 7256, 7261, 7268, 7274, 7280, 7281, 7282, 7283, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7292, 7293, 7294, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7330, 7331, 7332, 7333, 7334, 7335, 7336, 7337, 7338, 7339, 7340, 7341 and a eukaryotic genome comprising an attR sequence comprising a sequence selected from the group consisting of: 7385, 7290, 7295, 7300, 7305, 7310, 7315, 7320, 7325, 7330, 7335, 7340, 7345, 7350, 7355, 7360, 7365, 7370, 7375, 7380, 7385, 7390, 7395, 7400, and 7405.

34. 34. The eukaryotic genome of any one of claims 30, 32, or 33, wherein the attL sequence and the attR sequence are identical.

35. 35. The eukaryotic genome of any one of claims 29 to 34, wherein the attL sequence is a recombination sequence of a first binding site sequence and a second binding site sequence.

36. 36. The eukaryotic genome of any one of claims 30 or 32 to 35, wherein the attR sequence is a recombination sequence of a first binding site sequence and a second binding site sequence.

37. The eukaryotic genome of any one of claims 35 to 36, wherein the first binding site sequence is a bacterial genome recombination sequence (attB).

38. The eukaryotic genome of any one of claims 35 to 36, wherein the first binding site sequence is a phage genome recombination sequence (attP).

39. The eukaryotic genome of any one of claims 35 to 38, wherein the second binding site sequence is a bacterial genome recombination sequence (attB).

40. The eukaryotic genome of any one of claims 35 to 38, wherein the second binding site sequence is a phage genome recombination sequence (attP).

41. 41. The eukaryotic genome of any one of claims 37 to 40, wherein the attB sequence comprises from about 20 to about 500 nucleotides.

42. 42. The eukaryotic genome of any one of claims 38 to 41, wherein the attP sequence comprises from about 20 to about 500 nucleotides.

43. The attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188-7200, 7206-7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7277, 7282, 7287, 7288, 7290, 7291, 7292, 7293, 7294, 7295, 7296, 7297, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7400, 7410, 7411, 7412, 7413, 7414, 7415, 7416, 7417, 7418, 7419, 7420, 7421, 7422, 7423, 7424, 7425, 7426, 7427, 7428, 7429, 7430, 7431, 7432, 7433, 7434, 7435, 7436, 7437, 7438, 7439, 292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402. The eukaryotic genome of any one of claims 37-42, comprising at least about 80% sequence identity with any one of

44. 44. The eukaryotic genome of any one of claims 37 to 43, wherein the attB sequence comprises at least about 80% sequence identity with any one of SEQ ID NOs: 1, 5, 9, and 13.

45. The attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183 to 7187, 7201 to 7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7279, 7284, 7289, 7290, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7330, 7331, 7332, 7333, 7334, 7335, 7336, 7337, 7338, 7339, 7340, 7341, 7342, 7343, 7344, 7345, 7346, 7347, 734 294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399, and 7404. The eukaryotic genome of any one of claims 38 to 44, comprising at least about 80% sequence identity with any one of

46. 46. ​​The eukaryotic genome of any one of claims 38 to 45, wherein the attP sequence comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 2, 6, 10, and 14.

47. The attL sequence is selected from SEQ ID NOs: 3, 4, 7, 8, 11, 12, 15, 16, 7153, 7157, 7161, 7165, 7169, 7173, 7177, 7181, 7216, 7221, 7227, 7234, 7239, 7244, 7249, 7254, 7259, 7265, 7272, 7278, 7283, 7288, 7293, 7298, 7303, 7308, 7313, 7318, 7323, 7328, 7333, 7338, 7343, 7348, 7353, 7358, 7363, 7368, 7373, 7378, 7383, 7388, 7393, 7398, and 7403. The eukaryotic genome of any one of claims 29-46, comprising at least about 80% sequence identity with any one of

48. The attR sequence is selected from the group consisting of SEQ ID NOs: 3, 4, 7, 8, 11, 12, 15, 16, 7154, 7158, 7162, 7166, 7170, 7174, 7178, 7182, 7218, 7223, 7230, 7236, 7241, 7246, 7251, 7256, 7261, 7268, 7274, 7280, 7285, 7290, 7295, 7300, 7305, 7310, 7315, 7320, 7325, 7330, 7335, 7340, 7345, 7350, 7355, 7360, 7365, 7370, 7375, 7380, 7385, 7390, 7395, 7400, and 7405. The eukaryotic genome of any one of claims 29-47, comprising at least about 80% sequence identity with any one of

49. A mammalian cell comprising the genome of a eukaryotic organism according to any one of claims 29 to 48.

50. 50. The mammalian cell of claim 49, which is a human cell.

51. 51. The mammalian cell of any one of claims 49 to 50, further comprising a serine recombinase.

52. 52. The mammalian cell of claim 51, wherein the serine recombinase comprises at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214.

53. 52. The mammalian cell of claim 51, wherein the serine recombinase has at least about 80% sequence identity to SEQ ID NO:

21.

54. 52. The mammalian cell of claim 51, wherein the serine recombinase has at least about 80% sequence identity to SEQ ID NO:

22.

55. 52. The mammalian cell of claim 51, wherein the serine recombinase has at least about 80% sequence identity to SEQ ID NO:

23.

56. 52. The mammalian cell of claim 51, wherein the serine recombinase has at least about 80% sequence identity to SEQ ID NO:

24.

57. 52. The mammalian cell of claim 51, wherein the serine recombinase has a transduction efficiency of at least about 5%.

58. 52. The mammalian cell of claim 51, wherein the serine recombinase comprises a transduction efficiency of at least about 25%.

59. 52. The mammalian cell of claim 51, wherein the serine recombinase comprises a transduction efficiency of at least about 50%.

60. 52. The mammalian cell of claim 51, wherein the serine recombinase is capable of targeting a gene containing a catalase domain or a synthase domain.

61. 61. The mammalian cell of claim 60, wherein the catalytic enzyme is a manganese catalytic enzyme.

62. 62. The mammalian cell of any one of claims 60 to 61, wherein the synthase is queosynthase.

63. 63. The mammalian cell of any one of claims 60 to 62, wherein the serine recombinase is capable of targeting a gene containing a DUF4244 Pfam domain.

64. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214.

65. A eukaryotic cell comprising a serine recombinase having at least about 80% sequence identity to SEQ ID NO:

21.

66. A eukaryotic cell comprising a serine recombinase having at least about 80% sequence identity to SEQ ID NO:

22.

67. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO:

23.

68. A eukaryotic cell comprising a serine recombinase having at least about 80% sequence identity to SEQ ID NO:

24.

69. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO: 1848.

70. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO:7111.

71. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO:7115.

72. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO:7131.

73. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO:7136.

74. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO:7139.

75. A eukaryotic cell comprising a serine recombinase comprising at least about 80% sequence identity to SEQ ID NO:7140.

76. 76. A eukaryotic cell according to any one of claims 64 to 75, which is a mammalian cell.

77. 76. A eukaryotic cell according to any one of claims 64 to 75, which is a human cell.

78. A vector comprising: a) a nucleic acid encoding a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214; b) one or more regulatory elements.

79. 79. The vector of claim 78, wherein the one or more regulatory elements comprise a promoter, enhancer, intron, microRNA, linker, splicing element, or polyA signal.

80. 80. The vector of claim 79, wherein the promoter is selected from a constitutive promoter, an inducible promoter, a minipromoter, or derivatives thereof.

81. 80. The vector of claim 79, wherein the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, polH, EM7, OpIE1, and derivatives thereof.

82. 1. A vector comprising a nucleic acid encoding a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214, wherein the vector is selected from the group consisting of a plasmid, a nanoplasmid, a phagemid, a phage derivative, a bacmid, a bacterial artificial chromosome (BAC), a minicircle, a doggybone, a yeast artificial chromosome (YAC), and a cosmid.

83. 1. A method for gene editing, comprising: a) providing or identifying a first binding site sequence in a host genome; b) providing a host cell with nucleic acid comprising a donor polynucleotide and a second binding site sequence; c) contacting the host cell with a serine recombinase comprising at least about 80% sequence identity to any one of SEQ ID NOs: 21-7060, 7105-7142, and 7211-7214, or a nucleic acid encoding the serine recombinase; The method, wherein the first binding site sequence and the second binding site sequence are capable of recombination.

84. 84. The method of claim 83, wherein the first binding site sequence is endogenous in the host genome.

85. 84. The method of claim 83, wherein the first binding site sequence is provided using viral delivery.

86. 84. The method of claim 83, wherein the first binding site sequence is provided using a transpotase.

87. 84. The method of claim 83, wherein the first binding site sequence is provided using a nuclease.

88. 88. The method of claim 87, wherein the nuclease is a double-stranded nuclease.

89. 88. The method of claim 87, wherein the nuclease is a type II CRISPR endonuclease.

90. 88. The method of Claim 87, wherein the nuclease is a V-type CRISPR endonuclease.

91. 88. The method of Claim 87, wherein the nuclease is Cas9.

92. 77. The method of claim 76, wherein the first binding site sequence is provided using a reverse transcriptase enzyme.

93. 93. The method of any one of claims 83 to 92, wherein the second binding site sequence is 5' to the donor polynucleotide.

94. 94. The method of any one of claims 83 to 93, wherein the first binding site sequence is a bacterial genome recombination sequence (attB).

95. 95. The method of any one of claims 83 to 94, wherein the first binding site sequence is a phage genomic recombination sequence (attP).

96. 96. The method of any one of claims 83 to 95, wherein the second binding site sequence is a bacterial genome recombination sequence (attB).

97. 97. The method of any one of claims 83 to 96, wherein the second binding site sequence is a phage genomic recombination sequence (attP).

98. 98. The method of any one of claims 94 to 97, wherein the attB sequence comprises from about 20 to about 500 nucleotides.

99. 99. The method of any one of claims 95 to 98, wherein the attP sequence comprises from about 20 to about 500 nucleotides.

100. The attB sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7151, 7155, 7159, 7163, 7167, 7171, 7175, 7179, 7188 to 7200, and 7206 100. The method of any one of claims 94-99, comprising at least about 80% sequence identity with any one of the following: -7210, 7215, 7220, 7225, 7226, 7233, 7238, 7243, 7248, 7253, 7258, 7263, 7264, 7271, 7277, 7282, 7287, 7292, 7297, 7302, 7307, 7312, 7317, 7322, 7327, 7332, 7337, 7342, 7347, 7352, 7357, 7362, 7367, 7372, 7377, 7382, 7387, 7392, 7397, and 7402.

101. 101. The method of any one of claims 94-100, wherein the attB sequence comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 1, 5, 9, and 13.

102. The attP sequence is selected from the group consisting of SEQ ID NOs: 1, 2, 5, 6, 9, 10, 13, 14, 7152, 7156, 7160, 7164, 7168, 7172, 7176, 7180, 7183 to 7187, 7201 to 7205, 7217, 7222, 7228, 7229, 7235, 7240, 7245, 7250, 7255, 7260, 7266, 7267, 7273, 7279, 7284, 7285, 7286, 7287, 7288, 7289, 7290, 7291, 7292, 7293, 7294, 7295, 7296, 7297, 7298, 7300, 7301, 7302, 7303, 7304, 7305, 7306, 7307, 7308, 7309, 7310, 7311, 7312, 7313, 7314, 7315, 7316, 7317, 7318, 7319, 7320, 7321, 7322, 7323, 7324, 7325, 7326, 7327, 7328, 7329, 7400, 7401, 7402, 7403, 7404, 7405, 740 89, 7294, 7299, 7304, 7309, 7314, 7319, 7324, 7329, 7334, 7339, 7344, 7349, 7354, 7359, 7364, 7369, 7374, 7379, 7384, 7389, 7394, 7399, and 7404.

103. 103. The method of any one of claims 95 to 102, wherein the attP sequence has at least about 80% sequence identity to any one of SEQ ID NOs: 2, 6, 10 and 14.

104. 104. The method of any one of Claims 83-103, wherein the nucleic acid comprising the donor polynucleotide and the second binding site sequence is delivered by a plasmid, nanoplasmid, phagemid, phage derivative, virus, bacmid, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid.

105. 107. The method of any one of claims 83 to 106, wherein the nucleic acid encoding the serine recombinase is delivered by a plasmid, nanoplasmid, phagemid, phage derivative, virus, bactide, bacterial artificial chromosome (BAC), minicircle, doggybone, yeast artificial chromosome (YAC), or cosmid.

106. 106. The method of any one of claims 104-105, wherein the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus.

107. The AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AA V11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65 , AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK0 3, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, or AAV-HSC16, or a derivative thereof.

108. 107. The method of claim 106, wherein the herpes virus is HSV-1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

109. 109. The method of any one of claims 83-108, wherein the donor polynucleotide comprises a size of at least about 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 6 kb, 7 kb, 8 kb, 9 kb, 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, or greater than 120 kb.

110. 110. The method of any one of claims 83-109, wherein the donor polynucleotide encodes a therapeutic agent, a reporter, or a marker.

111. 111. The method of claim 110, wherein the reporter comprises a fluorescent protein.

112. 112. The method of claim 111, wherein the fluorescent protein is GFP, EBFP, EBFP2, azurite, mKalamal, ECFP, Cerulean, CyPet, YFP, citrine, venus, YPet, RFP, CFP, or a derivative thereof.

113. The method of claim 110, wherein the reporter is acetohydroxyacid synthase (AHAS), alkaline phosphatase (AP), beta-galactosidase (LacZ), beta-glucuronidase (GUS), chloramphenicol acetyltransferase (CAT), horseradish peroxidase (HRP), luciferase (Luc), nopaline synthase (NOS), octopine synthase (OCS), luciferase, or a derivative thereof.

114. 114. The method of any one of claims 110 to 113, wherein the marker is an antibiotic resistance marker.

115. 115. The method of claim 114, wherein the antibiotic resistance marker is kanamycin, spectinomycin, streptomycin, ampicillin, carbenicillin, bleomycin, erythromycin, polymyxin B, tetracycline, chloramphenicol, neomycin, zeocin, or a derivative thereof.

116. The method of any one of claims 110 to 113, wherein the marker is a cell surface marker.