High efficiency integrases for gene editing

Mutated Bxb1 integrases improve integration efficiency and reduce DNA damage in mammalian cells, enabling precise and scalable genetic engineering through PASTE systems.

WO2025193915A1PCT designated stage Publication Date: 2025-09-18MASSACHUSETTS INST OF TECH +2
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/019716
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-20
Filing Date
2025-03-13
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing integrases face inefficiencies in integration and recombination, leading to chromosomal rearrangements and DNA damage in mammalian cells, limiting their clinical application.

Method used

Development of high-efficiency integrases with specific mutations, such as Bxb1 integrases, that enhance integration and recombination efficiency by targeting site-specific integration sites in the human genome, using systems like PASTE for precise genetic engineering.

Benefits of technology

The mutated integrases achieve efficient and precise integration of exogenous nucleic acids into mammalian cells, reducing DNA damage and enabling scalable, multiplexed genetic engineering without relying on DNA damage responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000024_0001
    Figure IMGF000024_0001
  • Figure IMGF000025_0001
    Figure IMGF000025_0001
  • Figure IMGF000026_0001
    Figure IMGF000026_0001
Patent Text Reader

Abstract

This disclosure provides novel integrases for site-specific genetic engineering. Also provided are systems, methods, and compositions for site-specific genetic engineering using Programmable Addition via Site-Specific Targeting Elements (PASTE) with the novel integrases.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket: 761259: 083474-041PC HIGH EFFICIENCY INTEGRASES FOR GENE EDITING CROSS REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Patent Application Serial Nos.63 / 565,408, 63 / 672,398, and 63 / 722,658 filed respectively on March 14, 2024, July 17, 2024, and November 20, 2024. The entire content of the above-referenced patent applications is incorporated by reference in its entirety herein. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with Government support under Grant No. EB031957 awarded by the National Institutes of Health. The Government has certain rights in the invention. SEQUENCE LISTING

[0003] The instant application contains a Sequence Listing which has been submitted electronically in XML file format and is hereby incorporated by reference in its entirety. Said XML copy, created on March 11, 2025, is named 761259_083474-041PC_SL.xml and is 75,848,705 bytes in size. BACKGROUND

[0004] Integrases provide efficient modes for non-programmable genome integration. Site- specific integrases, such as large serine phage integrases, efficiently integrate their DNA cargo into sequence-defined landing sites that are about 30-50 nucleotides long and can be used to insert therapeutic transgenes at naturally occurring pseudo-sites in the human genome in pre-clinical models (Brown et al. Methods.2011.53: 372-379; Calos. Curr. Gene Ther. 2006.6: 633-645). Targeted integration can also be achieved by a two-step approach involving prior insertion of integrase landing sites at a desired location using homology- directed repair (HER) (Mulholland et al. Nucleic Acids Res.2015.43: e112). However, the inefficiency of two-step integration with HDR and the risks associated with double-strand breaks (DSBs) have limited this approach in mammalian cells. Furthermore, a major issue limiting clinical application of certain integrases, such as phiC31, is that chromosomal rearrangements between pseudo-sites can occur, leading to a significant DNA damageAttorney Docket: 761259: 083474-041PC response (Ehrhardt et al. Hum. Gene Ther.2006.17: 1077–1094; Liu et al. Gene Ther.2006. 13: 1188–1190).

[0005] Accordingly, there exists a need for novel integrases with increased efficiency of integration and / or recombination. SUMMARY

[0006] In one aspect, the disclosures provides a high efficiency integrase or fragment thereof comprising an amino acid sequence with one or more mutations relative to wild-type Bxb1 integrase amino acid sequence, wherein the amino acid sequence is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200-69221 and 69226-69340.

[0007] In certain embodiments, the one or more mutations are located at one or more amino acids at positions 7, 22, 23, 25, 58, 67, 71, 72, 79, 101, 110, 115, 129, 141, 146, 166, 182, 189, 194, 230, 238, 240, 245, 247, 253, 258, 271, 275, 292, 318, 345, 375, 376, 381, 385, 422, 439, 459, 460, and 482 relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO: 11.

[0008] In certain embodiments, the one or more mutations comprise 7L, 22V, 23K, 25A, 58T, 67C, 71G, 72A, 79K, 79R, 101G, 110M, 110N, 115K, 129W, 141D, 146A, 166A, 166C, 166E, 166Q, 166R, 182M, 189G, 189Q, 194A, 194C, 194F, 194G, 194M, 194W, 230R, 238M, 240C, 240G, 240H, 240L, 240M, 240Q, 240S, 240T, 240V, 245A, 247V, 253K, 258A, 271N, 275A, 275C, 275D, 275E, 275G, 275H, 275K, 275M, 275N, 275P, 275Q, 275R, 275S, 275Y, 292A, 292Q, 292S, 318F, 345H, 375A, 375C, 375D, 375E, 375F, 375G, 375Q, 375W, 375Y, 376Y, 381Q, 385A, 422P, 439A, 439C, 439D, 439E, 439G, 439H, 439I, 439K, 439N, 439P, 439Q, 439R, 439S, 439T, 439V, 459G, 460G, and / or 482L relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO: 11.

[0009] In certain embodiments, the one or more mutations are located at one or more amino acids at positions: 194, 166, and 110; 194, 166, 240, and 110; 194, 166, and 240; 194, 166, 240, 459, and 110; 194, 166, 240, and 459; 194, 166, 240, 375, and 110; 194, 166, and 275; 194, 166, 275, and 240; 194, 166, 275, 240, 375, and 110; 194, 166, 275, 240, and 375; 194, 166, 275, 375, and 110; 194, 166, 275, and 375; 194, 166, 459, and 110; 194, 166, 292, and 110; 194, 166, 292, 240, and 110; 194, 166, 292, and 240; 194, 166, 292, 240, 375, and 110; 194, 166, 292, 275, and 110; 194, 166, 292, and 275; 194, 166, 292, 275, 240, and 110; 194, 166, 292, 275, and 240; 194, 166, 292, 275, 240, and 375; 194, 166, 292, 275, 375, and 110; 166, 292, 275, and 240; 194, 166, and 240; 194, 166, 240, 459, and 110; 194, 166, 240, 375,Attorney Docket: 761259: 083474-041PC and 110; 194, 166, and 275; 194, 166, 275, and 375; 194, 166, 275, 375, and 110; 194, 166, 292, and 240; or 194, 166, 292, 240, 375, and 110, relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO: 11.

[0010] In certain embodiments, the one or more mutations comprise: 194G, 166R, and 110N; 194G, 166R, 240C, and 110N; 194G, 166R, and 240C; 194G, 166R, 240C, 459G, and 110N; 194G, 166R, 240C, and 459G; 194G, 166R, 240C, 375C, and 110N; 194G, 166R, and 275E; 194G, 166R, 275E, and 240C; 194G, 166R, 275E, 240C, 375C, and 110N; 194G, 166R, 275E, 240C, and 375C; 194G, 166R, 275E, 375C, and 110N; 194G, 166R, 275E, and 375C; 194G, 166R, 459G, and 110N; 194G, 166R, 292S, and 110N; 194G, 166R, 292S, 240C, and 110N; 194G, 166R, 292S, and 240C; 194G, 166R, 292S, 240C, 375C, and 110N; 194G, 166R, 292S, 275E, and 110N; 194G, 166R, 292S, and 275E; 194G, 166R, 292S, 275E, 240C, and 110N; 194G, 166R, 292S, 275E, and 240C; 194G, 166R, 292S, 275E, 240C, and 375C; 194G, 166R, 292S, 275E, 375C, and 110N; 166R, 292S, 275E, and 240C; 194G, 166R, and 240C; 194G, 166R, 240C, 459G, and 110N; 194G, 166R, 240C, 375C, and 110N; 194G, 166R, and 275E; 194G, 166R, 275E, and 375C; 194G, 166R, 275E, 375C, and 110N; 194G, 166R, 292S, and 240C; or 194G, 166R, 292S, 240C, 375C, and 110N, relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO: 11.

[0011] In certain embodiments, the high efficiency integrase or fragment thereof is capable of integrase, recombinase, or transposase activity.

[0012] In certain embodiments, the high efficiency integrase or fragment thereof is capable of binding to integration sites or pseudosites in a human genome sequence.

[0013] In certain embodiments, each integration site is between 12 and 60 nucleotides in length.

[0014] In certain embodiments, each integration site is between 18 and 50 nucleotides in length.

[0015] In certain embodiments, the integration sites comprise an attB and attP pair of integration sites.

[0016] In certain embodiments, the high efficiency integrase or fragment thereof binds to an attB integration site comprising at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, or 100% identity to an amino acid sequence set forth in any one of SEQ ID NOs: 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 55, and 69222.

[0017] In certain embodiments, the high efficiency integrase or fragment thereof binds to an attP integration site comprising at least 80% identity, at least 85% identity, at least 90%Attorney Docket: 761259: 083474-041PC identity, at least 95% identity, or 100% identity to an amino acid sequence set forth in any one of SEQ ID NOs: SEQ ID NOs: 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, and 69223.

[0018] In another aspect, the disclosures provides a polynucleotide comprising a nucleic acid sequence.

[0019] In another aspect, the disclosures provides a vector comprising a nucleic acid.

[0020] In another aspect, the disclosures provides a host cell comprising a vector.

[0021] In another aspect, the disclosures provides a system capable of site-specifically integrating an exogenous nucleic acid into a mammalian cell genome at a desired target site, wherein the system comprises: (a) a nucleic acid encoding a DNA binding nickase domain linked to a reverse transcriptase (RT) domain; (b) a nucleic acid encoding a guide RNA (gRNA) comprising: i. a primer binding sequence, ii. a sequence complementary to one strand of an integration recognition sequence, and iii. a target binding sequence, (c) a nucleic acid encoding a serine integration enzyme; and (d) the exogenous nucleic acid linked to a sequence that is a cognate of the integration recognition sequence, wherein the system site specifically integrates the exogenous nucleic acid into the mammalian cell genome at the desired target site, and wherein the serine integration enzyme comprises an amino acid sequence at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200-69221 and 69226-69340.

[0022] In certain embodiments, the integration recognition sequence is an attB or attP integration recognition sequence.

[0023] In another aspect, the disclosures provides a method of site-specific integration of a nucleic acid into a genome of a cell, the method comprising: (a) incorporating an integration site at a desired location in the genome by introducing into the cell: i. a DNA binding nuclease linked to a reverse transcriptase domain, wherein the DNA binding nuclease comprises a nickase activity; and ii. a guide RNA (gRNA) comprising a primer binding sequence linked to an integration recognition sequence, wherein the gRNA interacts with the DNA binding nuclease and targets the desired location in the genome, wherein the DNA binding nuclease nicks a strand of the genome and the reverse transcriptase domain incorporates the integration recognition sequence of the gRNA into the nicked site, thereby providing the integration site at the desired location of the genome; and (b) integrating the nucleic acid into the genome by introducing into the cell: i. a DNA or RNA strand comprising the nucleic acid linked to a sequence that is complementary or associated to theAttorney Docket: 761259: 083474-041PC integration site; and ii. an integration enzyme, wherein the integration enzyme incorporates the nucleic acid into the genome at the integration site by integration, recombination, or reverse transcription of the sequence that is complementary or associated to the integration site, thereby introducing the nucleic acid into the desired location of the cell genome of the cell, wherein the integration enzyme comprises an amino acid sequence at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200-69221 and 69226-69340.

[0024] In certain embodiments, the integration recognition sequence is an attB or attP integration recognition sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Aspects, features, benefits and advantages of the embodiments described herein will be apparent with regard to the following description, appended claims, and accompanying drawings.

[0026] FIG.1 shows a schematic diagram of a concept of Programmable Addition via Site- Specific Targeting Elements (PASTE).

[0027] FIG.2 shows a schematic representation of using Bxb1 to integrate a nucleic acid into the genome.

[0028] FIG.3 shows the percent integration of GFP or Gluc into the attB locus using Bxb1 Programmable Addition via Site-Specific Targeting Elements (PASTE).

[0029] FIG.4 shows the percent editing of various HEK3 targeting pegRNA Programmable Addition via Site-Specific Targeting Elements (PASTE).

[0030] FIG.5A shows a schematic of the integrase discovery pipeline from bacterial and metagenomic sequences (FIG.5A). FIG.5B shows the phylogenetic tree of discovered integrases showing distinct subfamilies. FIG.5C shows the phylogenetic tree of discovered integrases showing distinct subfamilies.

[0031] FIG.6A – FIG.6I show the activity of several integrases. FIG.6A shows an Integrase integration activity screen using reporters in HEK293FT cells compared to BxbINT and phiC31a. FIG.6B shows PASTE integration activity with the most active integrases compared to BxbINT. FIG.6C shows a characterization of integrase integration activity with truncated attachment sites using reporters in HEK293FT cells. FIG.6D shows PASTE integration activity with BceINT and BcyINT with truncated attachment sites compared to BxbINT. FIG.6E shows PASTE integration activity with SscINT and SacINT with truncatedAttorney Docket: 761259: 083474-041PC attachment sites compared to BxbINT. FIG.6F shows optimization BceINT and SacINT PASTE constructs via protein fusions for different sized attachment sites compared to BxbINT-based PASTE for EGFP integration at the ACTB locus. FIG.6G shows BceINT and INT2 PASTE protein constructs compared to BxbINT for EGFP integration at the ACTB locus. FIG.6H shows integration of EGFP at different endogenous genes for PASTE with either BceINT or BxbINT. FIG.6I shows PASTE integration activity with various integrases of EGFP at the ACTB locus.

[0032] FIG.7A – FIG.7B show the activity of several integrases. FIG.7A shows an Integrase integration activity screen using reporters in HEK293FT cells compared to BxbINT, SacINTd, and BceINT. FIG.7B shows PASTE integration activity with the integrase N352807_16_14 compared to BxbINT. N352807_16_14 was tested with various truncations of the attB sequence.

[0033] FIG.8A – FIG.8B show the percent integration of attB locus into the HEK293FT genome according to embodiments of the present teachings.

[0034] FIG.9 shows a schematic representation of using Bxb1 integrase and one plasmid with an attB site and a second plasmid with an attP site for recombination to make a recombined plasmid product.

[0035] FIG.10A shows a schematic describing the EVOLVE-Pro method. Proteins of interest go through iterative rounds of low-N screening. A foundational PLM generates embeddings for all mutants of a protein and the average embedding by pooling across all residues is used as input for the top layer model. Each mutant’s activity is experimentally determined and used to train a domain expert top layer model with PLM embedding as input. The top layer model then nominates the top-N mutants for the next round of testing and the weights are updated iteratively in an active learning format. Figure discloses SEQ ID NOS 69341-69343, 69344, 69342-69343 and 69341, respectively, in order of appearance. FIG. 10B shows the benchmarking of foundational models across a panel of 12 comprehensive deep mutational scanning (DMS) datasets. Each point is a unique protein and its DMS data. ESM2-15B has the highest average percent success in high activity variants prediction. FIG. 10C shows the comparison between EVOLVE-Pro in active learning format, in zero-shot pretraining format, and an existing zero-shot prediction method using protein language model across 12 DMS datasets. Each point is a unique protein using its DMS data. FIG.10D shows the performance over 10 rounds of EVOLVE-Pro with 16 mutants per round, compared to two different non-language model encoding schemes (One-hot encoding and integer encoding). Model performance is benchmarked on four datasets and compared to zero-shotAttorney Docket: 761259: 083474-041PC ESM2 nomination success rate and background random sampling. Error bar represents the standard deviation for n=10 random simulations. FIG.10E shows the engineering of REGN10987 over five rounds of EVOLVE-Pro. Data shows cumulative top 10 mutants’ fold improvement over wild-type binding affinity to the target antigen across 5 evolution rounds. Percentages show the percent of mutants that have higher activity than wild-type REGN10987 each round. FIG.10F shows the mapping of the top mutations on the structure of REGN10987 (PDB: 6XDG).

[0036] FIG.11 shows a table summarizing the parameters grid search for EVOLVE-Pro application.

[0037] FIG.12 shows a table describing 12 DMS datasets for EVOKV-Pro application.

[0038] FIG.13A shows a summary of parameter grid searches for EVOLVE-Pro with an ESM-215B foundational model, showing the random forest regressor combined with the top 10 active learning selection strategy returned the highest average binary top fitness success rate across 12 DMS datasets. FIG.13B shows the optimized EVOLVE-Pro model with n=16 mutants per round, showing both a higher median protein activity score and max protein activity score as the model progresses into later rounds of evolution, showing the utility of active learning. Error bars represent standard deviation with n=10 simulations. FIG.13C shows a comparison of the number of nominated mutants from n=10 to n=100 per round on the impact of EVOLVE-Pro evolution. Each line graph depicts the percent high fitness for one of the 12 DMS datasets and the error bar represents the standard error of the mean for 10 random simulations.

[0039] FIG.14A and FIG.14B shows Evolve-Pro model characterization using H3N2 hemagglutinin, GPCT, HIV envelope protein, infA, TP53, AsCas12f, PafA, and SARS- COVID2 spike protein. The performance is evaluated over 10 rounds of EVOLVE-Pro with 16 mutants per round, compared to two different non-language model encoding schemes (One-hot encoding, integer encoding). Model performance is benchmarked on 8 additional datasets and compared to both the zero-shot ESM2 nomination success rate and background random sampling (1). Error bars represent the standard deviations for 10 random simulations.

[0040] FIG.15A shows a schematic of the evolution strategy for evolving the Bxb1 serine integrase from the Mycobacteriophage. FIG.15B shows the engineering of the Bxb1 integrase over 8 rounds of EVOLVE-Pro. Data shows cumulative top 11 mutants from current and preceding rounds, as measured by fold improvement of plasmid integration over wild-type. FIG.15C shows the performance of top Bxb1 mutants for plasmid recombination with low Bxb1 expression in Hela cell. A two-sided Student’s t-test was run between WT andAttorney Docket: 761259: 083474-041PC each evolved Bxb1 integrase (***, p<0.001, ****, p<0.0001). Fold change over wild-type Bxb1 is shown for the best mutant. Error bars represent standard deviation with n=3 biological replicates. FIG.15D shows the validation of epBxb1 with PASTE at four genomic sites across human and mice genomes. A two-sided Student’s t-test was run between WT and each evolved Bxb1 integrase (*, p<0.05, **, p<0.01). Fold change over wild-type Bxb1 integrase is shown for each genomic locus. Error bars represent standard deviation with n=3 biological replicates. FIG.15E shows the mapping of the top mutations on the AlphaFold3 model of the Bxb1 monomer bound to DNA. Bxb1 forms a tetrameric synaptic complex during recombination between two DNA molecules. The active site is indicated by a red circle. FIG.15F shows a heatmap showing most common Bxb1 mutations explored by EVOLVE-Pro over rounds of evolution. Any position explored more than once is shown on a cumulative basis across rounds. FIG.15G shows a scatter plot comparing the predicted ESM-2 protein fitness score versus experimentally measured bxb1 integration efficiency (scaled) across evolution rounds. The correlation and linear regression line are shown in the plot. FIG.15H and FIG.15I show the comparison of the Bxb1 latent space with either predicted ESM-2 protein fitness (masked marginal score) or EVOLVE-Pro protein activity fold improvement. FIG.15J shows a kernel density estimate of protein fitness as predicted by ESM-2 versus protein function as predicted by EVOLVE-Pro. The correlation and linear regression line are shown in red and the R square of correlation is reported. FIG.15K shows the EVOLVEpro’s mutational trajectory from round 1 to round 7 on Bxb1.

[0041] FIG.16A shows individual mutant’s fold improvement in plasmid recombination efficiency in HEK293FT cells across 8 rounds of evolution. FIG.16B shows titration of input Bxb1 amount for WT, N194G, and T166R (epBxb1) mutants in a plasmid recombination assay in Hela cells. Error bars represent standard deviation of three biological replicates. FIG.16C shows comparison of epBxb1 with wild-type Bxb1 on integration of cargo into TRAC genomic locus in HEK293FT cells. Error bars represent standard deviation of three biological replicates. DETAILED DESCRIPTION

[0042] It will be appreciated that for clarity, the following discussion will describe various aspects of embodiments of the applicant’s teachings. It should be noted that the specific embodiments are not intended as an exhaustive description or as a limitation to the broader aspects discussed herein. One aspect described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced with any otherAttorney Docket: 761259: 083474-041PC embodiment(s). Reference throughout this specification to "one embodiment," "an embodiment," "an example embodiment," means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of the phrases "in one embodiment," "in an embodiment," or "an example embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular feature, structures or characteristics may be combined in any suitable manner, as would be apparent to a person skilled in the art from this disclosure, in one or more embodiments. Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the disclosure. For example, in the appended claims, any of the claimed embodiments can be used in any combination. General Definitions

[0043] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Definitions of common terms and techniques in molecular biology may be found in Molecular Cloning: A Laboratory Manual, 2nd edition (1989) (Sambrook, Fritsch, and Maniatis); Molecular Cloning: A Laboratory Manual, 4th edition (2012) (Green and Sambrook); Current Protocols in Molecular Biology (1987) (F.M. Ausubel et al. eds.); the series Methods in Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (1995) (M.J. MacPherson, B.D. Hames, and G.R. Taylor eds.): Antibodies, A Laboratory Manual (1988) (Harlow and Lane, eds.): Antibodies A Laboratory Manual, 2nd edition 2013 (E.A. Greenfield ed.); Animal Cell Culture (1987) (R.I. Freshney, ed.); Benjamin Lewin, Genes IX, published by Jones and Bartlet, 2008 (ISBN 0763752223); Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994 (ISBN 0632021829); Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 9780471185710); Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, N.Y.1994), March, Advanced Organic Chemistry Reactions, Mechanisms and Structure 4th ed., John Wiley & Sons (New York, N.Y.1992); and Marten H. Hofker and Jan van Deursen, Transgenic Mouse Methods and Protocols, 2nd edition (2011).Attorney Docket: 761259: 083474-041PC

[0044] As used herein, the singular forms "a," "an," and "the" include both singular and plural forms unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a plurality of such cells.

[0045] As used herein, the term "optional" or "optionally" means that the subsequent described event, circumstance or substituent may or may not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.

[0046] The recitation of numerical ranges by endpoints includes all numbers and fractions subsumed within the respective ranges, as well as the recited endpoints.

[0047] As used herein, the term "about" or "approximately" refers to a measurable value such as a parameter, an amount, a temporal duration, and the like, are meant to encompass variations of and from the specified value, such as variations of + / -10% or less, + / -5% or less, + / -1% or less, + / -0.5% or less, and + / -0.1% or less of and from the specified value, insofar such variations are appropriate to perform in the disclosure. It is to be understood that the value to which the modifier "about" or "approximately" refers is itself also specifically disclosed.

[0048] It is noted that all publications and references cited herein are expressly incorporated herein by reference in their entirety. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present disclosure is not entitled to antedate such publication. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed. Overview

[0049] The embodiments disclosed herein provide non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering using Programmable Addition via Site-Specific Targeting Elements (PASTE). A schematic diagram illustrating the concept of PASTE is shown in FIG.1. As discussed in more details below, the PASTE comprises the addition of an integration site into a target genome followed by the insertion of one or more genes of interest or one or more nucleic acid sequences of interest at the site. This process can be done as one or more reactions into a cell. The addition of the integration site into the target genome is done using gene editing technologies that include for example, without limitation, prime editing, recombinant adeno-associated virus (rAAV)-mediated nucleic acid integration, transcription activator-like effector nucleases (TALENS), and zinc finger nucleases (ZFNs). The integration of the transgene at theAttorney Docket: 761259: 083474-041PC integration site is done using integrase technologies that include for example, without limitation, integrases, recombinases, and reverse transcriptases. The necessary components for the site-specific genetic engineering disclosed herein comprise at least one or more nucleases, one or more guide RNA (gRNA), one or more integration enzymes, and one or more sequences that are complementary or associated to the integration site and linked to the one or more genes of interest or one or more nucleic acid sequences of interest to be inserted into the cell genome.

[0050] An advantage of the non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering disclosed herein is programmable insertion of large elements without reliance on DNA damage responses.

[0051] Another advantage of the non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering disclosed herein is facile multiplexing, enabling programmable insertion at multiple sites.

[0052] Yet another advantage of the non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering disclosed herein is scalable production and delivery through minicircle templates. Prime Editing

[0053] The present disclosure provides non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering using gene editing technologies such as prime editing to add an integration site into a target genome. Prime editing will be discussed in more detail below.

[0054] Prime editing is a versatile and precise genome editing method that directly writes new genetic information into a specified DNA site. Such method is explained fully in the literature. See, e.g., Anzalone, A.V., et al. “Search-and-replace genome editing without double-strand breaks or donor DNA,” Nature 576, 149–157 (2019). Prime editing uses a catalytically-impaired Cas9 endonuclease that is fused to an engineered reverse transcriptase (RT) and programmed with a prime-editing guide RNA (pegRNA). The skilled person in the art would appreciate that the pegRNA both specifies the target site and encodes the desired edit. The catalytically-impaired Cas9 endonuclease also comprises a Cas9 nickase that is fused to the reverse transcriptase. During genetic editing, the Cas9 nickase part of the protein is guided to the DNA target site by the pegRNA. The reverse transcriptase domain then uses the pegRNA to template reverse transcription of the desired edit, directly polymerizing DNA onto the nicked target DNA strand. The edited DNA strand replaces the original DNA strand,Attorney Docket: 761259: 083474-041PC creating a heteroduplex containing one edited strand and one unedited strand. Afterward, the prime editor (PE) guides resolution of the heteroduplex to favor copying the edit onto the unedited strand, completing the process.

[0055] The prime editors refer to a Moloney Murine Leukemia Virus (M-MLV) reverse transcriptase (RT) fused to a Cas9 H840A nickase. Fusing the RT to the C-terminus of the Cas9 nickase may result in higher editing efficiency. Such a complex is called PE1. The Cas9(H840A) can also be linked to a non-M-MLV reverse transcriptase such as a AMV-RT or XRT (Cas9(H840A)-AMV-RT or XRT). In some embodiments, Cas 9(H840A) can be replaced with Cas12a / b or Cas9(D10A). A Cas9 (wild type), Cas9(H840A), Cas9(D10A) or Cas 12a / b nickase fused to a pentamutant of M-MLV RT (D200N / L603W / T330P / T306K / W313F), having up to about 45-fold higher efficiency is called PE2. In some embodiments, the M-MLV RT comprise one or more of the mutations Y8H, P51L, S56A, S67R, E69K, V129P, L139P, T197A, H204R, V223H, T246E, N249D, E286R, Q291I, E302K, E302R, F309N, M320L, P330E, L435G, L435R, N454K, D524A, D524G, D524N, E562Q, D583N, H594Q, E607K, D653N, and L671P. In some embodiments, the reverse transcriptase can also be a wild-type or modified transcription xenopolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), Feline Immunodeficiency Virus reverse transcriptase (FIV- RT), FeLV-RT (Feline leukemia virus reverse transcriptase), HIV-RT (Human Immunodeficiency Virus reverse transcriptase), or Eubacterium rectale maturase RT (MarathonRT). PE3 involves nicking the non-edited strand, potentially causing the cell to remake that strand using the edited strand as the template to induce HR. The nicking of the non-edited strand can involve the use of a nicking guide RNA (ngRNA).

[0056] In certain embodiments, the reverse transcriptase contains a stabilization domain. In certain embodiments, the stabilization domain comprises the DNA-binding Sto7d protein from Sulfolobus tokodaii or the DNA-binding Sso7d protein. The DNA-binding proteins improves processivity and resistance to inhibitors of M-MuLV reverse transcriptase. The DNA-binding Sto7d protein from Sulfolobus tokodaii or the DNA-binding Sso7d protein are described in further detail in Oscorbin et al. (FEBS Letters.594(24): 4338-4356.2020), incorporated herein by reference.

[0057] Nicking the non-edited strand can increase editing efficiency. For example, nicking the non-edited strand can increase editing efficiency by about 1.1 fold, about 1.3 fold, about 1.5 fold, about 1.7 fold, about 1.9 fold, about 2.1 fold, about 2.3 fold, about 2.5 fold, about 2.7 fold, about 2.9 fold, about 3.1 fold, about 3.3 fold, about 3.5 fold, about 3.7 fold, aboutAttorney Docket: 761259: 083474-041PC 3.9 fold, 4.1 fold, about 4.3 fold, about 4.5 fold, about 4.7 fold, about 4.9 fold, or any range that is formed from any two of those values as endpoints.

[0058] Although the optimal nicking position varies depending on the genomic site, nicks positioned 3′ of the edit about 40–90 bp from the pegRNA-induced nick can generally increase editing efficiency without excess indel formation. The prime editing practice allows starting with non-edited strand nicks about 50 bp from the pegRNA-mediated nick, and testing alternative nick locations if indel frequencies exceed acceptable levels.

[0059] As used herein, the term “guide RNA” (gRNA) and the like refer to an RNA that guides the insertion or deletion of one or more genes of interest or one or more nucleic acid sequences of interest into a target genome. The gRNA can also refer to a prime editing guide RNA (pegRNA), a nicking guide RNA (ngRNA), and a single guide RNA (sgRNA). In some embodiments, the term “gRNA molecule” refers to a nucleic acid encoding a gRNA. In some embodiments, the gRNA molecule is naturally occurring. In some embodiments, a gRNA molecule is non-naturally occurring. In some embodiments, a gRNA molecule is a synthetic gRNA molecule. A gRNA can target a nuclease or a nickase such as Cas9, Cas 12a / b Cas9(H840A) or Cas9 (D10A) molecule to a target nucleic acid or sequence in a genome. In some embodiments, the gRNA can bind to a DNA nickase bound to a reverse transcriptase domain. A “modified gRNA,” as used herein, refers to a gRNA molecule that has an improved half-life after being introduced into a cell as compared to a non-modified gRNA molecule after being introduced into a cell. In some embodiments, the guide RNA can facilitate the addition of the insertion site sequence for recognition by integrases, transposases, or recombinases.

[0060] As used herein, the term “prime-editing guide RNA” (pegRNA) and the like refer to an extended single guide RNA (sgRNA) comprising a primer binding site (PBS), a reverse transcriptase (RT) template sequence, and an integration site sequence that can be recognized by recombinases, integrases, or transposases. For example, the PBS can have a length of at least about 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, or more nt. For example, the PBS can have a length of about 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, or any range that is formed from any two of those values as endpoints. For example, the RT template sequence can have a length of at least about 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, 31 nt, 32 nt, 33 nt, 34 nt, 35Attorney Docket: 761259: 083474-041PC nt, 36 nt, 37 nt, 38 nt, 39 nt, 40 nt, 41 nt, 42 nt, 43 nt, 44 nt, 45 nt, 46 nt, 47 nt, 48 nt, 49 nt, 50 nt, or more nt. For example, the RT template sequence can have a length of about 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, 31 nt, 32 nt, 33 nt, 34 nt, 35 nt, 36 nt, 37 nt, 38 nt, 39 nt, 40 nt, 41 nt, 42 nt, 43 nt, 44 nt, 45 nt, 46 nt, 47 nt, 48 nt, 49 nt, 50 nt, or any range that is formed from any two of those values as endpoints.

[0061] During genome editing, the primer binding site allows the 3’ end of the nicked DNA strand to hybridize to the pegRNA, while the RT template serves as a template for the synthesis of edited genetic information. The pegRNA is capable for instance, without limitation, of (i) identifying the target nucleotide sequence to be edited and (ii) encoding new genetic information that replaces the targeted sequence. In some embodiments, the pegRNA is capable of (i) identifying the target nucleotide sequence to be edited and (ii) encoding an integration site that replaces the targeted sequence.

[0062] As used herein, the term “nicking guide RNA” (ngRNA) and the like refer to an RNA sequence that can nick a strand such as an edited strand and a non-edited strand. The ngRNA can induce nicks at about 1 or more nt away from the site of the gRNA-induced nick. For example, the ngRNA can nick at least at about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, or more nt away from the site of the gRNA induced nick. As used herein, the terms “reverse transcriptase” and “reverse transcriptase domain” refer to an enzyme or an enzymatically active domain that can reverse a RNA transcribe into a complementary DNA. The reverse transcriptase or reverse transcriptase domain is a RNA dependent DNA polymerase. Such reverse transcriptase domains encompass, but are not limited, to a M-MLV reverse transcriptase, or a modified reverse transcriptase such as, without limitation, Superscript® reverse transcriptase (Invitrogen; Carlsbad, California), Superscript® VILO™ cDNA synthesis (Invitrogen; Carlsbad, California), RTX, AMV-RT, and Quantiscript Reverse Transcriptase (Qiagen, Hilden, Germany).

[0063] The pegRNA-PE complex disclosed herein recognizes the target site in the genome and the Cas9 for example nicks a protospacer adjacent motif (PAM) strand. The primer binding site (PBS) in the pegRNA hybridizes to the PAM strand. The RT template operablyAttorney Docket: 761259: 083474-041PC linked to the PBS, containing the edit sequence, directs the reverse transcription of the RT template to DNA into the target site. Equilibration between the edited 3′ flap and the unedited 5′ flap, cellular 5′ flap cleavage and ligation, and DNA repair results in stably edited DNA. To optimize base editing, a Cas9 nickase can be used to nick the non-edited strand, thereby directing DNA repair to that strand, using the edited strand as a template. Integrase Technologies

[0064] The present disclosure provides non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering using integrase technologies. Integrase technologies will be discussed in more detail below.

[0065] The integrase technologies used herein comprise proteins or nucleic acids encoding the proteins that direct integration of a gene of interest or nucleic acid sequence of interest into an integration site via an enzyme, such as an integrase or a prime editing nuclease. In certain embodiments, the protein directing the integration can be an enzyme such as an integration enzyme. In certain embodiments, the integration enzyme can be an integrase that incorporates the genome or nucleic acid of interest into the cell genome at the integration site by integration or recombination.

[0066] As used herein, the term “integration enzyme” refers to an enzyme or protein used to integrate a gene of interest or nucleic acid sequence of interest into a desired location or at the integration site, in the genome of a cell, in a single reaction or multiple reactions. In certain embodiments, the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200- 69221 and 69226-69340. In certain embodiments, the integrase or fragment thereof comprises an amino acid sequence that is about 80% identical, about 85% identical, about 90% identical, about 91% identical, about 92% identical, about 93% identical, about 94% identical, about 95% identical, about 96% identical, about 97% identical, about 98% identical, about 99% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200-69221 and 69226-69340.

[0067] In certain embodiments, the integrase fragment comprises (e.g., retains) integrase activity.

[0068] In certain embodiments, the integrase further comprises one or more mutations. Mutations include, but are not limited to, amino acid substitutions, amino acid deletions, and amino acid insertions.Attorney Docket: 761259: 083474-041PC

[0069] In some embodiments, the term “integration enzyme” refers to a nucleic acid (DNA or RNA) encoding the above-mentioned enzymes.

[0070] In some embodiments, the serine integrase φC31 from φC31 phage is used as an integration enzyme. The integrase φC31 in combination with a pegRNA can be used to insert the pseudo attP integration site (CCCCAACTGGGGTAACCTTTGAGTTCTCTCAGTTGGGG (SEQ ID NO: 54)). A DNA minicircle containing a gene or nucleic acid of interest and attB (GGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGG (SEQ ID NO: 55)) site can be used to integrate the gene or nucleic acid of interest into the genome of a cell. This integration can be aided by a co-transfection of an expression vector having the φC31 integrase.

[0071] In some embodiments, the integrase Bxb1 is used as an integration enzyme. The integrase Bxb1 can be used for integration / recombination using an attB sequence (GGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGG (SEQ ID NO: 69222)) and an attP sequence (GTGGTTTGTCTGGTCAACCACCGCGGTCTCAGTGGTGTACGGTACAAACCCA (SEQ ID NO: 69223)).

[0072] As used herein, the term “integrase” refers to a bacteriophage derived integrase, including wild-type integrase and any of a variety of mutant or modified integrases. As used herein, the term “integrase complex” may refer to a complex comprising integrase and integration host factor (IF). As used herein, the term “integrase complex” and the like may also refer to a complex comprising an integrase, an integration host factor, and a bacteriophage X-derived excisionase. EVOLVE-Pro Model for Integrase Enhancement

[0073] The EVOLVE-Pro model is used to enhance integrases, such as the Bxb1 integrase. EVOLVE-Pro is a frontier multi-modal protein design model. EVOLVE-Pro evolves high- activity protein variants with few-shot learning and minimal experimental testing, achieving accurate prediction of sequence-to-function for general properties. By applying EVOLVE- Pro in a few-shot, active learning framework, protein sequences with significantly higher activity can be efficiently nominated in a generalizable fashion with minimal effort. The modularity of the EVOLVE-Pro architecture allows this framework to scale with larger parameter Protein language models (PLMs).Attorney Docket: 761259: 083474-041PC

[0074] EVOLVE-Pro was benchmarked in silico across a panel of 12 different proteins, showing state-of-the-art performance, and the final model was then applied for genome editing with a Bxb1 integrase. EVOLVE-Pro yields mutants with significant improvement over initial proteins. EVOLVE-Pro’s utility was showcased in nominating multi-mutant protein designs out of a vast sequence space that enables final mutants that are much more active than naturally observed proteins. Development and Benchmarking of the EVOLVE-Pro Model

[0075] An ensemble model was designed to establish EVOLVE-Pro. The ensemble model involves: 1) a foundational protein language model to encode protein sequences into an information-rich latent space, and 2) a top-layer discrimination model to learn protein functional grammar in this evolutionary landscape and rank protein sequences according to a designed policy framework, and 3) an active learning framework using top layer discrimination model to nominate the next set of protein variants for experimental evaluation. This cycle is performed iteratively to evolve defined protein activities until they reach desired levels (Fig.10A).

[0076] EVOLVE-Pro was optimized across five parameters: 1) the strategy employed for the first round mutant selection, 2) the top layer discrimination model that learns the fitness landscape, 3) the active learning strategy for selecting mutants for the next round, 4) the evolution policy, and 5) the embedding vector transformation (FIG.11). To perform a grid search across this space, a panel of twelve unique deep mutagenesis scanning (DMS) datasets for in silico validation was curated (FIG.12). These twelve proteins represent diverse functions, including viral spike proteins, RNA-guided nucleases, lactases, and kinases, ensuring that the resulting model will be as generalizable as possible for learning diverse protein activity landscapes in PLM latent space.

[0077] ESM-2 protein language model was first assessed because of its large training data and available model size of >200M proteins and 15B parameters, respectively. Using the ESM-215B parameter model, the grid search found the optimal strategy was: 1) selecting a random set of first-round variants, 2) employing a random forest regressor discriminatory model to predict protein function, 3) using residue pooled average embeddings, and 4) using a top-N selection strategy in each round of evolution (FIG.13A). This policy nominated a high frequency of gain-of-function protein variants in only 5 rounds (FIG.13A). Since the focus was on the percent of activity passing a threshold as the evaluation metric in the grid search, increasing function during in silico evolution was next evaluated. Both the medianAttorney Docket: 761259: 083474-041PC activity and the activity of the nominated top mutant were found to increase monotonically from round to round across all DMS datasets, further validating the model’s performance in this low-N active learning setting (FIG.13B).

[0078] In general, 16 mutants per round of evolution for 10 rounds identified top mutants with fitness in the 50th percentile for eleven of the twelve DMS datasets. To understand how the number of variants per round affected performance, between 10 and 100 variants per round were tested, finding that larger rounds increased prediction accuracy (FIG.13C). This performance trade off indicates that EVOLVE-Pro can be used for both extremely low-N evolution (<20 mutants per round) for rapid and cheap experimental characterization and medium-N (~100 mutants per round) for quicker and more efficient evolution with fewer rounds.

[0079] After optimizing the top layer model and learning strategies, the PLM was optimized, comparing ESM-215B to a panel of foundational models. Using the optimal parameters from the grid search, performance was benchmarked against smaller versions of ESM-2 and ESM- 1, UniRep, ProtT5, ProteinBERT, Ankh, one-hot encoding, and integer encoded protein representations. ESM-215B parameter model outperformed all the other models for identifying the highest fitness proteins for all datasets except two, confirming its final selection for the EVOLVE-Pro latent space model (FIG.10B). Importantly, large parameter PLMs showed a significant boost in prediction accuracy compared to non-language model- based architectures, indicative of the powerful feature extraction present in transformer-based models (FIG.10B).

[0080] EVOLVE-Pro’s performance was then benchmarked relative to other PLM-based engineering approaches. As many methods require pre-training a discriminatory model on thousands of variants, tested versions of EVOLVE-Pro augmented with various amounts of pre-training were tested (FIG.10C). Reinforcement learning drastically reduced the overall number of mutants required: EVOLVE-Pro with only 5 rounds of evolution (16 mutants per round) was equivalent in performance to EVOLVE-Pro pre-trained with 160 mutants, while 10 rounds of evolution (16 mutants per round) was equivalent to pre-training with 500 mutants. Moreover, EVOLVE-Pro significantly outperformed zero-shot prediction methods. This comparison confirms that the few-shot nature of EVOLVE-Pro allows for efficient directed evolution with minimal effort and low-N testing per round (FIG.10C).

[0081] Lastly, the per-round evolution improvement for EVOLVE-Pro compared to one-hot and integer encoding and zero-shot prediction were evaluated, finding that by round 5 variants with significantly enhanced fitness could universally be found (at 16 mutations perAttorney Docket: 761259: 083474-041PC round) (FIG.10D, FIG.14A and FIG.14B). Moreover, in many cases, the one-hot and integer encoding frameworks saturated much earlier in the evolution process and never reached the fitness levels achieved by EVOLVE-Pro. Interestingly, for some proteins a non- linear increase in protein fitness after round 3 was observed, suggesting greater gains in mapping the protein fitness landscape as EVOLVE-Pro evolution proceeds. Bxb1 Integrase Evolution with EVOLVE-Pro

[0082] Large serine recombinases (LSRs) are enzymes that facilitate precise DNA rearrangements, making them crucial tools for genome editing. Their ability to recognize specific DNA sequences and catalyze targeted recombination events allows for efficient and accurate modifications of genetic material, which is essential for advanced gene therapy, synthetic biology, and genetic research. A gene insertion technology, PASTE, that leverages LSRs, specifically the Bxb1 integrase from Bxb1 phage, for programmable gene insertion in eukaryotic cells was recently developed. A limitation of Bxb1 integrase, however, is its activity saturates in the 20-60% range in cells, limiting the overall integration efficiency that can be achieved. Bxb1 was therefore enhanced using EVOLVE-Pro to improve its activity and demonstrate improved gene integration applications with PASTE in cells.

[0083] To evolve Bxb1, a simple integration assay in HEK293FT cells that involved the insertion of an AttP-containing DNA plasmid into an AttB target-containing plasmid was designed (FIG.15A). Integration can be measured by next-generation sequencing, and the evolution policy is designed to optimize this insertion efficiency. Evolution was strated with a round of 11 random Bxb1 point mutation variants and then over 8 rounds observed progressively increasing activity, resulting in mutants with over 2-fold higher activity than wild-type (FIG.15B, FIG.16A). As Bxb1 is already fairly active, this fold improvement is expected as near-saturating levels of insertion was reached. To validate the top hits from the evolution campaign, a Bxb1 plasmid titration experiment in a separate cell line (Hela cells) was performed and up to 4-fold improvement in recombination efficiency under low Bxb1 expression was observed (FIG.15C, FIG.16B). T166R was termed the final EVOLVE-Pro Bxb1 candidate as enhanced Bxb1, or epBxb1.

[0084] To test whether the epBxb1 variant’s improved activity can improve the programmable insertion of cargo DNA into the chromosome, this variant was tested in the context of PASTE and compared against the wild-type Bxb1 across five different genomic loci. Up to ~4-fold improvement was found in the final large cargo insertion rate into the genome, which highlights the generalizable gain in activity (FIG.15D, FIG.16C).Attorney Docket: 761259: 083474-041PC

[0085] An AlphaFold3-predicted model of Bxb1 bound to attachment site DNA indicates that the top beneficial EVOLVE-Pro mutations clustered in the Bxb1 DNA-binding domains, likely increasing the affinity to its DNA targets (FIG.15E). Of these residues, V292S could directly interact with the phosphate backbone of the target DNA based on its positioning relative to the attachment site, whereas the others likely modulate DNA binding via indirect interactions. Analysis of the residue exploration by the model revealed that multiple positions, including F439, V375, and L275, are visited up to 15 times; the DNA-interacting residue V292 was also visited multiple times, showing the model’s reasoning. Overall, this highlights EVOLVE-Pro's ability to recognize the functional importance of certain regions in the protein, much like structure-guided engineering approaches (FIG.15F).

[0086] The relationship between the fitness (pMMS) and function (observed fold improvement) for Bxb1 integrase was the calculated and a weakly positive correlation between the two metrics contrary to the other proteins reported evolves was found. This likely reflects a subset of protein families where protein stability and fitness as learned by the PLM can predict activity as previously reported (FIG.15G). However, given that the relationship is weak, a model like EVOLVE-Pro is still needed to efficiently and quickly reach high performing variants without encountering many false positives. Lastly, the global mutation landscape learned by EVOLVE-Pro was found to be still divergent from the predicted fitness (pMMS) by ESM2 with an even weaker correlation, further highlighting the ability of EVOLVE-Pro to learn protein function at a global scale and how stability / fitness prediction is not sufficient for rapid and efficient protein evolution (FIG.15H-J). Integration Site

[0087] The present disclosure provides non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering via the addition of an integration site into a target genome. The integration site will be discussed in more details below.

[0088] As used herein, the term “integration site” refers to the site where one or more genes of interest or one or more nucleic acid sequences of interest are inserted.

[0089] The integration site can be inserted into the genome or a fragment thereof of a cell using a nuclease, a gRNA, and / or an integration enzyme. The integration site can be inserted into the genome of a cell using a prime editor such as, without limitation, PE1, PE2, and PE3, wherein the integration site is carried on a pegRNA. The pegRNA can target any site that is known in the art. Examples of cites targeted by the pegRNA include, without limitation,Attorney Docket: 761259: 083474-041PC ACTB, SUPT16H, SRRM2, NOLC1, DEPDC4, NES, LMNB1, AAVS1 locus, CC10, CFTR, SERPINA1, ABCA4, and any derivatives thereof. The complementary integration site may be operably linked to a gene of interest or nucleic acid sequence of interest in an exogenous DNA or RNA. In some embodiments, one integration site is added to a target genome. In some embodiments, more than one integration sites are added to a target genome.

[0090] To insert multiple genes or nucleic acids of interest, two or more integration sites are added to a desired location. Multiple DNA comprising nucleic acid sequences of interest are flanked orthogonal to the integration sequences such as, without limitation, attB, attP, other recognition site pairs, or any pseudosites in the human genome. As used herein, a “pseudosite” is a nucleic acid sequence in the target genome (e.g., a human genome) that is similar to a wild type attB or attP sequences. The sequence similarity is sufficient to allow integration of a nucleic acid sequence with an integrase enzyme. An integration site is "orthogonal" when it does not significantly recognize the recognition site or nucleotide sequence of a recombinase. Thus, one attB site of a recombinase can be orthogonal to an attB site of a different recombinase. In addition, one pair of attB and attP sites of a recombinase can be orthogonal to another pair of attB and attP sites recognized by the same recombinase. A pair of recombinases are considered orthogonal to each other, as defined herein, when there is recognition of each other's attB or attP site sequences. In certain embodiments, the attB nucleic acid sequences selected from the group consisting of SEQ ID NOs: 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 55, and 69222. In certain embodiments, the attP nucleic acid sequences selected from the group consisting of SEQ ID NOs: 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, and 69223 In certain embodiments, the attB / attP nucleic acid pair is selected from the group consisting of: SEQ ID NO: 17 / SEQ ID NO: 18, SEQ ID NO: 19 / SEQ ID NO: 20, SEQ ID NO: 21 / SEQ ID NO: 22, SEQ ID NO: 23 / SEQ ID NO: 24, SEQ ID NO: 25 / SEQ ID NO: 26, SEQ ID NO: 27 / SEQ ID NO: 28, SEQ ID NO: 29 / SEQ ID NO: 30, SEQ ID NO: 31 / SEQ ID NO: 32, SEQ ID NO: 33 / SEQ ID NO: 34, SEQ ID NO: 35 / SEQ ID NO: 36, SEQ ID NO: 37 / SEQ ID NO: 38, SEQ ID NO: 39 / SEQ ID NO: 40, SEQ ID NO: 41 / SEQ ID NO: 42, SEQ ID NO: 43 / SEQ ID NO: 44, SEQ ID NO: 45 / SEQ ID NO: 46, SEQ ID NO: 47 / SEQ ID NO: 48, and SEQ ID NO: 69222 / SEQ ID NO: 69223.

[0091] In certain embodiments, the attB nucleic acid sequence is between 12 and 60 nucleotides in length or between 18 and 50 nucleotides in length. In certain embodiments, the attB nucleic acid sequence is 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27,Attorney Docket: 761259: 083474-041PC 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides in length.

[0092] In certain embodiments, the attP nucleic acid sequence is between 12 and 60 nucleotides in length or between 18 and 50 nucleotides in length. In certain embodiments, the attP nucleic acid sequence is 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides in length.

[0093] In certain embodiments, the attB and / or attP nucleic acid sequence comprises one or more truncations. The truncation may be at the 5’ end, 3’end, or both. The truncations to the attB and / or attP nucleic acids sequences may be made while still retaining the ability to bind an integrase.

[0094] In certain embodiments, the attB and / or attP nucleic acid sequence is truncated by 1 to 32 nucleotides from one or both of the 5’ end and 3’ end. In certain embodiments, the attB nucleic acid sequence is truncated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, or 32 nucleotides from one or both of the 5’ end and 3’ end. In certain embodiments, the attP nucleic acid sequence is truncated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, or 32 nucleotides from one or both of the 5’ end and 3’ end.

[0095] In certain embodiments, any one of the attB nucleic acid sequences selected from the group consisting of SEQ ID NOs: 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 55, and 69222 is truncated by 1 to 32 nucleotides from one or both of the 5’ end and 3’ end. In certain embodiments, any one of the attP nucleic acid sequences selected from the group consisting of SEQ ID NOs: 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, and 69223 is truncated by 1 to 32 nucleotides from one or both of the 5’ end and 3’ end.

[0096] The lack of recognition of integration sites can be less than about 30%. In some embodiments, the lack of recognition of integration sites or pairs of sites can be less than about 30%, less than about 28%, less than about 26%, less than about 24%, less than about 22%, less than about 20%, less than about 18%, less than about 16%, less than about 14%, less than about 12%, less than about 10%, less than about 8%, less than about 6%, less than about 4%, less than about 2%, about 1%, or any range that is formed from any two of those values as endpoints. The crosstalk can be less than about 30%. In some embodiments, the crosstalk is less than about 30%, less than about 28%, less than about 26%, less than about 24%, less than about 22%, less than about 20%, less than about 18%, less than about 16%, less than about 14%, less than about 12%, less than about 10%, less than about 8%, less thanAttorney Docket: 761259: 083474-041PC about 6%, less than about 4%, less than about 2%, less than about 1%, or any range that is formed from any two of those values as endpoints.

[0097] In some embodiments, the attB and / or attP site sequences comprise a central dinucleotide sequence. It has been shown that, for example, the central dinucleotide can be changed to GA from GT and that only GA containing attB / attP sites interact and will not cross react with GT containing sequences. In some embodiments, the central dinucleotide is selected from the group consisting of AG, AC, TG, TC, CA, CT, GA, AA, TT, CC, GG, AT, TA, GC, CG and GT.

[0098] As used herein, the term “pair of an attB and attP site sequences” and the like refer to attB and attP site sequences that share the same central dinucleotide and can recombine. This means that in the presence of one serine integrase as many as six pairs of these orthogonal att sites can recombine (attPTT will specifically recombine with attBTT, attPTC will specifically recombine with attBTC, and so on).

[0099] In some embodiments, the central dinucleotide is nonpalindromic. In some embodiments, the central dinucleotide is palindromic. In some embodiments, a pair of an attB site sequence and an attP site sequence are used in different DNA encoding genes of interest or nucleic acid sequences of interest for inducing directional integration of two or more different nucleic acids. In some embodiments, two integrases can be used for orthogonal insertion.

[0100] The Table 1 below shows examples of pairs of attB site sequence and attP site sequence with different central dinucleotide (CD). Table 1 – attB and attP sequences and central dinucleotides (CD)Attorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PC

[0101] In one aspect, the disclosure provides an integrase or fragment thereof, wherein: a) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 1, wherein the integraseAttorney Docket: 761259: 083474-041PC binds to the attB nucleic acid set forth in SEQ ID NO: 17 and the attP nucleic acid set forth in SEQ ID NO: 18; b) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 2, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 19 and the attP nucleic acid set forth in SEQ ID NO: 20; c) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 3, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 21 and the attP nucleic acid set forth in SEQ ID NO: 22; d) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 4, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 23 and the attP nucleic acid set forth in SEQ ID NO: 24; e) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 5, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 25 and the attP nucleic acid set forth in SEQ ID NO: 26; f) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 6, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 27 and the attP nucleic acid set forth in SEQ ID NO: 28; g) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 7, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 29 and the attP nucleic acid set forth in SEQ ID NO: 30; h) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 8, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 31 and the attP nucleic acid set forth in SEQ ID NO: 32; i) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 9, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 33 and the attP nucleic acid set forth in SEQ ID NO: 34;Attorney Docket: 761259: 083474-041PC j) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 10, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 35 and the attP nucleic acid set forth in SEQ ID NO: 36; k) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 11, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 37 and the attP nucleic acid set forth in SEQ ID NO: 38; l) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 12, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 39 and the attP nucleic acid set forth in SEQ ID NO: 40; m) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 13, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 41 and the attP nucleic acid set forth in SEQ ID NO: 42; n) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 14, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 43 and the attP nucleic acid set forth in SEQ ID NO: 44; o) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 15, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 45 and the attP nucleic acid set forth in SEQ ID NO: 46; or p) the integrase or fragment thereof comprises an amino acid sequence that is at least 80% identical to an amino acid sequence set forth in SEQ ID NO: 16, wherein the integrase binds to the attB nucleic acid set forth in SEQ ID NO: 47 and the attP nucleic acid set forth in SEQ ID NO: 48. PASTE

[0102] The present disclosure provides non-naturally occurring or engineered systems, methods, and compositions for site-specific genetic engineering using Programmable Addition via Site-Specific Targeting Elements (PASTE). PASTE will be discussed in more details below. The PASTE system is described in greater detail in U.S. Patent ApplicationAttorney Docket: 761259: 083474-041PC Nos.17 / 649,308, 18 / 066,233, 18 / 066,223, 18 / 487,610, and 18 / 487,744, and U.S. Patent Nos. 11,572,556, 11,827,881, 11,834,658, 11,952,571, and 12,195,744, each of which is incorporated herein by reference.

[0103] The site-specific genetic engineering disclosed herein is for the insertion of one or more genes of interest or one or more nucleic acid sequences of interest into a genome of a cell. In some embodiments, the gene of interest is a mutated gene implicated in a genetic disease such as, without limitation, a metabolic disease, cystic fibrosis, muscular dystrophy, hemochromatosis, Tay-Sachs, Huntington disease, Congenital Deafness, Sickle cell anemia, Familial hypercholesterolemia, adenosine deaminase (ADA) deficiency, X-linked SCID (X- SCID), and Wiskott-Aldrich syndrome (WAS). In some embodiments, the gene of interest or nucleic acid sequence of interest can be a reporter gene upstream or downstream of a gene for genetic analyses such as, without limitation, for determining the expression of a gene. In some embodiments, the reporter gene is a GFP template or a Gaussia Luciferase (G- Luciferase) template. In some embodiments, the gene of interest or nucleic acid sequence of interest can be used in plant genetics to insert genes to enhance drought tolerance, weather hardiness, and increased yield and herbicide resistance in plants. In some embodiments, the gene of interest or nucleic acid sequence of interest can be used for site-specific insertion of a protein (e.g., a lysosomal enzyme), a blood factor (e.g., Factor I, II, V, VII, X, XI, XII or XIII), a membrane protein, an exon, an intracellular protein (e.g., a cytoplasmic protein, a nuclear protein, an organellar protein such as a mitochondrial protein or lysosomal protein), an extracellular protein, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensory protein, a motor protein, a defense protein, or a storage protein, an anti-inflammatory signaling molecules into cells for treatment of immune diseases, including but not limited to arthritis, psoriasis, lupus, coeliac disease, glomerulonephritis, hepatitis, and inflammatory bowel disease.

[0104] The size of the inserted gene or nucleic acid can vary from about 1 bp to about 50,000 bp. In some embodiments, the size of the inserted gene or nucleic acid can be about 1 bp, 10 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 600 bp, 800 bp, 1000 bp, 1200 bp, 1400 bp, 1600 bp, 1800 bp, 2000 bp, 2200 bp, 2400 bp, 2600 bp, 2800 bp, 3000 bp, 3200 bp, 3400 bp, 3600 bp, 3800 bp, 4000 bp, 4200 bp, 4400 bp, 4600 bp, 4800 bp, 5000 bp, 5200 bp, 5400 bp, 5600 bp, 5800 bp, 6000 bp, 6200, 6400 bp, 6600 bp, 6800 bp, 7000 bp, 7200 bp, 7400 bp, 7600 bp, 7800 bp, 8000 bp, 8200 bp, 8400 bp, 8600 bp, 8800 bp, 9000 bp, 9200 bp, 9400 bp, 9600 bp, 9800 bp, 10,000 bp, 10,200 bp, 10,400 bp, 10,600 bp, 10,800 bp, 11,000 bp, 11,200 bp, 11,400 bp, 11,600 bp, 11,800 bp, 12,000 bp, 14,000 bp, 16,000 bp,Attorney Docket: 761259: 083474-041PC 18,000 bp, 20,000 bp, 30,000 bp, 40,000 bp, 50,000 bp, or any range that is formed from any two of those values as endpoints.

[0105] In some embodiments, the site-specific engineering using the gene of interest or nucleic acid sequence of interest disclosed herein is for the engineering of T cells and NKs for tumor targeting or allogeneic generation. These can involve the use of receptor or CAR for tumor specificity, anti-PD1 antibody, cytokines like IFN-gamma, TNF-alpha, IL-15, IL- 12, IL-18, IL-21, and IL-10, and immune escape genes.

[0106] In the present disclosure, the site-specific insertion of the gene of interest or nucleic acid of interest is performed through PASTE. Components for inserting a gene of interest or a nucleic acid of interest using PASTE are for example, without limitation, a nuclease, a gRNA adding the integration site, a DNA or RNA strand comprising the gene or nucleic acid linked to a sequence that is complementary or associated to the integration site, and an integration enzyme. Components for inserting a gene of interest or a nucleic acid of interest using PASTE are for example, without limitation, a prime editor expression, pegRNA adding the integration site, nicking guide RNA, integration enzyme (an integrase, such as an integrase of any one of SEQ ID NOs: 69200-69221 and 69226-69340), transgene vector comprising the gene of interest or nucleic acid sequence of interest with gene and integration signal. The nuclease and prime editor integrate the integration site into the genome. The integration enzyme integrates the gene of interest into the integration site. In some embodiments, the transgene vector comprising the gene or nucleic acid sequence of interest with gene and integration signal is a DNA minicircle devoid of bacterial DNA sequences. In some embodiments, the transgenic vector is a eukaryotic or prokaryotic vector.

[0107] As used herein, the term “vector” or “transgene vector” refers to a recombinant DNA molecule containing a desired coding sequence and appropriate nucleic acid sequences necessary for the expression of the operably linked coding sequence in a host organism. Nucleic acid sequences necessary for expression in prokaryotes usually include for example, without limitation, a promoter, an operator (optional), a ribosome binding site, and / or other sequences. Eukaryotic cells are generally known to utilize promoters (constitutive, inducible or tissue specific), enhancers, and termination and polyadenylation signals, although some elements may be deleted and other elements added without sacrificing the necessary expression. The transgenic vector may encode the PE and the integration enzyme, linked to each other via a linker. The linker can be a cleavable linker. In some embodiments, the linker can be a non-cleavable linker. In some embodiments the nuclease, prime editor, and / or integration enzyme can be encoded in different vectors.Attorney Docket: 761259: 083474-041PC

[0108] In one aspect, the disclosure provides a method of inserting multiple genes or nucleic acid sequences of interest into a single site. In some embodiments, multiplexing involves inserting multiple genes of interest in multiple loci using unique pegRNA (Merrick, C. A. et al., ACS Synth. Biol.2018, 7, 299−310). The insertion of multiple genes of interest or nucleic acids of interest into a cell genome, referred herein as “multiplexing,” is facilitated by incorporation of the complementary 5’ integration site to the 5’ end of the DNA or RNA comprising the first nucleic acid and 3’ integration site to the 3’ end of the DNA or RNA comprising the last nucleic acid. In some embodiments, the number of genome of interest or amino acid sequences of interest that are inserted into a cell genome using multiplexing can be about 1, 2, 3, 4, 5, 6, 7, 8, 910, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or any range that is formed from any two of those values as endpoints.

[0109] In some embodiments, multiplexing allows integration of for example, signaling cascade, over-expression of a protein of interest with its cofactor, insertion of multiple genes mutated in a neoplastic condition, or insertion of multiple CARs for treatment of cancer.

[0110] In some embodiments, the integration sites may be inserted into the genome using non-prime editing methods such as rAAV mediated nucleic acid integration, TALENS and ZFNs. A number of unique properties make AAV a promising vector for human gene therapy (Muzyczka, CURRENT TOPICS IN MICROBIOLOGY AND IMMUNOLOGY, 158:97-129 (1992)). Unlike other viral vectors, AAVs have not been shown to be associated with any known human disease and are generally not considered pathogenic. Wild type AAV is capable of integrating into host chromosomes in a site-specific manner M. Kotin et al., PROC. NATL. ACAD. SCI, USA, 87:2211-2215 (1990); R.J. Samulski, EMBO 10(12):3941-3950 (1991)). Instead of creating a double-stranded DNA break, AAV stimulates endogenous homologous recombination to achieve the DNA modification. Further, transcription activator-like effector nucleases (TALENs) and Zinc-finger nucleases (ZFNs) for genome editing and introducing targeted DSBs. The specificity of TALENs arises from two polymorphic amino acids, the so-called repeat variable diresidues (RVDs) located at positions 12 and 13 of a repeated unit. TALENS are linked to FokI nucleases, which cleaves the DNA at the desired locations. ZFNs are artificial restriction enzymes for custom site- specific genome editing. Zinc fingers themselves are transcription factors, where each finger recognizes 3-4 bases. By mixing and matching these finger modules, researchers can customize which sequence to target.

[0111] As used herein, the terms “administration,” “introducing,” or “delivery” into a cell, a tissue, or an organ of a plasmid, nucleic acids, or proteins for modification of the hostAttorney Docket: 761259: 083474-041PC genome refers to the transport for such administration, introduction, or delivery that can occur in vivo, in vitro, or ex vivo. Plasmids, DNA, or RNA for genetic modification can be introduced into cells by transfection, which is typically accomplished by chemical means (e.g., calcium phosphate transfection, polyethyleneimine (PEI) Or lipofection), physical means (electroporation or microinjection), infection (this typically means the introduction of an infectious agent such as a virus (e.g., a baculovirus expressing the AAV Rep gene)), transduction (in microbiology, this refers to the stable infection of cells by viruses, or the transfer of genetic material from one microorganism to another by viral factors (e.g., bacteriophages)). Vectors for the expression of a recombinant polypeptide, protein or oligonucleotide may be obtained by physical means (e.g., calcium phosphate transfection, electroporation, microinjection, or lipofection) in a cell, a tissue, an organ or a subject. The vector can be delivered by preparing the vector in a pharmaceutically acceptable carrier for the in vitro, ex vivo, or in vivo delivery to the carrier.

[0112] As used herein, the term “transfection” refers to the uptake of an exogenous nucleic acid molecule by a cell. A cell is “transfected” when an exogenous nucleic acid has been introduced into the cell membrane. The transfection can be a single transfection, co- transfection, or multiple transfection. Numerous transfection techniques are generally known in the art. See, for example, Graham et al. (1973) Virology, 52: 456. Such techniques can be used to introduce one or more exogenous nucleic acid molecules into a suitable host cell.

[0113] In some embodiments, the exogenous nucleic acid molecule and / or other components for gene editing are combined and delivered in a single transfection. In other embodiments, the exogenous nucleic acid molecule and / or other components for gene editing are not combined and delivered in a single transfection. In some embodiments, exogenous nucleic acid molecule and / or other components for gene editing are combined and delivered in a single transfection to comprise for example, without limitation, a prime editing vector, a landing site such as a landing site containing pegRNA, a nicking guide such as a nicking guide for stimulating prime editing, an expression vector such as an expression vector for a corresponding integrase or recombinase, a minicircle DNA cargo such as a minicircle DNA cargo encoding for green fluorescent protein (GFP), any derivatives thereof, and any combinations thereof. In some embodiments, the gene of interest or amino acid sequence of interest can be introduced using liposomes. In some embodiments, the gene of interest or amino acid sequence of interest can be delivered using suitable vectors for instance, without limitation, plasmids and viral vectors. Examples of viral vectors include, without limitation, adeno-associated viruses (AAV), lentiviruses, adenoviruses, other viral vectors, derivativesAttorney Docket: 761259: 083474-041PC thereof, or combinations thereof. The proteins and one or more guide RNAs can be packaged into one or more vectors, e.g., plasmids or viral vectors. In some embodiments, the delivery is via nanoparticles or exosomes. For example, exosomes can be particularly useful in delivery RNA.

[0114] In some embodiments, the prime editing inserts the landing site with efficiencies of at least about 1%, at least about 5%, at least about 10%, at least about 15%, at least about, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, or at least about 50%. In some embodiments, the prime editing inserts the landing site(s) with efficiencies of about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, or any range that is formed from any two of those values as endpoints. Sequences

[0115] Sequences of enzymes, guides, integration sites, and plasmids can be found in the Tables below. Table 2 – Integrase enzyme amino acid sequences and the AttB / AttP nucleic acid sequences recognized by said integrase enzymes.Attorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCTable 3 – Integrase enzyme nucleic acid sequencesAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCTable 4 – Linker sequencesTable 5 – Exemplary Cas9 nuclease and reverse transcriptaseAttorney Docket: 761259: 083474-041PCTable 6 – Sequences of atgRNAs, sgRNAs, and nicking guides (Spacers are labeled in bold, underlined text, RT regions in bold, underlined, and italicized text, AttB sites in bold text, and PBS in underlined and italicized text)Attorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCTable 7 – Mammalian expression vectorsAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PC5Attorney Docket: 761259: 083474-041PC Table 8 – Top scoring integrase enzyme amino acid sequences and the AttB / AttP nucleic acid sequences recognized by said integrase enzymesAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PC5Attorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PC_Attorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCEXAMPLESAttorney Docket: 761259: 083474-041PC

[0116] While several experimental Examples are contemplated, these Examples are intended to be non-limiting. Example 1 Bxb1 Integration Data Lenti Reporter

[0117] The PASTE system, including the description in Example 1 and Example 2, are described in greater detail in U.S. Patent Application Nos.18 / 487,610, and 18 / 487,744, and U.S. Patent Nos.11,572,556, 11,827,881, and 11,834,658, each of which is incorporated herein by reference.

[0118] Serine integrase Bxb1 has been shown to be more active than Cre recombinase and highly efficient in bacteria and mammalian cells for irreversible integration of target genes. FIG.1 and FIG.2 show schematics of PASTE methodology using Bxb1 (Merrick, C. A. et al., ACS Synth. Biol.2018, 7, 299−310).

[0119] To probe the efficiency of the Bxb1 integration system, a clonal HEK293FT cell line with attB Bxb1 site (GGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGG (SEQ ID NO: 3163)) integrated using lentivirus was developed. The modified HEK293FT cell line was then transferred with the following plasmids: (1) plus / minus Bxb1 expression plasmid and (2) plus / minus GFP or G-Luc minicircle template with attP Bxb1 site. After 72 hours, the integration of GFP or Gluc into the attB site in the HEK293FT genome was probed. The percent integrations of GFP or Gluc into the attB locus are shown in FIG.3. It was observed that GFP and Gluc showed efficient integration into the attB site in HEK293FT cells. Example 2 Addition of Bxb1 Site to Human Genome Using PRIME

[0120] The maximum length of attB that can be integrated into a HEK293FT cell line with the best efficiency was probed. To probe the best length of attB (GGCCGGCTTGTCGACGACGGCGGTCTCCGTCGTCAGGATCATCCGG (SEQ ID NO: 3164)) or its reverse complement attP (CCGGATGATCCTGACGACGGAGACCGCCGTCGTCGACAAGCCGGCC (SEQ ID NO: 3165)) for prime editing, pegRNAs having PBS length of 13 nt with varying RT homology length were used. The following plasmids were transfected in HEK293FT: (1) prime expression plasmid; (2) HEK3 targeting pegRNA design; and (3) HEK3 +90 nicking guide. After 72 hours, the percent integration of each of the attB construct was probed. FIG.4 shows the percent editing in each HEK3 targeting pegRNA. It was observed that attB withAttorney Docket: 761259: 083474-041PC 44, 34 and 26 base pairs and attB reverse complement with 34 and 26 base pairs showed the highest percent editing. Example 3 Integrase Discovery Platform & Use in PASTE System

[0121] Integrase choice can have implications for integration activity. To identify novel integrases with improved activity in the PASTE system, bacterial and metagenomic sequences were mined for new phage associated serine integrases (FIG.5A). Exploring over 10 TB worth of data from NCBI, JGI, and other sources, 27,399 novel integrases were found (FIG.5B, FIG.5C) and their associated attachment sites were annotated using a novel repeat finding algorithm that could predict potential 50 bp attachment sites with high confidence near phage boundaries. Table 8 above recites top scoring integrase enzyme amino acid sequences and the attB / attP nucleic acid sequences recognized by said integrase enzymes. Accordingly, for each row of Table 8, an integrase amino acid sequence is recited, followed by the corresponding attB and attP nucleic acid sequence to which said integrase binds. Analysis of the integrases sequences revealed that they fell into four distinct clusters: INTa, INTb, INTc, and INTd. About half of integrases (14,771) derive from metagenomic sequences, presumably from pro-phages, and 13,693 of the integrases specifically derive from human microbiome metagenomic samples. An initial screen of integrase activity using a reporter system revealed that a number of the integrases were highly active in HEK293FT cells with more activity than BxbINT, a member of the INTa family (FIG.6A). Using the predicted 50 bp sequences encoded in attachment site-containing guide RNAs (atgRNAs) along with minicircles containing the complementary AttP sites, it was found that these integrases were compatible with PASTE but with lower efficiency than BxbINTa-based PASTE (FIG.6B). It was hypothesized that this was because of their longer 50 bp AttB sequences and so truncations of these AttBs were explored in the hopes of finding more minimal attachment sites. Truncation screening on integrase reporters revealed that AttB truncations of all the integrases, including as short as 34 bp, were still active and many had more activity than BxbINTa (FIG.6C). Upon porting these new shorter AttBs to atgRNAs for PASTE, it was found that a number of integrases had more activity in the PASTE system than BxbINT-based PASTE at the ACTB locus, including the integrase from B. cereues (BceINTc), N191352_143_72 stool sample from China (SscINTd), and N684346_90_69 stool sample from adult in China (SacINTd), while others like the integrase from B. cytotoxicus (BcytINTd) and S. lugdunensis (SluINTd) did not (FIG.6A and FIG.6D-FIG.Attorney Docket: 761259: 083474-041PC 6E). Because of its superior efficiency when used with PASTE, BceINTc when used as PASTE is referred to as PASTEv4.1. Moreover, upon optimization of these integrases with different linkers and RT domains, it was found that BceINTc fused to SpCas9-RTSto7d or SpCas9-MLV-RTL139Pvariant had the most activity, even higher than BxbINTa-based PASTE (FIG.6G-FIG.6I). The construct SpCas9-MLV-RTL139P-BceINTc construct is referred to as PASTEv4.1. This optimized PASTEv4.1 was then evaluated and across a number of endogenous gene loci it was found that it performed better than BxbINTa-based PASTE (FIG.6H and FIG.6J). Example 4 Screening of Additional Integrases

[0122] Ten additional integrases were tested for activity that were identified from the original discovery platform described above. The integrase amino acid sequences, nucleic acid sequences, and their corresponding attB and attP sites are recited in Table 10 below. As shown in FIG.7A, several of the integrase displayed integrase activity in the HEK293 reporter assay described above. Integrases N352807_16_14, N362476_2_132, and N391614_1_4134 displayed measurable activity. The integrase N352807_16_14 was next tested in the PASTE system with integration at the ACTB gene locus, along with truncations of the attB site. The truncations tested were 2, 4, 6, 8, 10, 12, 14, or 16 base pair truncations from either the left or right of the attB sequence. As shown in FIG.7B, the integrase N352807_16_14 achieved higher integration levels at the ACTB gene locus with all attB truncations compared to the BxbINT PASTE system. Table 10 – Integrase amino acid sequences with corresponding attB / attP nucleic acid sequencesAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCExample 5 Screening of Additional Integrases

[0123] Various mutant integrases were generated and screened for their efficiency of integration and / or recombination.

[0124] Mutants of the WT Bxb1 integrase sequence were designed using computer modeling. In a first run of computer modeling (Round 1), the following Bxb1mutants were designed: 23K, 58T, 79R, 115K, 141D, 182M, 230R, 271N, 318F, 345H, 376Y, and 422P, relative to the WT sequence of Bxb1. In a second run of computer modeling (Round 2), the following Bxb1 mutants were designed: 166C, 189Q, 245A, 25A, 375C, 375Y, 385A, 238M, and 375E. The amino acid sequences of the Bxb1 mutants are provided in Table 11.

[0125] To evaluate the integration and / or recombination efficiency of the Bxb1 mutant integrases, the following method was used (as shown in FIG.9): 2ng of the mutant Bxb1 integrase was added to 20ng of WT attP-containing-plasmid and 20ng of attB-containing plasmid then transfected into HEK293FT cells. The nucleic acid sequences of the attB, attP, and corresponding plasmids are provided in Table 12. After 48 hours of incubation, genomic DNA from the HEK293FT cells was harvested and Next Generation Sequencing (NGS) was used to amplify a target amplicon located around the attB site. Recombination was then quantified using crispresso software. The individual Round 1 mutant’s integration efficiency is reported on the y axis of the graphs showed in FIG.8A and FIG.8B.

[0126] In a second round of computer modeling (Round 2), the 12 mutants designed in Round 1 were fed into the computer model, then 9 mutants were designed for Round 2 testing. As above, the integration and / or recombination efficiency of the mutant integrases were evaluated as follows: 2ng of the mutant integrase was added to 20ng of the WT attP- containing plasmid and 20ng of the attB-containing plasmid, then transfected into HEK293FT cells. The nucleic acid sequences of the attB, attP, and corresponding plasmids are provided in Table 12. After 48 hours of incubation, genomic DNA from the HEK293FT cells was harvested and Next Generation Sequencing (NGS) was used to amplify the targetAttorney Docket: 761259: 083474-041PC amplicon located around the attB site. Recombination was then quantified using crispresso software. The individual Round 2 mutant’s integration efficiency is reported on the y axis of the graph showed in FIG.8B.

[0127] The method of deep next-generation sequencing entails isolation using 50 µl of QuickExtract (Lucigen) per well, where target regions were PCR amplified with NEBNext High-Fidelity 2× PCR master mix (NEB) based on the manufacturer’s protocol. Barcodes and adapters for Illumina sequencing were added in a subsequent PCR amplification. Amplicons were pooled and prepared for sequencing on a MiSeq (Illumina). Reads were demultiplexed and analyzed with crispresso2. Table 11 – Wild type and mutated integrase amino acid sequencesAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCTable 12 – Nucleic acid sequences of attB, attP, and corresponding plasmidsAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PCAttorney Docket: 761259: 083474-041PC

[0128] The results of these experiments showed that several Round 2 Bxb1 mutant integrases had increased amounts of integration (% integration) relative to the WT Bxb1 integrase. The particular Round 2 Bxb1 mutant integrases with increased efficiency were the mutants having the following amino acid mutations relative to WT Bxb1: 166C, 189Q, 25A, 375C, 375Y, and 375E. Example 6 Mammalian cell culture and transfection

[0129] Assays using mammalian cell were performed.

[0130] Mammalian cell culture experiments were performed in the HEK293FT (Thermo Fisher Scientific), BJ Fibroblast (ATCC CRL-2522), and Hepa 1-6 (ATCC CRL-1830) cell lines, grown in Dulbecco’s Modified Eagle Medium with high glucose, sodium pyruvate, and GlutaMAX (Thermo Fisher Scientific), and supplemented with 1 × penicillin-streptomycin (Thermo Fisher Scientific) and 10% fetal bovine serum (VWR Seradigm). All cells were maintained at confluency below 80%. All transfections were performed with Lipofectamine 3000 (Thermo Fisher Scientific). Cells were plated 16–20 hours prior to transfection to ensure 90% confluency at the time of transfection. For 96-well plates, cells were plated at 2×104cells / well. For each well on the plate, transfection plasmids were combined with OptiMEM I Reduced Serum Medium (Thermo Fisher Scientific) to a final volume of 10 μL. Example 7 Mammalian genome editing

[0131] The activity of mutant integrases were assessed in mammalian genome editing.

[0132] To measure genome editing activity, 50 ng of protein expression construct, 50 ng of the corresponding guide construct, and optionally 20 ng of luciferase reporter were transfected in one well of a 96-well plate, using Lipofectamine 3000. After 72 h, the cellsAttorney Docket: 761259: 083474-041PC were washed once with 1 × DPBS (Sigma Aldrich). The cells were resuspended in 50 μL QuickExtract DNA Extraction Solution (Lucigen) and cycled at 65°C for 15 min, 68°C for 15 min, and then at 95°C for 10 min for lysis. A 2.5 μL aliquot of the cell lysate was used as input for each PCR reaction. For library amplification, target reporter regions were amplified with a 12-cycle PCR using NEBNext High Fidelity 2 × PCR Master Mix (NEB) with an annealing temperature of 63°C for 15 s, followed by a second 20-cycle round of PCR to add Illumina adapters and barcodes. The libraries were gel extracted and subject to paired-end sequencing on an Illumina MiSeq with Read 1220 cycles, Index 18 cycles, Index 28 cycles, and Read 280 cycles. Insertion / deletion (indel) frequency was analyzed using CRISPResso2. Example 8 Three Primer NGS for PASTE

[0133] The integration of attB / attP pairs in the Bxb1 assay and genome editing for prime editing and HDR integration in the genomic locus was assessed.

[0134] To quantify the integration of attB / attP pairs in the Bxb1 assay and genome editing for prime editing and HDR integration in the genomic locus, target regions were PCR amplified and analyzed by deep sequencing. Genomic DNA samples were isolated using 50 µL of QuickExtract (Lucigen) per well, and target regions were PCR amplified with NEBNext High-Fidelity 2× PCR master mix (NEB) based on the manufacturer’s protocol. Barcodes and adapters for Illumina sequencing were added in a subsequent PCR amplification. Amplicons were pooled and prepared for sequencing on a MiSeq (Illumina). Reads were demultiplexed and analyzed with appropriate pipelines. To analyze the Bxb1 integration assay, the frequency of recombined attB / attP sites was counted relative to intact attB sites. To analyze prime and HDR editing, amplicons were analyzed using CRISPresso2 to count the relative number of reads with the inserted sequence. A three-primer NGS assay to quantify left junction integration was also developed using a common forward primer, a reverse primer to detect the unintegrated genomic locus, and another reverse primer for detecting the insertion template. This assay was performed as above with each reverse primer at half concentration. Example 9 High throughput Cloning of mutantsAttorney Docket: 761259: 083474-041PC

[0135] Expression constructs for Bxb1 integrase were cloned for mammalian expression via Gibson cloning using Hifi Assembly mix (NEB) according to the manufacturer’s instructions. Overlapping reverse and forward primer-carrying mutations for desired amino acids are used to amplify the plasmid around the globe with 18bp of homology. Then DPNI is used to clean up the plasmid from PCR reactions followed by column cleanup.50 ng of the cleaned-up PCR product is then used to perform Gibson reactions according to the manufacture’s protocol. For all Gibson clonings, 2 μL of assembled reactions were transformed into 20 μL of competent Stbl3 cells generated by Mix and Go! competency kit (Zymo) and plated on agar plates supplemented with appropriate antibiotics. After growth overnight at 37 °C, colonies were picked into TB medium (Thermo Fisher Scientific) and incubated with shaking at 37 °C for 24 h. Cultures were collected using a QIAprep Spin Miniprep kit (Qiagen) according to the manufacturer’s instructions. References

[0136] A. Elnaggar, H. Essam, W. Salah-Eldin, W. Moustafa, M. Elkerdawy, C. Rochereau, B. Rost, Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, arXiv [cs.LG] (2023). http: / / arxiv.org / abs / 2301.06568.

[0137] N. Brandes, D. Ofer, Y. Peleg, N. Rappoport, M. Linial, ProteinBERT: a universal deep-learning model of protein sequence and function. Bioinformatics 38, 2102–2110 (2022).

[0138] L. Brenan, A. Andreev, O. Cohen, S. Pantel, A. Kamburov, D. Cacchiarelli, N. S. Persky, C. Zhu, M. Bagul, E. M. Goetz, A. B. Burgin, L. A. Garraway, G. Getz, T. S. Mikkelsen, F. Piccioni, D. E. Root, C. M. Johannessen, Phenotypic Characterization of a Comprehensive Set of MAPK1 / ERK2 Missense Mutants. Cell Rep.17, 1171–1183 (2016).

[0139] P. Notin, A. W. Kollasch, D. Ritter, L. van Niekerk, S. Paul, H. Spinner, N. Rollins, A. Shaw, R. Weitzman, J. Frazer, M. Dias, D. Franceschi, R. Orenbuch, Y. Gal, D. S. Marks, ProteinGym: Large-Scale Benchmarks for Protein Design and Fitness Prediction. bioRxiv, doi: 10.1101 / 2023.12.07.570727 (2023).

[0140] T. Hino, S. N. Omura, R. Nakagawa, T. Togashi, S. N. Takeda, T. Hiramoto, S. Tasaka, H. Hirano, T. Tokuyama, H. Uosaki, S. Ishiguro, M. Kagieva, H. Yamano, Y. Ozaki, D. Motooka, H. Mori, Y. Kirita, Y. Kise, Y. Itoh, S. Matoba, H. Aburatani, N. Yachie, T. Karvelis, V. Siksnys, T. Ohmori, A. Hoshino, O. Nureki, An AsCas12f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Cell 186, 4920–4935.e23 (2023).Attorney Docket: 761259: 083474-041PC

[0141] H. K. Haddox, A. S. Dingens, J. D. Bloom, Experimental Estimation of the Effects of All Amino-Acid Mutations to HIV’s Envelope Protein on Viral Replication in Cell Culture. PLoS Pathog.12, e1006114 (2016).

[0142] E. D. Kelsic, H. Chung, N. Cohen, J. Park, H. H. Wang, R. Kishony, RNA Structural Determinants of Optimal Codons Revealed by MAGE-Seq. Cell Syst 3, 563–571.e6 (2016).

[0143] M. A. Stiffler, D. R. Hekstra, R. Ranganathan, Evolvability as a function of purifying selection in TEM-1 β-lactamase. Cell 160, 882–892 (2015).

[0144] C. J. Markin, D. A. Mokhtari, F. Sunden, M. J. Appel, E. Akiva, S. A. Longwell, C. Sabatti, D. Herschlag, P. M. Fordyce, Revealing enzyme functional architecture via high- throughput microfluidic enzyme kinetics. Science 373 (2021).

[0145] A. O. Giacomelli, X. Yang, R. E. Lintner, J. M. McFarland, M. Duby, J. Kim, T. P. Howard, D. Y. Takeda, S. H. Ly, E. Kim, H. S. Gannon, B. Hurhula, T. Sharpe, A. Goodale, B. Fritchman, S. Steelman, F. Vazquez, A. Tsherniak, A. J. Aguirre, J. G. Doench, F. Piccioni, C. W. M. Roberts, M. Meyerson, G. Getz, C. M. Johannessen, D. E. Root, W. C. Hahn, Mutational processes shape the landscape of TP53 mutations in human cancer. Nat. Genet.50, 1381–1387 (2018).

[0146] E. M. Jones, N. B. Lubock, A. J. Venkatakrishnan, J. Wang, A. M. Tseng, J. M. Paggi, N. R. Latorraca, D. Cancilla, M. Satyadi, J. E. Davis, M. M. Babu, R. O. Dror, S. Kosuri, Structural and functional characterization of G protein-coupled receptors with deep mutational scanning. Elife 9 (2020).

[0147] M. B. Doud, J. D. Bloom, Accurate Measurement of the Effects of All Amino-Acid Mutations on Influenza Hemagglutinin. Viruses 8 (2016).

[0148] J. M. Lee, J. Huddleston, M. B. Doud, K. A. Hooper, N. C. Wu, T. Bedford, J. D. Bloom, Deep mutational scanning of hemagglutinin helps predict evolutionary fates of human H3N2 influenza variants. Proc. Natl. Acad. Sci. U. S. A.115, E8276–E8285 (2018).

[0149] A. Rives, J. Meier, T. Sercu, S. Goyal, Z. Lin, J. Liu, D. Guo, M. Ott, C. L. Zitnick, J. Ma, R. Fergus, Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. U. S. A.118 (2021).

[0150] E. C. Alley, G. Khimulya, S. Biswas, M. AlQuraishi, G. M. Church, Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods 16, 1315–1322 (2019).

[0151] A. Elnaggar, M. Heinzinger, C. Dallago, G. Rehawi, Y. Wang, L. Jones, T. Gibbs, T. Feher, C. Angerer, M. Steinegger, D. Bhowmik, B. Rost, ProtTrans: Toward UnderstandingAttorney Docket: 761259: 083474-041PC the Language of Life Through Self-Supervised Learning. IEEE Trans. Pattern Anal. Mach. Intell.44, 7112–7127 (2022).

[0152] B. L. Hie, V. R. Shanker, D. Xu, T. U. J. Bruun, P. A. Weidenbacher, S. Tang, W. Wu, J. E. Pak, P. S. Kim, Efficient evolution of human antibodies from general protein language models. Nat. Biotechnol.42, 275–283 (2024).

[0153] M. T. N. Yarnall, E. I. Ioannidi, C. Schmitt-Ulms, R. N. Krajeski, J. Lim, L. Villiger, W. Zhou, K. Jiang, S. K. Garushyants, N. Roberts, L. Zhang, C. A. Vakulskas, J. A. Walker, A. P. Kadina, A. E. Zepeda, K. Holden, H. Ma, J. Xie, G. Gao, L. Foquet, G. Bial, S. K. Donnelly, Y. Miyata, D. R. Radiloff, J. M. Henderson, A. Ujita, O. O. Abudayyeh, J. S. Gootenberg, Drag-and-drop genome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases. Nat. Biotechnol., 1–13 (2022).

[0154] J. Meier, R. Rao, R. Verkuil, J. Liu, T. Sercu, A. Rives, Language models enable zero-shot prediction of the effects of mutations on protein function, bioRxiv (2021)p. 2021.07.09.450648.

[0155] M. Sourisseau, D. J. P. Lawrence, M. C. Schwarz, C. H. Storrs, E. C. Veit, J. D. Bloom, M. J. Evans, Deep Mutational Scanning Comprehensively Maps How Zika Envelope Protein Mutations Affect Viral Growth and Antibody Escape. J. Virol.93 (2019).

[0156] K. Clement, H. Rees, M. C. Canver, J. M. Gehrke, R. Farouni, J. Y. Hsu, M. A. Cole, D. R. Liu, J. K. Joung, D. E. Bauer, L. Pinello, CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat. Biotechnol.37, 224–226 (2019).

Claims

Attorney Docket: 761259: 083474-041PC What is claimed:

1. A high efficiency integrase or fragment thereof comprising an amino acid sequence with one or more mutations relative to wild-type Bxb1 integrase amino acid sequence, wherein the amino acid sequence is at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200-69221 and 69226-69340.

2. The high efficiency integrase or fragment thereof of claim 1, wherein the one or more mutations are located at one or more amino acids at positions 7, 22, 23, 25, 58, 67, 71, 72, 79, 101, 110, 115, 129, 141, 146, 166, 182, 189, 194, 230, 238, 240, 245, 247, 253, 258, 271, 275, 292, 318, 345, 375, 376, 381, 385, 422, 439, 459, 460, and 482 relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO:

11.

3. The high efficiency integrase or fragment thereof of claim 1, wherein the one or more mutations comprise 7L, 22V, 23K, 25A, 58T, 67C, 71G, 72A, 79K, 79R, 101G, 110M, 110N, 115K, 129W, 141D, 146A, 166A, 166C, 166E, 166Q, 166R, 182M, 189G, 189Q, 194A, 194C, 194F, 194G, 194M, 194W, 230R, 238M, 240C, 240G, 240H, 240L, 240M, 240Q, 240S, 240T, 240V, 245A, 247V, 253K, 258A, 271N, 275A, 275C, 275D, 275E, 275G, 275H, 275K, 275M, 275N, 275P, 275Q, 275R, 275S, 275Y, 292A, 292Q, 292S, 318F, 345H, 375A, 375C, 375D, 375E, 375F, 375G, 375Q, 375W, 375Y, 376Y, 381Q, 385A, 422P, 439A, 439C, 439D, 439E, 439G, 439H, 439I, 439K, 439N, 439P, 439Q, 439R, 439S, 439T, 439V, 459G, 460G, and / or 482L relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO:

11.

4. The high efficiency integrase or fragment thereof of claim 1, wherein the one or more mutations are located at one or more amino acids at positions: 194, 166, and 110; 194, 166, 240, and 110; 194, 166, and 240; 194, 166, 240, 459, and 110; 194, 166, 240, and 459; 194, 166, 240, 375, and 110; 194, 166, and 275;Attorney Docket: 761259: 083474-041PC 194, 166, 275, and 240; 194, 166, 275, 240, 375, and 110; 194, 166, 275, 240, and 375; 194, 166, 275, 375, and 110; 194, 166, 275, and 375; 194, 166, 459, and 110; 194, 166, 292, and 110; 194, 166, 292, 240, and 110; 194, 166, 292, and 240; 194, 166, 292, 240, 375, and 110; 194, 166, 292, 275, and 110; 194, 166, 292, and 275; 194, 166, 292, 275, 240, and 110; 194, 166, 292, 275, and 240; 194, 166, 292, 275, 240, and 375; 194, 166, 292, 275, 375, and 110; 166, 292, 275, and 240; 194, 166, and 240; 194, 166, 240, 459, and 110; 194, 166, 240, 375, and 110; 194, 166, and 275; 194, 166, 275, and 375; 194, 166, 275, 375, and 110; 194, 166, 292, and 240; or 194, 166, 292, 240, 375, and 110, relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO:

11.

5. The high efficiency integrase or fragment thereof of claim 1, wherein the one or more mutations comprise: 194G, 166R, and 110N; 194G, 166R, 240C, and 110N; 194G, 166R, and 240C; 194G, 166R, 240C, 459G, and 110N; 194G, 166R, 240C, and 459G;Attorney Docket: 761259: 083474-041PC 194G, 166R, 240C, 375C, and 110N; 194G, 166R, and 275E; 194G, 166R, 275E, and 240C; 194G, 166R, 275E, 240C, 375C, and 110N; 194G, 166R, 275E, 240C, and 375C; 194G, 166R, 275E, 375C, and 110N; 194G, 166R, 275E, and 375C; 194G, 166R, 459G, and 110N; 194G, 166R, 292S, and 110N; 194G, 166R, 292S, 240C, and 110N; 194G, 166R, 292S, and 240C; 194G, 166R, 292S, 240C, 375C, and 110N; 194G, 166R, 292S, 275E, and 110N; 194G, 166R, 292S, and 275E; 194G, 166R, 292S, 275E, 240C, and 110N; 194G, 166R, 292S, 275E, and 240C; 194G, 166R, 292S, 275E, 240C, and 375C; 194G, 166R, 292S, 275E, 375C, and 110N; 166R, 292S, 275E, and 240C; 194G, 166R, and 240C; 194G, 166R, 240C, 459G, and 110N; 194G, 166R, 240C, 375C, and 110N; 194G, 166R, and 275E; 194G, 166R, 275E, and 375C; 194G, 166R, 275E, 375C, and 110N; 194G, 166R, 292S, and 240C; or 194G, 166R, 292S, 240C, 375C, and 110N, relative to wild-type Bxb1 integrase amino acid sequence SEQ ID NO:

11.

6. The high efficiency integrase or fragment thereof of claim 1, wherein the high efficiency integrase or fragment thereof is capable of integrase, recombinase, or transposase activity.Attorney Docket: 761259: 083474-041PC 7. The high efficiency integrase or fragment thereof of claim 1, wherein the high efficiency integrase or fragment thereof is capable of binding to integration sites or pseudosites in a human genome sequence.

8. The high efficiency integrase or fragment thereof of claim 1, wherein each integration site is between 12 and 60 nucleotides in length.

9. The high efficiency integrase or fragment thereof of claim 1, wherein each integration site is between 18 and 50 nucleotides in length.

10. The high efficiency integrase or fragment thereof of claim 1, wherein the integration sites comprise an attB and attP pair of integration sites.

11. The high efficiency integrase or fragment thereof of claim 1, wherein the high efficiency integrase or fragment thereof binds to an attB integration site comprising at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, or 100% identity to an amino acid sequence set forth in any one of SEQ ID NOs: 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 55, and 69222.

12. The high efficiency integrase or fragment thereof of claim 1, wherein the high efficiency integrase or fragment thereof binds to an attP integration site comprising at least 80% identity, at least 85% identity, at least 90% identity, at least 95% identity, or 100% identity to an amino acid sequence set forth in any one of SEQ ID NOs: SEQ ID NOs: 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, and 69223.

13. A polynucleotide comprising a nucleic acid sequence encoding the integrase or fragment thereof of claim 1.

14. A vector comprising the nucleic acid sequence of claim 13.

15. A host cell comprising the vector of claim 14.

16. A system capable of site-specifically integrating an exogenous nucleic acid into a mammalian cell genome at a desired target site, wherein the system comprises:Attorney Docket: 761259: 083474-041PC (a) a nucleic acid encoding a DNA binding nickase domain linked to a reverse transcriptase (RT) domain; (b) a nucleic acid encoding a guide RNA (gRNA) comprising: i. a primer binding sequence, ii. a sequence complementary to one strand of an integration recognition sequence, and iii. a target binding sequence, (c) a nucleic acid encoding a serine integration enzyme; and (d) the exogenous nucleic acid linked to a sequence that is a cognate of the integration recognition sequence, wherein the system site specifically integrates the exogenous nucleic acid into the mammalian cell genome at the desired target site, and wherein the serine integration enzyme comprises an amino acid sequence at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200-69221 and 69226-69340.

17. The system of claim 16, wherein the integration recognition sequence is an attB or attP integration recognition sequence.

18. A method of site-specific integration of a nucleic acid into a genome of a cell, the method comprising: (a) incorporating an integration site at a desired location in the genome by introducing into the cell: i. a DNA binding nuclease linked to a reverse transcriptase domain, wherein the DNA binding nuclease comprises a nickase activity; and ii. a guide RNA (gRNA) comprising a primer binding sequence linked to an integration recognition sequence, wherein the gRNA interacts with the DNA binding nuclease and targets the desired location in the genome, wherein the DNA binding nuclease nicks a strand of the genome and the reverse transcriptase domain incorporates the integration recognition sequence of the gRNA into the nicked site, thereby providing the integration site at the desired location of the genome; andAttorney Docket: 761259: 083474-041PC (b) integrating the nucleic acid into the genome by introducing into the cell: i. a DNA or RNA strand comprising the nucleic acid linked to a sequence that is complementary or associated to the integration site; and ii. an integration enzyme, wherein the integration enzyme incorporates the nucleic acid into the genome at the integration site by integration, recombination, or reverse transcription of the sequence that is complementary or associated to the integration site, thereby introducing the nucleic acid into the desired location of the cell genome of the cell, wherein the integration enzyme comprises an amino acid sequence at least 80% identical, at least 85% identical, at least 90% identical, at least 95% identical, or 100% identical to an amino acid sequence set forth in any one of SEQ ID NOs: 69200-69221 and 69226-69340.

19. The system of claim 18, wherein the integration recognition sequence is an attB or attP integration recognition sequence.

Citation Information

Patent Citations

  • Systems, methods, and compositions for site-specific genetic engineering using programmable addition via site-specific targeting elements (paste)

    US11572556B2

  • Systems, methods, and compositions for site-specific genetic engineering using programmable addition via site-specific targeting elements (paste)

    US11827881B2

  • Systems, methods, and compositions for site-specific genetic engineering using programmable addition via site-specific targeting elements (paste)

    US11834658B2

  • Systems, methods, and compositions for site-specific genetic engineering using programmable addition via site-specific targeting elements (paste)

    US11952571B2

  • Systems, methods, and compositions for site-specific genetic engineering using programmable addition via site-specific targeting elements (paste)

    US12195733B2