Serine recombinase
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2026-03-25
Smart Images

Figure 00000090_0000 
Figure 00000091_0000 
Figure 00000091_0001
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 275,288, filed November 3, 2021, U.S. Provisional Application No. 63 / 322,712, filed March 23, 2022, and U.S. Provisional Application No. 63 / 400,868, filed August 25, 2022, the contents of which are incorporated herein by reference in their entireties. Statement on Federally Sponsored Research
[0002] This invention was made with government support under Grant Nos. OD021369 and AI148623 awarded by the National Institutes of Health. The U.S. Government has certain rights in this invention.
[0003] Sequence Listing Description The contents of the Electronic Sequence Listing entitled 39817_601_SequenceListing.xml (Size: 3,888,144 bytes, and Created: November 3, 2022) are incorporated herein by reference in their entirety.
[0004] Technical Field The present invention relates to serine recombinases, and methods for their identification and use. [Background technology]
[0005] Despite recent advances in genome engineering, efficient methods for stably integrating multi-kilobase DNA cargoes into human and other eukaryotic cells remain necessary. Large serine recombinases (LSRs), such as BxB1 and ΦC31, have evolved to perform this task in microbial cells; however, the LSRs characterized to date have several limitations that make them unsuitable for use in eukaryotic genome engineering. Directed evolution and protein engineering efforts have yet to successfully convert these limited candidates into ideal molecular tools. Expanding the tools available for genetic engineering requires new recombinases and methods for identifying new recombinases. Summary of the Invention
[0006] Provided herein are systems for DNA modification. In selected embodiments, the systems are cell-free.
[0007] In some embodiments, the system includes a polypeptide comprising a recombinase, or an active fragment thereof, or a nucleic acid encoding the same, having an amino acid sequence at least 70% identical to any of SEQ ID NOs: 1-74. In some embodiments, the recombinase has an amino acid sequence at least 70% identical to any of SEQ ID NOs: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66. In certain embodiments, the recombinase has the amino acid sequence of SEQ ID NO: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66.
[0008] In some embodiments, the system comprises: 1) X 1a X 2a X 3a X 4a X 5a X 6a X 7a X 8a X 9a X 10a X 11a X 12a X13a X 14a X 15a X 16a X 17a X 18a X 19a X 20a X 21a X 22a X 23a X 24a X 25a X 26a X 27a X 28a X 29a X 30a X 31a X 32a X 33a X 34a (In the formula: X 1a is A, E, I, L, S, T, V, or Y; X 2a is A, D, E, G, K, Q, R, S, or T; X 6a is E or G; X 8a is A, C, F, L, M, or V; X 10a is A, F, I, L, M, T, or V; X 13a is F, H, I, L, M, N, or V; X 14a is A, G, S, or V; X 15a is A, D, I, L, S, T, or V; X 17a is A, G, or S; X 21a is K, R, S, or V; X 22a is A, D, E, G, K, N, S, or T; X 23a is A, E, I, K, M, N, Q, S, or T; X 24a is F, I, L, M, S, or T; X 26a is D, E, L, Q, S, or V; X27a is E, N, Q, or R; X 32a is A, F, H, I, K, L, M, N, Q, R, S, or V; X 34a is A, E, G, H, K, L, M, N, Q, R, S, or V; and X 3a , X 4a , X 5a , X 7a , X 9a , X 11a , X 12a , X 16a , X 18a , X 19a , X 20a , X 25a , X 28a , X 29a , X 30a , X 31a , and X 33a are each independently selected from any amino acid; 2) X 1b X 2b X 3b X 4b X 5b X 6b X 7b X 8b X 9b X 10b X 11b X 12b X 13b X 14b X 15b X 16b X 17b X 18b , (In the ceremony X 1b is A, G, or I; X 2b is D, E, G, N, P, S, T, or V; X 3b is D, G, N, Q, or S; X 4b is A, H, N, Q, R, T, V, or Y; X 6b is A, D, E, H, I, L, P, Q, R, T, or Y; X 7bis A, D, E, Q, or R; X 8b is F, I, K, or L; X 10b is D, E, F, G, N, Q, R, S, T, or V; X 11b is A, I, L, S, T, or V; X 12b is D, E, I, K, L, N, Q, R, S, T, or V; X 13b is A, D, E, K, M, N, R, S, T, or V; X 14b is A, G, Q, R, S, or T; X 16b is A, D, E, K, L, Q, R, or T, X 18b is A, L, M, or V; and X 5b , X 9b , X 15b , and X 17b are each independently selected from any amino acid; 3) X 1c X 2c X 3c X 4c X 5c X 6c X 7c X 8c X 9c X 10c X 11c X 12c X 13c ESX 16c X 17c KX 19c X 20c X 21c X 22c X 23c X 24c X 25c X 26c (In the formula: X 1c is A, D, F, I, L, M, N, S, or Y; X 4c is A, I, K, M, S, or V; X 6cis A, F, G, I, L, M, or V; X 10c is Q, R, or T; X 11c is A, G, or S; X 13c is D, E, G, N, Q, or S; X 17c is A, H, K, N, R, S, T, or V; X 21c is L, M, R, or Y; X 22c is A, I, N, Q, S, T, or V; X 23c is A, E, F, I, K, L, N, R, T, or V; X 25c is A, F, H, L, N, Q, S, T, or Y; X 26c is A, I, L, M, N, R, S, T, V, or Y; and X 2c , X 3c , X 5c , X 7c , X 8c , X 9c , X 12c , X 16c , X 19c , X 20c , and X 24c are each independently selected from any amino acid; 4) X 1d X 2d X 3d X 4d X 5d X 6d X 7d X 8d X 9d X 10d X 11d X 12d X 13d X 14d X 15d X 16d X 17d X 18d X 19d X 20d X 21d X 22d X 23d X24d X 25d X 26d X 27d X 28d (In the formula: X 1d is E, K, N, T, G, S, L, D, V, A, R, or P; X 2d is E, H, I, T, G, S, L, D, V, A, or P; X 4d is M, I, T, S, L, V, A, R or P; X 5d is E, K, N, I, T, G, S, D, Q, V, A, R, or P; X 6d is E, G, S, D, A, R, or P; X 7d is I, L, D, A, or R; X 8d is M, H, K, T, L, V, Q, D, A, or R; X 9d is E, K, I, T, G, S, L, D, Q, V, or A; X 10d is E, K, H, D, Q, V, A, or R; X 11d is M, H, I, S, L, V, Q, A, or R; X 12d is Q, E, K, N, M, S, L, D, V, A, or R; X 13d is E, K, H, G, S, L, D, Q, A, or R; X 14d is E, Y, K, N, I, H, L, V, or A; X 16d is E, K, I, T, G, S, L, D, Q, A, or R; X 17d is E, K, H, T, G, D, Q, A, or R; X 19d is Q, E, K, N, T, G, S, D, V, A, or R; X 20dis Q, E, K, N, T, G, S, V, D, A, or R; X 21d is I, S, W, L, V, F, A, or R; X 22d is Q, E, M, T, G, S, L, V, D, or A; X 23d is E, K, N, I, T, G, S, D, A, R, or P; X 24d is E, M, I, L, D, Q, or A; X 25d is E, Y, I, L, V, F, A, or R; X 26d is E, M, T, G, S, L, D, V, A, or R; X 27d is E, K, N, G, S, L, D, Q, A, or R; X 28d is Q, E, G, V, D, A, R, or P; and X 3d , X 15d , and X 18d are each independently selected from any amino acid; 5) X 1e X 2e X 3e X 4e X 5e X 6e X 7e X 8e X 9e X 10e X 11e X 12e X 13e X 14e X 15e X 16e X 17e X 18e (In the formula: X 1e is A, D, E, H, K, N, Q, R, or S; X 2e is A, D, E, F, G, H, K, M, N, Q, R, S, W, or Y; X 3e is E, F, or Y; X4e is F, H, L, W, or Y; X 6e is A, D, E, F, I, K, L, M, N, Q, R, S, T, or Y; X 7e is F, I, Q, S, T, or V; X 8e is A, G, K, L, N, R, S, T, or V; X 9e is A, D, E, H, K, N, Q, R, T, or Y; X 10e is I, N, Q, or R; X 11e is F, I, L, M, Q, or S; X 14e is A, G, K, N, or S; X 15e is K, M, Q, R, S, T, or V; X 18e is A, E, G, K, M, N, S, T, or Y; and X 5e , X 12e , X 13e , X 16e , and X 17e are each independently selected from any amino acid; 6)WX 2f X 3f X 4f X 5f X 6f X 7f X 8f X 9f X 10f X 11f X 12f X 13f X 14f X 15f X 16f GX 18f X 19f X 20f X 21f X 22f X 23f (In the formula: X 2f is A, E, H, N, R, S, T, or V; X4f is A, G, N, S, or T; X 5f is F, G, L, M, N, Q, S, T, or V; X 6f is I, L, P, or V; X 9f is I, L, T, or V; X 14f is A, C, G, M, Q, R, S, or T; X 16f is I, L, V, or Y; X 18f is D, E, H, N, Q, or S; X 20f is E, H, I, L, M, Q, R, or T; X 21f is A, E, F, H, L, N, P, or Y; X 22f is C, F, H, K, M, N, Q, R, T, or Y; X 23f is D, E, F, I, K, L, N, Q, R, S, T, or V; and X 3f , X 7f , X 8f , X 10f , X 11f , X 12f , X 13f , X 15f , and X 19f are each independently selected from any amino acid; 7) X 1g X 2g X 3g X 4g X 5g EX 7g X 8g X 9g X 10g X 11g X 12g RX 14g X 15g X 16g X 17g X 18g X 19g X 20g X 21g (In the formula: X 1g is A, G, I, N, S, T, or V; X 3g is A, I, or S; X 5g is F, I, L, M, or Y; X 7g is I or R; X 10g is D, I, L, or T; X 12g is A, E, I, K, M, Q, or S; X 14g is I, T, or V; X 16g is A, D, G, R, S, or T; X 18g is F, K, L, M, or Y; X 19g is A, E, H, I, K, L, M, N, Q, R, V, W, or Y; X 21g is A, I, K, L, M, or R; and X 2g , X 4g , X 8g , X 9g , X 11g , X 15g , X 17g , and X 20g are each independently selected from any amino acid; 8) X 1h X 2h X 3h X 4h X 5h X 6h X 7h X 8h X 9h X 10h X 11h (In the formula: X 1h is F or Y; X 2h is D, E, K, Q, or S; X 3h is E, K, L, M, or Q; X4h is K, L, or R; X 5h is K, L, or V; X 7h is G or N; X 8h is D, E, H, K, L, M, or R; X 9h is S or T; X 11h is F, H, I, Q, S, T, V, or W; and X 6h and X 10h are each independently selected from any amino acid; 9) X 1i X 2i X 3i X 4i X 5i X 6i X 7i X 8i X 9i X 10i X 11i SX 13i X 14i X 15i X 16i X 17i X 18i X 19i X 20i X 21i X 22i X 23i X 24i X 25i X 26i X 27i (In the formula: X 1i is I, L, or V; X 4i is A, D, F, H, I, L, M, N, Q, S, V, or Y; X 8i is A, G, or S; X 10i is D, E, I, K, N, Q, R, or S; X 11i is E or Q; X 15i is A or K; X16i is A, Q, R, or S; X 18i is L, M, or R; X 19i is I, L, Q, R, S, or V; X 21i is A, D, E, G, H, I, Q, R, or S; X 22i is A, K, N, Q, S, T, or V; X 23i is A, H, K, R, W, or Y; X 25i is A, G, H, I, K, Q, R, S, or T; X 27i is C, H, I, K, L, R, or V; and X 2i , X 3i , X 5i , X 6i , X 7i , X 9i , X 13i , X 14i , X 17i , X 20i , X 24i , and X 26i are each independently selected from any amino acid; 10)RX 2j X 3j X 4j W (In the formula: X 2j is L, M, Q, or R; X 3j is A, N, or S; X 4j is N, P, S, or T); 11) X 1k X 2k X 3k X 4k X 5k X 6k X 7k X 8k F (In the formula: X 1k is I, L, or V; X2k is A or V; X 4k is A, F, H, I, L, Q, W, or Y; X 5k is I, M, or V; X 7k is E, L, Q, or T; X 8k is A, I, or V; and X 3k and X 6k are each independently selected from any amino acid; 12)RX 2l X 3l X 4l X 5l X 6l X 7l X 8l X 9l X 10l X 11l X 12l X 13l (In the formula: X 2l is D, K, N, R, S, or V; X 3l is A, D, E, F, G, K, P, Q, or S; X 4l is A, E, I, K, L, S, T, or V; X 5l is any amino acid; X 6l is F, G, I, L, N, or V; X 7l is A, F, I, L, Q, R, V, or Y; X 8l is D, E, I, L, M, N, Q, S, T, or V; X 9l is D, E, F, I, L, M, Q, T, V, or Y; X 10l is I, K, L, R, or V; X 11l is D, E, K, N, Q, or R; X 12lis D, E, F, K, L, N, Q, W, or Y; and X 13l is F or L); and 13) X 1m X 2m X 3m X 4m X 5m X 6m X 7m X 8m X 9m X 10m X 11m X 12m X 13m X 14m X 15m X 16m X 17m X 18m X 19m X 20m X 21m X 22m X 23m X 24m (In the formula: X 1m is A, E, F, I, L, M, N, Q, S, T, V, or Y; X 2m is A, F, G, I, L, M, R, S, T, or V; X 6m is A, D, E, F, G, H, L, M, N, S, or T; X 9m is D, M, N, or S; X 10m is D, E, or Q; X 12m is C, F, H, L, T, V, or Y; X 14m is A, E, K, L, R, or Y; X 17m is A, L, or S; X 19m is D, E, K, N, Q, R, or S; X 20m is G, I, M, Q, R, T, or V; X 21m is D, H, K, N, Q, or R; X23m is A, G, I, L, N, S, T, or V; X 24m is F, H, I, K, L, M, N, Q, V, W, or Y; and X 3m , X 4m , X 5m , X 7m , X 8m , X 11m , X 13m , X 15m , X 16m , X 18m , and X 22m, are each independently selected from any amino acid), or an active fragment thereof, or a nucleic acid encoding the same; and a first polynucleotide comprising a donor recognition sequence for a recombinase.
[0009] In some embodiments, the system comprises a polypeptide comprising a recombinase having an amino acid sequence having at least 70% identity to SEQ ID NOs: 88-1183.
[0010] The system may further include a first polynucleotide comprising a donor recognition sequence for a recombinase. In some embodiments, the donor recognition sequence comprises a donor attachment site configured to bind to the recombinase. The recognition site is a polynucleotide sequence that includes any sequence element that facilitates recognition by the recombinase enzyme. The attachment site is a specific polynucleotide sequence at which recombination occurs.
[0011] In some embodiments, the first polynucleotide further comprises a cargo DNA sequence, which is the polynucleotide to be delivered or inserted into the target sequence. The cargo DNA sequence may be larger than 1 kilobase pair (e.g., larger than 2 kilobase pairs, larger than 4 kilobase pairs, larger than 6 kilobase pairs, larger than 8 kilobase pairs, larger than 10 kilobase pairs, larger than 15 kilobase pairs, larger than 20 kilobase pairs, or more). In selected embodiments, the cargo DNA sequence is larger than 5 kilobase pairs.
[0012] In some embodiments, the first polynucleotide further comprises a recipient recognition sequence for a recombinase. In some embodiments, the system further comprises a second polynucleotide comprising a recipient recognition sequence for a recombinase. In some embodiments, the recipient recognition sequence comprises a recipient attachment sequence configured to bind to the recombinase.
[0013] In some embodiments, the donor recognition sequence, the recipient recognition sequence, or both are pseudo-recognition sequences. A "pseudo-recognition sequence" or "pseudo-site" refers to a recognition sequence that need not be the native recognition sequence of a given recombinase, but rather is sufficient to promote recombination.
[0014] Also provided herein are compositions and cells that comprise the disclosed systems. In some embodiments, the cells are eukaryotic cells.
[0015] Further provided herein are methods for modifying target DNA.
[0016] In some embodiments, the methods include contacting the target DNA with a polypeptide comprising a recombinase, or an active fragment thereof, or a nucleic acid encoding the same, having an amino acid sequence having at least 70% identity to any of SEQ ID NOs: 1-74. In some embodiments, the recombinase has an amino acid sequence having at least 70% identity to any of SEQ ID NOs: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66. In certain embodiments, the recombinase has the amino acid sequence of SEQ ID NO: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66.
[0017] In some embodiments, the method comprises: 1) X 1a X 2a X 3a X 4a X 5a X 6a X 7a X 8a X 9a X 10a X 11a X 12a X 13a X 14a X 15a X 16a X 17a X 18a X 19a X 20a X 21a X 22a X 23a X 24a X 25a X 26a X 27a X 28a X 29a X 30a X 31a X 32a X 33a X 34a (In the formula: X 1a is A, E, I, L, S, T, V, or Y; X 2a is A, D, E, G, K, Q, R, S, or T; X 6a is E or G; X 8a is A, C, F, L, M, or V; X 10a is A, F, I, L, M, T, or V; X 13a is F, H, I, L, M, N, or V; X 14a is A, G, S, or V; X 15a is A, D, I, L, S, T, or V; X 17a is A, G, or S; X 21a is K, R, S, or V; X 22a is A, D, E, G, K, N, S, or T; X 23a is A, E, I, K, M, N, Q, S, or T; X 24a is F, I, L, M, S, or T; X 26a is D, E, L, Q, S, or V; X 27a is E, N, Q, or R; X 32a is A, F, H, I, K, L, M, N, Q, R, S, or V; X 34a is A, E, G, H, K, L, M, N, Q, R, S, or V; and X 3a , X 4a , X 5a , X 7a , X 9a , X 11a , X 12a , X 16a , X 18a , X 19a , X 20a , X 25a , X 28a , X 29a , X 30a , X 31a , and X 33a are each independently selected from any amino acid; 2) X 1bX 2b X 3b X 4b X 5b X 6b X 7b X 8b X 9b X 10b X 11b X 12b X 13b X 14b X 15b X 16b X 17b X 18b , (In the ceremony X 1b is A, G, or I; X 2b is D, E, G, N, P, S, T, or V; X 3b is D, G, N, Q, or S; X 4b is A, H, N, Q, R, T, V, or Y; X 6b is A, D, E, H, I, L, P, Q, R, T, or Y; X 7b is A, D, E, Q, or R; X 8b is F, I, K, or L; X 10b is D, E, F, G, N, Q, R, S, T, or V; X 11b is A, I, L, S, T, or V; X 12b is D, E, I, K, L, N, Q, R, S, T, or V; X 13b is A, D, E, K, M, N, R, S, T, or V; X 14b is A, G, Q, R, S, or T; X 16b is A, D, E, K, L, Q, R, or T, X 18b is A, L, M, or V; and X 5b , X 9b , X 15b, and X 17b are each independently selected from any amino acid; 3) X 1c X 2c X 3c X 4c X 5c X 6c X 7c X 8c X 9c X 10c X 11c X 12c X 13c ESX 16c X 17c KX 19c X 20c X 21c X 22c X 23c X 24c X 25c X 26c (In the formula: X 1c is A, D, F, I, L, M, N, S, or Y; X 4c is A, I, K, M, S, or V; X 6c is A, F, G, I, L, M, or V; X 10c is Q, R, or T; X 11c is A, G, or S; X 13c is D, E, G, N, Q, or S; X 17c is A, H, K, N, R, S, T, or V; X 21c is L, M, R, or Y; X 22c is A, I, N, Q, S, T, or V; X 23c is A, E, F, I, K, L, N, R, T, or V; X 25c is A, F, H, L, N, Q, S, T, or Y; X 26c is A, I, L, M, N, R, S, T, V, or Y; and X2c , X 3c , X 5c , X 7c , X 8c , X 9c , X 12c , X 16c , X 19c , X 20c , and X 24c are each independently selected from any amino acid; 4) X 1d X 2d X 3d X 4d X 5d X 6d X 7d X 8d X 9d X 10d X 11d X 12d X 13d X 14d X 15d X 16d X 17d X 18d X 19d X 20d X 21d X 22d X 23d X 24d X 25d X 26d X 27d X 28d (In the formula: X 1d is E, K, N, T, G, S, L, D, V, A, R, or P; X 2d is E, H, I, T, G, S, L, D, V, A, or P; X 4d is M, I, T, S, L, V, A, R or P; X 5d is E, K, N, I, T, G, S, D, Q, V, A, R, or P; X 6d is E, G, S, D, A, R, or P; X 7d is I, L, D, A, or R; X 8d is M, H, K, T, L, V, Q, D, A, or R; X 9d is E, K, I, T, G, S, L, D, Q, V, or A; X 10d is E, K, H, D, Q, V, A, or R; X 11d is M, H, I, S, L, V, Q, A, or R; X 12d is Q, E, K, N, M, S, L, D, V, A, or R; X 13d is E, K, H, G, S, L, D, Q, A, or R; X 14d is E, Y, K, N, I, H, L, V, or A; X 16d is E, K, I, T, G, S, L, D, Q, A, or R; X 17d is E, K, H, T, G, D, Q, A, or R; X 19d is Q, E, K, N, T, G, S, D, V, A, or R; X 20d is Q, E, K, N, T, G, S, V, D, A, or R; X 21d is I, S, W, L, V, F, A, or R; X 22d is Q, E, M, T, G, S, L, V, D, or A; X 23d is E, K, N, I, T, G, S, D, A, R, or P; X 24d is E, M, I, L, D, Q, or A; X 25d is E, Y, I, L, V, F, A, or R; X 26d is E, M, T, G, S, L, D, V, A, or R; X 27d is E, K, N, G, S, L, D, Q, A, or R; X 28d is Q, E, G, V, D, A, R, or P; and X 3d , X15d , and X 18d are each independently selected from any amino acid; 5) X 1e X 2e X 3e X 4e X 5e X 6e X 7e X 8e X 9e X 10e X 11e X 12e X 13e X 14e X 15e X 16e X 17e X 18e (In the formula: X 1e is A, D, E, H, K, N, Q, R, or S; X 2e is A, D, E, F, G, H, K, M, N, Q, R, S, W, or Y; X 3e is E, F, or Y; X 4e is F, H, L, W, or Y; X 6e is A, D, E, F, I, K, L, M, N, Q, R, S, T, or Y; X 7e is F, I, Q, S, T, or V; X 8e is A, G, K, L, N, R, S, T, or V; X 9e is A, D, E, H, K, N, Q, R, T, or Y; X 10e is I, N, Q, or R; X 11e is F, I, L, M, Q, or S; X 14e is A, G, K, N, or S; X 15e is K, M, Q, R, S, T, or V; X 18e is A, E, G, K, M, N, S, T, or Y; and X 5e , X 12e , X 13e , X 16e , and X 17e are each independently selected from any amino acid; 6)WX 2f X 3f X 4f X 5f X 6f X 7f X 8f X 9f X 10f X 11f X 12f X 13f X 14f X 15f X 16f GX 18f X 19f X 20f X 21f X 22f X 23f (In the formula: X 2f is A, E, H, N, R, S, T, or V; X 4f is A, G, N, S, or T; X 5f is F, G, L, M, N, Q, S, T, or V; X 6f is I, L, P, or V; X 9f is I, L, T, or V; X 14f is A, C, G, M, Q, R, S, or T; X 16f is I, L, V, or Y; X 18f is D, E, H, N, Q, or S; X 20f is E, H, I, L, M, Q, R, or T; X 21f is A, E, F, H, L, N, P, or Y; X 22f is C, F, H, K, M, N, Q, R, T, or Y; X 23fis D, E, F, I, K, L, N, Q, R, S, T, or V; and X 3f , X 7f , X 8f , X 10f , X 11f , X 12f , X 13f , X 15f , and X 19f are each independently selected from any amino acid; 7) X 1g X 2g X 3g X 4g X 5g EX 7g X 8g X 9g X 10g X 11g X 12g RX 14g X 15g X 16g X 17g X 18g X 19g X 20g X 21g (In the formula: X 1g is A, G, I, N, S, T, or V; X 3g is A, I, or S; X 5g is F, I, L, M, or Y; X 7g is I or R; X 10g is D, I, L, or T; X 12g is A, E, I, K, M, Q, or S; X 14g is I, T, or V; X 16g is A, D, G, R, S, or T; X 18g is F, K, L, M, or Y; X 19g is A, E, H, I, K, L, M, N, Q, R, V, W, or Y; X 21gis A, I, K, L, M, or R; and X 2g , X 4g , X 8g , X 9g , X 11g , X 15g , X 17g , and X 20g are each independently selected from any amino acid; 8) X 1h X 2h X 3h X 4h X 5h X 6h X 7h X 8h X 9h X 10h X 11h (In the formula: X 1h is F or Y; X 2h is D, E, K, Q, or S; X 3h is E, K, L, M, or Q; X 4h is K, L, or R; X 5h is K, L, or V; X 7h is G or N; X 8h is D, E, H, K, L, M, or R; X 9h is S or T; X 11h is F, H, I, Q, S, T, V, or W; and X 6h and X 10h are each independently selected from any amino acid; 9) X 1i X 2i X 3i X 4i X 5i X 6i X 7i X 8i X 9i X 10i X 11i SX13i X 14i X 15i X 16i X 17i X 18i X 19i X 20i X 21i X 22i X 23i X 24i X 25i X 26i X 27i (In the formula: X 1i is I, L, or V; X 4i is A, D, F, H, I, L, M, N, Q, S, V, or Y; X 8i is A, G, or S; X 10i is D, E, I, K, N, Q, R, or S; X 11i is E or Q; X 15i is A or K; X 16i is A, Q, R, or S; X 18i is L, M, or R; X 19i is I, L, Q, R, S, or V; X 21i is A, D, E, G, H, I, Q, R, or S; X 22i is A, K, N, Q, S, T, or V; X 23i is A, H, K, R, W, or Y; X 25i is A, G, H, I, K, Q, R, S, or T; X 27i is C, H, I, K, L, R, or V; and X 2i , X 3i , X 5i , X 6i , X 7i , X 9i , X 13i , X14i , X 17i , X 20i , X 24i , and X 26i are each independently selected from any amino acid; 10)RX 2j X 3j X 4j W (In the formula: X 2j is L, M, Q, or R; X 3j is A, N, or S; X 4j is N, P, S, or T; 11) X 1k X 2k X 3k X 4k X 5k X 6k X 7k X 8k F (In the formula: X 1k is I, L, or V; X 2k is A or V; X 4k is A, F, H, I, L, Q, W, or Y; X 5k is I, M, or V; X 7k is E, L, Q, or T, X 8k is A, I, or V; and X 3k and X 6k are each independently selected from any amino acid; 12)RX 2l X 3l X 4l X 5l X 6l X 7l X 8l X 9l X 10l X 11l X 12l X 13l (In the formula: X 2lis D, K, N, R, S, or V; X 3l is A, D, E, F, G, K, P, Q, or S; X 4l is A, E, I, K, L, S, T, or V; X 5l is any amino acid; X 6l is F, G, I, L, N, or V; X 7l is A, F, I, L, Q, R, V, or Y; X 8l is D, E, I, L, M, N, Q, S, T, or V; X 9l is D, E, F, I, L, M, Q, T, V, or Y; X 10l is I, K, L, R, or V; X 11l is D, E, K, N, Q, or R; X 12l is D, E, F, K, L, N, Q, W, or Y; and X 13l is F or L); and 13) X 1m X 2m X 3m X 4m X 5m X 6m X 7m X 8m X 9m X 10m X 11m X 12m X 13m X 14m X 15m X 16m X 17m X 18m X 19m X 20m X 21m X 22m X 23m X 24m (In the formula: X 1m is A, E, F, I, L, M, N, Q, S, T, V, or Y; X 2m is A, F, G, I, L, M, R, S, T, or V; X 6m is A, D, E, F, G, H, L, M, N, S, or T; X 9m is D, M, N, or S; X 10m is D, E, or Q; X 12m is C, F, H, L, T, V, or Y; X 14m is A, E, K, L, R, or Y; X 17m is A, L, or S; X 19m is D, E, K, N, Q, R, or S; X 20m is G, I, M, Q, R, T, or V; X 21m is D, H, K, N, Q, or R; X 23m is A, G, I, L, N, S, T, or V; X 24m is F, H, I, K, L, M, N, Q, V, W, or Y; and X 3m , X 4m , X 5m , X 7m , X 8m , X 11m , X 13m , X 15m , X 16m , X 18m , and X 22m, are each independently selected from any amino acid), or an active fragment thereof, or a nucleic acid encoding the same.
[0018] In some embodiments, the method includes contacting the target DNA with a polypeptide comprising a recombinase having an amino acid sequence having at least 70% identity to any of SEQ ID NOs: 88-1183, an active fragment thereof, or a nucleic acid encoding the same.
[0019] In some embodiments, the target DNA comprises a donor recognition sequence, a recipient recognition sequence, or both. In certain embodiments, the target DNA comprises a recipient attachment sequence configured to bind to a recombinase.
[0020] In some embodiments, the method further comprises contacting the target DNA with a first polynucleotide comprising a donor recognition sequence for a recombinase.
[0021] In some embodiments, the first polynucleotide further comprises a cargo DNA sequence. The cargo DNA sequence may be larger than 1 kilobase pair (e.g., larger than 2 kilobase pairs, larger than 4 kilobase pairs, larger than 6 kilobase pairs, larger than 8 kilobase pairs, larger than 10 kilobase pairs, larger than 15 kilobase pairs, larger than 20 kilobase pairs, or more). In selected embodiments, the cargo DNA sequence is larger than 5 kilobase pairs.
[0022] In some embodiments, the donor recognition sequence, the recipient recognition sequence, or both, are pseudorecognition sequences.
[0023] In some embodiments, the target DNA sequence encodes a gene product, hi certain embodiments, the target DNA sequence is a genomic DNA sequence.
[0024] In some embodiments, the target DNA is in a cell. In certain embodiments, the cell is a eukaryotic cell (e.g., a human cell or a plant cell). In certain embodiments, the cell is a prokaryotic cell.
[0025] In some embodiments, the contacting comprises introducing one or more components of the system into the cell, hi some embodiments, the recombinase or a nucleic acid encoding same is introduced into the cell before, simultaneously with, or after introduction of the donor polynucleotide.
[0026] In some embodiments, introducing into cells comprises administering one or more components of the system to a subject (e.g., a human). In certain embodiments, administering comprises in vivo administration. In certain embodiments, administering comprises transplantation of ex vivo treated cells comprising one or more components of the system.
[0027] Other aspects and embodiments of the present disclosure will become apparent in light of the following detailed description. [Brief explanation of the drawings]
[0028] [Figure 1A] Systematic identification of thousands of recombinases and their predicted attachment sites from site-specific and multitargeting / transposition families. Schematic diagram of the computational workflow for identifying LSRs and attachment sites. Briefly, protein sequences contained in RefSeq and GenBank bacterial isolate genomes were searched to identify sequences containing a "recombinase" (PF07508) domain. Genomes containing such proteins were compared with genomes lacking this protein to determine whether the recombinase resides on an integrated mobile genetic element. Once the boundaries of this MGE were identified, the original attachment site was reconstructed by examining the sequences flanking these boundaries. This workflow was an extension of a previous, smaller-scale computational method (Yang et al. 2014 Nat Methods. 11(12): 1261-1266, incorporated herein by reference in its entirety). [Figure 1B]This figure shows the systematic identification of thousands of recombinases and their predicted attachment sites in site-specific and multitargeting / transposition families. This figure shows a phylogenetic tree of amino acid sequences representing LSR families annotated according to the predicted target specificity of each LSR cluster. The figure legend "Unique Integration Targets" specifies the number of predicted target protein families found in the database to be targeted by each LSR cluster. Families labeled "1" were identified using the techniques described in Figure 1C. Families labeled "2," "3," or ">3" were identified as described in the panel in Figure 1F. The upper right portion of the phylogenetic tree shown here reveals prominent multitargeting families. The size of each dot indicates the number of unique sequences found in each LSR cluster. [Figure 1C] Systematic identification of thousands of recombinases and their predicted attachment sites from site-specific and multi-targeting / transposition families is shown. Schematic diagram of an exemplary technique for identifying site-specific LSRs. Briefly, all LSR families are considered site-specific if multiple LSR clusters (clustered at 50% identity) are integrated into a single gene cluster (clustered at 50% identity). A typical domain structure of a site-specific LSR is shown on the right and includes a resolvase (green), a recombinase (red), and a recombinase zinc beta ribbon domain (purple). [Figure 1D] Figure 1 shows the systematic identification of thousands of recombinases and their predicted attachment sites across site-specific and multi-targeting / transposition families. Figure 2 shows an exemplary network of predicted site-specific LSRs. Each node represents either an LSR cluster (red) or a target protein cluster (blue). Edges between nodes indicate that at least one member of the target protein cluster was found to integrate into at least one member of the target protein cluster. [Figure 1E]This figure shows the systematic identification of thousands of recombinases and their predicted attachment sites from site-specific and multitargeting / transposition families. This figure shows an exemplary hierarchical tree of diverse LSR sequences targeting a set of closely related attB sequences. The tree is constructed according to the distance between LSRs as a function of the percentage of identical amino acids after alignment. The alignment of related attB sequences is shown below in no particular order. At the end of the tree, numbers indicate the attB sequence targeted by each LSR. The attB alignment is colored according to consensus sequence similarity, with gray indicating matches to the consensus sequence, four unique colors indicating single-nucleotide mismatches from the consensus, and black indicating alignment gaps. [Figure 1F] Systematic identification of thousands of recombinases and their predicted attachment sites across site-specific and multitargeting / transposition families is shown. Schematic diagram of an exemplary technique for identifying multitargeting LSRs. Briefly, an LSR cluster is considered multitargeting if a single cluster of related LSRs (clustered at 90% identity) is integrated into multiple diverse target protein families (clustered at 50% identity). A typical domain architecture of a multitargeting LSR, including the addition of a domain of unknown function (yellow, DUF4368), is shown on the right. [Figure 1G] Systematic identification of thousands of recombinases and their predicted attachment sites from site-specific and multitargeting / transposition families. Example of observed network of predicted multitargeting LSRs. Node colors and sizes are the same as in Figure 1D. [Figure 1H]Systematic identification of thousands of recombinases and their predicted attachment sites across site-specific and multitargeting / transposition families is shown. Alignment of diverse attB sequences targeted by a single multitargeting LSR. Each target sequence is aligned against the core TT dinucleotide. A sequence logo is shown above the alignment to demonstrate conservation across sequence sites and suggest sequence specificity for this particular LSR. As in Figure 1E, the alignment is colored according to consensus. [Figure 2A] Characterization of the new landing pad LSR. Schematic of an exemplary plasmid recombination assay. Cells are co-transfected with LSR-2A-GFP, promoterless attP-mCherry, and EF1a-attB. Upon recombination, mCherry acquires the EF1a promoter and is expressed. [Figure 2B] Characterization of the new landing pad LSR. Plasmid recombination assay of the predicted LSR and att sites in HEK293FT cells. Fold change in mCherry mean fluorescence intensity (MFI) of all single cells compared to Bxb1 is shown. Points indicate the mean, and error bars indicate standard deviation (n=3 transfection replicates). [Figure 2C] Characterization of the new landing pad LSR. Exemplary mCherry distribution for all three plasmids (LSR+attB+attP) compared to the attP-only negative control. Cells are not gated for transfection delivery markers. [Figure 2D] Characterization of the new landing pad LSR: Plasmid recombination assay between all pairs of LSR+attP and attB in K562 cells (n=1). [Figure 2E]Characterization of the new landing pad LSR. Schematic diagram of an exemplary genomic landing pad assay. The EF1a promoter, attB, and LSR were integrated into the genome of K562 cells via low MOI lentivirus to form a single landing pad copy per cell. Next, a clonal cell line was electroporated with the attP-mCherry donor plasmid. Successful integration into the landing pad resulted in expression of mCherry and knockout of LSR and GFP. [Figure 2F] Characterization of the new landing pad LSR. Flow cytometry of mCherry+ cells 11 days after donor electroporation with 1000 ng of donor plasmid. Each dot represents a different clonal K562 cell line harboring the corresponding landing pad and LSR for the donor. When comparing donor electroporation conditions, PaO1 is significantly more efficient than BxB1 (**=P<0.005, one-way ANOVA). [Figure 2G] Characterization of the new landing pad LSR. Flow cytometry showing knockout of LSR-GFP and integration of mCherry in the same cells. Donor delivery was increased by electroporating the Pa01 clonal landing pad line twice, resulting in >70% mCherry+ cells. [Figure 2H] Characterization of the new landing pad LSR. Flow cytometry of mCherry+ cells 18 days after co-electroporation of the LSR and donor into WTK562 cells lacking a landing pad. The attD donor contains its own EF1a promoter, and the attD donor alone serves as a negative control. [Figure 2I]Characterization of new landing pad LSRs is shown. Genome-wide integration site mapping by next-generation sequencing is shown to measure the percentage of reads found within the genome outside of the expected landing pad. Raw (non-unique) reads found off-target are shown as a percentage of all reads (*=P<0.05, one-tailed t-test). For Kp03, Ec03, and Pa01, n=2 independent clonal landing pad lines with maximum mCherry yield 11 days after donor electroporation. For Bxb1, two technical replicates of a single clonal landing pad line with maximum mCherry yield 11 days after donor electroporation are shown. The number near the top of each bar indicates the total number of unique off-target reads divided by the total number of off-target loci. [Figure 2J] Characterization of the new landing pad LSR. Plasmid recombination assay of the second batch of predicted LSR and att sites in HEK293FT cells. Fold change in mCherry mean fluorescence intensity (MFI) of all single cells compared to Bxb1 is shown. Points indicate the mean, and error bars indicate standard deviation (n = 3 transfection replicates). [Figure 2K] Characterization of the new landing pad LSR. As shown, exemplary mCherry distributions for the three plasmids (LSR + attB + attP) compared to the attP-only negative control. Cells were not gated for transfection delivery markers. [Figure 2L]Characterization of the new landing pad LSR. Graph of promoterless mCherry donor integration efficiency into the polyclonal genomic landing pad (LP) K562 cell line, measured after 5 days (n=2, independently transduced and subsequently electroporated biological replicates). Asterisks indicate statistical significance of landing pad and donor conditions compared to Bxb1 (one-way ANOVA with Dunnett's multiple comparisons; *, P<0.05; ***, P<0.001; ****, P<0.0001; ns, not significant). [Figure 2M] Characterization of the new landing pad LSR is shown. Donor plasmid integration into clonal landing pad cell lines electroporated with 1000 ng of donor plasmid (10 days post-electroporation, left) or 3000 ng of donor plasmid (11 days post-electroporation, right) is shown. When comparing donor electroporation conditions, 1000 ng PaO1 is significantly more efficient than 1000 ng Bxb1 (P<0.005, one-way ANOVA; n = 3 clonal cell lines for PaO1, n = 4 clonal cell lines for the others, using one electroporation per clone at the 1000 ng dose; n = 2 clonal cell lines per LSR, using two electroporation replicates per clone at the 3000 ng dose; error = s.e.m.). Dots on the left represent individual clones, and dots on the right represent electroporation replicates; individual clones are vertically aligned separately. [Figure 2N] Characterization of the new landing pad LSR. Representative mCherry distributions of the three plasmids (LSR + attB + attP) compared to the attP-only negative control are shown, as indicated. [Figure 3A]This shows that genome-targeting LSRs can integrate into the human genome at the predicted target site. This is a schematic diagram of a computational strategy for identifying LSRs with a natural affinity for the human genome. Briefly, BLAST was used to search attB / attP candidates in the database against the human genome. The attachment site with the best match in the human genome was renamed attA (acceptor), and the target site in the human genome was renamed attH (human). Attachment sites that did not match the genome became attD (donor). [Figure 3B] This demonstrates that genome-targeting LSRs can integrate into the human genome at their predicted target sites. Figure 3B shows BLAST hits for attB / P sites homologous to sequences in the human genome. Attachment sites for quality-controlled LSR predictions were searched against the human genome using BLAST. All hits meeting E<0.01 are shown. Four candidates that were subsequently experimentally shown to integrate at their predicted target sites in integration site mapping assays are shown in red. The 22 autosomes are shown, starting with chromosome 1 in dark blue on the left and alternating with light blue on every other chromosome. [Figure 3C] This shows that genome-targeting LSRs can integrate into the human genome at their predicted target sites. Figure 4A shows the results of a plasmid recombination assay for LSRs with predicted pseudosites using their cognate predicted attachment sites. Candidates in red are considered active LSRs with predicted pseudosites (one-tailed t-test, P<0.05), while gray candidates are considered inactive candidates with predicted pseudosites (P>0.05). Controls and candidates validated in integration site mapping assays are highlighted. Some of these candidates did not meet the overall quality control filters but were selected due to the high similarity between their attachment sites and the human genome. An analysis of how validation rates vary with candidate quality is shown in Figure 4A. [Figure 3D]This shows that genome-targeting LSRs can integrate into the human genome at their predicted target sites. BLAST alignments of the microbial attachment site (attA) and the predicted human attachment site (attH) for three candidates are shown (SEQ ID NOs: 3494-3499 for attA and attH in Sp56, Pf80, and Enc3, respectively). attA is shown at the top of each alignment, and attH is shown at the bottom. [Figure 3E] This shows that genome-targeting LSRs can integrate into the human genome at the predicted target sites. This figure shows the results of an integration site mapping experiment to determine true integration at the predicted target sites. Integration sites are ranked based on the number of unique reads found at each site. For Sp56 and Pf80, the locus with the most reads matched the predicted locus. For Enc3, the predicted locus was not the most frequently targeted locus, but was still validated as the true integration site. [Figure 3F] Figure 1 shows that the genome-targeting LSR can integrate into the human genome at the predicted target site. Reads aligning to the integration site of Pf80 in the human genome (reads aligning in the forward (red) and reverse (blue) directions, with black lines connecting paired reads) are shown, with the predicted target site indicated. [Figure 3G] This shows that genome-targeted LSRs can integrate into the human genome at their predicted target sites. This graph shows the results of human integration assays of top candidates from the latest batch of LSR candidates. While on-target integration was detectable with previous genome-targeted candidates, overall integration efficiencies remained very low. With the new set of predicted genome-targeted candidates, Dn29 and Vp82 emerged as top candidates with corrected integration efficiencies of 4.5% (+ / - 0.13%) and 2.52% (+ / - 0.004%), respectively. PhiC31 is a known genome-targeted LSR used as a control, but its efficiency is below the detection limit (approximately 1% of cells). Bars represent means, and dots represent individual transfections. Error = standard deviation (* = P < .05, one-tailed t-test). [Figure 3H] This shows that genome-targeting LSRs can integrate into the human genome at predicted target sites. Integration site mapping results for Dn29 and Vp82 are shown. The top three targeted human genome sites are labeled in each panel. The most commonly targeted site in Dn29 accounted for approximately 17% of the detected reads, suggesting that this candidate has a favorable combination of efficiency and specificity. [Figure 3I] Figure 3A shows that genome-targeting LSRs can integrate into the human genome at predicted target sites. Figure 3B shows the target site motifs of the top 25 human genome target sites for genome targeting candidate Dn29. The attA sites are SEQ ID NOs: 3500-3503 from top to bottom. Figure 3C shows the target site motifs of the top 25 human genome target sites for genome targeting candidate Vp82. The attA sites are SEQ ID NOs: 3504-3507 from top to bottom. [Figure 3J] This shows that genome-targeting LSR can be integrated into the human genome at the predicted target site. The target site motifs of the top 25 human genome target sites of genome-targeting candidate Vp82 are shown. The attA sites are SEQ ID NOs: 3504 to 3507 from top to bottom. [Figure 3K] We demonstrate that genome-targeting LSRs can integrate into the human genome at their predicted target sites. Figure 3K shows LSR combination specificity versus efficiency. Black dots indicate integration in wild-type cells, while green dots indicate integration in cells with pre-established landing pads (Figure 2E). Selected LSRs are labeled. For wild-type cells, efficiency is estimated as the percent of mCherry+ cells 18 days after electroporation with mCherry-expressing donor plasmids modified by LSR and donor-only control transfections. For landing pad cells, efficiency is estimated as the average of mCherry+ cells across all clones (Figure 2G, right). To estimate specificity, we used UMI counts when available; otherwise, we used uniquely mapped read counts; counts were merged across replicates. [Figure 3L] The top three integration sites of Dn29 are shown in genomic context. Red lines indicate the exact location of the integration, and introns and exons of nearby genes are shown in blue. [Figure 4A] This shows that multi-targeting LSR is highly efficient and reusable. Figure 1 shows the co-transfection of LSR Cp36 and attD-mCherry donor plasmids into K562 cells without a landing pad. Bxb1 paired with the Cp36 attD donor was used as a negative control. The dose in ng refers to the LSR plasmid, and the attD donor plasmid was delivered at a 1:1 molar ratio. [Figure 4B] This demonstrates that multi-targeting LSR is highly efficient and reusable. Figure 4C shows the results of a Cp36 integration site mapping assay. In this experiment, an integration locus is defined as the integration of the donor cargo detected at a specific locus. The top 500 loci across two experiments are shown, one performed in HEK293FT cells and the other in K562 cells. Unique reads provide a conservative count estimate for loci with higher coverage. The sequences of the sites indicated by arrows are shown at the bottom of Figure 4C. [Figure 4C] This demonstrates that multi-targeting LSR is highly efficient and reusable. Examples of Cp36 target site motifs and target sequences. In HEK293FT and K562 experiments, the exact integration site and direction were predicted at all loci, and the nucleotide composition of the top 200 sites was calculated. The core dinucleotide is found in the center. Examples of integration sites, colored according to nucleotide, are shown below (SEQ ID NOs: 3508-3512). [Figure 4D]This demonstrates that multi-targeting LSRs are highly efficient and reusable. Figure 4D shows the efficiency of Cp36 versus PiggyBac (PB) for stable delivery of mCherry donor plasmids in K562 cells 10 days after transfection. The donor plasmid contains both the Cp36 attD and PiggyBac ITRs, and the Ec03 LSR is used as a negative control lacking the attachment site on this donor plasmid. [Figure 4E] 1 shows that multi-targeting LSR is highly efficient and reusable. 1 shows the mCherry incorporation efficiency of Cp36 with and without Cp36 re-administration at day 15. [Figure 4F] This demonstrates that multitargeting LSR is highly efficient and reusable. Figure 11D shows graphs of wild-type K562 or Cp36-treated mCherry and puromycin-selected cells transfected with a second fluorescent reporter (mTagBFP2) and analyzed by flow cytometry 13 days after electroporation with 2000 ng of BFP donor and an equimolar dose of 1600 ng of Cp36 plasmid. Bars indicate means, dots indicate replicates, error = s.e.m. (n = 2 electroporation replicates). Dashes indicate negative controls treated with BFP donor only. Corresponding mCherry levels are shown in Figure 11D. [Figure 4G] This demonstrates that multi-targeting LSR is highly efficient and reusable. Flow cytometry analysis 12 days after electroporation of both fluorescent donor and Cp36 plasmid into K562 cells. Negative control cells were transfected with donor and pUC19. Error = sem (n = 2 electroporation replicates). [Figure 5A]Phylogenetic tree of the 1,081 identified LSR clusters (50% identity). Tips are colored according to the phylum of the bacterial host species. The first heatmap ring, as in Figure 1B, is colored according to the number of unique target gene clusters each LSR cluster is predicted to incorporate. The second ring, annotated in green, indicates LSR clusters predicted to contain the DUF4368 Pfam domain. Clusters from the controls Bxb1 and PhiC31 are shown in bold, as are selected candidate clusters that were experimentally validated. [Figure 5B] The most commonly found Pfam domains in target genes are shown. Each target gene was annotated using the Pfam HMM model, and the total number of LSR clusters embedded in genes containing each Pfam domain was calculated. [Figure 5C] Figure 1E shows an alignment of the LSR sequences shown. The resolvase, recombinase, and Zn_recomb_ribbon Pfam domains are shown. The height and color of each bar above each aligned amino acid position indicates the average pairwise identity across all pairs in the column, with green indicating 100% identity across all sequences, green-brown indicating greater than 30% identity and less than 100% identity, and red indicating less than 30% identity. [Figure 5D] Exemplary predicted attB motifs are shown. Each column represents a different LSR attB motif. The first row shows motifs derived from various attB sequences targeted by a single unique LSR protein. The second row shows motifs derived from attB sequences targeted by LSR proteins that fall into a single 90% identity cluster. The third row shows motifs derived from attB sequences targeted by LSR proteins that fall into a single 50% identity cluster. [Figure 5E]Pfam domain enrichment analysis of target genes. Pfam domains that reach a significance cutoff of FDR<0.05 are shown. Pfam domains are ordered and displayed according to the -log10(P) value of Fisher's exact test. The number next to each point indicates the total number of target gene clusters containing the specified domain. [Figure 5F] Gene Ontology (GO) term enrichment analysis of target genes. All six terms that reach a significance cutoff of FDR<0.1 are shown. Term domains are ordered and displayed according to the -log10(P) value of Fisher's exact test. The number next to each point indicates the total number of target gene clusters that fall into the specified GO term. [Figure 5G] The distance between the target gene and the nearest phage defense gene is shown. For each target gene that appears on the flanking sequence with a defense gene, the distance is calculated, and then a random gene from the same flanking sequence is selected as a background control. Boxplots are shown with the median, first and third quartiles, 1.5xIQR as whiskers, and outliers as points. Significant differences between groups are tested using the Wilcoxon rank sum test. [Figure 6A]Characterization of the landing pad LSR. This graph shows the efficiency of promoterless mCherry donor integration into genomic landing pads (LPs) in K562 cells, as measured by flow cytometry. The landing pad and donor constructs are the same as those shown in Figure 2E, but here, polyclonal landing pad lines were derived by high MOI delivery of lentiviral landing pads without subsequent selection or sorting. 1.2 million K562 cells were electroporated with 600 ng of donor plasmid carrying the attP corresponding to the LSR and assayed 5 days later (n = 2 independently transduced and subsequently electroporated biological replicates). Asterisks indicate statistical significance of the landing pad and donor conditions compared to BxB1 (one-way ANOVA with Dunnett's multiple comparison test; *, P < 0.05; ***, P < 0.001; ****, P < 0.0001; ns, not significant). [Figure 6B] Characterization of landing pad LSR. Graph of the stability of polyclonal landing pads expressing LSR-GFP as measured over time by flow cytometry. These cells were not electroporated with a donor, and day 5 was the same measurement day as in Figure 6D (n=2 independently transduced biological replicates). [Figure 6C] Characterization of landing pad LSR. Flow cytometry measuring mCherry+ cells 10 days after electroporation of 2000 ng of donor plasmid. Each dot represents a different clonal K562 cell line harboring the corresponding landing pad and LSR for the donor. Error bars indicate standard deviation for conditions containing multiple clones. [Figure 6D] Characterization of landing pad LSR. Flow cytometry analysis of mCherry+ cells 12 days after electroporation of 2000 or 5000 ng of donor plasmid into a clonal K562 cell line carrying the landing pad. Error bars indicate standard deviation (n=3 electroporation replicates indicated by dots). [Figure 6E] Characterization of the landing pad LSR is shown. Minimization of the Pa01 attB sequence by trimming nucleotides from either end and using a plasmid recombination assay is shown. The arrow indicates the shortest attB that did not interfere with recombination activity. The estimated 33-bp minimal attB determined by this experiment is shown between the bottom vertical lines in the indicated sequence number 3513. The colored rectangle indicates the corrected average mCherry MFI (n = 3 transfection replicates in HEK293FT cells). The attB in the top rectangle extends in both directions and is the full-length attB obtained from the LSR database and used in Figures 2B-2C. [Figure 6F] Characterization of the landing pad LSR is shown. Minimization of the Kp03 attB sequence by trimming nucleotides from both ends using a plasmid recombination assay is shown. The shortest attB tested was 25 nucleotides. The colored rectangle indicates the average mCherry MFI normalized to the MFI of attD alone (n=3). The attB in the top rectangle extends in both directions and is the full-length attB obtained from the LSR database and used in Figures 2B-2C. The dinucleotide core determined by off-target integration site mapping is shown in bold within the indicated sequence number 3514. [Figure 6G] Characterization of landing pad LSR. Graph of Kp03 dinucleotide core exchange in a plasmid recombination assay to determine the ability to program specific matches between donor and acceptor attachment sites by altering the core. AC is the native dinucleotide core sequence. Values are the mean ± SD of n=3 transfection replicates in HEK293FT cells. [Figure 6H] Landing pad LSR characterization. Target site motifs of the top 25 human genomic target sites for landing pad candidates Kp03 (top) and Pa01 (bottom). The core dinucleotide is highly conserved between the integration sites of both candidates. [Figure 6I]Characterization of landing pad LSR. Schematic of the optimized integration site mapping assay, a modified version of UdiTaS. The addition of a series of amplifications using nested donor primers is expected to enrich for reads derived from the desired target, including both donor-only and donor-genome junction reads. [Figure 6J] Characterization of landing pad LSR is shown. Figure 6J is a graph showing the percentage of reads derived from different sources in an integration site mapping assay. The left side shows the percentage before assay optimization, and the right side shows the percentage after optimization. Both experiments are Cp36 circular donor experiments, but in two different cell types (HEK293FT on the left and K562 on the right). Target-derived reads are either donor-only reads (light green) or donor-genome integration junction reads (dark green). [Figure 6K] Characterization of the landing pad LSR. Flow cytometry measuring mCherry+ cells 18 days after co-electroporation of the LSR and donor into WTK562 cells lacking a landing pad. The attD donor contains its own EF1a promoter, and the attD donor alone serves as a negative control. [Figure 6L] Characterization of landing pad LSRs is shown. Results of a plasmid recombination assay of predicted LSRs and att sites in HEK293FT cells are shown as the percentage of mCherry+ cells gated on GFP-positive cells. mCherry and GFP gating was determined based on empty backbone transfection. Dots indicate each transfection replicate; error = sd (n = 3 transfection replicates). [Figure 6M]Characterization of landing pad LSRs. Graph showing GFP+ cell fraction in clonal cell lines 27 days after transduction. GFP+ cells were generated by sorting into wells as single cells, expanded for 2 weeks, and measured by flow cytometry. Populations were graded as GFP+ if they were >95% GFP+, suggesting a lack of transcriptional silencing. 16 wells were sorted for each LSR, and the number of wells with a viable cell population during flow analysis is indicated in the legend. For all LSRs, some wells were empty, likely due to missorting or cell death. [Figure 6N] Characterization of the landing pad LSR. Flow cytometry graph measuring mCherry+ cells 18 days after donor co-electroporation into WTK562 cells lacking the LSR and landing pad. The attD donor contains the EF-1α promoter driving mCherry expression, while an attD donor transfected with a mismatched LSR serves as a negative control (*=P<0.05, **=P<0.005, one-tailed t-test). (Errors = s.d. n=2 transfection replicates). [Figure 6O]Characterization of landing pad LSRs is shown. Genome-wide integration site mapping by next-generation sequencing is performed to measure the percentage of reads found within the genome outside the expected landing pad. For Kp03 and Ec03, n = 2 independent clonal landing pad lines were used, and for Pa01, n = 3 clonal landing pad lines were used, with mCherry maximal at 11 days after donor electroporation. For Bxb1, three technical replicates (started from different gDNA aliquots) of a single clonal landing pad line with maximal mCherry at 11 days after donor electroporation are shown. Raw (non-unique) reads found at off-target sites are shown as a percentage of all reads (* = P < 0.05, one-tailed t-test). The numbers near the top of each bar indicate the total number of off-target loci on the left, and the numbers in parentheses below indicate the subset of sites replicated in the landing pad cell line (left) and the subset replicated in the wild-type cell line (right). [Figure 7A] Genome targeting characteristics. Graph of the percentage of LSRs that mediate significant recombination in a plasmid recombination assay, with and without applying a quality control (QC) threshold for LSR candidate selection. The numbers above each bar indicate the number of candidates meeting P<0.05 in the plasmid recombination assay divided by the total number of candidates tested. [Figure 7B]
[0033] Figure 1 shows genome targeting characteristics. Graph of plasmid recombination assay for top genome targeting candidates using predicted attH sites. [Figure 7C] Genomic targeting features are shown. Reads aligning to the Sp56 integration site in the human genome are shown (reads aligning in forward (red) and reverse (blue) orientations; black lines connect paired reads). Using a linear donor alters the direction and location of integration, while a circular donor targets the exact predicted integration site. [Figure 7D]Genome targeting features are shown. Reads aligning to the Enc3 integration site in the human genome are shown (reads aligning in forward (red) and reverse (blue) orientations; black lines connect paired reads). Using a linear donor alters the direction and location of integration, while a circular donor targets the exact predicted integration site. [Figure 7E] Genome targeting features are shown. Target site motifs for Dn29 are shown. Each row shows motifs for a different subset of integration sites. [Figure 7F] Genome targeting features are shown. Target site motifs for Vp82 are shown. Each row shows motifs for a different subset of integration sites. [Figure 8A] 16 is a graph of Cp36 mCherry donor cargo integration in K562 cells without pre-installation of a landing pad or antibiotic selection utilizing both plasmid DNA and linear PCR amplicons as donor cargo. [Figure 8B] Figure 1 shows a graph of additional multitargeting LSRs validated using the pseudosite integration assay. Two additional candidates are shown: Pc01 and Enc9, both of which are found in the multitargeting family. [Figure 8C] FIG. 1 is a schematic diagram of the integration sites found for Cp36 using an integration site mapping assay. [Figure 8D] Schematic of the plasmid recombination assay with exchanged att sites and results for Cp36 compared to multiple landing pad LSRs. [Figure 8E] FIG. 1 is a schematic diagram of an exemplary plasmid used for direct comparison of Cp36 with PiggyBac, which contains both the PB inverted terminal repeats (ITRs) and Cp36 attD. [Figure 9]Schematic diagram of the canonical (can.) LSR integration mechanism. Briefly, the LSR protein (composed of three distinct domains and a coiled-coil structural motif) recognizes the attP sequence of nucleotides on the donor plasmid and the attB sequence on the target genome. Four LSR monomers assemble to catalyze recombination between the two attachment sites, resulting in a unidirectional reaction that forms the final integration product. [Figure 10] A phylogenetic tree of identified LSRs is shown, with phylogenetic clades containing two or more experimentally active LSRs descending from a common ancestor. [Figure 11A] Figure 1 shows that the multi-targeting recombinase is an efficient, unidirectional integrase. Correlation between read counts from Cp36 integration site mapping assays across HEK293FT and K562 cell lines is shown. The top 61 shared loci are shown, all of which are found in the top 200 most frequently targeted sites in the two cell types. Gray bands indicate 95% confidence intervals. [Figure 11B] Multitargeting recombinases are shown to be efficient, unidirectional integrases. Enrichment of target sites in DNase hypersensitive peaks for several multitargeters is shown. Statistical significance of each enrichment was calculated using Fisher's exact test. P values and the number of associated integration sites are shown above each associated lane. Error bars indicate 95% confidence intervals. [Figure 11C] We demonstrate that multitargeting recombinases are efficient, unidirectional integrases. The predicted target site motif is shown using 33 attB sequences in the LSR attachment site database that are targeted by LSRs and fall within the same 50% amino acid identity cluster as Cp36. The method used to construct this motif is the same as in Figures 1H and 5G. [Figure 11D]This demonstrates that multitargeting recombinases are efficient, unidirectional integrases. The schematic on the left shows p36 re-administration experiments in which mCherry+ cells were generated using Cp36 and mCherry donors, followed by re-administration of the Cp36 enzyme or an empty LSR-expressing backbone, followed by measuring the excision potential of the mCherry cargo by flow cytometry. The right shows the average percentage of mCherry+ cells at day 18, as measured by flow cytometry (n=2 transfection replicates). [Figure 11E] This demonstrates that the multitargeting recombinase is an efficient, unidirectional integrase. Delivery of the BFP donor alone is shown. K562 cells were electroporated with 2400 ng of Cp36 plasmid and 3000 ng of BFP donor plasmid, and BFP was measured by flow cytometry 12 days later. Dashes indicate non-electroporated cells; the Cp36 or donor-only conditions included the pUC19 stuffer plasmid, so the delivered mass was equal. Bars indicate averages, and dots indicate replicates. [Figure 11F] This demonstrates that the multitargeting recombinase is an efficient, unidirectional integrase. Cp36-treated mCherry+ and puromycin-selected cells were analyzed by flow cytometry 13 days after electroporation with 2000 ng of BFP donor and an equimolar dose of 1600 ng of Cp36 plasmid (or pUC19 stuffer plasmid). Bars represent the mean, and dots represent replicates (error = sem n = 2 electroporation replicates). Dashes represent non-electroporated controls. [Figure 12A]This figure shows the post-hoc identification of human genome integration sites using database sequence motifs. The performance of database-derived sequence motifs for predicting human genome integration sites, as measured by ROC curve analysis, is also shown. Sequence motifs for each LSR were automatically generated from bacterial sequence databases by selecting nonredundant (95% nucleotide identity) attB sequences of related LSR orthologs. These motifs were then searched against true integration sites and randomly selected background sequences using HOMER motif analysis software. ROC curves were generated by sliding across a relevant range of motif score cutoffs and calculating the false positive rate (x-axis) and true positive rate (y-axis) at each cutoff. The area under the curve (AUC) was then calculated as a single measure of predictive performance. Each ROC curve is labeled with the associated LSR name and the number of integration sites detected across all relevant experiments. [Figure 12B] Post-hoc identification of human genome integration sites using database sequence motifs is shown. Distribution of normalized HOMER motif scores for experimentally observed integration sites ("Obs.") versus randomly selected background sequences ("Rand.") is shown. Boxplots are shown with the median, first and third quartiles, 1.5xIQR as whiskers, and outliers as points. A one-sided Wilcoxon rank-sum test is used to test for significant differences between groups (**, P<0.01; ****, P<0.0001; ns, not significant). Red dots indicate the normalized HOMER motif score of the observed integration site with the most experimentally detected integration events compared to all other integration sites for each LSR. [Figure 12C] Figure 1 shows post-hoc identification of human genome integration sites using database sequence motifs. The final sequence motif used to predict the human genome integration site for each LSR is shown. Each sequence is labeled with the associated LSR, the number of attB sequences used to construct the motif, and the average percentage of amino acid identity of all LSR orthologs used to identify the associated attB sequence. DETAILED DESCRIPTION OF THE INVENTION
[0029] Herein, we describe large serine recombinases (LSRs) identified along with their cognate DNA attachment sites using a computational workflow. LSRs were characterized according to three distinct technical applications: 1) landing pad LSRs, which can efficiently integrate at pre-established integration sites; 2) multi-targeting LSRs, which can efficiently integrate at many different loci within a target genome; and 3) genome-targeting LSRs, which can integrate at one or more specific target sites within a given target genome. Several candidates from all three categories were validated in human cells. For landing pad LSRs, we identified many candidates that recombine with high efficiency at orthogonal attachment sites compared to the existing gold standard, Bxb1. For multi-targeting LSRs, which have not previously been developed as integration tools for human cells, we identified several that can integrate with high efficiency in human cell lines compared to ΦC31. For genome-targeting LSRs, we identified and validated several candidates that integrate DNA cargo into predicted human genome target sites without pre-established attachment sites.
[0030] Recombinases have wide applications as genome engineering tools. However, efficiently integrating large donor sequences into the human genome remains an unsolved problem in the field of human genome engineering. One major hurdle is the cargo size limitation of adeno-associated virus (AAV) vectors, the most successful vectors available for human genome engineering, which is approximately 4.7 kilobase pairs (kb). CRISPR-Cas9 can be used to introduce double-strand breaks at programmable locations, but subsequent homologous recombination to introduce new DNA results in an exponential decrease in integration efficiency as the insert size increases, although the reported maximum insert size is 3–6 kb. In contrast, recombinases have no apparent upper limit on the size of the integrated donor DNA, which is a major advantage of recombinases over other technologies.
[0031] The section headings used in this section and throughout this disclosure are for organizational purposes only and are not intended to be limiting.
[0032] 1.Definition As used herein, the terms "comprise," "include," "having," "has," "can," "containing," and variations thereof are intended to be open-ended transitional phrases, terms, or words that do not exclude the possibility of additional acts or structures. The singular forms "a," "and," and "the" include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments that "comprising," "consisting of," and "consisting essentially of" the embodiments or elements presented herein, whether explicitly stated or not. As used herein, comprising a particular sequence or a particular SEQ ID NO: typically means that at least one copy of the sequence is present in the recited peptide or polynucleotide. However, more than one copy is also contemplated.
[0033] For the recitation of numerical ranges herein, each intervening number therebetween to the same degree of precision is expressly contemplated. For example, for the range of 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.
[0034] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings commonly understood by those skilled in the art. The meaning and scope of terms shall be clear, but in the event of any potential ambiguity, the definitions provided herein shall take precedence over any dictionary or external definitions. Further, unless otherwise required by context, singular terms shall include the plural and plural terms shall include the singular.
[0035] As used herein, "nucleic acid" or "nucleic acid sequence" refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, at 793-800 (Worth Pub. 1982)). The present technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. The polymer or oligomer can be heterogeneous or homogeneous in composition and can be isolated from naturally occurring sources or can be artificially or synthetically produced. Furthermore, the nucleic acid can be DNA or RNA, or a mixture thereof, and can exist permanently or transiently in single- or double-stranded form, including homoduplexes, heteroduplexes, and hybrid states. In some embodiments, the nucleic acid or nucleic acid sequence comprises other types of nucleic acid structures, such as, for example, a DNA / RNA helix, a peptide nucleic acid (PNA), a morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002), and U.S. Patent No. 5,034,506), a locked nucleic acid (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97:5633-5638 (2000)), a cyclohexenyl nucleic acid (see Wang, J. Am. Chem. Soc., 122:8595-8602 (2000)), and / or a ribozyme.Thus, the term "nucleic acid" or "nucleic acid sequence" can encompass strands containing non-natural nucleotides, modified nucleotides, and / or non-nucleotide building blocks (e.g., "nucleotide analogs") that can perform the same function as natural nucleotides; furthermore, as used herein, the term "nucleic acid sequence" refers to oligonucleotides, nucleotides, or polynucleotides, and fragments or portions thereof, and DNA or RNA of genomic or synthetic origin, which may be single- or double-stranded and represent sense or antisense strands. The terms "nucleic acid," "polynucleotide," "nucleotide sequence," and "oligonucleotide" are used interchangeably. They refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs.
[0036] A "peptide" or "polypeptide" is a linked sequence of two or more amino acids linked by peptide bonds. A peptide or polypeptide can be natural, synthetic, or a modified or combination of natural and synthetic. Polypeptides include proteins such as binding proteins, receptors, and antibodies. Proteins may be modified by the addition of sugars, lipids, or other moieties not included in the amino acid chain. The terms "polypeptide" and "protein" are used interchangeably herein.
[0037] As used herein, the term "percent sequence identity" refers to the percentage of nucleotides or nucleotide analogs in a nucleic acid sequence, or amino acids in an amino acid sequence, that are identical to the corresponding nucleotides or amino acids in a reference sequence after aligning the two sequences and introducing gaps, if necessary, to achieve the maximum percent identity. Thus, if a nucleic acid obtained by this technique is longer than the reference sequence, additional nucleotides in the nucleic acid that do not align with the reference sequence are not considered in determining sequence identity. Many mathematical algorithms for obtaining optimal alignments and calculating identity between two or more sequences are known and are incorporated into many available software programs. Examples of such programs include CLUSTAL-W, T-Coffee, and ALIGN (for aligning nucleic acid and amino acid sequences), BLAST programs (e.g., BLAST2.1, BL2SEQ, and later versions thereof), and FASTA programs (e.g., FASTA3x, FAS™, and SSEARCH) (for sequence alignment and sequence similarity searches).Sequence alignment algorithms are also described in, for example, Altschul et al., J. Molecular Biol., 215(3): 403-410 (1990), Beigert et al., Proc. Natl. Acad. Sci. USA, 106(10): 3770-3775 (2009), Durbin et al., eds., Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press, Cambridge, UK (2009), Soding, Bioinformatics, 21(7): 951-960 (2005), Altschul et al., Nucleic Acids Res., 25(17): 3389-3402 (1997), and Gusfield, Algorithms on Strings, Trees and Sequences, Cambridge University Press, Cambridge UK. (1997)).
[0038] As used herein, the term "amino acid" or "any amino acid" refers to any and all amino acids, including naturally occurring amino acids (e.g., α-amino acids), unnatural amino acids, modified amino acids, and non-natural amino acids. Both D- and L-amino acids are included. Natural amino acids include those found in nature, such as the 23 amino acids that combine in peptide chains to form the building blocks of many proteins. These are primarily L-stereoisomers, although a small number of D-amino acids are present in bacterial envelopes and some antibiotics. "Non-standard" naturally occurring amino acids include, for example, pyrrolysine (present in methanogens and other eukaryotes), selenocysteine (present in many non-eukaryotes and most eukaryotes), and N-formylmethionine (encoded by the start codon AUG in bacteria, mitochondria, and chloroplasts). "Unnatural" or "non-natural" amino acids are non-proteinogenic amino acids (e.g., amino acids not naturally encoded or found in the genetic code), either naturally occurring or chemically synthesized. Over 140 unnatural amino acids are known, with thousands more possible combinations. Examples of "unnatural" amino acids include the β-amino acids (β 3 and β 2 ), homoamino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, straight-chain core amino acids, diamino acids, D-amino acids, alpha-methyl amino acids, and N-methyl amino acids. Unnatural or non-natural amino acids also include modified amino acids. "Modified" amino acids include amino acids (e.g., natural amino acids) that have been chemically modified to include groups or chemical moieties that do not naturally occur on the amino acid.
[0039] The names of naturally occurring and non-naturally occurring aminoacyl residues used herein largely follow the naming conventions proposed by the IUPAC Commission on Organic Chemical Nomenclature and the IUPAC-IUB Commission on Biochemical Nomenclature, as set forth in "Nomenclature of α-Amino Acids (Recommendations, 1974)" Biochemistry, 14(2), (1975). Where the names and abbreviations of amino acids and aminoacyl residues used in this specification and the appended claims differ from these proposals, this will be clarified.
[0040] Throughout this specification, naturally occurring amino acids are designated by their conventional three-letter or one-letter abbreviations (e.g., Ala or A for alanine, Arg or R for arginine, etc.) unless they are referred to by their full name (e.g., alanine, arginine, etc.). As used herein, the term "L-amino acid" refers to the "L" isomer of a peptide, and conversely, the term "D-amino acid" refers to the "D" isomer of a peptide (e.g., Dphe, (D)Phe, D-Phe, or DF for the D isomer of phenylalanine). D-isomer amino acid residues can be substituted for any L-amino acid residue, so long as the desired function is retained by the peptide.
[0041] For less common or non-naturally occurring amino acids, unless they are referred to by their full name (e.g., sarcosine, ornithine, etc.), the residues are designated by the frequently used three- or four-letter codes, including Sar or Sarc (sarcosine, i.e., N-methylglycine), Aib (α-aminoisobutyric acid), Dab (2,4-diaminobutanoic acid), Dapa (2,3-diaminopropanoic acid), γ-Glu (γ-glutamic acid), Gaba (γ-aminobutanoic acid), β-Pro (pyrrolidine-3-carboxylic acid), and 8Ado (8-amino-3,6-dioxaoctanoic acid), Abu (2-aminobutyric acid), βhPro (β-homoproline), βhPhe (β-homophenylalanine) and Bip (β,β-diphenylalanine), and Ida (iminodiacetic acid).
[0042] The term "pharmaceutically acceptable salt" in the context of the present invention means
[0043] The terms "non-naturally occurring," "engineered," and "synthetic" are used interchangeably and indicate the involvement of the hand of man. When referring to a nucleic acid molecule or polypeptide, these terms mean that the nucleic acid molecule or polypeptide is at least substantially free from at least one other component with which it is naturally associated and found in nature.
[0044] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, to which another DNA segment, e.g., an "insert," can be attached or incorporated so as to bring about the replication of the attached segment in a cell.
[0045] A cell has been "genetically modified," "transformed," or "transfected" by exogenous DNA, such as a recombinant expression vector, when that DNA has been introduced inside the cell. The presence of the exogenous DNA results in a permanent or transient genetic change. The transforming DNA may or may not be integrated (covalently linked) into the cell's genome. For example, in prokaryotes, yeast, and mammalian cells, the transforming DNA may be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming DNA has integrated into a chromosome and is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones, which comprise a population of daughter cells containing the transforming DNA. A "clone" is a population of cells derived from a single cell or common ancestor by mitosis. A "cell line" is a clone of a primary cell that can be stably grown in vitro for many generations.
[0046] The term "contacting" as used herein refers to bringing into contact, or causing contact, being in contact, or coming into contact. The term "contacting" as used herein refers to the state or situation of being in contact or in direct or local proximity. Contacting of the system with a target destination, such as, but not limited to, an organ, tissue, cell, or tumor, can be by any means of administration known to those skilled in the art.
[0047] As used herein, the terms "providing," "administering," and "introducing" are used interchangeably herein and refer to the placement of a system, recombinase, or nucleic acid of the present disclosure into a cell, organism, or subject by a method or route that results in at least partial localization of the system to a desired site. The system, recombinase, or nucleic acid can be administered by any suitable route that results in delivery to the desired location within the cell, organism, or subject.
[0048] A "subject" or "patient" can be human or non-human and can include, for example, animal strains or species used as "model systems" for research purposes, such as the mouse model described herein. Similarly, a patient can include an adult or adolescent (e.g., a child). Furthermore, a patient can refer to any living organism, preferably a mammal (e.g., human or non-human), that can benefit from the administration of the compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates, such as chimpanzees, and other ape and monkey species; livestock, such as cattle, horses, sheep, goats, and pigs; domestic animals, such as rabbits, dogs, and cats; and laboratory animals, including rodents, such as rats, mice, and guinea pigs. Examples of non-mammals include, but are not limited to, birds, fish, and the like. In one embodiment of the methods and compositions provided herein, the mammal is a human.
[0049] Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and not intended to be limiting.
[0050] 2. Recombinase system The present disclosure provides a system for DNA modification, comprising a polypeptide comprising, or a nucleic acid encoding, a recombinase (e.g., a large serine recombinase) having an amino acid sequence with at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) to any of SEQ ID NOs: 1-74; and a first polynucleotide comprising a donor recognition sequence for the recombinase. Also provided herein are enzymatically active fragments thereof (e.g., containing C- or N-terminal truncations or internal deletions but retaining the desired enzymatic activity). Active fragments may contain at least 20, 30, 40, 50, 100, or more amino acids of SEQ ID NOs: 1-74, or a sequence with at least 70% identity to at least 20, 30, 40, 50, 100, or more amino acids of SEQ ID NOs: 1-74. In some embodiments, the recombinase has an amino acid sequence that has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to any of SEQ ID NOs: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66, or an active fragment thereof. In selected embodiments, the recombinase has the amino acid sequence of SEQ ID NO: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66, or an active fragment thereof.
[0051] The present disclosure also provides a system for DNA modification, comprising a polypeptide comprising, or a nucleic acid encoding, a recombinase (e.g., a large serine recombinase), and a first polynucleotide comprising a donor recognition sequence for the recombinase, wherein the recombinase (e.g., a large serine recombinase) comprises one or more of the following amino acid motifs written in a consensus Prosite format, where the amino acids that may be at any one position are in square brackets, x is any amino acid, and x(n) represents n any amino acids (e.g., x(3) is xxx or three consecutive amino acids): Motif 1: [AEILSTVY]-[ADEGKQRST]-x(3)-[EG]-x-[ACFLMV]-x-[AFILMTV]-x(2)-[FHILMNV]-[AGSV]-[ADILSTV]-x-[AGS]-x (3)-[KRSV]-[ADEGKNST]-[AEIKMNQST]-[FILMST]-x-[DELQSV]-[ENQR]-x(4)-[AFHIKLMNQRSV]-x-[AEGHKLMNQRSV] Motif 2: [AGI]-[DEGNPSTV]-[DGNQS]-[AHNQRTVY]-x-[ADEHILPQRTY]-[ADEQR]-[FIKL]-x-[DEFGNQRSTV]-[AILSTV]-[DEIKLNQRSTV]-[ADEKMNRSTV]-[AGQRST]-x-[ADEKLQRT]-x-[ALMV] Motif 3: [ADFILMNSY]-x(2)-[AIKMSV]-x-[AFGILMV]-x(3)-[QRT]-[AGS]-x-[DEGNQS]-ESx-[AHKNRSTV]-Kx(2)-[LMRY]-[AINQSTV]-[AEFIKLNRTV]-x-[AFHLNQSTY]-[AILMNRSTVY] Motif 4: [EKNTGSLDVARP]-[EHITGSLDVAP]-x-[MITSLVARP]-[EKNITGSDQVARP]-[EGSDARP]-[ILDAR]-[MHKTLVQDAR]-[EKITGSLDQVA]-[EKHDQVAR]-[MHISLVQAR]-[QEKNMSLDVAR]-[EKHGSLDQAR]-[EYKNIHLVA]-x-[EKITGSLDQAR]-[EKHTGDQAR]-x-[QEKNTGSDVAR]-[QEKNTGSVDAR]-[ISWLVFAR]-[QEMTGSLVDA]-[EKNITGSDARP]-[EMILDQA]-[EYILVFAR]-[EMTGSLDVAR]-[EKNGSLDQAR]-[QEGVDARP] モチーフ5: [ADEHKNQRS]-[ADEFGHKMNQRSWY]-[EFY]-[FHLWY]-x-[ADEFIKLMNQRSTY]-[FIQSTV]-[AGKLNRSTV]-[ADEHKNQRTY]-[INQR]-[FILMQS]-x(2)-[AGKNS]-[KMQRSTV]-x(2)-[AEGKMNSTY] モチーフ6: W-[AEHNRSTV]-x-[AGNST]-[FGLMNQSTV]-[ILPV]-x(2)-[ILTV]-x(4)-[ACGMQRST]-x-[ILVY]-G-[DEHNQS]-x-[EHILMQRT]-[AEFHLNPY]-[CFHKMNQRTY]-[DEFIKLNQRSTV] モチーフ7: [AGINSTV]-x-[AIS]-x-[FILMY]-E-[IR]-x(2)-[DILT]-x-[AEIKMQS]-R-[ITV]-x-[ADGRST]-x-[FKLMY]-[AEHIKLMNQRVWY]-x-[AIKLMR] モチーフ8: [FY]-[DEKQS]-[EKLMQ]-[KLR]-[KLV]-x-[GN]-[DEHKLMR]-[ST]-x-[FHIQSTVW] モチーフ9: [ILV]-x(2)-[ADFHILMNQSVY]-x(3)-[AGS]-x-[DEIKNQRS]-[EQ]-Sx(2)-[AK]-[AQRS]-x -[LMR]-[ILQRSV]-x-[ADEGHIQRS]-[AKNQSTV]-[AHKRWY]-x-[AGHIKQRST]-x-[CHIKLRV] Motif 10: R-[LMQR]-[ANS]-[NPST]-W Motif 11: [ILV]-[AV]-x-[AFHILQWY]-[IMV]-x-[ELQT]-[AIV]-F Motif 12: R-[DKNRSV]-[ADEFGKPQS]-[AEIKLSTV]-x-[FGILNV]-[AFILQRVY]-[DEILMNQSTV]-[DEFILMQTVY]-[IKLRV]-[DEKNQR]-[DEFKLNQWY]-[FL] Motif 13: [AEFILMNQSTVY]-[AFGILMRSTV]-x(3)-[ADEFGHLMNST]-x(2)-[DMNS]-[DEQ]-x-[CFHLTVY]-x-[AEKLRY]-x(2)-[ALS]-x-[DEKNQRS]-[GIMQRTV]-[DHKNQR]-x-[AGILNSTV]-[FHIKLMNQVWY]
[0052] Alternatively, the motif can be written as follows, where each position is defined by a designated amino acid or X, where X is the amino acid choice in parentheses, or any amino acid shown: Motif 1:X 1a X 2a X 3a X 4a X 5a X 6a X 7a X 8a X 9a X 10a X 11a X 12a X 13a X 14a X 15a X 16a X17a X 18a X 19a X 20a X 21a X 22a X 23a X 24a X 25a X 26a X 27a X 28a X 29a X 30a X 31a X 32a X 33a X 34a , During the ceremony: X 3a , X 4a , X 5a , X 7a , X 9a , X 11a , X 12a , X 16a , X 18a , X 19a , X 20a , X 25a , X 28a , X 29a , X 30a , X 31a , and X 33a are each independently selected from any amino acid; X 1a is A, E, I, L, S, T, V, or Y; X 2a is A, D, E, G, K, Q, R, S, or T; X 6a is E or G; X 8a is A, C, F, L, M, or V; X 10a is A, F, I, L, M, T, or V; X 13a is F, H, I, L, M, N, or V; X 14a is A, G, S, or V; X 15a is A, D, I, L, S, T, or V; X 17a is A, G, or S; X21a is K, R, S, or V; X 22a is A, D, E, G, K, N, S, or T; X 23a is A, E, I, K, M, N, Q, S, or T; X 24a is F, I, L, M, S, or T; X 26a is D, E, L, Q, S, or V; X 27a is E, N, Q, or R; X 32a is A, F, H, I, K, L, M, N, Q, R, S, or V; and X 34a is A, E, G, H, K, L, M, N, Q, R, S, or V Motif 2: X 1b X 2b X 3b X 4b X 5b X 6b X 7b X 8b X 9b X 10b X 11b X 12b X 13b X 14b X 15b X 16b X 17b X 18b , During the ceremony: X 5b , X 9b , X 15b , and X 17b are each independently selected from any amino acid; X 1b is A, G, or I; X 2b is D, E, G, N, P, S, T, or V; X 3b is D, G, N, Q, or S; X 4b is A, H, N, Q, R, T, V, or Y; X 6bis A, D, E, H, I, L, P, Q, R, T, or Y; X 7b is A, D, E, Q, or R; X 8b is F, I, K, or L; X 10b is D, E, F, G, N, Q, R, S, T, or V; X 11b is A, I, L, S, T, or V; X 12b is D, E, I, K, L, N, Q, R, S, T, or V; X 13b is A, D, E, K, M, N, R, S, T, or V; X 14b is A, G, Q, R, S, or T; X 16b is A, D, E, K, L, Q, R, or T; and X 18b is A, L, M, or V Motif 3: X 1c X 2c X 3c X 4c X 5c X 6c X 7c X 8c X 9c X 10c X 11c X 12c X 13c ESX 16c X 17c KX 19c X 20c X 21c X 22c X 23c X 24c X 25c X 26c , During the ceremony: X 2c , X 3c , X 5c , X 7c , X 8c , X 9c , X 12c , X 16c , X 19c , X20c , and X 24c are each independently selected from any amino acid; X 1c is A, D, F, I, L, M, N, S, or Y; X 4c is A, I, K, M, S, or V; X 6c is A, F, G, I, L, M, or V; X 10c is Q, R, or T; X 11c is A, G, or S; X 13c is D, E, G, N, Q, or S; X 17c is A, H, K, N, R, S, T, or V; X 21c is L, M, R, or Y; X 22c is A, I, N, Q, S, T, or V; X 23c is A, E, F, I, K, L, N, R, T, or V; X 25c is A, F, H, L, N, Q, S, T, or Y; X 26c is A, I, L, M, N, R, S, T, V, or Y Motif 4: X 1d X 2d X 3d X 4d X 5d X 6d X 7d X 8d X 9d X 10d X 11d X 12d X 13d X 14d X 15d X 16d X 17d X 18d X 19d X 20d X 21d X 22d X 23d X 24dX 25d X 26d X 27d X 28d , During the ceremony: X 3d , X 15d , and X 18d are each independently selected from any amino acid; X 1d is E, K, N, T, G, S, L, D, V, A, R, or P; X 2d is E, H, I, T, G, S, L, D, V, A, or P; X 4d is M, I, T, S, L, V, A, R or P; X 5d is E, K, N, I, T, G, S, D, Q, V, A, R, or P; X 6d is E, G, S, D, A, R, or P; X 7d is I, L, D, A, or R; X 8d is M, H, K, T, L, V, Q, D, A, or R; X 9d is E, K, I, T, G, S, L, D, Q, V, or A; X 10d is E, K, H, D, Q, V, A, or R; X 11d is M, H, I, S, L, V, Q, A, or R; X 12d is Q, E, K, N, M, S, L, D, V, A, or R; X 13d is E, K, H, G, S, L, D, Q, A, or R; X 14d is E, Y, K, N, I, H, L, V, or A; X 16d is E, K, I, T, G, S, L, D, Q, A, or R; X 17d is E, K, H, T, G, D, Q, A, or R; X 19dis Q, E, K, N, T, G, S, D, V, A, or R; X 20d is Q, E, K, N, T, G, S, V, D, A, or R; X 21d is I, S, W, L, V, F, A, or R; X 22d is Q, E, M, T, G, S, L, V, D, or A; X 23d is E, K, N, I, T, G, S, D, A, R, or P; X 24d is E, M, I, L, D, Q, or A; X 25d is E, Y, I, L, V, F, A, or R; X 26d is E, M, T, G, S, L, D, V, A, or R; X 27d is E, K, N, G, S, L, D, Q, A, or R; and X 28d is Q, E, G, V, D, A, R, or P Motif 5: X 1e X 2e X 3e X 4e X 5e X 6e X 7e X 8e X 9e X 10e X 11e X 12e X 13e X 14e X 15e X 16e X 17e X 18e , During the ceremony: X 5e , X 12e , X 13e , X 16e , and X 17e are each independently selected from any amino acid; X 1e is A, D, E, H, K, N, Q, R, or S; X2e is A, D, E, F, G, H, K, M, N, Q, R, S, W, or Y; X 3e is E, F, or Y; X 4e is F, H, L, W, or Y; X 6e is A, D, E, F, I, K, L, M, N, Q, R, S, T, or Y; X 7e is F, I, Q, S, T, or V; X 8e is A, G, K, L, N, R, S, T, or V; X 9e is A, D, E, H, K, N, Q, R, T, or Y; X 10e is I, N, Q, or R; X 11e is F, I, L, M, Q, or S; X 14e is A, G, K, N, or S; X 15e is K, M, Q, R, S, T, or V; X 18e is A, E, G, K, M, N, S, T, or Y Motif 6: WX 2f X 3f X 4f X 5f X 6f X 7f X 8f X 9f X 10f X 11f X 12f X 13f X 14f X 15f X 16f GX 18f X 19f X 20f X 21f X 22f X 23f , During the ceremony: X 3f , X 7f , X 8f , X10f , X 11f , X 12f , X 13f , X 15f , and X 19f are each independently selected from any amino acid; X 2f is A, E, H, N, R, S, T, or V; X 4f is A, G, N, S, or T; X 5f is F, G, L, M, N, Q, S, T, or V; X 6f is I, L, P, or V; X 9f is I, L, T, or V; X 14f is A, C, G, M, Q, R, S, or T; X 16f is I, L, V, or Y; X 18f is D, E, H, N, Q, or S; X 20f is E, H, I, L, M, Q, R, or T; X 21f is A, E, F, H, L, N, P, or Y; X 22f is C, F, H, K, M, N, Q, R, T, or Y; and X 23f is D, E, F, I, K, L, N, Q, R, S, T, or V Motif 7: X 1g X 2g X 3g X 4g X 5g EX 7g X 8g X 9g X 10g X 11g X 12g RX 14g X 15g X 16g X 17g X 18g X 19g X 20g X21g , During the ceremony: X 2g , X 4g , X 8g , X 9g , X 11g , X 15g , X 17g , and X 20g are each independently selected from any amino acid; X 1g is A, G, I, N, S, T, or V; X 3g is A, I, or S; X 5g is F, I, L, M, or Y; X 7g is I or R; X 10g is D, I, L, or T; X 12g is A, E, I, K, M, Q, or S; X 14g is I, T, or V; X 16g is A, D, G, R, S, or T; X 18g is F, K, L, M, or Y; X 19g is A, E, H, I, K, L, M, N, Q, R, V, W, or Y; and X 21g is A, I, K, L, M, or R Motif 8: X 1h X 2h X 3h X 4h X 5h X 6h X 7h X 8h X 9h X 10h X 11h , During the ceremony: X 6h and X 10h are each independently selected from any amino acid; X 1h is F or Y; X 2h is D, E, K, Q, or S; X 3h is E, K, L, M, or Q; X 4h is K, L, or R; X 5h is K, L, or V; X 7h is G or N; X 8h is D, E, H, K, L, M, or R; X 9h is S or T; and X 11h is F, H, I, Q, S, T, V, or W Motif 9: X 1i X 2i X 3i X 4i X 5i X 6i X 7i X 8i X 9i X 10i X 11i SX 13i X 14i X 15i X 16i X 17i X 18i X 19i X 20i X 21i X 22i X 23i X 24i X 25i X 26i X 27i , During the ceremony: X 2i , X 3i , X 5i , X 6i , X 7i , X 9i , X 13i , X 14i , X 17i , X 20i , X 24i , and X 26i are each independently selected from any amino acid; X1i is I, L, or V; X 4i is A, D, F, H, I, L, M, N, Q, S, V, or Y; X 8i is A, G, or S; X 10i is D, E, I, K, N, Q, R, or S; X 11i is E or Q; X 15i is A or K; X 16i is A, Q, R, or S; X 18i is L, M, or R; X 19i is I, L, Q, R, S, or V; X 21i is A, D, E, G, H, I, Q, R, or S; X 22i is A, K, N, Q, S, T, or V; X 23i is A, H, K, R, W, or Y; X 25i is A, G, H, I, K, Q, R, S, or T; and X 27i is C, H, I, K, L, R, or V Motif 10: RX 2j X 3j X 4j W, During the ceremony: X 2j is L, M, Q, or R; X 3j is A, N, or S; and X 4j is N, P, S, or T Motif 11: X 1k X 2k X 3k X 4k X 5k X 6k X 7k X 8kF, During the ceremony: X 3k and X 6k are each independently selected from any amino acid; X 1k is I, L, or V; X 2k is A or V; X 4k is A, F, H, I, L, Q, W, or Y; X 5k is I, M, or V; X 7k is E, L, Q, or T; X 8k is A, I, or V Motif 12: RX 2l X 3l X 4l X 5l X 6l X 7l X 8l X 9l X 10l X 11l X 12l X 13l , During the ceremony: X 2l is D, K, N, R, S, or V; X 3l is A, D, E, F, G, K, P, Q, or S; X 4l is A, E, I, K, L, S, T, or V; X 5l is any amino acid; X 6l is F, G, I, L, N, or V; X 7l is A, F, I, L, Q, R, V, or Y; X 8l is D, E, I, L, M, N, Q, S, T, or V; X 9l is D, E, F, I, L, M, Q, T, V, or Y; X 10l is I, K, L, R, or V; X 11l is D, E, K, N, Q, or R; X 12l is D, E, F, K, L, N, Q, W, or Y; X 13l is F or L Motif 13: X 1m X 2m X 3m X 4m X 5m X 6m X 7m X 8m X 9m X 10m X 11m X 12m X 13m X 14m X 15m X 16m X 17m X 18m X 19m X 20m X 21m X 22m X 23m X 24m , During the ceremony: X 3m , X 4m , X 5m , X 7m , X 8m , X 11m , X 13m , X 15m , X 16m , X 18m , and X 22m, are each independently selected from any amino acid, X 1m is A, E, F, I, L, M, N, Q, S, T, V, or Y; X 2m is A, F, G, I, L, M, R, S, T, or V; X 6m is A, D, E, F, G, H, L, M, N, S, or T; X 9m is D, M, N, or S; X 10m is D, E, or Q; X12m is C, F, H, L, T, V, or Y; X 14m is A, E, K, L, R, or Y; X 17m is A, L, or S; X 19m is D, E, K, N, Q, R, or S; X 20m is G, I, M, Q, R, T, or V; X 21m is D, H, K, N, Q, or R; X 23m is A, G, I, L, N, S, T, or V; X 24m is F, H, I, K, L, M, N, Q, V, W, or Y.
[0053] In some embodiments, a recombinase can comprise an amino acid sequence having at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) to any of amino acid motifs 1-13. Recombinases can also comprise enzymatically active fragments of the listed amino acid motifs (e.g., containing C- or N-terminal truncations, or internal deletions, but retaining the desired enzymatic activity).
[0054] In some embodiments, the system includes a polypeptide comprising a recombinase having an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to any of SEQ ID NOs: 88-1183 (listed in Tables 4 and 5). Also provided herein are enzymatically active fragments of SEQ ID NOs: 88-1183 from the sequences listed in Tables 4 and 5 (e.g., containing C- or N-terminal truncations, or internal deletions, but retaining the desired enzymatic activity). Active fragments may contain at least 20 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 100 amino acids, or more of SEQ ID NOs: 88-1183 (Tables 4 and 5), or a sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to at least 20 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 100 amino acids, or more of SEQ ID NOs: 88-1183 (Tables 4 and 5).
[0055] As used herein, the term "recombinase" refers to a site-specific enzyme that mediates the recombination of DNA between recombinase recognition sequences, resulting in the excision, integration, inversion, or exchange (e.g., translocation) of the DNA fragment between the recombinase recognition sequences. In some embodiments, the recombinase is a large serine recombinase.
[0056] Large serine recombinases (LSRs) are site-specific recombinases commonly found on microbial mobile genetic elements and within phage genomes, allowing invading phages to insert into the host genome and thereby enter its prophage state. A typical LSR is composed of distinct domains: an N-terminal "resolvase" domain containing the active site; a "recombinase" domain that determines the enzyme's DNA-binding specificity; and a zinc-beta ribbon domain and coiled-coil motif that are responsible for further binding specificity and the irreversibility of the forward integration reaction without an excision cofactor. Based on detailed studies of the ΦC31 LSR, the following mechanism has been proposed: two LSR monomers bind to the donor attachment site and two to the acceptor attachment site—four of these monomers assemble to form a tetramer (Figure 9). This complex then cleaves both DNA strands and rejoins them at the attachment site to form a stably integrated final product.
[0057] The first polynucleotide may be part of a bacterial plasmid, a bacteriophage, a plant virus, a retrovirus, a DNA virus, an autonomously replicating extrachromosomal DNA element, a linear plasmid, mitochondrial or other organelle DNA, chromosomal DNA, etc. In some embodiments, the first polynucleotide comprises a human nucleic acid sequence. In some embodiments, the first polynucleotide is an exogenous or synthetic polynucleotide (e.g., a vector or an engineered plasmid).
[0058] The first polynucleotide can include a donor recognition site for a recombinase. A recognition site is a specific polynucleotide sequence recognized by a recombinase enzyme described herein. The terms "attB" and "attP" refer to the attachment (or recombination) sites derived from the bacterial target and phage donor, respectively; although recombination sites for specific enzymes may have different names (e.g., "attD" and "attA"), these are used herein. Recombination sites typically include a left arm and a right arm separated by a core or spacer region.
[0059] In some embodiments, the first polynucleotide further comprises a cargo nucleic acid. The cargo nucleic acid may encode a gene product, including, but not limited to, RNA (e.g., non-coding RNA such as tRNA, rRNA, microRNA (miRNA), small interfering RNA (siRNA), coding RNA such as messenger RNA (mRNA)), or a protein, or a polypeptide. The cargo nucleic acid may encode a transcriptional or translational control element (e.g., a promoter element, a response element (e.g., an activator / repressor sequence)). In some embodiments, the cargo nucleic acid encodes a therapeutic protein. In some embodiments, the cargo nucleic acid encodes a therapeutic RNA.
[0060] The donor DNA, and therefore the cargo nucleic acid, may be of any suitable length to facilitate recombination and delivery of the complete cargo nucleic acid, for example, about 50 to 100 bp (base pairs), about 100 to 1000 bp, at least or about 10 bp, at least or about 20 bp, at least or about 25 bp, at least or about 30 bp, at least or about 35 bp, at least or about 40 bp, at least or about 45 bp, at least or about 50 bp, at least or about 55 bp, at least or about 60 bp, at least or about 65 bp, at least or about 70 bp, at least or about 75 bp, at least or about 80 bp, at least or about 85 bp, at least or about 90 bp, or at least The donor DNA and cargo nucleic acid may be at least about 95 bp, at least about 100 bp, at least about 200 bp, at least about 300 bp, at least about 400 bp, at least about 500 bp, at least about 600 bp, at least about 700 bp, at least about 800 bp, at least about 900 bp, at least about 1 kb (kilobase pair), at least about 2 kb, at least about 3 kb, at least about 4 kb, at least about 5 kb, at least about 6 kb, at least about 7 kb, at least about 8 kb, at least about 9 kb, at least about 10 kb, or less than 10 kb in length, or greater than 10 kb in length. The donor DNA and cargo nucleic acid may be at least about 10 kb, at least about 50 kb, at least about 100 kb, between 20 kb and 60 kb, or between 20 kb and 100 kb.
[0061] Essentially, by contacting a set of corresponding recombination recognition sites with a corresponding recombinase, the recombinase mediates recombination between the sites. In some embodiments, the first polynucleotide further comprises a recipient recognition sequence for the recombinase.
[0062] In some embodiments, the system further comprises a second polynucleotide comprising a recipient recognition sequence for a recombinase. The second polynucleotide may be part of a bacterial plasmid, a bacteriophage, a plant virus, a retrovirus, a DNA virus, an autonomously replicating extrachromosomal DNA element, a linear plasmid, mitochondrial or other organelle DNA, chromosomal DNA, etc. In some embodiments, the second polynucleotide comprises a human nucleic acid sequence.
[0063] The type of recognition site varies depending on the recombinase. In some embodiments, the recombinase is a landing pad LSR that can efficiently integrate into a pre-installed recognition site. Examples of landing pad LSRs are shown in Table 1 along with the corresponding recombination attachment sites. In some embodiments, the recombinase is a multi-targeting LSR that can efficiently integrate into many different loci within a target genome. Examples of multi-targeting LSRs are shown in Table 3 along with the corresponding recombination attachment sites. In some embodiments, the recombinase is a genome-targeting LSR that can integrate into one or more target sites within a given target (e.g., a target genome). Examples of genome-targeting LSRs are shown in Table 2 along with the corresponding recombination attachment sites. As described herein, attachment sites can be determined by mapping the edges of the mobile genetic element.
[0064] In some embodiments, the donor recognition sequence, the recipient recognition sequence, or both are pseudorecognition sequences or pseudosites. A "pseudorecognition sequence" or "pseudosite" does not necessarily refer to the original recognition sequence of a particular recombinase, but rather to a recognition sequence sufficient to promote recombination. A pseudorecognition sequence differs from the corresponding native recombinase recognition sequence by one or more nucleotides (e.g., by an insertion, deletion, or substitution). In some embodiments, a pseudorecognition sequence may be less than 50% identical to the native sequence. A pseudorecognition sequence may be a sequence that exists as an endogenous sequence in a genome that differs from the sequence in the genome in which the wild-type recognition sequence of the recombinase is present. Identification of a pseudorecognition sequence can be achieved, for example, by using sequence alignment and analysis, as described herein, where the query sequence is the recognition sequence of interest.
[0065] Depending on the relative locations of the recombination attachment sites, any of several events can occur as a result of recombination. For example, if the recombination attachment sites are located on different nucleic acid molecules, recombination can result in the integration of one nucleic acid molecule into the second molecule.
[0066] Recombination attachment sites can also be present on the same nucleic acid molecule. In such cases, the resulting product typically depends on the relative orientation of the attachment sites. For example, recombination between sites in a parallel or direct orientation generally results in excision of any DNA between the recombination attachment sites. In contrast, recombination between attachment sites in a reverse orientation can result in inversion of the intervening DNA.
[0067] The present disclosure also provides nucleic acids encoding the recombinases disclosed herein. The present disclosure further provides nucleic acids encoding a first polynucleotide and a second polynucleotide. The recombinase and the first polynucleotide can be encoded by the same or different nucleic acids (e.g., vectors). In some embodiments, the nucleic acid sequence encoding the recombinase is transiently or stably integrated into a cell, tissue, or organism such that the cell, tissue, or organism expresses the heterologous recombinase.
[0068] The nucleic acids of the present disclosure can comprise any of a number of promoters known in the art, which can be constitutive, regulatable or inducible, cell type-specific, tissue-specific, or species-specific. In addition to sequences sufficient to induce transcription, promoter sequences of the present invention can also include sequences for other regulatory elements involved in regulating transcription (e.g., enhancers, Kozak sequences, and introns). Many promoters / regulatory sequences useful for driving constitutive expression of genes are available in the art, including, but not limited to, CMV (cytomegalovirus promoter), EF1a (human elongation factor 1 alpha promoter), SV40 (simian vacuolar virus 40 promoter), PGK (mammalian phosphoglycerate kinase promoter), Ubc (human ubiquitin C promoter), human beta-actin promoter, rodent beta-actin promoter, CBh (chicken beta-actin promoter), CAG (hybrid promoter includes CMV enhancer, chicken beta-actin promoter, and rabbit beta-globin splice acceptor), TRE (tetracycline response element promoter), H1 (human polymerase III RNA promoter), U6 (human U6 small nuclear promoter), etc. Additional promoters that can be used to express the components of this system include, but are not limited to, the cytomegalovirus (CMV) intermediate-early promoter, viral LTRs, such as Rous sarcoma virus LTR, HIV-LTR, HTLV-1 LTR, Moloney murine leukemia virus (MMLV) LTR, myeloproliferative sarcoma virus (MPSV) LTR, spleen-limited focus-forming virus (SFFV) LTR, simian virus 40 (SV40) early promoter, herpes simplex tk virus promoter, and the elongation factor 1-alpha (EF1-α) promoter with or without the EF1-α intron. Additional promoters include constitutively active promoters. Alternatively, any regulatable promoter can be used, the expression of which can be regulated within the cell.
[0069] Furthermore, inducible expression can be achieved by placing a nucleic acid encoding such a molecule under the control of an inducible promoter / regulatory sequence. Promoters known in the art that can be induced in response to inducers such as metals, glucocorticoids, tetracycline, hormones, etc. are also contemplated for use in the present invention. Thus, it will be understood that the present disclosure encompasses the use of any promoter / regulatory sequence known in the art that is capable of driving expression of a desired protein operably linked thereto.
[0070] The present disclosure also provides vectors containing the nucleic acids or systems, and cells containing the nucleic acids or vectors. Accordingly, the present disclosure further provides cells comprising the serine recombinases or systems disclosed herein.
[0071] Vectors can be used to propagate and / or effect expression from a nucleic acid in an appropriate cell (e.g., expression vectors). Those skilled in the art will be aware of the variety of vectors available for propagating and expressing nucleic acid sequences.
[0072] To construct cells expressing the systems described herein, expression vectors for stable or transient expression of the systems can be constructed by conventional methods and introduced into cells. For example, the nucleic acid can be cloned into a suitable expression vector, such as a plasmid or viral vector, operably linked to a suitable promoter. The expression vector / plasmid / viral vector selected must be suitable for integration and replication in eukaryotic cells.
[0073] In certain embodiments, the vectors of the present disclosure can drive expression of one or more sequences in mammalian cells using mammalian expression vectors. Examples of mammalian expression vectors include pCDM8 (Seed, Nature (1987) 329:840, incorporated herein by reference) and pMT2PC (Kaufman, et al., EMBO J. (1987) 6:187, incorporated herein by reference). When used in mammalian cells, the expression vector's control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other expression systems suitable for both prokaryotic and eukaryotic cells, see, e.g., Chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL. 2nd eds., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989, which is incorporated herein by reference.
[0074] The vectors of the present disclosure can direct expression of a nucleic acid in a specific cell type (e.g., expressing a nucleic acid using tissue-specific regulatory elements). Such regulatory elements include promoters that can be tissue-specific or cell-specific. The term "tissue-specific" as applied to a promoter refers to a promoter that is capable of directing the selective expression of a nucleotide sequence of interest in a specific tissue type (e.g., a seed) and is relatively free of expression of the same nucleotide sequence of interest in different tissue types. The term "cell-type specific" as applied to a promoter refers to a promoter that is capable of directing the selective expression of a nucleotide sequence of interest in a specific cell type and is relatively free of expression of the same nucleotide sequence of interest in different cell types within the same tissue. The term "cell-type specific," when applied to a promoter, also refers to a promoter that can promote the selective expression of a nucleotide sequence of interest in a region within a single tissue. The cell-type specificity of a promoter can be assessed using methods well known in the art, such as immunohistochemical staining.
[0075] Additionally, vectors can contain, for example, some or all of the following: a selectable marker gene for selection of stable or transient transfectants in host cells; transcription termination and RNA processing signals; 5'- and 3'-untranslated regions; an internal ribosome binding site (IRES), a versatile multiple cloning site; and a reporter gene for assessing expression of the chimeric receptor. Suitable vectors and methods for generating vectors containing transgenes are well known and available in the art. Selectable markers include chloramphenicol resistance, tetracycline resistance, spectinomycin resistance, neomycin, streptomycin resistance, erythromycin resistance, rifampicin resistance, bleomycin resistance, heat-adapted kanamycin resistance, gentamicin resistance, hygromycin resistance, trimethoprim resistance, dihydrofolate reductase (DHFR), GPT; and the URA3, HIS4, LEU2, and TRP1 genes of S. cerevisiae.
[0076] Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into cells, tissues, or subjects. Such methods can be used to administer nucleic acids to cultured cells or cells within a host organism. Non-viral vector delivery systems include DNA plasmids, cosmids, RNA (e.g., transcripts of the vectors described herein), nucleic acids, and nucleic acids complexed with a delivery vehicle.
[0077] The nucleic acid can be delivered by any suitable means. In certain embodiments, the nucleic acid or its protein is delivered in vivo. In other embodiments, the nucleic acid or its protein is delivered to isolated / cultured cells in vitro or ex vivo to provide modified cells useful for in vivo delivery to patients suffering from a disease or condition.
[0078] Vectors according to the present disclosure can be transformed, transfected, or otherwise introduced into a wide variety of host cells. Transfection refers to the uptake of a vector by a cell, regardless of whether a coding sequence is actually expressed. Many methods of transfection are known to those of skill in the art, such as lipofectamine, calcium phosphate co-precipitation, electroporation, DEAE-dextran treatment, microinjection, viral infection, and other methods known in the art. Transduction refers to the entry of a virus into a cell and the expression (e.g., transcription and / or translation) of sequences delivered by the viral vector genome. In the case of recombinant vectors, "transduction" generally refers to the entry of a recombinant viral vector into a cell and the expression of a nucleic acid of interest delivered by the vector genome.
[0079] Methods for delivering vectors to cells are well known in the art and include DNA or RNA electroporation, transfection reagents such as liposomes or nanoparticles for delivering DNA or RNA; delivery of DNA, RNA, or proteins by mechanical deformation (see, e.g., Sharei et al. Proc. Natl. Acad. Sci. USA (2013) 110(6): 2082-2087, incorporated herein by reference). Nucleic acids can be delivered as part of a larger construct, such as a plasmid or viral vector, or directly by, for example, electroporation, lipid vesicles, viral transporters, microinjection, and biolistics (high velocity particle bombardment).
[0080] In addition, delivery vehicles such as nanoparticle-based and lipid-based delivery systems can also be used. Further examples of delivery vehicles include lentiviral vectors, ribonucleoprotein (RNP) complexes, lipid-based delivery systems, gene guns, hydrodynamics, electroporation or nucleofection microinjection, and biolistics. Various gene delivery methods are discussed in detail by Nayerossadat et al. (Adv Biomed Res. 2012;1: 27) and Ibraheem et al. (Int J Pharm. 2014 Jan 1;459(1-2):70-83), which are incorporated herein by reference.
[0081] Accordingly, the present disclosure provides isolated cells comprising the vector(s) or nucleic acid(s) disclosed herein. Preferred cells are those that can be grown easily and reliably, have a reasonably fast growth rate, have a well-characterized expression system, and can be easily and efficiently transformed or transfected. Examples of suitable prokaryotic cells include, but are not limited to, cells from the genera Bacillus (such as Bacillus subtilis and Bacillus brevis), Escherichia (such as E. coli), Pseudomonas, Streptomyces, Salmonella, and Envinia. Suitable eukaryotic cells are known in the art and include, for example, yeast cells, insect cells, and mammalian cells. Examples of suitable yeast cells include those from the genera Kluyveromyces, Pichia, Rhino-sporidium, Saccharomyces, and Schizosaccharomyces. Exemplary insect cells include Sf-9 and HIS (Invitrogen, Carlsbad, Calif.), and are described, for example, in Kitts et al., Biotechniques, 14: 810-817 (1993); Lucklow, Curr. Opin. Biotechnol., 4: 564-572 (1993); and Lucklow et al., J. Virol., 67: 4566-4579 (1993), which are incorporated herein by reference. Desirably, the cell is a mammalian cell, and in some embodiments, the cell is a human cell. Numerous suitable mammalian and human host cells are known in the art, and many are available from the American Type Culture Collection (ATCC, Manassas, Va.).Examples of suitable mammalian cells include, but are not limited to, Chinese hamster ovary cells (CHO) (ATCC No. CCL61), CHO DHFR cells (Urlaub et al., Proc. Natl. Acad. Sci. USA, 97: 4216-4220 (1980)), human embryonic kidney (HEK) 293 or 293T cells (ATCC No. CRL1573), and 3T3 cells (ATCC No. CCL92). Other suitable mammalian cell lines are the monkey COS-1 (ATCC No. CRL1650) and COS-7 (ATCC No. CRL1651) cell lines, and the CV-1 (ATCC No. CCL70) cell line. Further exemplary mammalian host cells include primate, rodent, and human cell lines, including transformed cell lines. Normal diploid cells, cell lines derived from in vitro culture of primary tissues, and primary explants are also suitable. Other suitable mammalian cell lines include, but are not limited to, mouse neuroblastoma N2A cells, HeLa, HEK, A549, HepG2, mouse L-929 cells, and BHK or HaK hamster cell lines.
[0082] Methods for selecting suitable mammalian cells and for transforming, culturing, amplifying, screening, and purifying the cells are known in the art.
[0083] The present invention also relates to compositions comprising the recombinases, systems, nucleic acids, vectors, or cells described herein.
[0084] Further disclosed herein are methods for identifying recombinases for use in the systems and methods disclosed herein. In some embodiments, the method includes: obtaining a bacterial genome sequence; identifying putative recombinase genes in the bacterial genome sequence based on predicted recombinase domains; comparing the genome encoding the putative recombinase gene with a genome that does not contain the putative recombinase gene; mapping the boundaries of the mobile genetic element containing the putative recombinase gene; and determining the recombinase recognition sequence and / or attachment site. In some embodiments, the predicted recombinase domain is a Pfam domain. In some embodiments, the method further includes isolating the mobile genetic element from the bacterial genome sequence before identifying the putative recombinase gene. Mapping the boundaries of the mobile genetic element can include determining the 3' and 5' flanking sequences of the mobile genetic element ends and, if present, the duplication site created upon insertion of the mobile genetic element.
[0085] 3. Methods for modifying DNA Genetic engineering applications through DNA modification have led to impactful results, such as CAR-T cell therapy, genetically modified crops, and cells that produce diverse compounds and pharmaceuticals. For many of these applications, genomic integration is strongly preferred over plasmid-based methods for maintaining heterologous genes within engineered cells due to increased genome stability, improved copy number control, and regulatory concerns regarding the biological containment of recombinant DNA. However, generating modified cells with genome-wide modifications spanning several kilobases remains a practical challenge and often requires inefficient, multi-step processes that are time- and resource-intensive. The systems and methods described herein enable the integration of large (e.g., kilobases or larger) exogenous donor polynucleotides into DNA sequences. These methods can be used in vitro, ex vivo, or in vivo, allowing for the modification of target DNA strands in solution, in cells, tissues, or within a subject.
[0086] The present disclosure provides a method for modifying a target nucleic acid sequence. As used herein, the phrase "modifying a DNA sequence" or "modifying a target DNA" refers to modifying at least one physical characteristic of a target DNA sequence. DNA modification includes, for example, single-stranded or double-stranded DNA breaks, deletions, or insertions of one or more nucleotides, and other modifications that affect the structural integrity or nucleotide sequence of a DNA sequence.
[0087] In some embodiments, the method includes contacting a target nucleic acid sequence with a system disclosed herein or a polypeptide comprising a recombinase having an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to any of SEQ ID NOs: 1-74, an enzymatically active fragment thereof, or a nucleic acid encoding same.
[0088] In some embodiments, the recombinase has an amino acid sequence that has at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to any of SEQ ID NOs: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66. In selected embodiments, the recombinase has the amino acid sequence of SEQ ID NO: 2, 6, 10, 12, 18, 19, 26, 29, 61, 65, or 66.
[0089] In some embodiments, the method includes contacting a target nucleic acid sequence with a system disclosed herein or a polypeptide comprising a recombinase having an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to any of motifs 1-13 disclosed above, or an enzymatically active fragment thereof, or a nucleic acid encoding same.
[0090] In some embodiments, the system includes a polypeptide comprising a recombinase having an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to any of SEQ ID NOs: 88-1183 listed in Tables 4 and 5. Also provided herein are enzymatically active fragments (e.g., containing C-terminal or N-terminal truncations, or internal deletions, but retaining desired enzymatic activity) of SEQ ID NOs: 88-1183 listed in Tables 4 and 5. Active fragments may contain at least 20 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 100 amino acids, or more of SEQ ID NOs: 88-1183 (Tables 4 and 5), or a sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%) identity to at least 20 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 100 amino acids, or more of SEQ ID NOs: 88-1183 (Tables 4 and 5).
[0091] In some embodiments, the target DNA comprises a donor recognition sequence, a recipient recognition sequence, or both.
[0092] In some embodiments, the method further comprises contacting the target DNA with a first polynucleotide comprising a donor recognition sequence for a recombinase. In some embodiments, the first polynucleotide further comprises a cargo DNA sequence. In some embodiments, the donor recognition sequence, the recipient recognition sequence, or both are pseudorecognition sequences.
[0093] The descriptions and embodiments provided above for the disclosed system, recombinases, first and second polynucleotides, donor and recipient recognition sequences, and cargo DNA sequences are applicable to the methods described herein.
[0094] In some embodiments, the method can include introducing a disclosed system or recombinase, or a nucleic acid encoding it, and a donor polynucleotide into a cell. In some embodiments, the recombinase or a nucleic acid encoding it is introduced into the cell before introduction of the donor polynucleotide. In some embodiments, the recombinase or a nucleic acid encoding it is introduced into the cell after introduction of the donor polynucleotide. In some embodiments, the recombinase or a nucleic acid encoding it and the donor polynucleotide can be introduced in any order, with a time interval between each introduction.
[0095] In some embodiments, the recombinase is part of a system that includes a Cas protein, reverse transcriptase, or an active fragment or combination thereof. In some embodiments, the recombinase is in a fusion protein with a Cas protein (e.g., Cas9) and reverse transcriptase, or an active fragment thereof. For example, the programmable addition by site-specific targeting element (PASTE) system integrates large cargo in a single delivery. See Eleonora I. Ioannidi, et al., bioRxiv 2021.11.01.466786, incorporated herein by reference in its entirety.
[0096] In some embodiments, the recombinase or nucleic acid encoding it is introduced into the cell simultaneously with the introduction of the donor polynucleotide, e.g., the recombinase or nucleic acid encoding it and the donor polynucleotide are introduced simultaneously or near simultaneously.
[0097] The cell can be any eukaryotic cell or eukaryote (e.g., a cell of a unicellular eukaryote, a plant cell, an algae cell, a fungal cell (e.g., a yeast cell), an animal cell, a cell from an invertebrate (e.g., a Drosophila, a Cnidarian, an Echinoderm, a nematode, an insect, an arachnid, etc.), a cell from a vertebrate (e.g., a fish, an amphibian, a reptile, a bird, a mammal), a cell from a mammal, a cell from a rodent, a cell from a human, etc.), or a protozoan cell, a mitotic and / or post-mitotic cell from the cell. Any type of cell can be of interest (e.g., stem cells, e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells; somatic cells, e.g., fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells, liver cells, lung cells, skin cells, in vitro or in vivo embryonic cells of any stage of embryo, e.g., zebrafish embryos at the 1-cell, 2-cell, 4-cell, 8-cell, etc. stage, etc.). Cells may be derived from an established cell line or may be primary cells, where "primary cells," "primary cell line," and "primary culture" are used interchangeably herein and refer to cells and cell cultures derived from a subject and grown in vitro for a limited number of passages.
[0098] In some embodiments, the one or more cells are animal cells. The present disclosure provides modified animal cells produced by the present systems and methods, animals comprising the animal cells, cell populations comprising the cells, tissues of the animals, and at least one organ. The present disclosure further encompasses progeny, clones, cell lines, or cells of the genetically modified animals. The cells of the present invention can be used for transplantation (e.g., hematopoietic stem cells or bone marrow).
[0099] Non-limiting examples of animal cells that can be genetically modified using the systems and methods include, but are not limited to, cells derived from mammals, such as primates (e.g., apes, chimpanzees, macaques), rodents (e.g., mice, rabbits, rats), dogs, livestock (cattle / cows, donkeys, sheep / sheep, goats, or pigs), poultry (e.g., chickens), and fish (e.g., zebrafish). The methods and systems of the present invention can be used with cells from other eukaryotic model organisms, such as Drosophila, C. elegans, etc. In certain embodiments, the mammal is a human, a non-human primate (e.g., marmoset, rhesus monkey, chimpanzee), a rodent (e.g., mouse, rat, gerbil, guinea pig, hamster, cotton rat, naked mole rat), rabbit, a livestock animal (e.g., goat, sheep, pig, cow, cattle, buffalo, horse, camelid), a pet mammal (e.g., dog, cat), a zoo mammal, a marsupial, an endangered mammal, and outbred or random-breeding populations thereof.
[0100] In some embodiments, the one or more cells comprise plant cells. Suitable plant cells can be derived from many different plants, including, but not limited to, monocotyledonous and dicotyledonous plants from crops, including grain crops (e.g., wheat, corn, rice, millet, barley), fruit crops (e.g., tomatoes, apples, pears, strawberries, oranges), forage crops (e.g., alfalfa), root crops (e.g., carrots, potatoes, sugar beets, yams), leafy vegetable crops (e.g., lettuce, spinach); flowering plants (e.g., petunias, roses, chrysanthemums), coniferous and pine trees (e.g., pine, fir, spruce); plants used in phytoremediation (e.g., heavy metal accumulating plants); oil crops (e.g., sunflower, rapeseed), and plants used for experimental purposes (e.g., Arabidopsis thaliana). Thus, the disclosed methods and compositions have use in a variety of plants, including, but not limited to, species from the genera Asparagus, Avena, Brassica, Citrus, Citrullus, Capsicum, Cucurbita, Daucus, Glycine, Hordeum, Lactuca, Lycopersicon, Malus, Manihot, Nicotiana, Oryza, Persea, Pisum, Pyrus, Prunus, Raphanus, Secale, Solanum, Sorghum, Triticum, Vitis, Vigna, and Zea.
[0101] In some embodiments, the one or more cells comprise microbial cells. In some embodiments, the microbial cells are gram-negative bacterial cells, gram-positive bacterial cells, or a combination thereof. In some embodiments, the microbial cells are pathogenic bacterial cells. In some embodiments, the microbial cells are non-pathogenic bacterial cells (e.g., probiotic and / or commensal bacterial cells). In some embodiments, the microbial cells form a microbiota (e.g., a natural human microbiota). In some embodiments, the microbial cells are used in industrial or environmental bioprocesses (e.g., bioremediation).
[0102] The cell may be a cancer cell. Suitable cancer cells may be derived from breast cancer, lung cancer, colon cancer, pancreatic cancer, kidney cancer, stomach cancer, liver cancer, bone cancer, blood cancer (e.g., leukemia or lymphoma), neural tissue cancer, melanoma, ovarian cancer, testicular cancer, prostate cancer, cervical cancer, vaginal cancer, or bladder cancer.
[0103] These systems and methods can be used to modify stem cells. The term "stem cell" is used herein to refer to cells that have both the capacity for self-renewal and the ability to generate differentiated cell types (see Morrison et al. (1997) Cell 88:287-298, incorporated herein by reference). Stem cells may be characterized by both the presence and absence of certain markers (proteins, RNA, etc.). Stem cells can also be identified by both in vitro and in vivo functional assays, particularly assays relating to the stem cell's ability to give rise to multiple differentiated progeny. Examples of stem cells include pluripotent stem cells, multipotent stem cells, and unipotent stem cells. Examples of pluripotent stem cells include embryonic stem cells, embryonic germ cells, embryonic carcinoma cells, and induced pluripotent stem cells (iPSCs). The cells may be induced pluripotent stem cells (iPSCs), e.g., derived from a subject's fibroblasts. In another embodiment, the cells may be fibroblasts. In some embodiments, the cells may be cancer stem cells.
[0104] The present disclosure further provides progeny of the genetically modified cells, which can contain the same genetic modification as the genetically modified cell from which they were derived. The present invention also provides compositions comprising the genetically modified cells. In some embodiments, the genetically modified host cells are capable of generating genetically modified organisms. For example, the genetically modified host cells are pluripotent stem cells and can generate genetically modified organisms. Methods for producing genetically modified organisms are known in the art.
[0105] In some embodiments, the cell is in an organism or host, and introducing a disclosed recombinase, system, composition, nucleic acid, or vector into the cell comprises administration to the subject. The method may include providing or administering a recombinase, nucleic acid, vector, composition, or system described herein to a subject in vivo or by transplantation of ex vivo treated cells.
[0106] Cell replacement therapy can be used to prevent, correct, or treat a disease or condition, and the methods of the disclosure are applied to isolated cells of a subject (ex vivo), after which the genetically modified cells are administered to a patient.
[0107] The cells can be autologous or allogeneic to the subject to which the cells are administered. As described herein, the genetically modified cells can be autologous to the subject, for example, the cells are obtained from a subject in need of treatment, genetically engineered, and then administered to the same subject. Alternatively, the host cells can be allogeneic, for example, the cells are obtained from a first subject, genetically engineered, and administered to a second subject that is different from the first subject but of the same species. In some embodiments, the genetically modified cells are allogeneic cells and are further genetically engineered to reduce graft-versus-host disease.
[0108] A "subject" can be human or non-human and can include, for example, animal strains or species used as "model systems" for research purposes, such as the mouse model described herein. Similarly, a subject can include an adult or adolescent (e.g., a child). Furthermore, a subject can refer to any living organism, preferably a mammal (e.g., human or non-human), that can benefit from the administration of the compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates, such as chimpanzees, and other ape and monkey species; livestock, such as cattle, horses, sheep, goats, and pigs; domestic animals, such as rabbits, dogs, and cats; and laboratory animals, including rodents, such as rats, mice, and guinea pigs. Examples of non-mammals include, but are not limited to, birds, fish, and the like. In one embodiment of the methods and compositions provided herein, the mammal is a human.
[0109] These methods are used to inactivate genes of interest or delete nucleic acid sequences. In some embodiments, the disclosed methods modify target genomic DNA sequences in host cells, tissues, or subjects to modulate expression of the target DNA sequence, e.g., to increase, decrease, or completely eliminate expression of the target DNA sequence (e.g., by deleting the gene, inserting or inverting a promoter element). In some embodiments, the systems and methods described herein can be used to introduce exogenous donor polynucleotides into the target DNA sequence.
[0110] In some embodiments, the target DNA encodes a gene product. As used herein, the term "gene product" refers to any biochemical product resulting from the expression of a gene. A gene product may be RNA or protein. RNA gene products include non-coding RNA, such as tRNA, rRNA, microRNA (miRNA), and small interfering RNA (siRNA), and coding RNA, such as messenger RNA (mRNA). In some embodiments, the target genomic DNA sequence encodes a protein or polypeptide. However, the present invention is not limited to editing gene products. Any target DNA sequence can be edited as desired. For example, in some embodiments, the target DNA includes non-coding DNA or a region responsible for the production of RNA. In some embodiments, the gene of interest is located on a chromosome. In some embodiments, the gene of interest is located on an episome, for example, within a bacterial cell.
[0111] The method for inactivating a gene of interest includes introducing a recombinase, system, nucleic acid, or vector described herein into one or more cells, wherein the target nucleic acid sequence comprises at least a portion of the gene of interest. The gene of interest can include any gene of interest to be inactivated. In some embodiments, the gene of interest includes an antibiotic resistance gene, a virulence gene, a metabolic gene, a toxin gene, a remodeling gene, a disease-causing gene or gene variant, or a mutated gene.
[0112] In selected embodiments, the systems and methods described herein can be used to correct one or more defects or mutations in a gene (referred to as "gene correction"). In such cases, the cell or target sequence encodes a defective version of the gene, and the disclosed system further includes a cargo nucleic acid molecule encoding a wild-type or corrected version of the gene. In other words, the cell expresses a "disease-associated" gene. The term "disease-associated gene" refers to any gene or polynucleotide whose gene product is expressed at an abnormal level or in an abnormal form in cells obtained from an individual affected with a disease compared to tissues or cells obtained from an individual not affected with the disease. A disease-associated gene may be expressed at an abnormally high level or an abnormally low level, and changes in expression correlate with the development and / or progression of the disease. A disease-associated gene also refers to a gene whose mutation or genetic variant is directly involved in the pathogenesis of the disease or is in linkage disequilibrium with a gene or genes involved in the pathogenesis of the disease. Examples of such "single gene" or "monogenic" disease-causing genes include, but are not limited to, adenosine deaminase, alpha-1 antitrypsin, cystic fibrosis transmembrane conductance regulator (CFTR), beta-hemoglobin (HB), oculocutaneous albinism II (OCA2), huntingtin (HTT), myotonic dystrophy protein kinase (DMPK), low-density lipoprotein receptor (LDLR), apolipoprotein B (APOB), neurofibromin 1 (NF1), polycystic kidney disease 1 (PKD1), polycystic kidney disease 2 (PKD2), coagulation factor VIII (F8), dystrophin (DMD), phosphate-regulated endopeptidase homolog, X-linked (PHEX), methyl-CpG-binding protein 2 (MECP2), and ubiquitin-specific peptidase 9Y, Y-linked (USP9Y).Other monogenic or monogenic diseases are known in the art and are described, for example, in Chial, H. Rare Genetic Disorders: Learning About Genetic Disease Through Gene Mapping, SNPs, and Microarray Data, Nature Education 1(1):192 (2008); Online Mendelian Inheritance in Man (OMIM); and the Human Gene Mutation Database (HGMD). In another embodiment, the target genomic DNA sequence can contain genes whose mutations, in combination with mutations in other genes, contribute to a particular disease. Diseases caused by the contribution of multiple genes that lack a simple (i.e., Mendelian) inheritance pattern are referred to in the art as "multifactorial" or "polygenic" diseases. Examples of multifactorial or polygenic diseases include, but are not limited to, asthma, diabetes, epilepsy, hypertension, bipolar disorder, and schizophrenia. Certain developmental abnormalities may also be inherited in a multifactorial or polygenic pattern, including, for example, cleft lip / palate, congenital heart disease, and neural tube defects.
[0113] 4. Kit Also within the scope of this disclosure are kits comprising a recombinase, or a nucleic acid encoding same, a donor or first polynucleotide, a composition or system described herein, or cells comprising a system described herein or a recombinase described herein.
[0114] The kit can also include instructions for using the components of the kit. The instructions are materials or methodologies related to the kit. The materials can include any combination of background information, a list of components, brief or detailed protocols for using the compositions, troubleshooting, reference materials, technical support, and other related documentation. The instructions can be provided with the kit or as a separate component, in paper or electronic form, on a computer-readable memory device, downloaded from an internet website, or provided as a recorded presentation.
[0115] It is understood that the disclosed kits can be used in conjunction with the disclosed methods. The kits may include instructions for use in any of the methods described herein. The instructions may include instructions for using the components for a method of identifying a recombinase or a method of modifying DNA.
[0116] The kits provided herein are in suitable packaging, including but not limited to vials, bottles, jars, flexible packaging, and the like.
[0117] Kits can optionally provide additional components, such as buffers and interpretive information. Typically, kits include a container and a label or package insert(s) on or associated with the container. In some embodiments, the disclosure provides an article of manufacture comprising the contents of the kit.
[0118] The kits may further include devices for holding or administering the recombinases, nucleic acids, systems, or compositions of the invention, which may include infusion devices, intravenous solution bags, hypodermic needles, vials, and / or syringes.
[0119] The present disclosure also provides kits for carrying out the methods or producing components in vitro. The kits may include the components of the system. Optional components of the kits include one or more of the following: (1) buffer components, (2) control plasmids, and (3) transfection or transduction reagents. [Example]
[0120] Cell lines and cell culture. K562 (ATCC CCL-243) cells were cultured for 37 days in RPMI 1640 (Gibco) medium supplemented with 10% FBS (Hyclone), penicillin (10,000 IU / mL), streptomycin (10,000 μg / mL), and L-glutamine (2 mM). o The cells were cultured in a humidified incubator at 37°C and 5% CO2. HEK-293T cells, as well as HEK-293FT and HEK-293T-LentiX cells used for lentivirus production described below, were grown in DMEM (Gibco) medium supplemented with 10% FBS (Hyclone), penicillin (10,000 IU / mL), and streptomycin (10,000 μg / mL).
[0121] Selection of large serine recombinases (LSRs) for initial pilot experiments. LSRs for pilot experiments were identified by searching for recombinase Pfam domains among previously identified mobile genetic elements (MGEs) (see Durrant et al. (2020) Cell Host & Microbe 28(5): 767 and El-Gebali et al., Nucleic Acids Res. 47, D427-D432 (2019), the entire contents of which are incorporated herein by reference). The identity of the attachment site was inferred from the boundary of the MGE containing each LSR. For example, if the sequence has the following structure: B1-D-P1-E-P2-D-B2 where B1 denotes the sequence adjacent to the 5' end of the MGE insertion, D denotes the target site overlap created upon insertion (if present), P1 denotes the sequence adjacent to the 5' integration boundary contained in the MGE, E is the intervening MGE, P2 denotes the sequence adjacent to the 3' integration boundary contained in the MGE, B2 denotes the sequence adjacent to the 3' end of the MGE insertion, and the attB and attP sequences can be reconstructed as follows: attB=B1+D+B2 attP=P2+D+P1 where the "+" operator in this case indicates the concatenation of nucleotide sequences.
[0122] Candidates were then annotated to determine the following characteristics: 1) whether the element was predicted to be a phage element, 2) the number of isolates containing the integrated MGE, and 3) the frequency with which MGEs containing different LSRs integrated into the same location in the genome. Candidates were then given high priority if they were contained within a predicted phage element, appeared in multiple isolates, and their attachment site was targeted by multiple different LSRs.
[0123] Computational workflow for identifying thousands of LSRs and cognate attachment sites. The LSR identification workflow was implemented as outlined in Figure 9. 146,028 bacterial isolate genomes available in the NCBI RefSeq database were identified. Genomes were then clustered at the species level using NCBI taxon IDs and the TaxonKit tool. Genomes within each species were randomized and batched into sets of 50 and 20 genomes, with the first batch containing 50 genomes and all subsequent batches containing 20 genomes. Each batch was then processed by downloading all relevant genomes from NCBI, annotating the coding sequences of each genome using Prodigal, and then searching for all coded proteins containing predicted recombinase Pfam domains using HMMER (El-Gebali et al., 2019; HMMER, n.d.). Genomes containing predicted LSRs were then compared to genomes lacking the same LSRs using the MGEfinder command whole genome, which was developed by adapting the default MGEfinder to work with draft genomes. When an LSR-containing MGE boundary was identified, all associated sequence data were saved and stored in a database. The workflow was parallelized using Google Cloud virtual machines.
[0124] After this first round of LSR mining was completed, a modified approach was adopted to further expand the database and avoid redundant searches. First, bacterial species with a large number of isolated genomes available in the first round of LSR mining were analyzed to determine whether further mining of these genomes was necessary. Rarefaction curves, representing the number of new LSR families identified for each additional genome analyzed, were estimated for these common species. Species that appeared saturated (e.g., fewer than one new cluster per 1,000 genomes analyzed) were considered "complete," meaning that no further genomes belonging to this species would be analyzed. Next, 48,557 genomes that met these filtering criteria were downloaded from the GenBank database and prepared for further analysis. The analysis was very similar to Round 1, with some notable differences. First, a database of over 496,133 isolated genomes from RefSeq and GenBank genomes was constructed. Next, PhyloPhlAn marker genes were extracted from all of these genomes. Next, for each genome found to contain a given LSR, closely related isolates found in the database were selected according to marker gene homology and chosen for comparative genomic analysis and further LSR discovery. This marker gene search approach was made available in a public GitHub repository (github.com / bhattlab / GenomeSearch). This second round of LSR and attachment site mining increased the total number of candidates by approximately 32%.
[0125] Prediction of LSR target site specificity. LSR protein sequences were clustered at 90% and 50% identity using MMseqs2. Protein sequences overlapping with the predicted attachment sites were extracted from their genomes of origin and clustered with all other target proteins at 50% identity using MMseqs2. LSR-attachment site combinations found to satisfy intermediate quality control filters were considered. To identify site-specific LSRs, only LSRs clustered at 50% identity and target proteins clustered at 50% identity were considered. LSR-target pairs were then filtered to include only target protein clusters targeted by three or more LSR clusters. Next, only LSR clusters targeting a single target protein cluster were considered. The remaining set of LSR clusters was considered monotargeting, meaning that they likely site-specifically targeted only one protein cluster. Multitargeting or transposition LSRs with minimal site specificity were identified. Only LSRs clustered at 90% identity and target proteins clustered at 50% identity were considered. Next, the total number of target protein clusters targeted by each LSR cluster was counted, and LSR clusters targeting only one protein cluster were excluded from consideration. Next, the remaining LSRs were binned according to the number of target protein clusters, with "2" indicating two target proteins, "3" indicating three target proteins, and ">3" indicating more than three target proteins. As referred to herein, "2" and "3" are considered to be moderate multitargeting, while ">3" is considered to be complete multitargeting. Next, each 50% identity cluster was assigned to a multitargeting bin according to the highest bin achieved by any one 90% cluster found within the 50% identity cluster.
[0126] Phylogenetic analysis of site-specific integrases targeting conserved attachment sites. Examples of several site-specific integrases targeting conserved attachment sites are shown in Figure 1E. All attB attachment sites were clustered at 80% identity using MMseqs2. Candidates were then filtered to include only sites that met the QC threshold, and then ranked by the number of LSR clusters in which the attB site was found to target. An example attB cluster was selected for further analysis. All LSRs targeting this attB cluster were extracted from the database and aligned using the MAFFT-LINSI algorithm. Amino acid identity distances between all LSRs were calculated, and the distance matrix was used to create a hierarchical tree in R. LSRs with 99% or more amino acid identity were clustered into a single cluster. This hierarchical tree, along with all attB sites targeted by the LSR, is visualized in Figure 1E.
[0127] Target site motifs were identified from attachment sites within the LSR database. Multitargeting LSRs in the database were analyzed at the individual protein level, the 90% amino acid identity cluster level, and the 50% amino acid identity cluster level. For each of these levels, only candidates found to target more than 10 unique attB sequences or 10 target genes clustered at 50% amino acid identity were retained. To avoid redundancy, only one attachment site per target gene cluster was extracted, and all corresponding attB sequences were extracted. These attB sequences were then first aligned using MAFFT-LINSI. Potential core dinucleotides in each alignment were then identified by extracting all dinucleotides within the alignment and ranking them by their most frequent nucleotide conservation and proximity to the attB sequence center using a custom score that equally weighted high nucleotide conservation and normalized distance to the attB center. Candidates were then realigned only with respect to these predicted dinucleotide cores, rather than using an alignment algorithm such as MAFFT. These alignments were then visualized using ggseqlogo to identify conserved target site motifs.
[0128] LSR quality control and selection criteria. LSRs with large attachment site cores greater than 20 base pairs in length were removed. The attachment site cores are the portions of attB and attP predicted to be fully homologous. LSRs with attachment sites in which more than 5% of the nucleotides were ambiguous in the original genome assembly were excluded. Only LSRs between 400 and 650 amino acids were retained. Next, only predicted LSRs containing at least one of the three major LSR Pfam domains (resolvase, recombinase, and Zn_ribbon_recom) were retained. Next, LSRs were removed from consideration if their sequence contained more than 5% ambiguous amino acids. Only LSRs found in integrated mobile genetic elements less than 200 kilobases in length were retained. Finally, only LSRs within 500 nucleotides of the predicted attachment site were retained. Candidates that met all of these filters were considered to meet the quality control threshold.
[0129] Plasmid recombination assay to validate LSR-attD-attA predictions. Three plasmids were designed for each LSR candidate. The effector plasmid contained an EF1a promoter followed by a recombinase coding sequence (codon-optimized for human cells), a 2A self-cleaving peptide, and an eGFP coding sequence. The attA plasmid contained an EF1a promoter followed by an attA sequence followed by an mTagBFP2 coding sequence, which should constitutively express the mTagBFP2 protein in human cells. The attD plasmid contained only the attD sequence followed by the mCherry coding sequence, which should not produce fluorescent mCherry before integration. HEK-293T cells were plated in 96-well plates and transfected one day later with 200 ng of the effector plasmid, 70 ng of the attA plasmid, and 50 ng of the attD plasmid using Lipofectamine 2000 (Invitrogen). Two to three days after transfection of all three plasmids into cells, cells were analyzed using flow cytometry on an Attune NxT flow cytometer (ThermoFisher). HEK-293T cells were detached from the plate using TrypLE (Gibco) and resuspended in staining buffer (BD). These experiments were performed on triplicate transfections. Cells were gated on single cells using forward and side scatter, and then on cells expressing fluorescent eGFP. Next, mTagBFP2 fluorescence was measured to indicate the amount of unrecombined attD plasmid, and mCherry fluorescence was measured to indicate the amount of recombinant plasmid.
[0130] Experiments testing the recombinase with congruent and mismatched attD plasmids were similarly performed according to the protocol described above for K562 cells. Three days after transfection, cells were measured by flow cytometry on a BD Accuri C6 cytometer.
[0131] Generation of landing pad cell lines. Landing pad LSR candidates were cloned into lentiviral plasmids under the expression of the strong pEF1a promoter, with an attB site between the promoter and start codon, and a 2A-EGFP fluorescent marker downstream of the LSR coding sequence. Lentiviral production and spinfection of K562 cells were performed as follows: HEK-293T cells were plated in 6-well tissue culture plates. 5 × 10 5 HEK-293T cells were plated in 2 mL of DMEM, grown overnight, and then transfected with 0.75 μg of an equimolar mixture of three third-generation packaging plasmids (pMD2.G, psPAX2, and pMDLg / pRRE) and 0.75 μg of LSR vectors using 10 μl of polyethyleneimine (PEI, Polysciences #23966) and 200 μl of cold serum-free DMEM: pMD2.G (Addgene Plasmid #12259; RRID:Addgene_12259), psPAX2 (Addgene Plasmid #12260; RRID:Addgene_12260), and pMDLg / pRRE (Addgene Plasmid #12251; RRID:Addgene_12251). After 24 h, 3 mL of DMEM was added to the cells, and lentivirus was harvested after 72 h of incubation. The pooled lentivirus was filtered through a 0.45 μm PVDF filter (Millipore) to remove cell debris. 5 K562 cells were centrifuged at 1000 x g for 33 min. oLentivirus was infected by spinfection at 37°C for 2 hours. To find conditions with a low multiplicity of infection, where each transduced cell was likely to contain only a single integrated copy of the landing pad, lentivirus doses of 50, 100, and 200 μl were used for each vector. Infected cells were grown for 3 days and EGFP (BD Accuri C6) was measured using flow cytometry to determine infection efficiency; doses yielding 5–15% EGFP+ cells for each LSR were selected for further experiments. After 10 days, these EGFP+ cells were sorted into 96-well plates containing a single cell per well to derive clonal lines with a single landing pad location. After 2 weeks, four clones for each LSR with highly unimodal EGFP expression levels were selected for expansion and subsequent experiments.
[0132] Landing pad integration efficiency assay. Clonal landing pad strains were electroporated with a promoterless mCherry donor containing a matching attP at a dose of either 1,000 or 2,000 ng of donor plasmid. At 3–11 days post-electroporation, cells were subjected to flow cytometry to measure mCherry (BD Accuri C6).
[0133] Pseudosite integration efficiency assay to measure percent integration into the WT genome. To determine the percentage of integration of the attD donor into pseudosites in the human genome, the attD sequence was cloned into a plasmid containing the Ef1a promoter followed by mCherry, the p2a self-cleaving peptide, and a puromycin resistance marker. 1.0 × 10 6 K562 cells were electroporated with 3000 ng of LSR plasmid and 2000 ng of mock-site attD plasmid in Amaxa solution (Lonza Nucleofector SF, program FF-120). As a mismatched LSR control, 3000 ng of Bxb1 was used instead of the corrected LSR plasmid. Cells were cultured at 2 × 10 5 cells / mL~1×10 6The cells were cultured at 100 μL / mL for 2-3 weeks. 100 μL of each sample was analyzed every 3-4 days using an Attune NxT flow cytometer to measure mCherry signal. After 2-3 weeks, the transiently transfected plasmid was almost completely diluted in the non-matched LSR control, and the efficiency of LSR was determined by the difference in mCherry percentage between the non-matched LSR control and the experimental conditions.
[0134] Integration site mapping assay to determine human genome integration specificity. Using the same protocol as above, K562 cells were electroporated with the LSR and pseudosite attD plasmids. After 5 days of culture, puromycin was added to the medium at 1 μg / mL. Cells were cultured for an additional 1.5 weeks, and gDNA was recovered using a Quick-DNA Miniprep Kit (Zymo) and quantified using the Qubit HS dsDNA Assay (Thermo). A modified version of the UDiTaS sequencing assay was used, as described in Giannoukos et al. BMC Genomics 19, 212 (2018) and Danner, 2020 Protocols.io. (doi.org / 10.17504 / protocols.io.7k2hkye). Tn5 was purified and stored at 7.5 mg / mL. Adapters were assembled by combining 50 μL of 100 μM top and bottom strands, heating to 95°C for 2 minutes, and slowly cooling to 25°C over 12 hours. Transposomes were then assembled by mixing 85.7 μL of Tn5 transposase with 14.3 μL of preannealed oligos and incubating at room temperature for 60 minutes. Tagmentation was performed by adding 150 ng gDNA, 4 μL of 5× TAPS-DMF (50 mM TAPS NaOH, 25 mM MgCl2, 50% v / v DMF (pH 8.5) at 25°C), 3 μL of assembled transposomes, and water in a final reaction volume of 20 μL. The reaction was incubated at 55°C for 10–15 minutes and purified using Zymo DNA Clean and Concentrator-5. The tagmented products were run on an Agilent Bioanalyzer HS DNA kit to confirm an average fragment size of approximately 2 kb. Next, 12 cycles of PCR were performed with the outer primer using 12.5 μL of Platinum Superfi PCR Master Mix (Thermo), 1.5 μL of 0.5 M TMAC, 0.5 μL of 10 μM outer nest GSP primer, 0.25 μL of 10 μM outer i5 primer, 9 μL of tagmented DNA, and 1.25 μL of DMSO.After cleanup with Ampure XP 0.9x beads, a second PCR was performed using the following inner primers for 18 cycles: PCR contained 25 μL Platinum Superficial MasterMix (Thermo), 3 μL 0.5M TMAC, 2.5 μL DMSO, 2.5 μL 10 μM i5 primer, 5 μL 10 μM i7 GSP primer, 10 μL purified first-round PCR product, and 2 μL water in a final reaction volume of 50 μL. The final library was size-selected for fragments between 300 and 800 bases on a 2% agarose gel, extracted from the gel using the Monarch DNA Gel Extraction Kit (NEB), quantified using the Qubit HS dsDNA Assay (Thermo) and the KAPA Library Quantification Kit, fragment analysis using the Agilent Bioanalyzer HS DNA Kit, and sequenced on a MiSeq (Illumina).
[0135] Computational analysis of integration site mapping sequence assays. We developed a Snakemake workflow and used it to analyze NGS data from the UDiTaS pseudosite sequencing assay. First, we used a custom Python script to add and remove stagger sequences (filler sequences added to better identify samples during sequencing) to primers. Next, we used fastp to trim nextera adapters from reads and remove reads with low PHRED scores. Next, we used BWA to align reads in single-end mode to both the human genome (GRCh38) and the donor plasmid sequence containing the LSR-specific attD sequence. We individually analyzed reads using a custom Python script to identify: 1) whether the read aligned to the donor plasmid, the human genome, or both; 2) whether the read started with the predicted primer; and 3) whether the pre-integration attachment site was intact. Next, the reads were filtered to include only reads that mapped to both the donor plasmid and the human genome, reads beginning at the primer site, and reads without an intact attD sequence (if this could be determined from the length of the specific read). This filtered set of reads was aligned to the human genome in paired-end mode using the default settings of BWA MEM. Alignments with a mapping quality score below 30 were removed, along with supplemental and paired-read alignments with insert sizes greater than 1500 bp. The samtools markdup tool was used to remove potential PCR duplicates and identify unique reads for downstream analysis. Next, clipped end sequences were extracted from the reads aligned to the human genome using MGEfinder, and consensus sequences of clipped ends representing crossovers from the human genome to the integrated attD sequence were generated. Using a custom Python script, k-mers 9 base pairs in length were extracted from these consensus sequences and compared to the partial sequence of the attD plasmid extending from the original primer to 25 bp after the end of the attD attachment site. If there is no shared 9-mer, the candidate is discarded.Otherwise, the consensus sequence was clipped to begin at the primer site. These consensus sequences were then aligned to the original attD subsequence using the biopython local alignment tool. Two aligned portions were extracted: the complete local alignment of the consensus sequence to attD (termed the "complete local alignment"), and the longest subset of the alignment that contained no ambiguous bases or gaps (termed the "continuous alignment"). To filter the final set of true insertion sites, only sites with at least 80% nucleotide identity shared between the consensus sequence and the attD subsequence in either the complete local alignment or the continuous alignment were retained. Finally, only sites with crossover points within 15 base pairs of the predicted dinucleotide core were retained.
[0136] While this approach can accurately predict integration sites, errors in sequencing reads lead to some variability in these predictions. To account for this, we used bedtools to combine integration sites into integration "locuses" by merging all sites within 500 base pairs of each other. This approach merges integration events that occurred in opposite directions at the same site, for example. When pooling reads across biological or technical replicates, we also merge these loci if they overlap. To measure the relative frequency of insertions across different loci, we counted all uniquely aligned reads (de-duplicated using samtools markdup) found within each locus. These were then converted to a percentage for each locus by dividing by the total number of unique reads aligned to all integration loci.
[0137] The target site motifs for different LSRs could be determined from accurate predictions of the dinucleotide core for all integration sites. For each integration locus, if multiple integration sites existed, only one was selected, prioritizing integration sites with a higher number of reads supporting them. Using bedtools, we extracted up to 30 base pairs of human genome sequence surrounding the predicted dinucleotide core, selecting either the forward or reverse strand depending on the integration direction. All such target sites, or a subset of these target sites, if desired, were then analyzed for conservation at each nucleotide position using the ggseqlogo package in R.
[0138] Phylogenetic tree construction. A phylogenetic tree was constructed using representative amino acid sequences from each quality-controlled 50% identity LSR cluster. LSRs were aligned using MAFFT in G-INS-i mode, and then a consensus tree was generated using IQ-TREE with 1000 bootstrap replicates and automatic model selection.
[0139] Example 1 Systematic identification of recombinases and predicted attachment sites reveals site-specific and multitargeting / transposition families LSRs, such as Bxb1 and PhiC31, catalyze integration reactions that recombine two DNA sequences, called attP (a DNA sequence found in phages) and attB (a DNA sequence found in bacteria), at specific attachment sites. Using a comparative genomics approach designed to identify the precise boundaries of integration elements (Figure 1A), we identified thousands of LSRs in public databases of clinical and environmental bacterial isolate genomes. Once an LSR was identified, we searched for closely related genomes (average nucleotide identity (ANI) >95%) lacking the given LSR and aligned the entire genome with and without the LSR using a previously developed bioinformatics tool, MGEfinder (Durrant et al. (2020) Cell Host & Microbe 28(5): 767, incorporated herein by reference in its entirety), which allowed us to identify integrated prophage or mobile genetic element sequences (Figure 1A). The boundaries of these predicted sequences represent the attL and attR sites formed when attP recombines with attB (Figure 1A, box) and flank the integrated prophage genome or mobile genetic element containing the LSR. Using this approach on 194,585 bacterial isolate genomes, we identified 12,638 candidate LSRs and reconstructed their original attP and attB attachment sites. After applying various quality control filters and clustering protein sequences at 50% identity, the final dataset of LSR attachment site predictions contained 1,081 LSR clusters recovered from genomes belonging to 20 host phyla (Figure 5A), demonstrating a good representation of published bacterial assemblies.
[0140] To predict the site specificity of candidate LSRs using only the constructed database, we examined the network of LSRs and associated attachment sites and recovered LSRs from a diverse set of 20 host phyla (Figure 5A), which shows a good representation of publicly available bacterial assemblies. We compared integration patterns across LSR clusters. If many distantly related LSRs appear to target similar integration sites, these LSRs may be site-specific. Conversely, if an LSR cluster targets many different integration sites, it becomes "multitargeting," meaning that the LSR cluster has relaxed sequence specificity or evolved to target sequences occurring at multiple different sites within the host organism. Target similarity was measured by mapping attB integration sites to nearby ORF predictions, allowing us to group attB sites by ORF sequence; these are referred to as "target genes." The protein sequences of these target genes were then clustered at 50% amino acid identity, further grouping more distantly related integration sites. Clustering by target gene rather than by attB sequence alone facilitated the use of protein homology rather than DNA homology to group more distantly related target sites.
[0141] For each LSR cluster, the number of associated target gene clusters was estimated and visualized at the amino acid level on a phylogenetic tree of representatives of each LSR cluster. LSRs were binned into two groups: "site-specific integrases" or "multi-targeting integrases" (Figure 1B). 82.8–88.3% of LSR clusters were predicted to be site-specific or have intermediate site specificity, where the total number of unique target genes is one, two, or three, depending on the stringency of the criteria used. A phylogenetic cluster emerged from a large number of multi-targeting LSRs, or LSRs predicted to require integration into three or more target protein families, suggesting this is an evolved strategy inherited from a single ancestor. This phylogenetic cluster was strongly correlated with DUF4368, a Pfam domain of unknown function (Figure 5A). This includes previously characterized LSRs of the Tnd-like transposase subfamily (H. Wang and Mullany 2000 Journal of Bacteriology 182 (23): 6577-83; Adams et al. 2004 Molecular Microbiology 53 (4): 1195-1207, each of which is incorporated herein by reference in its entirety).
[0142] Many examples of distantly related LSRs targeted the same gene cluster (Figures 1D and 1E). Figure 1D shows an example of a network of diverse LSR clusters that primarily target a single gene cluster, a gene with homologs annotated as ATP-dependent proteases / Mg(2+) chelatase family proteins / ComM-like proteins, containing the predicted Pfam domains ChlI (Mg chelatase subunit ChlI), Mg_chelatase (magnesium chelatase, subunit ChlI), and Mg_chelatase_C (magnesium chelatase, subunit ChlI C-terminus). Homologues of this particular gene are among the most commonly targeted genes (Figure 5E) and are targeted by 12.4% of all predicted site-specific integrases (Figure 5B). Figure 1E also shows an example of a diverse set of LSRs found to target a single conserved site, the CDS sequence of a prolyl isomerase. When we aligned LSR candidates targeting this site, we found that the DNA-binding resolvase, recombinase, and Zn_ribbon_recom domains were significantly more conserved than the C-terminus, which is not thought to play a key role in DNA binding (Figure 5C). DNA competence genes were more globally enriched, with no enrichment within or near anti-phage defense genes (Figures 5E-5G).
[0143] Figure 1G shows an example network of multi-targeting LSRs. Some multi-targeting LSRs have numerous related attB target sites, allowing computational inference of their sequence specificity from databases. As shown in Figure 1H, a single multi-targeting integrase was found to integrate into 21 different sites. Alignment of the target sites revealed a conserved TT dinucleotide core, enriched in T and A nucleotides at the 5' and 3' ends, respectively. This suggested that in this particular example, the TT central dinucleotide was the most important feature for integration, most likely resulting in relaxed overall sequence specificity. Other examples of multi-targeting LSRs with different target site motifs are shown in Figure 5D, including some with more complex motifs than the AT-rich one shown in Figure 1H.
[0144] Example 2 Landing Pad LSR Characterization One valuable application of LSRs in biotechnology is the specific delivery of genetic cargo to introduction sites, so-called "landing pads," that are not present elsewhere in the target genome. An ideal landing pad LSR would be highly specific for attB, which is not present in the target genome, yet would be able to integrate efficiently once attB is installed.
[0145] Using MGE for previously identified LSRs (Durrant et al. (2020) Cell Host & Microbe 28(5):767, incorporated herein by reference in its entirety), we curated a set of 17 LSR candidates with evidence of site specificity as an initial proof-of-concept. To verify that these recombinases are active in mammalian cells, we developed an inter-plasmid recombination assay in HEK293FT cells by synthesizing three plasmids: one for expression of the human codon-optimized LSRs and a separate plasmid containing their putative attP and attB sequences (Figure 2A). In this plasmid recombination assay, the attP plasmid contains a promoterless mCherry gene, which acquires a promoter upon recombination with the attB plasmid, resulting in expression of a fluorescent protein that can be read by flow cytometry. In this initial set of 17 candidates, 15 candidates were identified with mCherry+ MFI values greater than the attD-only control (one-tailed t-test, P < 0.05), demonstrating functional recombination (Figures 2B, 2C, and 6L). Compared to the positive control, 13 candidates had mCherry+ MFI greater than PhiC31, and 3 candidates had mCherry+ MFI greater than Bxb1. For a subset of LSRs, we tested the orthogonality of the attachment sites using assays with different attachment site combinations, and found them to be highly specific and orthogonal to each other (Figure 2D).
[0146] We also tested integration into attB-containing landing pads previously placed in the human genome (Figure 2F). A construct containing the Ef1a promoter, attB, a matching LSR, and GFP was integrated into the genome of K562 cells via high MOI lentivirus, resulting in a polyclonal cell population with a potential landing pad at a different chromosomal location in each cell. Successful integration of the promoterless mCherry donor into the landing pad results in mCherry expression while GFP is knocked out. Using this landing pad assay, five of the new LSRs were found to integrate into the human genome with measurable efficiency, with Ec04, Ec07, Kp03, and Pa01 being significantly more efficient than BxB1 (Figures 6A and 6L). The stability of these polyclonal landing pads expressing LSR-GFP was assessed over time by flow cytometry. For some landing pads, such as Ec07 and Ec03, the majority of cells lost GFP expression, suggesting that the landing pads were transcriptionally silenced or genetically unstable (Figure 6B). LSRs can function on human chromosomal DNA, and Kp03 and Pa01 emerged as the top candidates in terms of efficiency.
[0147] Landing pad integration may be most useful when the landing pad is known to be located at a single genomic site in all cells. To develop single-location landing pad lines, we integrated the landing pad LSR-GFP construct via low MOI lentivirus, resulting in a single copy of the landing pad per cell. Clonal cell lines containing a single landing pad site were then sorted, expanded, and electroporated with the attP-mCherry donor plasmid. Using this landing pad assay, we tested four integrase candidates (Ec03, Ec04, Kp03, and Pa01). Pa01 outperformed Bxb1 in terms of the percentage of stably fluorescent cells after 11 days (Figure 2F). At a threefold higher donor DNA dose (3000 ng), Pa01 reached 52% efficiency, while Bxb1 only achieved 3% integration (Figure 2M). In one Pa01 experiment, electroporation of cells with the donor plasmid increased integration efficiency two-fold to over 70% (Figure 2G). The difference in efficiency decreased with increasing doses of donor DNA (Figures 6C-6D), suggesting variability in integration rates for different LSRs.
[0148] Previous characterization of Bxb1 attB identified a sequence as short as 38 bp required for integration, but our computational pipeline initially conservatively predicted a 100 bp attB sequence. A minimum of 33 bp attB was determined for efficient PaO1 recombination, but efficient recombination for KpO3 was observed down to 25 bp attB (Figure 6F). Due to its short length, the attachment site can be easily installed by a variety of methods during cloning and cell engineering.
[0149] Efficient landing pads may be particularly useful for multiple gene integration, which can be achieved by using several LSRs in parallel, given that they do not interact with each other's attachment sites (Figure 2D). Interestingly, other well-studied LSRs, Bxb1 and PhiC31, contain modular dinucleotide cores in their attachment sites that can be altered to enable orthogonal integration (Ghosh, Kim, and Hatfull 2003 Molecular Cell 12(5):1101-11, incorporated herein by reference in its entirety). Thus, the same LSR can be applied to direct multiple cargoes to specific landing pads with different core dinucleotides. The ability to replace the core dinucleotide was tested using a plasmid recombination assay for one of the LSRs, Kp03 (Figure 6G). Changing either nucleotide in the dinucleotide core of one attachment site dramatically reduced integration efficiency, whereas subsequently altering the other attachment site to match the first restored integration efficiency. This suggested that this LSR could be used to orthogonally incorporate different cargoes at the landing pads of up to 10 different attachment sites.
[0150] The specificity of these LSRs was tested by transfecting the attP-pEF1a-mCherry donor with or without a cotransfected LSR into wild-type K562 cells and measuring mCherry expression after 18 days, by which time the episomal donor plasmid was no longer detectable. Pa01 showed no evidence of mCherry integration above background, whereas Kp03 had elevated mCherry fluorescence, suggesting that it had off-target spurious sites (Figure 2H). To identify these sites, we modified the UDiTaS™ genome-wide single-sided PCR-based sequencing assay for use as an LSR integration site mapping assay. After optimizing this assay, the percentage of on-target reads increased from 1.6% to 73.2% (Figure 6J). This assay was first performed on landing pad cell lines, allowing for an estimation of the percentage of off-target integration relative to on-target integration (Figure 2I).
[0151] This assay detected off-target integrations for all LSRs, including Bxb1 (3.48% + / - 2.98%, 9 unique reads across 9 integration loci) and Pa01 (0.47% + / - 0.46%, 13 unique reads across 10 loci), but Kp03 showed a significantly higher off-target integration rate (15.5% + / - 2.43%), with 312 unique reads across 83 distinct loci, confirming a relatively high percentage of off-target integrations. Wild-type cells transfected with Kp03 and Pa01 were sequenced using a high-coverage integration site mapping assay, which detected 79 off-target genomic integration loci for Pa01 and 2,415 off-target integration loci for Kp03. From these integration sites, we identified target site motifs targeted by these LSRs, which showed conservation in the dinucleotide core and flanking sequences, indicating bona fide integrations rather than random plasmid integrations (Figure 2H). Collectively, these results establish PaO1 as a more efficient and relatively specific landing pad LSR compared to BxB1.
[0152] A second batch of 21 LSRs was selected from the database, prioritizing those with low BLAST similarity between their attB / P sites and the human genome, and a strict quality threshold was applied. Seventeen of the 21 (81%) were functional in plasmid recombination assays, validating our computational pipeline for identifying functional candidates. Encouragingly, 16 candidates had mCherry+ MFI values higher than PhiC31, and 11 had MFI values higher than Bxb1 (Figure 2J). Integration fluorescence assays in wild-type cells using the top candidates identified three with low percentages of off-target integration (Figure 6K), with Si74 being the top candidate with favorable performance in terms of both plasmid recombination efficiency and off-target integration (Figure 2K).
[0153] Example 3 Genome-targeting LSRs integrate into the human genome at predicted target sites Particularly useful LSRs would be those that integrate directly into only one or a very few pseudosites in safe locations within the human genome and do so with considerable efficiency. Historically, LSRs with pseudosites, such as PhiC31, had to be discovered experimentally by transfecting the LSR into human cells and searching for integration sites. While this approach is effective for proof-of-concept, highly efficient and specific human genome-targeting LSRs have not been obtained. Using BLAST, we searched all attB / P sequences against the GRCh38 human genome assembly (Figure 3A), and identified 856 LSRs within the human genome with highly significant matches to at least one site (BLAST E-value <1e-3, Figure 3B). Although many of these LSR attachment site predictions did not meet our quality control threshold, we prioritized BLAST match quality when selecting candidates and synthesized 103 LSRs of varying quality. The attP and attB regions were renamed according to the BLAST hits. The attachment site matching the human genome was renamed attA (acceptor) and the other attD (donor). The predicted target site in the human genome was renamed attH (human) (Figures 3A and 3D).
[0154] All 103 candidates were tested in a plasmid recombination assay, and 27 candidates recombined at the predicted attachment sites (one-tailed t-test, P < 0.05; Figure 3C), with 4 of 64 (6.25%) low-quality candidates recombining as predicted and 21 of 37 (56.75%) high-quality candidates recombining as predicted (Figure 7A). Subsequent batches of genomic target candidates utilized only high-quality LSR attachment site predictions, which included 201 unique LSRs with attachment sites significantly matching sites in the human genome (BLAST E-value < 1e-3).
[0155] To determine whether these LSRs could directly target the chromosome of human K562 cells, we performed another plasmid recombination assay, substituting the native attA with a human pseudosite (or attH) instead of the native attachment site. We found that four of the candidates recombined with the predicted attH: Sp56, Pf80, Ps45, and Enc3 (Figure 3D). This was followed by human genome integration and integration site mapping assays. Multiple integration sites were detected for all of these candidates when using both circular donor plasmids and linear PCR amplicons. For Sp56 and Pf80, the integration sites with the most unique reads across experiments (presumably the most frequently targeted loci) were the target sites predicted by BLAST alignment: an exon in SPATA20 and an exon in FKBP2, respectively (Figure 3E). For Enc3, the predicted target site had the 12th most reads among all loci containing detected integrations. Ps45 detected reads at the predicted target site in one experiment, but coverage was too low to estimate relative specificity. Examples of reads from integration site mapping assays aligned to predicted sites are shown in Figures 3F, 7C, and 7D. These four examples demonstrate that candidates can be selected prior to experimental validation based on BLAST similarity to the human genome, and that 4 of 27 functional candidates tested (14.8%) were capable of recombining with the predicted site.
[0156] Of these four candidates, Pf80 had the highest predicted specificity, with 34.3% of unique reads mapping to the predicted target site, i.e., an exon of the gene FKBP2 at position 64,243,293 on chromosome 11 (Figure 3F). However, in efficiency assays, Pf80, Sp56, and Ps45 did not exhibit mCherry+ fluorescence above background, suggesting low overall efficiency (Figure 3G). Enc3 was the most efficient of these candidates, with 6% of cells being mCherry+ at 18 days posttransfection. We subsequently tested other genome-targeting candidates, Dn29 and Vp82, which had 4.5% and 2.5% mCherry+ cells, respectively, in efficiency assays, but no integration was detected at the predicted target site in integration site mapping assays (Figure 3G-3H). Dn29 had relatively high specificity, with 17.4% of unique reads mapping to its top target site and 33.0% of unique reads mapping to its top three target sites. Analysis of Dn29 and Vp82 integration sites revealed distinct sequence profiles of their targets, which may inform future efforts to engineer and optimize these candidates (Figures 3I-3J and 7E-7F). Several of these candidates outperform PhiC31 in terms of efficiency and specificity, and Dn29 possesses a favorable combination of both, making it a promising genome-targeting candidate. It combines ideal genome-targeting LSRs in a site-specific manner with robust efficiency. The genome-targeting candidates tested showed varying levels of efficiency, with Enc3 and Dn29 in particular having significantly higher efficiencies (6% and 5%, respectively) than PhiC31 or Pf80 (both <1%; Figure 3G). For Dn29, 61.9% of integrations occurred exclusively at the top five target sites, which were found in introns or intergenic regions (Figures 3K-3L).
[0157] Example 4 Direct integration of DNA into the human genome by multi-targeting LSR An LSR is considered a good multitargeting candidate if it has relaxed specificity requirements, if it appears in a multitargeting phylogenetic cluster (Figure 1B), and / or if it has DUF4368, a Pfam domain found to correlate with a multitargeting phylogenetic cluster (Figure 5A).
[0158] One such multitargeting LSR, called Cp36, found in Clostridium perfringens was characterized. This LSR is 544 amino acids long and contains a predicted DUF4368 domain at its C-terminus. This LSR can integrate mCherry donor cargo into the genome of K562 cells with up to 40% efficiency without prior installation of a landing pad or antibiotic selection (Figure 4A). This high level of integration efficiency was verified in HEK293FT cells using both plasmid DNA and linear PCR amplicons as donor cargo (Figure 8A). Using an integration site mapping assay, over 2,000 unique integration sites were found, with a strong bias toward specific sites (Figures 4B and 8C). The locus with the most integration events, i.e., chr1:101,429,889 (for GRCh38), was the target of approximately 2% of all integration events. There was high concordance between the two cell types, with a Jaccard similarity of 20% between the top 100 sites and 17.8% between the top 200 sites in both cell types. The number of unique reads at the top 61 sites found in both cell types was highly correlated (Pearson's r = 0.45, P = 0.0002; Figures 8D and and11A),11A), suggesting that the relative efficiency of integration at these sites is highly consistent between cell types.
[0159] Using these accurate predictions of the human integration site, we reconstructed the sequence motif targeted by Cp36 (Figure 4C). This sequence motif consists of an A-rich 5' region followed by an AA dinucleotide core, followed by a 3' T-rich region. Comparison of the native attB in the C. perfringens genome with three commonly targeted human genome target sites revealed that the three human genome integration sites closely matched the motif. One target site with low integration efficiency in both cell types also closely matched the motif, despite having short stretches of A and T nucleotides at the 5' and 3' ends. The poly(A) and poly(T) flanking regions were consistent with previous descriptions of the native attB in TndX, a previously characterized LSR that is 35.4% identical to Cp36 at the amino acid level.
[0160] To compare the efficiency of Cp36 with that of PiggyBac (PB) transposase, a commonly used tool for randomly delivering DNA cargo to TTAA tetranucleotides found within target genomes, we designed a plasmid construct containing Cp36 attD (donor attachment site), a PBITR sequence, and an mCherry reporter (Figure 8E). Cp36 functioned with similar efficiency to PB (Figure 4D). While Cp36 catalyzed unidirectional integration, similar to other site-specific LSRs (Figures 4E, 8F, and 8G), PB has been shown to be bidirectional, resulting in both cargo excision and local hopping upon PB re-administration.
[0161] To test whether Cp36 could be reused to integrate a second gene, we generated a pure population of mCherry+ cells via Cp36-mediated integration and puromycin selection and re-electroporated them with a donor containing Cp36 and BFP. After 13 days, 9% of the cells were double-positive (mCherry+ and BFP+) (Figures 4F and 11E), with no decrease in mCherry (Figure 11F), demonstrating that the second gene was delivered without loss of the first cargo. Furthermore, we found that co-delivery of Cp36 with both mCherry and BFP fluorescent reporter donors resulted in a stable population expressing both markers (Figure 4G), suggesting that Cp36 can be used to generate cells with multipart genetic circuits in a single transfection.
[0162] Furthermore, two other orthologs (Pc01 and Enc9) were also found in the database, and these also functioned as multitargeters with efficiencies of 13% and 35% in human cells (Figure 8B). These results reveal the existence of a subset of LSRs with highly efficient unilateral integration activity and longer target DNA motifs (≥20 bp) compared to lentivirus or transposase systems (2-4 bp), which have not previously been tested in eukaryotic cells. Example 5
[0163] Biological roles of LSR target genes The genes targeted and disrupted during LSR integration may represent an evolved strategy of LSR-harboring MGEs. Enriched Pfam domains were identified among the target genes (Figure 5E). Enriched domains were found in magnesium chelatase, competence proteins, type II / IV secretion system proteins, and HNH endonucleases, among others. Gene Ontology (GO) pathway analysis of the target genes identified six significantly enriched pathways (FDR < 0.1; Figure 5F). In particular, the GO term "establishment of transformation competence" (GO:0030420) was the most significantly enriched pathway, with 15 target gene clusters annotated with this term. Among these target genes were the ComK transcription factor and other ComG operon proteins, suggesting that disrupting competence and DNA transformation is a common strategy of LSR-harboring MGEs. We reasoned that @LSR may also have evolved to target host antiphage defense systems upon integration. We used DefenseFinder to annotate the relevant genomes and search for genes occurring within or near these identified systems. We identified several defense genes targeted by integrases, including the CRISPR spacer acquisition gene cas2, the CASCADE complex helicase cas3, a type I restriction-modification enzyme, the Hachiman defense gene hamA, and a UvrD-like helicase gene. However, defense genes are rarely targeted by LSR, and we did not find enrichment of target genes near defense genes, suggesting that this is not a common strategy (Figure 5G). These findings support an evolved strategy employed by LSR-carrying MGEs that primarily limits further horizontal gene transfer through the disruption of competence.
[0164] Example 6 Post-mortem identification of human genome integrations A post-hoc analysis of the genome-targeting and multi-targeting candidates in this study was performed to determine the feasibility of motif-based searches. Starting with each experimentally characterized candidate, sequence motifs were constructed by iteratively adding the native attB sequences of the most closely related LSR orthologs. Additional attB sequences were added only if they shared 95% or less identity with the already selected attB sequence. Motifs of 20, 50, and 100 such attB sequences were constructed. These motifs were then searched against experimentally observed human integration sites and approximately 30,000 randomly selected human genome sequences. These sequences were then iterated across motif score cutoffs, and the true and false positive rates were calculated at each cutoff to generate ROC curves (Figure 12A). For each LSR, the motif with the largest AUC was selected.
[0165] Sequence motifs belonging to the multitargeting candidates performed very well, with AUC values ranging from 0.94 for the Cp36 motif to 0.68 for the Bt24 motif. For genome-targeting candidates, the performance of the sequence motifs varied, with AUC values ranging from 0.65 for Dn29 to 0.44 for Enc3. All of these motifs assigned significantly higher scores to the observed integration sites than randomly selected controls, except for Sp56 and Enc3, for which there was no significant difference (Wilcoxon rank-sum test; P<0.0001 for Cp36, Enc9, Pc01, Bt24, and Dn29, P<0.01 for Pf80, P>0.05 for Sp56 and Enc3). Despite the relatively poor performance of the Pf80 and Sp56 motifs, we assigned the highest motif scores to the most frequently targeted human genome integration sites, suggesting that these database-derived sequence motifs have predictive value (Figure S12B). Visual inspection of the motifs revealed various patterns, with the Cp36 and Enc9 motifs possessing a characteristic AT-rich motif typical of many multitargeting LSRs, while others, such as Dn29 and Bt24, showed less variability and less clearly defined boundaries (Figure S12C).
[0166] These results suggest that motif-based sequence searches are valuable when prioritizing multitargeting and genomic targeting candidates. The potential targeting profiles of multitargeters, such as Cp36 and Enc9, can be better understood before experimental validation, and genomic targeting candidates can be selected based on high matched outlier motifs, such as Pf80, which may exhibit higher specificity. The difference in performance between motifs can be explained by the differential selection pressure on multitargeting and single-targeting LSRs. Multitargeting LSRs are more likely to maintain relaxed sequence specificity over longer evolutionary distances due to the greater abundance of potential target sites, resulting in more precise sequence motifs. These results may also be influenced by the efficiency of LSRs in human cells or by epigenetic modifications, such as those affecting chromatin accessibility (Figure 8H).
[0167] All references cited in this specification, including publications, patent applications, and patents, are herein incorporated by reference to the same extent as if each reference was individually and specifically indicated to be incorporated by reference.
[0168] Preferred embodiments of this invention are described herein, including the best mode known to the inventors for carrying out the invention. Variations of these preferred embodiments may become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventors anticipate that skilled artisans will employ such variations as they see fit, and the inventors intend for the invention to be practiced otherwise than as specifically described herein. Accordingly, this invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is included in the invention unless otherwise indicated herein or clearly contradicted by context.
[0169]
Table 1
Table 2
Table 3
Table 4-1
Table 4-2
Table 4-3
Table 4-4
Table 4-5
Table 4-6
Table 4-7
Table 4-8
Table 4-9
Table 4-10
Table 4-11
Table 4-12
Table 4-13
Table 4-14
Table 4-15
Table 4-16
Table 5
Claims
1. Polypeptides comprising a recombinase having an amino acid sequence having at least 70% identity with any of SEQ ID NOs: 2, 6, 29, 61, 66, 1, 3-5, 7-28, 30-60, 62-65, 67-74, and 88-1183, or nucleic acids encoding the same; and A first polynucleotide containing the donor recognition sequence of the recombinase A system for DNA modification, including...
2. The system according to claim 1, wherein the recombinase has an amino acid sequence that is at least 70% identical to any of SEQ ID NOs: 2, 6, 29, 61, 66, 10, 12, 18, 19, 26, or 65.
3. The system according to claim 1, wherein the donor recognition sequence includes a donor attachment site configured to bind to the recombinase.
4. The system according to claim 1, wherein the first polynucleotide further comprises a cargo DNA sequence.
5. The system according to claim 4, wherein the cargo DNA sequence is greater than 1 kilobase pair.
6. The system according to claim 1, wherein the first polynucleotide further comprises a recipient recognition sequence for the recombinase, or the system further comprises a second polynucleotide comprising a recipient recognition sequence for the recombinase.
7. The system according to claim 6, wherein the recipient recognition sequence includes a recipient attachment sequence configured to bind to the recombinase.
8. The system according to claim 6, wherein the donor recognition sequence, the recipient recognition sequence, or both are false recognition sequences.
9. The system according to any one of claims 1 to 8, wherein the system is a cell-free system.
10. A composition comprising the system described in any one of claims 1 to 8.
11. A cell comprising the system described in any one of claims 1 to 8.
12. A method for modifying target DNA in vitro or ex vivo, wherein the target DNA is: The method comprising contacting a polypeptide having an amino acid sequence having at least 70% identity with any of SEQ ID NOs: 1 to 74 and 88 to 1183, or a nucleic acid encoding the same.
13. The method according to claim 12, wherein the target DNA comprises a donor recognition sequence, a recipient recognition sequence, or both.
14. The method according to claim 12, further comprising contacting the target DNA with a first polynucleotide comprising the donor recognition sequence of the recombinase.
15. The method according to claim 14, wherein the first polynucleotide further comprises a cargo DNA sequence.
16. The method according to claim 15, wherein the cargo DNA sequence is greater than 1 kilobase pair.
17. The method according to claim 14, wherein the target DNA includes a recipient attachment sequence configured to bind to the recombinase.
18. The method according to claim 12, wherein the target DNA sequence encodes a gene product.
19. The method according to claim 12, wherein the target DNA is located inside a cell, and the contact includes introducing it into the cell.
20. The composition according to claim 10 for use in modifying target DNA, wherein the target DNA comprises a donor recognition sequence, a recipient recognition sequence, or both, or the modification of the target DNA comprises inserting a first polynucleotide comprising the donor recognition sequence of the recombinase into the target DNA.