Engineered large serine recombinases
Engineered Dn29 recombinases with targeted mutations enhance genome editing efficiency and specificity in eukaryotic cells by improving site-specific recombination and integration, addressing the limitations of existing LSRs in eukaryotic genome manipulation.
Patent Information
- Application Number
- PCT/US2025/012631
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-31
AI Technical Summary
There is a need for effective tools to manipulate eukaryotic genomes, as existing large serine recombinases (LSRs) have limitations in genome engineering in eukaryotic cells, necessitating the discovery of new recombinases and methods to broaden the range of available genetic engineering tools.
Engineered Dn29 recombinases with specific mutations, such as E70G, F138L, A224P, N341Q, and others, are developed to enhance on-target integration efficiency and specificity, allowing for site-specific recombination and integration of DNA sequences into eukaryotic genomes, including the use of engineered Dn29-fused DNA binding domains like dCas9 for targeted genome editing.
The engineered Dn29 recombinases improve on-target integration efficiency and specificity, enabling precise genetic manipulation in eukaryotic cells, including integration, excision, inversion, and translocation of DNA fragments, and reducing off-target effects.
Smart Images

Figure US2025012631_31072025_PF_FP_ABST
Abstract
Description
ENGINEERED LARGE SERINE RECOMBINASES
[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 623,718 filed January 22, 2024, the contents of which is hereby incorporated by reference in its entirety.
[0002] All patents, patent applications and publications cited herein are hereby incorporated by reference in their entirety. The disclosures of these publications in their entireties are hereby incorporated by reference into this application.
[0003] This patent disclosure contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure as it appears in the U.S. Patent and Trademark Office patent file or records, but otherwise reserves any and all copyright rights.SEQUENCE LISTING
[0004] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on January 21, 2025, is named 2220476-0127W01_SL.xml and is 230,688 bytes in size.BACKGROUND OF THE INVENTION
[0005] Despite recent progress in genome engineering, there is still a demand for an effective approach to manipulate eukaryotic genomes. Large serine recombinases (LSRs) have evolved for this purpose in microbial cells. However, the previously studied LSRs possess various limitations on their use for genome engineering in eukaryotic cells. There is a need for the discovery of new recombinases and methods to identify them, aiming to broaden the range of tools available for genetic engineering.SUMMARY OF THE INVENTION
[0006] It is understood that any of the embodiments described below can be combined in any desired way, and that any embodiment or combination of embodiments can be applied to each of the aspects described below, unless the context indicates otherwise.
[0007] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising at least 70% identity to SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from E70, F138, A224, N341,L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
[0008] In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0009] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
[0010] In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0011] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising two or more amino acid mutations, wherein at least one mutation is selected from E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H and at least one mutation selected is selected I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0012] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, or G227. In some embodiments, the mutation comprises M6I, E70G, A224P, or G227V. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, and G227V. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234. In some embodiments, the engineered Dn29 further comprises a mutation at amino acid residue N341. In some embodiments, the mutation comprises N341K. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selectedfrom E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0013] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q332 or N341.
[0014] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, Q332K, or N341K. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q332K, and N341K. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341 V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from 1233, K248, Q390, L388, L393, or D503. In some embodiments, the mutation comprises I233K, K248R, Q390P, L388P, L393P, or D503N. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from D503, L388, Q390, L393, W373, R389, E452, F138, or V386. In some embodiments, the mutation comprises D503N, L388P, Q390P, L393P, W373R, R389S, E452G, F138L, or V386A. In some embodiments, the engineered Dn29 further comprises a mutation of one or more amino acid residue 1233, K248, Q390, L388, L393, or D503 and a mutation of one or more amino acid residue D503, L388, Q390, L393, W373, R389, E452, F138, or V386. In some embodiments, the mutation comprises I233K, K248R, Q390P, L388P, L393P, D503N, W373R, R389S, E452G, F138L, or V386A.
[0015] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, or N341
[0016] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, or N341Q. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, and N341Q. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from W373, L388, L393, or D503. In some embodiments, the mutation comprises W373R, L388P, L393P, or D503N. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from 1233, 1303, K248, or Q332. In some embodiments, the mutation comprises I233K, I303K, K248R, or Q332K. In some embodiments, the engineered Dn29 further comprises mutation of one or more amino acid residue W373, L388, L393, or D503 and a mutation of one or more amino acid residue 1233, 1303, K248, or Q332. In some embodiments, the mutation comprises W373R, L388P, L393P, D503N, I233K, I303K, K248R, or Q332K.
[0017] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, 1303, N341, or L388.
[0018] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, I303K, N341Q or L388P. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, I303K, N341Q and L388P. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224,N341, L388, V386, Q390, L393, D503, F138, V202, W373, S322, F374, W373, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, S428, or D437. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0019] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, 1303, N341, or L393.
[0020] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, I303K, N341Q, or L393P. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, I303K, N341Q, and L393P. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, F138, V202, W373, S322, F374, W373, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, S428, or D437. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0021] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q322, N341, L393, or D503.
[0022] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, Q322K, N341K, L393P, or D503N. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, and D503N. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404,L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, N341I, N341H, N341L, N341A, N341K, N341C, L388P, V386A, Q390P, L393P, D503N, F138L, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, I303K, Q332K, I233K, K248R, I98V, R246G, V150K, S428I, or D437G.
[0023] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q322, N341, L388, or R389.
[0024] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, Q322K, N341K, L388P, or R389S. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L388P, and R389S. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0025] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q322, N341, L393, D503, 198, 1233, 1303, or K248.
[0026] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, or K248R. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233,K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341 V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0027] In certain aspects the subject matter described herein provides an engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, 1233, K248, 1303, Q332, N341, L388, or Y404.
[0028] In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P, or Y404C. In some embodiments, the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P, or Y404C and optionally one or more mutations selected from D503N, W373R, L449P, N214D, I98V, or K288E. In some embodiments, a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318. In some embodiments, the engineered Dn29 further comprises a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, F138, V202, W373, S322, F374, W373, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, S428, or D437. In some embodiments, the mutation comprises E70G, F138L, A224P, N341Q, N341 V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
[0029] In certain aspects the subject matter described herein provides a fusion polypeptide, wherein the fusion polypeptide comprises an engineered Dn29 portion comprising the engineered Dn29 disclosed herein and a DNA binding domain (DBD) portion.
[0030] In some embodiments, the engineered Dn29 portion is fused N-terminal to the DBD portion. In some embodiments, the fusion polypeptide further comprises a peptide linker positioned between the engineered Dn29 portion and the DBD portion. In some embodiments, the peptide linker comprises 2 to 100 amino acids. In some embodiments, the peptide linker comprises glycine and serine residues, one or more XTEN16 repeats, or a combination thereof. In some embodiments, the peptide linker comprises SEQ ID NOs: 18- 26. In some embodiments, the DBD portion comprises Cas9, Cpfl, Casl2b, Casl2c, Casl2d,Casl2e, Casl2f, Casl2h, Casl2i, or Casl2g. In some embodiments, the Cas9, Cpfl, Cast 2b, Cast 2c, Cast 2d, Casl2e, Casl2f, Casl2h, Casl2i, or Cast 2g lack nuclease and / or nickase activity. In some embodiments, the DBD portion comprises dCas9. In some embodiments, the DBD portion comprises an amino acid sequence at least 90% identical to dCas9 (SEQ ID NO: 10), dCas9-HFl (SEQ ID NO: 11), dCas9-SpG (SEQ ID NO: 12), or dCas9-SpG-HFl (SEQ ID NO: 13). In some embodiments, the DBD portion comprises an amino acid sequence of dCas9 (SEQ ID NO: 10), dCas9-HFl (SEQ ID NO: 11), dCas9-SpG (SEQ ID NO: 12), or dCas9-SpG-HFl (SEQ ID NO: 13). In some embodiments, the engineered Dn29 disclosed herein or the fusion polypeptide disclosed herein further comprising one or more nuclear localization signals (NLSs). In some embodiments, the DBD portion of the fusion polypeptide binds to a guide RNA (gRNA).
[0031] In certain aspects the subject matter described herein provides a nucleic acid encoding any engineered Dn29 disclosed herein.
[0032] In certain aspects the subject matter described herein provides a nucleic acid encoding any fusion polypeptide disclosed herein.
[0033] In certain aspects the subject matter described herein provides a vector comprising any of the nucleic acids encoding any engineered Dn29 disclosed herein.
[0034] In certain aspects the subject matter described herein provides a vector comprising any of the nucleic acids encoding any fusion polypeptide disclosed herein.
[0035] In certain aspects the subject matter described herein provides a host cell comprising any vector disclosed herein.
[0036] In certain aspects the subject matter described herein provides a nucleic acid editing system comprising a first nucleic encoding any engineered Dn29 disclosed herein and a second nucleic acid encoding a gRNA.
[0037] In some embodiments, the gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracr RNA portion, wherein the nucleic acid sequence of the spacer sequence portion is the same as a target nucleic acid sequence, except that T in the target nucleic acid sequence is U in the spacer sequence portion, and wherein the target nucleic acid sequence is within 80 nucleotides upstream or downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest.
[0038] In some embodiments, the spacer sequence portion is 16 to 20 nucleotides long. In some embodiments, the gRNA encoded by the nucleic acid is an sgRNA. In someembodiments, the nucleic acid editing system wherein immediately 3’ to the target nucleic acid sequence on the DNA of interest is a PAM sequence. In some embodiments, the target nucleic acid sequence is within 80 nucleotides upstream or downstream of a dinucleotide core of an attA site of the engineered Dn29 portion of the fusion polypeptide on a target DNA of interest. In some embodiments, the attA site is a pseudosite in a mammalian target DNA of interest. In some embodiments, the attA site is a pseudosite in the human genome (attH). In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attH site is chrl0:21130404-21130406:-, chrl 1:77367459- 77367461 :-, chrl :230490334-230490336:+, chr2: 14280297-14280299:+, chr9: 116464427- 116464429:+, chr20:38982599-38982601:+, chr5:3553012-3553014:-, chr7: 134676315- 134676317:-, chrl0:58514255-58514257:+, or chr4:92338934-92338936:+. In some embodiments, the attA pseudosite is a pseudosite in a non-human primate genome. In some embodiments, the attA site is a pseudosite in a mouse (Mus musculus) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chrX:98,518,458. In some embodiments, the attA site is a pseudosite in a common marmoset (Callithrix jacchus) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr7: 19,911,959. In some embodiments, the attA site is a pseudosite in a Rhesus monkey (Macaco mulatto) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr9:22, 186,336. In some embodiments, the attA site is a pseudosite in a Cynomolgus monkey (Macaca fascicularis) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr9:21,946,090.
[0039] In some embodiments, the tracr RNA portion comprises SEQ ID NO: 108. In some embodiments, the target nucleic acid sequence is within 80 nucleotides upstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest. In some embodiments, the target nucleic acid sequence is within 80 nucleotides downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest. In some embodiments, the nucleic acid editing system further comprises a third nucleic acid encoding a second gRNA. In some embodiments, the second gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracr RNA portion, wherein the nucleic acid sequence of the spacer sequence portion is the same as a target nucleic acid sequence, exceptthat T in the target nucleic acid sequence is U in the spacer sequence portion, and wherein the target nucleic acid sequence is within 80 nucleotides downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest.
[0040] In some embodiments, the spacer sequence portion of the second gRNA is 16 to 20 nucleotides long. In some embodiments, the second gRNA encoded by the nucleic acid is an sgRNA. In some embodiments, the nucleic acid editing system wherein immediately 3’ to the target nucleic acid sequence on the DNA of interest is a PAM sequence. In some embodiments, the nucleic acid editing system further comprises a third nucleic acid comprising a donor DNA sequence which comprises an attD attachment site of the engineered Dn29 portion of the fusion polypeptide and a nucleic acid sequence for insertion into the target DNA of interest.
[0041] In some embodiments, the third nucleic acid further comprises a portion that has the same target nucleic acid sequence for the gRNA as the target DNA of interest. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises: dCas9 (SEQ ID NO: 10), the attH site on the target DNA of interest is chromosomal locus chrl0:21130404- 21130406:-, chrl 1 :77367459-77367461 :-, chr 1:230490334-230490336:+, chr2: 14280297- 14280299:+, chr9: 116464427-116464429:+, chr20:38982599-38982601:+, chr5:3553012- 3553014:-, chr7:134676315-134676317:-, chrl0:58514255-58514257:+, or chr4:92338934- 92338936:+ or comprises the attH sequence found at said chromosomal locus, and the attD attachment site of the donor DNA sequence comprises SEQ ID NO: 109, or a sequence 90% identical to SEQ ID NO: 109. In some embodiments, the third nucleic acid is a plasmid. In some embodiments, the third nucleic acid is a linear amplicon. In some embodiments, the nucleic acid encoding the fusion polypeptide, the nucleic acid encoding the gRNA, or both, and / or, where present, the third nucleic acid encoding the second gRNA are expressed from an inducible promoter.
[0042] In certain aspects the subject matter described herein provides a nucleic acid editing system comprising a first nucleic acid encoding any engineered Dn29 disclosed herein.
[0043] In some embodiments, an attA site of the engineered Dn29 on a target DNA of interest is a pseudosite in a mammalian target DNA of interest. In some embodiments, the attA site is a pseudosite in the human genome (attH). In some embodiments, the attH site is chrl0:21130404-21130406:-, chrl 1 :77367459-77367461 :-, chrl :230490334-230490336:+,chr2: 14280297-14280299:+, chr9: 116464427-116464429:+, chr20:38982599-38982601 :+, chr5:3553012-3553014:-, chr7: 134676315-134676317:-, chrl0:58514255-58514257:+, or chr4:92338934-92338936:+. In some embodiments, the attA pseudosite is a pseudosite in a non-human primate genome. In some embodiments, the attA site is a pseudosite in a mouse (Mus musculus) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chrX:98,518,458. In some embodiments, the attA site is a pseudosite in a common marmoset (Callithrix jacchus) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr7: 19,911,959. In some embodiments, the attA site is a pseudosite in a Rhesus monkey (Macaca mulatto) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr9:22, 186,336. In some embodiments, the attA site is a pseudosite in a Cynomolgus monkey (Macaca fascicularis) genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr9:21,946,090
[0044] In some embodiments, the nucleic acid editing system further comprises a second nucleic acid comprising a donor DNA sequence which comprises an attD attachment site of the engineered Dn29 and a nucleic acid sequence for insertion into the target DNA of interest. In some embodiments, the attH site on the target DNA of interest is chromosomal locus chrl0:21130404-21130406:-, chrl 1 :77367459-77367461 :-, chrl :230490334-230490336:+, chr2: 14280297-14280299:+, chr9: 116464427-116464429:+, chr20:38982599-38982601 :+, chr5:3553012-3553014:-, chr7: 134676315-134676317:-, chrl0:58514255-58514257:+, or chr4:92338934-92338936:+ or comprises the attH sequence found at said chromosomal locus, and the attD attachment site of the donor DNA sequence comprises SEQ ID NO: 109, or a sequence 90% identical to SEQ ID NO: 109.
[0045] In some embodiments, the second nucleic acid is a plasmid. In some embodiments, the second nucleic acid is a linear amplicon. In some embodiments, the nucleic acid encoding the engineered Dn29 is expressed from an inducible promoter.
[0046] In certain aspects the subject matter described herein provides a vector comprising any of the nucleic acids of any nucleic acid editing system disclosed herein.
[0047] In certain aspects the subject matter described herein provides a host cell comprising any of the vector(s) comprising any of the nucleic acids of any nucleic acid editing system disclosed herein.
[0048] In certain aspects the subject matter described herein provides a method of integrating a donor DNA sequence into a target DNA of interest of a cell, the method comprising introducing into the cell: any nucleic acid editing system disclosed herein.
[0049] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a human embryonic stem cell, a primary T cell, or a non-dividing cell. In some embodiments, the cell is a non-human primate cell. In some embodiments, the cell is a mouse cell. In some embodiments, the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site. In some embodiments, the donor DNA comprises an engineered Dn29 attD attachment site which is integrated into the target DNA of interest. In some embodiments, the target DNA of interest of the cell is the genome of the cell. In some embodiments, the target DNA of interest of the cell is a plasmid.
[0050] In certain aspects the subject matter described herein provides a method of inverting a DNA sequence of a target DNA of interest, the method comprising introducing into a cell: any nucleic acid editing system disclosed herein, wherein attD and attA attachment sites of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) are present on the same DNA target molecule of interest in reverse orientation.
[0051] In some embodiments, the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site. In some embodiments, the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attD attachment site. In some embodiments, the target DNA of interest of the cell is the genome of the cell.
[0052] In certain aspects the subject matter described herein provides a method of excising a DNA sequence of a target DNA of interest, the method comprising introducing into a cell: any nucleic acid editing system disclosed herein, wherein attD and attA attachment sites of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) are present on the same DNA target molecule of interest in the same orientation.
[0053] In some embodiments, the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site. In some embodiments, the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attD attachment site. In some embodiments, the target DNA of interest of the cell is the genome of the cell.
[0054] In certain aspects the subject matter described herein provides a method of translocating DNA sequences between two linear target DNA molecules of interest, the method comprising introducing into a cell: any nucleic acid editing system disclosed herein, wherein an attD attachment site of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) portion of the fusion polypeptide is present on a first linear target DNA molecule and an attA attachment site of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) is present on a second linear target DNA molecule.
[0055] In some embodiments, the first target DNA molecules of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site. In some embodiments, the second target DNA molecules of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attD attachment site. In some embodiments, the linear target DNA molecules of interest of the cell are chromosomes of the cell.
[0056] Other embodiments of the invention are further described in the following sections of the application, including the Detailed Description, Examples, and Claims. Still other objects and advantages of the invention will become apparent by those of skill in the art from the disclosure herein, which are simply illustrative and not restrictive. Thus, other embodiments will be recognized by the ordinarily skilled artisan without departing from the spirit and scope of the invention.BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The patent or application file contains at least one drawing executed in color. To conform to the requirements for PCT patent applications, many of the figures presented herein are black and white representations of images originally created in color.
[0058] FIG. 1 shows a schematic of substrate-linked Directed Evolution selection method. The pEVO plasmid expresses a library of Dn29 variants from an arabinose inducible promoter, araBAD. This plasmid backbone also contains the two recombinase attachment sites, attP and attHl, flanking an Ndel restriction site. Upon recombination, the attP and attHl sequences recombine, removing the Ndel restriction site from the plasmid. The plasmid library is digested with Ndel, resulting in digestion of only non-active variants. PCR amplifying across the Ndel digest site results in amplification of non-digested, active variants, which can be re-cloned into the original backbone for subsequent rounds of evolution.
[0059] FIG. 2 shows a schematic of substrate linked directed evolution cycling and methodological overview. The addition of steps of cycling to the DE methodology utilized herein are shown in Figure 2. Cycling iteratively increases selection pressure for greater enrichment of beneficial mutations with each round, and DNA shuffling allows for stacking of mutations. The library of recombinase variants is electroporated into E. coli, and expression is induced with L-arabinose addition to the culture. The cells are grown overnight and then the plasmid library is recovered via midiprep. The selection digest and PCR as described in Figure 1 is performed to recover active variants, and an optional DNA shuffling step is performed. Finally, this new variant library is restriction enzyme digested and ligated into the pEVO backbone for subsequent rounds of evolution.
[0060] FIG. 3 shows a schematic of the generation of a library using NKK deep mutational scanning cloning strategy. For each position in the CDS, two primers are designed with NNK at the specified position, and stitched together with overlap extension PCRs to generate the 16,576 variant library of single mutations.
[0061] FIG. 4 shows the quality control assessment of the input library. As shown in Figure 4 the input library generated by the methodology described herein has only a few dropouts. The library of variants is sequenced and mutations are quantified. All possible amino acids for all positions in the CDS appear in the library, except for dropouts at 10 positions. A * indicates a stop codon (shown in black). The analysis reveals minimal dropouts, which may be attributed to inherent sequencing errors.
[0062] FIG. 5 shows QC of the input library. The input library mutations generated herein are evenly distributed within a position but more uneven across positions as shown in Figure 5. The percent of reads containing a given mutation is colored, dropouts are shown in black, and the wildtype base is shown in white. If the library was evenly distributed, the expected percent of reads for a given mutation would be 0.006%, shown in gray. Each row is normalized by the number of codons that encode that amino acid when limiting codons to NNK.
[0063] FIG. 6 shows a schematic of progression through directed evolution of Dn29. Each library is labeled with a number corresponding to (Shuffle Number. Selection Cycle Number). Within each box is a set of selection cycles, with the number of cycles performed indicated underneath the arrow. Upon shuffling the library, the selection cycle number restarts to 0. Arabinose concentration used for that set of selection cycles is indicated at the bottom of each box.
[0064] FIG. 7 shows a comparison of distribution of mutations in the input library and the final library after 2 rounds of shuffling and 12 rounds of selection. After DE, the output library generated here has an altered distribution of mutations as shown in Figure 7.
[0065] FIG. 8 shows mutations present in the output library (2.6 Replicate 1). A white box indicates the wildtype base at each position, a black box indicates a mutation is fully dropped out from the library, and a gray box indicates the presence of at least one variant containing that mutation. A * indicates a stop codon. The output library has high rates of non-functional mutant dropout as shown in Figure 8.
[0066] FIG. 9 shows the distribution and percentage of mutations present in the output library (2.6 Replicate 1). A white box indicates the wildtype base at each position, a black box indicates a mutation is fully dropped out from the library, and a colored box indicates the presence of at least one variant containing that mutation, with gray indicating the expected percent of reads if all mutations were evenly distributed, and red and blue indicating an enrichment and depletion of that mutation, respectively. The output library also has high rates of mutant depletion, as shown in Figure 9.
[0067] FIG. 10 shows mutation enrichment plots of the first 84 residues throughout the progression of the directed evolution. Enrichment score is calculated comparing each evolved library to the input library. As shown in Figure 10, the methodology described herein leads to evolution and enrichment progress with increasing cycles of evolution. Enrichment mutational maps across all fragments show mutational hotspots in the clones.
[0068] FIG. 11 shows mutation enrichment plots of the full CDS after 2 shuffles and 12 cycles. Enrichment score is calculated comparing each evolved library to the input library.
[0069] FIG. 12 shows attHl integration efficiency in HEK293FT cells of randomly picked clones from directed evolution libraries after 5 cycles, 1 shuffle and 7 cycles, and 2 shuffles and 12 cycles, normalized to wildtype Dn29 efficiency. Right panel shows the ratio of attHl integration efficiency / attH3 integration efficiency of the same clones as shown in the left panel, normalized to wildtype Dn29 attHl / attH3 ratio.
[0070] FIG. 13 shows attHl integration efficiency of 248 mutants, normalized to wildtype Dn29. Further details of the clones shown in this figure are provided in Table 2.
[0071] FIG. 14 shows the specificity (attHl / attH3) and efficiency (attHl) of all mutants in Figure 13, normalized by wildtype Dn29. Dotted lines indicate thresholds used to classify variants as specificity or efficiency enhancing. Further details of the clones shown in this figure are provided in Table 2.
[0072] FIG. 15 shows attHl integration efficiency and attHl / attH3 ratio, normalized to wildtype Dn29, of variant 127, which contains all mutations in variant 62 (M6I / E70G / A224P / G227V / A234) and variant 93 (A318 / N341K). “A” indicates a silent mutation in the nucleic acid sequence encoding the Dn29 variant.
[0073] FIG. 16 shows efficiency determining driver mutations from assayed variants.Further details of the clones shown in this figure are provided in Table 2. As shown in Figure16, all single mutations from variants with 1.5x WT integration efficiency or 2x WT attHl / attH3 ratio were layered on top of variant 127 to determine which mutations in a mutant are driving efficiency-related phenotypic changes and which are passenger mutations. Figure 16 shows the attHl integration efficiency.
[0074] FIG. 17 shows specificity determining driver mutations from assayed variants.Further details of the clones shown in this figure are provided in Table 2. As shown in Figure17, all single mutations from variants with 1.5x WT integration efficiency or 2x WT attHl / attH3 ratio were layered on top of variant 127 to determine which mutations in a mutant are driving specificity-related phenotypic changes and which are passenger mutations. Figure 17 shows the attHl / attH3 ratio.
[0075] FIG. 18 shows all specificity driver mutations mapped to the Alphafold3 structure ofDn29 bound to attB-R. Statistically significant mutations are: I303K, Q332K N341K, I233K, K248R, I98V, V150K, S288A, R246G, K288E, N341R, D437G.
[0076] FIG. 19 shows all efficiency driver mutations mapped to the Alphafold3 structure of Dn29 bound to attB-R.. Statistically significant mutations are: E70G, F138L, A224P, N341Q, N341V, N341I, N341H, L388P, V386A, Q390P, L393P, D503N, F138L, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D.
[0077] FIG. 20 shows integration efficiency (attHl integration) and specificity (attHl / attH3 ratio), plotted as fold change to wildtype, of variant 127 with position 341 saturation mutagenesis. Position 341 is a driver mutation of both specificity and efficiency. Further details of the clones shown in this figure are provided in Table 2.
[0078] FIG. 21 shows integration efficiency vs. specificity of 341K and 341Q lineages with driver mutations. Dots and error bars represent the mean ± SD of n=2 biological replicates for each experimental condition (n=3 for WT and variant 127).
[0079] FIG. 22 shows round three of stacking driver mutations. Further details of the clones shown in this figure are provided in Table 2. Each variant plotted contains the 388 base mutations (M6I / E70G / A224P / G227V / A234 / N341Q / A318), plus one of the fourefficiency mutations in the set (W373R, L388P, L393P, D503N) and one, three, or four specificity mutations in the set (I233K, I303K, K248R, Q332K). The same data is plotted on the top and bottom plots, with the shapes indicating efficiency mutation (top plot) or specificity mutation (bottom plot).
[0080] FIG. 23 shows round three of stacking additional driver mutations. Further details of the clones shown in this figure are provided in Table 2. Each variant plotted contains the 381 base mutations (M6I / E70G / A224P / G227V / A234 / Q332K / N341K / A318), plus one mutation in set 1 : {I233K, K248R, Q390P, L388P, L393P, D503N}, and 1 mutation in set 2: {D503N, L388P, Q390P, L393P, W373R, R389S, E452G, F138L, V386A}. The legend indicates the mutation in set 1. On the bottom is a zoomed in version of the plot on the top. LOD is indicated as the attH3 percent detected by ddPCR in wells transfected with a mismatched LSR negative control.
[0081] FIG. 24 shows variants in Figure 22 assayed for off-target integration at a second off-target site on chromosome 10 (a pseudosite).
[0082] FIG. 25 shows on-target efficiency of WT (n=10 biological replicates), superDn29 (n=5 biological replicates, previously variant 608), goldDn29 (n=5 biological replicates, previously variant 538), and hifiDn29 (n=3 biological replicates, previously variant 637). Bars and error bars represent the mean ± SD, and dots represent each replicate. Data presented is the same as shown in Figure 1H. Asterisks show t-test significance compared to WT. *=two-tailed p<0.05, ****=two-tailed pO.OOOl.
[0083] FIG. 26 shows efficiency of integration into a single off-target (attH3) of WT (n=10 biological replicates), superDn29 (n=2 biological replicates), goldDn29 (n=5 biological replicates), and hifiDn29 (n=3 biological replicates). Bars and error bars represent the mean ± SD, and dots represent each replicate. Asterisks show t-test significance compared to WT. ***=two-tailed p<0.001, ****=two-tailed p<0.0001.
[0084] FIG. 27 shows genome-wide specificity of on-target integration compared to all genomic insertions of WT (n=3 biological replicates) superDn29 (n=4 biological replicates), goldDn29 (n=2 biological replicates) and hifiDn29 (n=2 biological replicates). Bars and error bars represent the mean ± SD, and dots represent each replicate. Asterisks show t-test significance compared to WT. *=one-tailed p<0.05, ***=one-tailed p<0.001.
[0085] FIG. 28 shows integration efficiencies of Dn29 variants at attHl, with and without the dCas9 fusion and guide Dn29-Hl-3. The bars and error bars represent the mean ± SD of n=3 biological replicates, shown as dots.
[0086] FIG. 29 shows attHl and attH3 integration percentage of many efficiency and specificity variants, with dCas9 fusion and without dCas9 fusion and guide Dn29-Hl-3. As Figure 29 shows, integration at attHl is much lower without dCas9 fusion.
[0087] FIG. 30 shows genome-wide specificity profile for wildtype Dn29, Dn29-dCas9, superDn29-dCas9, goldDn29-dCas9, and hifiDn29-dCas9. As shown in Figure 30, for WT Dn29 UMI specificity for attHl is about 12% while the variant-dCas9 fusions exhibit about 60%-97% specificity for attHl.
[0088] FIG. 31 shows round four of stacking additional driver mutations. Each variant plotted contains the 511 base mutations M6I / E70G / A224P / G227V / I233K / A234 / K248R / I303K / A318 / Q332K / N341Q / L388P), plus Y404C and one mutation in the set {D503N, W373R, L449P, N214D, I98V, K288E}.
[0089] FIG. 32 shows round five of stacking additional driver mutations in an LSR- dCas9 fusion context (top) and unfused context (bottom). Each variant plotted contains the 538 base mutations (M6I / E70G / A224P / G227V / A234 / A318 / Q322K / N341K / L393P / D503N) plus 1-4 of the specificity mutations in the set {I98V, I233K, I303K, K248R}.
[0090] FIG. 33 shows the genome-wide specificity for attHl for multiple replicates of wildtype Dn29, Dn29-dCas9, superDn29-dCas9, goldDn29-dCas9, and hifiDn29-dCas9.
[0091] FIG. 34 shows exemplary attD sequences (SEQ ID NOs: 109) and corresponding attH pseudosites (provided as chromosomal locus according to human genome assembly GRCh38, available at www.ncbi.nlm.nih.gov / genome / guide / human / ) for Dn29.
[0092] FIG. 35 shows the efficiency and specificity of significant variants harboring driver mutations (one-tailed p<0.05), shown as fold change to Variant 127. Variants are generated by adding individual mutations from the enhanced variants in Figure 14 on top of variant 127. Dots represent the mean of n=2 biological replicates.
[0093] FIG. 36 shows the integration efficiency (orange, left y-axis) and specificity (teal, right y-axis) of variant 127 with lysine scan mutations of putative DNA binding residues. Bars and error bars represent the mean ± SD of n=2 biological replicates.
[0094] FIG. 37 shows the integration efficiency (orange, left y-axis) and specificity (orange, right y-axis) of variant 381 with significant (one-tailed p<0.05) mutations from second validation round. Bars and error bars represent the mean ± SD of n=6 (variant 381) or n=2 (other variants) biological replicates.
[0095] FIG. 38 shows the efficiency (top, variants 639-647) and specificity (bottom, variants 649-651) of top model-guided combinatorial mutants, generated by predicting theactivity of combining two mutations on top of superDn29. Each dot represents the mean of n=6 biological replicates for WT and superDn29, and n=2 biological replicates for all model- guided variants.
[0096] FIG. 39 shows the integration efficiencies of Dn29 variants and dCas9 fusions at attHl, with and without cell cycle arrest by aphi dicolin treatment. The dots represent the mean of n=3 biological replicates
[0097] FIG.40 shows the integration efficiencies of Dn29 variants and dCas9 fusions in human embryonic stem cells (hESCs). Bars and error bars represent the mean ± SD of n=3 biological replicates, shown as dots.
[0098] FIG.41 shows the integration efficiency of 12 kb CRISPRi donor by LSR-dCas9s at attHl. Bars and error bars represent the mean ± SD of n=3 biological replicates, shown as dots.
[0099] FIG.42 shows CRISPRi-BFP cassette expression in engineered hESCs after selection, pre- and post- differentiation into HPCs. Bars and error bars represent the mean ± SD of n=3 biological replicates, shown as dots.
[0100] FIG.43 shows the genotyping of hESC single cell clones engineered with goldDn29-dCas9. N of clones analyzed per sample is indicated in the legend.
[0101] FIG.44 shows HPC cell surface marker expression after guide transduction and selection, relative to non-targeting guide, in CRISPRi cells engineered by goldDn29-dCas9. Left histograms are representative cell surface marker expression plots. Right bar plots show the knockdown quantification of 4 biological replicates, calculated as target / non-target median fluorescence intensity, represented as a percentage.
[0102] FIG.45 shows the integration efficiencies of Dn29 variants and dCas9 fusions at attHl in primary human T cells using scAAV donor. Bars and error bars represent the mean ± SD of n=4 biological replicates, each originating from a different blood donor.
[0103] FIG.46 shows (top) a schematic of plasmid recombination assay for attachment site recombination. mCherry expresses upon recombination between attachment sites X and Y. (Bottom) recombination of Dn29, key variants, and mismatching LSR control between attP, attB, attL, and attR, measured by mCherry median fluorescence intensity (MFI). Dotted line indicates the background fluorescence associated with the mismatching LSR control. Bars and error bars represent the mean ± SD of n=3 biological replicates, shown as dots.
[0104] FIG.47 shows the specificity analysis of HEK293FT single cell clones engineered with Dn29 and hifiDn29, with and without dCas9 fusions. N of clones analyzed per sample is indicated in the x-axis.
[0105] FIG.48 shows the on-target insertion copy number per clone for hifiDn29 and Dn29, with and without dCas9 fusion.
[0106] FIG.49 shows the on-target insertion copy number per clone for hifiDn29-dCas9 and Dn29-dCas9. N of clones is labeled above each bar.
[0107] FIG.50 shows the specificity of Dn29 and variants in hESCs, measured as attH3 off-target integration efficiency by ddPCR. Bars and error bars represent the mean + SD of n=3 biological replicates, shown as dots.
[0108] FIG.51 shows Hl hESC clones (n=37) edited with goldDn29-dCas9: BFP expression (top) and genotyping (bottom). Integration / reference ratio of 0.5 indicates heterozygous insertions, 1 indicates homozygous insertions. Single clone per bar / dot.
[0109] FIG.52 shows the integration efficiencies of Dn29 variants and dCas9 fusions at attHl in primary human T cells using ssAAV donor. Bars and error bars represent the mean ± SD of n=4 biological replicates, each originating from a different blood donor.
[0110] FIG.53 shows (Top) alignment of attHl-like pseudosites in human, marmoset, rhesus monkey, cynomolgus monkey, mouse and a sequence logo of the top 100 pseudosites in HEK293FTs. The attHl-like sequences used for alignment are disclosed in SEQ ID NOS: 209-213. (Bottom) schematic of plasmid recombination assay for testing attHl-like pseudosites in HEK293FTs. (Right) plasmid recombination efficiency between attP and each pseudosite, using Dn29 and goldDn29, in HEK293FTs. For the mouse pseudosite, the cognate attP plasmid is modified to contain the matching TA dinucleotide core sequence.
[0111] FIG. 54 shows integration efficiencies of superDn29-dCas9 at attHl in primary human T cells. Effector is in vitro transcribed and electroporated at 1.5 or 3 pg mRNA per condition. Donor is electroporated as a plasmid, at 2 or 4 pg per condition. sgRNA is delivered as a plasmid, electroporated at 1.5 pg per condition. Bars and error bars represent the mean ± SD of n=2 biological replicates, each originating from a different blood donor.DETAILED DESCRIPTION
[0112] The present invention relates to engineered large serine recombinases (LSRs). The engineered LSR recognizes two DNA sequences, also known as attachment sites, one of which is the target site and the other is a DNA sequence often found on a separate DNAmolecule. The engineered LSR performs site-specific recombination, integrating the DNA found on the separate DNA molecule into the target site. And in cases where the attachment sites are on the same molecule, depending on their relative orientation, the engineered LSR can perform excision or inversion recombination reactions. Further, translocation may occur when the attachment sites are on different molecules in a particular relative orientation. The engineered LSRs can increase on-target integration efficiency and / or specificity as compared to an LSR without any modifications. In some embodiments, the engineered LSR is an engineered Dn29.
[0113] In some embodiments, the engineered LSR may be used as part of an LSR-DNA binding domain (DBD) fusion as described in International Publication No. WO2024 / 097747, the contents of which is hereby incorporated by reference in its entirety. The DNA binding domain is targeted, via direct protein-DNA binding or RNA-guided targeting, to a site proximal to, overlapping with, or within the engineered LSR target site, directing the engineered LSR to a single, specific DNA attachment site, such as a pseudosite in a mammalian genome. In some embodiments, the engineered LSR is an engineered Dn29 fused to a DNA binding domain. In some embodiments, the engineered LSR is an engineered Dn29 fused to dCas9.
[0114] The terms “polynucleotide”, “nucleotide sequence”, “nucleic acid” and “oligonucleotide” are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof.Polynucleotides may have any three-dimensional structure, and may perform any function, known or unknown. The following are non-limiting examples of polynucleotides: coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, guide RNA (gRNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A polynucleotide may comprise one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.
[0115] The terms “polypeptide”, “peptide” and “protein” are used interchangeably herein to refer to polymers of amino acids of any length. The polymer may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The terms also encompass an amino acid polymer that has been modified; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component. As used herein the term “amino acid” includes natural and / or unnatural or synthetic amino acids, including glycine and both the D or L optical isomers, and amino acid analogs and peptidomimetics.
[0116] The term “percent sequence identity” refers to the percentage of nucleotides or nucleotide analogs in a nucleic acid sequence, or amino acids in an amino acid sequence, that is identical with the corresponding nucleotides or amino acids in a reference sequence after aligning the two sequences and introducing gaps, if necessary, to achieve the maximum percent identity. Hence, in case a nucleic acid according to the technology is longer than a reference sequence, additional nucleotides in the nucleic acid, that do not align with the reference sequence, are not taken into account for determining sequence identity. A number of mathematical algorithms for obtaining the optimal alignment and calculating identity between two or more sequences are known and incorporated into a number of available software programs. Examples of such programs include CLUSTAL-W, T-Coffee, and ALIGN (for alignment of nucleic acid and amino acid sequences), BLAST programs (e.g., BLAST 2.1, BL2SEQ, and later versions thereof) and FASTA programs (e.g., FASTA3x, FAS™, and SSEARCH) (for sequence alignment and sequence similarity searches). Sequence alignment algorithms also are disclosed in, for example, Altschul et al., J.Molecular Biol., 215(3): 403-410 (1990), Beigert et al., Proc. Natl. Acad. Sci. USA, 106(10): 3770-3775 (2009), Durbin et al., eds., Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press, Cambridge, UK (2009), Soding, Bioinformatics, 21(7): 951-960 (2005), Altschul et al., Nucleic Acids Res., 25(17): 3389- 3402 (1997), and Gusfield, Algorithms on Strings, Trees and Sequences, Cambridge University Press, Cambridge UK (1997)).
[0117] The practice of aspects of the present invention can employ, unless otherwise indicated, conventional techniques of cell biology, cell culture, molecular biology, transgenic biology, microbiology, recombinant DNA, and biochemistry, which are within the skill of the art. Such techniques are explained fully in the literature. See, e.g., Molecular Cloning A Laboratory Manual, 3rd Ed., ed. by Sambrook (2001), Fritsch and Maniatis (Cold SpringHarbor Laboratory Press: 1989); DNA Cloning, Volumes I and II (D. N. Glover ed., 1985); Oligonucleotide Synthesis (M. J. Gait ed., 1984); Mullis et al. U.S. Pat. No: 4,683,195; Nucleic Acid Hybridization (B. D. Hames & S. J. Higgins eds. 1984); Transcription and Translation (B. D. Hames & S. J. Higgins eds. 1984); Culture Of Animal Cells (R. I. Freshney, Alan R. Liss, Inc., 1987); Immobilized Cells and Enzymes (IRL Press, 1986); B. Perbal, A Practical Guide To Molecular Cloning (1984); the series, Methods In Enzymology (Academic Press, Inc., N.Y.), specifically, Methods In Enzymology, Vols. 154 and 155 (Wu et al. eds.); Gene Transfer Vectors For Mammalian Cells (J. H. Miller and M. P. Calos eds., 1987, Cold Spring Harbor Laboratory); Immunochemical Methods In Cell And Molecular Biology (Caner and Walker, eds., Academic Press, London, 1987); Handbook Of Experimental Immunology, Volumes I-FV (D. M. Weir and C. C. Blackwell, eds., 1986); Manipulating the Mouse Embryo, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1986) and subsequent versions thereof, the contents of each of which are hereby incorporated by reference in their entireties.
[0118] One skilled in the art can obtain a protein in several ways, which include, but are not limited to, isolating the protein via biochemical means or expressing a nucleotide sequence encoding the protein of interest by genetic engineering methods.
[0119] A protein is encoded by a nucleic acid (including, for example, genomic DNA, messenger RNA (mRNA), complementary DNA (cDNA), synthetic DNA, as well as any form of corresponding RNA). Nucleic acids encoding a protein can be produced via recombinant DNA technology and such recombinant nucleic acids can be prepared by conventional techniques, including chemical synthesis, genetic engineering, enzymatic techniques, or a combination thereof.Engineered Large Serine Recombinases
[0120] The present invention relates to an engineered large serine recombinase (LSR). In some embodiments, the engineered LSR is an engineered Dn29. In certain embodiments, the invention provides an engineered Dn29 or nucleic acid encoding an engineered Dn29 comprising one or more of the mutations as described in Tables 1 and 2. In some embodiments, the mutations described herein can be incorporated into an amino acid sequence of Dn29 (SEQ ID NO: 1). In some embodiments, a nucleic acid encoding an amino acid sequence of Dn29 (e.g., SEQ ID NO: 2) can be mutated so the nucleic acid encodes an amino acid sequence comprising any of the mutations described herein. The engineered Dn29 can increase on-target integration efficiency and / or specificity as compared to SEQ IDNO: 1. In some embodiments, the mutations described herein can be incorporated into an amino acid sequence having at least 70% identity to SEQ ID NO: 1, or a nucleic acid encoding said amino acid sequences. The engineered Dn29 can increase on-target integration efficiency and / or specificity as compared to the amino acid sequence without the mutations. The engineered Dn29-DBD fusion (e.g. engineered Dn29-dCas9 fusions (SEQ ID NOs:3-7)) can increase on-target integration efficiency and / or specificity as compared to a Dn29-DBD fusion wherein the Dn29 portion comprises SEQ ID NO: 1 (e.g. as compared to Dn29-dCas9 fusions (SEQ ID NO:3-7)). In some embodiments, the mutations described herein can be incorporated into a Dn29-DBD fusion wherein the Dn29 portion comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 2, or a nucleic acid encoding said amino acid sequences. In some embodiments, the mutations described herein can be incorporated into a Dn29-dCas9 fusions having at least 70% identity to SEQ ID NOs: 3-7, or a nucleic acid encoding said amino acid sequences.
[0121] The engineered Dn29 can mediate the recombination of DNA between recombinase recognition sequences, which results in the excision, integration, inversion, or exchange (e.g., translocation) of DNA fragments between the recombinase recognition sequences. See for example WO2023 / 081762 titled “SERINE RECOMBINASES” which content is herein incorporated by reference in its entirety. Recombinases have numerous applications, including the creation of gene knockouts / knock-ins and gene therapy applications. In particular, engineered Dn29 is useful in the nucleic acids, polypeptides, compositions, systems, and methods disclosed herein.
[0122] The engineered Dn29 recognizes two DNA sequences, also known as attachment sites, one of which is the target site and the other is a DNA sequence found on a separate DNA molecule (for integration embodiments). Engineered Dn29 performs a site-specific recombination between the two attachment sites. The native attachment sites targeted by engineered Dn29 are termed “attP” (phage) and “attB” (bacteria) sites wherein each of the attP and attB sites comprises two half-sites joined at a central sequence. The central sequence consists of a central dinucleotide sequence, described further herein. In general, the recombination reaction is performed by a tetramer of the recombinase, in which each subunit is bound to a half-site of the attP or attB site. During the recombination reaction, each of the attP and attB sites is cut into two half-sites, in which each half-site has an overhang region comprising the central sequence (e.g. the central dinucleotide). When applied for genomic integration of donor cargos into a genome, the terms attD (donor) and attA (acceptor) may beused to refer to the two attachment sites. Either an attP or an attB can be the attD or attA, depending on which sequence is chosen to be present on the donor molecule (e.g., if attP is attD, then attB is attA; if attB is attD, then attP is attA). In another embodiment, the attD integrates directly into an endogenous pseudosite natively found in the target genome. As described in Durrant MG, Fanton A, Tycko J, et al., Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome, Nat Biotechnol., 2023, 41(4):488-499, pseudosites can be experimentally determined by analyzing the sequences adjacent to successful integration of a donor molecule with an attD site - where the pseudosites will be adjacent to the attD half-sites. In the case of mammalian genome integration(s), for example, human genome integration endogenous pseudosite(s) is (are) termed attH; and therefore an attH site is a type of attA.
[0123] In some aspects, an engineered Dn29 or engineered Dn29-DBD fusion (e.g. Dn29- dCas9 fusion) is used for site-specific recombination, wherein DNA strand exchange takes place between DNA sequences possessing attB and attP sites (or attD and attA sites), and wherein the recombinase rearranges DNA segments by recognizing and binding to the attB and attP sites, at which they cleave the DNA backbone, exchange the two DNA helices involved and rejoin the DNA strands.
[0124] An engineered Dn29 or engineered Dn29-DBD fusion (e.g. Dn29-dCas9 fusion) can also site-specifically integrate DNA sequences of interest containing an attD into a DNA target of mammalian cells, both at pre-installed integration sites (e.g., a pre-installed attA) or at endogenous genomic pseudosites (e.g., attH). For example, a donor DNA sequence of interest containing a native attP site can be integrated into a DNA target with the corresponding native attB acceptor attachment site (also referred to as a “landing pad”). A donor DNA sequence of interest containing a native attB site can be integrated into a DNA target with the corresponding attP acceptor attachment site (also referred to as a “landing pad”). Mammalian DNA may also contain endogenous genomic pseudosites which have high sequence similarity to an attA site, and can functionally recombine with an attD. If the attA sequence is found in a mammalian genome, for example the human genome, it is termed an attH sequence. For example, a donor DNA sequence of interest containing a native attP site can be integrated into a DNA target with an attH pseudosite with high sequence similarity to the corresponding native attB acceptor attachment site. A donor DNA sequence of interest containing a native attB site can be integrated into a DNA target with an attH pseudosite with high sequence similarity to the corresponding native attP acceptor attachmentsite. Thus, an engineered Dn29 or engineered Dn29-DBD fusion (e.g. Dn29-dCas9 fusion) can be used to integrate a DNA sequence of interest into a target DNA, such as a cellular DNA. Despite their sequence specificity, an engineered Dn29 or engineered Dn29-DBD fusion (e.g. Dn29-dCas9 fusion) may integrate into numerous sites in a mammalian genome, such as the human genome, due to the presence of multiple loci with sufficient “attH” integration site sequences.
[0125] The systematic discovery of recombinases for integrating DNA into the human genome is described in WO2023 / 081762 and Durrant MG, Fanton A, Tycko J, et al., Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome, Nat Biotechnol., 2023, 41(4):488-499, the entire contents of both of which are hereby incorporated by reference in their entireties.
[0126] An exemplary engineered LSR as described herein is engineered Dn29 (SEQ ID NO: 1). The native attP and attB sequences for Dn29 are provided as SEQ ID NOs: 8 (attP Dn29) and 9 (attB Dn29). In some embodiments, the attachment site for the engineered Dn29 comprises a sequence that follows the consensus sequence logo motifs provided in Supplemental Figure 6C of Durrant, M.G., Fanton, A., Tycko, J. et al. Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome, Nat Biotechnol 41, 488-499 (2023), the content of which is hereby incorporated by reference in its entirety.
[0127] In certain aspects, described herein is an engineered LSR comprising the amino acid sequence of Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In certain aspects, described herein is a nucleic acid encoding an engineered LSR comprising the amino acid sequence of Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2.
[0128] In certain aspects, described herein is an engineered Dn29 comprising an amino acid sequence having 70% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In some embodiments, the amino acid sequence has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29 comprising an amino acid sequence having 70% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In some embodiments, the amino acid sequence has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%,98%, or 99% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2.
[0129] In certain aspects, any of the engineered Dn29 sequences described herein can be used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion). Thus, in certain aspects, described herein is a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion), wherein the Dn29 portion comprises the amino acid sequence of Dn29, (SEQ ID NO: 1) and comprises one or more mutations provided in Tables 1 and 2. In certain aspects, described herein is a nucleic acid encoding a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion), wherein the Dn29 portion comprises the amino acid sequence of Dn29, (SEQ ID NO: 1) and comprises one or more mutations provided in Tables 1 and 2.
[0130] In certain aspects, described herein is a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion), wherein the Dn29 portion comprises an amino acid sequence having 70% identity to amino acid sequence of Dn29, (SEQ ID NO: 1) and comprises one or more mutations provided in Tables 1 and 2. In certain aspects, described herein is a nucleic acid encoding a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion), wherein the Dn29 portion comprises an amino acid sequence having 70% identity to Dn29 (SEQ ID NO: 1) and comprises one or more mutations provided in Tables 1 and 2. In some embodiments, the amino acid sequence of the Dn29 portion has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2.
[0131] In certain aspects, described herein is an engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences. In some embodiments, the engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences is Dn29 (SEQ ID NOs: 1) comprising one or more mutations provided in Tables 1 and 2.
[0132] In some embodiments, the engineered Dn29 described herein include, without limitation, an engineered Dn29 comprising one or more of the following amino acid motifs, written in the common Prosite format, where x is any amino acid and x(n) represents n number of any amino acid (e.g., x(3) is xxx or 3 consecutive amino acids):
[0133] Motif 5: [ADEHKNQRS]-[ADEFGHKMNQRSWY]-[EFY]-[FHLWY]-x- [ADEFIKLMNQRSTY]-[FIQSTV]-[AGKLNRSTV]-[ADEHKNQRTY]-[INQR]-[FILMQS]-x(2)-[AGKNS]-[KMQRSTV]-x(2)-[AEGKMNSTY] (SEQ ID NO: 196); “ADEHKNQRS” is disclosed as SEQ ID NO: 197, “ADEFGHKMNQRSWY” is disclosed as SEQ ID NO: 198, “FHLWY” is disclosed as SEQ ID NO: 199, “ADEFIKLMNQRSTY” is disclosed as SEQ ID NO: 200, “FIQSTV” is disclosed as SEQ ID NO: 201, “AGKLNRSTV” is disclosed as SEQ ID NO: 202, “ADEHKNQRTY” is disclosed as SEQ ID NO: 203, “INQR” is disclosed as SEQ ID NO: 204, “FILMQS” is disclosed as SEQ ID NO: 205, “AGKNS” is disclosed as SEQ ID NO: 206, “KMQRSTV” is disclosed as SEQ ID NO: 207, and “AEGKMNSTY” is disclosed as SEQ ID NO: 208.
[0134] Thus, in certain aspects, described here is an engineered Dn29 comprising the amino acid sequence of the above motif and an amino acid sequence having 70% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In some embodiments, the amino acid sequence has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In certain aspects, described here is a nucleic acid encoding an engineered Dn29 comprising the amino acid sequence of the above motif and an amino acid sequence having 70% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In some embodiments, the amino acid sequence has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Dn29 (SEQ ID NOs: 1) and comprising one or more mutations provided in Tables 1 and 2. In certain aspects, the engineered Dn29 sequences comprising the above motif and described herein can be used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).Engineered Dn29 Mutations
[0135] Various methods for generating genetic diversity in protein evolution are known in the art. For example, point mutagenesis, combinatorial cassette mutagenesis, DNA shuffling or substrate-linked protein evolution can be employed to generate genetic diversity for directed protein evolution. Alteration of Cre recombinase site specificity by substrate- linked protein evolution, Nat Biotechnol., 19, 1047-1052 (2001), the entire content is hereby incorporated by reference in its entirety. In some embodiments, substrate-linked protein evolution is employed to generate engineered Dn29. See Example 1.
[0136] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation at one or more amino acids selected from E70, F138,A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, K288 or any combination thereof. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0137] In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and one or more mutations selected from E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, K288E, or any combination thereof. In some embodiments, the mutation is E70G. In some embodiments, the mutation is F138L. In some embodiments, the mutation is A224P. In some embodiments, the mutation is N341Q. In some embodiments, the mutation is N341 V. In some embodiments, the mutation is L388P. In some embodiments, the mutation is N341I. In some embodiments, the mutation is N341H. In some embodiments, the mutation is V386A. In some embodiments, the mutation is Q390P. In some embodiments, the mutation is L393P. In some embodiments, the mutation is D503N. In some embodiments, the mutation is V202A. In some embodiments, the mutation is W373R. In some embodiments, the mutation is S322G. In some embodiments, the mutation is F374S. In some embodiments, the mutation is W373C. In some embodiments, the mutation is R389S. In some embodiments, the mutation is Y404C. In some embodiments, the mutation is L449P. In some embodiments, the mutation is N214D. In some embodiments, the mutation is I303K. In some embodiments, the mutation is Q332K. In some embodiments, the mutation is N341K. In some embodiments, the mutation is I233K. In some embodiments, the mutation is K248R. In some embodiments, the mutation is I98V. In some embodiments, the mutation is R246G. In some embodiments, the mutation is V150K. In some embodiments, the mutation is N341R. In some embodiments, the mutation is D437G. In some embodiments, the mutation is S228A. In some embodiments, the mutation is K288E. In some embodiments, the mutation is a combination of any mutation described herein. In some embodiments, the mutation is a combination of any mutation described herein. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0138] In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and at least two or more mutations, wherein at least one mutation is selected from the “efficiency”mutations of Table 1 (E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H) and at least one mutation selected is selected from the “specificity” mutations of Table 1 (I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, K288E). In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and at least two or more mutations, wherein at least one mutation is selected from the “efficiency” mutations of Table 1 (E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H) and at least one mutation selected is selected from the “specificity” mutations of Table 1 (I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, K288E). In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0139] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acid residue M6, E70, A224, or G227. In some embodiments, the mutations are M6I, E70G, A224P, or G227V. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codon encoding amino acid 234. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, and G227. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codon encoding amino acid 234. In some embodiments, the mutations are M6I, E70G, A224P, and G227V. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, and G227V and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silentmutations in the codon encoding amino acid 234. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises a mutation at amino acid residue N341. In some embodiments, the mutation is N341K. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codon encoding amino acid 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0140] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of N341. In some embodiments, the mutation is N341K. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codon encoding amino acid 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, VI 50, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0141] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acid residue M6, E70, A224, G227, Q332 or N341. In some embodiments, the mutations are M6I, E70G, A224P, G227V, Q332K, or N341K. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, Q332 and N341. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P, G227V, Q332K, and N341K. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150,D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q332K, and N341K and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0142] In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q332K, and N341K and further comprises a mutation of one or more amino acid residue 1233, K248, Q390, L388, L393, or D503. In some embodiments, the mutation is I233K, K248R, Q390P, L388P, L393P, or D503N. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q332K, and N341K and further comprises a mutation of one or more amino acid residue D503, L388, Q390, L393, W373, R389, E452, F138, or V386. In some embodiments, the mutation is D503N, L388P, Q390P, L393P, W373R, R389S, E452G, F138L, or V386A. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q332K, and N341K and further comprises a mutation of one or more amino acid residue 1233, K248, Q390, L388, L393, or D503 and a mutation of one or more amino acid residue D503, L388, Q390, L393, W373, R389, E452, F138, or V386. In some embodiments, the mutation is I233K, K248R, Q390P, L388P, L393P, D503N, W373R, R389S, E452G, F138L, or V386A. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q332K, and N341K and further comprises a mutation of one of amino acid residue 1233, K248, Q390, L388, L393, or D503 and a mutation of one of amino acid residue D503, L388, Q390, L393, W373, R389, E452, F138, or V386. In some embodiments, the mutation is I233K, K248R, Q390P, L388P, L393P, D503N, W373R, R389S, E452G, F138L, or V386A. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprisessilent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29- dCas9 fusion).
[0143] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acid residue M6, E70, A224, G227, or N341. In some embodiments, the mutations are M6I, E70G, A224P, G227V, or N341Q. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, and N341. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P, G227V, and N341Q. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, and N341Q and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0144] In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, and N341Q and further comprises a mutation of one or more amino acid residue W373, L388, L393, or D503. In some embodiments, the mutation is W373R, L388P, L393P, or D503N. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V,and N341Q and further comprises a mutation of one or more amino acid residue 1233, 1303, K248, or Q332. In some embodiments, the mutation is I233K, I303K, K248R, or Q332K. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, and N341Q and further comprises a mutation of one or more amino acid residue W373, L388, L393, or D503 and a mutation of one or more amino acid residue 1233, 1303, K248, or Q332. In some embodiments, the mutation is W373R, L388P, L393P, D503N, I233K, I303K, K248R, or Q332K. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, and N341Q and further comprises a mutation of one of amino acid residue W373, L388, L393, or D503 and a mutation of one, two, three or four of amino acid residue 1233, 1303, K248, or Q332. In some embodiments, the mutation is W373R, L388P, L393P, D503N, I233K, I303K, K248R, or Q332K. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0145] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acid residue M6, E70, A224, G227, 1303, N341, or L388. In some embodiments, the mutations are M6I, E70G, A224P, G227V, I303K, N341Q or L388P. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, 1303, N341, and L388. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P, G227V, I303K, N341Q and L388P. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, VI 50, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acidencoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, I303K, N341Q and L388P and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0146] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acid residue M6, E70, A224, G227, 1303, N341, or L393. In some embodiments, the mutations are M6I, E70G, A224P, G227V, I303K, N341Q, or L393P. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, 1303, N341, and L393. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P, G227V, I303K, N341Q, and L393P. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, VI 50, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, I303K, N341Q, and L393P and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0147] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acid residue M6, E70, A224, G227, Q322, N341, L393, or D503. In some embodiments, the mutations are M6I, E70G, A224P, G227V, Q322K, N341K, L393P, or D503N. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, Q322, N341, L393, and D503. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P, G227V, Q322K, N341K, L393P, and D503N. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, and D503N and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29- dCas9 fusion).
[0148] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acid residue M6, E70, A224, G227, Q322, N341, L388, or R389. In some embodiments, the mutations are M6I, E70G, A224P, G227V, Q322K, N341K, L388P, or R389S. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) andcomprises mutations of amino acid residues M6, E70, A224, G227, Q322, N341, L388, and R389. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P, G227V, Q322K, N341K, L388P, and R389S. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L388P, and R389S and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29- dCas9 fusion).
[0149] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q322, N341, L393, D503, 198, 1233, 1303, or K248. In some embodiments, the mutations are M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, or K248R. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, Q322, N341, L393, D503, 198, 1233, 1303, and K248. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, Q322, N341, L393, D503, and at least one mutation selected from 198, 1233, 1303, and K248. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P,G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R and further comprises one or more additional mutations of Table 1. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0150] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation of one or more amino acids selected from M6, E70, A224, G227, 1233, K248, 1303, Q332, N341, L388, or Y404. In some embodiments, the mutation comprises M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P, or Y404C. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations of amino acid residues M6, E70, A224, G227, 1233, K248, 1303, Q332, N341, L388 and Y404. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In some embodiments, the mutations are M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P and Y404C. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations of Table 1. In some embodiments, theengineered Dn29, or nucleic acid encoding an engineered Dn29, further comprises one or more additional mutations selected from D503N, W373R, L449P, N214D, I98V, K288E, V386A, Q390P, L393P, F138L, V202A, S228A, or S428I. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P and Y404C and further comprises one or more additional mutations of Table 1. In some embodiments, the engineered Dn29, or nucleic acid encoding an engineered Dn29, comprises the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises mutations M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P and Y404C and further comprises one or more additional mutations selected from D503N, W373R, L449P, N214D, I98V, K288E, V386A, Q390P, L393P, F138L, V202A, S228A, or S428I. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0151] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation at amino acids M6, E70, A224, G227, Q322, N341, L393, and D503. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, and D503N and silent mutations in the codons encoding amino acid 234 and 318. In some embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, consisting of mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, and D503N and silent mutations in the codons encoding amino acid 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29- dCas9 fusion).
[0152] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation at amino acids M6, E70, A224, G227, 1233, K248, 1303, Q322, N341, L388, and Y404. In some embodiments, the nucleic acid encoding theengineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising mutations M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q322K, N341Q, L388P, and Y404C and silent mutations at in the codons encoding amino acid 234 and 318. In some embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, consisting of mutations M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q322K, N341Q, L388P, and Y404C and silent mutations at the codons encoding amino acid 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0153] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising a mutation at amino acids M6, E70, A224, G227, Q322, N341, L393, D503, 198, 1233, 1303, and K248. In some embodiments, the nucleic acid encoding the engineered Dn29 further comprises silent mutations in the codons encoding amino acids 234 and / or 318. In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising an amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R and silent mutations at in the codons encoding amino acid 234 and 318. In some embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, consisting of mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R and silent mutations at the codons encoding amino acid 234 and 318. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion).
[0154] In certain embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the clones provided in Table 2. Thus, in some embodiments, the invention provides an engineered Dn29, or nucleic acid encoding an engineered Dn29, comprising the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprising the mutations described for each clone in Table 2. In certain aspects, the engineered Dn29 sequence is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29- dCas9 fusion).
[0155] In some embodiments, the engineered Dn29 disclosed herein can have enhanced efficiency in recombination activity compared to wild-type Dn29 (SEQ ID NO: 1). Efficiencyrefers to the percent of recombination events at a specific DNA sequence. In some embodiments, the efficiency is the percent of recombination events at attHl site of Dn29. In some embodiments, the mutations that confer enhanced efficiency are one or more of the “efficiency” mutations of Table 1. Where the engineered Dn29 is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion), in certain embodiments, the engineered Dn29-DBD fusion has enhanced efficiency in recombination activity compared to wild-type Dn29-DND fusion (e g., SEQ ID NOs: 3-7).
[0156] In certain embodiments, the engineered Dn29 disclosed herein can have enhanced specificity in recombination activity compared to wild-type Dn29 (SEQ ID NO: 1).Specificity refers to the ability of a LSR to target a specific DNA sequence. In some embodiments, the specificity is the specificity for recombination event at attHl site compared to attH3 site. In some embodiments, the specificity is the specificity for recombination event at attHl site compared to attH chrlO site. In some embodiments, the specificity is the specificity for recombination event at attHl site compared to all other pseudosites. In some embodiments, the mutations that confer enhanced efficiency is one or more of the “specificity” mutations of Table 1. Where the engineered Dn29 is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion), in certain embodiments, the engineered Dn29-DBD fusion has enhanced specificity in recombination activity compared to wild-type Dn29-DND fusion (e g., SEQ ID NOs: 3-7).
[0157] In certain embodiments, the engineered Dn29 disclosed herein can have enhanced efficiency and enhanced specificity in recombination activity compared to wild-type Dn29 (SEQ ID NO: 1). Where the engineered Dn29 is used as the Dn29 portion of a Dn29-DBD fusion (e.g. Dn29-dCas9 fusion), in certain embodiments, the engineered Dn29-DBD fusion has enhanced efficiency and specificity in recombination activity compared to wild-type Dn29-DND fusion (e g., SEQ ID NOs: 3-7).
[0158] In any of the embodiments described herein, a nucleotide sequence encoding the engineered Dn29 can be codon-optimized.LSR Donor and Acceptor Attachment Sites
[0159] Recombination sites for the engineered Dn29 described herein are typically between 30 and 200 nucleotides in length and comprising two motifs with a partial inverted- repeat symmetry, which flank a central crossover sequence at which the recombination takes place. The engineered Dn29 can bind to these inverted-repeated sequences, which are specific to Dn29, and may be referred to as “recombinase recognition sequences,” “recombinaserecognition sites,” “attP sites,” “attB sites,” “attD sites,” “attH sites,” “attA sites,” “attachment sites,” “pseudosites,” “genomic pesudosites,” or “genomic insertion sites”. In some embodiments, an attB site is present in the target DNA sequence (such as cellular DNA) and an attP site is present in the DNA sequence to be integrated into the target DNA sequence. In some embodiments, an attP site is present in the target DNA sequence (such as cellular DNA) and an attB site is present in the DNA sequence to be integrated into the target DNA sequence. As disclosed herein, “attD” refers to a donor attachment site, which could be an attP or an attB site, “attA” refers to the cognate acceptor site and “attH” refers to integration sites found natively in a mammalian genome, for example the human genome. A “landing pad,” is an exogenous DNA sequence that includes an attachment site of a LSR integrated into a location of the target DNA. A landing pad can be integrated into a target DNA using any method known in the art, such as by using a zinc finger nuclease, TALEN, or the CRISPR-Cas system, or by using an LSR-DBD fusion described herein.
[0160] During recombination, crossover occurs at the central dinucleotide of the attB / attP sites. The sequence of the central dinucleotide is the sole determinant of the directionality of the recombination. For the recombination to be directional, the central dinucleotide needs to be non-palindromic. For example, the central dinucleotide sequence found in the attB / attP sites for large serine recombinases, which are strictly directional, can be AA, TT, GG, CC, AG, GA, AC, CA, TG, GT, TC, or CT.
[0161] The outcome of recombination depends, in part, on the location and orientation of the attachment sites. For example, inversion recombination happens between two inverted attachment sites located on the same DNA molecule. A DNA loop formation brings the two attachment sites together, at which point DNA cleavage, strand exchange, and ligation occur. This reaction is ATP independent. The end result of such an inversion recombination event is that the stretch of DNA between the repeated site inverts (i.e., the stretch of DNA reverses orientation) such that what was the coding strand is now the non-coding strand and vice versa. Such reactions, the DNA is conserved with no net gain or no loss of DNA. Conversely, excisive recombination occurs between two attachment sites that are oriented in the same direction on the same DNA molecule. In this case, the intervening DNA is excised / removed. Integrative recombination can occur between two attachment sites that are located on different DNA molecules, where one of the DNA molecules is circular (for integration of the entire circular molecule). If the other DNA molecule is cellular or genomic DNA, the two molecules are combined into one molecule, with the circular DNA integrated into the cellularor genomic DNA. Finally, translocation occurs upon recombination of two attachment sites found on different, linear DNA molecules.
[0162] Engineered Dn29 has two attachment sites to which it binds and recombines sequence-specifically. In some embodiments, target DNA, with an introduced attachment site is targeted. In another embodiment, to target sequences endogenously present in a target DNA, a sequence similar to the desired attachment site sequence must be present in the target DNA, such as in a genome or other cellular DNA. Another factor that may be relevant is the number of endogenous sites that the engineered Dn29 can integrate into. Having fewer (but not 0) integration sites may increase efficiency of integration into a single pseudosite, since there will be fewer potential off-target sites which may act as a sink for LSRs thus reducing on-target efficiency. Thus, in some embodiments, an engineered Dn29 that has the ability to target a single or up to thousands of endogenous sequences can be used.LSR-DBD Fusions
[0163] In some embodiments, the engineered Dn29 may be fused to a DNA binding domain (DBD). In some embodiments, the engineered Dn29 portion is fused directly to the DBD portion. In some embodiments, the engineered Dn29-DBD fusion comprises a linker between the engineered Dn29 and DBD portions of the fusion protein. The use of “engineered Dn29-DBD” is intended to encompass both embodiments unless specified otherwise (i.e., in “engineered Dn29-DBD” indicates both a direct bond or a linker between the engineered Dn29 and DBD portions of the engineered Dn29-DBD fusion protein). The fusions can direct an engineered Dn29 to a specific target site via DNA binding domain fusions to increase efficiency and specificity of the engineered Dn29. Without being bound by theory these fusions will increase the local concentration of engineered Dn29 monomers at target DNA attachment sites, cause longer duration of engineered Dn29 residence at target DNA attachment sites, provide for improved target DNA scanning efficiency or kinetics and / or provide increased chromatin accessibility by dual protein- mediated binding to two sites.
[0164] The DBD, and where used, linker portions, for use in engineered Dn29-DBD fusions are described below. The engineered Dn29 portion is described above and throughout.DNA Binding Domains (DBDs)
[0165] As an RNA-guided nuclease, Cas proteins have been adapted for targeted gene editing and selection in a variety of organisms. Nuclease-null Cas variants that have nosubstantial nuclease activity are useful to localize proteins and RNA to nearly any set of dsDNA sequences.
[0166] In some embodiments, the DNA binding domain of the engineered Dn29-DBD fusion described herein comprises a modified form of a Cas protein, for example, without limitation, Cas9, Cpfl, Cas 12b, Cas 12c, Cas 12d, Casl2e, Casl2f, Cas 12g, Casl2h, Casl2i, Cas3, Cas8a-c, CaslO, Csel, Csyl, Csnl, Csn2, Cas4, Csm2, Cm5, Casl, Cas2, Cas7, C2c3, C2c2, C2cl, or Cas5, which forms a complex with a guide RNA. When the DBD is a Cas protein in complex with a guide RNA, the Cas protein can bind a target DNA via the guide RNA spacer sequence, which base pairs with a complementary target DNA sequence proximal to, overlapping with, or within the recombinase target site. In some instances, the modified form of the Cas protein comprises an amino acid change (e.g., deletion, insertion, or substitution) that reduces the naturally-occurring nuclease activity of the Cas protein. For example, in some instances, the modified form of the Cas protein has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas protein. In preferred embodiments, the modified form of the Cas protein has no substantial nuclease activity. When the DBD of the engineered Dn29-DBD fusion is a modified form of a Cas protein that has no substantial nuclease activity, it can be referred to as a “dead Cas” or “dCas”. In some embodiments, a Cas protein may have nickase activity. In some embodiments, the modified form of the Cas protein has no substantial nickase activity. In some embodiments, the modified form of the Cas protein has no substantial nickase activity and no substantial nuclease activity
[0167] A person of skill in the art recognizes that Cas proteins can be isolated from different bacterial species. In some embodiments, the DNA binding domain of the engineered Dn29-DBD fusion described herein comprises a Cas protein from Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Campylobacter jejuni, Streptococcus thermophilus, Lachnospiraceae bacterium, Acidaminococcus sp. , Alicyclobacillus acidiphilus, or Bacillus hisashii. In some embodiments, the DNA binding domain of the engineered Dn29-DBD fusion described herein comprises Cas9 from Streptococcus pyogenes or dCas9 form thereof. In some embodiments, the DNA binding domain of the engineered Dn29-DBD fusion described herein comprises Cas9 from Staphylococcus aureus or dCas9 form thereof.
[0168] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising the amino acid sequence of dCas9, Cas9, Cpfl, Casl2b, Casl2c, Casl2d, Casl2e,Casl2f, Casl2g, Casl2h, Casl2i, Cas3, Cas8a-c, CaslO, Csel, Csyl, Csnl, Csn2, Cas4, Csm2, Cm5, Casl, Cas2, Cas7, C2c3, C2c2, C2cl, or Cas5. In certain aspects, described herein is a nucleic acid encoding an engineered LSR-DBD fusion comprising the amino acid sequence of dCas9, Cas9, Cpfl, Casl 2b, Casl 2c, Casl 2d, Casl2e, Casl2f, Casl 2g, Casl2h, Casl2i, Cas3, Cas8a-c, CaslO, Csel, Csyl, Csnl, Csn2, Cas4, Csm2, Cm5, Casl, Cas2, Cas7, C2c3, C2c2, C2cl, or Cas5. In some embodiments, the DNA binding domain of the engineered LSR-DBD fusion described herein comprises Streptococcus pyogenes dCas9. In some embodiments, the DNA binding domain of the engineered LSR-DBD fusion described herein comprises Staphylococcus aureus dCas9.
[0169] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively). In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively). In some embodiments, the nucleic acid sequence encoding the DBD portion comprises SEQ ID NOs: 14-17.
[0170] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising an amino acid sequence having 70% identity to dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively). In some embodiments, the amino acid sequence has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively). In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising an amino acid sequence having 70% identity to dCas9, dCas9- HF1, dCas9-SpG, dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively). In some embodiments, the amino acid sequence has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG- HF1(SEQ ID NOs: 10-13, respectively).
[0171] In certain aspects, described herein is an engineered Dn29-DBD fusion, wherein the DBD portion consists of the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively). In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion, wherein the DBD portion consists of the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl (SEQ IDNOs: 10-13, respectively). In some embodiments, the nucleic acid sequence encoding the DBD portion consists of SEQ ID NOs: 14-17.
[0172] In certain aspects, described herein is an engineered Dn29-DBD fusion, wherein the DBD portion comprises DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion, wherein the DBD portion comprises DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site. In some embodiments, the DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site is dCas9,dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl, Cas9, Cpfl, Casl2b, Casl2c, Casl2d, Casl2e, Casl2f, Casl2g, Casl2h, Casl2i, Cas3, Cas8a-c, CaslO, Csel, Csyl, Csnl, Csn2, Cas4, Csm2, Cm5, Casl, Cas2, Cas7, C2c3, C2c2, C2cl, or Cas5.
[0173] In other embodiments, other DNA binding domains may be used (e.g., ZFPs or TALEs) that bind to a DNA target site proximal to, overlapping with, or within the recombinase target site. In some embodiments, the DNA binding domain binds to a DNA target nucleic acid sequence within 200 nucleotides upstream or downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest. In some embodiments, the DNA binding domain binds to a DNA target nucleic acid sequence within 100 nucleotides upstream or downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest. In some embodiments, the DNA binding domain binds to a DNA target nucleic acid sequence within 80 nucleotides upstream or downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest. In some embodiments, the DNA binding domain binds to a DNA target nucleic acid sequence within 50 nucleotides upstream or downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest. In some embodiments, one of the two or more domains is a zinc finger (ZF) or TALE DNA binding domain. A “zinc finger DNA binding protein” (or binding domain) is a protein, or a domain within a larger protein, that binds DNA in a sequence-specific manner through one or more zinc fingers, which are regions of amino acid sequence within the binding domain whose structure is stabilized through coordination of a zinc ion. The term zinc finger DNA binding protein is often abbreviated as zinc finger protein or ZFP. A “TALE DNA binding domain” or “TALE” is a polypeptide comprising one ormore TALE repeat domains / units. The repeat domains are involved in binding of the TALE to its cognate target DNA sequence. A single “repeat unit” (also referred to as a “repeat”) is typically 33-35 amino acids in length and exhibits at least some sequence homology with other TALE repeat sequences within a naturally occurring TALE protein. Each TALE repeat unit includes 1 or 2 DNA-binding residues making up the Repeat Variable Diresidue (RVD), typically at positions 12 and / or 13 of the repeat. Zinc finger and TALE binding domains can be “engineered” to bind to a predetermined nucleotide sequence, for example via engineering (altering one or more amino acids) of the recognition helix region of a naturally occurring zinc finger or TALE protein. Therefore, engineered DNA binding proteins (zinc fingers or TALEs) are proteins that are non-naturally occurring.Linkers
[0174] In some embodiments, the fusion between the engineered Dn29 and DBD protein may include a linker. The term “linker,” as used herein, refers to a chemical group or a molecule linking two molecules or moieties, e.g., engineered Dn29 and Cas protein. Typically, the linker is positioned between, or flanked by, two groups, molecules, or other moieties and connected to each one via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker may comprise a peptide or a non-peptide moiety. In some embodiments, the linker is 2-100 amino acids in length, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.
[0175] Exemplary linkers include, for example, flexible, GlySer linkers for use in the engineered Dn29-DBD fusions described herein. In some embodiments a “GGS” linker is used, which can be used in various repeats, for example in repeats of 1 (GGS), 2 ((GGS)?) (SEQ ID NO: 36), 3 ((GGS)3) (SEQ ID NO: 37), 4 ((GGS)4) (SEQ ID NO: 38), 5 ((GGS)s)(SEQ ID NO: 39), 6 ((GGS)6) (SEQ ID NO: 40), 7 ((GGS)7) (SEQ ID NO: 41), 8 ((GGS)s)(SEQ ID NO: 18), 9 ((GGS)9) (SEQ ID NO: 42), 10 ((GGS)io) (SEQ ID NO: 43), 11((GGS)n) (SEQ ID NO: 44), 12 ((GGS)n) (SEQ ID NO: 45), or more, to provide suitable lengths, as required. In some embodiments a “GGGS” linker (SEQ ID NO: 46) is used, which can be used in various repeats, for example in repeats of 1 (GGGS) (SEQ ID NO: 47), 2 ((GGGS)2) (SEQ ID NO: 48), 3 ((GGGS)3) (SEQ ID NO: 49), 4 ((GGGS)4) (SEQ ID NO:50), 5 ((GGGS)s) (SEQ ID NO: 51), 6 ((GGGS)6) (SEQ ID NO: 52), 7 ((GGGS)7) (SEQ ID NO: 53), 8 ((GGGS)s) (SEQ ID NO: 54), 9 ((GGGS)9) (SEQ ID NO: 55), 10 ((GGGS)io) (SEQ ID NO: 56), 11 ((GGGS)n) (SEQ ID NO: 57), 12 ((GGGS)n) (SEQ ID NO: 58), or more, to provide suitable lengths, as required. In some embodiments a “GGSS” linker (SEQ ID NO: 59) is used, which can be used in various repeats, for example in repeats of 1 (GGSS) (SEQ ID NO: 60), 2 ((GGSS)2) (SEQ ID NO: 61), 3 ((GGSS)3) (SEQ ID NO: 62), 4 ((GGSS)4) (SEQ ID NO: 63), 5 ((GGSS)s) (SEQ ID NO: 64), 6 ((GGSS)6) (SEQ ID NO: 65), 7 ((GGSS)7) (SEQ ID NO: 66), 8 ((GGSS)s) (SEQ ID NO: 67), 9 ((GGSS)9) (SEQ ID NO: 68), 10 ((GGSS)io) (SEQ ID NO: 69), 11 ((GGSS)n) (SEQ ID NO: 70), 12 ((GGSS)i2) (SEQ ID NO: 71), or more, to provide suitable lengths, as required. In some embodiments a “GGGGS” linker (SEQ ID NO: 72) is used, which can be used in various repeats, for example, they can be used in repeats of 3 ((GGGGS)3) (SEQ ID NO: 73), or 6 ((GGGGS)e) (SEQ ID NO: 74), 9 ((GGGGS)9) (SEQ ID NO: 75) or 12 ((GGGGS)I2) (SEQ ID NO: 76) or more, to provide suitable lengths, as required. Other alternatives are (GGGGS)i (SEQ ID NO: 77), (GGGGS)2(SEQ ID NO: 78), (GGGGS)4, (SEQ ID NO: 79) (GGGGS)s (SEQ ID NO: 80), (GGGGS)7(SEQ ID NO: 81), (GGGGS)s (SEQ ID NO: 82), (GGGGS)io (SEQ ID NO: 83), or (GGGGS)n (SEQ ID NO: 84). Additional glycine and / or serine residues can be included at the ends of the linker or between the various repeats, for example, S(GGGGS)eS (SEQ ID NO: 19).
[0176] In some embodiments, XTEN linkers are used in the engineered Dn29-DBD fusions described herein. For example, in some embodiments, XTEN16 (SGSETPGTSESATPESS (SEQ ID NO: 20)) is used. In some embodiments, XTEN32, or XTEN48, which have two and three repeats of XTEN16, respectively are used. In some embodiments, additional XTEN16 repeats can be used to provide suitable lengths, as required.
[0177] In some embodiments, an alpha-helical linker such as (Ala(GluAlaAlaAlaLys)Ala) (SEQ ID NO: 85) is also contemplated for use in the engineered Dn29-DBD fusions described herein.
[0178] In some embodiments, rigid linkers are contemplated, such as (EAAAK)3(SEQ ID NO: 86), (EAAAK)n(n=l-3) (SEQ ID NO: 87), A(EAAK)4(ALEA(EAAAK)4A (SEQ ID NO: 88), PAPAP (SEQ ID NO: 89), AEAAAKEAAAKA (SEQ ID NO: 90), (Ala-Pro)n (n=10-34) (SEQ ID NO: 91) for use in the engineered Dn29-DBD fusions described herein. In some embodiments, cleavable linkers are contemplated, such as, disulfide bonds,VSQTSKLTR|AETVFPDV (SEQ ID NO: 92), PLG|LWA (SEQ ID NO: 93), RVL|AEA (SEQ ID NO: 94), EDVVCQSMSY (SEQ ID NO: 95), GGIER|GS (SEQ ID NO: 96), TRHRQPR|GWE (SEQ ID NO: 97), AGNRVRR|SVG (SEQ ID NO: 98), RRRRRRR|R|R (SEQ ID NO: 99) for use in the engineered Dn29-DBD fusions described herein.
[0179] In some embodiments, 2A self-cleaving peptides are used in the engineered Dn29- DBD fusions described herein. These peptides share a core sequence motif of DXEXNPGP (SEQ ID NO: 100). In some embodiments, T2A linker (GSG)EGRGSLLTCGDVEENPGP(S) (SEQ ID NO: 101) is used. In some embodiments, P2A linker (GSG)ATNFSLLKQAGDVEENPGP(S) (SEQ ID NO: 102) is used. In some embodiments, E2A linker (GSG)QCTNYALLKLAGDVESNPGP(S) (SEQ ID NO: 103) is used. In some embodiments, F2A linker (GSG)VKQTLNFDLLKLAGDVESNPGP(S) (SEQ ID NO: 104) is used. The linkers can comprise optional “GSG” residues at the N-terminus and optional “S” residue at the C-terminus as indicated in parentheses.
[0180] In some embodiments, a linker for use in the engineered Dn29-DBD fusions described herein can comprise a combination of one or more of a GlySer linker, an XTEN linker, and / or a 2A self-cleaving peptides described above. Exemplary, non-limiting linkers for use in the Dn29-DBD fusions described herein are provided in Figure 34.
[0181] In other embodiments, the linker is at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 6 amino acids, at least 7 amino acids, at least 8 amino acids, at least 9 amino acids, at least 10 amino acids, at least 11 amino acids, at least 12 amino acids, at least 13 amino acids, at least 14 amino acids, at least 15 amino acids, at least 16 amino acids, at least 17 amino acids, at least 18 amino acids, at least 19 amino acids, at least 20 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 60 amino acids, at least 70 amino acids, at least 80 amino acids, at least 90 amino acids, at least 100 amino acids, at least 200 amino acids, at least 300 amino acids, at least 400 amino acids or at least 500 amino acids in length.
[0182] In some embodiments, the engineered Dn29 is fused directly to a DBD by a covalent bond. In certain embodiments, the covalent bond is a carbon-carbon bond, disulfide bond, carbon-heteroatom bond, a carbon-nitrogen bond of an amide linkage, etc. In certain embodiments, the engineered Dn29 is fused to a DBD by a linker that is a peptide or based on amino acids. In other embodiments, the linker is not peptide-like. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic or heteroaliphatic linker. In certain embodiments, the linker is polymeric (e.g., polyethylene,polyethylene glycol, polyamide, polyester, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminoalkanoic acid. In certain embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3- aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In certain embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In certain embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In certain embodiments, the linker comprises an aryl or heteroaryl moiety. In certain embodiments, the linker is based on a phenyl ring. The linker may include functionalized moieties to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates. In other embodiments, the linker comprises amino acids. In certain embodiments, the linker comprises a peptide.
[0183] In certain aspects, described herein is an engineered Dn29-DBD fusion wherein the engineered Dn29is fused directly to the DBD. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion wherein the engineered Dn29is fused directly to the DBD. In certain aspects, described herein is an engineered Dn29-DBD fusion comprising an engineered Dn29portion, DBD portion, fused together via a peptide linker. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising an engineered Dn29 portion, DBD portion, fused together via a peptide linker. In some embodiments, the peptide linker is 2 to 100 amino acids long. In some embodiments, the peptide linker is 2 to 50 amino acids long. In some embodiments, the peptide linker is 2 to 30 amino acids long. In some embodiments, the peptide linker comprises glycine and serine residues. In some embodiments, the peptide linker comprises only glycine and serine residues. In some embodiments, the peptide linker is 2 to 30 amino acids long and comprises only glycine and serine residues. In some embodiments, the peptide linker is 24 amino acids long and comprises only glycine and serine residues. In some embodiments, the peptide linker is 30 amino acids long and comprises only glycine and serine residues. In some embodiments, the peptide linker comprises GGS repeats. In some embodiments, the peptide linker comprises 2-12 GGS repeats (SEQ ID NO: 105). In some embodiments, the peptide linker consists of 2-12 GGS repeats (SEQ ID NO: 105). In some embodiments, the peptide linker comprises 8 GGS repeats (SEQ ID NO: 18). In some embodiments, the peptide linkerconsists of 8 GGS repeats (SEQ ID NO: 18). In some embodiments, the peptide linker comprises GGSS repeats (SEQ ID NO: 60). In some embodiments, the peptide linker comprises 2-12 GGSS repeats (SEQ ID NO: 106). In some embodiments, the peptide linker consists of 2-12 GGSS repeats (SEQ ID NO: 106). In some embodiments, the peptide linker comprises 2 GGSS repeats (SEQ ID NO: 61). In some embodiments, the peptide linker comprises GGGGS repeats (SEQ ID NO: 72). In some embodiments, the peptide linker comprises 2-12 GGGGS repeats (SEQ ID NO: 107). In some embodiments, the peptide linker consists of 2-12 GGGGS repeats (SEQ ID NO: 107). In some embodiments, the peptide linker comprises 6 GGGGS repeats (SEQ ID NO: 74). In some embodiments, the peptide linker consists of 6 GGGGS repeats (SEQ ID NO: 74). In some embodiments, the peptide linker comprises an XTEN16 sequence. In some embodiments, the peptide linker consists of an XTEN16 sequence. In some embodiments, the peptide linker comprises an XTEN32 sequence. In some embodiments, the peptide linker consists of an XTEN32 sequence. In some embodiments, the peptide linker comprises an XTEN48 sequence. In some embodiments, the peptide linker consists of an XTEN48 sequence. In some embodiments, the peptide linker comprises an F2A, E2A, P2A or T2A sequence. In some embodiments, the peptide linker consists of an F2A, E2A, P2A or T2A sequence. In some embodiments, the peptide linker comprises an XTEN16 sequence and one or more glycine or serine residues at the N- or C-terminus of the XTEN16 sequence. In some embodiments, the peptide linker comprises an XTEN32 sequence and one or more glycine or serine residues at the N- or C-terminus of the XTEN32 sequence. In some embodiments, the peptide linker comprises an XTEN48 sequence and one or more glycine or serine residues at the N- or C- terminus of the XTEN48 sequence. In some embodiments, the peptide linker comprises one or more XTEN16 sequences (e.g., XTEN16, XTEN32, XTEN48) and one or more GGSS (SEQ ID NO: 60), GGS, or GGGGS (SEQ ID NO: 72) repeats. In some embodiments, the peptide linker comprises one or more XTEN16 sequences (e.g., XTEN16, XTEN32, XTEN48) and one or more F2A, E2A, P2A or T2A sequence. In some embodiments, the peptide linker comprises one or more GGSS (SEQ ID NO: 59), GGS, or GGGGS (SEQ ID NO: 72) repeats and one or more F2A, E2A, P2A or T2A sequence. In some embodiments, the peptide linker comprises the amino acid sequence of SEQ ID NOs: 18-26. In some embodiments, the nucleic acid sequence encoding the peptide linker portion comprises SEQ ID NOs: 27-35.
[0184] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising a peptide linker means for fusing together the engineered Dn29portion and DBD portion. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29- DBD fusion comprising a peptide linker means for fusing together the engineered Dn29 portion and DBD portion.
[0185] In some embodiments, the fusion protein further comprises or consists essentially of or consists of a localization (nuclear import or export) signal as, or as part of, the linker between the DBD (e.g., Cas enzyme) portion and the engineered Dn29 portion. HA or Flag tags are also within the ambit of the invention as linkers. The linkers allow the user to engineer appropriate amounts of “mechanical flexibility”.
[0186] Contemplated herein are fusions oriented in either orientation. Thus, in some embodiments, the engineered Dn29 is fused to the C-terminus of a DBD. Alternatively, the engineered Dn29 is fused to the N-terminus of a DBD. In another instance, the engineered Dn29 is fused to a position other than the C-terminus or the N-terminus of a DBD, e.g., an internal residue of a DBD. Fusions oriented with the engineered Dn29 at the N-terminus, e.g., Dn29-linker-dCas9, are preferable to fusions oriented with the Dn29 at the C-terminus, e.g., dCas9-linker-Dn29. Thus, in certain aspects, described herein is an engineered Dn29- DBD fusion wherein the engineered Dn29 portion is N-terminal to the DBD portion. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion wherein the engineered Dn29 portion is N-terminal to the DBD portion.
[0187] Longer linkers are preferable as well, for example Dn29-XTEN32-(GGSS)2- XTEN-dCas9 is preferable to Dn29-XTEN16-dCas9. Dn29-(GGGGS)e-dCas9 is preferable to Dn29-(GGS)8-dCas9. “(GGSS)2”, “(GGGGS)6” and “(GGS)s” are disclosed as SEQ ID NOS 61, 74 and 18, respectively. Linker flexibility is also a factor, as more flexible linkers (GGS and GGGGS (SEQ ID NO: 72)) are preferable than more rigid linkers (XTEN16) in the dCas9-linker-Dn29 fusions.
[0188] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising any of the engineered Dn29 and DBD portions described herein. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising any of the engineered Dn29 and DBD portions described herein. In some embodiments, the engineered Dn29 portion comprises: (a) the amino acid sequence of Dn29 (SEQ ID NO: 1 comprising one or more of the mutations described herein), (b) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%,or 99% identity to the amino acid sequence of (a) and comprising one or more of the mutations described herein, (c) an amino acid sequence of Motif 5 and comprising one or more of the mutations described herein, (d) an amino acid sequence of Motif 5 and an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of (a) comprising one or more of the mutations described herein, or (e) engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences; and the DBD portion comprises: (f) an amino acid sequence of Cas9, Cpfl, Cast 2b, Cast 2c, Cast 2d, Casl2e, Casl2f, Cast 2g, Casl2h, Casl2i, Cas3, Cas8a-c, CaslO, Csel, Csyl, Csnl, Csn2, Cas4, Csm2, Cm5, Cast, Cas2, Cas7, C2c3, C2c2, C2cl, or Cas5, (g) the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl (SEQ ID NOs: 10-13), (h) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of (g), or (i) DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site.
[0189] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising an engineered Dn29 portion comprising the amino acid sequence of Dn29 (SEQ ID NO: 1 comprising one or more of the mutations described herein) and a DBD portion comprising the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, or dCas9-SpG-HFl (SEQ ID NOs: 10-13). In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising an engineered Dn29 portion comprising the amino acid sequence of Dn29 (SEQ ID NO: 1 comprising one or more of the mutations described herein) and a DBD portion comprising the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, or dCas9-SpG-HFl (SEQ ID NOs: 10-13). In some embodiments, the engineered Dn29 portion comprises SEQ ID NO: 1 comprising one or more of the mutations described herein and the DBD portion comprises dCas9 (SEQ ID NO: 10). In some embodiments, the amino acid sequence of the engineered Dn29 portion has 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of Dn29 (SEQ ID NO:1 comprising one or more of the mutations described herein). In some embodiments, the amino acid sequence of the DBD portion has 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, or dCas9-SpG-HFl (SEQ ID NOs: 10-13).
[0190] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising an engineered Dn29 portion comprising engineered Dn29 means for mediatingrecombination of DNA between recombinase recognition sequences and a DBD portion comprising DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising an engineered Dn29 portion comprising engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences and a DBD portion comprising DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site.
[0191] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising any of the engineered Dn29, DBD, and linker portions described herein. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising any of the engineered Dn29, DBD, and linker portions described herein. In some embodiments, the engineered Dn29 portion comprises a) the amino acid sequence of Dn29 (SEQ ID NO: 1 comprising one or more of the mutations described herein), b) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of a) and comprising one or more of the mutations described herein, c) an amino acid sequence of Motif 5 and comprising one or more of the mutations described herein, d) an amino acid sequence of Motif 5 and an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of a) and comprising one or more of the mutations described herein, or e) engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences; the DBD portion comprises f) an amino acid sequence of Cas9, Cpfl, Cast 2b, Cast 2c, Cast 2d, Casl2e, Casl2f, Cast 2g, Casl2h, Casl2i, Cas3, Cas8a-c, CaslO, Csel, Csyl, Csnl, Csn2, Cas4, Csm2, Cm5, Casl, Cas2, Cas7, C2c3, C2c2, C2cl, or Cas5, g) the amino acid sequence of dCas9, dCas9-HFl, dCas9- SpG, or dCas9-SpG-HFl (SEQ ID NOs: 10-13, repectively), h) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of g), or i) DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site; and the linker portion comprising j) a peptide linker, k) a peptide linker comprising one or more glycineserine repeats, 1) a peptide linker comprising one or more XTEN linkers, m) a peptide linker comprising one or more glycine- serine repeats and one or more XTEN linkers, n) an amino acid comprising SEQ ID NOs: 18-26, o) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acidsequence of k), or p) peptide linker means for fusing together the engineered Dn29 portion and DBD portion.
[0192] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising an engineered Dn29 portion comprising the amino acid sequence of engineered Dn29 (SEQ ID NO: 1 comprising one or more of the mutations described herein), a DBD portion comprising the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, or dCas9- SpG-HFl (SEQ ID NOs: 10-13, respectively), and a linker portion comprising the amino acid sequence of SEQ ID NOs: 18-26. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising an engineered Dn29 portion comprising the amino acid sequence of engineered Dn29 (SEQ ID NO: 1 comprising one or more of the mutations described herein) and a DBD portion comprising the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, or dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively), and a linker portion comprising the amino acid sequence of SEQ ID NOs: 18- 26. In some embodiments, the engineered Dn29 portion comprises engineered Dn29 (SEQ ID NO: 1 comprising one or more of the mutations described herein), the DBD portion comprises dCas9 (SEQ ID NO: 10), and the linker portion comprises the amino acid sequence of SEQ ID NOs: 18-26. In some embodiments, the amino acid sequence of the engineered Dn29 portion has 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of Dn29 (SEQ ID NO: 1) and comprises one or more of the mutations described herein). In some embodiments, the amino acid sequence of the DBD portion has 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, or dCas9-SpG-HFl (SEQ ID NOs: 10-13, respectively). In some embodiments, the amino acid sequence of the linker portion has 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of SEQ ID NOs: 18-26.
[0193] In certain aspects, described herein is an engineered Dn29-DBD fusion comprising an engineered Dn29 portion comprising engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences, a DBD portion comprising DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site, and peptide linker means for fusing together the engineered Dn29 portion and DBD portion. In certain aspects, described herein is a nucleic acid encoding an engineered Dn29-DBD fusion comprising an engineered Dn29 portioncomprising engineered Dn29 means for mediating recombination of DNA between recombinase recognition sequences, a DBD portion comprising DBD means for binding a target DNA sequence proximal to, overlapping with, or within the recombinase target site, and peptide linker means for fusing together the engineered Dn29 portion and DBD portion.
[0194] In any of the embodiments described herein, a nucleotide sequence encoding the engineered Dn29-DBD fusion polypeptide or the engineered Dn29, DBD, and / or linker portions thereof can be codon-optimized. This type of optimization is known in the art and entails the mutation of foreign-derived DNA to mimic the codon preferences of the intended host organism or cell while encoding the same protein. Thus, the codons are changed, but the encoded protein remains unchanged. For example, if the intended target cell was a human cell, a human codon-optimized Cas protein (or variant, e.g., dCas) would be a suitable DBD. Any suitable DBD can be codon optimized. As another non-limiting example, if the intended host cell were a mouse cell, then a mouse codon-optimized Cas protein (or variant, e.g., dCas) would be a suitable DBD. While codon optimization is not required, it is acceptable and may be preferable in certain cases.
[0195] Protein-mediated recruitment refers to the fusion of the DBD and engineered Dn29 to two interacting protein domains that can allow trans expression of each protein and subsequent recruitment to create the fusion. Some example systems include, but are not limited to, SunTag (a protein scaffold containing peptide epitopes fused to the dCas9 protein). The engineered Dn29 can be fused to single-chain variable fragment (scFV) antibodies, which when delivered in trans, are recruited to the peptide epitopes), SpyTag (a 13 residue peptide called Spytag and a 116 residue complementary domain) are fused to the DBD and engineered Dn29 respectively, which when delivered in trans, spontaneously assemble creating a covalent isopeptide bond), coiled-coil peptide heterodimers, or SnoopTag and SnoopCatcher can also be used.
[0196] Inducible recruitment refers to a DBD and an engineered Dn29 fused to inducible binding proteins, whereupon stimulus such as small molecules or light, cause dimerization, recruiting the engineered Dn29 to the DBD (e.g., dCas9). Examples of the system include FK506 binding protein 12 (FKBP) and FKBP rapamycin binding (FRB) domains, that dimerize upon rapamycin induction, pMag and nMag, which dimerize upon exposure to blue light, and DmrA / DmrC, which dimerize in the presence of rapamycin analog known as the A / C heterodimerizer.Guide polynucleotides
[0197] When a Cas protein domain is used as the DBD portion of the engineered Dn29- DBD fusion, the Cas portion is capable of binding one or more guide RNAs (gRNAs), in which the spacer sequences directs or targets the engineered Dn29-DBD fusion to a target nucleic acid of interest. In some embodiments, a guide RNA is used that targets a target sequence present on an acceptor target DNA of interest. In some embodiments, a guide RNA is used that targets a target sequence present on a donor DNA of interest. In some embodiments, the system described herein uses two guide RNAs, one that targets a target sequence present on an acceptor target DNA of interest and a second that targets a target sequence present on a donor DNA of interest. In some embodiments, the system described herein uses two guide RNAs, one that targets a target sequence present on an acceptor target DNA of interest and a second that targets a second target sequence present on the acceptor target DNA of interest. In some embodiments, the first and second target sequences on the acceptor target DNA of interest are on either side of the LSR attachment site in the target DNA of interest. In some embodiments, a guide RNA is used that targets a target sequence present on an acceptor target DNA of interest and a target sequence present on a donor DNA of interest, wherein the target sequences are the same. For example, the target sequence targeted by the guide in the acceptor target DNA of interest is included on the donor DNA molecule proximal to, overlapping with, or within the attD site. In some embodiments, more than two guide RNA sequences are used, for example one or more guide RNA sequences that target(s) one or more target sequences present on a donor DNA molecule of interest and one or more guide RNA sequences that target(s) one or more target sequences present on an acceptor target DNA of interest.
[0198] As used herein, the term “guide polynucleotide” or “guide RNA” or “gRNA”, relates to a polynucleotide sequence that can form a complex with a Cas protein and enables the Cas protein to recognize, bind to, and optionally cleave a DNA target site. The guide RNA is a specific RNA sequence that recognizes a target DNA region of interest and directs the Cas protein, and thus the engineered Dn29-DBD fusion, to that site. The gRNA is typically made up of two parts: CRISPR RNA (crRNA) (also referred to as a gRNA spacer or spacer sequence), a nucleotide sequence that binds to a complement of a target DNA sequence, and a trans-activating CRISPR RNA (tracr RNA), which serves as a binding scaffold for the Cas protein. In the context of CRISPR, hybridization between the complementary sequence of a target sequence and a gRNA spacer sequence promotes theformation of a CRISPR complex. Full complementarity is not necessarily required, provided there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex.
[0199] While crRNAs and tracrRNAs exist as two separate RNA molecules in nature, a single RNA molecule can contain both the crRNA sequence fused to the scaffold tracrRNA sequence, referred to as a single guide RNA (sgRNA). In some embodiments, the gRNA is a sgRNA. In some embodiments, the gRNA comprises two separate RNA molecules. The guide polynucleotide sequence can be a RNA sequence, a DNA sequence, or a combination thereof (a RNA-DNA combination sequence), such as the Caribou Biosciences system that uses a “chRDNA” system where the guide polynucleotide is a hybrid RNA / DNA system. Optionally, the guide polynucleotide can comprise at least one nucleotide, phosphodiester bond or linkage modification such as, but not limited, to Locked Nucleic Acid (LNA), 5- methyl dC, 2,6-Diaminopurine, 2'-Fluoro A, 2'-Fluoro U, 2'-O-Methyl RNA, phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 (hexaethylene glycol chain) molecule, or 5' to 3' covalent linkage resulting in circularization. See also U.S. Patent Application US 2015-0082478 Al, published on Mar. 19, 2015 and US 2015-0059010 Al, published on Feb. 26, 2015, both are hereby incorporated in its entirety by reference.
[0200] In one embodiment of the disclosure, the guide polynucleotide is a sgRNA capable of forming a guide RNA / protein RNP complex with the DBD of the engineered Dn29-DBD fusions disclosed herein, wherein said RNP complex can recognize and bind to a complement of a target sequence. One or more target sequences may be present in the acceptor target DNA of interest, the donor DNA of interest, or both.
[0201] In one embodiment of the disclosure, the guide polynucleotide is a sgRNA capable of forming a guide RNA / protein RNP complex with the DBD of the engineered Dn29-DBD fusions disclosed herein, wherein said complex can recognize and bind to a complement of a target sequence, wherein said sgRNA comprises a “crRNA” or “spacer” or “spacer sequence” linked to a “scaffold” or “scaffold sequence” or “tracrRNA.” One or more target sequences may be present in the acceptor target DNA of interest, the donor DNA of interest, or both.
[0202] In one embodiment of the disclosure, the guide polynucleotide is a gRNA capable of forming a guide RNA / protein RNP complex with the DBD of the engineered Dn29-DBD fusions disclosed herein, wherein said complex can recognize and bind to a complement of atarget sequence, wherein said guide RNA is a duplex molecule comprising a spacer and a scaffold, wherein said spacer comprises a sequence capable of hybridizing to a complement of a target DNA sequence. One or more target sequences may be present in the acceptor target DNA of interest, the donor DNA of interest, or both.
[0203] The guide polynucleotide can be a double molecule (also referred to as duplex guide polynucleotide) comprising a spacer sequence and a scaffold sequence. The spacer includes a first nucleotide sequence domain that can hybridize to a nucleotide sequence in a target DNA (i.e., to a nucleotide sequence complementary to a target sequence) and a second nucleotide sequence (also referred to as a “tracr mate” sequence) that is part of a Cas protein recognition (CPR) domain. The tracr mate sequence can be hybridized to a scaffold along a region of complementarity and together form a Cas protein recognition domain or CPR domain. The CPR domain is capable of interacting with a Cas protein. The spacer and the scaffold of the duplex guide polynucleotide can be RNA, DNA, and / or RNA-DNA- combination sequences. In some embodiments, the spacer molecule of the duplex guide polynucleotide is referred to as “spacer DNA” or “crDNA” (when composed of a contiguous stretch of DNA nucleotides) or “spacer RNA” or “crRNA” (when composed of a contiguous stretch of RNA nucleotides), or “spacer DNA-RNA” or “crDNA-RNA” (when composed of a combination of DNA and RNA nucleotides). The size of the fragment of the spacer naturally occurring in Bacteria and Archaea that can be present in a spacer disclosed herein can range from, but is not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 or more nucleotides. In some embodiments the scaffold is referred to as “scaffold RNA” or “tracrRNA” (when composed of a contiguous stretch of RNA nucleotides) or “scaffold DNA” or “tracrDNA” (when composed of a contiguous stretch of DNA nucleotides) or “scaffold DNA-RNA” or “tracrDNA-RNA” (when composed of a combination of DNA and RNA nucleotides. In one embodiment, the RNA that guides the RNA / Cas9 RNP complex of the LSR-DBD fusion is a duplexed RNA comprising a duplex spacer-scaffold. The scaffold or tracrRNA contains, in the 5 '-to-3 ' direction, (i) a sequence that anneals with the repeat region of CRISPR type II crRNA and (ii) a stem loop-containing portion (Deltcheva et al., Nature 471 :602-607). The duplex guide polynucleotide can form a complex with a Cas protein portion of the LSR-DBD fusion, wherein said guide polynucleotide / Cas RNP complex (also referred to as a guide polynucleotide / Cas RNP system) can direct the DBD of the engineered Dn29-DBD fusion proteins described herein to a target site, enabling the DBD protein to recognize and bind to the target site. See also U.S. Patent Application US 2015-0082478 Al, published on Mar. 19, 2015 and US 2015-0059010 Al, published on Feb. 26, 2015, both are hereby incorporated in its entirety by reference. In some embodiments, the spacer sequence is fused to the 5’ end of the scaffold sequence. Alternatively, the spacer sequence is fused to the 3’ end of the scaffold sequence.
[0204] The guide polynucleotide can also be a single molecule (also referred to as single guide polynucleotide) comprising a spacer sequence linked to a scaffold sequence. The single guide polynucleotide comprises a first nucleotide sequence domain that can hybridize to a nucleotide sequence in a target DNA (i.e., to a nucleotide sequence complementary to a target sequence) and comprises a Cas protein recognition domain (CPR domain), that interacts with a Cas protein. By “domain” as used in this context it is meant a contiguous stretch of nucleotides that can be RNA, DNA, and / or RNA-DNA-combination sequence. The spacer domain and / or the CPR domain of a single guide polynucleotide can comprise a RNA sequence, a DNA sequence, or a RNA-DNA-combination sequence. The single guide polynucleotide being comprised of sequences from the spacer and the scaffold may be referred to as “single guide RNA” (when composed of a contiguous stretch of RNA nucleotides) or “single guide DNA” (when composed of a contiguous stretch of DNA nucleotides) or “single guide RNA-DNA” (when composed of a combination of RNA and DNA nucleotides). The single guide polynucleotide can form a complex with a Cas protein portion of the engineered Dn29-DBD fusion, wherein said guide polynucleotide / Cas RNP complex (also referred to as a guide polynucleotide / Cas RNP system) can direct the DBD of the engineered Dn29-DBD fusion proteins described herein to a target site, enabling the DBD to recognize and bind to the target site. See also U.S. Patent Application US 2015-0082478 Al, published on Mar. 19, 2015 and US 2015-0059010 Al, published on Feb. 26, 2015, both are hereby incorporated in its entirety by reference.
[0205] In some embodiments, the gRNA comprises a sgRNA comprising a spacer RNA sequence portion and a tracr RNA portion, wherein the nucleic acid sequence of the spacer RNA sequence portion is the same as a target sequence on a DNA target of interest, and thus is complementary to, and hybridizes with the complement of the target sequence on the DNA target of interest. One or more target sequences may be present in the acceptor target DNA of interest, the donor DNA of interest, or both.
[0206] In some embodiments, immediately 3’ to the target sequence on the DNA target of interest is a protospacer adjacent motif (“PAM”) sequence. The PAM is a short DNA sequence (usually 2-6 base pairs in length) that, in a CRISPR-Cas9 system, follows the DNAregion targeted for cleavage by the CRISPR system. In some embodiments, the DBD portion of the engineered Dn29-DBD fusion comprises Streptococcus pyogenes dCas9 which recognizes the PAM sequence 5'-NGG-3' (where “N” can be any nucleotide base). Thus, in some embodiments, the DNA target of interest comprises a nucleotide sequence that is the same as the spacer sequence of the guide polynucleotide immediately followed in the 3’ direction by “NGG”. A person of skill in the art recognizes that there are different Cas endonucleases isolated from different bacterial species, each of which recognizes a different PAM. In some embodiments, the DBD portion of the engineered Dn29-DBD fusion comprises Staphylococcus aureus dCas9 which recognizes the PAM sequence 5'-NGRRT-3' or 5’ - or NGRRN-3’ (where “N” can be any nucleotide base). In some embodiments, the DBD portion of the engineered Dn29-DBD fusion comprises Neisseria meningitidis dCas9 which recognizes the PAM sequence 5'-NNNNGATT-3' (where “N” can be any nucleotide base). In some embodiments, the DBD portion of the engineered Dn29-DBD fusion comprises Campylobacter jejuni dCas9 which recognizes the PAM sequence 5'- NNNNRYAC-3' (where “N” can be any nucleotide base). In some embodiments, the DBD portion of the engineered Dn29-DBD fusion comprises Streptococcus thermophilus dCas9 which recognizes the PAM sequence 5'-NNAGAAW-3' (where “N” can be any nucleotide base). Cas9 mutants that have altered specificity, relaxed PAM requirements, or recognize novel PAM sequences can also be used as a DBD portion of the engineered Dn29-DBD fusion. In some embodiments, the DBD portion of the engeineered Dn29-DBD fusion comprises dCas9-SpG which recognizes the PAM sequence 5'-NGN-3' (where “N” can be any nucleotide base).
[0207] In some embodiments, the guide polynucleotide comprises a spacer sequence portion, wherein the nucleic acid sequence of the spacer sequence portion is the same as a target sequence on a target or donor DNA of interest (except in RNA spacer sequences “T” is “U”), wherein the target sequence is proximal to, overlapping with, or within the attachment site (e.g., attA or attD) of the engineered Dn29 on a target DNA of interest. In some embodiments, the target sequence on a target or donor DNA of interest is within 300 nucleotides upstream or downstream of an attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest, wherein distance is measured from the center of the dinucleotide core of the attachment site to the position between the spacer sequence and the PAM. In some embodiments, the target sequence on a target or donor DNA of interest within 200 nucleotides upstream ordownstream of an attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. In some embodiments, the target sequence on a target or donor DNA of interest is within 100 nucleotides upstream or downstream of an attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. In some embodiments, the target sequence on a target or donor DNA of interest is within 80 nucleotides upstream or downstream of an attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. In some embodiments, the target sequence on a target or donor DNA of interest is within 50 nucleotides upstream or downstream of an attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. A target sequence can be on either strand of target or donor DNA of interest. In some embodiments, the guide polynucleotide is a sgRNA. Generally, spacers that are directly proximal to the target integration attachment site, e.g., attH, have the highest integration rates, the spacers farther away have reduced integration, and spacers that overlap with the dinucleotide core of an attachment site greatly reduce or fully ablate integration.
[0208] In certain aspects, described herein is a nucleic acid encoding a guide polynucleotide for use with the engineered Dn29-DBD fusions described herein. The guide polynucleotide may be encoded on the same nucleic acid molecule as the engineered Dn29- DBD fusion and / or as a donor polynucleotide, or may be encoded on a separate nucleic acid molecule. In some embodiments, the guide polynucleotide is a gRNA comprising a spacer sequence portion and a tracr RNA portion. In some embodiments, the guide polynucleotide is a sgRNA comprising a spacer sequence portion and a tracr RNA portion. In some embodiments, the spacer sequence portion is about 20 nucleotides in length. In some embodiments, the spacer sequence portion is 16 nucleotides in length. In some embodiments, the spacer sequence portion is 20 nucleotides in length. In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a target or donor DNA of interest, wherein the target sequence is proximal to, overlapping with, or within the attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29- DBD fusion on a target or donor DNA of interest. In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a target or donor DNA of interest, wherein the target sequence is within 300 nucleotides of the attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBDfusion on a target or donor DNA of interest. In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a target or donor DNA of interest, wherein the target sequence is within 200 nucleotides of the attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a target or donor DNA of interest, wherein the target sequence is within 100 nucleotides of the attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a target or donor DNA of interest, wherein the target sequence is within 80 nucleotides of the attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. In some embodiments, the DNA sequence immediately 3’ to the target sequence on a target or donor DNA of interest comprises a PAM sequence. In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a target or donor DNA of interest, wherein the target sequence is within 50 nucleotides of the attachment site (e.g., attA or attD) of the engineered Dn29 of the engineered Dn29-DBD fusion on a target or donor DNA of interest. In some embodiments, the DNA sequence immediately 3’ to the targl06et sequence on a target or donor DNA of interest comprises a PAM sequence. In some embodiments, the DNA sequence immediately 3’ to the target sequence on a target or donor DNA of interest comprises a PAM sequence NGG. In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a target DNA of interest (e.g., proximal to, overlapping with, or within an attA site). In some embodiments, the spacer sequence portion comprises the same nucleotide sequence as a target sequence on a donor DNA of interest (e.g., proximal to, overlapping with, or within an attD site). In some embodiments, the spacer sequence portion of the gRNA or sgRNA comprises a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195). In some embodiments, the spacer sequence portion of the gRNA or sgRNA comprises a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195) with an additional “G” nucleotide present on the 5’ end. In some embodiments, the spacer sequence portion of the gRNA or sgRNA consists of a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195). In some embodiments, the spacer sequence portion of the gRNA or sgRNA consists of a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195) with an additional “G”nucleotide present on the 5’ end. In some embodiments, the tracr RNA portion of the gRNA or sgRNA comprises SEQ ID NO: 108. In some embodiments, the tracr RNA portion of the gRNA or sgRNA consists of SEQ ID NO: 108. In some embodiments, the spacer sequence portion of the gRNA or sgRNA comprises a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195) and the tracr RNA portion of the gRNA or sgRNA comprises SEQ ID NO: 108. In some embodiments, the spacer sequence portion of the gRNA or sgRNA comprises a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195) with an additional “G” nucleotide present on the 5’ end and the tracr RNA portion of the gRNA or sgRNA comprises SEQ ID NO: 108. In some embodiments, the spacer sequence portion of the gRNA or sgRNA consists of a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195) and the tracr RNA portion of the gRNA or sgRNA consists of SEQ ID NO: 108. In some embodiments, the spacer sequence portion of the gRNA or sgRNA consists of a nucleotide sequence selected from Figure 34 (SEQ ID NOs: 153-195) with an additional “G” nucleotide present on the 5’ end and the tracr RNA portion of the gRNA or sgRNA consists of SEQ ID NO: 108. In some embodiments the gRNA or sgRNA comprises SEQ ID NOs: 153-195 immediately followed by SEQ ID NO: 108. In some embodiments the gRNA or sgRNA comprises SEQ ID NOs: 153-105 with an additional “G” nucleotide present on the 5’ end immediately followed by SEQ ID NO: 108. In some embodiments the gRNA or sgRNA consists of SEQ ID NOs: 153-195 immediately followed by SEQ ID NO: 108. In some embodiments the gRNA or sgRNA consists of SEQ ID NOs: 153-195 with an additional “G” nucleotide present on the 5’ end immediately followed by SEQ ID NO: 108.Donor DNAs
[0209] Certain aspects of the present application are directed to a nucleic acid for use in site-specific insertion of an exogenous nucleic acid, e.g., a gene of interest (GOI), into a target DNA, e.g., a genome. In some embodiments, the exogenous nucleic acid for insertion (e.g., the GOI) can be up to about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 140, 150, 160, 170, 180, 190, 200, or 250 kilobases or higher in length. The GOI can include non-coding sequences, including cis regulatory regions and introns.
[0210] The donor DNA can contain from 15 bases (b) or base pairs (bp) to about 250 kilobases (kb) or kilobase pairs (kbp) in length (e.g., from about 50, 75, or 100 b or bp to about 110, 120, 125, 150, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1250, 1500, 1750, 2000, 2250,2500, 2750, 3000, 3250, 3500, 3750, 4000, 4250, 4500, 4750, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000 (and increasing by 1,000 increments) up to 250,000 b or bp in length. Longer donor DNA molecules can be provided in the form of a circular or linearized plasmid or as a component of a vector (e.g., as a component of a viral vector, such as a single-stranded AAV (ssAAV) or self-complementary AAV (scAAV)), or an amplification or polymerization product thereof. Shorter donor DNA molecules can be provided as double stranded oligonucleotides. Exemplary double-stranded template oligonucleotides are, or are least about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26,27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51,52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76,77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100,110, 115, 120, 125, 150, 175, 200, 225, or 250 b or bp in length. The donor DNA can be provided in the reaction mixture for introduction into the cell at a concentration of from about 1 pM to about 200 pM, from about 2 pM to about 190 pM, from about 2 pM to about 180 pM, from about 5 pM to about 180 pM, from about 9 pM to about 180 pM, from about 10 pM to about 150 pM, from about 20 pM to about 140 pM, from about 30 pM to about 130 pM, from about 40 pM to about 120 pM, or from about 45 or 50 pM to about 90 or 100 pM. In some cases, the donor DNA can be provided in the reaction mixture for introduction into the cell at a concentration of, or of about, 1 pM, 2 pM, 3 pM, 4 pM, 5 pM, 6 pM, 7 pM, 8 pM, 9 pM, 10 pM, 11 pM, 12 pM, 13 pM, 14 pM, 15 pM, 16 pM, 17 pM, 18 pM, 19 pM, 20 pM, 25 pM, 30 pM, 35 pM, 40 pM, 45 pM, 50 pM, 55 pM, 60 pM, 70 pM, 80 pM, 90 pM, 100 pM, 110 pM, 115 pM, 120 pM, 130 pM, 140 pM, 150 pM, 160 pM, 170 pM, 180 pM, 190 pM, 200 pM, or more.
[0211] In some embodiments, wherein the engineered Dn29 is used in a Dn29-DBD domain fusion (e.g. Dn29-dCas9), the donor DNA comprises a target sequence which is the same nucleotide sequence as the spacer sequence portion of a guide polynucleotide (e.g, gRNA, sgRNA). In some embodiments, the donor DNA comprises a target sequence which is the same as the target sequence of the target DNA of interest so that the same guide polynucleotide sequence can be used to target the engineered Dn29-DBD fusion to the donor and target DNA of interest.
[0212] The donor DNA can contain a wide variety of different sequences. In some cases, the donor DNA encodes a stop codon, or frame shift, as compared to the target genomic region prior to cleavage and recombination. Such a donor DNA can be useful for knocking out or inactivating a gene or portion thereof. In some cases, the donor DNA encodes one or more missense mutations or in-frame insertions or deletions as compared to the target genomic region. Such a donor DNA can be useful for altering the expression level or activity (e.g., ligand specificity) of a target gene or portion thereof.
[0213] As another example, the donor DNA can encode a wild-type sequence for rescuing the expression level or activity of a target endogenous gene or protein. For instance, T cells containing a mutation in the FoxP3 gene, or a promoter region thereof, can be rescued to treat X-linked IPEX or systemic lupus erythematous. Alternatively, the donor DNA can encode a sequence that results in lower expression or activity of a target gene. For example, an increased immunotherapeutic response can be achieved by deleting or reducing the expression or activity of FoxP3 in T cells prepared for immunotherapy against a cancer or infectious disease target.
[0214] As another example, the donor DNA can encode a mutation that alters the function of a target gene. For instance, the donor DNA can encode a mutation of a cell surface protein necessary for viral recognition or entry. The mutation can reduce the ability of the virus to recognize or infect the target cell. For example, mutations of CCR5 or CXCR4 can confer increased resistance to HIV infection in CD4+ T cells.
[0215] In some cases, the donor DNA encodes a sequence that, although adjacent to, is entirely orthogonal to the endogenous sequence. For example, the donor DNA can encode an inducible promoter or repressor element unrelated to the endogenous promoter of a target gene. The inducible promoter or repressor element can be inserted into the promoter region of a target gene to provide temporal and / or spatial control of the target gene expression or activity.
[0216] In some instances, the donor DNA sequence includes an attD attachment site, such as an attB or an attP site, of an engineered Dn29, a constitutive promoter operably linked to a nucleotide sequence encoding a detectable marker, followed by a nucleotide sequence encoding a first selectable marker.Target DNAs
[0217] Target DNA can be any type of DNA molecule, in vitro or in vivo, including but not limited to genomic DNA, mitochondrial DNA, eukaryotic DNA, prokaryotic DNA,cDNA, and synthesized DNA. The key requirement for the target DNA is that it contains an engineered Dn29 attachment site, including but not limited to an attB site, an attP site, an attH site, or a pseudosite. In some embodiments, the target DNA is human DNA. In some embodiments, the target DNA is non-human primate DNA. In some embodiments, the target DNA is mouse DNA. In some embodiments, the DNA is Common Marmoset (Callithrix jacchus) DNA. In some embodiments, the DNA is Rhesus Macaque (Macaca mulatto) DNA. In some embodiments, the DNA is Cynomolgus Macaque (Macaca fascicularis) DNA.
[0218] The target DNA (or target genome) can contain multiple engineered Dn29 attachment sites. Through the use of engineered Dn29-DBD fusions, the DNA-binding domain of the fusion can direct the engineered Dn29 domain to a single attachment site thereby substantially mitigating off-target recombination.
[0219] In some instances, the target DNA sequence includes an attA attachment site, such as an attB or an attP site, of an engineered Dn29, a constitutive promoter operably linked to a nucleotide sequence encoding a detectable marker, followed by a nucleotide sequence encoding a first selectable marker. In certain types of landing pads, the attachment site is between the promoter and the nucleotide sequence encoding the detectable protein. When there are more than one landing pads used in a given cell, it is preferred that an attachment site of one landing pad is orthogonal to an attachment site of the same large serine recombinase in any other landing pad. The landing pad is used for further genetic engineering and integration of a nucleic acid molecule of interest via site-specific recombination.Vectors and Cell Lines
[0220] Several aspects of the invention relate to vector systems comprising one or more vectors, or vectors as such comprising nucleic acid sequences encoding the engineered Dn29 described herein, and / or comprising donor or target DNA sequences. Several aspects of the invention relate to vector systems comprising one or more vectors, or vectors as such comprising nucleic acid sequences encoding the engineered Dn29-DBD fusions described herein, encoding guide polynucleotides described herein, and / or comprising donor or target DNA sequences. Vectors can be designed for expression of transcripts (e.g. nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, transcripts can be expressed in bacterial cells such as Escherichia co . insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells. Suitable host cells are discussed further in Goeddel, Gene Expression Technology: Methods In Enzymology 185, Academic Press, San Diego, Calif. (1990), the contents of which is hereby incorporated by reference in itsentirety. Alternatively, the recombinant expression vector can be transcribed and translated in vitro, for example using T7 promoter regulatory sequences and T7 polymerase.
[0221] Vectors may be introduced and propagated in a prokaryote. In some embodiments, a prokaryote is used to amplify copies of a vector to be introduced into a eukaryotic cell or as an intermediate vector in the production of a vector to be introduced into a eukaryotic cell (e.g. amplifying a plasmid as part of a viral vector packaging system). In some embodiments, a prokaryote is used to amplify copies of a vector and express one or more nucleic acids, such as to provide a source of nucleic acid constructs or one or more proteins for delivery to a host cell or host organism. Expression of proteins in prokaryotes is most often carried out in Escherichia coli with vectors containing constitutive or inducible promoters directing the expression of either fusion or non-fusion proteins. Fusion vectors add a number of amino acids to a protein encoded therein, such as to the amino terminus of the recombinant protein (in this case engineered Dn29 or engineered Dn29-DBD fusions). Such fusion vectors may serve one or more purposes, such as: (i) to increase expression of recombinant protein; (ii) to increase the solubility of the recombinant protein; and (iii) to aid in the purification of the recombinant protein by acting as a ligand in affinity purification. Often, in fusion expression vectors, a proteolytic cleavage site is introduced at the junction of the fusion moiety and the recombinant protein to enable separation of the recombinant protein from the fusion moiety subsequent to purification of the fusion protein. Such enzymes, and their cognate recognition sequences, include Factor Xa, thrombin and enterokinase. Example fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith and Johnson, 1988. Gene 67: 31-40), pMAL (New England Biolabs, Beverly, Mass.) and pRIT5 (Pharmacia, Piscataway, N.J.) that fuse glutathione S-transferase (GST), maltose E binding protein, or protein A, respectively, to the target recombinant protein, the content of each of which are hereby incorporated by reference in thier entireties.
[0222] Minicircles are small circular plasmids or DNA vectors that are episomal and are produced as a circular expression cassette devoid of any bacterial plasmid backbone. They can be generated from a parental bacterial plasmid that contains a heterologous nucleic acid and two recombinase target sites by intramolecular (cis-) recombination using a site-specific recombinase, such as PhiC31 integrase. Recombination between the two sites generates a minicircle and a leftover miniplasmid. The minicircle can be recovered via separation from the miniplasmid.
[0223] Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pET lid (Studier et al., Gene Expression Technology: Methods In Enzymology 185, Academic Press, San Diego, Calif. (1990) 60-89), the contents of each of which are hereby incorporated by reference in their entireties.
[0224] In some embodiments, a vector is a yeast expression vector. Examples of vectors for expression in yeast Saccharomyces cerivisae include pYepSecl (Baldari, et al., 1987. EMBO J. 6: 229-234), pMFa (Kuijan and Herskowitz, 1982. Cell 30: 933-943), pJRY88 (Schultz et al., 1987. Gene 54: 113-123), pYES2 (Invitrogen Corporation, San Diego, Calif.), and picZ (InVitrogen Corp, San Diego, Calif.), the contents of each of which are hereby incorporated by reference in their entireties.
[0225] In some embodiments, a vector drives protein expression in insect cells using baculovirus expression vectors. Baculovirus vectors available for expression of proteins in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith, et al., 1983. Mol. Cell. Biol. 3: 2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170: 31-39), the contents of each of which are hereby incorporated by reference in their entireties.
[0226] In some embodiments, a vector is capable of driving expression of one or more sequences in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329: 840) and pMT2PC (Kaufman, et al., 1987. EMBO J. 6: 187-195), the contents of each of which are hereby incorporated by reference in their entireties. When used in mammalian cells, the expression vector's control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other suitable expression systems for both prokaryotic and eukaryotic cells see, e.g., Chapters 16 and 17 of Sambrook, et al., Molecular Cloning: A Laboratory Manual. 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989, the contents of each of which are hereby incorporated by reference in their entireties.
[0227] In some embodiments, a vector is capable of driving expression of one or more sequences in plant cells using a plant cell expression vector.
[0228] In some embodiments, the recombinant mammalian expression vector is capable of directing expression of the nucleic acid preferentially in a particular cell type (e.g., tissuespecific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promotersinclude the albumin promoter (liver-specific; Pinkert, et al., 1987. Genes Dev. 1 : 268-277), lymphoid-specific promoters (Calame and Eaton, 1988. Adv. Immunol. 43: 235-275), in particular promoters of T cell receptors (Winoto and Baltimore, 1989. EMBO J. 8: 729-733) and immunoglobulins (Baneiji, et al., 1983. Cell 33: 729-740; Queen and Baltimore, 1983. Cell 33: 741-748), neuron-specific promoters (e.g., the neurofilament promoter; Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86: 5473-5477), pancreas-specific promoters (Edlund, et al., 1985. Science 230: 912-916), and mammary gland-specific promoters (e.g., milk whey promoter; U.S. Pat. No. 4,873,316 and European Application Publication No. 264,166). Developmentally-regulated promoters are also encompassed, e.g., the murine hox promoters (Kessel and Gruss, 1990. Science 249: 374-379) and the a-fetoprotein promoter (Campes and Tilghman, 1989. Genes Dev. 3: 537-546), the contents of each of which are hereby incorporated by reference in their entireties.
[0229] In some embodiments, methods for introducing engineered Dn29 or engineered Dn29-DBD fusion-gRNA ribonucleoprotein complex into a cell (e.g., a hematopoietic cell or hematopoietic stem cell, including, e.g., such cells from humans, human embryonic stem cells, T cells including primary T cells, non-dividing cells) include forming a reaction mixture containing the protein or ribonucleoprotein complex and introducing transient holes in the extracellular membrane of the cell. Such transient holes can be introduced by a variety of methods, including, but not limited to, electroporation, cell squeezing, or contacting with nanowires or nanotubes. Generally, the transient holes are introduced in the presence of the protein or ribonucleoprotein complex and the protein or ribonucleoprotein complex is allowed to diffuse into the cell.
[0230] Methods, compositions, and devices for electroporating cells to introduce a protein or ribonucleoprotein complex can include those described in WO / 2006 / 001614 or Kim, J. A. et al. Biosens. Bioelectron. 23, 1353-1360 (2008), the contents of each of which are hereby incorporated by reference in their entireties. Additional or alternative methods, compositions, and devices for electroporating cells to introduce a protein or ribonucleoprotein complex can include those described in U.S. Patent Appl. Pub. Nos. 2006 / 0094095;2005 / 0064596; or 2006 / 0087522, the contents of each of which are hereby incorporated by reference in their entireties. Additional or alternative methods, compositions, and devices for electroporating cells to introduce a protein or ribonucleoprotein complex can include those described in Li, L. H. et al. Cancer Res. Treat. 1, 341-350 (2002); U.S. Pat. Nos. 6,773,669; 7,186,559; 7,771,984; 7,991,559; 6,485,961; 7,029,916; and U.S. Patent Appl. Pub. Nos:2014 / 0017213; and 2012 / 0088842 and Geng, T. et al. J. Control Release 144, 91-100 (2010); and Wang, J., et al. Lab. Chip 10, 2057-2061 (2010), the contents of each of which are hereby incorporated by reference in their entireties.
[0231] In some cases, the methods or compositions described in the patents or publications cited herein are modified for protein or ribonucleoprotein delivery. Such modification can include increasing or decreasing voltage, pulse length, and / or the number of pulses. Such modification can further include modification of buffers, media, electrolytic solutions, or components thereof. Electroporation can be performed using devices known in the art, such as a Bio-Rad Gene Pulser Electroporation device, an Invitrogen Neon transfection system, a MaxCyte transfection system, a Lonza Nucleofection device, a NEPA Gene NEPA21 transfection device, a flow through electroporation system containing a pump and a constant voltage supply, or other electroporation devices or systems known in the art.
[0232] Methods, compositions, and devices for squeezing or deforming a cell to introduce a protein or ribonucleoprotein complex can include those described herein. Additional or alternative methods, compositions, and devices can include those described in Nano Lett. 2012 Dec. 12; 12(12):6322-7; Proc Natl Acad Sci USA. 2013 Feb. 5;110(6):2082-7; J Vis Exp. 2013 Nov. 7; (81):e50980; and Integr Biol (Camb). 2014 April; 6(4):470-5, the contents of each of which are hereby incorporated by reference in their entireties. Additional or alternative methods, compositions, and devices can include those described in U.S. Patent Appl. Publ. No. 2014 / 0287509, the content of which is hereby incorporated by reference in its entirety. Generally, the protein or ribonucleoprotein complex is provided in a reaction mixture containing the cell and the reaction mixture is forced through a cell deforming orifice or constriction. In some cases, the constriction is smaller than the diameter of the cell. In some cases, the constriction contains cell-deforming components such as regions of strong electrostatic charge, regions of hydrophobicity, or regions containing nanowires or nanotubes. The forcing can introduce transient pores into a cell membrane of the cell allowing the protein or ribonucleoprotein complex to enter the cell through the transient pores. In some cases, squeezing or deforming a cell to introduce the protein or ribonucleoprotein can be effective even when the cell is in a non-dividing state.
[0233] Methods for introducing a protein or ribonucleoprotein complex into a cell include forming a reaction mixture containing the protein or ribonucleoprotein complex and contacting the cell with the protein or ribonucleoprotein complex to induce receptor-mediated internalization. Compositions and methods for receptor mediated internalization aredescribed, e.g., in Wu et al., J. Biol. Chem. 262, 4429-4432 (1987); and Wagner et al., Proc. Natl. Acad. Sci. USA 87, 3410-3414 (1990), the contents of each of which are hereby incorporated by reference in their entireties. Generally, the receptor-mediated internalization is mediated by interaction between a cell surface receptor and a ligand fused to the protein or fused to the ribonucleoprotein complex (e.g., covalently attached or fused to an RNA in the ribonucleoprotein complex). The ligand can be any protein, small molecule, polymer, or fragment thereof that binds to, or is recognized by, a receptor on the surface of the cell. An exemplary ligand is an antibody or an antibody fragment (e.g., scFv).
[0234] In some embodiments, the reaction mixture for introducing the protein or ribonucleoprotein complex into the cell can contain a nucleic acid for directing binding to the target genomic region.
[0235] In some embodiments, delivery is via a nucleic acid (e.g., plasmid(s)) transfected into a cell. The transfected nucleic acids (e.g., plasmid(s)) can comprise an expression vector for an engineered Dn29 or an engineered Dn29-DBD fusion, a nucleic acid (e.g., plasmid) comprising a donor molecule for integration into the cell’s genome, and in the case of engineered Dn29-DBD fusions, an expression vector for guide polynucleotides (e.g., gRNA or sgRNA).
[0236] The nucleic acids may be delivered using adeno associated virus (AAV), lentivirus, adenovirus or other viral vector types, or combinations thereof. The nucleic acids can be packaged into virions using appropriate packaging cells lines as known in the art. In some embodiments, the engineered Dn29 or engineered Dn29-DBD fusion protein and one or more exogenous nucleic acids are delivered to a cell using a lentivirus particle.
[0237] In some cases, expression of the engineered Dn29 or engineered Dn29-DBD fusions described herein and / or the guide polynucleotides are under the control of an inducible promoter or repressor element. The inducible promoter or repressor element can be inserted into the promoter region of a nucleic acid sequence encoding the engineered Dn29 or engineered Dn29-DBD fusions described herein and / or the guide polynucleotides to provide temporal and / or spatial control of the expression or activity.
[0238] Upon delivery of a nucleic acid encoding an engineered Dn29 or engineered Dn29-DBD fusion to a cell, the nucleic acid can be transcribed and translated into an engineered Dn29 or engineered LSR-DBD protein. The engineered Dn29 or engineered Dn29-DBD protein can form a tetrameric complex inside the cell. In some embodiments, theengineered Dn29 and engineered Dn29-DBD form a tetrameric complex which can comprise one, two, or three engineered Dn29-DBD fusion proteins.Applications
[0239] Described herein are several applications of the engineered Dn29 or engineered Dn29-DBD fusion system described herein including, but not limited to a method for amplicon library installation at genomic landing pads, delivery of cargos without a landing pad with sufficient efficiency to integrate multiple constructs in the same cell simultaneously, and direct targeting of specific sites in a mammalian genome with significantly higher efficiency than PhiC31 (which has ~1% genome-targeted LSR integration efficiency).
[0240] Site-specific nucleases and site-specific recombinases are powerful tools for targeted genome modification in vitro and in vivo. It has been reported that nuclease cleavage in living cells triggers a DNA repair mechanism that frequently results in a modification of the cleaved and repaired genomic sequence, for example, via homologous recombination. Accordingly, the targeted cleavage of a specific unique sequence within a genome using the engineered Dn29 or engineered Dn29-DBD fusions described herein opens up new avenues for gene targeting and gene modification in living cells, including cells that are hard to manipulate with conventional gene targeting methods, such as many human somatic or embryonic stem cells. Site-specific recombinases possess all the functionality required to bring about efficient, precise integration, deletion, inversion, or translocation of specified DNA segments without exposed DNA double-stranded breaks.
[0241] In some cases, the efficiency of genome-targeted integration using the engineered Dn29 or engineered Dn29-DBD fusion proteins described herein can be at least about, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%. 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65% 70%, 75%, 80%, 85%, 90%, 95%, 99%, or higher than integration using a wildtype Dn29 (SEQ ID NO: 1). In some cases, the efficiency of incorporation of the sequence of the donor DNA can be at least, or at least about, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65% 70%, 75%, 80%, 85%, 90%, 95%, 99%, or higher than incorporation using a wildtype Dn29 (SEQ ID NO: 1).
[0242] In some embodiments, the one or more nucleic acids encoding engineered Dn29 described herein, or engineered Dn29-DBD fusion and guide polynucleotide(s) described herein, are used to produce a non-human transgenic animal or transgenic plant or transgenic organoid. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat,rabbit, non-human primate, common marmoset, Rhesus macaque, or cynomolgus macaque. In certain embodiments, the organism or subject is a plant. In certain embodiments, the organism or subject or plant is algae or crops. In certain embodiments, the subject is an organoid. Methods for producing transgenic plants, organoids, and animals are known in the art, and generally begin with a method of cell transfection, such as described herein. Transgenic animals are also provided, as are transgenic plants, especially crops and algae. The transgenic animal or plant may be useful in applications outside of providing a disease model. These may include food or feed production through expression of, for instance, higher protein, carbohydrate, nutrient or vitamins levels than would normally be seen in the wildtype. In this regard, transgenic plants, especially pulses and tubers, and animals, especially mammals such as livestock (cows, sheep, goats and pigs), but also poultry and edible insects, are preferred.
[0243] The invention comprehends the use of the nucleic acids, polypeptides, compositions, systems, and methods disclosed herein to establish and utilize transgenic cells / animals / organoids. Disclosed herein is a non-naturally occurring or engineered composition, or one or more polynucleotides encoding components of said composition, or vector or delivery systems comprising one or more polynucleotides encoding components of said composition for use in a modifying a target cell in vivo, ex vivo or in vitro and, may be conducted in a manner alters the cell such that once modified the progeny or cell line of the modified cell retains the altered phenotype. The modified cells and progeny may be part of a multicellular organism such as a plant or animal with ex vivo or in vivo application of the engineered Dn29 described herein, or engineered Dn29-DBD fusion system to desired cell types.
[0244] The invention may be a therapeutic method of treatment. The therapeutic method of treatment may comprise gene or genome editing, or gene therapy. A method of the invention may be used to create a plant, an animal or cell that may be used to model and / or study genetic or epigenetic conditions of interest, such as through a model of mutations of interest or as a disease model. A method of the invention may be used to create an animal or cell that comprises a modification in one or more nucleic acid sequences associated with a disease, or a plant, animal or cell in which the expression of one or more nucleic acid sequences associated with a disease are altered. Such a nucleic acid sequence may encode a disease associated protein sequence or may be a disease associated control sequence. Accordingly, it is understood that in embodiments of the invention, a plant, subject, patient,organism, or cell can be a non-human subject, patient, organism or cell. Thus, the invention provides a plant, animal or cell, produced by the present methods, or a progeny thereof. The progeny may be a clone of the produced plant or animal, or may result from sexual reproduction by crossing with other individuals of the same species to introgress further desirable traits into their offspring. The cell may be in vivo or ex vivo in the cases of multicellular organisms, particularly animals or plants. In the instance where the cell is cultured, a cell line may be established if appropriate culturing conditions are met and preferably if the cell is suitably adapted for this purpose (for instance a stem cell). Bacterial cell lines produced by the invention are also envisaged. Hence, cell lines are also envisaged.
[0245] In a further embodiment, the present invention provides an ex vivo method for transfecting the engineered Dn29, or engineered Dn29-DBD fusion system described herein in relevant host cells (e.g. e.g., a hematopoietic cell or hematopoietic stem cell, including, e.g., such cells from humans, human embryonic stem cells, T cells including primary T cells, non-dividing cells). In one embodiment, suitable cells are isolated from the mammal, eventually differentiated in vitro and incubated with an effective amount of a pharmaceutical composition of the present invention. Thereafter, the treated (transfected) cells are reintroduced into the organism.
[0246] The gene therapy composition of the invention comprises, in addition to adequate salts (alkali metal as counter ion and dications in formulation) and eventually other therapeutic or immunosuppressive agents, a pharmaceutically acceptable carrier and / or a pharmaceutically acceptable vehicle and / or pharmaceutically acceptable diluent. Appropriate routes for suitable formulation and preparation of the gene therapy vehicle according to the invention are disclosed in Remington: “The Science and Practice of Pharmacy,” 20th Edn., A. R. Gennaro, Editor, Mack Publishing Co., Easton, Pa. (2003), the content of which is hereby incorporated by reference in its entirety. Possible carrier substances for parenteral administration are e.g. sterile water, Ringer, Ringer lactate, sterile sodium chloride solution, polyalkylene glycols, hydrogenated naphthalenes and, in particular, biocompatible lactide polymers, lactide / glycolide copolymers or polyoxy ethylene / polyoxy-propylene copolymers. The particular embodiments of the gene therapy formulation are chosen according to the physical properties, for example in respect of solubility, stability, bioavailability or degradability.
[0247] The mode and method of administration and the dosage of the gene therapy according to the invention depend on the nature of the disease to be treated, where appropriate the stage thereof, and also the body weight, the age and the sex of the patient.
[0248] In some methods, the disease model can be used to study the effects of mutations on the animal or cell and development and / or progression of the disease using measures commonly used in the study of the disease. Alternatively, such a disease model is useful for studying the effect of a pharmaceutically active compound on the disease.
[0249] In some methods, the disease model can be used to assess the efficacy of a potential gene therapy strategy. That is, a disease-associated gene or polynucleotide can be modified such that the disease development and / or progression is inhibited or reduced. In particular, the method comprises modifying a disease-associated gene or polynucleotide such that an altered protein is produced and, as a result, the animal or cell has an altered response. Accordingly, in some methods, a genetically modified animal may be compared with an animal predisposed to development of the disease such that the effect of the gene therapy event may be assessed.
[0250] In another embodiment, this invention provides a method of developing a biologically active agent that modulates a cell signaling event associated with a disease gene. The method comprises contacting a test compound with a cell comprising one or more vectors that drive expression of the engineered Dn29, or engineered Dn29-DBD fusion system of the present invention; and detecting a change in a readout that is indicative of a reduction or an augmentation of a cell signaling event associated with, e.g., a mutation in a disease gene contained in the cell.
[0251] A cell model or animal model can be constructed in combination with the method of the invention for screening a cellular function change. Such a model may be used to study the effects of a genome sequence modified by the engineered Dn29, or engineered Dn29- DBD fusion of the invention on a cellular function of interest. For example, a cellular function model may be used to study the effect of a modified genome sequence on intracellular signaling or extracellular signaling. Alternatively, a cellular function model may be used to study the effects of a modified genome sequence on sensory perception. In some such models, one or more genome sequences associated with a signaling biochemical pathway in the model are modified.
[0252] A transgenic cell in which one or more nucleic acids encoding one or more of the components of the present invention are provided or introduced can be operably connected inthe cell with a regulatory element comprising a promoter of one or more gene of interest. In some embodiments, a cell, such as a eukaryotic cell, can have an engineered Dn29, or engineered Dn29-DBD fusion, genomically integrated. The nature, type, or origin of the cell are not particularly limiting according to the present invention. Also the way in which the engineered Dn29, or engineered Dn29-DBD fusion, transgene is introduced in the cell may vary and can be any method as is known in the art. In certain embodiments, the engineered Dn29, or engineered Dn29-DBD fusion, transgenic cell is obtained by introducing the engineered Dn29, or engineered Dn29-DBD fusion, transgene in an isolated cell. In certain other embodiments, the engineered Dn29, or engineered Dn29-DBD fusion, transgenic cell is obtained by isolating cells from an engineered Dn29, or engineered Dn29-DBD fusion, transgenic organism. By means of example, and without limitation, the engineered Dn29, or engineered Dn29-DBD fusion, transgenic cell as referred to herein may be derived from an engineered Dn29, or engineered Dn29-DBD fusion, transgenic eukaryote, such as an engineered Dn29, or engineered Dn29-DBD fusion, knock-in eukaryote. Reference is made to WO 2014 / 093622 (PCT / US 13 / 74667), the contents of which is hereby incorporated by reference in its entirety. Methods of US Patent Nos. 8,771,985 and 9,567,573 assigned to Sangamo Biosciences, Inc. (and the contents of each of which are hereby incorporated by reference in their entireties) directed to targeting the Rosa locus may be modified to utilize the engineered LSR-DBD fusion system of the present invention. Methods of US Patent Publication No. 20130236946 assigned to Cellectis (the contents of which is hereby incorporated by reference in its entirety) directed to targeting the Rosa locus may also be modified to utilize the engineered Dn29, or engineered Dn29-DBD fusion, system of the present invention. The engineered Dn29, or engineered Dn29-DBD fusion, transgene can further comprise a Lox-Stop-poly A-Lox(LSL) cassette thereby rendering engineered Dn29, or engineered Dn29-DBD fusion, expression inducible by Cre recombinase. Alternatively, the engineered Dn29, or engineered Dn29-DBD fusion, transgenic cell may be obtained by introducing the engineered Dn29, or engineered Dn29-DBD fusion, transgene in an isolated cell. Delivery systems for transgenes are well known in the art. By means of example, the engineered Dn29, or engineered Dn29-DBD fusion, protein transgene may be delivered in for instance eukaryotic cell by means of vector (e.g., AAV, adenovirus, lentivirus) and / or particle and / or nanoparticle delivery, as also described herein elsewhere.
[0253] In certain aspects, described herein is a cell comprising a nucleic acid encoding any of the engineered Dn29, or engineered Dn29-DBD fusions, disclosed herein. In someembodiments, the genome of the cell comprises an attachment site for the engineered Dn29, or engineered Dn29 portion of the engineered Dn29-DBD fusion. Such a cell line can be used in a method wherein a nucleic acid comprising a donor attachment site and a nucleic acid for insertion is introduced into the cell to generate an engineered cell line comprising the nucleic acid of interest inserted into the engineered Dn29 attachment site. In some embodiments, described herein is a kit comprising a cell, the cell comprising a nucleic acid encoding any of the engineered Dn29, or engineered Dn29-DBD fusions, disclosed herein. In some embodiments, the genome of the cell of the kit comprises an attachment site for the engineered Dn29 or engineered Dn29 portion of the engineered Dn29-DBD fusion. In some embodiments, the kit further comprises a nucleic acid vector (e.g. plasmid) comprising a donor attachment site. In some embodiments, the nucleic acid vector (e.g. plasmid) of the kit further comprises a multicloning site for insertion of a nucleic acid of interest.
[0254] Several further aspects of the invention relate to modeling defects associated with a wide range of genetic diseases in plants or animals, which are further described on the website of the National Institutes of Health under the topic subsection Genetic Disorders (website at health.nih.gov / topic / GeneticDisorders).SEQUENCES AND TABLES
[0255] The amino acid sequence of Dn29 is provided as SEQ ID NO: 1 (Figure 34).
[0256] Exemplary nucleic acid sequence encoding Dn29 is provided as SEQ ID NO: 2(Figure 34).
[0257] The amino acid sequence of Dn29-dCas9 fusions is provided as SEQ ID NOs: 3-7 (Figure 34).
[0258] The native attP sequences for Dn29 is provided as SEQ ID NO: 8: TGATTTTACGCTGGTGCTATATCCTAAACTCCCACAGATAAACAGTTAATGGTAA TGAAATAACATTAACTGTTTATCTGTGTTTAATGCCTTAACTTAATCTAGTAGGAG GG (SEQ ID NO: 8).
[0259] The native attB sequences for Dn29 is provided as SEQ ID NO: 9: AAGTTAAAGCGGAGGTTTCTCTGTACGACCCCATTGGTGTAGACAAGGAAGGTA ATGAAATAAGTTTGATAGATATTTTGGGTACCGACCCGGAAGTGGTGGCGGACA TGGTG (SEQ ID NO: 9).
[0260] The amino acid sequence of dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl is provided as SEQ ID NOs: 10-13 (Figure 34).
[0261] Exemplary nucleic acid sequences encoding dCas9, dCas9-HFl, dCas9-SpG, dCas9-SpG-HFl are provided as SEQ ID NOs: 14-17 (Figure 34).
[0262] The amino acid sequence of exemplary linkers are provided as SEQ ID NOs: 18- 26 (Figure 34). Exemplary nucleic acid sequences encoding exemplary linkers are provided as SEQ ID NOs: 27-35 (Figure 34).
[0263] Figure 34 discloses exemplary attD sequences (SEQ ID NOs: 109) and corresponding attH pseudosites (provided as chromosomal locus according to human genome assembly GRCh38, available at www.ncbi.nlm.nih.gov / genome / guide / human / ) for Dn29. Exemplary pseudosites for model organisms for Dn29 are provided as chromosomal locus according to their genome assembly (Mouse) Mus musculus chrX:98,518,458; (Common Marmoset) Callithrix jacchus chr7: 19,911,959; (Rhesus Macaque) Macaca mulatta chr9:22,186,336; (Cynomolgus Macaque) Macaca fascicularis chr9:21,946,090.
[0264] Figure 34 also discloses exemplary gRNA sequences with target site (provided as chromosomal locus according to human genome assembly GROG 8, available at www.ncbi.nlm.nih.gov / genome / guide / human / ) that the gRNA spacer is proximal to, overlapping with, or within, the target DNA sequence (SEQ ID NOs: 110-152), the corresponding gRNA spacer (SEQ ID NOs: 153-195, and an exemplary gRNA scaffold (SEQ ID NO: 108) for use with the LSR-DBD fusions described herein.
[0265] Figure 53 discloses a multiple sequence alignment between attHl and other attHl pseudosites across found in different model organisms. Nucleic acid sequences encoding these pseudosites are provided as SEQ ID NOS: 209-213 (Figure 53).Table 1: MutationsTable 2: All Variants. “A” indicates a silent mutation in the codon of the nucleic acid encoding the amino acid position shown.
[0266] The invention is further described by the following non-limiting Examples.EXAMPLES
[0267] Examples are provided below to facilitate a more complete understanding of the invention. The following examples serve to illustrate the exemplary modes of making and practicing the invention. However, the scope of the invention is not to be construed as limited to specific embodiments disclosed in these Examples, which are illustrative only. EXAMPLE 1:
[0268] Manipulation of eukaryotic genomes, particularly the integration of multi-kilobase DNA sequences, remains challenging and limits the rapidly growing fields of synthetic biology and cell engineering. Programmable nucleases such as Cas9 coupled with homologous repair templates revolutionized the field by enabling programmable integration of DNA cargos (Mali et al. 2013, Cong et al. 2013, Jinek et al. 2013). However, this system relies on generating double stranded breaks and exploiting host repair machinery and homology directed repair, limiting its application to dividing cells and raising potential safety concerns for gene therapy applications. Lentiviral and piggybac-based approaches for cell engineering are popular due to their broad tropism and ease of use, however their lack of target specificity increases outcome variability and insertional mutagenesis risk. Recent developments, such as CRISPR-associated transposases, show promise for programmable DNA integration (Lampe et al 2023, Tou et al. 2023). However, their mechanisticdisadvantages, including the need for the association of numerous protein subunits into large complexes, significantly limit integration efficiency in human cells.
[0269] Many of these limitations could be addressed by large serine recombinases (LSRs), which are small (-500 AA) single subunit effectors that mediate site-specific and unidirectional integration of DNA through recombination of two attachment sequences (Van Duyne and Rutherford, 2013). Natively, these enzymes facilitate mobile genetic element integration into bacterial genomes, classically originating from bacteriophage integration systems. These enzymes recognize and recombine two DNA attachment sites (attP and attB), resulting in site-specific payload integration without generating double stranded breaks or relying on host repair mechanisms. These enzymes have been shown to site-specifically integrate DNA payloads containing a donor attachment site (attD) into mammalian cells, both at pre-installed integration sites called landing pads (attA) or at endogenous genomic pseudosites with high sequence similarity to their native integration site (attH) (Thyagarajan et al. 2001, Matreyek et al, 2017, Durrant et al. 2023). However, the efficiency and specificity of pseudosite integration is severely limited by various factors, including the biochemical and enzymatic properties of LSRs, the degree of sequence similarity between pseudosites and attachment sites, and the number and quality of pseudosites present within the genome.Development of LSRs which can precisely and efficiently integrate into endogenous sites will bypass the current two step process of pre-inserting a landing pad, reducing effort, time, and complexity of the genetic engineering process and permitting easy translation of genome engineering to new cell types or organisms.
[0270] We have implemented various engineering strategies to enhance the efficiency and specificity of LSR-mediated integration into endogenous pseudosites. We conducted a comprehensive deep mutational scan-directed evolution campaign to generate new LSR variants, which were screened in mammalian cells for on- and off-target integration to identify driver mutations for efficiency and specificity. Through rounds of rational mutation combinations, we fine tuned the variants for either higher efficiency, improved specificity, or both. As a result, we have developed mutants that exhibit up to an eight-fold increase in integration efficiency while reducing off-target effects.Methods:Dn29 Deep Mutational Scanning Cloning
[0271] We generated an NNK deep mutational scanning library of the entire Dn29 CDS using NNK oligos and overlap extension PCRs. Frst, we used a custom script to generateNNK oligos targeting each amino acid position in the Dn29 protein with a primer Tm of 65°C. We ordered these oligos as 500 picomole DNA oligo plates from IDT, with the forward and reverse primers for the same position in the same well location in different plates. We also designed universal forward and reverse primers, which anneal upstream or downstream, respectively, of the protein CDS, to be paired with each NNK oligo, also at a Tm of 65°C. We used these oligos to amplify the regions upstream of the desired NNK mutation (Fragment 1) with the following protocol: A master mix of consisting of 2.5 pL Q5 Mastermix (NEB), 0.01 pL Dn29 plasmid template (100 ng / pL), 0.025 pL Universal forward primer (lOOpM), and 1 ,465uL water per reaction was combined with 1 pL unique NNK reverse primer (2.5 pM) in a 96 well plate. We similarly generated the protein fragment downstream of the NNK mutation (Fragment 2), but with the universal reverse primer and the NNK forward primer. The PCR reaction was cycled as follows: 1 cycle of 98°C for 30 seconds, 30 cycles of 98°C for 10 seconds, 60°C for 30 seconds, and 72°C for 1 minute, and a final extension cycle of 72°C for 2 minutes.
[0272] Next, for each pair of Fragment 1 and Fragment 2 reactions, we pooled 2.5 pL of each into a new 96 well plate. The reaction was cleaned up with 2 pL of ExoSAP-IT (Thermo-Fisher) and 0.5 pL of Dpnl (NEB) per well, and incubated at 37°C for 30 minutes followed by 80°C for 15 minutes. To perform the next overlap extension PCR, we took 1 pL of the cleaned Fragment 1 and Fragment 2 PCR pools, 2.5 pL of Q5 2x mastermix, 0.025 pL of each universal primer (100 pM), and 1.45 pL of water, and ran the same PCR cycling protocol as described above.
[0273] Finally, to generate an entire mutant pool, we pooled 2.5 pL of each overlap extension PCR into a single microfuge tube. We gel purified only the band corresponding to the full length Dn29 fragment, and then digested the library with Xbal and Hindlll-HF, whose digest sites were located flanking the Dn29 CDS. We also pre-digested the pEVO backbone with the same enzymes. Next we performed a T4 ligation reaction with 100 ng of total DNA, a 3:1 molar ratio of library to backbone, 2 pL of T4 ligase (NEB), 4 pL of lOx T4 ligase buffer (NEB), and water to 40 pL total volume. This reaction was split into 2x20 pL reactions, and ligated for 30 minutes at room temperature, then inactivated at 65°C for 10 minutes. Next, this library was purified with Zymo Clean and Concentrate -5 Kit, electroporated into XL-1 Blue cells according to manufacturer’s instructions, recovered for 1 hour at 37°C, and plated onto four Bioassay dishes. An approximate colony count of IMcolonies was calculated with serial dilutions. The bioassay dishes were scraped and the plasmids were purified using Macherey Nagel Midiprep Kits.Substrate Linked Directed Evolution
[0274] 4 pL of the mutant library was electroporated into 50 pL XL-1 Blue cells in a1mm Biorad Pulse Cuvette using the following conditions: 1700 V, 25 pF, 200 . Cultures were recovered for 1 hour at 37°C in ImL SOC media, and then seeded into lOOmL LB media with desired concentration of L-arabinose and grown at 37°C overnight. Serial dilutions were plated on LB+Carb plates to determine the library coverage, which was maintained above IM colonies, and grown overnight at 37°C. The next day, the cultures were spun down and plasmid was extracted with Qiagen Plasmid Midi kit, with 0.3 g wet bacteria pellet per column. Next, 500 ng of plasmid was digested with Ndel (NEB) to digest inactive variants. To amplify active variants, the digested material is amplified using the following conditions: 25 pL 2x Platinum Superfi II Mastermix (Thermo Fisher), 19 pL water, 2 pL forward primer, 2 pL reverse primer and 2 pL of Ndel digested material, with the following cycling protocol: 98°C for 30 seconds, 30 cycles of 98°C for 10 seconds, 52°C for 10 seconds, and 72°C for 55 seconds, and a final extension of 72°C for 5 minutes. The PCR reaction was run on a 1% agarose gel, and the correct size band was gel extracted with the Monarch DNA Gel Extraction Kit (NEB).
[0275] Next, the amplified active variants were cloned into the pEVO plasmid backbone. First the gel extracted plasmid library and the plasmid backbone were digested with Xbal and Hindlll, which directly flanks the CDS and creates overhangs for ligation cloning into the plasmid backbone. This digest was run at 37°C for 30 minutes, and then heat inactivated at 80°C for 20 minutes. Next, the active variant digest was purified with DNA Clean and Concentrator -5 (Zymo), and the digested plasmid backbone was purified with DNA Clean and Concentrator -25 (Zymo). Five T4 ligation reactions were set up, using a 3: 1 ratio of amplified variants to backbone, a total DNA input of 100 ng per 20 pL reaction, NEB T4 ligase, and NEB T4 ligase buffer, and ligated for 30 minutes at room temperature, then heat inactivated for 10 minutes at 65°C. All ligation reactions were pooled and purified with the DNA Clean and Concentrator -5 Kit (Zymo), eluted in 6 pl, and electroporated into XL-1 Blue cells, starting the N+l cycle of directed evolution.DNA Shuffling
[0276] Shuffling the active variants between rounds of cycling involved a uridine exchange PCR to swap out some thymidines for uridine, USER enzyme fragmentation atI l luridines, and fragment reassembly with two rounds of PCR. For the uridine exchange PCR, the ratio of dUTP / dTTP was optimized to be 3 / 7 for optimal yield and fragmentation size for Dn29. The following PCR recipe was used: 5 pL lOx Thermopol Buffer, 1 pL lOmM dNTPs,I pL Forward Primer (10 pM), luL Reverse Primer (10 pM), 1 pL plasmid library, 1 pL Taq Polymerase, and 40 pL water. The PCR was cycled with the following protocol: 1 cycle of 95°C for 30 seconds, 30 cycles of 95°C for 20 seconds, 60°C for 30 seconds, and 68°C for 1 min / kb, and a final extension of 68°C for 5 minutes. The PCR was run on an agarose gel, and the DNA band of the full length gene was extracted using the Monarch Gel Extraction Kit (NEB). The product was split into 500 ng aliquots, and digested with USER Enzyme (NEB), which is a uracil specific excision reagent that generates a single nucleotide gap at each uracil. The digest is run for 3 hours at 37°C, and run on a gel to determine fragment size distribution.
[0277] To reassemble these fragments, the DNA was first purified using DNA Clean and Concentrator -5 (Zymo). Next, a primerless PCR, containing 25 pL purified DNA fragments and 25 pL 2x Q5 High Fidelity Master Mix (NEB), was cycled with the following protocol: 1 Cycle at 98°C for 30 seconds, 30-50 cycles at 98°C for 10 seconds, 30°C for 30 seconds with +1°C per cycle, and 72°C for 1 minute +4s per cycle, and a final extension of 1 cycle at 72°C for 10 minutes. Efficiency of reassembly can be assessed on an agarose gel, which should yield a strong band at the expected full length gene size. A final PCR was run to recover the full length gene with universal primers outside the CDS that include the restriction digest sites used for cloning. The PCR recipe is as follows: 25 pL Platinum Superfi II 2x Mastermix (Thermo), 10 pL reassembled fragment DNA, 2 pL forward primer, 2 pL reverse primer, andI I pL water. This reaction was cycled with the following protocol: 1 cycle at 98°C for 30 seconds, 35 cycles of 98°C for 10 seconds, 60°C for 10 seconds, and 72°C for 55 seconds, and finally 1 cycle of 72°C for 5 minutes. After gel extraction, this new library of shuffled and reassembled genes can be cloned back into the plasmid backbone utilizing the Xbal+Hindlll digest and T4 ligation protocol described above.Variant Library Next Generation Sequencing and Analysis
[0278] To sequence the input and output variant libraries with Illumina Next Generation Sequencing, we designed 6 sets of primers that annealed to the Dn29 CDS in -260 bp segments, and contained Illumina adapter overhangs. We performed two rounds of PCR to amplify each region of the protein and add P5 and P7 adapters and i5 and i7 indices. The amplicons were Ampure XP (Beckman Coulter) bead cleaned in between rounds of PCR andafter the final PCR, pooled by equimolar ratios, quantified with Qubit dsDNA High Sensitivity Kit (Thermo Fisher), and sequenced on the Illumina NextSeq 600 cycle kit. For each amplicon, we ensured full overlap between read 1 and read 2, so that each read has a higher Q score, resulting in higher confidence of single mutations versus sequencing error. We generated a custom python script to count and plot the mutations at each position, and calculated enrichment between the input library and an output library with the following formul a : ((% A Aoutput) / ( 1 -% A Aoutput)) / ((% A Ainput) / ( 1 -% A Ainput)) .Transfection of HEK293FTs
[0279] One day before transfection, 12-18K HEK293FT cells were plated per well of a 96 well plate, with a target of 60% confluency at the time of transfection. Each well was transfected with 725 ng of DNA, containing a 5: 1 molar ratio of donor plasmid to effector plasmid, using 0.5 pL of Lipofectamine 2000 (Thermo) per well. The cells were incubated and monitored for three days for mCherry expression (from the donor plasmid) and GFP expression (from the effector plasmid), before harvesting cells for flow or genomic DNA for ddPCR.Cell Harvest, ddPCR and Flow Cytometry
[0280] After three days, cells were trypsinized with 50 pL TrypLE (Gibco) for 10 minutes, and then quenched with 50 pL Stain Buffer (BD). The 100 pL cell suspension was split into two 50 pL aliquots in U Bottom 96 well plates. The two plates were centrifuged at 300 g for 5 minutes, then the supernatant was aspirated, and one plate was resuspended in 200 pL Stain buffer for flow while the other plate was resuspended in 50 pL QuickExtract DNA I (Biosearch Technologies) for genomic DNA extraction. The cells for flow were analyzed with the Attune NxT Flow Cytometer with Autosampler (Thermo Fisher). The cells in QuickExtract were vortexed for 15 seconds, and then run on a thermocycler with the following protocol: 65°C for 15 minutes, 68°C for 15 minutes, 98°C for 10 minutes. Next, the genomic DNA was cleaned with a 0.9x Ampure Bead cleanup. ddPCR samples were prepared with the following reaction: 11 pL ddPCR Supermix for Probes (no dUTP) (BioRad), 1.98 pL of each primer at 10 pM, 0.55 pL of each probe at 10 pM, 1.65 pL of cleaned gDNA, 0.22 pL SacI-HF (NEB), and water up to a final volume of 22 pL. Each reaction contains two sets of primers and 2 probes, amplifying the target site (attHl or attH3) with a FAM probe (IDT) and amplifying a reference on the same chromosome with a HEX probe (IDT). This reaction is run on the QX200 AutoDG Droplet Digital PCR System (Biorad).Site-directed mutagenesis for combinatorial mutant cloning
[0281] Site-directed mutagenesis primers were designed using the script provided in Bi et al. 2019, choosing the primers with Tm closest to 65°C. For each mutation, two primers are generated: a forward and a reverse primer, each containing the desired mutation centered in the middle of the primer sequence. To generate single mutants, we set up PCRs combining each forward SDM primer with a universal reverse primer that anneals to the plasmid backbone, and each reverse SDM primer with a universal forward primer that is the reverse complement of the universal reverse primer. The PCRs contained 6.25 pL Platinum Superfi II Mastermix, 0.5 pL each primer at lOuM, 1 pL plasmid template DNA at 1 ng / pL, and water up to 12.5 pL. The PCRs were run following the standard Platinum Superfi II Mastermix protocol, but with the annealing temperature set to 65°C. The PCRs were then Ampure XP bead cleaned at 0.5x. 1 pL of each PCR for the desired mutation was combined with 5 pL gibson mastermix and 3 pL of water, and gibson assembled at 50°C for 15 minutes before transformation into e.coli and plating. For cloning two or more mutations simultaneously, the universal primers were replaced with the other mutation’s forward and reverse primers, so that installing two mutations required a two piece gibson, three mutations required a three piece gibson etc.Results:
[0282] Informed by the efficiency and specificity data of LSR orthologs in Durrant et al. 2022, we chose Dn29 as the engineering scaffold. This ortholog was functional for integration into a few primary genomic pseudosites, with 62% of integrations occurring in the top 5 sites, and with an overall integration efficiency of ~5% into wildtype cells. The top pseudosite, termed attHl, receives around 38% of all integrations, so it was chosen as our on- target site for engineering.
[0283] We next developed a bacterial selection system to enrich functional LSR variants. Adapted from Buchholz et al 2001 and Lansing et al. 2020, we employed substrate-linked directed evolution to evolve Dn29 to improve recombination of attP and attHl. In this system, a Dn29 variant library is expressed from an L-arabinose-inducible promoter. The plasmid backbone contains the two attachment sites flanking a Ndel restriction enzyme digest site (Figure 1 and 2). Electroporation of this library into e.coli and induction with arabinose will lead to effector expression, and functional protein variants will recombine the two attachment sites, removing the Ndel digest site from the plasmid. Digestion of the library with Ndel and recovery of Dn29 variants with primers that span the digest site result in PCRamplification of only functional variants, which can be cloned back into the SLIDE backbone for subsequent rounds of evolution.
[0284] To systematically explore the mutational landscape and enhance our understanding of protein function, a comprehensive deep mutational scanning library was generated. Employing NNK primers for each residue in the 519 amino acid protein, we employed overlap extension to generate every possible single amino acid mutation variant of Dn29 (Figure 3). Through NGS, we determined that this method generated a high diversity library with evenly dispersed mutations, with only 10 positional dropouts. (Figure 4 and 5) Additionally, interspersed between evolution cycles, we introduced new mutation combinatorial diversity by performing DNA shuffling. We evaluated different L-arabinose conditions, noting that we saw recombination at both 10 pg / mL and 0 pg / mL, likely due to leaky expression from the pBAD promoter. Consequently, we used 10 pg / mL L-arabinose for the first 7 rounds of selection and 0 pg / mL L-arabinose for the last 5 rounds, aiming to progressively heighten the selective pressure over time (Figure 6).
[0285] After twelve cycles and two rounds of shuffling, we again sequenced the library to explore the mutational landscape. The resulting library displayed high rates of mutation dropout and a reshaped distribution of mutations across the protein (Figure 7, 8, and 9). We calculated an enrichment score based on the frequency a mutation occurs in the output library normalized to its frequency in the input library to visualize which specific mutations are being enriched over the evolution time course (Figure 10 and 11). Promisingly, we noted that enrichment scores increase as the evolution progresses, and there is a depletion of both stop codons and mutations of the catalytic serine, indicating that our selection process is depleting non-functional mutants. However, since our primary objective was to identify mutations to engineer Dn29 in the context of human genome manipulation, it was imperative to validate whether these mutations indeed imparted advantageous traits within a mammalian cell environment.
[0286] To validate whether the library of enriched variants generated mutants with improved activity in mammalian cells, we cloned the library at various stages of cycling into a mammalian expression vector, picked colonies, and transfected them into HEK293FTs with an attP-Efla-mCherry donor plasmid. Using ddPCR, we read out the percent integration into attHl and attH3, a common off-target site. After 5 cycles, there was only one mutant with on- target efficiency above wildtype, but as we performed more DNA shuffling and cycling we were able to generate tens of mutants with improved efficiency. (Figure 12 and 13) As aproxy for specificity, we calculated the ratio between attHl and attH3 integration. Although we found that high efficiency doesn't correlate with high specificity, we were able to identify numerous mutants with relatively reduced attH3 integration (Figure 13). Given the number of mutants we were able to identify that confer improved integration properties, we transitioned away from the library -based directed evolution method and towards identifying driver mutations from the assayed clones to rationally design combinatorial mutants.
[0287] Our approach involved identifying a top-performing mutant and incorporating all possible single mutations from any mutant exhibiting 1.5 times the wildtype efficiency or 2x wildtype attHl / attH3 ratio. From the arrayed ddPCR assay, we identified clone 62 (M6I / E70G / A224P / G227V / A234), which had 2.5x WT activity at attHl, and clone 93 (A318 / N341K) which had an 8x improved on / off target ratio. Combining all 7 mutations, we generated clone 127, our new top-performing mutant, which demonstrated 3x wildtype efficiency and 13x attHl / attH3 specificity (Figure 15). Using SDM primers and gibson cloning, we generated 105 new mutants, with one for each mutation in the set of beneficial mutants. Given residue 341 seemed to be a critical nucleotide for conferring specificity, we also generated all possible residue mutations at 341. Finally, we included in this set reversion mutations of all of the clone 62 mutations back to wildtype. We transfected these clones and measured attHl and attH3 integration with ddPCR, and identified 24 mutations that drive efficiency and 11 mutations that drive specificity (Figure 16 and 17, Table 1).
[0288] Mapping these mutations to the Dn29 Alphafold structure provides insight into how these mutations may affect protein function. Many of the specificity driver mutations occur in the zinc-ribbon domain and recombinase domain, which are the main DNA binding regions. (Figure 18). The efficiency driver mutations are dispersed throughout the protein, with a notable cluster of mutations occurring in the coiled-coiled domain. As this region is not thought to be DNA interacting, these mutations may be affecting the protein conformation, stability, or protein dimerization / tetramerization. (Figure 19) We identified residue 341 appears to be a key residue for protein function, with many amino acids at that position either greatly enhancing efficiency or specificity. (Figure 20)
[0289] We next began progressive rounds of stacking single mutations on top of the best clone from the previous round. In general, combining more efficiency mutations increased overall efficiency at attHl but also increased integration rates into attH3. Adding specificity variants decreased attH3 integration, but often at the expense of attHl efficiency. Over 3 rounds of combinations, we generated variants that combined various specificity andefficiency mutations to try to optimize for both characteristics (Figure 21, 22, 23). We additionally screened colonies for off-target integration into a second off target site, attH chrlO, noting that two specificity mutations tested, I303K and K248R, only seemed to decrease attH3 integration but not attH chrlO, while two other specificity mutations tested, Q332K and I233K, reduced both attH3 and attH chrlO integration (Figure 24). We identified two highly efficient variants, 508 and 515, with 25% integration efficiency into attHl, over 8-fold wildtype Dn29 (Figure 25). These variants had similar specificity as wildtype Dn29, with -35% of integrations occurring at attHl and a long tail of very rare integration sites (Figure 25, Figure 26). We generated two specificity focused variants, 538 and 570, which increased on-target specificity up to 45%, while also increasing efficiency up to 3-5 fold at attHl (Figure 27, Figure 28).
[0290] We’ve previously shown that Dn29-dCas9 fusions, targeted via gRNA to a genomic pseudosite, increases the integration efficiency and specificity to that site. These new variants in a dCas9 fusion system showed enhanced attHl efficiency compared to wildtype Dn29-dCas9, although to a lesser fold change than that seen in the non-fused context (Figure 29). These variants also improved Dn29-dCas9 specificity, increasing the percent of integrations at attHl from 60% (WT Dn29) to 73% (Variant 538) (Figure 30).
[0291] In subsequent rounds of mutation combinations, we generated Variant 637, which contains all the mutations in 538 plus 4 more specificity mutations (Figure 32). This variant exhibited the most specific integration profile, with 84% of integrations occurring at attHl and comparable efficiency to wildtype Dn29 (Figure 33). Additionally, we generated another highly efficient variant, 608, which is equivalent to variants 515 and 508 in attHl efficiency, but with 20-fold reduced integration into attH3. (Figure 31).References:
[0292] Bi, Chuyun, et al. "A python script to design site-directed mutagenesis primers." Protein Science 29.4 (2020): 1040-1045.
[0293] Buchholz, F., & Stewart, A. F. (2001). Alteration of Cre recombinase site specificity by substrate-linked protein evolution. Nature biotechnology, 79(11), 1047-1052.
[0294] Cong, L., Ran, F. A., Cox, D., Lin, S., Barretto, R., Habib, N., ... & Zhang, F. (2013). Multiplex genome engineering using CRISPR / Cas systems. Science, 339(6121), 819- 823.
[0295] Durrant, M. G., Fanton, A., Tycko, J., Hinks, M., Chandrasekaran, S. S., Perry, N. T., ... & Hsu, P. D. (2023). Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome. Nature Biotechnology, 41(4), 488-499.
[0296] Jinek, M., East, A., Cheng, A., Lin, S., Ma, E., & Doudna, J. (2013). RNA- programmed genome editing in human cells, elife, 2, e00471.
[0297] Lansing, F., Paszkowski-Rogacz, M., Schmitt, L. T., Schneider, P. M., Rojo Romanos, T., Sonntag, J., & Buchholz, F. (2020). A heterodimer of evolved designerrecombinases precisely excises a human genomic DNA locus. Nucleic Acids Research, 48( ), 472-485.
[0298] Lampe, G. D., King, R. T., Halpin-Healy, T. S., Klompe, S. E., Hogan, M. I., Vo, P. L. H., ... & Sternberg, S. H. (2023). Targeted DNA integration in human cells without double-strand breaks using CRISPR-associated transposases. Nature Biotechnology, 1-12.
[0299] Mali, P., Yang, L., Esvelt, K. M., Aach, J., Guell, M., DiCarlo, J. E., ... & Church, G. M. (2013). RNA-guided human genome engineering via Cas9. Science, 339(6121), 823- 826.
[0300] Matreyek, K. A., Stephany, J. J., & Fowler, D. M. (2017). A platform for functional assessment of large variant libraries in mammalian cells. Nucleic acids research, 45(11), el02-el02.
[0301] Muller, K. M., Stebel, S. C., Knall, S., Zipf, G., Bemauer, H. S., & Arndt, K. M. (2005). Nucleotide exchange and excision technology (NExT) DNA shuffling: a robust method for DNA fragmentation and directed evolution. Nucleic acids research, 33(13), el 17- el l7.Thyagarajan, B., Olivares, E. C., Hollis, R. P., Ginsburg, D. S., & Calos, M. P. (2001). Site-specific genomic integration in mammalian cells mediated by phage cpC31 integrase.Molecular and cellular biology, 27(12), 3926-3934.
[0302] Tou, C. J., Orr, B., & Kleinstiver, B. P. (2023). Precise cut-and-paste DNA insertion using engineered type VK CRISPR-associated transposases. Nature Biotechnology, 1-12.
[0303] Van Duyne, G. D., & Rutherford, K. (2013). Large serine recombinase domain structure and attachment site binding. Critical reviews in biochemistry and molecular biology, 48(5), 476-491.EXAMPLE 2
[0304] Technologies for precisely inserting large DNA sequences into the genome are critical for diverse research and therapeutic applications. Large serine recombinases (LSRs)can mediate direct, site-specific genomic integration of multi-kilobase DNA sequences without a pre-installed landing pad, but current approaches suffer from low insertion rates and high off-target activity. Here, we present a comprehensive engineering roadmap for the joint optimization of DNA recombination efficiency and specificity. We combined directed evolution, structural analysis, and computational models to rapidly identify additive mutational combinations. We further enhanced performance through donor DNA optimization and dCas9 fusions, enabling simultaneous target and donor recruitment. Top engineered LSR variants achieved up to 53% integration efficiency and 97% genome-wide specificity at an endogenous human locus, and effectively integrated large DNA cargoes (up to 12 kb tested) for stable expression in challenging cell types, including non-dividing cells, human embryonic stem cells, and primary human T cells. This blueprint for rational engineering of DNA recombinases enables precise genome engineering without the generation of double-stranded breaks.A framework for recombinase engineering to enable site-specific genome insertion
[0305] We next assessed all point mutations from variants with 1.5-fold WT efficiency (n = 47 mutations) and 2-fold WT specificity (n = 28 mutations). We also included mutations at putative DNA-interfacing residues, chosen based on alignment with the crystal structure of Listeria integrase C-terminal domain bound to attP (PDB: 4KIS)23, hypothesizing that positively charged mutations could modify DNA binding (Fig. 36). These mutations were individually installed into variant 127, identifying 12 additional efficiency and 7 additional specificity driver mutations, each contributing 1.2 to 2.5-fold efficiency and 1.1 to 6.8-fold specificity improvements over variant 127 (Fig. 35). A final round of mutation validation (Fig. 37) yielded a final list of 21 efficiency and 12 specificity driver mutations for further rational engineering (Table 1).Computational modeling and structural analysis of recombinase mutation stacking
[0306] We demonstrated that LSR directed evolution followed by classification across our efficiency and specificity metrics enables iterative mutational combination to generate LSRs with a desired functional profile. To expedite experimental testing of larger numbers of variants, we sought to develop a computational model for predicting combinatorial variant activity from single mutant data. The variants were divided into distinct rounds, with the individual mutation validation in round 1 (Fig. 35) and the iterative mutation stacking experiments comprising rounds 2-5. We trained two linear models (linear regression, ridgeregression) and two gradient boosting models (XGBoost, CatBoost) on one-hot encoded variant sequences from round 1, and then tested these models on the data from rounds 2-5.
[0307] From a dataset of individual mutations, measured for specificity and efficiency by ddPCR, we generated ridge regression models and examined their coefficients to quantify the impact of key mutations. By predicting efficiency and specificity activity of combinations of key mutations, we can prioritize mutants for experimental testing. To demonstrate this approach, we predicted the activities of all double mutants combined on top of superDn29, skipping round 6 to directly test round 7 of iteration. We tested the top 10 efficiency and top 3 specificity variants for genome integration, and found that 8 / 10 efficiency variants and 3 / 3 specificity variants performed better than superDn29 (Fig. 38).Unifying engineering strategies for maximal LSR efficiency and specificity
[0308] To better understand the single cell variation of insertional mutagenesis, including on / off-target co-occurrence and integration copy number, we mapped integrations in ~50 clonal HEK293FT populations edited with hifiDn29 or WT Dn29, with and without dCas9 fusions (Fig. 47). dCas9 fusion dramatically improved performance for both Dn29 and hifiDn29, resulting in over 95% of clones containing on-target edits. However, the most striking difference was observed in off-target insertions - 91% of Dn29 clones contained off- target insertions, compared to 46% of hifiDn29 clones. Ultimately, with dCas9 fused to hifiDn29, off-target insertions were reduced to only 9% of clones, compared to 52% for dCas9-Dn29.
[0309] Beyond targeting accuracy, measuring the number of integrations per cell is crucial for fully understanding editing outcomes. Because dCas9 increases efficiency, it also increases the rate of multiple on-target insertions. HifiDn29 showed the highest rate of single on-target insertion events, with half of the clones exhibiting this genotype. In contrast, including the dCas9 fusion decreased the rate of single on-target insertion events to 38%, and increased the rate of multiple on-target insertions from 4% to 53% of clones (Fig. 48). Due to HEK293FT’s pseudo-triploid genome and copy number variation / instability 33, on-target insertions ranged from 0-5 per cell, with a median of 2 for both hifiDn29-dCas9 and Dn29- dCas9 clones (Fig. 49). Overall, these single cell results demonstrate hifiDn29-dCas9's enhanced precision and efficiency, nominate hifiDn29 for generating clonal cell lines containing single on-target integrations, and highlight the value of single-cell analysis in evaluating gene editing outcome heterogeneityEngineered LSR systems insert multi-kilobase DNA cargo into the genome of nondividing cells, human embryonic stem cells, and primary T cells
[0310] Next, we sought to benchmark our engineered LSRs in a diverse set of genome insertion tasks across non-dividing cells and dividing primary cells. LSRs offer an advantage in genome engineering applications involving non-dividing cells due to their independence from DNA repair machinery and homologous recombination 34. We treated HEK293FT cells with aphidicolin to induce cell cycle arrest and observed that Dn29 and key variants showed largely equivalent integration rates to untreated cells. dCas9 fusions experienced decreased integration efficiency but still achieved up to 30% on-target integration (Fig. 39).
[0311] We then tested DNA cargo installation in Hl human embryonic stem cells (hESCs) and observed that on-target integration efficiency with engineered recombinases increased up to 6-fold relative to wildtype Dn29, from 4.1% to 24.5% (Fig. 40). These improvements are 25-50% of the integration efficiencies seen in HEK293FT, likely due to the 50-80% reduction in plasmid transfection efficiency in stem cells. Insertion specificity also significantly improved, as off-target integration at attH3 for the engineered variants approached the ddPCR detection limit (Fig. 50).
[0312] To evaluate our optimized LSR variants' capacity for larger cargo installation and enable functional genomics applications, we designed a large 12kb CRISPRi construct encoding for the dCas9-ZIM3 fusion and multiple regulatory elements and marker genes, including a BFP, neomycin resistance marker, Woodchuck post-transcriptional regulatory element (WPRE), and ubiquitous chromatin opening element (UCOE). We achieved robust insertion efficiencies up to 13% with standard lipid transfection (Fig. 41).
[0313] After neomycin selection, -60% of the bulk population was BFP+ by flow cytometry (Fig. 42), indicating successful CRISPRi construct integration and expression. Clonal analysis of the goldDn29-dCas9-edited hESCs demonstrated that the majority (84%) of the clones were 100% BFP+, while some clones had significant levels of BFP silencing at the stem cell stage (Fig. 51). Genotyping revealed that 95% of clones possessed precise heterozygous attHl insertions, with only one homozygous clone and one clone showing both on- and off-target integrations (Fig. 43), These results underscore the high clonal consistency achieved, demonstrating a predominance of accurate, single-copy integrations and strong maintenance of cargo gene expression in Hl stem cells.
[0314] Transgene silencing is a persistent challenge in stem cell engineering, likely due to the extensive chromatin restructuring that occurs during stem cell differentiation 35. Toassess whether attHl supports stable cargo expression during differentiation, we differentiated the LSR-dCas9-edited hESCs into hematopoietic progenitor cells (HPCs) under neomycin selection. Post-differentiation, edited HPCs maintained robust cargo expression (-70% BFP+) with -80% of cells expressing canonical HPC markers such as CD34 and CD43 (Fig. 42). Next, we introduced sgRNAs targeting CD63, CD81, and CD147 via lentiviral transduction into the goldDn29-dCas9-edited HPCs and observed 91-98% knockdown of these cell surface markers compared to a non-targeting guide (Fig. 44). Taken together, we demonstrate that engineered recombinases efficiently produce bulk hESC lines with near-clonal genotypes and stable large cargo expression throughout hESC-to-HPC differentiation, making them suitable for generating stable cell lines for CRISPR screens or potentially introducing therapeutic genetic cargoes.
[0315] Site-specific transgene insertions show great promise in immune cell engineering, enabling integration of functional cargos like chimeric antigen receptors (CARs) and additional immune regulators for therapeutic applications. However, high plasmid DNA toxicity in primary T cells limits use of conventional donor and effector expression plasmids. To address this, we electroporated T cells with effector mRNA and delivered donor templates via single-stranded AAV (ssAAV) or self-complementary AAV (scAAV) expressing mCherry (Fig. 45, Fig. 52). LSR-dCas9 fusions achieved up to 17% integration efficiency into attHl using AAV donors, which maintained high cell viability even at the highest doses. The scAAV donor yielded 2.3 to 3.3-fold higher integration rates compared to the ssAAV donor, likely due to the requirement of double-stranded DNA for LSR-mediated integration. The efficiency and viability we observe strongly support further development of this approach for engineering primary T cells.
[0316] Finally, we investigated the cross-reactivity of our engineered LSRs across model organisms, including mice and various non-human primates. An attHl -like sequence is present in the NEBL intron in marmosets, rhesus monkeys, and cynomolgus monkeys, with 1-2 point mutations compared to attHl, and is located intergenically in the mouse X chromosome with 6 point mutations. Dn29 and goldDn29 could robustly recombine attP with the model organism pseudosites in a plasmid recombination assay (Fig. 53), enabling future advancement in preclinical animal studies that bridge the gap between laboratory research and human clinical trials.Discussion
[0317] Here, we report a framework that combines directed evolution, protein engineering, and ML models for engineering DNA recombinases to efficiently and specifically insert large genetic cargos directly into the human genome, overcoming the need to preinstall attachment site landing pads. As a proof-of-concept, we report integration into a single genomic locus using optimized Dn29 LSRs, achieving a 13- to 17-fold improvement in insertion efficiency. Combining mutants with dCas9 fusions and optimized donor sequences (e-attP and sgRNA target sites) yielded recombinases with 40-53% efficiency and 90-97% genome-wide specificity for an endogenous locus.
[0318] Our engineering efforts provide numerous mechanistic insights into LSR function during genome integration. Notably, our dCas9 fusion experiments demonstrate that improved genome search and DNA binding are crucial areas for increased integration efficiency. To further interrogate DNA binding, we employed structural modeling and attachment site screening to identify specific protein and DNA regions critical for target recognition. Directed evolution revealed an inherent trade-off between efficiency- and specificity-improving mutations, which we overcame by strategically pairing mutations across distinct LSR domains and increasing the protein-DNA interface through fusions with DNA-binding proteins.
[0319] Our current system incorporates CRISPR components, which include both protein (dCas9) and RNA (sgRNA) elements. Although this design improves efficiency by 7-fold and specificity by 5-fold, it also increases the overall size of the system and introduces an additional RNA component. CRISPR components could be replaced with smaller, protein- only DNA binding domains such as zinc fingers. Such modifications would preserve the delivery advantages of these compact recombinases and maintain a streamlined system of a single protein and single DNA donor.
[0320] Our study presents multiple orthogonal engineering strategies to enhance an LSR's ability to recognize and integrate at endogenous genomic sequences. We demonstrate the generalizability of these approaches beyond Dn29 to Nm60, improving on-target genomic insertion efficiency to 73%.
[0321] These advancements have utility across diverse research and therapeutic applications of LSRs. The current paradigm for functional genomics employs lentiviral engineering of cell lines and single copy installation of pooled libraries, which can lead to unpredictable effects on gene expression, potential insertional mutagenesis, and silencing of transgenes. Our engineered recombinases overcome these limitations by targeting a definedintegration locus at high efficiencies, which are essential for large-scale and uniform functional genomics studies. In hESCs, we demonstrate that 95% of cells have single copy, on-target insertions, enabling the generation of homogenous bulk cell populations without single-clone selection. Furthermore, we previously demonstrated the utility of LSRs for virus-free library screening in landing pad cell lines. Our work described herein extends this capability, showing the feasibility of integrating guide or protein libraries directly into an endogenous human genomic locus. In hESCs, these integrations occur at copy numbers comparable to low multiplicity of infection (MOI) lentivirus (MOI=0.1), well within the standard guidelines for approximating one integrant per cell.
[0322] In the therapeutic space, our approach offers advantages over prevailing CRISPR- based gene therapies, which require a new guide RNA to target a distinct disease-causing mutation. By contrast, multi-kilobase insertions enable replacement of entire corrective open reading frames, providing a "one size fits all" approach for correcting genetic diseases with mutational heterogeneity across patient populations. Additionally, these corrective transgenes can include critical non-coding regulatory elements for enhanced control of gene expression. Furthermore, Dn29 can cross-reactively integrate into attHl-like sequences in diverse model organisms, an important consideration for future IND-enabling studies.
[0323] The strategies outlined in this work can be adapted to target diverse genomic loci beyond Dn29 attHl. To target a different locus, such as validated genomic safe harbors like AAVS1 or therapeutic targets like TRAC, LSRs can be mined from the thousands of naturally occurring orthologs to find a recombinase with a closer match to the desired target sequence. These candidate LSRs can then be subjected to our joint optimization approach, combining directed evolution, machine learning predictions, and DNA-binding protein fusions to enhance both efficiency and specificity. Overall, the distinct LSR engineering strategies presented in this study provide a comprehensive blueprint for the development of next-generation recombinases capable of direct, site-specific genome integrations.MethodsCell lines and cultureExperiments were conducted in HEK293FT cells (Thermo Fisher), Hl human embryonic stem cells (hESCs), and primary human T cells (STEMCELL Technologies) from deidentified healthy donors (Cat #200-0092). HEK293FT cells were cultured in DMEM with 10% FBS (Gibco) and lx Penicillin-Streptomycin (Thermo Fisher), and dissociated using TrypLE Express (Gibco). Hl hESCs were maintained in MTeSR Plus (STEMCELLTechnologies) supplemented with lx Antibiotic-Antimycotic (Thermo Fisher) and cultured on Cultrex (Bio-Techne) or Matrigel (Coming) coated plates. For routine passaging, hESCs were dissociated with ReLeSR (STEMCELL Technologies). For 96-well plating prior to transfections, single-cell dissociation was performed using Accutase (STEMCELL Technologies). Hl hESCs were supplemented with 10 pMRock inhibitor for 24 hours postdissociation. Primary human T cells were cultured in complete X-VIVO 15 (cXVIVO 15) (Lonza Bioscience, Visp, Switzerland #04-418Q) which consists of 5% FCS (R&D systems, Cat #M19187), 5 ng pl-1 IL-7, and 5 ng pl-1 IL-15. Dn29 deep mutational scan library construction
[0324] An NNK deep mutational scanning library of the entire Dn29 CDS was generated using NNK oligos and overlap extension PCRs. First, forward and reverse oligos with NNK mixed bases at each codon were designed with a Tm of 65°C. Each NNK forward primer was paired with Dn29 DMS universal reverse that binds downstream of the CDS, and each NNK reverse primer with DMS universal forward primer, generating amplicons flanking the mutated codon. PCR reactions contained 2.5 pL Q5 Mastermix (NEB), 0.01 pL Dn29 plasmid template (100 ng / pL), 0.025 pL universal primer (100 pM), 1.465 pL water, and 1 pL unique NNK primer (2.5 pM). Cycling conditions: 98°C for 30 seconds; 30 cycles of 98°C for 10 seconds, 60°C for 30 seconds, and 72°C for 1 minute; final extension of 72°C for 2 minutes.
[0325] Upstream and downstream amplicons (2.5 pL each) were pooled and cleaned with 2 pL ExoSAP -IT (Thermo Fisher) and 0.5 pL of Dpnl (NEB), incubating at 37°C for 30 minutes then 80°C for 15 minutes. For the overlap extension PCR, 1 pL cleaned PCR pool was mixed with 2.5 pL Q5 2x Mastermix, 0.025 pL of each universal primer (100 uM), and 1.45 pL of water, using the same cycling conditions
[0326] The full mutant pool was created by combining 2.5 pL of each overlap extension PCR. The full-length Dn29 fragment was gel-extracted (Monarch DNA Gel Extraction Kit, NEB). The library and pEVO backbone were digested with Xbal and Hindlll-HF (NEB). Ligation used 100 ng total DNA(3:1 molar ratio of library to backbone), 2 pL T4 ligase (NEB), 4 pL lOx T4 ligase buffer (NEB), and water to 40 pL. The reaction was split into two 20 pL reactions, ligated for 30 minutes at room temperature, inactivated at 65°C for 10 minutes, and purified (Clean and Concentrator-5 Kit, Zymo).
[0327] The ligation product was electroporated into XL-1 Blue cells (Agilent) according to the manufacturer’s instructions, recovered for 1 hour at 37°C in 1 mL SOC media, andplated onto four 245 mm x 245 mm Bioassay dishes. Approximately IM colonies were obtained. Plasmids were purified using NucleoBond Xtra Midi EF kit (Macherey Nagel) and sequenced with Illumina NextSeq2000 600 cycle Pl kit .Substrate-linked directed evolution
[0328] Library transformation, induction, and growth: 4 pL pEVO plasmid library was electroporated into 50 pL XL-1 Blue competent cells (Agilent), recovered in 1 mL SOC media (37°C, 1 hour), and then seeded into 100 mL LB media with carbenicillin and L- arabinose (10 pg / mL or 0 pg / mL). Cultures were grown overnight at 37°C. Library coverage (>1M colonies) was confirmed by plating serial dilutions. Plasmids were extracted using Qiagen Plasmid Midi kit (0.3 g wet bacteria pellet per column).
[0329] Selection of active variants: 500 ng plasmid was digested with Ndel (NEB) to eliminate inactive variants. Active variants were amplified using: 25 pL 2x Platinum Superfi II Mastermix (Thermo Fisher), 19 pL water, 2 pL each SLiDE recovery forward and SLiDE recovery reverse primers (10 pM), and 2 pL NdeLdigested material. PCR conditions: 98°C for 30 seconds; 30 cycles of 98°C for 10 seconds, 52°C for 10 seconds, 72°C for 55 seconds; final extension at 72°C for 5 minutes. The correct-size band was gel- extracted (Monarch DNA Gel Extraction Kit, NEB).
[0330] Cloning for next evolution cycle: Amplified active variants and pEVO backbone were digested with Xbal and Hindlll-HF (NEB) at 37°C for 30 minutes, then heat-inactivated at 80°C for 20 minutes. Digested variants were purified using DNA Clean and Concentrator- 5 (Zymo), and backbone with DNA Clean and Concentrator-25 (Zymo). Five ligation reactions (20 pL each) were set up using 100 ng DNA (3: 1 ratio of library to backbone) and T4 ligase (NEB). Ligation occurred at room temperature for 30 minutes, followed by heat inactivation at 65°C for 10 minutes. Pooled reactions were purified (DNA Clean and Concentrator-5 Kit, Zymo), eluted in 6 pL water, and electroporated into XL-1 Blue cells to start the next evolution cycle.DNA shuffling and fragment reassembly
[0331] Shuffling the active variants between rounds of cycling involved a uridine exchange PCR to partially exchange thymidines for uridine, USER enzyme fragmentation at uridine sites, primerless PCR fragment reassembly, and PCR for full-length gene recovery.
[0332] Uridine exchange PCR: Fragment size and yield was optimized by modifying dUTP / dTTP ratio, with the optimal ratio being 3 / 7. PCR mixture: 5 pL lOx Thermopol Buffer, 1 pL lOmM dNTPs, 1 pL each SLiDE recovery forward andSLiDE recovery reverse primers (10 pM), 1 pL plasmid library, 1 pL Taq Polymerase, and 40 pL water. Cycling conditions: 95°C for 30 seconds; 30 cycles of 95°C for 20 seconds, 60°C for 30 seconds, 68°C for 1 min / kb; final extension at 68°C for 5 minutes. Full-length gene band was gel-extracted (Monarch Gel Extraction Kit, NEB). USERenzyme digestion: 500 ng aliquots were digested with 2 pL USER Enzyme (NEB) at 37°C for 3 hours. Gel electrophoresis confirmed fragment distribution (lOO-lOOObp).
[0333] Fragment reassembly: Fragments were purified (DNA Clean and Concentrator-5, Zymo) and reassembled in a Primerless PCR reaction using the following conditions: 25 pL purified fragments, 25 pL 2x Q5 High Fidelity Master Mix (NEB). Cycling conditions: 98°C for 30 seconds; 30-50 cycles of 98°C for 10 seconds, 30°C for 30 seconds (+l°C / cycle), 72°C for 1 minute (+4s / cycle); final extension at 72°C for 10 minutes. A final PCR was performed to recover only the full length Dn29 CDS for further rounds of directed evolution. The following conditions were used for full-length gene recovery: PCR mixture: 25 pL Platinum Superfi II 2x Mastermix (Thermo), 10 pL reassembled fragments, 2 pL eachDMS universal forward and DMS universal reverse primers (10 pM), and 11 pL water. Cycling conditions: 98°C for 30 seconds; 35 cycles of 98°C for 10 seconds, 60°C for 10 seconds, 72°C for 55 seconds; final extension at 72°C for 5 minutes.
[0334] The gel-extracted, shuffled, and reassembled genes were cloned into the plasmid backbone using Xbal and Hindlll digest and T4 ligation as previously described.Variant library next generation sequencing and analysis
[0335] Six primer sets (DMS_NGS primers) were designed to amplify -260 bp segments of the Dn29 CDS with Illumina adapter overhangs. Two rounds of PCR were performed to add P5 / P7 adapters and i5 / i7 indexes (FLAP2 primers). Amplicons were cleaned with Ampure XP beads (Beckman Coulter) between PCR rounds and after the final PCR.Amplicons were pooled in equimolar ratios, quantified using Qubit dsDNA High Sensitivity Kit (Thermo Fisher), and sequenced on Illumina NextSeq2000 (600 cycle kit). Full overlap between read 1 and read 2 was ensured for higher confidence in mutation calling.
[0336] Paired-end reads were merged using BBMerge (v.39.06) and analyzed with a custom Python script. The script converted Phred Quality scores to error probabilities using the formula P = 10(Q / -10), where P is the probability of error and Q is the Phred Quality Score. Reads with a summed error probability greater than 0.5 or containing frameshifts were filtered out. Nucleotide and amino acid mutations at each position were then counted and plotted. Enrichment for each amino acid (AA) between the input and output libraries wascalculated using the formula: ((%AAoutput) / (l-%AAoutput)) / ((%AAinput) / (l-%AAinput)). To distinguish library construction-based dropouts from selection-based dropouts in the enrichment heat maps, any amino acids with zero reads in the output library were assigned a single read.Nanopore sequencing and analysis
[0337] Variants were cloned into a vector containing a 100 nucleotide random UMI barcode with a BHVDrepeat pattern. The plasmid library was linearized by Ecol05I digestion. Nanopore libraries were prepared using the barcoded nanopore sequencing kit (SQK-NBD114.24) with 1 pg linearized plasmid library and sequenced on a MinlON flow cell (RIO.4.1) for 72 hours. Sequencing reads were filtered using nanoq (v.0.9.0) with settings— min len 4500— max len 5500— min-qual 10 (i.e. minimum q score of 10, a minimum read length equivalent to 90% of the expected read length, and a maximum read length equivalent to 110% of the expected read length). The UMI sequence was extracted using cutadapt (v.1.18) with settings -g "GGCGGTCACCATCACCACCACCACGCTACACG;max_error_rate=0.2...ACTGTAC;ma x error _rate=0.2"— trimmed-only— revcomp— minimum length 95. All UMI sequences were trimmed to 95 nt using seqkit (v.1 ,3-rl06) with the command seqkit subseq-r 1 :95.
[0338] Reads were clustered by UMI with mmseqs easy-linclust (v.14.7e284) with setting— min-seq-id 0.5. For each UMI cluster bin with at least 15 reads, a representative cluster sequence was generated by using usearch (v.l 1) with settings-cluster_fast-id 0.75- strand both-sizeout-centroids and taking the first representative sequence of the output 40. A final consensus sequence was generated by 1 round of polishing with Medaka (v.1.9.1) with settings-m rl041_e82_260bps_hac_g632. Counts for each unique variant were determined by tallying the total consensus sequences.Cloning variant library into a mammalian expression vector
[0339] Primers (DE mammalian forward and DE mammalian reverse) were designed to amplify the Dn29 CDS from the active variant PCR library, adding overhangs for Esp3i- compatible Golden Gate cloning. PCR conditions: 25 pL 2x Platinum Superfi II Mastermix (Thermo Fisher), 19 pL water, 2 pL purified active variant library, 2 pL each primer. Cycling: 98°C for 60 seconds; 30 cycles of 98°C for 10 seconds, 60°C for 10 seconds, 72°C for 55 seconds; final extension at 72°C for 5 minutes. The product was purified (DNA Clean and Concentrator-5, Zymo) and quantified by Nanodrop.
[0340] A mammalian expression vector was designed with the EFla promoter upstream of an Esp3i golden gate landing pad, used as the destination for the protein variant library. The landing pad was followed by a T2A self-cleaving peptide sequence and an EGFP CDS.
[0341] Golden Gate reaction mixture: 75 ng mammalian expression vector, amplified variant library (3 : 1 molar ratio to vector), 1 pL T4 DNA Ligase Buffer (NEB), 0.5 pL T4 DNA Ligase (NEB), 0.5 pL Esp3i (ThermoFisher), and up to 10 pL nuclease-free water. Cycling: 35 cycles of 37°C for 1 minute, 16°C for 1 minute; 37°C for 30 minutes; 80°C for 20 minutes. Five Golden Gate reactions were performed, pooled, and purified (DNA Clean and Concentrator-5, Zymo). The library was transformed into Maehl E. coli and plated for overnight growth. Random colonies were picked, grown in 4 mL TB-Carbenicillin, and miniprepped (NucleoSpin Plasmid Transfection Grade Mini kit, Machery-Nagel).Transfection of HEK293FTs for assessing genomic integration
[0342] One day before transfection, 12-18K HEK293FT cells were plated per well of a 96-well plate, aiming for 60%-80% confluency at the time of transfection.
[0343] Standard LSR + donor transfection, for transfections containing an LSR effector plasmid and a donor plasmid, each well was transfected with 725 ng of DNA, containing a 5: 1 molar ratio of donor plasmid to effector plasmid, using 0.5 pL of Lipofectamine 2000 (Thermo) per well.
[0344] Standard LSR-dCas9 + donor + guide transfection: LSR-dCas9 effector plasmid, donor plasmid, and guide plasmid were transfected with 725 ng total DNA, containing a 5:1 : 1 molar ratio of donor:effector:guide plasmid with 0.5 pL of Lipofectamine 2000 per well, unless specified otherwise in figure legends.
[0345] The cells were incubated and monitored for three days for mCherry (donor plasmid) and GFP (effector plasmid) expression. Cells were then harvested for flow cytometry (Attune NxT Flow Cytometer, Thermo Fisher) or genomic DNA extraction for downstream analyses.Cell Harvest, ddPCR, qPCR and flow cytometry
[0346] Three days post-transfection, cells were trypsinized with 50 pL TrypLE (Gibco) for 10 minutes and then quenched with 50 pL Stain Buffer (BD). The 100 pL cell suspension was split into two 50 pL aliquots in U-bottom 96-well plates, centrifuged (300 x g, 5 minutes), and supernatant aspirated. One plate was resuspended in 200 pL Stain Buffer (BD) and analyzed with Attune NxT Flow Cytometer with Autosampler (Thermo Fisher).
[0347] The other plate was resuspended in 50 pL QuickExtract DNA Solution (Biosearch Technologies), vortexed for 15 seconds, and thermocycled: 65°C for 15 minutes, 68°C for 15 minutes, 98°C for 10 minutes. DNA was cleaned with 0.9x AmpureXP (Beckman Coulter) beads.
[0348] To assess integration efficiency and specificity, qPCR / ddPCR primers and probes were designed to span the left integration junction of attHl and attH3, using a constant primer that binds to the donor plasmid sequence (ddPCR donor reverse l) a genome binding primer near the pseudosite (ddPCR attHl forward l, ddPCR_attH3_forward) and a FAM probe within the amplicon (ddPCR_attHl_probe_l, ddPCR_attH3_probe). For attHl, a second set of primers / probes was designed to target the right junction to verify measurement accuracy (ddPCR_attHl_2 primers / probe). Genomic reference primers and probes located nearby each attachment site were designed to measure pseudosite copy number for efficiency percentage calculations. ddPCR reaction mix (22 pL total): 11 pL ddPCR Supermix for Probes (no dUTP) (Bio-Rad), 1.98 pL of each primer (10 pM), 0.55 pL of each probe (10 pM), 1.65 pL cleaned gDNA, 0.22 pL SacI-HF (NEB), water to volume. Each reaction contained primers and probes for the target site (FAM probe) and a nearby reference locus (HEX probe). Reactions were run on QX200 AutoDG Droplet Digital PCR System (Biorad). For off-target detection or low concentration samples, primers were increased to 20 pM and volume halved, and gDNA volume was increased to 4.95 pL.
[0349] qPCR reaction mix (40 pL total): 1 pL of each primer, 0.8 pL of each probe, 20 pL TaqMan Fast Advanced Mastermix (Thermo Fisher), 2.4 pL genomic DNA, and 12 pL water. Mastermix was split into three 10 pL technical replicates in a 384-well plate and run on LightCycler 480 (Roche). Primer pairs for ddPCR and qPCR are provided in Table S5.Three plasmid recombination assay in HEK293FT cells
[0350] A fluorescent reporter assay was used to assess episomal plasmid recombination in HEK293FT cells. One day before transfection, 12-18K HEK293FT cells were plated per well of a 96-well plate, aiming for 60%-80% confluency at the time of transfection. Three plasmids at a 1 : 1 : 1 molar ratio were transfected into the cells using Lipofectamine 2000: 1. 200 ng of the effector plasmid expressing the Dn29 variants and GFP, 2. 50.5 ng of the donor plasmid containing the attP attachment sequence and mCherry, and 3. 70.6 ng of the acceptor plasmid containing an Efl a promoter and the cognate attB attachment sequence. Upon recombination of the two attachment sequences, the Efl a promoter will drive expression of the mCherry CDS, which is read out by flow cytometry (Fig. S6D). To assess the excisionreaction, the attP in the donor plasmid is replaced with the left post-recombination attachment site (attB-L:attP-R), called attL, and the attB is replaced with the right post-recombination attachment site (attP-L:attB-R), called attR. To assess attP recombination with model organism pseudosites, the attB sequence is replaced with the pseudosite sequences. Mismatching LSR (Bxbl) controls with each donor and acceptor plasmid is used to correct for the leaky mCherry background expression, defining the flow cytometry gating boundaries. Three days post-transfection, the cells were trypsinized with 50 pL TrypLE (Gibco) for 10 minutes, quenched with 50 pL Stain Buffer (BD), transferred to U-bottom 96- well plates, centrifuged (300 x g, 5 minutes), and supernatant aspirated. Plates were resuspended in 200 pL Stain Buffer (BD) and analyzed with Attune NxT Flow Cytometer with Autosampler (Thermo Fisher).Site-directed mutagenesis for combinatorial mutant cloning
[0351] Site-directed mutagenesis (SDM) primers were designed using the script from Bi et al. 2020, selecting primers with Tm closest to 65°C 41. For each mutation, a forward and reverse primer were generated, each containing the desired mutation at the center. PCR reactions were set up combining: forward SDM primer with DMS universal reverse primer or reverse SDM primer with DMS universal reverse primer. PCR mixture (12.5 pL total): 6.25 pL Platinum Superfi II Mastermix, 0.5 pL each primer (10 pM), 1 pL plasmid template DNA (1 ng / pL), water to volume. PCRwasrun using the standard Platinum Superfi II Mastermix protocol with annealing temperature at 65°C. Products were cleaned with 0.5x Ampure XP beads. For Gibson assembly, 1 pL each cleaned PCR product, 5 pL Gibson mastermix, and 3 pL water was incubated at 50°C for 15 minutes, then transformed into Maehl E. coli and plated. For simultaneous cloning of two or more mutations, universal primers were replaced with other mutation's forward and reverse primers. Two mutations required a two-piece Gibson assembly, three mutations required a three-piece assembly, and so forth.Genome-wide integration site mapping
[0352] HEK293FTs were transfected as previously described, with a non-matching LSR (Bxbl) plasmid replacing the effector plasmid as a control for donor plasmid dilution. Cells were cultured for 2-3 weeks, passaging and analyzing by flow cytometry every 2-3 days at 80% confluency, until the non-matching LSR control was <1% mCherry+, indicating the plasmid had nearly completely diluted out. Genomic DNA was extracted using Quick-DNA Miniprep Plus Kit (Zymo), quantified by Qubit HS dsDNA Assay (Thermo), and 1 pg ofgDNA per sample was DpnI-digested (NEB) to remove residual donor plasmid. Tn5 tagmentation, nested PCR enrichment of integration sites, NGS sequencing, and computational analysis were performed as described in Durrant et al. 2023. To reduce occurrence of index hopping, unique dual i7 and i5 barcodes were utilized for the attHl targeted samples. To directly compare specificity of samples with different numbers of measured integration events, samples were downsampled to the same total UMI count.LSR-dCas9 and gRNA plasmid design and cloning
[0353] Fusion proteins consisting of a catalytically dead Cas9 fused to an LSR and a P2A-GFP were constructed by Gibson assembly into a pUC19-derived plasmid containing the Efla promoter and a SV40 poly-A tail. Variable flexible linkers, including a (GGS)8, (GGGGS)6 XTEN16, XTEN32-(GGSS)2, and XTEN48-(GGSS)2were used to link the dCas9 and LSR. Spacers targeting loci proximal to the LSR integration site and non-targeting controls were cloned into an sgRNA expressing plasmid via oligo ligation and Golden Gate cloning. Spacer selection was based on PAM sequence and pseudosite proximity.Designing and cloning the attP library
[0354] Two plasmid libraries (attP-L and attP-R) were constructed to determine nucleotide preference within the attP, with each 26 bp half-site mutagenized separately. IDT- synthesized oligo pools contained 79% WT base and 7% each of the other bases at each position. Single stranded oligo pools were subjected to second strand synthesis. First, an oligo anneal reaction containing 2 pL Library Oligo (100 pM), 4 pL klenow primer (100 pM), 3.4 pL lOx STE Buffer, and 24.6 pL water was heated at 95°C for 5 minutes, then cooled to room temperature. Next, a Klenow extension reaction containing 34 pL annealed libraries, 8 pL water, 5 pL lOx NEBuffer2, 2 pL 10 mM dNTPs (NEB), and 1 pL DNA Polymerase I, Large (Klenow) Fragment (NEB, 5000 U / mL) was incubated at 37°C for 30 minutes, purified (DNA Clean and Concentrator- 5, Zymo), and eluted in 20 pL nuclease-free water.
[0355] The purified product was cloned by Esp3i Golden Gate cloning into pCB235: 75 ng pre-digested backbone, 3: 1 molar ratio of attP library to backbone, 0.5 pL each of T4 DNA ligase (NEB) and Esp3i (Thermo Fisher), 1 pL T4 DNA Ligase Buffer (NEB), and water to 10 pL was incubated at 37°C for 1 hour, purified, and eluted in 6 pL nuclease-free water. 1 pL purified library was electroporated into Endura Electrocompetent Cells (Biosearch Technologies) at 10 pF, 600 Q, 1800 V, recovered in 2 mL Lucigen Recovery Media (37°C, 1 hour), plated on 245 mm x 245 mm Bioassay dishes, and incubated at 30°Covernight. Final libraries were scraped, purified (Nucleobond Xtra Maxi EF kit, Machery- Nagel), and sequenced with Illumina NextSeq2000. attP library transfection, harvest, and library preparation
[0356] 2.2 x 106 HEK293FT cells were plated on 10 cm dishes one day before transfection to achieve 70% confluence at transfection. 24 pg total plasmid DNA was prepared at a 5: 1 :1 molar ratio (attP library :LSR effector: sgRNA). DNA and 72 pL Lipofectamine 2000 were separately mixed with 1.5 mL OMEM, incubated for 5 minutes, then combined and incubated for 10 minutes before adding dropwise to cells. After three days, cells were harvested with TrypLE (Gibco) and genomic DNA extracted using Quick DNA Midiprep Plus Kit (Zymo). Integration events were amplified by single-step PCR with i5 / i7 index-adding primers using all available genomic DNA. Biological replicates had 1 bp staggered amplicons to increase nucleotide diversity. PCR conditions: 25 pL NEBNext High Fidelity PCR Master Mix, 2.5 pg genomic DNA, 1.25 pL each of the attL or attR i5 or i7 primers (Table S5), water to 50 pL. Cycling: 25 cycles of 98°C for 10s, 63°C for 10s, 72°C for 25s. PCR products were pooled, run on 2%agarose gel, and correct-size bands were extracted (Monarch® DNA Gel Extraction Kit, NEB). Libraries were quantified (Qubit lx dsDNA High Sensitivity assay, Thermo Fisher), pooled equimolar with 35% PhiX spike-in, and sequenced on Illumina NextSeq2000 (150bp paired-end reads). attP Library enrichment analysis
[0357] Libraries were demultiplexed using Illumina Basespace automatic demultiplexing workflow. Paired-end reads were merged using BBMerge (v.39.06) and analyzed with a custom Python script. Reads were filtered for exact amplicon length and QScore > 30. Next, percent abundance of each nucleotide at each attP position was calculated for input and output libraries. Enrichment scores were computed using the equation: r = AI ~A B / l-B , where A and B represent the read counts for selected nucleotides in output and input libraries, respectively, normalized to the total number of reads. Enrichment scores were converted to sequence logos, generated using Logomaker 42 and matplotlib packages. Unique library members recovered as integration events were assessed by generating the set of unique reads. The number of unique integration events from NGS analysis was compared to ddPCR analysis of bulk genomic DNA for validation.Stem cell transfection
[0358] Hl hESCs were cultured in MTeSR Plus medium (STEMCELL Technologies) on Cultrex-coated (Bio-Techne) or Matrigel-coated (Coming) 6-well plates. Cells were routinelysubcultured at a 1 : 12 ratio using ReLeSR Passaging Reagent (STEMCELL Technologies) every four days or at 70-80% confluency. Three days after splitting (60% confluency), the cells were dissociated for 10 minutes with Accutase (STEMCELL Technologies), and plated in Cultrex-coated 96-well plates at 25-30K cells per well with 10 pM Rock inhibitor. The next day (at 70% confluency), media was changed to include 50 pM Rock inhibitor two hours pre-transfection. 3 pg plasmid DNA containing a 1 : 1 molar ratio of combined effector / guide plasmid to donor plasmid in 10 pL volume was diluted in 81 pL MTeSR Plus and thoroughly pipette mixed. 9 pL of Fugene HD Transfection Reagent (Promega) was added to the DNA / MTeSR mix, thoroughly mixed, and incubated for 12 minutes. After another thorough pipette mix, 7 pL of the DNA was added dropwise to each well. The cells were incubated at 37°C for 1-2 days, splitting 1 :2 if 90% confluency was reached. After 3 days, the cells were dissociated with Accutase and split into two V-bottom plates, one for flow cytometry and one for gDNA harvest with QuickExtract DNA Solution (BioSearch Technologies).HPC differentiation and surface marker staining
[0359] hESCs were differentiated into hematopoietic progenitor cells using the STEMdiff Hematopoietic Kit (STEMCELL Technologies). On day 10 of differentiation, 250 pL of nonadherent cells were collected from the supernatant using wide bore P1000 tips and transferred to a V-bottom 96 well plate. Next, the cells were pelleted at 400g for 5 minutes, supernatant discarded, and resuspended in 95 mL Stain Buffer (BD) containing 1 pL of each antibody with a wide bore pipette. The following antibodies were used: APC CD81 (BD, Cat:551112), APC CD147 (Thermo Fisher, Ref: A15706), Alexa Fluor® 647 CD63 (BD, Cat: 561983), APC / Cyanine7 CD34 (BioLegend, Cat: 343514), PE CD43 (BioLegend, Cat: 343204). The cells were incubated in the dark for 20 minutes to 1 hour, washed once with Stain Buffer, and flowed on the Attune Flow Cytometer (Thermo Fisher). hESC single cell dilution and genotyping
[0360] hESCs were diluted to 1 cell / 100 pL in MTeSR Plus media supplemented with lx CloneR (STEMCELL Technologies) and plated into two 96-well plates per sample. Cells were maintained until colonies were visible, then wells with multiple colonies were removed. Single colonies were expanded to 24-well dishes when they covered half the surface area of the 96-well. At the next split, one quarter of each well was pelleted for gDNA extraction using QuickExtract DNA Solution (BioSearch Technologies). The extracted gDNA was cleaned with 0.9x AmpureXP beads and genotyped by ddPCR. Primers and probes weredesigned to target the attHl junction (ddPCR attHl l set), the donor sequence(Amp forward, Amp reverse, Amp_probe), and a nearby genomic reference sequence. On- target zygosity was determined by the attHl / reference ratio, while total zygosity was measured by the donor / reference ratio.HEK293FT single cell sorting and genotyping
[0361] HEK293FT cells were transfected as previously described. Eight days posttransfection, cells were placed under puromycin selection (0.5 pg / mL) for 10 days. On day 18, cells were trypsinized, strained through a 35 pm filter, and single mCherry+ cells were sorted into four 96-well plates per sample using the FACSAria Fusion (BD). Single cell colonies were expanded for two weeks until >50% confluent, with visual inspection to ensure single colony growth. Wells with zero or multiple colonies were excluded from analysis.
[0362] Confluent colonies were harvested with QuickExtract DNA Solution (BioSearch Technologies) and amplified in two separate PCRs: PCR 1 using primers UMI reverse and ddPCR attHl forward l, flanking the UMI and attHl donor / genome junction, and PCR 2 using primers UMI reverse and UMI forward, flanking the UMI on the donor plasmid. Amplicons were sequenced via Sanger and / or NGS to determine on-target UMI count (PCR 1) and total UMI count (PCR 2), allowing calculation of on-target and off-target insertion counts per colony.Lentivirus production and HPC transduction
[0363] sgRNA spacers targeting cell surface markers CD81, CD 147, and CD63 were cloned into the LentiGuide-Puro construct (Addgene #52963). Lentivirus was generated using the LV-MAX Lentiviral Production Kit (Invitrogen) according to manufacturer's instructions and concentrated lOOx with Lenti-X Concentrator (Takara). HPCs were diluted to 50,000 cells per well in 100 pL of Media B (STEMdiff Hematopoietic Kit, STEMCELL Technologies) in a 96-well plate. Each well received 1 pL of LentiBOOST (SIRION Biotech) and 1 pL of lentivirus. Media was changed the following day. Four days post-transduction, a subset of cells was stained for cell surface markers. Remaining cells were treated with 1 pg / mL puromycin for four days to select for transduced cells, followed by cell surface marker staining.Generating Alphafold3 models of Dn29 bound to attB
[0364] The full-length wildtype Dn29 protein sequence and minimal attB-L or attB-R sequence (attB-L: GTAGACAAGGAAGGTAATGA; attB-R: GAAATAAGTTTGATAGATAT) were input into the Alphafold3 web server with the seedset to "auto". Five models were generated for each query of Dn29 bound to an attB half-site. Outputs were manually inspected to ensure correct orientation of Dn29 bound to the half-site, with the dinucleotide core of the DNA proximal to the NTD. One model (Dn29 x attB-R) out of the 10 generated models met this criterion and was selected for further analysis. The chosen model was compared to the Listeria Integrase crystal structure of the LSR CTD and attP complex (pdb:4KIS). Despite 4KIS being bound to attP instead of attB, domain-wise comparisons showed strong alignment: RMSDs were 1.341 and 1.707 for the zinc-ribbon domain, and recombinase domain,. Protein / DNA interface residues were identified with the InterfaceResidues pymol script using default settings.Predicting combinatorial mutations and feature importance with machine learning
[0365] The efficiency and specificity data of all Dn29 variants were split into a training and test set based on what round of experimentation they were generated in. The training set, called round 1, contained all variants from the two single mutation validation experiments, where mutations were tested individually on top of variant or variant 381. The testing set contained all higher order combinations from the iterative rounds of driver mutation stacking (rounds 2-5). The efficiency (percent of integrations at attHl) was normalized to wildtype and specificity (ratio of attHl / attH3 activity) was log-transformed. The full amino acid sequences of the protein variants were one-hot encoded, and activity in the training set was modeled using linear regression, ridge regression, XGBoost, and CatBoost with the scikit- learn, xgboost, and catboost Python libraries.
[0366] For the ridge regression, optimal alpha was identified through minimization of the testing set R2 (alpha=0.8 for efficiency model, alpha=1.3 for specificity model). Hyperparameter optimizations were conducted for XGBoost and CatBoost by performing a randomized search, evaluating on negative mean squared error, using the following parameters: XGBoost: 'n_estimators': [100, 500, 1000], 'leaming_rate': [0.01, 0.05, 0.1], 'max_depth': [3, 5, 7], 'subsample': [0.5, 0.6, 0.7, 0.8, 1.0], 'colsample_bytree': [0.7, 0.8, 1.0]; CatBoost: 'iterations': [100, 200, 500, 1000], 'leaming_rate': [0.01, 0.05, 0.1, 0.2], 'depth': [4, 6, 8, 10], '12_leaf_reg': [1, 3, 5, 7, 9], 'bagging_temperature': [0, 1, 2, 3], 'random_strength': [1, 1.5, 2, 3], 'border count' : [32, 64, 128], 'grow_policy': ['SymmetricTree', 'Depthwise', 'Lossguide'].
[0367] The following parameters were chosen for each model: XGBoost, specificity: 'subsample - 0.5, 'n_estimators'= 1000, 'max_depth'= 7, 'leaming_rate'= 0.1, 'colsample_bytree'= 0.8; XGBoost, efficiency: 'subsample - 0.7, 'n_estimators'= 100,'max_depth'= 7, 'learning_rate'= 0.05, 'colsample_bytree'= 0.7; CatBoost, specificity: 'random_strength'= 1.5, 'learning_rate'= 0.1, '12_leaf_reg'= 1, 'iterations - 1000, 'grow_policy'= 'Depthwise', 'depth - 4, 'border_count'= 128, 'bagging_temperature'= 1; CatBoost; efficiency: 'random_strength'= 1.5, 'learning_rate'= 0.1, '12_leaf_reg'= 7, 'iterations - 500, 'grow_policy'= 'Lossguide', 'depth - 4, 'border_count'= 32, 'bagging_temperature'= 2.In vitro transcription and purification of mRNA
[0368] Effector constructs were cloned into an in vitro transcription (IVT) plasmid as previously described 43. This plasmid contained a mutated T7 promoter, 5' UTR, P2A EGFP, and 3' UTR followed by a 145 bp poly A sequence. IVT templates were generated by PCR using primers 0GXOO6 and oLGR009, which incorporate a polyA tail and correct the T7 promoter mutation. PCRreactions were performed using KAPA-HiFi HotStart 2x (Roche) master mix with 6.25 ng of plasmid template per 25 pL reaction. The PCR protocol involved annealing at 63°C, extending for 45 seconds per kb, and running for 18 cycles. The reactions were purified using 0.8x volume of SPRI beads and eluted into water. The purified PCRs were analyzed by gel electrophoresis and nanodrop to ensure correct size and determine concentration.
[0369] The IVT reactions were set up using the HiScribe T7 High-Yield RNA Synthesis Kit (New England Biolabs, Cat #E2040S), modified with full pseudo-UTP substitution using N1 -Methyl -Pseudo-U (TriLink Biotechnologies, Cat #N-1081) and co-transcriptionally capped with CleanCap AG (TriLink Biotechnologies, Cat #N-7113). Each IVT reaction contained 5 mM ATP, CTP, GTP, and pseudo-UTP, 4 mM CleanCAP AG, lx Transcription Buffer, 3.75 ng / pL DNA template, 1 U / pL Murine RNAse Inhibitor (New England Biolabs, Cat #M0314L), 0.002 U / pL yeast inorganic pyrophosphatase (NEB. Cat # M2403L), and 5 U / pL T7 RNA polymerase. Reactions were incubated for 2.5 hours at 37°C.
[0370] Next, the mRNA was purified using lithium chloride. To each reaction, 1.5X water and 1.25X 7.5M LiCl was added. The solution was chilled at-20°C for 30 minutes and then spun at max speed (16,000xg) for 15 minutes at 4°C. The supernatant was discarded, and the pellet was rinsed with 70% ice cold ethanol to remove residual salts. After another max speed spin for 10 minutes at 4°C, the mRNA was resuspended in water and stored at- 80°C. The mRNA was analyzed on the Agilent Tapestation and by Qubit RNA High Sensitivity (Thermo) to ensure correct size and determine concentration.Electroporation of primary human T cells
[0371] Two days before electroporation, T cells were seeded at 1 x 106 fresh cells / mL and activated with a 1 : 1 bead-to-cell ratio with anti-CD3 / CD28 Dynabeads (Life Technologies, #40203D). On the day of electroporation, the beads were magnetically removed and the T cells were electroporated with 2 pg LSR-dCas9-P2A-EGFP mRNA and 2 pg sgRNA (Synthego) for LSR-dCas9 samples or 1 pg LSR-P2A-EGFP mRNA for LSR samples using the Lonza P3 Primary Cell Kit. Each electroporation contained between 0.5 x 106- 1 x 106 cells in 20 pL total volume, and was electroporated using the 4D Nucleofector system and the DS-137 pulse code. Immediately after electroporation, 80 pL pre-warmed culture media was added to the Nucleocuvette strip, which was then incubated at 37°C for 15-30 minutes. Next, 2 x 105cells per condition were split into 96 well U-bottom plates in 100 pL of serum free media (TheraPEAK® X-VIV0®-15 Serum-free Hematopoietic Cell Medium, Cat #BEBP04-744Q) supplemented with 5 ng pl-1 IL-7, and 5 ng pl-1 IL-15 . Cells were then transduced at an MOI of 1 x 105 genome copies / cell with ssAAV or scAAV vectors of serotype 6 (AAV6) containing the e-attP sequence, attHl sgRNA target sequence, and an mCherry expression cassette which were ordered from VectorBuilder. The next morning, cells were spun down at 300 x g for 5 minutes, the serum free media was removed, and cells were resuspended in 200 pl of fresh cX-VIVO. Cells were maintained and passaged as needed by the addition of cX-VIVO every 2-3 days.T cell staining, flow cytometry, and genomic harvesting
[0372] Three days post-electroporation, 50 pL of T cells were collected for staining and flow cytometry. Briefly, cells were centrifuged, washed once with 200 pl cell staining buffer, and stained with Ghost Dye™ Red 780 at a 1 : 1000 dilution (Tonbo, Cat #13-0865-T500) for 20 minutes in the dark at 4°C. The cells were measured using an Attune NxT Cytometer with a 96-well autosampler (Invitrogen) and analyzed using FlowJo for viability, mCherry fluorescence (expressed on the AAV) and GFP fluorescence (effector expression). The remaining 150 pL of T cells in culture were centrifuged at 300 x g for 5 minutes and the gDNA was harvested using QuickExtract DNA Solution (BioSearch Technologies) and analyzed by ddPCR as described above.Generative Al
[0373] Al language models (ChatGPT and Claude) were utilized for generating custom Python scripts for data analysis and visualization, assistance with copy-editing, and infillingpreliminary drafts of some sections based on an author-provided outline. All Al-generated content was thoroughly reviewed, edited, and verified by the authors.Example 3: Integration Efficiency of super Dn29-dCas9 at attHl in primary human T cells
[0374] Figure 54 shows in vitro integration using superDn29-dCas9 for T cell engineering using plasmid donors.
Claims
CLAIMSWhat is claimed is:
1. An engineered Dn29 comprising an amino acid sequence comprising at least 70% identity to SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
2. The engineered Dn29 of claim 1, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
3. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
4. The engineered Dn29 of claim 3, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
5. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising two or more amino acid mutations, wherein at least one mutation is selected from E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H and at least one mutation selected is selected I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
6. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, or G2277. The engineered Dn29 of claim 6, wherein the mutation comprises M6I, E70G, A224P, or G227V.
8. The engineered Dn29 of claim 6, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, and G227V.
9. The engineered Dn29 of claims 6-8, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234.
10. The engineered Dn29 of claims 6-9, further comprising a mutation at amino acid residue N341.
11. The engineered Dn29 of claim 10, wherein the mutation comprises N341K.
12. The engineered Dn29 of claims 6-11, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 318.
13. The engineered Dn29 of claims 6-12, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
14. The engineered Dn29 of claim 13, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
15. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q332 or N34116. The engineered Dn29 of claim 15, wherein the mutation comprises M6I, E70G, A224P, G227V, Q332K, or N341K.
17. The engineered Dn29 of claim 15, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q332K, and N341K.
18. The engineered Dn29 of claims 15-17, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or19. The engineered Dn29 of claims 15-18, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
20. The engineered Dn29 of claim 19, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
21. The engineered Dn29 of claims 15-18, further comprising a mutation at one or more additional amino acids selected from 1233, K248, Q390, L388, L393, or D503.
22. The engineered Dn29 of claim 21, wherein the mutation comprises I233K, K248R, Q390P, L388P, L393P, or D503N.
23. The engineered Dn29 of claims 15-18, further comprising a mutation at one or more additional amino acids selected from D503, L388, Q390, L393, W373, R389, E452, F138, or V386.
24. The engineered Dn29 of claim 23, wherein the mutation comprises D503N, L388P, Q390P, L393P, W373R, R389S, E452G, F138L, or V386A.
25. The engineered Dn29 of claims 15-18, further comprising a mutation of one or more amino acid residue 1233, K248, Q390, L388, L393, or D503 and a mutation of one or more amino acid residue D503, L388, Q390, L393, W373, R389, E452, F138, or V386.
26. The engineered Dn29 of claim 25, wherein the mutation comprises I233K, K248R, Q390P, L388P, L393P, D503N, W373R, R389S, E452G, F138L, or V386A.
27. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, or N34128. The engineered Dn29 of claim 27, wherein the mutation comprises M6I, E70G, A224P, G227V, or N341Q.
29. The engineered Dn29 of claim 27, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, and N341Q.
30. The engineered Dn29 of claims 27-29, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
31. The engineered Dn29 of claims 27-30, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
32. The engineered Dn29 of claim 31, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
33. The engineered Dn29 of claims 27-32, further comprising a mutation at one or more additional amino acids selected from W373, L388, L393, or D503.
34. The engineered Dn29 of claim 33, wherein the mutation comprises W373R, L388P, L393P, or D503N.
35. The engineered Dn29 of claims 27-32, further comprising a mutation at one or more additional amino acids selected from 1233, 1303, K248, or Q332.
36. The engineered Dn29 of claim 35, wherein the mutation comprises I233K, I303K, K248R, or Q332K.
37. The engineered Dn29 of claims 27-32, further comprising a mutation of one or more amino acid residue W373, L388, L393, or D503 and a mutation of one or more amino acid residue 1233, 1303, K248, or Q332.
38. The engineered Dn29 of claim 37, wherein the mutation comprises W373R, L388P, L393P, D503N, I233K, I303K, K248R, or Q332K.
39. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, 1303, N341, or L388.
40. The engineered Dn29 of claim 39 wherein the mutation comprises M6I, E70G, A224P, G227V, I303K, N341Q or L388P.
41. The engineered Dn29 of claim 39, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, I303K, N341Q and L388P.
42. The engineered Dn29 of claims 39-41, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
43. The engineered Dn29 of claims 39-42, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
44. The engineered Dn29 of claim 43, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
45. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, 1303, N341, or L39346. The engineered Dn29 of claim 45, wherein the mutation comprises M6I, E70G, A224P, G227V, I303K, N341Q, or L393P.
47. The engineered Dn29 of claim 45, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, I303K, N341Q, and L393P.
48. The engineered Dn29 of claims 45-47, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
49. The engineered Dn29 of claims 45-47, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
50. The engineered Dn29 of claim 49, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
51. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q322, N341, L393, or D503.
52. The engineered Dn29 of claim 51, wherein the mutation comprises M6I, E70G, A224P, G227V, Q322K, N341K, L393P, or D503N.
53. The engineered Dn29 of claim 51, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, and D503N.
54. The engineered Dn29 of claims 51-53, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
55. The engineered Dn29 of claim 54, consisting of mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, and D503N and silent mutations in the codons encoding amino acid 234 and 318.
56. The engineered Dn29 of claims 51-54, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
57. The engineered Dn29 of claim 56, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
58. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q322, N341, L388, or R389.
59. The engineered Dn29 of claim 58, wherein the mutation comprises M6I, E70G, A224P, G227V, Q322K, N341K, L388P, or R389S.
60. The engineered Dn29 of claim 58, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L388P, and R389S.
61. The engineered Dn29 of claims 58-60, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
62. The engineered Dn29 of claims 58-61, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
63. The engineered Dn29 of claim 62, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
64. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, Q322, N341, L393, D503, 198, 1233, 1303, or K248.
65. The engineered Dn29 of claim 64, wherein the mutation comprises M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, or K248R.
66. The engineered Dn29 of claim 64, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R.
67. The engineered Dn29 of claims 64-66, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
68. The engineered Dn29 of claims 64-67, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
69. The engineered Dn29 of claim 68, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
70. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected from M6, E70, A224, G227, 1233, K248, 1303, Q332, N341, L388, or Y404.
71. The engineered Dn29 of claim 70, wherein the mutation comprises M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P, or Y404C.
72. The engineered Dn29 of claim 70, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q332K, N341Q, L388P and Y404C.
73. The engineered Dn29 of claim 72, wherein the amino acid sequence further comprises one or more mutations selected from D503N, W373R, L449P, N214D, I98V, or K288E.
74. The engineered Dn29 of claims 70-73, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
75. The engineered Dn29 of claim 74, consisting of mutations M6I, E70G, A224P, G227V, I233K, K248R, I303K, Q322K, N341Q, L388P, and Y404C and silent mutations at in the codons encoding amino acid 234 and 318.
76. The engineered Dn29 of claims 70-74, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
77. The engineered Dn29 of claim 76, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G, F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
78. An engineered Dn29 comprising an amino acid sequence comprising SEQ ID NO: 1 and comprising a mutation at one or more amino acids selected M6, E70, A224, G227, Q322, N341, L393, D503, 198, 1233, 1303, or K248.
79. The engineered Dn29 of claim 78, wherein the mutation comprises M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, or K248R.
80. The engineered Dn29 of claim 78, wherein the amino acid sequence comprises mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R.
81. The engineered Dn29 of claims 78-80, wherein a nucleic acid encoding the engineered Dn29 comprises a silent mutation in the codon encoding amino acid 234 and / or 318.
82. The engineered Dn29 of claim 81, consisting of mutations M6I, E70G, A224P, G227V, Q322K, N341K, L393P, D503N, I98V, I233K, I303K, and K248R and silent mutations at in the codons encoding amino acid 234 and 318.
83. The engineered Dn29 of claims 78-81, further comprising a mutation at one or more additional amino acids selected from E70, F138, A224, N341, L388, V386, Q390, L393, D503, V202, W373, S322, F374, R389, Y404, L449, N214, 1303, Q332, 1233, K248, 198, R246, V150, D437, S228, or K288.
84. The engineered Dn29 of claim 83, wherein the mutation comprises E70G, F138L, A224P, N341Q, N341V, L388P, V386A, Q390P, L393P, D503N, V202A, W373R, S322G,F374S, W373C, R389S, Y404C, L449P, N214D, N341I, N341H, I303K, Q332K, N341K, I233K, K248R, I98V, R246G, V150K, N341R, D437G, S228A, or K288E.
85. A fusion polypeptide, wherein the fusion polypeptide comprises an engineered Dn29 portion comprising the engineered Dn29 of claim 1-84 and a DNA binding domain (DBD) portion.
86. The fusion polypeptide of claim 85, wherein the engineered Dn29 portion is fused N- terminal to the DBD portion.
87. The fusion polypeptide of claims 85-86, wherein the fusion polypeptide further comprises a peptide linker positioned between the engineered Dn29 portion and the DBD portion.
88. The fusion polypeptide of claim 87, wherein the peptide linker comprises 2 to 100 amino acids.
89. The fusion polypeptide of claim 87, wherein the peptide linker comprises glycine and serine residues, one or more XTEN16 repeats, or a combination thereof.
90. The fusion polypeptide of claim 87, wherein the peptide linker comprises SEQ ID NOs: 18-26.
91. The fusion polypeptide of claims 85-90, wherein the DBD portion comprises Cas9, Cpfl, Cast 2b, Cast 2c, Cast 2d, Casl2e, Casl2f, Casl2h, Casl2i, or Cast 2g.
92. The fusion polypeptide of claim 91, wherein the Cas9, Cpfl, Casl2b, Casl2c, Casl2d, Casl2e, Casl2f, Casl2h, Casl2i, or Casl2g lack nuclease and / or nickase activity.
93. The fusion polypeptide of claims 85-90, wherein the DBD portion comprises dCas9.
94. The fusion polypeptide of claim 93, wherein the DBD portion comprises an amino acid sequence at least 90% identical to dCas9 (SEQ ID NO: 10), dCas9-HFl (SEQ ID NO: 11), dCas9-SpG (SEQ ID NO: 12), or dCas9-SpG-HFl (SEQ ID NO: 13).
95. The fusion polypeptide of claim 93, wherein the DBD portion comprises an amino acid sequence of dCas9 (SEQ ID NO: 10), dCas9-HFl (SEQ ID NO: 11), dCas9-SpG (SEQ ID NO: 12), or dCas9-SpG-HFl (SEQ ID N0:13).
96. The engineered Dn29 of claims 1-84 or the fusion polypeptide of claims 85-95 further comprising one or more nuclear localization signals (NLSs).
97. The fusion polypeptide of claims 85-96, wherein the DBD portion of the fusion polypeptide binds to a guide RNA (gRNA).
98. A nucleic acid encoding the engineered Dn29 of claims 1-84, or 96.
99. A nucleic acid encoding the fusion polypeptide of claims 85-96.
100. A vector comprising any of the nucleic acids of claim 98.
101. A vector comprising any of the nucleic acids of claim 99.
102. A host cell comprising the vector of claim 100 or 101.
103. A nucleic acid editing system comprising a first nucleic acid according to claim 100 and a second nucleic acid encoding a gRNA.
104. The nucleic acid editing system of claim 103, wherein the gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracr RNA portion, wherein the nucleic acid sequence of the spacer sequence portion is the same as a target nucleic acid sequence, except that T in the target nucleic acid sequence is U in the spacer sequence portion, and wherein the target nucleic acid sequence is within 80 nucleotides upstream or downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest.
105. The nucleic acid editing system of claim 104, wherein the spacer sequence portion is 16 to 20 nucleotides long.
106. The nucleic acid editing system of claims 103-105, wherein the gRNA encoded by the nucleic acid is an sgRNA.
107. The nucleic acid editing system of claims 104-106 wherein immediately 3’ to the target nucleic acid sequence on the DNA of interest is a PAM sequence.
108. The nucleic acid editing system of claims 104-107, wherein the target nucleic acid sequence is within 80 nucleotides upstream or downstream of a dinucleotide core of an attA site of the engineered Dn29 portion of the fusion polypeptide on a target DNA of interest.
109. The nucleic acid editing system of claim 108, wherein the attA site is a pseudosite in a mammalian target DNA of interest.
110. The nucleic acid editing system of claim 109, wherein the attA site is a pseudosite in the human genome (attH).
111. The nucleic acid editing system of claim 110, wherein the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attH site is chrl0:21130404- 21130406:-, chrl 1 :77367459-77367461 :-, chr 1:230490334-230490336:+, chr2: 14280297- 14280299:+, chr9: 116464427-116464429:+, chr20:38982599-38982601 :+, chr5:3553012- 3553014:-, chr7:134676315-134676317:-, chrl0:58514255-58514257:+, or chr4:92338934- 92338936:+.
112. The nucleic acid editing system of claim 109, wherein the attA site is a pseudosite in a non-human primate genome.
113. The nucleic acid editing system of claim 109, wherein the attA site is a pseudosite in a mouse (Mus musculus) genome.
114. The nucleic acid editing system of claim 113, wherein the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chrX:98,518,458.
115. The nucleic acid editing system of claim 109, wherein the attA site is a pseudosite in a common marmoset (Callithrix jacchus) genome.
116. The nucleic acid editing system of claim 115, wherein the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr7: 19,911,959117. The nucleic acid editing system of claim 109, wherein the attA site is a pseudosite in a Rhesus monkey (Macaca mulatto) genome.
118. The nucleic acid editing system of claim 117, wherein the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr9:22, 186,336.
119. The nucleic acid editing system of claim 109, wherein the attA site is a pseudosite in a Cynomolgus monkey (Macaca fascicularis) genome.
120. The nucleic acid editing system of claim 119, wherein the fusion polypeptide encoded by the nucleic acid comprises dCas9 (SEQ ID NO: 10) and the attA site is chr9:21,946,090.
121. The nucleic acid editing system of claims 104-120, wherein the tracr RNA portion comprises SEQ ID NO: 108.
122. The nucleic acid editing system of claims 104-121, wherein the target nucleic acid sequence is within 80 nucleotides upstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest.
123. The nucleic acid editing system of claims 104-121, wherein the target nucleic acid sequence is within 80 nucleotides downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest.
124. The nucleic acid editing system of claims 122, further comprising a third nucleic acid encoding a second gRNA.
125. The nucleic acid editing system of claim 124, wherein the second gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracr RNA portion, wherein the nucleic acid sequence of the spacer sequence portion is the same as a target nucleic acid sequence, except that T in the target nucleic acid sequence is U in the spacer sequence portion, and wherein the target nucleic acid sequence is within 80 nucleotides downstream of a dinucleotide core of an attachment site of the engineered Dn29 portion of the fusion polypeptide on a DNA of interest.
126. The nucleic acid editing system of claim 125, wherein the spacer sequence portion of the second gRNA is 16 to 20 nucleotides long.
127. The nucleic acid editing system of claims 124-126, wherein the second gRNA encoded by the nucleic acid is an sgRNA.
128. The nucleic acid editing system of claims 125-127, wherein immediately 3’ to the target nucleic acid sequence on the DNA of interest is a PAM sequence.
129. The nucleic acid editing system of claims 103-123, further comprising a third nucleic acid comprising a donor DNA sequence which comprises an attD attachment site of theengineered Dn29 portion of the fusion polypeptide and a nucleic acid sequence for insertion into the target DNA of interest.
130. The nucleic acid editing system of claim 129, wherein the third nucleic acid further comprises a portion that has the same target nucleic acid sequence for the gRNA as the target DNA of interest.
131. The nucleic acid editing system of claims 129-130, wherein the fusion polypeptide encoded by the nucleic acid comprises: dCas9 (SEQ ID NO: 10), the attH site on the target DNA of interest is chromosomal locus chrl0:21130404-21130406:-, chrl 1 :77367459-77367461 :-, chrl :230490334-230490336:+, chr2: 14280297-14280299:+, chr9: 116464427-116464429:+, chr20:38982599-38982601 :+, chr5:3553012-3553014:-, chr7: 134676315-134676317:-, chrl0:58514255-58514257:+, or chr4:92338934-92338936:+ or comprises the attH sequence found at said chromosomal locus, and the attD attachment site of the donor DNA sequence comprises SEQ ID NO: 109, or a sequence 90% identical to SEQ ID NO: 109.
132. The nucleic acid editing system of claims 129-131, wherein the third nucleic acid is a plasmid.
133. The nucleic acid editing system of claims 129-131, wherein the third nucleic acid is a linear amplicon.
134. The nucleic acid editing system of claims 103-133, wherein the nucleic acid encoding the fusion polypeptide, the nucleic acid encoding the gRNA, or both, and / or, where present, the third nucleic acid encoding the second gRNA are expressed from an inducible promoter.
135. A nucleic acid editing system comprising a first nucleic acid of claim 98.
136. The nucleic acid editing system of claim 128, wherein an attA site of the engineered Dn29 on a target DNA of interest is a pseudosite in a mammalian target DNA of interest.
137. The nucleic acid editing system of claim 129, wherein the attA site is a pseudosite in the human genome (attH).
138. The nucleic acid editing system of claim 130, wherein the attH site is chrl0:21130404-21130406:-, chrl 1 :77367459-77367461 :-, chrl :230490334-230490336:+,chr2: 14280297-14280299:+, chr9: 116464427-116464429:+, chr20:38982599-38982601 :+, chr5:3553012-3553014:-, chr7: 134676315-134676317:-, chrl0:58514255-58514257:+, or chr4:92338934-92338936:+.
139. The nucleic acid editing system of claim 129, wherein the attA site is a pseudosite in a non-human primate genome.
140. The nucleic acid editing system of claim 129, wherein the attA site is a pseudosite in a mouse (Mus musculus) genome.
141. The nucleic acid editing system of claim 140, wherein the attA site is chrX:98,518,458.
142. The nucleic acid editing system of claim 129, wherein the attA site is a pseudosite in a common marmoset (Callithrix jacchus) genome.
143. The nucleic acid editing system of claim 142, wherein the attA site is chr7: 19,911,959.
144. The nucleic acid editing system of claim 129, wherein the attA site is a pseudosite in a Rhesus monkey (Macaca mulatto) genome.
145. The nucleic acid editing system of claim 144, wherein the attA site is chr9:22,186,336.
146. The nucleic acid editing system of claim 129, wherein the attA site is a pseudosite in a Cynomolgus monkey (Macaca fascicularis) genome.
147. The nucleic acid editing system of claim 146, wherein the attA site is chr9:21,946,090.
148. The nucleic acid editing system of claims 135-147, further comprising a second nucleic acid comprising a donor DNA sequence which comprises an attD attachment site of the engineered Dn29 and a nucleic acid sequence for insertion into the target DNA of interest.
149. The nucleic acid editing system of claims 135-137, wherein the attH site on the target DNA of interest is chromosomal locus chrl0:21130404-21130406:-, chrl 1 :77367459- 77367461 :-, chrl :230490334-230490336:+, chr2: 14280297-14280299:+, chr9: 116464427-116464429:+, chr20:38982599-38982601:+, chr5:3553012-3553014:-, chr7: 134676315- 134676317:-, chrlO:58514255-58514257:+, or chr4:92338934-92338936:+ or comprises the attH sequence found at said chromosomal locus, and the attD attachment site of the donor DNA sequence comprises SEQ ID NO: 109, or a sequence 90% identical to SEQ ID NO: 109.
150. The nucleic acid editing system of claims 134-135, wherein the second nucleic acid is a plasmid.
151. The nucleic acid editing system of claims 148-149, wherein the second nucleic acid is a linear amplicon.
152. The nucleic acid editing system of claims 135-151 , wherein the nucleic acid encoding the engineered Dn29 is expressed from an inducible promoter.
153. A vector comprising any of the nucleic acids of the nucleic acid editing system of claims 103-152.
154. A host cell comprising any of the vector(s) of claim 153.
155. A method of integrating a donor DNA sequence into a target DNA of interest of a cell, the method comprising introducing into the cell: a nucleic acid editing system according to claims 103-155.
156. The method of claim 155, wherein the cell is a mammalian cell.
157. The method of claim 156, wherein the cell is a human cell.
158. The method of claim 157, wherein the cell is a human embryonic stem cell, a primaryT cell, or a non-dividing cell.
159. The method of claim 156, wherein the cell is a non-human primate cell.
160. The method of claim 159, wherein the cell is a mouse cell161. The method of claims 141-146, wherein the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site.
162. The method of claims 155-161, wherein the donor DNA comprises an engineered Dn29 attD attachment site which is integrated into the target DNA of interest.
163. The method of claims 155-162, wherein the target DNA of interest of the cell is the genome of the cell.
164. The method of claims 155-162, wherein the target DNA of interest of the cell is a plasmid.
165. A method of inverting a DNA sequence of a target DNA of interest, the method comprising introducing into a cell: a nucleic acid editing system according to claims 103-152, wherein attD and attA attachment sites of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) are present on the same DNA target molecule of interest in reverse orientation.
166. The method of claim 165, wherein the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site.
167. The method of claims 165-166, wherein the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attD attachment site.
168. The method of claims 165-167, wherein the target DNA of interest of the cell is the genome of the cell.
169. A method of excising a DNA sequence of a target DNA of interest, the method comprising introducing into a cell: a nucleic acid editing system according to claims 103-152, wherein attD and attA attachment sites of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) are present on the same DNA target molecule of interest in the same orientation.
170. The method of claim 169, wherein the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site.
171. The method of claims 169-170, wherein the target DNA of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attD attachment site.
172. The method of claims 169-171, wherein the target DNA of interest of the cell is the genome of the cell.
173. A method of translocating DNA sequences between two linear target DNA molecules of interest, the method comprising introducing into a cell: a nucleic acid editing system according to claims 103-152, wherein an attD attachment site of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) portion of the fusion polypeptide is present on a first linear target DNA molecule and an attA attachment site of the engineered Dn29 (or engineered Dn29 portion of the Dn29-DBD fusion) is present on a second linear target DNA molecule.
174. The method of claim 173, wherein the first target DNA molecules of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attA attachment site.
175. The method of claims 173-174, wherein the second target DNA molecules of interest of the cell was engineered before introduction of the nucleic acid editing system to contain an attD attachment site.
176. The method of claims 173-175, wherein the linear target DNA molecules of interest of the cell are chromosomes of the cell.
Citation Information
Patent Citations
Serine recombinases mediating stable integration into plant genomes
US20210395761A1
Serine recombinases
WO2023081762A2
Serine recombinase systems for site-specific gene editing
WO2023147507A1