Recombinase fusions
Fusion polypeptides combining LSR and DBD enhance DNA integration efficiency and specificity in mammalian genomes, addressing limitations of LSRs by enabling precise manipulation of larger DNA sequences.
Patent Information
- Application Number
- JP2025525706
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-28
- Filing Date
- 2023-11-01
- Publication Date
- 2025-11-14
AI Technical Summary
Large serine recombinases (LSRs) face challenges in efficiently integrating DNA payloads into mammalian genomes due to low insertion efficiency, high indel rates, and cargo size limitations, particularly for sequences larger than 1 kilobase, hindering advancements in synthetic biology and cell engineering.
Development of nucleic acids encoding fusion polypeptides comprising a large serine recombinase (LSR) portion fused N-terminally to a DNA binding domain (DBD) portion, often with a peptide linker and nuclear localization signals, and optionally guided by guide RNA, to enhance integration specificity and efficiency.
The fusion polypeptides significantly improve integration efficiency and specificity, enabling precise insertion, inversion, excision, and translocation of DNA sequences up to several kilobases in mammalian cells, including human cells.
Smart Images

Figure 2025537169000001_ABST
Abstract
Description
[Technical Field]
[0001] This international patent application claims the benefit of and priority to U.S. Patent Application No. 63 / 421,480, filed November 1, 2022, entitled "RECOMBINASES FOR INTEGRATING DNA," and U.S. Patent Application No. 63 / 516,424, filed July 28, 2023, entitled "DNA RECOMBINASE FUSIONS," the entire contents of which are incorporated herein by reference.
[0002] All patents, patent applications, and publications cited herein are incorporated by reference in their entirety, and the disclosures of these publications are incorporated by reference in their entirety into this application.
[0003] This patent disclosure contains material that is subject to copyright protection. The copyright owner has no objection to anyone reproducing this patent document or this patent disclosure, as it appears in the U.S. Patent and Trademark Office patent files or records, but otherwise reserves all copyright rights whatsoever.
[0004] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, and is incorporated herein by reference in its entirety. The XML copy, created on October 30, 2023, is named 2220476_00123WO1_SL.xml and is 620,017 bytes in size. [Background technology]
[0005] Large serine recombinases (LSRs) are DNA integrases encoded by bacteriophages that can promote site-specific and unidirectional integration of phage DNA into bacterial genomes through recombination of two binding sites, designated attP (phage) and attB (bacterial). These enzymes have been shown to integrate DNA payloads containing a donor binding site (attD, which can correspond to the native attP or attB) into mammalian cells either at pre-introduced integration sites or at endogenous genomic pseudosites with high sequence similarity to their corresponding acceptor binding sites (attA). When the attA sequence is present in the human genome, it is referred to as the attH sequence. However, despite its sequence specificity, LSRs can integrate into multiple sites in the human genome because multiple loci with the appropriate integration site sequence exist.
[0006] Engineering eukaryotic genomes, especially the integration of DNA sequences spanning several kilobases, remains challenging, limiting progress in rapidly developing fields such as synthetic biology and cell engineering. Challenges include low insertion efficiency, high indel rates, and cargo size limitations, with limited success for cargoes larger than 1 kilobase (kb). In particular, LSR can be limited by low integration efficiency. Therefore, improvements in genetic engineering systems and methods remain necessary. Summary of the Invention
[0007] It is understood that any of the embodiments described below can be combined in any desired manner, and that any embodiment or combination of embodiments can be applied to each of the aspects described below, unless the context indicates otherwise.
[0008]
[0010] In certain aspects, described herein are nucleic acids comprising a sequence encoding a fusion polypeptide, the fusion polypeptide comprising a large serine recombinase (LSR) portion and a DNA binding domain (DBD) portion. In some embodiments, the nucleic acid sequence encodes a fusion polypeptide in which the LSR portion is fused N-terminally to the DBD portion. In some embodiments, the nucleic acid sequence encoding the fusion polypeptide further comprises a nucleic acid sequence encoding a peptide linker positioned between the nucleic acid sequence encoding the LSR portion and the nucleic acid sequence encoding the DBD portion. In some embodiments, the nucleic acid sequence encodes a fusion polypeptide in which the LSR portion is fused N-terminally to the DBD portion by the peptide linker.
[0009] In some embodiments, the peptide linker encoded by the nucleic acid comprises at least one amino acid. In some embodiments, the peptide linker encoded by the nucleic acid comprises 2 to 100 amino acids. In some embodiments, the peptide linker encoded by the nucleic acid comprises 15 to 70 amino acids. In some embodiments, the peptide linker encoded by the nucleic acid comprises glycine and serine residues. In some embodiments, the peptide linker encoded by the nucleic acid comprises GGS repeats, GGSS (SEQ ID NO: 584) repeats, GGGS (SEQ ID NO: 572) repeats, or GGGGS (SEQ ID NO: 596) repeats. In some embodiments, the peptide linker encoded by the nucleic acid comprises one or more XTEN16 repeats. In some embodiments, the polypeptide linker encoded by the nucleic acid comprises one XTEN16 repeat, two XTEN16 repeats, or three XTEN16 repeats. In some embodiments, the polypeptide linker encoded by the nucleic acid comprises the amino acid sequence of SEQ ID NOs: 11-15. In some embodiments, the nucleic acid sequence encoding the polypeptide linker comprises SEQ ID NOs: 20-24.
[0010] In some embodiments, the LSR portion encoded by the nucleic acid comprises an amino acid sequence having at least 90% identity to SEQ ID NO: 1-5, 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, or 291. In some embodiments, the LSR portion encoded by the nucleic acid comprises the amino acid sequence of SEQ ID NO: 1-5, 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, or 291. In some embodiments, the LSR portion encoded by the nucleic acid comprises Dn29 (SEQ ID NO: 1), Pf80 (SEQ ID NO: 2), Cp36 (SEQ ID NO: 3), Nm60 (SEQ ID NO: 4), or Si74 (SEQ ID NO: 5). In some embodiments, the nucleic acid sequence encoding the LSR portion comprises a nucleic acid sequence having at least 90% identity to SEQ ID NOs: 6-10. In some embodiments, the nucleic acid sequence encoding the LSR portion comprises the nucleic acid sequence of SEQ ID NOs: 6-10.
[0011] In some embodiments, the fusion polypeptide encoded by the nucleic acid further comprises one or more nuclear localization signals (NLS). In some embodiments, the DBD portion encoded by the nucleic acid comprises Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12h, Cas12i, or Cas12g. In some embodiments, Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12h, Cas12i, or Cas12g does not have nuclease and / or nickase activity. In some embodiments, the DBD portion encoded by the nucleic acid comprises dCas9. In some embodiments, the DBD portion encoded by the nucleic acid comprises an amino acid sequence having at least 90% identity to dCas9 (SEQ ID NO:29), dCas9-HF1 (SEQ ID NO:30), dCas9-SpG (SEQ ID NO:31), or dCas9-SpG-HF1 (SEQ ID NO:32). In some embodiments, the DBD portion encoded by the nucleic acid comprises the amino acid sequence of dCas9 (SEQ ID NO:29), dCas9-HF1 (SEQ ID NO:30), dCas9-SpG (SEQ ID NO:31), or dCas9-SpG-HF1 (SEQ ID NO:32).
[0012] In some embodiments, the nucleic acid sequence encoding the DBD portion comprises a nucleic acid sequence having at least 90% identity to SEQ ID NOs: 33-36. In some embodiments, the nucleic acid sequence encoding the DBD portion comprises the nucleic acid sequence of SEQ ID NOs: 33-36.
[0013] In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), Pf80 (SEQ ID NO: 2) and dCas9 (SEQ ID NO: 29), Cp36 (SEQ ID NO: 3) and dCas9 (SEQ ID NO: 29), Nm60 (SEQ ID NO: 4) and dCas9 (SEQ ID NO: 29), or Si74 (SEQ ID NO: 5) and dCas9 (SEQ ID NO: 29). In some embodiments, the fusion polypeptide encoded by the nucleic acid further comprises a peptide linker located between the nucleic acid sequence encoding the LSR portion and the nucleic acid sequence encoding the DBD portion, wherein the LSR portion is fused to the N-terminus of the DBD portion by the peptide linker, and the peptide linker encoded by the nucleic acid comprises (GGS)8 (SEQ ID NO:11), (GGGGS)6 (SEQ ID NO:598), S(GGGGS)6S (SEQ ID NO:12), XTEN16 (SEQ ID NO:13), XTEN32-(GGSS)2 (SEQ ID NO:14), or XTEN48-(GGSS)2 (SEQ ID NO:15). In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises an amino acid sequence having at least 90% identity to SEQ ID NOs:37-42. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises the amino acid sequence of SEQ ID NOs:37-42. In some embodiments, the DBD portion of the fusion polypeptide encoded by the nucleic acid binds to a guide RNA (gRNA).
[0014] In certain aspects, described herein are vectors comprising any of the nucleic acids of the invention. In certain aspects, described herein are host cells comprising the vectors of the invention.
[0015] In one aspect, a nucleic acid editing system is described herein, comprising a first nucleic acid encoding an LSR-DBD described herein and a second nucleic acid encoding a gRNA. In some embodiments, the gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracrRNA portion, the nucleic acid sequence of the spacer sequence portion being identical to the target nucleic acid sequence except that T in the target nucleic acid sequence is replaced with U in the spacer sequence portion, and the target nucleic acid sequence is located within 80 nucleotides upstream or downstream of the dinucleotide core of the binding site for the LSR portion of the fusion polypeptide on the target DNA. In some embodiments, the spacer sequence portion is 16-20 nucleotides in length. In some embodiments, the gRNA encoded by the nucleic acid is an sgRNA. In some embodiments, a PAM sequence is present immediately 3' to the target nucleic acid sequence on the target DNA.
[0016] In some embodiments, the target nucleic acid sequence is located within 80 nucleotides upstream or downstream of the dinucleotide core at the attA site for the LSR portion of the fusion polypeptide on the target DNA of interest. In some embodiments, the attA site is a pseudo site in the mammalian target DNA of interest. In some embodiments, the attA site is a pseudo site (attH) in the human genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), where the attH site is chr10:21130404-21130406:-, chr11:77367459-77367461:-, chr1:230490334-230490336:+, chr2:14280297- In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises Pf80 (SEQ ID NO:2) and dCas9 (SEQ ID NO:29), and the attH site is chr11:64243293-64243295.
[0017] In some embodiments, the tracrRNA portion comprises SEQ ID NO: 153. In some embodiments, the target nucleic acid sequence is located on the DNA of interest within 80 nucleotides upstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide. In some embodiments, the target nucleic acid sequence is located on the DNA of interest within 80 nucleotides downstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide.
[0018] In some embodiments, the nucleic acid editing system further includes a third nucleic acid encoding a second gRNA. In some embodiments, the second gRNA encoded by the nucleic acid includes a spacer sequence portion and a tracr RNA portion, and the nucleic acid sequence of the spacer sequence portion is identical to the target nucleic acid sequence except that T in the target nucleic acid sequence is replaced with U in the spacer sequence portion, and the target nucleic acid sequence is located within 80 nucleotides downstream of the dinucleotide core of the binding site for the LSR portion of the fusion polypeptide on the target DNA. In some embodiments, the spacer sequence portion of the second gRNA is 16 to 20 nucleotides in length. In some embodiments, the second gRNA encoded by the nucleic acid is an sgRNA. In some embodiments, a PAM sequence is present immediately 3' to the target nucleic acid sequence on the target DNA.
[0019] In some embodiments, the nucleic acid editing system further comprises a third nucleic acid comprising a donor DNA sequence comprising an attD binding site for the LSR portion of the fusion polypeptide and a nucleic acid sequence for insertion into a target DNA of interest. In some embodiments, the third nucleic acid further comprises a portion of identity between the gRNA and the target nucleic acid sequence of the target DNA of interest.
[0020] In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises (a) Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), wherein the attH sites on the target DNA of interest are located at chromosomal loci chr10:21130404-21130406:-, chr11:77367459-77367461:-, chr1:230490334-230490336:+, chr2:14280297-14280299:+, chr9:116464427-116464429:+, chr20:38982599-38982599 or chr4:92338934-92338936:+, or the attH sequence found at the chromosomal locus, wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO: 154 or a sequence having 90% identity to SEQ ID NO: 154; (b) Pf80 (SEQ ID NO: 2) and dCas9 (SEQ ID NO: 29), wherein the attH site on the target DNA of interest is , chromosomal loci chr11:64243293-64243295:+, chr1:162878224-162878226:+, chr11:92763120-92763122:-, chr9:103309977-103309979:-, chr13:9114 5766-91145768:+, chr2:102467361-102467363:+, chr13:99865454-998 65456:+, chr9:113640780-113640782:-, chr9:123986548-123986550:- , chr15:53565450-53565452:-, or the attH sequence found at the chromosomal locus, wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO:265 or a sequence having 90% identity to SEQ ID NO:265; (c) comprising Cp36 (SEQ ID NO:3) and dCas9 (SEQ ID NO:29), wherein the attH site on the desired target DNA is at chromosomal locus chr16:2789124-2789126:+, chr22:43958465-43958467:-, chr10:117762740-117762742:+;chr7:157294532-157294534:-, chr13:20558930-20558932:-, chr6:151120348-151120350:-, chr10:101429887-101429889:+, chr1:20686551-20686553:+, chr19:50987430-50987432:+, chr4:183226741-183226743:-, or contains the attH sequence found at the chromosomal locus (d) comprising Nm60 (SEQ ID NO: 4) and dCas9 (SEQ ID NO: 29), wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO: 267 or a sequence having 90% identity to SEQ ID NO: 267, and wherein the attH site on the target DNA of interest is located at chromosomal loci chr9:83308042-83308044:-, chr13:79497139-79497141:-, chr9:131409759-131409761:+, chr4:55980785-55980787: +, chr5:96968267-96968269:+, chr6:37700280-37700282:-, chr19:17495840-17495842:-, chr5:126546219-126546221:+, chr10:15703649-15703651:-, chr10:395348-395350:+, or comprises the attH sequence found at the chromosomal locus, and the attD binding site of the donor DNA sequence is SEQ ID NO: 234 or SEQ ID NO: 2 or (e) Si74 (SEQ ID NO: 5) and dCas9 (SEQ ID NO: 29), wherein the attH site on the target DNA of interest is chromosomal locus chr7:155557356-155557358:+, chr9:77155112-77155114:-, or comprises the attH sequence found at the chromosomal locus; and the attD binding site of the donor DNA sequence comprises SEQ ID NO: 266 or a sequence having 90% identity to SEQ ID NO: 266.
[0021] In some embodiments, the third nucleic acid is a plasmid. In some embodiments, the third nucleic acid is a linear amplicon.
[0022] In certain aspects, described herein are vectors comprising any of the nucleic acids of the invention. In certain aspects, described herein are host cells comprising any of the vectors of the invention. In some embodiments, the nucleic acid encoding the fusion polypeptide, the nucleic acid encoding the gRNA, or both, and / or, if included, a third nucleic acid encoding a second gRNA, is expressed from an inducible promoter.
[0023] In some aspects, described herein are methods for integrating a donor DNA sequence into a target DNA of interest in a cell, the method comprising introducing a nucleic acid editing system of the invention into the cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a human embryonic stem cell. In some embodiments, the cell is a hepatocellular carcinoma cell. In some embodiments, the cell is a HEK cell. In some embodiments, the target DNA of interest in the cell is modified to contain an attA binding site prior to introduction of the nucleic acid editing system. In some embodiments, the donor DNA contains an LSR attD binding site that is integrated into the target DNA of interest. In some embodiments, the target DNA of interest in the cell is the genome of the cell. In some embodiments, the target DNA of interest in the cell is a plasmid.
[0024] In one aspect, described herein is a method for inverting a DNA sequence of a target DNA of interest, comprising introducing a nucleic acid editing system of the invention into a cell, wherein the attD binding site and the attA binding site for the LSR portion of the fusion polypeptide are located in opposite orientations on the same target DNA molecule of interest. In some embodiments, the target DNA of interest of the cell is modified to contain an attA binding site prior to introduction of the nucleic acid editing system. In some embodiments, the target DNA of interest of the cell is modified to contain an attD binding site prior to introduction of the nucleic acid editing system. In some embodiments, the target DNA of interest of the cell is the genome of the cell.
[0025] In certain aspects, described herein are methods for excising a DNA sequence in a target DNA of interest, comprising introducing a nucleic acid editing system of the invention into a cell, wherein the attD binding site and the attA binding site for the LSR portion of the fusion polypeptide are located in the same orientation on the same target DNA molecule of interest. In some embodiments, the target DNA of interest in the cell is modified to contain an attA binding site prior to introduction of the nucleic acid editing system. In some embodiments, the target DNA of interest in the cell is modified to contain an attD binding site prior to introduction of the nucleic acid editing system. In some embodiments, the target DNA of interest in the cell is the genome of the cell.
[0026] In one aspect, described herein is a method for translocating a DNA sequence between two linear target DNA molecules of interest, comprising introducing a nucleic acid editing system of the present invention into a cell, wherein an attD binding site for the LSR portion of the fusion polypeptide is located on a first linear target DNA molecule and an attA binding site for the LSR portion of the fusion polypeptide is located on a second linear target DNA molecule. In some embodiments, the first target DNA molecule of interest in the cell is modified to contain an attA binding site before introducing the nucleic acid editing system. In some embodiments, the second target DNA molecule of interest in the cell is modified to contain an attD binding site before introducing the nucleic acid editing system. In some embodiments, the linear target DNA molecule of interest in the cell is a chromosome of the cell.
[0027] Other embodiments of the present invention are further described in the following sections of this application, including the drawings, detailed description, examples, and claims. Furthermore, other objects and advantages of the present invention will be apparent to those skilled in the art from the disclosure herein, which is illustrative only and not limiting. Accordingly, other embodiments will be apparent to those skilled in the art without departing from the spirit and scope of the present invention.
[0028] This patent or application file contains at least one drawing executed in color. To conform to PCT patent application requirements, many of the figures presented herein are black-and-white representations of images originally executed in color. [Brief explanation of the drawings]
[0029] [Figure 1A] FIG. 1A is a schematic diagram of LSR-mediated, irreversible, kilobase-scale, site-specific genomic insertion between two DNA-binding sequences, attP and attB.
[0030] [Figure 1B] FIG. 1B shows that LSR can mediate integration into pre-introduced landing pads or endogenous pseudo-sites. Pseudo-sites can be empirically identified by expressing LSR and delivering a DNA cargo (e.g., a cargo containing a reporter gene) with a binding site into cells. If the DNA cargo is integrated into the genome, this genomic locus is determined to contain a pseudo-site. The genomic locus can be sequenced using techniques known in the art. For example, sequencing primers can be designed to target the integrated DNA cargo sequence so that sequence information of the genomic locus near the cargo can be obtained and analyzed for similarity to the binding site sequence of the DNA cargo construct that mediates integration.
[0031] [Figure 2] FIG. 2 shows that the RNA-guided DNA-binding domain co-localizes integrase to the genomic pseudosite (attH), leading to targeted integration of donor DNA via integrase-mediated recombination.
[0032] [Figure 3A] FIG. 3A shows that the LSR "Dn29" is a genome-targeted LSR with favorable efficiency and specificity. [Figure 3B]Figure 3B shows that the LSR "Dn29" is a genome-targeting LSR with favorable efficiency and specificity. 62% of integrations occurred at the top five sites. Figure 3B is an excerpt from Supplementary Figure 4E of Durrant, M.G., Fanton, A., Tycko, J. et al., Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome, Nat Biotechnol 41, 488-499 (2023).
[0033] [Figure 4] Figure 4 shows that LSR binds to attP and attB in a tetrameric complex. Figure adapted from Rutherford et al. Curr Opin Struct Biol. 2014.
[0034] [Figure 5] Figure 5 shows that the N-terminus of the LSR is important for tetrameric complex formation and subunit rotation. Figure modified from Rutherford et al. Curr Opin Struct Biol. 2014.
[0035] [Figure 6] Figure 6 shows exemplary designs of Dn29-dCas9 fusion constructs (see Figures 33-36 for sequences) and pseudo-site integration efficiencies at attH1 as measured by qPCR with untargeted guides. The data demonstrate that fusions with N-terminal Dn29 and C-terminal dCas9 have improved integration efficiencies compared to fusions with wild-type Dn29 and N-terminal dCas9.
[0036] [Figure 7]Figure 7 shows that the construct configuration is generalizable to another LSR, "Cp36." Figure 7 shows the pseudo-site integration efficiency at attH1 as measured by qPCR with a non-targeted guide. The data show that fusions with N-terminal Cp36 and C-terminal dCas9 have improved integration efficiency compared to fusions with wild-type Cp36 and N-terminal dCas9.
[0037] [Figure 8] Figure 8 shows a model of the LSR-dCas9 fusion construct, a tetrameric complex targeting a pseudosite in the genome with a single guide RNA. The guide RNA (shown as lines within the four outer lobes) is complementary to a genomic region adjacent to the integration site, resulting in only one dCas9 monomer binding to the genomic DNA (the outer lobe at the bottom left shows the gRNA hybridizing to the upstream sequence of the integration site), while the other three monomers are unbound.
[0038] [Figure 9] Figure 9 shows Dn29-dCas9 targeting attH1. The top diagram shows the location of the gRNA spacer and its targeting sequence relative to attH1. The bottom diagram shows the fold change in pseudo-site integration efficiency at attH1 as measured by qPCR compared to two non-targeting guide (NTG) controls.
[0039] [Figure 10] Figure 10 shows Dn29-dCas9-mediated cargo integration targeted to attH1, validated by orthogonal detection methods. The top panel shows attH1 integration measured by ddPCR. The bottom panel shows the total integration efficiency (at any genomic locus) by integration of an mCherry expression plasmid and flow cytometric detection of stable mCherry expression.
[0040] [Figure 11]Figure 11 shows Dn29-dCas9 targeting attH3. The top panel shows qPCR detection expressed as fold change compared to two untargeted guide controls. The bottom panel shows absolute efficiency measured by ddPCR.
[0041] [Figure 12] Figure 12 shows that another LSR ortholog (Pf80) can be targeted to pseudosites via dCas9 fusions. The upper left panel shows the relative integration efficiency of Pf80 into pseudosites in the human genome, with the most frequently integrated site (attH1) located at locus chr11:64,243,293. The upper right panel compares Pf80-dCas9 fusions with Pf80, showing the integration efficiency of attH1 when the gRNA is located adjacent to, overlapping, or within attH1. The lower panel shows SEQ ID NO: 534, which includes the location of the spacer sequence of each gRNA relative to the attH1 pseudosite. Generally, the spacer of the gRNA can be designed to target a sequence within 200, 175, 150, 125, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 10, or 5 nucleotides of the dinucleotide core sequence of the target binding site.
[0042] [Figure 13-1] Figure 13 shows that another LSR ortholog (Nm60) can be targeted to the pseudosite via dCas9 fusion. The top panel shows the integration efficiency of Nm60-dCas9 with each gRNA into the most frequently integrated pseudosite, located at chr9:83308042. The bottom panel shows SEQ ID NOS: 535-536 and the position of the spacer sequence of each gRNA relative to the attH1 pseudosite. [Figure 13-2] As described for Figure 13-1.
[0043] [Figure 14]Figure 14 shows that dCas9 fusions increase integration efficiency by up to 30% with attH1 and up to 8% with attH3 (left). The fold change relative to the untargeted guide ranges from 3 to 11 (right).
[0044] [Figure 15] Figure 15 is a schematic diagram of non-limiting embodiments of plasmids that can be used to achieve DNA insertion (top). The bottom diagram shows various molar ratios and integration rates of the three plasmids.
[0045] [Figure 16] Figure 16 is a schematic diagram of the introduction of a mixed population of target LSR-dCas9 fusions and unfused LSR monomers capable of forming a tetrameric complex.
[0046] [Figure 17] FIG. 17 shows that partial or complete separation of LSR and dCas9 reduces integration efficiency.
[0047] [Figure 18] Figure 18 shows the incorporation efficiency as a function of distance from the dinucleotide core. Distance is measured from the center of the dinucleotide core to the position (NGG) between the spacer and PAM. Data are cumulative for five experiments: Dn29-XTEN32-(GGSS)2-dCas9 for three pseudosites; Si74-XTEN32-(GGSS)2-dCas9 for attB, a landing pad located in AAVS1; and Pf80-XTEN32-(GGSS)2-dCas9 for attH1. "(GGSS)2" is disclosed as SEQ ID NO: 585.
[0048] [Figure 19]Figure 19 is a schematic diagram of one embodiment of a design change to optimize integration efficiency, showing that two dCas9s are targeted by two guide RNAs, one on each side of a pseudo-site, to promote recruitment and dimerization of the LSR at the genomic binding site.
[0049] [Figure 20] Figure 20 shows the integration rates measured by ddPCR for single-guide and multi-guide Dn29-dCas9 targeting of attH3. The last column of each plot represents the hypothetical value obtained by additively combining the single-guide integration efficiencies. The location of the spacer sequence of each gRNA relative to SEQ ID NO: 537 and the attH3 pseudosite is shown in the schematic diagram.
[0050] [Figure 21] Figure 21 shows single-guide and multi-guide Dn29-dCas9 targeting of attH1 as measured by qPCR. The locations of the gRNA binding sites are shown in a schematic diagram.
[0051] [Figure 22] Figure 22 is a schematic diagram of one embodiment of a design change to optimize integration efficiency. We demonstrate here that guide RNAs target the donor plasmid and facilitate its recruitment into the nucleus. In some embodiments, multiple guide RNAs can be used, including one or more different gRNAs targeting sequences adjacent to or adjacent and overlapping with the pseudosite, as shown in Figure 19, and one or more gRNAs targeting the donor plasmid.
[0052] [Figure 23-1]Figure 23 shows the integration efficiency as a fold change relative to a non-targeting guide when two guide RNAs are introduced, one targeting the pseudosite and the other targeting the donor plasmid. The only donor-targeting gRNA with a significant effect is guide 8. The location of the spacer sequence of each donor-targeting gRNA relative to SEQ ID NOs: 538-539 and attD is shown in the schematic diagram. [Figure 23-2] As described for Figure 23-1.
[0053] [Figure 24] Figure 24 shows the specificity of Dn29 and Dn29-dCas9 fusions. The left panel shows a plot of all integration sites detected in Dn29, ranked based on the number of UMI sequences at each locus. The most frequently integrated site is attH1, located at chr10:21,130,404. When Dn29-dCas9 targeting attH1 and guide 3 are used, the proportion of integrations occurring at attH1 increases to approximately 78% of all integrations (right panel).
[0054] [Figure 25] Figure 25 shows the specificity of Dn29-(GGGGS)6-dCas9 and Dn29-XTEN32-(GGSS)2-dCas9 targeting attH3, shown as the percentage of unique integrations (UMIs) occurring at that locus. "(GGGGS)6" and "(GGSS)2" are disclosed as SEQ ID NOs: 598 and 585, respectively.
[0055] [Figure 26]Figure 26 shows the correlation between specificity and efficiency across multiple guides for Dn29-dCas9 targeting. In the left panel, efficiency is measured by ddPCR for six guides targeting attH3, and specificity is measured by the percentage of UMIs occurring at the target pseudosite (attH3). Two different fusion construct designs are shown: Dn29-(GGGGS)6-dCas9 and Dn29-XTEN32-(GGSS)2-dCas9. "(GGGGS)6" and "(GGSS)2" are disclosed as SEQ ID NOs: 598 and 585, respectively. In the right panel, efficiency is measured by ddPCR for two guides targeting attH1 and a non-targeting guide, and specificity is measured by the percentage of UMIs occurring at attH1.
[0056] [Figure 27] Figure 27 shows a schematic diagram of a productive recombination reaction between attP and attB when the dinucleotide core is matched between the two sequences, compared to a non-productive recombination reaction between mismatched dinucleotide cores (bottom). In a non-productive reaction, ligation between the half-sites cannot occur, so the binding site returns to a second subunit rotation step, religating the original attP and attB. For recombination to be directional, the dinucleotide core must be non-palindromic.
[0057] [Figure 28]Figure 28 is a schematic diagram of the orientation of binding sites that result in integration, inversion, deletion, chromosomal translocation, and linear donor integration. In some embodiments, inversion or excision can be achieved using LSR fusions, including LSR-dCas9 fusions, by integrating a binding site near an endogenous binding site (including pseudosites). For inversion, the binding site must be integrated in the opposite orientation relative to the binding site in the target nucleic acid. For excision, the binding site must be integrated in the same orientation relative to the binding site in the target nucleic acid. In some embodiments, chromosomal translocation can be achieved using LSR fusions, including LSR-dCas9 fusions, by integrating a binding site on a different chromosome than the endogenous binding site (including pseudosites). In other embodiments, either circular or linear exogenous DNA fragments can be introduced with LSR fusions to achieve integration or linear donor integration. When linear donor integration occurs, the double-strand break that occurs after recombination with the linear amplicon is repaired by endogenous DNA repair pathways, such as non-homologous end joining.
[0058] [Figure 29] Figure 29 shows the integration efficiency of attH1 when the PAM-flexible dCas9 mutant, dCas9-SpG, is fused to Dn29. Guides targeting various NGG PAMs, which are likely targetable by both dCas9 and dCas9-SpG, and NGN PAMs, which are likely targetable only by dCas9-SpG, are shown. qPCR data normalized to dCas9 with a non-targeting guide are shown. For each guide, data for "Dn29-dCas" and then "Dn29-dCas9-SpG" are shown. The third data point for the guide "NTG" is the "LSR control with mismatch," and the fourth data point for the guide "NTG" is "Dn29."
[0059] [Figure 30]Figure 30 shows the same data set as Figure 29, highlighting the SpG-specific effect using fold changes normalized to the Dn29-dCas9 fusion construct with each guide. For each guide, data is shown for "Dn29-dCas" and then "Dn29-dCas9-SpG." The third data point for guide "NTG" is the "LSR control with mismatch," and the fourth data point for guide "NTG" is "Dn29."
[0060] [Figure 31] Figure 31 shows a schematic (top) and results (bottom) of a single-guide dual-targeting design in which a genomic protospacer (the DNA sequence targeted by the gRNA spacer) is included in the donor DNA molecule adjacent to attD, allowing for single-guide targeting of both the genome and the donor binding site. qPCR data normalized to a donor with attD and no protospacer are shown.
[0061] [Figure 32-1] Figure 32 shows examples of sequence logos for the Nm60 attB, Fm04 attB, Bt24 attB, and Dn29 attB binding sites. These motifs were generated by aligning the top 100 or top 300 genomic integration sites of the corresponding attP sequences. The height of the letter at each position indicates the enrichment of that nucleotide at that position. Other attB sequence motifs for the LSRs Cp36, Enc9, Pc01, Bt24, Dn29, Pf80, Sp36, and Enc3 are shown and described in Supplementary Figure 6C of Durrant, M.G., Fanton, A., Tycko, J. et al., Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome, Nat Biotechnol 41, 488-499 (2023), the entire contents of which are incorporated herein by reference. [Figure 32-2]As described for Figure 32-1.
[0062] [Figure 33-1] FIG. 33 is a diagram disclosing various sequences described herein. [Figure 33-2] As described for Figure 33-1. [Figure 33-3] As described for Figure 33-1. [Figure 33-4] As described for Figure 33-1. [Figure 33-5] As described for Figure 33-1. [Figure 33-6] As described for Figure 33-1. [Figure 34-1] FIG. 34 is a diagram disclosing various sequences described herein. [Figure 34-2] As described for Figure 34-1. [Figure 35-1] FIG. 35 is a diagram disclosing various sequences described herein. [Figure 35-2] As described for Figure 35-1. [Figure 35-3] As described for Figure 35-1. [Figure 35-4] As described for Figure 35-1. [Figure 35-5] As described for Figure 35-1. [Figure 35-6] As described for Figure 35-1. [Figure 35-7] As described for Figure 35-1. [Figure 35-8] As described for Figure 35-1. [Figure 35-9] As described for Figure 35-1. [Figure 36-1] FIG. 36 is a diagram disclosing various sequences described herein. [Figure 36-2] As described for Figure 36-1. [Figure 36-3] As described for Figure 36-1. [Figure 36-4] As described for Figure 36-1. [Figure 36-5] As described for Figure 36-1. [Figure 36-6] As described for Figure 36-1. [Figure 37-1] FIG. 37 is a diagram disclosing various sequences described herein. [Figure 37-2] As described for Figure 37-1. [Figure 37-3] As described for Figure 37-1. [Figure 37-4] As described for Figure 37-1. [Figure 37-5] As described for Figure 37-1. [Figure 37-6] As described for Figure 37-1. [Figure 38-1] FIG. 38 is a diagram disclosing various sequences described herein. [Figure 38-2] As described for Figure 38-1. [Figure 38-3] As described for Figure 38-1. [Figure 38-4] As described for Figure 38-1. [Figure 38-5] As described for Figure 38-1. [Figure 38-6] As described for Figure 38-1. [Figure 38-7] As described for Figure 38-1. [Figure 38-8] As described for Figure 38-1. [Figure 38-9] As described for Figure 38-1. [Figure 38-10] As described for Figure 38-1. [Figure 38-11] As described for Figure 38-1. [Figure 38-12] As described for Figure 38-1. [Figure 38-13] As described for Figure 38-1. [Figure 38-14] As described for Figure 38-1. [Figure 38-15] As described for Figure 38-1. [Figure 38-16] As described for Figure 38-1. [Figure 38-17] As described for Figure 38-1. [Figure 38-18] As described for Figure 38-1. [Figure 38-19] As described for Figure 38-1. [Figure 38-20] As described for Figure 38-1. [Figure 39-1] FIG. 39 is a diagram disclosing various sequences described herein. [Figure 39-2] As described for Figure 39-1. [Figure 39-3] As described for Figure 39-1. [Figure 39-4] As described for Figure 39-1. [Figure 39-5] As described for Figure 39-1. [Figure 39-6] As described for Figure 39-1. [Figure 40-1] FIG. 40 is a diagram disclosing various sequences described herein. [Figure 40-2] As described for Figure 40-1. [Figure 40-3] As described for Figure 40-1. [Figure 40-4] As described for Figure 40-1. [Figure 40-5] As described for Figure 40-1. [Figure 40-6] As described for Figure 40-1. [Figure 40-7] As described for Figure 40-1. [Figure 40-8] As described for Figure 40-1. [Figure 40-9] As described for Figure 40-1. [Figure 40-10] As described for Figure 40-1. [Figure 40-11] As described for Figure 40-1. [Figure 40-12] As described for Figure 40-1. [Figure 40-13] As described for Figure 40-1. [Figure 40-14] As described for Figure 40-1. [Figure 40-15] As described for Figure 40-1. [Figure 40-16] As described for Figure 40-1. [Figure 40-17] As described for Figure 40-1. [Figure 40-18] As described for Figure 40-1. [Figure 40-19] As described for Figure 40-1.
[0063] [Figure 41-1] Figure 41 shows Dn29-dCas9-mediated integration of a plasmid donor at attH1 in H1 human embryonic stem cells. [Figure 41-2] As described for Figure 41-1. [Figure 41-3] As described for Figure 41-1.
[0064] [Figure 42] Figure 42 shows Dn29-dCas9-mediated integration of the plasmid donor at attH1 in the HepG2 hepatocellular carcinoma cell line. DETAILED DESCRIPTION OF THE INVENTION
[0065] The present invention relates to the fusion of a large serine recombinase (LSR) to a DNA-binding domain (DBD). LSR recognizes two DNA sequences, also known as binding sites, one of which is a target site and the other is a DNA sequence that is often present on another DNA molecule. LSR performs site-specific recombination, integrating DNA present on another DNA molecule into the target site. Furthermore, when the binding sites are on the same molecule, LSR can perform excision or inversion recombination reactions depending on their relative orientation. Furthermore, when the binding sites are on different molecules in a specific relative orientation, translocation can occur. The DNA-binding domain is targeted to sites located adjacent to, overlapping with, or internal to the LSR target site via direct protein-DNA binding or RNA-guided targeting, and guides the LSR to a single specific DNA-binding site, such as a pseudosite in a mammalian genome. This design increases on-target integration efficiency by up to 30-fold compared to LSRs not fused to a DNA-binding domain, significantly increasing the ratio of on-target to off-target integration.
[0066] The term "cellular DNA" refers to, but is not limited to, genomic or non-genomic DNA present in a cell, or isolated forms of such DNA. Genomic or non-genomic DNA includes, but is not limited to, chromosomal DNA or non-chromosomal DNA such as episomal DNA, viral DNA, plasmid DNA, mitochondrial DNA, or chloroplast DNA.
[0067] The terms "polynucleotide," "nucleotide sequence," "nucleic acid," and "oligonucleotide" are used interchangeably. They refer to a polymeric form of nucleotides of any length, composed of either deoxyribonucleotides or ribonucleotides, or their analogs. Polynucleotides may have any three-dimensional structure and may perform any function, known or unknown. Coding or non-coding regions of a gene or gene fragment, loci defined by linkage analysis, exons, introns, guide RNA (gRNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers are non-limiting examples of polynucleotides. Polynucleotides may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. When modified nucleotides are included, the nucleotide structure may be modified before or after polymer formation. The sequence of nucleotides may be interrupted by non-nucleotide components. A polynucleotide may be further modified after polymerization, such as by conjugation with a labeling component.
[0068] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acids of any length. The polymers may be linear or branched, may comprise modified amino acids, and may be interrupted by non-amino acids. The term also encompasses amino acid polymers that have been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component. As used herein, the term "amino acid" includes natural and / or unnatural or synthetic amino acids, including glycine and both the D- and L-optical isomers, as well as amino acid analogs and peptidomimetics.
[0069] The practice of each aspect of the present invention may employ, unless otherwise indicated, conventional techniques of cell biology, cell culture, molecular biology, transgenic biology, microbiology, recombinant DNA, and biochemistry, which are within the skill of one of ordinary skill in the art. Such techniques are fully explained in the literature, see, for example, Molecular Cloning A Laboratory Manual, 3rd Edition, pp. 111-114, 1999, each of which is incorporated herein by reference in its entirety. rdEd., ed. by Sambrook (2001), Fritsch and Maniatis (Cold Spring Harbor Laboratory Press: 1989), DNA Cloning, Volumes I and II (D. N. Glover ed., 1985), Oligonucleotide Synthesis (M. J. Gait ed., 1984), Mullis et al. U.S. Patent No. 4,683,195; Nucleic Acid Hybridization (B. D. Hames & S. J. Higgins eds. 1984), Transcription and Translation (B. D. Hames & S. J. Higgins eds. 1984), Culture Of Animal Cells (R. I. Freshney, Alan R. Liss, Inc., 1987), Immobilized Cells and Enzymes (IRL Press, 1986), B. Perbal, A Practical Guide To Molecular Cloning (1984), Methods In Enzymology (Academic Press, Inc., N.Y.) series, particularly Methods In Enzymology, Vols. 154 and 155 (Wu et al. eds.), Gene Transfer Vectors For Mammalian Cells (J. H. Miller and M. P. Calos eds., 1987, Cold Spring Harbor Laboratory), Immunochemical Methods In Cell And Molecular Biology (Caner and Walker, eds., Academic Press, London, 1987), Handbook Of Experimental Immunology, Volumes I-IV (D.M. Weir and C.C. Blackwell, eds., 1986), Manipulating the Mouse Embryo, (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1986) and subsequent editions.
[0070] One skilled in the art can obtain proteins in several ways, including, but not limited to, isolating the protein by biochemical means or expressing a nucleotide sequence encoding the protein of interest by genetic engineering techniques, including, but not limited to, cell-based and cell-free techniques.
[0071] Proteins are encoded by nucleic acids (including, for example, genomic DNA, messenger RNA (mRNA), complementary DNA (cDNA), synthetic DNA, and any corresponding form of RNA). Nucleic acids encoding proteins can be produced by recombinant DNA technology, and such recombinant nucleic acids can be prepared by conventional techniques including chemical synthesis, genetic engineering, enzymatic technology, or a combination thereof.
[0072] LSR-DBD fusion
[0073] The present invention relates to the fusion of a large serine recombinase (LSR) to a DNA-binding domain (DBD), also referred to herein as an "LSR-DBD" fusion. In some embodiments, the LSR portion is fused directly to the DBD portion. In some embodiments, the LSR-DBD fusion includes a linker between the LSR and DBD portions of the fusion protein. The use of "LSR-DBD" is intended to encompass both embodiments unless otherwise specified (i.e., the "-" in "LSR-DBD" refers to either a direct bond or a linker between the LSR and DBD portions in an LSR-DBD fusion protein). Without being bound by theory, the fusions of the present invention target the LSR to a specific target site via the DNA-binding domain fusion, increasing the efficiency and specificity of the LSR. Without being bound by theory, these fusions may result in an increase in the local concentration of LSR monomers at the target DNA binding site, an increase in the residence time of LSR at the target DNA binding site, improved efficiency or kinetics of scanning the target DNA, and / or increased chromatin accessibility due to binding by two proteins at two sites.
[0074] The LSR, DBD, and linker moieties, if used, for use in LSR-DBD fusions are described below.
[0075] Large serine recombinase (LSR)
[0076] Recombinases (also called integrases) are a family of enzymes that mediate site-specific recombination between specific DNA sequences recognized by the enzyme. The original purpose of recombinases is to insert DNA, such as viral genomes or nonviral mobile genetic elements, into host cells, establishing the transition between the lytic and lysogenic cycles. Recombinases can be classified into two groups: tyrosine recombinases and serine recombinases, based on the active amino acid (tyrosine or serine) involved in the catalytic domain of the enzyme. Serine recombinases create a double-stranded break in DNA by forming a covalent 5'-phosphoserine bond with DNA, followed by strand exchange and ligation. On the other hand, tyrosine recombinases act by cleaving a single strand of DNA and forming a covalent 3'-phosphotyrosine bond with DNA, followed by a Holliday junction-like intermediate state.
[0077] As used herein, the term "recombinase" refers to a site-specific enzyme that mediates DNA recombination between recombinase recognition sequences, resulting in the excision, integration, inversion, or exchange (e.g., translocation) of the DNA fragment between the recombinase recognition sequences. Non-limiting examples of serine recombinases include Dn29, Pf80, Cp36, Nm60, Si74, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, and Ef0. 2, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01 , Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb27, Rh64, Rl09, Sa01 , Sa02, Sa10, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, and Sa11 (SEQ ID NOs: 1 to 5, 432 to 443, 445 to 446, 448 to 467, 469 to 476, and 478 to 492, respectively). Recombinases include, but are not limited to, large and small serine recombinases such as (494-501, 276, 279, 282, 285, 288, 291), Hin, Gin, Tn3, β-six, CinH, ParA, γδ, φC31, TP901, TG1, φBT1, R4, φRV1, φFC1, MR11, A118, U153, and gp29. Recombinases have many uses, including the generation of gene knockouts or knockins and gene therapy applications.
[0078] Large serine recombinases are efficient, directional, and specific recombinases for DNA integration in mammalian cells. See, e.g., Figure 1A. Examples of large serine recombinases useful in the nucleic acids, polypeptides, compositions, systems, and methods provided or disclosed herein include KSSJEB, PattyP, Doom, Scowl, Lockley, Switzer, Bob3, Trouble, Abrogate, Anglerfish, Sarfire, SkiPole, ConceptII, and Mu from recently sequenced mycobacteriophages. seum, Severus, Rey, Bongo, Airmid, Benedict, Theia, Hinder, Icleared, Sheen, Mundrea, Veracruz, and Rebeuca, as well as the already characterized Peaches, PhiC31, and BxZ2, as well as Dn29, Pf80, Cp36, Nm60, Si74, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, and Cd16 , Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma 05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb27, Rh64, Rl09, Sa01, Sa02 , Sa10, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, or Sal11 (SEQ ID NOS: 1-5, 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, and 291, respectively). In particular, LSR Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5) are useful in the nucleic acids, polypeptides, compositions, systems, and methods disclosed herein.
[0079] LSRs recognize two DNA sequences, also known as binding sites, one of which is a target site and the other is a DNA sequence present on another DNA molecule (an embodiment of integration). See Figure 28. LSRs perform site-specific recombination between two binding sites, as shown in Figure 1A. The natural binding sites targeted by LSRs are called "attP" (phage) and "attB" (bacterial) sites, and each of the attP and attB sites contains two half-sites connected by a central sequence. The central sequence consists of a dinucleotide core sequence, which is further described herein. Generally, recombination reactions are carried out by a tetramer of recombinase, with each subunit binding to either an attP or attB half-site, as shown in Figure 4. During the recombination reaction, each of the attP and attB sites is cleaved into two half-sites, each with an overhanging region (e.g., a dinucleotide core) that includes the central sequence. When applied to genomic integration of a donor cargo into a genome, the terms attD (donor) and attA (acceptor) are sometimes used to refer to the two binding sites. attP and attB can be attD or attA, depending on the sequences selected to be included on the donor molecule (e.g., if attP is attD, then attB is attA; if attB is attD, then attP is attA). In another embodiment, attD is integrated directly into an endogenous pseudosite naturally present in the target genome. As described in the Examples and in Durrant et al., NBT 2022, pseudosites can be experimentally identified by analyzing the sequences flanking successful integration of donor molecules bearing attD sites; i.e., the pseudosites will be adjacent to attD half-sites. When integration into a mammalian genome, e.g., a human genome, occurs, the endogenous pseudosite is referred to as attH. Thus, the attH site is a type of attA.
[0080] In some embodiments, LSR is used for site-specific recombination to exchange DNA strands between DNA sequences containing attB and attP sites (or attD and attA sites). Recombinases recognize and bind to the attB and attP sites, cleaving the DNA backbone, exchanging the two DNA helices involved, and rejoining the DNA strands, thereby rearranging the DNA fragments.
[0081] LSR can also site-specifically integrate target DNA sequences containing attD into DNA targets in mammalian cells, both at pre-introduced integration sites (e.g., pre-introduced attA) or endogenous genomic pseudo-sites (e.g., attH). For example, as shown in Figure 1B, a target donor DNA sequence containing a native attP site can be integrated into a DNA target with a corresponding native attB acceptor binding site (also called a "landing pad"). A target donor DNA sequence containing a native attB site can be integrated into a DNA target with a corresponding attP acceptor binding site (also called a "landing pad"). Mammalian DNA may also contain endogenous genomic pseudo-sites that have high sequence similarity to attA sites and are functionally recombinable with attD. When an attA sequence is present in a mammalian genome, e.g., the human genome, it is referred to as an attH sequence. For example, as shown in Figure 1B, a target donor DNA sequence containing a native attP site can be integrated into a DNA target with an attH pseudo-site that has high sequence similarity to the corresponding native attB acceptor binding site. A donor DNA sequence of interest containing a natural attB site can be integrated into a target DNA containing an attH pseudosite, which has high sequence similarity to the corresponding natural attP acceptor binding site. Therefore, LSRs can be used to integrate a target DNA sequence into target DNA, such as cellular DNA. Despite their sequence specificity, LSRs can integrate into numerous sites in mammalian genomes, such as the human genome, because multiple loci with the appropriate "attH" integration site sequence exist.
[0082] A systematic search for recombinases for integrating DNA into the human genome is described in WO 2023 / 081762 and Durrant MG, Fanton A, Tycko J, Hinks M, Chandrasekaran SS, Perry NT, Schaepe J, Du PP, Lotfy P, Bassik MC, Bintu L, Bhatt AS, Hsu PD. Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome. Nat Biotechnol., 41(4):488-499, Epub Oct 10, 2022, the contents of each of which are incorporated by reference in their entirety.
[0083] Exemplary LSRs that may be used in the LSR-DBD fusions described herein include, but are not limited to, the LSRs (Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively)) in Figure 33. The native attP and attB sequences for the LSRs in Figure 33 are provided as SEQ ID NOS: 304 (attP Cp36), 307 (attP Dn29), 328 (attP Nm60), 337 (attP Pf80), 353 (attP Si74), 374 (attB Cp36), 377 (attB Dn29), 398 (attB Nm60), 407 (attB Pf80), and 423 (attB Si74). In some embodiments, the binding site for the LSR portion of the LSR-DBD fusion, shown in Figure 32, i.e., Supplementary Figure 6C of Durrant, MG, Fanton, A., Tycko, J. et al. Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome, Nat Biotechnol 41, 488-499 (2023), the entire contents of which are incorporated herein by reference, comprises a sequence following the motif indicated by the consensus sequence logo for the corresponding LSR.
[0084] In certain aspects, described herein are LSR-DBD fusions comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOs: 1-5, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOs: 1-5, respectively). In some embodiments, the nucleic acid sequence encoding the LSR portion comprises SEQ ID NOs: 6-10.
[0085] In certain aspects, described herein are LSR-DBD fusions comprising an amino acid sequence having 70% identity to Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising an amino acid sequence having 70% identity to Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOs: 1-5, respectively).
[0086] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion consists of the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOs: 1-5, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion consists of the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOs: 1-5, respectively). In some embodiments, the nucleic acid sequence encoding the LSR portion consists of SEQ ID NOs: 6-10.
[0087] Other exemplary LSRs that may be used in the LSR-DBD fusions described herein include, but are not limited to, LSRs Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82.
[0088]
[0013] In one aspect, described herein are LSR-DBD fusions comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (SEQ ID NOs: 1-5, 433, 435, 437, 438, 445, 448, 457, 459, 462, 467, 469, 471, 479, 482, 495, 498, 499, 500, 501, respectively).
[0013] In one aspect, described herein are nucleic acids encoding an LSR-DBD fusion comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (SEQ ID NOs: 1-5, 433, 435, 437, 438, 445, 448, 457, 459, 462, 467, 469, 471, 479, 482, 495, 498, 499, 500, 501, respectively). In some embodiments, the nucleic acid sequence encoding the LSR portion comprises SEQ ID NOs: 6-10 or 515-533.
[0089]
[0013] In one aspect, described herein are LSR-DBD fusions comprising an amino acid sequence having 70% identity to Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (SEQ ID NOs: 1-5, 433, 435, 437, 438, 445, 448, 457, 459, 462, 467, 469, 471, 479, 482, 495, 498, 499, 500, 501, respectively). In some embodiments, the amino acid sequence is Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (or any of its equivalents). and 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NOs: 1 to 5, 433, 435, 437, 438, 445, 448, 457, 459, 462, 467, 469, 471, 479, 482, 495, 498, 499, 500, and 501, respectively.
[0013] In one aspect, described herein are nucleic acids encoding an LSR-DBD fusion comprising an amino acid sequence having 70% identity to Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (SEQ ID NOs: 1-5, 433, 435, 438, 445, 448, 457, 459, 462, 467, 469, 471, 482, 495, 498, 499, 500, 501, respectively).In some embodiments, the nucleic acid encoding the LSR-DBD fusion is selected from the group consisting of Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (or any of its variants). Each of these amino acid sequences has 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NOs: 1 to 5, 433, 435, 437, 438, 445, 448, 457, 459, 462, 467, 469, 471, 479, 482, 495, 498, 499, 500, and 501).
[0090]
[0013] In one aspect, described herein are LSR-DBD fusions, wherein the LSR portion consists of the amino acid sequence of Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (SEQ ID NOS: 1-5, 433, 435, 437, 438, 445, 448, 457, 459, 462, 467, 469, 471, 479, 482, 495, 498, 499, 500, and 501, respectively).
[0013] In one aspect, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion consists of the amino acid sequence of Dn29, Pf80, Cp36, Nm60, Si74, Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82 (SEQ ID NOS: 1-5, 433, 435, 437, 438, 445, 448, 457, 459, 462, 467, 469, 471, 479, 482, 495, 498, 499, 500, and 501, respectively). In some embodiments, the nucleic acid sequence encoding the LSR portion consists of SEQ ID NOs: 6-10 or 515-533.
[0091] Other exemplary LSRs that may be used in the LSR-DBD fusions described herein include the LSRs Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps 40, Ps45, Rb27, Rh64, Rl09, Sa01, Sa02, Sal0, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, or Sal11 (SEQ ID NOs: 432 to 443, 445 to 446, 448 to 467, 469 to 476, 478 to 492, 494 to 501, 276, 279, 282, 285, 288, and 291, respectively).
[0092] In one embodiment, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Described herein are LSR-DBD fusions comprising the amino acid sequence of Ps45, Rb27, Rh64, RlO9, SaOl, SaO2, SalO, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, TdOl, TdO8, uCb4, Vh19, Vh73, Vp82, CdO8, CMp1, ElOl, Pa19, Pg17, or Sal1 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, 291, respectively). In some embodiments, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45 , Rb27, Rh64, RlO9, SaOl, SaO2, SalO, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, TdOl, TdO8, uCb4, Vh19, Vh73, Vp82, CdO8, CMp1, ElOl, Pa19, Pg17, or Sal11 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, 291, respectively).
[0093] In some embodiments, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, R Described herein are LSR-DBD fusions comprising an amino acid sequence having 70% identity to b27, Rh64, RlO9, SaOl, SaO2, SalO, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, TdOl, TdO8, uCb4, Vh19, Vh73, Vp82, CdO8, CMp1, ElOl, Pa19, Pg17, or Sal1 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, and 291, respectively).In some embodiments, the amino acid sequence is selected from the group consisting of Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb 27, Rh64, Rl09, Sa01, Sa02, Sa10, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, or Sa11 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, 291, respectively).In one embodiment, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb27, Described herein are nucleic acids encoding LSR-DBD fusions comprising an amino acid sequence having 70% identity to Rh64, RlO9, SaOl, SaO2, SalO, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, TdOl, TdO8, uCb4, Vh19, Vh73, Vp82, CdO8, CMp1, ElOl, Pa19, Pg17, or Sal1 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, and 291, respectively).In some embodiments, the nucleic acid encoding the LSR-DBD fusion is selected from the group consisting of Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb2 7, Rh64, Rl09, Sa01, Sa02, Sa10, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, or Sa11 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, and 291, respectively).
[0094] In some embodiments, the LSR moiety is selected from the group consisting of Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps4 Described herein are LSR-DBD fusions consisting of the amino acid sequence of SEQ ID NOs: 0, Ps45, Rb27, Rh64, RlO9, SaOl, SaO2, SalO, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, TdOl, TdO8, uCb4, Vh19, Vh73, Vp82, CdO8, CMp1, ElOl, Pa19, Pg17, or Sal11 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, 291, respectively). In some embodiments, the LSR moiety is selected from the group consisting of Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps Described herein are nucleic acids encoding LSR-DBD fusions consisting of the amino acid sequence of SEQ ID NOs: 45, Rb27, Rh64, RlO9, Sa01, Sa02, SalO, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, ElO1, Pa19, Pg17, or Sal1 (SEQ ID NOs: 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, and 291, respectively).
[0095] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises an LSR means for mediating DNA recombination between recombinase recognition sequences. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises an LSR means for mediating DNA recombination between recombinase recognition sequences. In some embodiments, the LSR means for mediating DNA recombination between recombinase recognition sequences is selected from the group consisting of Dn29, Pf80, Cp36, Nm60, Si74, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No67, PaO1, PaO3, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb27, Rh64, Rl09, Sa01, Sa02, Sal0, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, or Sal11 (SEQ ID NOs: 1 to 5, 432 to 443, 445 to 446, 448 to 467, 469 to 476, 478 to 492, 494 to 501, 276, 279, 282, 285, 288, and 291, respectively).
[0096] Serine recombinases typically have a catalytic domain consisting of approximately 150 amino acid residues at the N-terminus. Several amino acids within the catalytic domain are highly conserved and are known to contribute to the structure of the active site. Serine recombinases also contain a C-terminal binding site for the catalytic domain, the size of which may vary. In LSRs, the site group may be a complex multi-domain region with both regulatory and DNA-binding functions.
[0097] In some embodiments, the LSR-DBD fusion comprises the catalytic domain of a large serine recombinase. By "catalytic domain of a large serine recombinase" is meant that the LSR-DBD fusion protein comprises a domain comprising amino acid sequences from (for example) a large serine recombinase, such that appropriate recombination occurs when the domain contacts a target nucleic acid (alone or in conjunction with other factors, including other large serine recombinase catalytic domains, which may or may not form part of the LSR-DBD fusion protein). In some embodiments, the catalytic domain of a large serine recombinase does not comprise the DNA-binding domain of a large serine recombinase. In some embodiments, the catalytic domain of a large serine recombinase comprises part or all of a large serine recombinase; for example, the catalytic domain may comprise a large serine recombinase domain and a DNA-binding domain or a portion thereof, or the catalytic domain may comprise a large serine recombinase domain and a DNA-binding domain mutated or truncated to abolish DNA-binding activity. Large serine recombinases and catalytic domains of large serine recombinases are known to those of skill in the art and include, for example, those described herein. In some embodiments, the catalytic domain is from any large serine recombinase.
[0098] In some embodiments, the LSRs used in the LSR-DBD fusions described herein include, but are not limited to, LSRs that contain one of the following amino acid motifs, written in the general Prosite format, where x is any amino acid and x(n) represents n any amino acids (e.g., x(3) represents xxx, i.e., 3 consecutive amino acids):
[0099] Motif 1:
[0100] [AEILSTVY]-[ADEGKQRST]-x(3)-[EG]-x-[ACFLMV]-x-[AFILMTV]-x(2)-[FHILMNV]-[AGSV]-[ADILSTV]-x-[AGS]-x (3)-[KRSV]-[ADEGKNST]-[AEIKMNQST]-[FILMST]-x-[DELQSV]-[ENQR]-x(4)-[AFHIKLMNQRSV]-x-[AEGHKLMNQRSV]
[0101] モチーフ2:
[0102] [AGI]-[DEGNPSTV]-[DGNQS]-[AHNQRTVY]-x-[ADEHILPQRTY]-[ADEQR]-[FIKL]-x-[DEFGNQRSTV]-[AILSTV]-[DEIKLNQRSTV]-[ADEKMNRSTV]-[AGQRST]-x-[ADEKLQRT]-x-[ALMV]
[0103] モチーフ3:
[0104] [ADFILMNSY]-x(2)-[AIKMSV]-x-[AFGILMV]-x(3)-[QRT]-[AGS]-x-[DEGNQS]-ESx-[AHKNRSTV]-Kx(2)-[LMRY]-[AINQSTV]-[AEFIKLNRTV]-x-[AFHLNQSTY]-[AILMNRSTVY]
[0105] モチーフ4:
[0106] [EKNTGSLDVARP]-[EHITGSLDVAP]-x-[MITSLVARP]-[EKNITGSDQVARP]-[EGSDARP]-[ILDAR]-[MHKTLVQDAR]-[EKITGSLDQVA]-[EKHDQVAR]-[MHISLVQAR]-[QEKNMSLDVAR]-[EKHGSLDQAR]-[EYKN IHLVA]-x-[EKITGSLDQAR]-[EKHTGDQAR]-x-[QEKNTGSDVAR]-[QEKNTGSVDAR]-[ISWLVFAR]-[QEM TGSLVDA]-[EKNITGSDARP]-[EMILDQA]-[EYILVFAR]-[EMTGSLDVAR]-[EKNGSLDQAR]-[QEGVDARP]
[0107] モチーフ5:
[0108] [ADEHKNQRS]-[ADEFGHKMNQRSWY]-[EFY]-[FHLWY]-x-[ADEFIKLMNQRSTY]-[FIQSTV]-[AGK LNRSTV]-[ADEHKNQRTY]-[INQR]-[FILMQS]-x(2)-[AGKNS]-[KMQRSTV]-x(2)-[AEGKMNSTY]
[0109] モチーフ6:
[0110] W-[AEHNRSTV]-x-[AGNST]-[FGLMNQSTV]-[ILPV]-x(2)-[ILTV]-x(4)-[ACGMQRST]-x-[ILVY]-G-[DEHNQS]-x-[EHILMQRT]-[AEFHLNPY]-[CFHKMNQRTY]-[DEFIKLNQRSTV]
[0111] Letter7:
[0112] [AGINSTV]-x-[AIS]-x-[FILMY]-E-[IR]-x(2)-[DILT]-x-[AEIKMQS]-R-[ITV]-x-[ADGRST]-x-[FKLMY]-[AEHIKLMNQRVWY]-x-[AIKLMR]
[0113] モチーフ8:
[0114] [FY]-[DEKQS]-[EKLMQ]-[KLR]-[KLV]-x-[GN]-[DEHKLMR]-[ST]-x-[FHIQSTVW]
[0115] モチーフ9:
[0116] [ILV]-x(2)-[ADFHILMNQSVY]-x(3)-[AGS]-x-[DEIKNQRS]-[EQ]-Sx(2)-[AK]-[AQRS]-x-[LMR]-[ILQRSV]-x-[ADEGHIQRS]-[AKNQSTV]-[AHKRWY]-x-[AGHIKQRST]-x-[CHIKLRV]
[0117] モチーフ10:
[0118] R-[LMQR]-[ANS]-[NPST]-W
[0119] モチーフ11:
[0120] [ILV]-[AV]-x-[AFHILQWY]-[IMV]-x-[ELQT]-[AIV]-F
[0121] モチーフ12:
[0122] R-[DKNRSV]-[ADEFGKPQS]-[AEIKLSTV]-x-[FGILNV]-[AFILQRVY]-[DEILMNQSTV]-[DEFILMQTVY]-[IKLRV]-[DEKNQR]-[DEFKLNQWY]-[FL]
[0123] モチーフ13:
[0124] [AEFILMNQSTVY]-[AFGILMRSTV]-x(3)-[ADEFGHLMNST]-x(2)-[DMNS]-[DEQ]-x-[CFHLTVY]-x-[AEKLRY]-x(2)-[ALS]-x-[DEKNQRS]-[GIMQRTV]-[DHKNQR]-x-[AGILNSTV]-[FHIKLMNQVWY]
[0125] Thus, in certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of one or more motifs selected from motif 1 through motif 13. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of one or more motifs selected from motif 1 through motif 13.
[0126] In one aspect, described herein is an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 2 and comprises an amino acid sequence having 70% identity to Si74 (SEQ ID NO: 5). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Si74 (SEQ ID NO: 5). In one aspect, described herein is a nucleic acid encoding an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 2 and comprises an amino acid sequence having 70% identity to Si74 (SEQ ID NO: 5). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Si74 (SEQ ID NO: 5).
[0127] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 3 and comprises an amino acid sequence having 70% identity to Bm99, Cs56, or Vp82 (SEQ ID NOs: 433, 445, and 501, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Bm99, Cs56, or Vp82 (SEQ ID NOs: 433, 445, and 501, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 3 and comprises an amino acid sequence having 70% identity to Bm99, Cs56, or Vp82 (SEQ ID NOs: 433, 445, and 501, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Bm99, Cs56, or Vp82 (SEQ ID NOs: 433, 445, 501, respectively).
[0128] In one aspect, described herein is an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 4 and comprises an amino acid sequence having 70% identity to Me99 (SEQ ID NO: 467). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Me99 (SEQ ID NO: 467). In one aspect, described herein is a nucleic acid encoding an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 4 and comprises an amino acid sequence having 70% identity to Me99 (SEQ ID NO: 467). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Me99 (SEQ ID NO: 467).
[0129] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 5 and comprises an amino acid sequence having 70% identity to Dn29, Nm60, or Bt24 (SEQ ID NOs: 1, 4, and 435, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Dn29, Nm60, or Bt24 (SEQ ID NOs: 1, 4, and 435, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 5 and comprises an amino acid sequence having 70% identity to Dn29, Nm60, or Bt24 (SEQ ID NOs: 1, 4, and 435, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Dn29, Nm60, or Bt24 (SEQ ID NOs: 1, 4, and 435, respectively).
[0130]
[0013] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 6 and comprises an amino acid sequence having 70% identity to Vh19 or Vh73 (SEQ ID NOs: 499 and 500, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Vh19 or Vh73 (SEQ ID NOs: 499 and 500, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 6 and comprises an amino acid sequence having 70% identity to Vh19 or Vh73 (SEQ ID NOs: 499 and 500, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Vh19 or Vh73 (SEQ ID NOs: 499 and 500, respectively).
[0131] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 7 and comprises an amino acid sequence having 70% identity to Fm04, uCb4, or Cb16 (SEQ ID NOs: 459, 498, 438, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Fm04, uCb4, or Cb16 (SEQ ID NOs: 459, 498, 438, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 7 and comprises an amino acid sequence having 70% identity to Fm04, uCb4, or Cb16 (SEQ ID NOs: 459, 498, 438, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Fm04, uCb4, or Cb16 (SEQ ID NOs: 459, 498, 438, respectively).
[0132] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 8 and comprises an amino acid sequence having 70% identity to Ec03 or Kp03 (SEQ ID NOs: 448, 462, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Ec03 or Kp03 (SEQ ID NOs: 448, 462, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 8 and comprises an amino acid sequence having 70% identity to Ec03 or Kp03 (SEQ ID NOs: 448, 462, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Ec03 or Kp03 (SEQ ID NOs: 448 and 462, respectively).
[0133] In one aspect, described herein is an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 9 and comprises an amino acid sequence having 70% identity to Pa03 (SEQ ID NO: 471). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Pa03 (SEQ ID NO: 471). In one aspect, described herein is a nucleic acid encoding an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 9 and comprises an amino acid sequence having 70% identity to Pa03 (SEQ ID NO: 471). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Pa03 (SEQ ID NO: 471).
[0134] In one aspect, described herein is an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 11 and comprises an amino acid sequence having 70% identity to Pf80 or Ps45 (SEQ ID NOs: 2, 482, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Pf80 or Ps45 (SEQ ID NOs: 2, 482, respectively). In one aspect, described herein is a nucleic acid encoding an LSR-DBD fusion, wherein the LSR portion comprises the amino acid sequence of motif 11 and comprises an amino acid sequence having 70% identity to Pf80 or Ps45 (SEQ ID NOs: 2, 482, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Pf80 or Ps45 (SEQ ID NOs: 2 and 482, respectively).
[0135]
[0010] In certain aspects, described herein are LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 13 and comprises an amino acid sequence having 70% identity to Cp36 (SEQ ID NO: 3). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to Cp36 (SEQ ID NO: 3). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the LSR portion comprises the amino acid sequence of motif 13 and comprises an amino acid sequence having 70% identity to Cp36 (SEQ ID NO: 3). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to Cp36 (SEQ ID NO: 3).
[0136] DNA binding domain (DBD)
[0137] As RNA-guided nucleases, Cas proteins have been applied to target gene editing and selection in various organisms. Nuclease-deficient Cas mutants that lack substantial nuclease activity are useful for localizing proteins and RNA to almost any dsDNA sequence.
[0138] In some embodiments, the DNA-binding domain of the LSR-DBD fusions described herein comprises a modified form of a Cas protein, such as, but not limited to, Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn1, Csn2, Cas4, Csm2, Cm5, Cas1, Cas2, Cas7, C2c3, C2c2, C2c1, or Cas5, that is complexed with a guide RNA. When the DBD is a Cas protein that is complexed with a guide RNA, the Cas protein can bind to the target DNA via the guide RNA spacer sequence, which base pairs with a complementary target DNA sequence located adjacent to, overlapping with, or within the recombinase target site. In some instances, the modified form of a Cas protein comprises amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the nuclease activity of the Cas protein. For example, in some instances, the modified form of a Cas protein has less than 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nuclease activity of the corresponding wild-type Cas protein. In preferred embodiments, the modified form of a Cas protein has no substantial nuclease activity. When the DBD of an LSR-DBD fusion is a modified form of a Cas protein that has no substantial nuclease activity, it can be referred to as "dead Cas" or "dCas." In some embodiments, the Cas protein may have nickase activity. In some embodiments, the modified form of a Cas protein has no substantial nickase activity. In some embodiments, the modified form of a Cas protein has no substantial nickase activity and no substantial nuclease activity.
[0139] Those skilled in the art will appreciate that Cas proteins can be isolated from a variety of bacterial species. In some embodiments, the DNA-binding domain of the LSR-DBD fusions described herein comprises a Cas protein from Streptococcus pyogenes, Staphylococcus aureus, Neisseria meningitidis, Campylobacter jejuni, Streptococcus thermophilus, Lachnospiraceae bacterium, Acidaminococcus sp., Alicyclobacillus acidiphilus, or Bacillus hisashii. In some embodiments, the DNA-binding domain of the LSR-DBD fusions described herein comprises Cas9 from Streptococcus pyogenes or its dCas9 form. In some embodiments, the DNA-binding domain of the LSR-DBD fusions described herein comprises Cas9 from Staphylococcus aureus or its dCas9 form.
[0140] In certain aspects, described herein are LSR-DBD fusions comprising the amino acid sequence of dCas9, Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn1, Csn2, Cas4, Csm2, Cm5, Cas1, Cas2, Cas7, C2c3, C2c2, C2c1, or Cas5. Described herein, in certain aspects, are nucleic acids encoding LSR-DBD fusions comprising the amino acid sequence of dCas9, Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas3, Cas8a-c, Cas10, Csel, Csy1, Csn1, Csn2, Cas4, Csm2, Cm5, Cas1, Cas2, Cas7, C2c3, C2c2, C2c1, or Cas5. In some embodiments, the DNA-binding domain of the LSR-DBD fusions described herein comprises dCas9 from Streptococcus pyogenes. In some embodiments, the DNA-binding domain of the LSR-DBD fusions described herein comprises dCas9 from Staphylococcus aureus.
[0141] In certain aspects, described herein are LSR-DBD fusions comprising the amino acid sequences of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOs: 29-32, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising the amino acid sequences of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOs: 29-32, respectively). In some embodiments, the nucleic acid sequence encoding the DBD portion comprises SEQ ID NOs: 33-36.
[0142] In certain aspects, described herein are LSR-DBD fusions comprising an amino acid sequence that is 70% identical to dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively). In some embodiments, the amino acid sequence is 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising an amino acid sequence having 70% identity to dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively). In some embodiments, the nucleic acid encoding the LSR-DBD fusion comprises an amino acid sequence having 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively).
[0143] Described herein in certain aspects are LSR-DBD fusions in which the DBD portion consists of the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively). Described herein in certain aspects are nucleic acids encoding LSR-DBD fusions in which the DBD portion consists of the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1. In some embodiments, the nucleic acid sequence encoding the DBD portion consists of SEQ ID NOS: 33-36.
[0144] In certain aspects, described herein are LSR-DBD fusions, wherein the DBD portion comprises DBD means for binding to a target DNA sequence located adjacent to, overlapping, or within a recombinase target site. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions, wherein the DBD portion comprises DBD means for binding to a target DNA sequence located adjacent to, overlapping, or within a recombinase target site. In some embodiments, the DBD means that binds to a target DNA sequence located adjacent to, overlapping with, or within a recombinase target site is dCas9, dCas9-HF1, dCas9-SpG, dCas9-SpG-HF1 (SEQ ID NOs: 29-32, respectively), Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas3, Cas8a-c, Cas10, Cse1, Csy1, Csn1, Csn2, Cas4, Csm2, Cm5, Cas1, Cas2, Cas7, C2c3, C2c2, C2c1, or Cas5.
[0145] In other embodiments, other DNA-binding domains (e.g., ZFPs or TALEs) may be used that bind to DNA target sites located adjacent to, overlapping with, or within the recombinase target site. In some embodiments, the DNA-binding domain binds to a DNA target nucleic acid sequence located on the DNA of interest within 200 nucleotides upstream or downstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide. In some embodiments, the DNA-binding domain binds to a DNA target nucleic acid sequence located on the DNA of interest within 100 nucleotides upstream or downstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide. In some embodiments, the DNA-binding domain binds to a DNA target nucleic acid sequence located on the DNA of interest within 80 nucleotides upstream or downstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide. In some embodiments, the DNA-binding domain binds to a DNA target nucleic acid sequence located on the DNA of interest within 50 nucleotides upstream or downstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide. In some embodiments, one of the two or more domains is a zinc finger (ZF) or TALE DNA-binding domain. A "zinc finger DNA-binding protein" (i.e., binding domain) is an amino acid sequence region within a protein, i.e., a binding domain, present within a larger protein that binds DNA in a sequence-specific manner via one or more zinc fingers whose structure is stabilized by coordination with zinc ions. The term zinc finger DNA-binding protein is often abbreviated as zinc finger protein or ZFP. A "TALE DNA-binding domain" or "TALE" is a polypeptide that contains one or more TALE repeat domains / units. The repeat domains are responsible for the binding of the TALE at its corresponding target DNA sequence. One "repeat unit" (also called a "repeat") typically consists of 33-35 amino acids and shows at least some sequence homology to other TALE repeat sequences in naturally occurring TALE proteins.Each TALE repeat unit contains one or two DNA-binding residues, typically located at positions 12 and / or 13 of the repeat, constituting a variable residue dimer (RVD). Zinc finger and TALE binding domains can be "engineered" to bind to a given nucleotide sequence, for example, by modifying (changing one or more amino acids in) the recognition helix region of a naturally occurring zinc finger or TALE protein. Engineered DNA-binding proteins (zinc finger or TALE) are therefore non-naturally occurring proteins.
[0146] Linker
[0147] In some embodiments, the fusion between the LSR and the DBD protein may include a linker. As used herein, the term "linker" refers to a chemical group or molecule that links two molecules or moieties, such as an LSR and a Cas protein. Typically, a linker is located between or sandwiched between two groups, molecules, or other moieties and is covalently bonded to each, thereby linking them together. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker may include a peptide or non-peptide moiety. In some embodiments, the linker is 2 to 100 amino acids in length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.
[0148] Exemplary linkers include flexible glycine-serine (GlySer or GS) linkers, e.g., for use in the LSR-DBD fusions described herein. In some embodiments, a "GGS" linker is used, with various repeats, e.g., 1 repeat (GGS), 2 repeats ((GGS)2) (SEQ ID NO: 562), 3 repeats ((GGS)3) (SEQ ID NO: 563), 4 repeats ((GGS)4) (SEQ ID NO: 564), 5 repeats ((GGS)5) (SEQ ID NO: 565), 6 repeats ((GGS)6) (SEQ ID NO: 566), 7 repeats ((GGS)7) (SEQ ID NO: 567), 8 repeats ((GGS)8) (SEQ ID NO: 11), 9 repeats ((GGS)9) (SEQ ID NO: 568), 10 repeats ((GGS) 10 ) (SEQ ID NO: 569), which is 11 repeats ((GGS) 11 ) (SEQ ID NO: 570), 12 repeats ((GGS) 12 ) (SEQ ID NO:571), or more repeats. In some embodiments, a "GGGS" linker (SEQ ID NO:572) is used, and various repeats can be used to provide the appropriate length as needed, for example, 1 repeat (GGGS) (SEQ ID NO:572), 2 repeats ((GGGS)2) (SEQ ID NO:573), 3 repeats ((GGGS)3) (SEQ ID NO:574), 4 repeats ((GGGS)4) (SEQ ID NO:575), 5 repeats ((GGGS)5) (SEQ ID NO:576), 6 repeats ((GGGS)6) (SEQ ID NO:577), 7 repeats ((GGGS)7) (SEQ ID NO:578), 8 repeats ((GGGS)8) (SEQ ID NO:579), 9 repeats ((GGGS)9) (SEQ ID NO:580), 10 repeats ((GGGS) 10 ) (SEQ ID NO: 581), 11 repeats ((GGGS) 11 ) (SEQ ID NO: 582), 12 repeats ((GGGS) 12) (SEQ ID NO: 583), or more repeats. In some embodiments, a "GGSS" linker (SEQ ID NO: 584) is used, and various repeats, such as 1 repeat (GGSS) (SEQ ID NO: 584), 2 repeats ((GGSS)2) (SEQ ID NO: 585), 3 repeats ((GGSS)3) (SEQ ID NO: 586), 4 repeats ((GGSS)4) (SEQ ID NO: 587), 5 repeats ((GGSS)5) (SEQ ID NO: 588), 6 repeats ((GGSS)6) (SEQ ID NO: 589), 7 repeats ((GGSS)7) (SEQ ID NO: 590), 8 repeats ((GGSS)8) (SEQ ID NO: 591), 9 repeats ((GGSS)9) (SEQ ID NO: 592), 10 repeats ((GGSS) 10 ) (SEQ ID NO: 593), which is 11 repeats ((GGSS) 11 ) (SEQ ID NO: 594), which is a 12-repeat ((GGSS) 12 ) (SEQ ID NO: 595), or more repeats. In some embodiments, a "GGGGS" linker (SEQ ID NO: 596) is used and can be used for various repeats to provide the appropriate length as needed, for example, 3 repeats ((GGGGS)3) (SEQ ID NO: 597), 6 repeats ((GGGGS)6) (SEQ ID NO: 598), 9 repeats ((GGGGS)9) (SEQ ID NO: 599), or 12 repeats ((GGGGS) 12 ) (SEQ ID NO: 600), or more repeats can be used. Other options include (GGGGS)1 (SEQ ID NO: 596), (GGGGS)2 (SEQ ID NO: 601), (GGGGS)4, (SEQ ID NO: 602), (GGGGS)5 (SEQ ID NO: 603), (GGGGS)7 (SEQ ID NO: 604), (GGGGS)8 (SEQ ID NO: 605), (GGGGS) 10 (SEQ ID NO: 606), or (GGGGS) 11(SEQ ID NO: 607). Additional glycine and / or serine residues can be included at the ends of the linker or between the various repeats, for example S(GGGGS)6S (SEQ ID NO: 12).
[0149] In some embodiments, an XTEN linker is used in the LSR-DBD fusions described herein. For example, in some embodiments, XTEN16 (SGSETPGTSESATPESS (SEQ ID NO: 13)) is used. In some embodiments, XTEN32, which has two XTEN16 repeats, or XTEN48, which has three XTEN16 repeats, is used. In some embodiments, additional XTEN16 repeats may be used to provide the appropriate length, if necessary.
[0150] In some embodiments, an alpha helical linker such as (Ala(GluAlaAlaAlaLys)Ala) (SEQ ID NO: 608) is contemplated for use in the LSR-DBD fusions described herein.
[0151] In some embodiments, (EAAAK)3 (SEQ ID NO: 609), (EAAAK) n (n=1-3) (SEQ ID NO: 610), A(EAAK)4(ALEA(EAAAK)4A (SEQ ID NO: 611), PAPAP (SEQ ID NO: 612), AEAAAKEAAAKA (SEQ ID NO: 613), (Ala-Pro) n (n=10-34) (SEQ ID NO: 614) are contemplated for use in the LSR-DBD fusions described herein. In some embodiments, cleavable linkers such as the disulfide bonds VSQTSKLTR|AETVFPDV (SEQ ID NO: 615), PLG|LWA (SEQ ID NO: 616), RVL|AEA (SEQ ID NO: 631), EDVVCC|SMSY (SEQ ID NO: 617), GGIER|GS (SEQ ID NO: 618), TRHRQPR|GWE (SEQ ID NO: 619), AGNRVRR|SVG (SEQ ID NO: 620), RRRRRRR|R|R (SEQ ID NO: 621) are contemplated for use in the LSR-DBD fusions described herein.
[0152] In some embodiments, a 2A self-cleaving peptide is used in the LSR-DBD fusions described herein. This peptide shares the core sequence motif of DXEXNPGP (SEQ ID NO: 622). In some embodiments, a T2A linker (GSG)EGRGSLLTCGDVEENPGP(S) (SEQ ID NO: 623) is used. In some embodiments, a P2A linker (GSG)ATNFSLLKQAGDVEENPGP(S) (SEQ ID NO: 624) is used. In some embodiments, an E2A linker (GSG)QCTNYALLKLAGDVESNPGP(S) (SEQ ID NO: 625) is used. In some embodiments, an F2A linker (GSG)VKQTLNFDLLKLAGDVESNPGP(S) (SEQ ID NO: 626) is used. The linker may optionally include a "GSG" residue at the N-terminus and an optional "S" residue at the C-terminus, as indicated in parentheses.
[0153] In some embodiments, linkers for use in the LSR-DBD fusions described herein may comprise a combination of one or more of the GlySer linkers, XTEN linkers, and / or 2A self-cleaving peptides described above. Exemplary, non-limiting linkers for use in the LSR-DBD fusions described herein are shown in Figure 34.
[0154] In other embodiments, the linker is at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 6 amino acids, at least 7 amino acids, at least 8 amino acids, at least 9 amino acids, at least 10 amino acids, at least 11 amino acids, at least 12 amino acids, at least 13 amino acids, at least 14 amino acids, at least 15 amino acids, at least 16 amino acids, at least 17 amino acids, at least 18 amino acids, at least 19 amino acids, at least 20 amino acids, at least 30 amino acids, at least 40 amino acids, at least 50 amino acids, at least 60 amino acids, at least 70 amino acids, at least 80 amino acids, at least 90 amino acids, at least 100 amino acids, at least 200 amino acids, at least 300 amino acids, at least 400 amino acids, or at least 500 amino acids in length.
[0155] In some embodiments, the LSR is fused directly to the DBD by a covalent bond. In some embodiments, the covalent bond is a carbon-carbon bond, a disulfide bond, a carbon-heteroatom bond, a carbon-nitrogen bond of an amide bond, or the like. In some embodiments, the LSR is fused to the DBD by a peptide or amino acid-based linker. In other embodiments, the linker is not peptidomimetic. In some embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, branched or unbranched, aliphatic or heteroaliphatic linker. In some embodiments, the linker is comprised of a polymer (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). In some embodiments, the linker comprises a monomer, dimer, or polymer of an aminoalkanoic acid. In some embodiments, the linker comprises an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta-alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.). In some embodiments, the linker comprises a monomer, dimer, or polymer of aminohexanoic acid (Ahx). In some embodiments, the linker is based on a carbocyclic moiety (e.g., cyclopentane, cyclohexane). In other embodiments, the linker comprises a polyethylene glycol moiety (PEG). In some embodiments, the linker comprises an aryl or heteroaryl moiety. In some embodiments, the linker is based on a phenyl ring. The linker may include a functionalized moiety to facilitate attachment of a nucleophilic group (e.g., thiol, amino) from the peptide to the linker. Any electrophile can be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, Michael acceptors, alkyl halides, aryl halides, acyl halides, and isothiocyanates. In other embodiments, the linker comprises an amino acid. In some embodiments, the linker comprises a peptide.
[0156] In some aspects, described herein are LSR-DBD fusions in which the LSR is fused directly to the DBD. In some aspects, described herein are nucleic acids encoding LSR-DBD fusions in which the LSR is fused directly to the DBD. In some aspects, described herein are LSR-DBD fusions comprising an LSR portion and a DBD portion fused to each other via a peptide linker. In some aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising an LSR portion and a DBD portion fused to each other via a peptide linker. In some embodiments, the peptide linker is 2-100 amino acids in length. In some embodiments, the peptide linker is 2-50 amino acids in length. In some embodiments, the peptide linker is 2-30 amino acids in length. In some embodiments, the peptide linker comprises glycine and serine residues. In some embodiments, the peptide linker comprises only glycine and serine residues. In some embodiments, the peptide linker is 2-30 amino acids in length and comprises only glycine and serine residues. In some embodiments, the peptide linker is 24 amino acids long and comprises only glycine and serine residues. In some embodiments, the peptide linker is 30 amino acids long and comprises only glycine and serine residues. In some embodiments, the peptide linker comprises GGS repeats. In some embodiments, the peptide linker comprises 2 to 12 GGS repeats (SEQ ID NO: 627). In some embodiments, the peptide linker consists of 2 to 12 GGS repeats (SEQ ID NO: 627). In some embodiments, the peptide linker comprises 8 GGS repeats (SEQ ID NO: 11). In some embodiments, the peptide linker consists of 8 GGS repeats (SEQ ID NO: 11). In some embodiments, the peptide linker comprises GGSS (SEQ ID NO: 584) repeats. In some embodiments, the peptide linker comprises 2 to 12 GGSS repeats (SEQ ID NO: 629). In some embodiments, the peptide linker consists of 2 to 12 GGSS repeats (SEQ ID NO: 629). In some embodiments, the peptide linker comprises two GGSS repeats (SEQ ID NO: 585).In some embodiments, the peptide linker comprises a GGGGS repeat (SEQ ID NO: 596). In some embodiments, the peptide linker comprises 2 to 12 GGGGS repeats (SEQ ID NO: 630). In some embodiments, the peptide linker consists of 2 to 12 GGGGS repeats (SEQ ID NO: 630). In some embodiments, the peptide linker comprises 6 GGGGS repeats (SEQ ID NO: 598). In some embodiments, the peptide linker consists of 6 GGGGS repeats (SEQ ID NO: 598). In some embodiments, the peptide linker comprises an XTEN16 sequence. In some embodiments, the peptide linker consists of an XTEN16 sequence. In some embodiments, the peptide linker comprises an XTEN32 sequence. In some embodiments, the peptide linker consists of an XTEN32 sequence. In some embodiments, the peptide linker comprises an XTEN48 sequence. In some embodiments, the peptide linker consists of an XTEN48 sequence. In some embodiments, the peptide linker comprises an F2A, E2A, P2A, or T2A sequence. In some embodiments, the peptide linker consists of an F2A, E2A, P2A, or T2A sequence. In some embodiments, the peptide linker comprises an XTEN16 sequence and one or more glycine or serine residues located at the N-terminus or C-terminus of the XTEN16 sequence. In some embodiments, the peptide linker comprises an XTEN32 sequence and one or more glycine or serine residues located at the N-terminus or C-terminus of the XTEN32 sequence. In some embodiments, the peptide linker comprises an XTEN48 sequence and one or more glycine or serine residues located at the N-terminus or C-terminus of the XTEN48 sequence. In some embodiments, the peptide linker comprises one or more XTEN16 sequences (e.g., XTEN16, XTEN32, XTEN48) and one or more GGSS (SEQ ID NO: 584), GGS, or GGGGS (SEQ ID NO: 596) repeats. In some embodiments, the peptide linker comprises one or more XTEN16 sequences (e.g., XTEN16, XTEN32, XTEN48) and one or more F2A, E2A, P2A, or T2A sequences.In some embodiments, the peptide linker comprises one or more GGSS (SEQ ID NO:584), GGS, or GGGGS (SEQ ID NO:596) repeats and one or more F2A, E2A, P2A, or T2A sequences. In some embodiments, the peptide linker comprises the amino acid sequence of SEQ ID NOs:11-19. In some embodiments, the nucleic acid sequence encoding the peptide linker portion comprises SEQ ID NOs:20-28.
[0157] In certain embodiments, described herein are LSR-DBD fusions that include a peptide linker means fusing the LSR portion and the DBD portion. In certain embodiments, described herein are nucleic acids that encode LSR-DBD fusions that include a peptide linker means fusing the LSR portion and the DBD portion.
[0158] In some embodiments, the fusion protein further comprises, consists essentially of, or consists of a localization (nuclear import or export) signal as, or as part of, a linker between the DBD (e.g., Cas enzyme) portion and the LSR portion. HA tags or Flag tags are also included within the scope of the present invention as linkers. The linker allows the user to design an appropriate degree of "mechanical flexibility."
[0159] Fusions constructed in either orientation are contemplated herein. Thus, in some embodiments, the LSR is fused C-terminal to the DBD. Alternatively, the LSR is fused N-terminal to the DBD. In another example, the LSR is fused to a position other than the C- or N-terminal end of the DBD, such as an internal residue of the DBD. Fusions in which the LSR is located N-terminally (e.g., LSR-dCas9 or LSR-linker-dCas9) are preferred over fusions in which the LSR is located C-terminally (e.g., dCas9-LSR or dCas9-linker-LSR). Thus, in certain aspects, LSR-DBD fusions in which the LSR portion is located N-terminally to the DBD portion are described herein. In certain aspects, nucleic acids encoding LSR-DBD fusions in which the LSR portion is located N-terminally to the DBD portion are described herein.
[0160] Longer linkers are also preferred; for example, Dn29-XTEN32-(GGSS)2-XTEN-dCas9 is preferred over Dn29-XTEN16-dCas9. Dn29-(GGGGS)6-dCas9 is preferred over Dn29-(GGS)8-dCas9. "(GGSS)2," "(GGGGS)6," and "(GGS)8" are disclosed as SEQ ID NOs: 585, 598, and 11, respectively. Linker flexibility is also a factor, with more flexible linkers (GGS and GGGGS (SEQ ID NO: 596)) being preferred over more rigid linkers (XTEN16) in dCas9-linker-Dn29 fusions.
[0161] In certain aspects, described herein are LSR-DBD fusions comprising any of the LSR and DBD portions described herein. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising any of the LSR and DBD portions described herein. In some embodiments, the LSR moiety is (a) Dn29, Pf80, Cp36, Nm60, Si74, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, No 67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb27, Rh64, Rl09, Sa01, Sa02, Sa10, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, or Sa11 (SEQ ID NOs: 1 to 5, 432 to 443, 445 to 446, 448 to 449, 450 to 451, 452 to 453, 454 to 455, 456 to 457, 458 to 459, 460 to 461, 462 to 463, 464 to 465, 466 to 467, 468 to 469, 470 to 471, 472 to 473, 474 to 475, 476 to 477, 478 to 479, 479 to 480, 481 to 482, 483 to 484, 485 to 486, 487 to 488, 489 to 490, 491 to 500, 501 to 502, 503 to 504, 505 to 506, 507 to 508, 509 to 510, 511 to 512, 513 to 514, 515 to (b) amino acid sequences having 70%, 75%, 80%, 85%, 90%, 70%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of (a), (c) amino acid sequences of motif 1 to motif 13, (d) amino acid sequences of motif 1 to motif 13, and amino acid sequences having 70%, 75%, 80%, 85%, 90%, 70%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of (a). 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity, or (e) an LSR means for mediating DNA recombination between recombinase recognition sequences, wherein the DBD portion comprises (f) Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas3, Cas8a-c, Cas10, Cse1, Csy1, Csn1, Csn2, Cas4,(g) the amino acid sequence of Csm2, Cm5, Cas1, Cas2, Cas7, C2c3, C2c2, C2c1, or Cas5; (g) the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, dCas9-SpG-HF1 (SEQ ID NOs: 29 to 32, respectively); (h) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of (f); or (i) a DBD means that binds to a target DNA sequence located adjacent to, overlapping, or within a recombinase target site.
[0162] Described herein in certain aspects are LSR-DBD fusions comprising an LSR portion comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively), and a DBD portion comprising the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively). Described herein in certain aspects are nucleic acids encoding LSR-DBD fusions comprising an LSR portion comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively), and a DBD portion comprising the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively). In some embodiments, the LSR portion comprises Dn29 (SEQ ID NO: 1) and the DBD portion comprises dCas9 (SEQ ID NO: 29). In some embodiments, the LSR portion comprises Pf80 (SEQ ID NO: 2) and the DBD portion comprises dCas9 (SEQ ID NO: 29). In some embodiments, the LSR portion comprises Cp36 (SEQ ID NO: 3) and the DBD portion comprises dCas9 (SEQ ID NO: 29). In some embodiments, the LSR portion comprises Nm60 (SEQ ID NO: 4) and the DBD portion comprises dCas9 (SEQ ID NO: 29). In some embodiments, the LSR portion comprises Si74 (SEQ ID NO: 5) and the DBD portion comprises dCas9 (SEQ ID NO: 29). In some embodiments, the amino acid sequence of the LSR portion is 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively). In some embodiments, the amino acid sequence of the DBD portion is 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively).
[0163] In certain aspects, described herein are LSR-DBD fusions comprising an LSR portion comprising an LSR means for mediating DNA recombination between recombinase recognition sequences, and a DBD portion comprising DBD means for binding to a target DNA sequence located adjacent to, overlapping, or within a recombinase target site. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising an LSR portion comprising an LSR means for mediating DNA recombination between recombinase recognition sequences, and a DBD portion comprising DBD means for binding to a target DNA sequence located adjacent to, overlapping, or within a recombinase target site.
[0164] In certain aspects, described herein are LSR-DBD fusions comprising any of the LSR portions, DBD portions, and linker portions described herein. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising any of the LSR portions, DBD portions, and linker portions described herein. In some embodiments, the LSR moiety is a) Dn29, Pf80, Cp36, Nm60, Si74, Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cs56, Ct03, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99 , No67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb27, Rh64, Rl09, Sa01, Sa02, Sa10, Sa34, Sa51, Se37, Sh25, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, Vp82, Cd08, CMp1, El01, Pa19, Pg17, or Sa11 (SEQ ID NOs: 1 to 5, 432 to 443, 445 to 446, respectively). 446, 448 to 467, 469 to 476, 478 to 492, 494 to 501, 276, 279, 282, 285, 288, 291), b) amino acid sequences having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of a), c) amino acid sequences of motif 1 to motif 13, d) amino acid sequences of motif 1 to motif 13, and amino acid sequences having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of a). , 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity, or e) an LSR means for mediating DNA recombination between recombinase recognition sequences, and the DBD portion comprises f) Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas3, Cas8a to c, Cas10, Cse1, Csy1, Csn1, Csn2,g) the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOs: 29-32, respectively); h) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of f); or i) a DBD means for binding to a target DNA sequence located adjacent to, overlapping with, or within a recombinase target site, and a linker The moiety comprises j) a peptide linker, k) a peptide linker comprising one or more glycine-serine repeats, l) a peptide linker comprising one or more XTEN linkers, m) a peptide linker comprising one or more glycine-serine repeats and one or more XTEN linkers, n) an amino acid sequence comprising SEQ ID NOs: 11-19, o) an amino acid sequence having 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity to the amino acid sequence of k), or p) a peptide linker means fusing the LSR moiety and the DBD moiety.
[0165] In certain aspects, described herein are LSR-DBD fusions comprising an LSR portion comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOS: 1-5, respectively), a DBD portion comprising the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOS: 29-32, respectively), and a linker portion comprising the amino acid sequence of SEQ ID NOS: 11-19. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising an LSR portion comprising the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NOs: 1-5, respectively), a DBD portion comprising the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOs: 29-32, respectively), and a linker portion comprising the amino acid sequence of SEQ ID NOs: 11-19. In some embodiments, the LSR portion comprises Dn29 (SEQ ID NO: 1), the DBD portion comprises dCas9 (SEQ ID NO: 29), and the linker portion comprises the amino acid sequence of SEQ ID NOs: 11-19. In some embodiments, the LSR portion comprises Pf80 (SEQ ID NO: 2), the DBD portion comprises dCas9 (SEQ ID NO: 29), and the linker portion comprises the amino acid sequence of SEQ ID NOs: 11-19. In some embodiments, the LSR portion comprises Cp36 (SEQ ID NO:3), the DBD portion comprises dCas9 (SEQ ID NO:29), and the linker portion comprises the amino acid sequence of SEQ ID NO:11-19. In some embodiments, the LSR portion comprises Nm60 (SEQ ID NO:4), the DBD portion comprises dCas9 (SEQ ID NO:29), and the linker portion comprises the amino acid sequence of SEQ ID NO:11-19. In some embodiments, the LSR portion comprises Si74 (SEQ ID NO:5), the DBD portion comprises dCas9 (SEQ ID NO:29), and the linker portion comprises the amino acid sequence of SEQ ID NO:11-19. In some embodiments, the amino acid sequence of the LSR portion is 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of Dn29, Pf80, Cp36, Nm60, or Si74 (SEQ ID NO:1-5, respectively).In some embodiments, the amino acid sequence of the DBD portion is 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of dCas9, dCas9-HF1, dCas9-SpG, or dCas9-SpG-HF1 (SEQ ID NOs: 29-32, respectively). In some embodiments, the amino acid sequence of the linker portion is 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of SEQ ID NOs: 11-19.
[0166] In certain aspects, described herein are LSR-DBD fusions comprising an LSR portion comprising an LSR means for mediating DNA recombination between recombinase recognition sequences, a DBD portion comprising DBD means for binding to a target DNA sequence located adjacent to, overlapping, or within a recombinase target site, and a peptide linker means fusing the LSR portion and the DBD portion. In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising an LSR portion comprising an LSR means for mediating DNA recombination between recombinase recognition sequences, a DBD portion comprising DBD means for binding to a target DNA sequence located adjacent to, overlapping, or within a recombinase target site, and a peptide linker means fusing the LSR portion and the DBD portion.
[0167] In certain aspects, described herein are LSR-DBD fusions comprising the amino acid sequence set forth in Figure 36 (SEQ ID NOS: 37-42). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions comprising the amino acid sequence set forth in Figure 36 (SEQ ID NOS: 37-42). In some embodiments, the amino acid sequence is 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the amino acid sequence of SEQ ID NOS: 37-42. In certain aspects, described herein are LSR-DBD fusions consisting of the amino acid sequence set forth in Figure 36 (SEQ ID NOS: 37-42). In certain aspects, described herein are nucleic acids encoding LSR-DBD fusions consisting of the amino acid sequence set forth in Figure 36 (SEQ ID NOS: 37-42).
[0168] In any embodiment described herein, the nucleotide sequence encoding the LSR-DBD fusion polypeptide, or its LSR portion, DBD portion, and / or linker portion, may be codon-optimized. This type of optimization is known in the art and involves introducing mutations into exogenous DNA to resemble the codon usage preferences of the intended host organism or cell while still encoding the same protein. Thus, the codons are altered, but the encoded protein remains the same. For example, if the intended target cell is a human cell, a Cas protein (or variant, e.g., dCas) codon-optimized for human use would be a suitable DBD. Any suitable DBD may be codon-optimized. As another non-limiting example, if the intended host cell is a mouse cell, a Cas protein (or variant, e.g., dCas) codon-optimized for mouse use would be a suitable DBD. While codon optimization is not required, it is permissible and may even be preferred in some cases.
[0169] Protein-mediated recruitment refers to the process of fusing the DBD and LSR to two interacting protein domains, expressing each protein in trans, and then assembling them to function as a fusion protein. Some exemplary systems include, but are not limited to, SunTag (a protein scaffold containing a peptide epitope fused to a dCas9 protein). The LSR can be fused to a single-chain variable fragment (scFV) antibody, which recruits to the peptide epitope when introduced in trans. Alternatively, SpyTag (a 13-residue peptide called Spytag) and a 116-residue complementary domain can be fused to the DBD and LSR, respectively, and spontaneously associate to form a covalent isopeptide bond when introduced in trans. Furthermore, coiled-coil peptide heterodimers, SnoopTag, or SnoopCatcher, can also be used.
[0170] Inducible recruitment refers to the dimerization of the DBD and LSR by fusion with an inducible binding protein, such as a small molecule or light, resulting in the recruitment of the LSR to the DBD (e.g., dCas9). Examples of such systems include FK506-binding protein 12 (FKBP) and the FKBP rapamycin-binding (FRB) domain, which dimerize upon rapamycin induction; pMag and nMag, which dimerize upon blue light irradiation; and DmrA / DmrC, which dimerize in the presence of a rapamycin analog (known as the A / C heterodimerizer).
[0171] LSR donor and acceptor binding sites
[0172] The recombination sites for the LSRs of the LSR-DBD fusions described herein typically consist of 30-200 nucleotides and contain two motifs with partial inverted repeat symmetry flanking a central crossover sequence where recombination occurs. These inverted repeat sequences, specific for each recombinase, bind to the recombinase and are referred to herein as "recombinase recognition sequences," "recombinase recognition sites," "attP sites," "attB sites," "attD sites," "attH sites," "attA sites," "binding sites," "pseudo sites," "genomic pseudo sites," or "genomic insertion sites." In some embodiments, the attB site is present in the target DNA sequence (e.g., cellular DNA) and the attP site is present in the DNA sequence to be integrated into the target DNA sequence. In some embodiments, the attP site is present in the target DNA sequence (e.g., cellular DNA) and the attB site is present in the DNA sequence to be integrated into the target DNA sequence. As disclosed herein, "attD" refers to a donor binding site, which can be an attP or attB site; "attA" refers to the corresponding acceptor site; and "attH" refers to an integration site naturally occurring in a mammalian genome, e.g., the human genome. A "landing pad" is an exogenous DNA sequence containing a binding site for an LSR that is integrated into a target DNA location. Landing pads can be incorporated into target DNA using any technique known in the art, such as by using zinc finger nucleases, TALENs, or CRISPR-Cas systems, or by using the LSR-DBD fusions described herein.
[0173] During recombination, crossover occurs at the dinucleotide core of the attB / attP site. The dinucleotide core sequence is the only factor that determines the directionality of recombination. For directional recombination, the dinucleotide core must be non-palindromic. See Figure 26. For example, the dinucleotide core sequence present within the attB / attP site of a strictly directional large serine recombinase can be AA, TT, GG, CC, AG, GA, AC, CA, TG, GT, TC, or CT. A schematic diagram is shown in Figure 27.
[0174] The outcome of recombination depends, in part, on the location and orientation of the binding sites. For example, inversion recombination occurs between two oppositely oriented binding sites located on the same DNA molecule. DNA loop formation brings the two binding sites into close proximity, at which point DNA cleavage, strand exchange, and ligation occur. This reaction is ATP-independent. The net result of such an inversion recombination event is that the DNA fragment flanked by the repeats is inverted (i.e., the orientation of the DNA fragment is reversed), thereby converting the original coding strand to a non-coding strand, and vice versa. In such a reaction, DNA is preserved and no substantial gain or loss of DNA occurs. Conversely, excision recombination occurs between two binding sites arranged in the same orientation on the same DNA molecule. In this case, the intervening DNA is excised / removed. Integrative recombination can occur between two binding sites located on different DNA molecules, one of which is circular (due to integration of the entire circular molecule). If the other DNA molecule is cellular or genomic DNA, the two molecules combine to form a single molecule, and the circular DNA is integrated into the cellular or genomic DNA. Translocations also occur upon recombination of two binding sites present on different linear DNA molecules. A schematic representation of insertion / integration, excision, inversion, and translocation is shown in Figure 28.
[0175] The LSR has two binding sites to which the LSR binds and performs sequence-specific recombination. In some embodiments, target DNA into which the binding sites have been introduced is targeted. In another embodiment, to target a sequence that is endogenously present in the target DNA, a sequence similar to the desired binding site sequence must be present in the target DNA, such as genomic DNA or other cellular DNA. In some embodiments, an LSR capable of targeting an endogenous sequence may be used in an LSR-DBD fusion. Another potentially relevant factor is the number of endogenous sites into which the LSR can integrate. Reducing (but not eliminating) the number of integration sites may improve the integration efficiency per pseudo-site by reducing off-target sites (which can act as sinks for the LSR) that can impair on-target efficiency. Thus, in some embodiments, an LSR capable of targeting one or up to several thousand endogenous sequences may be used in an LSR-DBD fusion.
[0176] Guide polynucleotide
[0177] When a Cas protein domain is used as the DBD portion of an LSR-DBD fusion, the Cas portion can bind to one or more guide RNAs (gRNAs) having spacer sequences, including but not limited to, those depicted in Figure 37, thereby guiding or directing the LSR-DBD fusion to a target nucleic acid of interest. In some embodiments, a guide RNA is used that targets a target sequence located on a desired acceptor target DNA. In some embodiments, a guide RNA is used that targets a target sequence located on a desired donor DNA. In some embodiments, the systems described herein use two guide RNAs, one that targets a target sequence located on a desired acceptor target DNA and one that targets a target sequence located on a desired donor DNA. In some embodiments, the systems described herein use two guide RNAs, one that targets a target sequence located on a desired acceptor target DNA and one that targets a target sequence located on a second desired acceptor target DNA. In some embodiments, the LSR binding site in the target DNA of interest is flanked by a first and a second target sequence on the desired acceptor target DNA. In some embodiments, guide RNAs are used that target a target sequence located on the intended acceptor target DNA and a target sequence located on the intended donor DNA, and the target sequences are identical. For example, the target sequence targeted by the guide in the intended acceptor target DNA is present on the donor DNA molecule and is located adjacent to, overlapping with, or within an attD site. In some embodiments, three or more guide RNA sequences are used, for example, one or more guide RNA sequences that target one or more target sequences located on the intended donor DNA molecule and one or more guide RNA sequences that target one or more target sequences located on the intended acceptor target DNA.
[0178] As used herein, the term "guide polynucleotide" or "guide RNA" or "gRNA" refers to a polynucleotide sequence that can form a complex with a Cas protein and enable the Cas protein to recognize, bind to, and optionally cleave a DNA target site. A guide RNA is a specific RNA sequence that recognizes a target DNA region of interest and guides the Cas protein, and thus the LSR-DBD fusion, to that site. A gRNA typically consists of two parts: a CRISPR RNA (crRNA) (also called a gRNA spacer or spacer sequence), a nucleotide sequence that binds to the complementary sequence of the target DNA sequence, and a trans-activating CRISPR RNA (tracrRNA), which serves as a scaffold for binding to the Cas protein. In the context of CRISPR, hybridization between the complementary sequence of the target sequence and the gRNA spacer sequence promotes the formation of a CRISPR complex. Perfect complementarity is not necessarily required, as long as there is sufficient complementarity to cause hybridization and promote the formation of a CRISPR complex.
[0179] While crRNA and tracrRNA exist naturally as two separate RNA molecules, a single RNA molecule can also exist that can contain both the crRNA sequence and the scaffold tracrRNA sequence fused together (called a single guide RNA (sgRNA)). In some embodiments, the gRNA is an sgRNA. In some embodiments, the gRNA comprises two separate RNA molecules. The guide polynucleotide sequence can be an RNA sequence, a DNA sequence, or a combination thereof (a mixed RNA-DNA sequence), such as the system from Caribou Biosciences that uses the "chRDNA" system, in which the guide polynucleotide is an RNA / DNA hybrid system. Optionally, the guide polynucleotide can include at least one nucleotide modification, phosphodiester bond modification, or phosphodiester linkage modification, such as, but not limited to, locked nucleic acid (LNA), 5-methyl dC, 2,6-diaminopurine, 2'-fluoro A, 2'-fluoro U, 2'-O-methyl RNA, phosphorothioate bond, linkage to a cholesterol molecule, linkage to a polyethylene glycol molecule, linkage to a spacer 18 (hexaethylene glycol chain) molecule, or circularization by covalent linkage between the 5' and 3' ends. See also U.S. Patent Application Publication No. 2015-0082478(A1), published March 19, 2015, and U.S. Patent Application Publication No. 2015-0059010(A1), published February 26, 2015, the contents of each of which are incorporated herein by reference in their entirety.
[0180] In one embodiment of the present disclosure, the guide polynucleotide is an sgRNA capable of forming a guide RNA / protein RNP complex with the DBD of an LSR-DBD fusion disclosed herein, and the RNP complex is capable of recognizing and binding to a complementary sequence of a target sequence. The target sequence or sequences may be contained in the intended acceptor target DNA, the intended donor DNA, or both.
[0181] In one embodiment of the present disclosure, the guide polynucleotide is an sgRNA capable of forming a guide RNA / protein RNP complex with the DBD of the LSR-DBD fusion disclosed herein, wherein the complex is capable of recognizing and binding to the complementary sequence of the target sequence, and the sgRNA comprises a "crRNA" or "spacer" or "spacer sequence" linked to a "scaffold" or "scaffold sequence" or "tracrRNA." The one or more target sequences may be contained in the intended acceptor target DNA, the intended donor DNA, or both.
[0182] In one embodiment of the present disclosure, the guide polynucleotide is an sgRNA capable of forming a guide RNA / protein RNP complex with the DBD of the LSR-DBD fusion disclosed herein, the complex being capable of recognizing and binding to a complementary sequence of the target sequence, the guide RNA being a double-stranded molecule comprising a spacer and a scaffold, the spacer comprising a sequence capable of hybridizing to a complementary sequence of the target DNA sequence, and the one or more target sequences may be contained in the intended acceptor target DNA, the intended donor DNA, or both.
[0183] The guide polynucleotide may be a double-stranded molecule (also called a double-stranded guide polynucleotide) comprising a spacer sequence and a scaffold sequence. The spacer comprises a first nucleotide sequence domain capable of hybridizing with a nucleotide sequence in the target DNA (i.e., a nucleotide sequence complementary to the target sequence), and a second nucleotide sequence (also called a "tracr-mate" sequence) that is part of a Cas protein recognition (CPR) domain. The tracr-mate sequence can hybridize with the scaffold along the complementary region to form a Cas protein recognition domain or CPR domain. The CPR domain can interact with the Cas protein. The spacer and scaffold of the double-stranded guide polynucleotide may be RNA sequences, DNA sequences, and / or mixed RNA-DNA sequences. In some embodiments, the spacer molecule of the double-stranded guide polynucleotide is referred to as "spacer DNA" or "crDNA" (when composed of contiguous DNA nucleotide fragments) or "spacer RNA" or "crRNA" (when composed of contiguous RNA nucleotide fragments) or "spacer DNA-RNA" or "crDNA-RNA" (when composed of a mixture of DNA and RNA nucleotides). The size of naturally occurring spacer fragments in bacteria and archaea that can be included in the spacers disclosed herein can range from, but is not limited to, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 nucleotides or more. In some embodiments, the scaffold is referred to as "scaffold RNA" or "tracrRNA" (when composed of contiguous RNA nucleotide fragments) or "scaffold DNA" or "tracrDNA" (when composed of contiguous DNA nucleotide fragments) or "scaffold DNA-RNA" or "tracrDNA-RNA" (when composed of a mixture of DNA and RNA nucleotides). In one embodiment, the RNA that guides the RNA / Cas9 RNP complex of the LSR-DBD fusion is a double-stranded RNA that includes a double-stranded spacer scaffold.The scaffold or tracrRNA contains, from 5' to 3', (i) a sequence that anneals to the repeat region of the CRISPR type II crRNA, and (ii) a portion containing a stem-loop structure (Deltcheva et al., Nature 471:602-607). A double-stranded guide polynucleotide can form a complex with the Cas protein portion of the LSR-DBD fusion, and the guide polynucleotide / Cas RNP complex (also referred to as the guide polynucleotide / Cas RNP system) can guide the DBD of the LSR-DBD fusion protein described herein to a target site, allowing the DBD protein to recognize and bind to the target site. See also U.S. Patent Application Publication Nos. 2015-0082478 (A1), published March 19, 2015, and 2015-0059010 (A1), published February 26, 2015, each of which is incorporated herein by reference in its entirety. In some embodiments, the spacer sequence is fused to the 5' end of the scaffold sequence, or the spacer sequence is fused to the 3' end of the scaffold sequence.
[0184] The guide polynucleotide can also be a single molecule (also referred to as a single guide polynucleotide) comprising a spacer sequence linked to a scaffold sequence. A single guide polynucleotide comprises a first nucleotide sequence domain capable of hybridizing to a nucleotide sequence in a target DNA (i.e., a nucleotide sequence complementary to the target sequence) and a Cas protein recognition domain (CPR domain) that interacts with a Cas protein. As used herein, "domain" refers to a contiguous nucleotide fragment that can be an RNA sequence, a DNA sequence, and / or a mixed RNA-DNA sequence. The spacer domain and / or CPR domain of a single guide polynucleotide may comprise an RNA sequence, a DNA sequence, or a mixed RNA-DNA sequence. A single guide polynucleotide comprised of a spacer and a scaffold-derived sequence can be referred to as a "single guide RNA" (when comprised of a contiguous RNA nucleotide fragment), a "single guide DNA" (when comprised of a contiguous DNA nucleotide fragment), or a "single guide RNA-DNA" (when comprised of a mixture of RNA and DNA nucleotides). A single guide polynucleotide can form a complex with the Cas protein portion of an LSR-DBD fusion, and the guide polynucleotide / Cas RNP complex (also referred to as a guide polynucleotide / Cas RNP system) can guide the DBD of the LSR-DBD fusion protein described herein to a target site, allowing the DBD to recognize and bind to the target site. See also U.S. Patent Application Publication Nos. 2015-0082478(A1), published March 19, 2015, and 2015-0059010(A1), published February 26, 2015, each of which is incorporated by reference in its entirety.
[0185] In some embodiments, the gRNA comprises an sgRNA that includes a spacer RNA sequence portion and a tracrRNA portion, wherein the nucleic acid sequence of the spacer RNA sequence portion is identical to a target sequence on the desired DNA target and is therefore complementary to and hybridizes with the complementary sequence of the target sequence on the desired DNA target. The one or more target sequences may be contained in the desired acceptor target DNA, the desired donor DNA, or both.
[0186] In some embodiments, a protospacer adjacent motif (PAM) sequence is present immediately 3' from the target sequence on the DNA target. In the CRISPR-Cas9 system, a PAM is a short DNA sequence (typically consisting of 2-6 base pairs) that follows the DNA region targeted for cleavage by the CRISPR system. In some embodiments, the DBD portion of the LSR-DBD fusion comprises dCas9 derived from Streptococcus pyogenes, which recognizes the PAM sequence 5'-NGG-3' (where "N" can be any nucleotide base). Thus, in some embodiments, the DNA target of interest contains a nucleotide sequence identical to the spacer sequence of the guide polynucleotide, immediately followed by "NGG" in the 3' direction. Various Cas endonucleases isolated from various bacterial species, each recognizing a different PAM, are known to those skilled in the art. In some embodiments, the DBD portion of the LSR-DBD fusion comprises dCas9 derived from Staphylococcus aureus, which recognizes the PAM sequence 5'-NGRRT-3' or 5'-NGRRRN-3' (where "N" can be any nucleotide base). In some embodiments, the DBD portion of the LSR-DBD fusion comprises dCas9 from Neisseria meningitidis, which recognizes the PAM sequence 5'-NNNNGATT-3' (where "N" can be any nucleotide base). In some embodiments, the DBD portion of the LSR-DBD fusion comprises dCas9 from Campylobacter jejuni, which recognizes the PAM sequence 5'-NNNNRYAC-3' (where "N" can be any nucleotide base). In some embodiments, the DBD portion of the LSR-DBD fusion comprises dCas9 from Streptococcus thermophilus, which recognizes the PAM sequence 5'-NNAGAAW-3' (where "N" can be any nucleotide base). Cas9 variants with altered specificity, relaxed PAM requirements, or recognition of novel PAM sequences can also be used as the DBD portion of the LSR-DBD fusion. In some embodiments, the DBD portion of the LSR-DBD fusion comprises dCas9-SpG, which recognizes the PAM sequence 5'-NGN-3' (where "N" can be any nucleotide base).
[0187] In some embodiments, the guide polynucleotide comprises a spacer sequence portion, the nucleic acid sequence of which is identical to the target sequence on the target DNA or donor DNA of interest (except for "U" instead of "T" in the RNA spacer sequence), and the target sequence is located adjacent to, overlapping with, or within the binding site for the LSR (e.g., attA or attD) on the target DNA of interest. In some embodiments, the target sequence on the target DNA or donor DNA of interest is located within 300 nucleotides upstream or downstream of the binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA of interest or donor DNA, and the distance is measured from the center of the dinucleotide core of the binding site to the position between the spacer sequence and the PAM. In some embodiments, the target sequence on the target DNA or donor DNA of interest is located within 200 nucleotides upstream or downstream of the binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA of interest or donor DNA. In some embodiments, the target sequence on the target DNA of interest or donor DNA is located within 100 nucleotides upstream or downstream of the binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA of interest or donor DNA. In some embodiments, the target sequence on the target DNA of interest or donor DNA is located within 80 nucleotides upstream or downstream of the binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA of interest or donor DNA. In some embodiments, the target sequence on the target DNA of interest or donor DNA is located within 50 nucleotides upstream or downstream of the binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA of interest or donor DNA. The target sequence may be on either strand of the target DNA of interest or donor DNA. In some embodiments, the guide polynucleotide is an sgRNA.In general, spacers immediately adjacent to the target integration binding site, e.g., attH, have the highest incorporation rates, spacers at more distant locations have reduced incorporation, and spacers overlapping the dinucleotide core of the binding site have significantly reduced incorporation or no incorporation at all.
[0188] In certain aspects, described herein are nucleic acids encoding guide polynucleotides for use in the LSR-DBD fusions described herein. The guide polynucleotide may be encoded on the same nucleic acid molecule as the LSR-DBD fusion and / or donor polynucleotide, or may be encoded on a separate nucleic acid molecule. In some embodiments, the guide polynucleotide is a gRNA comprising a spacer sequence portion and a tracr RNA portion. In some embodiments, the guide polynucleotide is an sgRNA comprising a spacer sequence portion and a tracr RNA portion. In some embodiments, the spacer sequence portion is about 20 nucleotides in length. In some embodiments, the spacer sequence portion is 16 nucleotides in length. In some embodiments, the spacer sequence portion is 20 nucleotides in length. In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on a target DNA of interest or donor DNA, where the target sequence is located adjacent to, overlapping with, or within the binding site (e.g., attA or attD) for the LSR of the LSR-DBD fusion on the target DNA of interest or donor DNA. In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on the target DNA or donor DNA of interest, where the target sequence is located within 300 nucleotides of a binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA or donor DNA of interest. In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on the target DNA or donor DNA of interest, where the target sequence is located within 200 nucleotides of a binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA or donor DNA of interest. In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on the target DNA or donor DNA of interest, where the target sequence is located within 100 nucleotides of a binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA or donor DNA of interest.In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on the target DNA of interest or donor DNA, where the target sequence is located within 80 nucleotides of a binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA of interest or donor DNA. In some embodiments, the DNA sequence immediately 3' to the target sequence on the target DNA of interest or donor DNA comprises a PAM sequence. In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on the target DNA of interest or donor DNA, where the target sequence is located within 50 nucleotides of a binding site for the LSR of the LSR-DBD fusion (e.g., attA or attD) on the target DNA of interest or donor DNA. In some embodiments, the DNA sequence immediately 3' to the target sequence on the target DNA of interest or donor DNA comprises a PAM sequence. In some embodiments, the DNA sequence immediately 3' to the target sequence on the target DNA of interest or donor DNA comprises the PAM sequence NGG. In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on a target DNA of interest (e.g., located adjacent to, overlapping, or within an attA site). In some embodiments, the spacer sequence portion comprises a nucleotide sequence identical to a target sequence on a donor DNA of interest (e.g., located adjacent to, overlapping, or within an attD site). In some embodiments, the spacer sequence portion of a gRNA or sgRNA comprises a nucleotide sequence selected from Figure 37 (SEQ ID NOs: 98-152, 551-561). In some embodiments, the spacer sequence portion of a gRNA or sgRNA comprises a nucleotide sequence selected from Figure 37 (SEQ ID NOs: 98-152, 551-561), with an additional "G" nucleotide at the 5' end. In some embodiments, the spacer sequence portion of a gRNA or sgRNA consists of a nucleotide sequence selected from Figure 37 (SEQ ID NOs: 98-152, 551-561).In some embodiments, the spacer sequence portion of the gRNA or sgRNA consists of a nucleotide sequence selected from Figure 37 (SEQ ID NOS: 98-152, 551-561) with an additional "G" nucleotide at the 5'-end. In some embodiments, the tracr RNA portion of the gRNA or sgRNA comprises SEQ ID NO: 153. In some embodiments, the tracr RNA portion of the gRNA or sgRNA consists of SEQ ID NO: 153. In some embodiments, the spacer sequence portion of the gRNA or sgRNA comprises a nucleotide sequence selected from Figure 37 (SEQ ID NOS: 98-152, 551-561) with an additional "G" nucleotide at the 5'-end, and the tracr RNA portion of the gRNA or sgRNA comprises SEQ ID NO: 153. In some embodiments, the spacer sequence portion of the gRNA or sgRNA comprises a nucleotide sequence selected from Figure 37 (SEQ ID NOS: 98-152, 551-561) with an additional "G" nucleotide at the 5'-end, and the tracr RNA portion of the gRNA or sgRNA comprises SEQ ID NO: 153. In some embodiments, the spacer sequence portion of the gRNA or sgRNA consists of a nucleotide sequence selected from Figure 37 (SEQ ID NOS: 98-152, 551-561), and the tracr RNA portion of the gRNA or sgRNA consists of SEQ ID NO: 153. In some embodiments, the spacer sequence portion of the gRNA or sgRNA consists of a nucleotide sequence selected from Figure 37 (SEQ ID NOS: 98-152, 551-561), with an additional "G" nucleotide at the 5' end, and the tracr RNA portion of the gRNA or sgRNA consists of SEQ ID NO: 153. In some embodiments, the gRNA or sgRNA comprises SEQ ID NOS: 98-152, 551-561, immediately followed by SEQ ID NO: 153. In some embodiments, the gRNA or sgRNA comprises SEQ ID NOS: 98-152, 551-561, with an additional "G" nucleotide at the 5' end, immediately followed by SEQ ID NO: 153. In some embodiments, the gRNA or sgRNA consists of SEQ ID NOs: 98-152, 551-561, immediately followed by SEQ ID NO: 153. In some embodiments, the gRNA or sgRNA consists of SEQ ID NOs: 98-152, 551-561, with an additional "G" nucleotide at the 5' end, immediately followed by SEQ ID NO: 153.
[0189] Donor DNA
[0190] Certain aspects of the present application relate to nucleic acids for use in site-specific insertion of an exogenous nucleic acid, e.g., a gene of interest (GOI), into a target DNA, e.g., a genome. In some embodiments, the exogenous nucleic acid (e.g., GOI) used for insertion can be up to about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 140, 150, 160, 170, 180, 190, 200, or 250 kilobases in length, or longer. The GOI can include non-coding sequences, including cis-regulatory regions and introns.
[0191] The donor DNA may be from 15 bases (b) or base pairs (bp) to about 250 kilobases (kb) or kilobase pairs (kbp) (e.g., from about 50, 75, or 100 b or bp to about 110, 120, 125, 150, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 650, 700, 7 ... 50, 800, 850, 900, 950, 1000, 1250, 1500, 1750, 2000, 2250, 2500, 2750, 3000, 3250, 3500, 3750, 4000, 4250, 4500, 4750, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000 0, 11,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000 (and 1,0 The donor DNA molecules may be up to 250,000 b or bp in length (in increments of 00). Longer donor DNA molecules can be provided in the form of circular or linear plasmids, or as a component of a vector (e.g., as a component of a viral vector), or as an amplification or polymerization product thereof. Shorter donor DNA molecules can be provided as double-stranded oligonucleotides.Exemplary double-stranded template oligonucleotides include those having about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60 , 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 115, 120, 125, 150, 175, 200, 225, or 250b or bp in length or a minimum of about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63 , 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 110, 115, 120, 125, 150, 175, 200, 225, or 250 b or bp in length. For introduction into cells, donor DNA can be provided in the reaction mixture at a concentration of about 1 μM to about 200 μM, about 2 μM to about 190 μM, about 2 μM to about 180 μM, about 5 μM to about 180 μM, about 9 μM to about 180 μM, about 10 μM to about 150 μM, about 20 μM to about 140 μM, about 30 μM to about 130 μM, about 40 μM to about 120 μM, or about 45 μM or 50 μM to about 90 μM or 100 μM.In some cases, the reaction mixture may contain 1 μM, 2 μM, 3 μM, 4 μM, 5 μM, 6 μM, 7 μM, 8 μM, 9 μM, 10 μM, 11 μM, 12 μM, 13 μM, 14 μM, 15 μM, 16 μM, 17 μM, 18 μM, 19 μM, 20 μM, 25 μM, 30 μM, 35 μM, 40 μM, 45 μM, 50 μM, 55 μM, 60 μM, 70 μM, 80 μM, 90 μM, 100 μM, 110 μM, 115 μM, 120 μM, 130 μM, 140 μM, 150 μM, 160 μM, 170 μM, 180 μM, 190 μM, 200 μM or more for introduction into cells. Donor DNA can be provided at a concentration of about or greater than about 1 μM, 2 μM, 3 μM, 4 μM, 5 μM, 6 μM, 7 μM, 8 μM, 9 μM, 10 μM, 11 μM, 12 μM, 13 μM, 14 μM, 15 μM, 16 μM, 17 μM, 18 μM, 19 μM, 20 μM, 25 μM, 30 μM, 35 μM, 40 μM, 45 μM, 50 μM, 55 μM, 60 μM, 70 μM, 80 μM, 90 μM, 100 μM, 110 μM, 115 μM, 120 μM, 130 μM, 140 μM, 150 μM, 160 μM, 170 μM, 180 μM, 190 μM, 200 μM, or more.
[0192] In some embodiments, the donor DNA comprises a target sequence having a nucleotide sequence identical to the spacer sequence portion of a guide polynucleotide (e.g., gRNA, sgRNA). In some embodiments, the donor DNA comprises a target sequence that is identical to the target sequence of a target DNA of interest, such that the same guide polynucleotide sequence can be used to target the LSR-DBD fusion to the donor DNA and target DNA of interest.
[0193] Donor DNA can contain a wide variety of different sequences. In some cases, the donor DNA encodes to introduce a stop codon or a frameshift compared to the target genomic region before cleavage and recombination. Such donor DNA can be useful for knocking out or inactivating a gene or a portion thereof. In some cases, the donor DNA encodes one or more missense mutations or in-frame insertions or deletions compared to the target genomic region. Such donor DNA can be useful for altering the expression level or activity (e.g., ligand specificity) of a target gene or a portion thereof.
[0194] As another example, the donor DNA can encode a wild-type sequence for restoring the expression level or activity of a target endogenous gene or protein. For example, T cells containing mutations in the FoxP3 gene or its promoter region can be restored to treat X-linked IPEX syndrome or systemic lupus erythematosus. Alternatively, the donor DNA can encode a sequence that results in reduced expression or activity of a target gene. For example, an increased immunotherapeutic response can be achieved by deleting or reducing FoxP3 expression or activity in T cells prepared for immunotherapy against cancer or infectious disease targets.
[0195] As another example, the donor DNA can encode a mutation that alters the function of the target gene. For example, the donor DNA can encode a mutation in a cell surface protein required for viral recognition or entry. The mutation may reduce the virus's ability to recognize or infect target cells. For example, a mutation in CCR5 or CXCR4 may increase the resistance of CD4+ T cells to HIV infection.
[0196] In some cases, the donor DNA encodes a sequence adjacent to, but unrelated to, the endogenous sequence. For example, the donor DNA can encode an inducible promoter or repressor element that is unrelated to the endogenous promoter of the target gene. The inducible promoter or repressor element can be inserted into the promoter region of the target gene to provide temporal and / or spatial control of the expression or activity of the target gene.
[0197] In some examples, the donor DNA sequence includes an attD binding site, such as an attB site or an attP site, for the LSR, a constitutive promoter operably linked to a nucleotide sequence encoding a detectable marker, followed by a nucleotide sequence encoding a first selectable marker.
[0198] target DNA
[0199] The target DNA may be any type of DNA molecule, in vitro or in vivo, including, but not limited to, genomic DNA, mitochondrial DNA, eukaryotic DNA, prokaryotic DNA, cDNA, and synthetic DNA. The key requirement for the target DNA is that it contains an LSR binding site, including, but not limited to, an attB site, an attP site, an attH site, or a pseudo site.
[0200] The target DNA (target genome) may contain multiple LSR binding sites. By using an LSR-DBD fusion, the DNA binding domain of the fusion can direct the LSR domain to a single binding site, thereby substantially reducing off-target recombination.
[0201] In some instances, the target DNA sequence comprises an attA binding site, such as an attB or attP site for LSR, a constitutive promoter operably linked to a nucleotide sequence encoding a detectable marker, followed by a nucleotide sequence encoding a first selectable marker. In one type of landing pad, the binding site is between the promoter and the nucleotide sequence encoding the detectable protein. When two or more landing pads are used in a given cell, it is preferred that the binding site of one landing pad be orthogonal to the binding site of the same large serine recombinase in any other landing pad. The landing pad is used for further genetic modification and integration of the nucleic acid molecule of interest by site-specific recombination.
[0202] Nucleic Acid Editing System
[0203] In one aspect, described herein is a nucleic acid editing system comprising a first nucleic acid encoding an LSR-DBD described herein and a second nucleic acid encoding a gRNA. In some embodiments, the gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracrRNA portion, the nucleic acid sequence of the spacer sequence portion being identical to the target nucleic acid sequence except that T in the target nucleic acid sequence is replaced with U in the spacer sequence portion, and the target nucleic acid sequence is located within 80 nucleotides upstream or downstream of a dinucleotide core at the binding site for the LSR portion of the fusion polypeptide on the target DNA. In some embodiments, the spacer sequence portion is 16-20 nucleotides in length. In some embodiments, the gRNA encoded by the nucleic acid is an sgRNA. In some embodiments, a PAM sequence is present immediately 3' to the target nucleic acid sequence on the target DNA. In some embodiments, the first and second nucleic acids are present on the same molecule, for example, but not limited to, the same plasmid or vector. In some embodiments, the first and second nucleic acids are present on different molecules, for example, but not limited to, different plasmids or vectors.
[0204] In some embodiments, the target nucleic acid sequence is located within 80 nucleotides upstream or downstream of the dinucleotide core at the attA site for the LSR portion of the fusion polypeptide on the target DNA of interest. In some embodiments, the attA site is a pseudo site in the mammalian target DNA of interest. In some embodiments, the attA site is a pseudo site (attH) in the human genome. In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), where the attH site is chr10:21130404-21130406:-, chr11:77367459-77367461:-, chr1:230490334-230490336:+, chr2:14280297- In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises Pf80 (SEQ ID NO:2) and dCas9 (SEQ ID NO:29), where the attH site is chr11:64243293-64243295.
[0205] In some embodiments, the tracrRNA portion comprises SEQ ID NO: 153. In some embodiments, the target nucleic acid sequence is located on the DNA of interest within 80 nucleotides upstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide. In some embodiments, the target nucleic acid sequence is located on the DNA of interest within 80 nucleotides downstream of the dinucleotide core at the binding site for the LSR portion of the fusion polypeptide.
[0206] In some embodiments, the nucleic acid editing system further includes a third nucleic acid encoding a second gRNA. In some embodiments, the second gRNA encoded by the nucleic acid includes a spacer sequence portion and a tracrRNA portion, and the nucleic acid sequence of the spacer sequence portion is identical to the target nucleic acid sequence except that T in the target nucleic acid sequence is replaced by U in the spacer sequence portion, and the target nucleic acid sequence is located within 80 nucleotides downstream of the dinucleotide core of the binding site for the LSR portion of the fusion polypeptide on the target DNA. In some embodiments, the spacer sequence portion of the second gRNA is 16 to 20 nucleotides in length. In some embodiments, the second gRNA encoded by the nucleic acid is an sgRNA. In some embodiments, a PAM sequence is present immediately 3' to the target nucleic acid sequence on the target DNA. In some embodiments, the first, second, and third nucleic acids are present on the same molecule, for example, but not limited to, the same plasmid or vector. In some embodiments, the first and second nucleic acids are contained in the same molecule, for example, but not limited to, the same plasmid or vector, and the third nucleic acid is present on a different molecule, for example, but not limited to, a different plasmid or vector. In some embodiments, the second and third nucleic acids are contained in the same molecule, for example, but not limited to, the same plasmid or vector, and the first nucleic acid is present on a different molecule, for example, but not limited to, a different plasmid or vector. In some embodiments, the first, second, and third nucleic acids are present on different molecules, for example, but not limited to, a different plasmid or vector.
[0207] In some embodiments, the nucleic acid editing system further comprises a third nucleic acid comprising a donor DNA sequence comprising an attD binding site for the LSR portion of the fusion polypeptide and a nucleic acid sequence for insertion into a target DNA of interest. In some embodiments, the third nucleic acid further comprises a portion of the gRNA that is identical to the target nucleic acid sequence of the target DNA of interest. In some embodiments, the first, second, and third nucleic acids are present on the same molecule, for example, but not limited to, the same plasmid or vector. In some embodiments, the first and second nucleic acids are present on the same molecule, for example, but not limited to, the same plasmid or vector, and the third nucleic acid is present on a different molecule, for example, but not limited to, a different plasmid or vector. In some embodiments, the second and third nucleic acids are present on the same molecule, for example, but not limited to, the same plasmid or vector, and the first nucleic acid is present on a different molecule, for example, but not limited to, a different plasmid or vector. In some embodiments, the first, second, and third nucleic acids are present on different molecules, for example, but not limited to, a different plasmid or vector.
[0208] In some embodiments, the fusion polypeptide encoded by the nucleic acid comprises (a) Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), wherein the attH sites on the target DNA of interest are located at chromosomal loci chr10:21130404-21130406:-, chr11:77367459-77367461:-, chr1:230490334-230490336:+, chr2:14280297-14280299:+, chr9:116464427-116464429:+, chr20:38982599-38982599 or chr4:92338934-92338936:+, or the attH sequence found at the chromosomal locus, wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO: 154 or a sequence having 90% identity to SEQ ID NO: 154; (b) Pf80 (SEQ ID NO: 2) and dCas9 (SEQ ID NO: 29), wherein the attH site on the target DNA of interest is , chromosomal loci chr11:64243293-64243295:+, chr1:162878224-162878226:+, chr11:92763120-92763122:-, chr9:103309977-103309979:-, chr13:9114 5766-91145768:+, chr2:102467361-102467363:+, chr13:99865454-998 65456:+, chr9:113640780-113640782:-, chr9:123986548-123986550:- , chr15:53565450-53565452:-, or the attH sequence found at the chromosomal locus, wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO:265 or a sequence having 90% identity to SEQ ID NO:265; (c) comprising Cp36 (SEQ ID NO:3) and dCas9 (SEQ ID NO:29), wherein the attH site on the desired target DNA is at chromosomal locus chr16:2789124-2789126:+, chr22:43958465-43958467:-, chr10:117762740-117762742:+;chr7:157294532-157294534:-, chr13:20558930-20558932:-, chr6:151120348-151120350:-, chr10:101429887-101429889:+, chr1:20686551-20686553:+, chr19:50987430-50987432:+, chr4:183226741-183226743:-, or contains the attH sequence found at the chromosomal locus (d) comprising Nm60 (SEQ ID NO: 4) and dCas9 (SEQ ID NO: 29), wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO: 267 or a sequence having 90% identity to SEQ ID NO: 267, and wherein the attH site on the target DNA of interest is located at chromosomal loci chr9:83308042-83308044:-, chr13:79497139-79497141:-, chr9:131409759-131409761:+, chr4:55980785-55980787: +, chr5:96968267-96968269:+, chr6:37700280-37700282:-, chr19:17495840-17495842:-, chr5:126546219-126546221:+, chr10:15703649-15703651:-, chr10:395348-395350:+, or comprises the attH sequence found at the chromosomal locus, and the attD binding site of the donor DNA sequence is SEQ ID NO: 234 or SEQ ID NO: 2 or (e) Si74 (SEQ ID NO: 5) and dCas9 (SEQ ID NO: 29), wherein the attH site on the target DNA of interest is chromosomal locus chr7:155557356-155557358:+, chr9:77155112-77155114:-, or comprises the attH sequence found at the chromosomal locus; and the attD binding site of the donor DNA sequence comprises SEQ ID NO: 266 or a sequence having 90% identity to SEQ ID NO: 266.
[0209] In some embodiments, the third nucleic acid is a plasmid. In some embodiments, the third nucleic acid is a linear amplicon.
[0210] In some embodiments, the ratio of donor DNA to target DNA is controlled within the nucleic acid editing system and by methods described herein using the nucleic acid editing system. In some embodiments, the ratio of donor DNA to target DNA is 5:1. In some embodiments, the ratio of donor DNA to target DNA is 4:1. In some embodiments, the ratio of donor DNA to target DNA is 3:1. In some embodiments, the ratio of donor DNA to target DNA is 2:1. In some embodiments, the ratio of donor DNA to target DNA is 1:1. In some embodiments, the ratio of donor DNA to target DNA is 1:2.
[0211] Vectors and cell lines
[0212] Some aspects of the present invention relate to vector systems comprising one or more vectors, wherein the vectors comprise a nucleic acid sequence encoding an LSR-DBD fusion described herein, a nucleic acid sequence encoding a guide polynucleotide described herein, and / or a donor or target DNA sequence. Vectors can be designed for expression of transcripts (e.g., nucleic acid transcripts, proteins, or enzymes) in prokaryotic or eukaryotic cells. For example, transcripts can be expressed in bacterial cells such as Escherichia coli, insect cells (using baculovirus expression vectors), yeast cells, or mammalian cells. Suitable host cells are further discussed in Goeddel, Gene Expression Technology: Methods In Enzymology 185, Academic Press, San Diego, Calif. (1990), the entire contents of which are incorporated herein by reference. Alternatively, recombinant expression vectors can be transcribed and translated in vitro, for example, using T7 promoter regulatory sequences and T7 polymerase.
[0213] Vectors can be introduced and propagated in prokaryotes. In some embodiments, prokaryotes are used to amplify copies of vectors introduced into eukaryotic cells or as intermediate vectors in the production of vectors introduced into eukaryotic cells (e.g., amplifying plasmids as part of a viral vector packaging system). In some embodiments, prokaryotes are used to amplify copies of vectors and express one or more nucleic acids, such as to provide a source of a nucleic acid construct or one or more proteins used for delivery to a host cell or host organism. Protein expression in prokaryotes is most commonly carried out in E. coli using vectors containing constitutive or inducible promoters that drive the expression of either fusion or non-fusion proteins. Fusion vectors add multiple amino acids to a protein encoded within the fusion vector, such as to the amino terminus of the recombinant protein (in this case, an LSR-DBD fusion). Such fusion vectors can serve one or more purposes, including (i) increased expression of the recombinant protein, (ii) increased solubility of the recombinant protein, and (iii) aiding in the purification of the recombinant protein by acting as a ligand in affinity purification. Fusion expression vectors often incorporate a proteolytic cleavage site at the junction between the fusion moiety and the recombinant protein, allowing for separation of the fusion moiety from the recombinant protein after purification. Examples of such enzymes and their corresponding recognition sequences include factor Xa, thrombin, and enterokinase. Examples of fusion expression vectors include pGEX (Pharmacia Biotech Inc., Smith and Johnson, 1988. Gene 67:31-40), pMAL (New England Biolabs, Beverly, MA), and pRIT5 (Pharmacia, Piscataway, NJ), which fuse glutathione S-transferase (GST), maltose E-binding protein, or protein A, respectively, to the target recombinant protein; the entire contents of each are incorporated herein by reference.
[0214] Minicircles are small circular plasmids or DNA vectors that are episomal and are generated as circular expression cassettes without a bacterial plasmid backbone. They can be generated from a parent bacterial plasmid containing heterologous nucleic acid and two recombinase target sites by intramolecular (cis) recombination using a site-specific recombinase, such as PhiC31 integrase. Recombination between the two sites generates a minicircle and the remaining miniplasmid. The minicircle can be recovered by separation from the miniplasmid.
[0215] Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amrann et al., (1988) Gene 69:301-315) and pETlid (Studier et al., Gene Expression Technology: Methods In Enzymology 185, Academic Press, San Diego, Calif. (1990) 60-89), each of which is incorporated herein by reference in its entirety.
[0216] In some embodiments, the vector is a yeast expression vector. Examples of vectors used for expression in the yeast Saccharomyces cerivisae include pYepSec1 (Baldari, et al., 1987. EMBO J. 6:229-234), pMFa (Kuijan and Herskowitz, 1982. Cell 30:933-943), pJRY88 (Schultz et al., 1987. Gene 54:113-123), pYES2 (Invitrogen Corporation, San Diego, CA), and picZ (InVitrogen Corp, San Diego, CA), each of which is incorporated herein by reference in its entirety.
[0217] In some embodiments, the vector drives protein expression in insect cells using a baculovirus expression vector. Baculovirus vectors available for protein expression in cultured insect cells (e.g., SF9 cells) include the pAc series (Smith, et al., 1983. Mol. Cell. Biol. 3:2156-2165) and the pVL series (Lucklow and Summers, 1989. Virology 170:31-39), the entire contents of each of which are incorporated herein by reference.
[0218] In some embodiments, the vector can drive expression of one or more sequences in mammalian cells (e.g., but not limited to, human embryonic stem cells, HEK cells, hepatocellular carcinoma cells) using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, 1987. Nature 329:840) and pMT2PC (Kaufman, et al., 1987. EMBO J.6:187-195), each of which is incorporated herein by reference in its entirety. When used in mammalian cells, the expression vector's control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. In some embodiments, the promoter is an Ef1a promoter. In some embodiments, the promoter is a U6 promoter. For other suitable expression systems for both prokaryotic and eukaryotic cells, see, for example, Sambrook, et al., Molecular Cloning: A Laboratory Manual. 2, each of which is incorporated herein by reference in its entirety. nd See Chapters 16 and 17 of "Theory of Chemical Biology," ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
[0219] In some embodiments, the vector is capable of driving expression of one or more sequences in a plant cell using a plant cell expression vector.
[0220] In some embodiments, the recombinant mammalian expression vector is capable of directing expression of the nucleic acid preferentially in a particular cell type (e.g., tissue-specific regulatory elements are used to express the nucleic acid). Tissue-specific regulatory elements are known in the art. Non-limiting examples of suitable tissue-specific promoters include the albumin promoter (liver-specific, Pinkert et al., 1987. Genes Dev. 1:268-277), lymphocyte-specific promoters (Calame and Eaton, 1988. Adv. Immunol. 43:235-275), in particular promoters of T cell receptors (Winoto and Baltimore, 1989. EMBO J. 8:729-733) and immunoglobulins (Baneiji et al., 1983. Cell 33:729-740; Queen and Baltimore, 1983. Cell 33:741-748), neuron-specific promoters (e.g., the neurofilament promoter, Byrne and Ruddle, 1989. Proc. Natl. Acad. Sci. USA 86:5473-5477), pancreatic-specific promoters (Edlund, et al., 1985, Science 230:912-916), and mammary gland-specific promoters (e.g., whey promoters, U.S. Pat. No. 4,873,316 and European Patent Application Publication No. 264,166). Developmentally regulated promoters, such as murine hox promoters (Kessel and Gruss, 1990, Science 249:374-379) and the alpha-fetoprotein promoter (Campes and Tilghman, 1989, Genes Dev. 3:537-546), each of which is incorporated herein by reference in its entirety, are also encompassed.
[0221] In some embodiments, a method for introducing an LSR-DBD fusion-gRNA ribonucleoprotein complex into a cell (e.g., a hematopoietic cell or hematopoietic stem cell, including, e.g., a human-derived hematopoietic cell or hematopoietic stem cell) includes forming a reaction mixture containing the protein or ribonucleoprotein complex and introducing a transient pore into the outer cell membrane of the cell. Such a transient pore can be introduced by a variety of techniques, including, but not limited to, electroporation, cell squeezing, or contact with a nanowire or nanotube. Generally, the transient pore is introduced in the presence of the protein or ribonucleoprotein complex, and the protein or ribonucleoprotein complex is allowed to diffuse into the cell.
[0222] Techniques, compositions, and devices for electroporating cells to introduce proteins or ribonucleoprotein complexes may include those described in WO 2006 / 001614 or Kim, JA et al. Biosens.Bioelectron.23, 1353-1360 (2008), the contents of each of which are incorporated herein by reference in their entirety. Additional or alternative techniques, compositions, and devices for electroporating cells to introduce proteins or ribonucleoprotein complexes may include those described in U.S. Patent Application Publication Nos. 2006 / 0094095; 2005 / 0064596; or 2006 / 0087522, the contents of each of which are incorporated herein by reference in their entirety. Additional or alternative techniques, compositions, and devices for electroporating cells to introduce proteins or ribonucleoprotein complexes are described in Li, L. et al. Cancer Res. Treat. 1, 341-350 (2002), U.S. Patent Nos. 6,773,669; 7,186,559; 7,771,984; 7,991,559; 6,485,961; 7,029,916, and U.S. Patent Application Publication Nos. 2014 / 0017213 and 2012 / 0088842, and Geng, T. et al. J. Control Release 144, 91-100 (2010) and Wang, J., et al. Lab. Chip 10, 2057-2061, each of which is incorporated herein by reference in its entirety. (2010).
[0223] In some cases, the techniques or compositions described in the patents or publications cited herein are modified for use in introducing proteins or ribonucleoproteins. Such modifications may include increasing or decreasing the voltage, pulse length, and / or number of pulses. Such modifications may further include modifying buffers, media, electrolytes, or their components. Electroporation can be performed using equipment known in the art, such as the Bio-Rad Gene Pulser electroporation apparatus, the Invitrogen Neon transfection system, the MaxCyte transfection system, the Lonza Nucleofection apparatus, the NEPA Gene NEPA21 transfection apparatus, a flow-through electroporation system equipped with a pump and constant voltage power supply, or other electroporation apparatus or systems known in the art.
[0224] Methods, compositions, and devices for compressing or deforming cells to introduce proteins or ribonucleoprotein complexes may include those described herein. Additional or alternative methods, compositions, and devices may include those described in Nano Lett. 2012 Dec. 12; 12(12):6322-7, Proc Natl Acad Sci USA. 2013 Feb. 5; 110(6):2082-7, J Vis Exp. 2013 Nov. 7; (81):e50980, and Integr Biol (Camb). 2014 Apr. 6(4):470-5, the entire contents of each of which are incorporated herein by reference. Additional or alternative methods, compositions, and devices may include those described in U.S. Patent Application Publication No. 2014 / 0287509, the entire contents of which are incorporated herein by reference. Generally, a protein or ribonucleoprotein complex is provided in a reaction mixture containing cells, and the reaction mixture is forced through a cell-deforming opening or constriction. In some cases, the constriction is smaller than the diameter of the cell. In some cases, the constriction contains a cell-deforming component, such as a region with a strong electrostatic charge, a hydrophobic region, or a region containing a nanowire or nanotube. Forcing the reaction mixture through the constriction can introduce a transient pore in the cell membrane, allowing the protein or ribonucleoprotein complex to enter the cell through the transient pore. In some cases, introducing the protein or ribonucleoprotein by squeezing or deforming the cell can be effective even when the cell is in a non-dividing state.
[0225] A method for introducing a protein or ribonucleoprotein complex into a cell includes forming a reaction mixture containing the protein or ribonucleoprotein complex and contacting cells with the protein or ribonucleoprotein complex to induce receptor-mediated cellular uptake. Compositions and methods for receptor-mediated cellular uptake are described, for example, in Wu et al., J.Biol.Chem. 262, 4429-4432 (1987); and Wagner et al., Proc. Natl. Acad. Sci. USA 87, 3410-3414 (1990), the entire contents of each of which are incorporated herein by reference. Generally, receptor-mediated cellular uptake is mediated by the interaction between a cell surface receptor and a ligand fused to the protein or ribonucleoprotein complex (e.g., covalently bound or fused to RNA in the ribonucleoprotein complex). The ligand can be any protein, small molecule, polymer, or fragment thereof that binds to or is recognized by a receptor on the cell surface. An exemplary ligand is an antibody or antibody fragment (eg, scFV).
[0226] In some embodiments, the reaction mixture for introducing a protein or ribonucleoprotein complex into a cell can include a nucleic acid for directing binding to a target genomic region.
[0227] In some embodiments, introduction is via a nucleic acid (e.g., a plasmid) transfected into the cell. The transfected nucleic acid (e.g., a plasmid) may include an expression vector for an LSR-DBD fusion, a nucleic acid (e.g., a plasmid) comprising a donor molecule for integration into the genome of the cell, and an expression vector for a guide polynucleotide (e.g., a gRNA or sgRNA).
[0228] The nucleic acid can be introduced using adeno-associated virus (AAV), lentivirus, adenovirus, or other types of viral vectors, or combinations thereof. The nucleic acid can be packaged into virions using appropriate packaging cell lines known in the art. In some embodiments, the LSR-DBD fusion protein and one or more exogenous nucleic acids are delivered to cells using lentiviral particles.
[0229] In some cases, expression of the LSR-DBD fusions and / or guide polynucleotides described herein is under the control of an inducible promoter or repressor element. Insertion of an inducible promoter or repressor element into the promoter region of a nucleic acid sequence encoding an LSR-DBD fusion and / or guide polynucleotide described herein allows for temporal and / or spatial control of expression or activity.
[0230] When a nucleic acid encoding an LSR-DBD fusion is delivered to a cell, the nucleic acid can be transcribed and translated into an LSR-DBD protein. The LSR-DBD protein can form a tetrameric complex within the cell. In some embodiments, a nucleic acid encoding an LSR-DBD fusion can be delivered to a cell along with a nucleic acid encoding an LSR. In some embodiments, the LSR and LSR-DBD form a tetrameric complex that can include one, two, or three LSR-DBD fusion proteins.
[0231] Purpose
[0232] Described herein are several applications of the LSR-DBD fusion system described herein, including, but not limited to, a method for introducing amplicon libraries into genomic landing pads, a method for introducing cargo without a landing pad with sufficient efficiency to simultaneously integrate multiple constructs into the same cell, and a method for directly targeting specific sites in mammalian genomes with significantly higher efficiency compared to PhiC31 (which has a genome-targeted LSR integration efficiency of approximately 1%).
[0233] Site-specific nucleases and site-specific recombinases are powerful tools for targeted genome modification in vitro and in vivo. Nuclease cleavage in living cells has been reported to trigger DNA repair mechanisms, frequently resulting in the modification of the cleaved and repaired genome sequence, for example, via homologous recombination. Therefore, targeted cleavage of specific unique sequences in the genome using the LSR-DBD fusions described herein opens up new avenues for gene targeting and gene modification in living cells, including cells that are difficult to manipulate using conventional gene targeting methods, such as many human somatic stem cells or embryonic stem cells. Site-specific recombinases possess all the functions necessary to efficiently and precisely integrate, delete, invert, or translocate specific DNA fragments without causing exposed DNA double-strand breaks.
[0234] In some cases, genome-targeted integration efficiency using the LSR-DBD fusion proteins described herein may be at least about 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more. In some cases, the efficiency of integration of the donor DNA sequence is at least 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, or at least about 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more.
[0235] In some embodiments, one or more nucleic acids encoding the LSR-DBD fusions and guide polynucleotides described herein are used to generate non-human transgenic animals, transgenic plants, or transgenic organoids. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat, or rabbit. In some embodiments, the organism or subject is a plant. In some embodiments, the organism or subject or plant is algae or crops. In some embodiments, the subject is an organoid. Methods for generating transgenic plants, organoids, and animals are known in the art and generally begin with cell transfection methods as described herein. Transgenic animals are also provided, as are transgenic plants, particularly crops and algae. Transgenic animals or plants may be useful for purposes other than providing disease models. These may include, for example, food or feed production by expressing higher levels of proteins, carbohydrates, nutrients, or vitamins than normally found in wild-type organisms. In this regard, transgenic plants, especially pulses and tubers, and animals, especially mammals such as livestock (cattle, sheep, goats and pigs), but also poultry and edible insects, are preferred.
[0236] Transgenic algae or other plants, such as rapeseed, may be particularly useful for the production of biofuels, such as vegetable oils or alcohols (especially methanol and ethanol), which can be engineered to express or overexpress high levels of oils or alcohols for use in the biooil or biofuel industries.
[0237] Plant pathogens are often host-specific. For example, Fusarium oxysporum f. sp. Lycopersici, which causes tomato wilt, infects only tomatoes, while F. oxysporum f. dianthii and Puccinia graminis f. sp. Tritici, which cause wheat stem rust, infect only wheat. Plants have pre-existing and induced defense mechanisms to resist most pathogens. Mutation and recombination events across generations in plants create genetic diversity that leads to susceptibility. This tendency is particularly pronounced because pathogens multiply more frequently than plants. Non-host resistance can exist in plants, for example, when the host and pathogen are incompatible. Horizontal resistance (e.g., partial resistance to all pathogen races, typically controlled by many genes) and vertical resistance (e.g., true resistance to some pathogen races but not others, typically controlled by a few genes) may also exist. At the gene-for-gene level, plants and pathogens coevolve, with genetic changes in one balancing the other. Thus, breeders exploit the diversity in nature to combine the most useful genes involved in yield, quality, uniformity, tolerance, and resistance. Sources of resistance genes include native or exotic varieties, traditional varieties, wild plant relatives, and induced mutations (e.g., treating plant material with mutagens). The present invention provides plant breeders with a new tool for inducing mutations. Thus, those skilled in the art can perform genomic analysis of sources of resistance genes and use the present invention to induce the appearance of resistance genes in varieties with desired characteristics or traits with greater precision than traditional mutagens, thereby facilitating and improving plant breeding programs.
[0238] The present invention encompasses the use of nucleic acids, polypeptides, compositions, systems, and methods disclosed herein for establishing and utilizing transgenic cells / animals / organoids.Disclosed herein are non-naturally occurring or modified compositions, or one or more polynucleotides encoding the components of the compositions, or vectors or delivery systems comprising one or more polynucleotides encoding the components of the compositions, for use in modifying target cells in vivo, ex vivo, or in vitro, and can be carried out in a manner that alters cells, so that once modified, daughter cells or cell lines of the modified cells retain their altered phenotype.By applying the LSR-DBD fusion system to desired cell types ex vivo or in vivo, the modified cells and daughter cells can become part of a multicellular organism, such as a plant or an animal.
[0239] The present invention may be a therapeutic treatment method. The therapeutic treatment method may include gene or genome editing or gene therapy. The methods of the present invention can be used to generate plants, animals, or cells that can be used to model and / or study genetic or epigenetic conditions of interest, for example, through targeted mutation models or as disease models. As used herein, "disease" refers to a disease, disorder, or indication in a subject. For example, the methods of the present invention can be used to generate animals or cells containing modifications in one or more nucleic acid sequences associated with a disease, or plants, animals, or cells in which expression of one or more nucleic acid sequences associated with a disease is altered. Such nucleic acid sequences may encode disease-associated protein sequences or may be disease-associated regulatory sequences. Accordingly, in embodiments of the present invention, it is understood that the plant, subject, patient, organism, or cell can be a non-human subject, patient, organism, or cell. Accordingly, the present invention provides plants, animals, or cells generated by the methods of the present invention, or their progeny. The progeny may be clones of the generated plant or animal, or may be obtained by sexual reproduction by mating with other individuals of the same species to further introduce desirable traits into the progeny. Cells can be in vivo or ex vivo in the case of multicellular organisms, particularly animals or plants.When culturing cells, if suitable culture conditions are met, preferably if the cells are suitable for this purpose (e.g., stem cells), cell lines can be established.The bacterial cell lines produced by the present invention are also contemplated.Therefore, cell lines are also contemplated.
[0240] To further enhance efficacy, the gene therapy vehicle can include one or more immunosuppressants. As used herein, the term "immunosuppressant" includes any compound that suppresses an immune response. Particularly preferred immunosuppressants are cyclosporine, cyclophosphamide, anti-lymphocyte antibodies (e.g., anti-CD20), or anti-cytokine antibodies (e.g., anti-TNF alpha).
[0241] In a preferred embodiment, the gene therapy vehicle according to the present invention may be used in combination with another therapeutic reagent. An effective amount of a pharmaceutical composition according to the present invention is administered, optionally in combination with another therapeutic treatment or agent, such as an immunosuppressant.
[0242] In a further embodiment, the present invention provides a method for ex vivo transfection of relevant host cells (e.g., stem cells) with the LSR-DBD system described herein. In one embodiment, suitable cells are isolated from a mammal, then differentiated in vitro and incubated with an effective amount of a pharmaceutical composition of the present invention. The treated (transfected) cells are then reintroduced into the organism.
[0243] The gene therapy compositions of the present invention contain a pharmaceutically acceptable carrier, a pharmaceutically acceptable vehicle, and / or a pharmaceutically acceptable diluent, in addition to an appropriate salt (alkali metal as counterion and dication in the formulation) and, if necessary, other therapeutic or immunosuppressive agents. Suitable formulations and routes of administration for the preparation of gene therapy vehicles according to the present invention are described in Remington: "The Science and Practice of Pharmacy" 2004, the entire contents of which are incorporated herein by reference. th Edn., ARGennaro, Editor, Mack Publishing Co., Easton, Pa. (2003). Carrier substances that can be used for parenteral administration are, for example, sterile water, Ringer's solution, lactated Ringer's solution, sterile sodium chloride solution, polyalkylene glycols, hydrogenated naphthalenes, particularly biocompatible lactide polymers, lactide-glycolide copolymers, or polyoxyethylene-polyoxypropylene copolymers. Specific embodiments of gene therapy formulations are selected for their physical properties, such as solubility, stability, bioavailability, or degradability.
[0244] The controlled or sustained release of active drug (drug-like) components according to the present invention includes formulations based on lipid-soluble depots (e.g., fatty acids, waxes, or oils). The present invention also discloses coatings of vaccine substances according to the present invention, i.e., coatings with polymers (e.g., poloxamers or poloxamines). Gene therapy substances or compositions according to the present invention can further comprise protective coatings, such as protease inhibitors or permeability enhancers. Preferred carriers are typically aqueous carrier materials, such as water for injection (WFI), water buffered with phosphate, citrate, HEPES, or acetate, or Ringer's solution or lactated Ringer's solution, with a pH typically adjusted to 5.0-8.0, preferably 6.5-7.5. The carrier or vehicle further preferably contains a salt component, such as sodium chloride or potassium chloride, or other components that render the solution isotonic, for example. Furthermore, the carrier or vehicle can contain additional components, such as human serum albumin (HSA), polysorbate 80, sugars, or amino acids, in addition to the above components.
[0245] The mode and method of administration of the gene therapy preparations according to the invention and the dosage will also depend on the nature of the disease to be treated, its stage, if appropriate, and the weight, age, and sex of the patient.
[0246] The gene therapy preparations of the present invention may be administered to patients preferably parenterally, for example, intravenously, intraarterially, subcutaneously, intradermally, intralymphatically, or intramuscularly. The gene therapy preparations may also be administered topically, orally, or intranasally. Furthermore, administration may be by injection into tumor tissue or tumor cavity (e.g., in the case of brain tumors, after surgical removal of the tumor).
[0247] In some methods, disease models can be used to study the effects of mutations on animals or cells and the development and / or progression of disease using measures commonly used in disease research. Alternatively, such disease models are useful for studying the effects of pharmaceutically active compounds on disease.
[0248] In some methods, disease models may be used to evaluate the efficacy of potential gene therapy strategies. That is, disease-associated genes or polynucleotides may be modified to inhibit or reduce the onset and / or progression of the disease. In particular, the methods involve modifying disease-associated genes or polynucleotides such that an altered protein is produced, resulting in an altered response in the animal or cell. Thus, in some methods, genetically modified animals may be compared to animals prone to developing the disease so that the effectiveness of gene therapy events can be evaluated.
[0249] In another embodiment, the invention provides a method for developing biologically active agents that modulate cell signaling events associated with a disease gene, the method comprising contacting a test compound with cells comprising one or more vectors that drive expression of an LSR-DBD fusion system of the invention, and detecting a change in detection indicative of, for example, a decrease or increase in a cell signaling event associated with a mutation in the disease gene contained in the cells.
[0250] In combination with the method of screening for changes in cell function of the present invention, a cellular or animal model can be constructed. Such a model can be used to study the effect of a genomic sequence modified by an LSR-DBD fusion of the present invention on a cell function of interest. For example, a cellular function model can be used to study the effect of a modified genomic sequence on intracellular or extracellular signaling. Alternatively, a cellular function model can be used to study the effect of a modified genomic sequence on sensory perception. In some such models, one or more genomic sequences associated with a signal transduction biochemical pathway within the model are modified.
[0251] Transgenic cells into which one or more nucleic acids encoding one or more components of the present invention are provided or introduced may be operably linked to regulatory elements, including promoters of one or more genes of interest, within the cell. As used herein, the term "LSR-DBD fusion transgenic cell" refers to a cell, such as a eukaryotic cell, into which an LSR-DBD fusion has been genomically integrated. The nature, type, or origin of the cell is not particularly limited according to the present invention. Furthermore, the method for introducing an LSR-DBD fusion transgene into a cell may vary and may be any method known in the art. In certain embodiments, an LSR-DBD fusion transgenic cell is obtained by introducing an LSR-DBD fusion transgene into an isolated cell. In certain other embodiments, an LSR-DBD fusion transgenic cell is obtained by isolating cells from an LSR-DBD fusion transgenic organism. By way of example and not limitation, the LSR-DBD fusion transgenic cells referred to herein may be derived from an LSR-DBD fusion transgenic eukaryotic organism, such as an LSR-DBD fusion knock-in eukaryotic organism. See International Publication No. WO 2014 / 093622 (PCT / US13 / 74667), the entire contents of which are incorporated herein by reference. Methods relating to targeting the Rosa locus described in U.S. Patent Nos. 8,771,985 and 9,567,573, assigned to Sangamo Biosciences, Inc. (each of which is incorporated herein by reference in its entirety), can be modified to utilize the LSR-DBD fusion system of the present invention. Methods relating to targeting the Rosa locus described in U.S. Patent Application Publication No. 20130236946, assigned to Cellectis (the entire contents of which are incorporated herein by reference), can also be modified to utilize the LSR-DBD fusion system of the present invention. The LSR-DBD fusion transgene can further comprise a Lox-Stop-PolyA-Lox (LSL) cassette, which allows expression of the LSR-DBD fusion to be inducible by Cre recombinase.Alternatively, LSR-DBD fusion transgenic cells can be obtained by introducing an LSR-DBD fusion transgene into isolated cells. Transgene delivery systems are well known in the art. For example, the LSR-DBD fusion protein transgene can be delivered into eukaryotic cells, for example, by vector (e.g., AAV, adenovirus, lentivirus) and / or particle and / or nanoparticle delivery, as described elsewhere herein.
[0252] In certain aspects, described herein are cells comprising a nucleic acid encoding any of the LSR-DBD fusions disclosed herein. In some embodiments, the genome of the cell comprises a binding site for the LSR portion of the LSR-DBD fusion. Such cell lines can be used in methods of introducing a nucleic acid comprising a donor binding site and a nucleic acid for insertion into the cell to generate an engineered cell line in which a nucleic acid of interest is inserted into the LSR binding site. In some embodiments, described herein are kits comprising cells, wherein the cells comprise a nucleic acid encoding any of the LSR-DBD fusions disclosed herein. In some embodiments, the genome of the cells of the kit comprises a binding site for the LSR portion of the LSR-DBD fusion. In some embodiments, the kit further comprises a nucleic acid vector (e.g., a plasmid) comprising the donor binding site. In some embodiments, the nucleic acid vector (e.g., a plasmid) of the kit further comprises a multiple cloning site for insertion of a nucleic acid of interest. In some embodiments, the cell is a human cell. In some embodiments, the cell is a human embryonic stem cell. In some embodiments, the cell is an H1 human embryonic stem cell. In some embodiments, the cell is a human cancer cell. In some embodiments, the cell is a human cancer cell line. In some embodiments, the cells are human hepatoma cell lines. In some embodiments, the cells are hepatocellular carcinoma cell lines. In some embodiments, the cells are HepG2 hepatocellular carcinoma cell lines. In some embodiments, the cells are HEK cells.
[0253] Some additional aspects of the present invention relate to modeling abnormalities associated with various genetic disorders in plants or animals, which are further described in the subsection of the topic Genetic Disorders on the National Institutes of Health website (health.nih.gov / topic / GeneticDisorders). Genetic brain disorders include, but are not limited to, adrenoleukodystrophy, corpus callosum hypoplasia, Aicardi syndrome, Alpers disease, Alzheimer's disease, Barth syndrome, Batten disease, CADASIL, cerebellar degeneration, Fabry disease, Gerstmann-Straussler-Scheinker disease, Huntington's disease and other triplet repeat diseases, Leigh disease, Lesch-Nyhan syndrome, Menkes disease, mitochondrial myopathies, and NINDS-defined posterior horn dilation of the lateral ventricles (colpocephaly). These disorders are further described in the subsection of the Genetic Brain Disorders on the National Institutes of Health website.
[0254] In some embodiments, the condition may be neoplasia. In some embodiments, the condition may be age-related macular degeneration. In some embodiments, the condition may be schizophrenia. In some embodiments, the condition may be a trinucleotide repeat disease. In some embodiments, the condition may be fragile X syndrome. In some embodiments, the condition may be a secretase-associated disease. In some embodiments, the condition may be a prion-associated disease. In some embodiments, the condition may be ALS. In some embodiments, the condition may be drug addiction. In some embodiments, the condition may be autism. In some embodiments, the condition may be Alzheimer's disease. In some embodiments, the condition may be inflammation. In some embodiments, the condition may be Parkinson's disease.
[0255] Examples of proteins associated with Parkinson's disease include, but are not limited to, alpha-synuclein, DJ-1, LRRK2, PINK1, Parkin, UCHL1, Synphilin-1, and NURR1.
[0256] An example of an addiction-related protein is ABAT.
[0257] Examples of inflammation-related proteins include monocyte chemoattractant protein-1 (MCP1) encoded by the Ccr2 gene, CC chemokine receptor type 5 (CCR5) encoded by the Ccr5 gene, IgG receptor IIB (FCGR2b, also known as CD32) encoded by the Fcgr2b gene, or Fc epsilon R1g (FCER1g) protein encoded by the Fcer1g gene.
[0258] Examples of cardiovascular disease-related proteins include interleukin 1 beta (IL1B), xanthine dehydrogenase (XDH), tumor protein p53 (TP53), prostaglandin I2 (prostacyclin) synthase (PTGIS), myoglobin (MB), interleukin 4 (IL4), angiopoietin 1 (ANGPT1), ATP-binding cassette, subfamily G (WHITE), member 8 (ABCG8), or cathepsin K (CTSK).
[0259] Examples of Alzheimer's disease-related proteins include the very low density lipoprotein receptor protein (VLDLR) encoded by the VLDLR gene, the ubiquitin-like modifier activating enzyme 1 (UBA1) encoded by the UBA1 gene, or the NEDD8 activating enzyme E1 catalytic subunit protein (UBE1C) encoded by the UBA3 gene.
[0260] Examples of autism spectrum disorder-associated proteins include benzodiazapine receptor (peripheral)-associated protein 1 (BZRAP1) encoded by the BZRAP1 gene, AF4 / FMR2 family member 2 protein (AFF2) encoded by the AFF2 gene (also known as MFR2), fragile X mental retardation autosomal homolog 1 protein (FXR1) encoded by the FXR1 gene, or fragile X mental retardation autosomal homolog 2 protein (FXR2) encoded by the FXR2 gene.
[0261] Examples of macular degeneration-associated proteins include the ATP-binding cassette subfamily A (ABC1) member 4 protein (ABCA4) encoded by the ABCR gene, the apolipoprotein E protein (APOE) encoded by the APOE gene, or the chemokine (CC motif) ligand 2 protein (CCL2) encoded by the CCL2 gene.
[0262] Examples of schizophrenia-related proteins include NRG1, ErbB4, CPLX1, TPH1, TPH2, NRXN1, GSK3A, BDNF, DISC1, GSK3B, and combinations thereof.
[0263] Examples of proteins involved in tumor suppression include ataxia telangiectasia mutated (ATM), ataxia telangiectasia and Rad3 related (ATR), epidermal growth factor receptor (EGFR), v-erb-b2 erythroblastic leukemia viral oncogene homolog 2 (ERBB2), v-erb-b2 erythroblastic leukemia viral oncogene homolog 3 (ERBB3), v-erb-b2 erythroblastic leukemia viral oncogene homolog 4 (ERBB4), Notch1, Notch2, Notch3, or Notch4.
[0264] Examples of secretase disorder-associated proteins include presenilin enhancer 2 homolog (PSENEN) (Caenorhabditis elegans (C. elegans)), cathepsin B (CTSB), presenilin 1 (PSEN1), amyloid beta (A4) precursor protein (APP), anterior pharyngeal deficiency 1 homolog B (aAPH1B) (Caenorhabditis elegans), presenilin 2 (PSEN2) (Alzheimer's disease 4), or beta-site APP-cleaving enzyme 1 (BACE1).
[0265] Examples of amyotrophic lateral sclerosis-related proteins include superoxide dismutase 1 (SOD1), amyotrophic lateral sclerosis type 2 (ALS2), fusion in sarcoma (FUS), TAR DNA-binding protein (TARDBP), vascular endothelial growth factor A (VAGFA), vascular endothelial growth factor B (VAGFB), and vascular endothelial growth factor C (VAGFC), and any combination thereof.
[0266] Examples of prion disease-associated proteins include superoxide dismutase 1 (SOD1), amyotrophic lateral sclerosis type 2 (ALS2), fusion in sarcoma (FUS), TAR DNA-binding protein (TARDBP), vascular endothelial growth factor A (VAGFA), vascular endothelial growth factor B (VAGFB), and vascular endothelial growth factor C (VAGFC), and any combination thereof.
[0267] Examples of proteins associated with neurodegenerative conditions in prion diseases include alpha-2-macroglobulin (A2M), apoptosis antagonistic transcription factor (AATF), prostatic acid phosphatase (ACPP), aortic smooth muscle alpha-2 actin (ACTA2), ADAM metallopeptidase domain (ADAM22), adenosine A3 receptor (ADORA3), or alpha-1D adrenergic receptor (ADRA1D), or alpha-1D adrenergic receptor (ADRA1D).
[0268] Examples of immunodeficiency-associated proteins include, for example, alpha-2-macroglobulin (A2M), arylalkylamine N-acetyltransferase (AANAT), ATP-binding cassette subfamily A (ABC1) member 1 (ABCA1), ATP-binding cassette subfamily A (ABC1) member 2 (ABCA2), or ATP-binding cassette subfamily A (ABC1) member 3 (ABCA3).
[0269] Examples of trinucleotide repeat disease-associated proteins include androgen receptor (AR), fragile X mental retardation 1 (FMR1), huntingtin (HTT), or myotonic dystrophy protein kinase (DMPK), frataxin (FXN), and ataxin 2 (ATXN2).
[0270] Examples of proteins associated with impaired neurotransmission include somatostatin (SST), nitric oxide synthase 1 (NOS1) (neuron), alpha-2A adrenergic receptor (ADRA2A), alpha-2C adrenergic receptor (ADRA2C), tachykinin receptor 1 (TACR1), or 5-hydroxytryptamine (serotonin) receptor 2C (HTR2c).
[0271] Examples of neurodevelopment-related sequences include ataxin 2-binding protein 1 (A2BP1), aminoadipate aminotransferase (AADAT), arylalkylamine N-acetyltransferase (AANAT), 4-aminobutyrate aminotransferase (ABAT), ATP-binding cassette subfamily A (ABC1) member 1 (ABCA1), or ATP-binding cassette subfamily A (ABC1) member 13 (ABCA13).
[0272] Further examples of preferred conditions treatable with the present system include Aicardi-Gutierrez syndrome, Alexander disease, Allan-Herndon-Dudley syndrome, POLG-related disorders, alpha-mannosidosis (types II and III), Alström syndrome, Angelman syndrome, ataxia-telangiectasia, neuronal ceroid lipofuscinoses, beta-thalassemia, bilateral optic atrophy and (infantile) optic atrophy type 1, retinoblastoma (bilateral), Canavan disease, cerebro-oculofacial-skeletal syndrome 1 (COFS1), cerebrotendinous xanthomatosis, Cornelia de Lange syndrome, and MAP. T-linked disorders, genetic prion diseases, Dravet syndrome, early-onset familial Alzheimer's disease, Friedreich's ataxia (FRDA), Freyne's syndrome, fucositosis, Fukuyama-type congenital muscular dystrophy, galactosialitis, Gaucher disease, organic acidemia, hemophagocytic lymphohistiocytosis, Hutchinson-Gilford progeria syndrome, mucolipidosis II, infantile free sialic acid storage disease, PLA2G6-associated neurodegeneration, Jerber-Lang-Nielsen syndrome, junctional epidermolysis bullosa, Huntington's disease, Krabbe disease (infancy), mitochondrial DNA-associated Leigh syndrome, and and NARP, Lesch-Nyhan syndrome, LIS1-associated squirrel's brain, Lowe syndrome, maple syrup urine disease, MECP2 duplication syndrome, ATP7A-associated copper transport disorder, LAMA2-associated muscular dystrophy, arylsulfatase A deficiency, mucopolysaccharidosis type I, II, or III, peroxisome biogenesis disorder, Zellweger syndrome spectrum disorder, neurodegeneration with impaired brain iron storage, acid sphingomyelinase deficiency, Niemann-Pick disease type C, glycine encephalopathy, ARX-related disorders, urea cycle disorders, COL1A1 / 2-associated osteogenesis imperfecta, mitochondrial DNA deletion disorders syndrome, PLP1-related disorder, Perry syndrome, Phelan-McDermid syndrome, glycogen storage disease type II (Pompe disease) (infancy), MAPT-related disorder, MECP2-related disorder, proximal chondrodysplasia punctata type 1, Roberts syndrome, Sandhoff disease, Schindler disease type 1, adenosine deaminase deficiency, Smith-Lemli-Opitz syndrome, spinal muscular atrophy, infantile spinocerebellar ataxia, hexosaminidase A deficiency, squamous dysplasia type 1, collagen type VI-related disorder, Usher syndrome type 1, congenital muscular dystrophy, Wolf-Hirschhorn syndrome, lysosomal acid lipase deficiency,and xeroderma pigmentosum. It is clearly contemplated that any polynucleotide sequence of interest can be targeted using this system.
[0273] The nucleic acids, polypeptides, compositions, systems, and methods disclosed herein can be used to introduce nucleic acid sequences encoding chimeric antigen receptors into cells. Chimeric antigen receptor molecules are recombinant molecules characterized by their ability to bind to antigens and transmit activation signals via immunoreceptor activation motifs (ITAMs) located in their cytoplasmic tails. Receptor constructs utilizing antigen-binding portions (e.g., made from single-chain antibodies (scFv)) offer the added advantage of being "universal" in that they bind to native antigens on the surface of target cells in an HLA-independent manner. In some embodiments, the chimeric antigen receptor comprises a) an intracellular signaling domain, b) a transmembrane domain, and c) an extracellular domain comprising an antigen-binding region.
[0274] In certain embodiments, the intracellular receptor signaling domain in the CAR includes a domain from the T cell antigen receptor complex, such as the CD3 zeta chain, an Fcy RIII costimulatory signaling domain, CD28, CD27, DAP10, CD137, OX40, CD2, e.g., alone or in tandem with CD3 zeta. In certain embodiments, the intracellular domain (which may be referred to as the cytoplasmic domain) includes part or all of one or more of the TCR zeta chain, CD28, CD27, OX40 / CD134, 4-1BB / CD137, FcsRIy, ICOS / CD278, IL-2R beta / CD122, IL-2R alpha / CD 132, DAP 10, DAP 12, and CD40. In some embodiments, one skilled in the art will use any portion of the endogenous T cell receptor complex in the intracellular domain. One or more cytoplasmic domains can be used, and in so-called third generation CARs, for example, at least two or three signaling domains are fused to obtain additive or synergistic effects.
[0275] For example, donor DNA can be used to replace one or more complementarity-determining regions or portions thereof of a T cell receptor chain or antibody gene. Such donor DNA can thus alter the antigen specificity of a target cell. For example, the target cell can be altered to recognize a tumor antigen or an infectious disease antigen, thereby eliciting an immune response against the tumor antigen or infectious disease antigen.
[0276] In certain embodiments of the present invention, CAR cells are introduced into an individual in need thereof, such as an individual with cancer or an infectious disease. The cells then enhance the individual's immune system to attack the respective cancer or pathogenic cells. In some cases, the individual is administered one or more doses of antigen-specific CAR T cells. If an individual is administered more than one dose of antigen-specific CAR T cells, the interval between doses must be long enough to allow for proliferation within the individual; in certain embodiments, the interval between doses is 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, or more.
[0277] The source of allogeneic T cells that are modified to contain a chimeric antigen receptor and not have a functional TCR can be of any type, but in certain embodiments, the cells are obtained, for example, from umbilical cord blood, peripheral blood, human embryonic stem cells, or a bank of induced pluripotent stem cells. A suitable dose for therapeutic efficacy is, for example, at least 10 per dose, preferably in a series of administration cycles. 5 cells, or approximately 10 5 ~about 10 10 An exemplary dosing regimen consists of four ascending weekly dosing cycles, with at least about 10 cells on day 0. 5 cells and achieve a target dose of approximately 10% within a few weeks of the initiation of an intrapatient dose escalation scheme. 10 Suitable modes of administration include intravenous, subcutaneous, intracavitary (e.g., using a reservoir access device), intraperitoneal, and direct injection into the tumor mass.
[0278] The compositions of the present invention can be provided in unit dosage forms, such as injections, each containing a predetermined amount of the composition, either alone or in appropriate combination with other active substances. As used herein, the term unit dosage form refers to a physically discrete unit suitable for use as a single dose in human and animal subjects, each containing a predetermined amount of the composition of the present invention calculated to be sufficient to produce the desired effect, either alone or in combination with other active substances, optionally in combination with a pharmaceutically acceptable diluent, carrier, or vehicle. The specifications of the novel unit dosage forms of the present invention depend on the specific pharmacodynamic properties associated with the pharmaceutical composition in a particular subject.
[0279] Therefore, the dose of transgenic T cells should take into account the route of administration and should be such that a sufficient number of transgenic T cells are introduced to achieve the desired therapeutic response. Furthermore, the amount of each active agent included in the compositions described herein (e.g., the amount per contacted cell or the amount per specific body weight) may vary depending on the application. Generally, the concentration of transgenic T cells is desirably at least about 1 x 10 6 ~Approx. 1×10 9 transfected T cells, and even more preferably about 1 x 10 7 ~Approx. 5×10 8 The concentration should be sufficient to administer 10 ... 8 cells) or less (e.g., 1 x 10 7 Any suitable amount of cells may be employed. Administration schedules may be based on well-established cell-based therapies (see, e.g., Topalian and Rosenberg, 1987; U.S. Pat. No. 4,690,915, each of which is incorporated herein by reference in its entirety), or alternatively, a continuous infusion strategy may be employed.
[0280] These values provide general guidance for the range of transgenic T cells to be utilized by practitioners when optimizing the methods of the present invention for practicing the present invention. The recitation of such ranges herein in no way precludes the use of greater or lesser amounts of components as may be warranted in a particular application. For example, the actual dose and schedule may vary depending on whether the composition is administered in combination with other pharmaceutical compositions, or on individual differences in pharmacokinetics, disposition, and metabolism. Those skilled in the art can easily make any necessary adjustments as required by a particular situation.
[0281] In some embodiments, the donor DNA encodes a recombinant antigen receptor, a portion thereof, or a component thereof. Recombinant antigen receptors, portions thereof, and components thereof include those described in U.S. Patent Application Publication Nos. 2003 / 0215427; 2004 / 0043401; 2007 / 0166327; 2012 / 0148552; 2014 / 0242701; 2014 / 0274909; 20140314795; 2015 / 0031624, and International Publication Nos. 2000 / 023573 and 2014 / 134165, the entire contents of each of which are incorporated herein by reference. Such recombinant antigen receptors can be used in immunotherapy targeting specific tumor-associated antigens or infectious disease-associated antigens. In some cases, the methods described herein may be used to knock out an endogenous antigen receptor, such as a T cell receptor, a B cell receptor, or a portion or component thereof. Also, the methods described herein may be used to knock in a recombinant antigen receptor, a portion thereof, or a component thereof. In some embodiments, the endogenous receptor is knocked out and replaced with a recombinant receptor (e.g., a recombinant T cell receptor or a recombinant chimeric antigen receptor). In some cases, the recombinant receptor is inserted at the genomic location of the endogenous receptor. In some cases, the recombinant receptor is inserted at a different genomic location than the endogenous receptor.
[0282] As another example, the donor DNA may encode a suicide gene, a reporter gene, or a rheostat gene, or portions thereof. Suicide genes can be used to remove antigen-specific immunotherapy cells from the host after successful treatment. Rheostat genes can be used to regulate the activity of the immune response during immunotherapy. Reporter genes can be used to monitor cell number, location, and activity in vitro or in vivo after introduction into the host. In a preferred embodiment, the donor DNA contains an attD site that allows site-specific integration of the donor DNA into cellular DNA.
[0283] Exemplary rheostat genes are immune checkpoint genes. Increasing or decreasing the expression or activity of one or more immune checkpoint genes can be used to regulate the activity of an immune response during immunotherapy. For example, expression of an immune checkpoint gene can be increased to decrease the immune response. Alternatively, an immune checkpoint gene can be inactivated to increase the immune response. Exemplary immune checkpoint genes include, but are not limited to, CTLA-4 and PD-1. Other rheostat genes can include any gene that regulates target cell proliferation or effector function. Such rheostat genes include transcription factors, chemokine receptors, cytokine receptors, or genes involved in co-inhibitory pathways such as TIGIT or TIM. In some cases, the rheostat gene is a synthetic or recombinant rheostat gene that interacts with cell signaling mechanisms. For example, a synthetic rheostat gene can be a drug- or light-dependent molecule that inhibits or activates cell signaling. Such synthetic genes are described, for example, in Cell 155(6):1422-34 (2013) and Proc Natl Acad Sci USA. 2014 Apr. 22; 111(16):5896-901, the contents of each of which are incorporated herein by reference in their entirety.
[0284] Exemplary suicide genes include, but are not limited to, thymidine kinase, herpes simplex virus type 1 thymidine kinase (HSV-tk), cytochrome P450 isoenzyme 4B1 (cyp4B1), cytosine deaminase, human folylpolyglutamate synthase (fpgs), or inducible casp9. In some embodiments, the suicide gene is selected from the group consisting of genes encoding HSV-1 thymidine kinase (abbreviated as HSV-tk), splice-corrected HSV-tk (abbreviated as cHSV-tk; see Fehse B et al., Gene Ther (2002) 9(23):1633-1638), which encode ganciclovir-hypersensitive HSV-tk mutants (mutants in which the residues at positions 75 and / or 39 are mutated; see Black Me. Et al. Cancer Res (2001) 61(7):3022-3026 and Qasim W et al., Gene Ther (2002) 9(12):824-827), each of which is incorporated herein by reference in its entirety. Suicide genes other than thymidine kinase-based genes can alternatively be utilized.For example, human CD20 (the target of clinical-grade monoclonal antibodies such as Rituximab®, see Serafini M et al., Hum Gene Ther. 2004; 15:63-76), inducible caspases (e.g., modified human caspase 9 fused to human FK506-binding protein (FKBP) to allow conditional dimerization by small molecule drugs, Di Stasi A et al., N Engl J Med. 2011 Nov. 3; 365(18):1673-83, Tey SK et al., Biol Blood Marrow Transplant. 2007 August) '3(8):9) '3-24. Epub 2007 May Genes encoding 5-fluorocytosine or 5-FC (which modifies the non-toxic prodrug 5-fluorocytosine or 5-FC into its highly cytotoxic derivatives 5-fluorouracil or 5-FU and 5'-fluorouridine-5' monophosphate or 5'-FUMP; Breton E et al., CR Biol. 2010 March; 333(3):220-5. Epub 2010 January 25) can be used as suicide genes. The contents of each of these documents are incorporated herein by reference in their entirety.
[0285] array FIG. 33 discloses the amino acid (SEQ ID NOs: 1-5) and corresponding nucleotide sequences (SEQ ID NOs: 6-10) of exemplary LSRs (Dn29, Pf80, Cp36, Nm60, Si74) for use in the LSR-DBD fusions described herein.
[0286] Other LSRs for use in LSR-DBD fusions include the list of experimentally characterized large serine recombinases provided in Supplementary Table 2 of Durrant, M.G., Fanton, A., Tycko, J. et al. Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome, Nat Biotechnol 41, 488-499 (2023), the entire contents of which are incorporated herein by reference. Any of these LSRs (Bc30, Bm99, Bs46, Bt24, Bu30, Bxb1, Cb16, Cc91, Cd04, Cd15, Cd16, Cd31, Cp36, Cs56, Ct03, Dn29, Ec03, Ec04, Ec05, Ec06, Ec07, Ef01, Ef02, Efs2, Em12, Enc3, Enc9, Fm04, Fp10, Kp01, Kp03, Kp04, Kp05, Ma05, Ma37, Me99, Nm60, No The following LSRs (SEQ ID NOs: 67, Pa01, Pa03, Pc01, Pc64, Pf13, Pf15, Pf48, Pf80, Ph43, PhiC31, Pp20, Ps40, Ps45, Rb27, Rh64, Rl09, Sa01, Sa02, Sal0, Sa34, Sa51, Se37, Sh25, Si74, Sm18, Sp56, Td01, Td08, uCb4, Vh19, Vh73, and Vp82) can be used in the LSR-DBD fusions described herein. The amino acid sequences of these LSRs are provided as SEQ ID NOs: 432 to 501, respectively, in the Sequence Listing attached to this application. The attP binding sites corresponding to these LSRs are provided as SEQ ID NOs: 292 to 361, respectively, in the Sequence Listing attached to this application. The attB binding sites corresponding to these LSRs are provided as SEQ ID NOS: 362 to 431, respectively, in the sequence listing attached to this application.
[0287] Figure 39 discloses the amino acids (SEQ ID NOS: 276, 279, 282, 285, 288, and 291) of exemplary LSRs (Cd08, CMp1, E101, Pa19, Pg17, Sal1) for use in the LSR-DBD fusions described herein, and the corresponding attP binding sites (SEQ ID NOS: 274, 277, 280, 283, 286, and 289), and the corresponding attB binding sites (SEQ ID NOS: 275, 278, 281, 284, 287, 290), respectively. Figure 40 discloses the nucleic acid sequences (SEQ ID NOS: 515-533) of exemplary LSRs (Bm99, Bt24, Bxb1, Cb16, Cs56, Ec03, Enc3, Fm04, Kp03, Me99, No67, Pa03, PhiC31, Ps45, Sp56, uCb4, Vh19, Vh73, or Vp82) for use in the LSR-DBD fusions described herein.
[0288] FIG. 34 discloses the amino acid sequences (SEQ ID NOS: 11-19) and corresponding nucleotide sequences (SEQ ID NOS: 20-28) of exemplary linkers for use in the LSR-DBD fusions described herein.
[0289] Figure 35 discloses the amino acid (SEQ ID NOs: 29-32) and corresponding nucleotide sequences (SEQ ID NOs: 33-36) of exemplary DBDs (dCas9, dCas9-HF1, dCas9-SpG, dCas9-Spg-HF1) for use in the LSR-DBD fusions described herein.
[0290] FIG. 36 discloses the amino acid sequences of exemplary LSR-DBD fusions described herein (SEQ ID NOs: 37-42).
[0291] Figure 37 discloses exemplary gRNA sequences with target sites (shown as chromosomal loci based on the human genome assembly GRCh38, available at www.ncbi.nlm.nih.gov / genome / guide / human / ) where the gRNA spacer is located adjacent to, overlapping, or internal to the target DNA sequence (SEQ ID NOS: 43-97, 540-550). Also disclosed are the corresponding gRNA spacers (SEQ ID NOS: 98-152, 551-561), an exemplary gRNA scaffold (SEQ ID NOS: 153) for use with the LSR-DBD fusions described herein.
[0292] Figure 38 discloses exemplary attD sequences (SEQ ID NOs: 154, 164, 174, 184, 194, 204, 214, 224, 234, 244, 254, 264-267) and corresponding attH pseudosites (shown as chromosomal loci based on the human genome assembly GRCh38, available at www.ncbi.nlm.nih.gov / genome / guide / human / ) for various LSRs described.
[0293] The present invention is further illustrated by the following non-limiting examples. [Example]
[0294] In order to facilitate a more complete understanding of the present invention, the following examples are set forth. The following examples serve to illustrate exemplary modes for making and practicing the invention. However, the scope of the invention should not be construed as being limited to the specific embodiments disclosed in these examples, which are merely illustrative.
[0295] Example 1: Methods
[0296] LSR integration site mapping:
[0297] To determine the location and efficiency of integration into the pseudosites in the human genome, donor plasmids containing various LSR orthologs and their corresponding attD genes were transfected into cells and subjected to next-generation sequencing to map their integration sites. For K562 cells, 1.0 × 10 chromosomes were transfected into 100 μL of Amaxa solution (Lonza Nucleofector SF, program FF-120) containing 3,000 ng of LSR plasmid and 2,000 ng of pseudosite attD plasmid. 6 Cells were electroporated. As a non-matching LSR control, 3000 ng of Bxb1 was used in place of the correct LSR plasmid. 2 × 10 5 cells / mL~1×10 6 Cells were cultured at 1000 cells / mL for 2–3 weeks. Genomic DNA was extracted using the Quick-DNA Miniprep Kit (Zymo) and quantified using the Qubit HS dsDNA Assay (Thermo). Tn5 tagmentation, nested PCR enrichment of integration sites, NGS sequencing, and computational analysis of integration sites were performed as described in Durrant et al., NBT 2022.
[0298] Construct design and cloning:
[0299] A fusion protein in which catalytically inactive Cas9 was fused to LSR-P2A-GFP was constructed by Gibson cloning of each portion into a pUC19-derived plasmid containing the Ef1a promoter and an SV40 polyA tail. Various linkers, including (GGS)8 (SEQ ID NO:11), (GGGGS)6 (SEQ ID NO:598), XTEN16, XTEN32-(GGSS)2 (SEQ ID NO:14), and XTEN48-(GGSS)2 (SEQ ID NO:15), were tested to link dCas9 to the LSR as N- and C-terminal fusions. Spacers targeting loci adjacent to the experimentally determined LSR integration site with an NGG PAM and a non-targeting control were cloned into a U6 promoter-driven sgRNA expression plasmid. The donor plasmid contained the attD sequence corresponding to the LSR, the Ef1a promoter, mCherry, and a puromycin resistance gene.
[0300] Transfection and genomic DNA extraction:
[0301] 20,000 HEK293FT cells were plated in each well of a 96-well plate. One day later, they were transfected with 375 ng of effector plasmid, 100 ng of sgRNA plasmid, and 250 ng of donor plasmid per well using Lipofectamine 2000. In some transfections, 389 ng of donor plasmid, 259 ng of effector plasmid, and 76 ng of sgRNA plasmid were introduced using a 5:1:1 ratio of donor, effector, and guide plasmids. Three days after transfection, genomic DNA was harvested. The medium was aspirated from each well, 50 μL of QuickExtract DNA I was added, the cells were mixed by pipetting, transferred to a qPCR plate, vortexed, and thermal cycled at 65°C for 15 minutes, 68°C for 15 minutes, and 98°C for 10 minutes. Genomic DNA was purified using AmpureXP with 0.9 volumes of beads according to the manufacturer's instructions.
[0302] Quantification of on-target integration by ddPCR:
[0303] PCR primers and a FAM-BHQ1 TaqMan® probe were designed to span the donor-genome junction at attH1. For each target site, a control set of primers and a HEX-BHQ1 probe was designed targeting an adjacent site on the same chromosome. ddPCR droplets were generated, amplified, and measured using a QX200 AutoDG Droplet Digital PCR System (Biorad). Integration efficiency was calculated by calculating the ratio of FAM-positive droplets to HEX-positive droplets.
[0304] Quantification of on-target integration by qPCR:
[0305] PCR primers and a FAM-BHQ1 TaqMan probe were designed to span the donor-genome junction at attH1. For each target site, a control set of primers and a HEX-BHQ1 probe targeting an adjacent site on the same chromosome was designed. To quantify integration efficiency, multiplex qPCR was performed using TaqMan Fast Advanced MasterMix (Thermo). Delta Ct values were calculated relative to the control primer / probe set.
[0306] Quantification of total incorporation efficiency by flow cytometry:
[0307] 20,000 HEK293FT cells were plated in each well of a 96-well plate. One day later, 375 ng of effector plasmid (LSR or dCas9-LSR fusion), 100 ng of sgRNA plasmid, and 250 ng of donor plasmid were transfected using Lipofectamine 2000. A nonmatching LSR (Bxb1) and empty sgRNA plasmid were transfected with the donor plasmid as a control using only the donor plasmid. Cells were cultured for 17 days to dilute out unintegrated donor plasmid. Flow cytometry measurements were performed at various time points using an Attune NxT Flow Cytometer (Thermo).
[0308] Example 2: Design and optimization of Dn29-dCas9 fusion constructs
[0309] The LSR binds to attP and attB in the tetrameric complex. Figure 4.
[0310] The N-terminal side of the LSR is important for tetrameric complex formation, subunit rotation, cleavage, and ligation (Figure 5).
[0311] The design of an exemplary Dn29-dCas9 fusion construct is shown in Figure 6. N-terminal and C-terminal fusions were tested, with a preference for long, flexible linkers. Examples of linkers tested include (GGS)8 (SEQ ID NO: 11), (GGGGS)6 (SEQ ID NO: 598), XTEN16, XTEN32-(GGSS)2 ("(GGSS)2" disclosed as SEQ ID NO: 585), and XTEN48-(GGSS)2 ("(GGSS)2" disclosed as SEQ ID NO: 585). To confirm whether the resulting fusion products were competent for recombination, plasmids expressing each fusion construct were co-transfected into HEK293FT cells with donor plasmids containing attD and a non-targeting guide RNA expression plasmid. After 3 days, integration efficiency at attH1 was measured by qPCR. The results show that the Dn29-linker-dCas9 fusion is recombinationally active at levels comparable to or higher than wild-type Dn29, and the dCas9-linker-Dn29 fusion construct has reduced recombination potential.
[0312] In each condition, 725 ng of DNA was transfected. The size of the wild-type Dn29 effector plasmid was 6 kb, while the fusion effector plasmid was 10 kb. Since the same mass of effector plasmid was used in both conditions, the cells transfected with the fusion effector had a lower molar concentration of effector plasmid and a higher molar ratio of donor plasmid to effector plasmid. This factor may explain why the fusion construct had a higher integration efficiency than the wild-type construct, even when transfected with a non-targeting gRNA. The dCas9-linker-Dn29 fusion may have reduced recombination due to steric hindrance caused by the bulky dCas9 domain, which interferes with tetrameric complex formation or subunit rotation.
[0313] Although all Dn29-linker-dCas9 conditions yielded recombination-competent proteins, higher integration rates were achieved with longer, more flexible linkers, such as the XTEN32-(GGSS)2 linker ("(GGSS)2," disclosed as SEQ ID NO: 585) (Figure 6).
[0314] The construction of the construct is versatile for other LSRs (Cp36) (Figure 7).
[0315] Example 3: Proof of concept of pseudo-site targeting with a single guide RNA
[0316] A single guide RNA complementary to DNA adjacent to the pseudosite can guide the LSR-dCas9 monomer to the pseudosite, increasing integration efficiency at this site (Figure 8). Proof-of-concept for this system was demonstrated using fusions of Dn29 and dCas9 with various guide RNAs targeting attH1 and attH3. attH1 is the most efficient pseudosite for Dn29 and is located within the NEBL (cardiac neblet) intron, at positions 21, 130, and 404 on chromosome 10. attH3 is the third most efficient pseudosite. It is located between genes on chromosome 1. The nearest genes are LOC105373164 (non-coding RNA) and PGDB5 (piggyBac transposable element-derived 5).
[0317] Figure 9 shows Dn29-dCas9 targeting attH1. Six gRNAs were designed to target closely to attH1, as shown in the schematic diagram above. HEK293FT cells were transfected with the Dn29-dCas9 fusion effector plasmid, the donor plasmid containing attD, and the gRNA plasmid. After three days, integration efficiency was detected by qPCR. Two gRNAs (2 and 3) were identified that significantly increased integration efficiency compared to the non-targeting guide. This integration efficiency was verified by orthogonal detection methods, including ddPCR (Figure 10, top) and flow cytometry of stably integrated mCherry expression (Figure 10, bottom).
[0318] This method of targeting Dn29-dCas9 to pseudosites was further validated with another pseudosite, attH3. Eight gRNAs were designed to target attH3, and six of these gRNAs increased integration efficiency by up to sixfold compared to non-targeting gRNAs, as demonstrated by qPCR (Figure 11, top) and ddPCR (Figure 11, bottom). Finally, we demonstrated that this method of targeting LSR-dCas9 pseudosites can be generalized to other LSR orthologs, Pf80 and Nm60. Another human genomic plasmid targeting LSR, Pf80, was fused to dCas9 and delivered into HEK293FT cells along with an attD donor plasmid and various attH1-targeting gRNAs (the spacer locations are shown in the schematic diagram at the bottom of Figure 12). attH1 was located at loci 64, 243, and 293 on chromosome 11, as determined by integration site mapping assays (Figure 12, left). The qPCR results (right) show that various gRNAs can increase the integration efficiency of Pf80 in attH1. Similarly, when using various gRNAs whose spacer positions are shown in the schematic diagram at the bottom of Figure 13, the Nm60-dCas9 fusion shown in Figure 13 increases integration efficiency in attH1 by up to 25%. The dCas9 fusion increased integration efficiency by up to 30% in attH1 and up to 8% in attH3, with the fold change in guides successfully integrating into untargeted guides ranging from 3 to 11 (Figure 14). The difference in absolute integration efficiency between attH1 and attH3 indicates that maximum integration efficiency may be limited by the initial insertion efficiency.
[0319] Example 4: Mechanism of Action
[0320] We attempted to determine the limiting reagent in the reaction. Figure 15 shows a schematic diagram of a non-limiting embodiment of a plasmid that can be used to achieve DNA insertion (top). The bottom diagram shows the integration rate when three plasmids are transfected at different molar ratios.
[0321] The donor plasmid is the limiting reagent, and strategies to increase the molar concentration of the donor plasmid in the nucleus, including the use of minicircles, bDNA nuclear import signals, and donor gRNA targeting, can improve efficiency.
[0322] We investigated whether non-targeted dCas9 monomers sterically hinder recombination. We hypothesized that the three non-targeted dCas9 monomers might sterically hinder the rotation / recombination mechanism. Figure 16 shows a schematic representation of this hypothesis when the three non-targeted LSR proteins in the tetrameric complex are not fused to dCas9. We tested this hypothesis by introducing LSR-dCas9 fusion proteins linked to various self-cleaving 2a peptides. These constructs would generate mixed populations of LSR monomers and LSR-dCas9 fusion proteins ranging from approximately 50% cleavage (F2a) to nearly complete cleavage (P2a). Figure 17 shows that introducing mixed populations of LSR and LSR-dCas9 at the indicated ratios did not improve recombination efficiency; in fact, partial cleavage of the fusion constructs reduced recombination efficiency, indicating that these extra dCas9 domains do not sterically hinder recombination. However, this does not exclude the possibility that a mixed population of unfused LSR and LSR-dCas9 fusions may enhance efficiency in untested proportions. The data also indicate that direct fusion of LSR to dCas9 is required to confirm efficacy, as complete cleavage (P2a) was not significant compared to untargeted guides. This supports the hypothesis that direct LSR binding in proximity to the pseudosite enhances efficiency compared to other effects of dCas9 / DNA binding, such as effects on chromatin, such as loosening of chromatin structure or localized nucleosome respiration.
[0323] We investigated how gRNA distance to the pseudosite affects integration. Figure 18 shows integration efficiency as a function of distance from the dinucleotide core, measured from the center of the dinucleotide core to the position between the spacer and the PAM. In some embodiments, including those in which the functional guide is located adjacent to or just outside the pseudosite sequence, the distance from the core is less than 80 bp. This data indicates that the spacing between the PAM and pseudosite affects the ability to find a functional guide to target a new pseudosite. With this insight in mind, we next tested LSR fusions with a PAM-flexible dCas9 mutant called SpG, which has NGN PAM specificity (Figure 29).
[0324] In summary, we found that the donor plasmid is the limiting reagent for transfection, that direct binding of the LSR to dCas9 is required, that there appears to be no steric hindrance caused by untargeted dCas9 in the tetrameric complex, and that the gRNA is preferably located in close proximity to the pseudosite.
[0325] In some embodiments, PAM-flexible Cas mutants may be used to expand guide RNA target selection.
[0326] Example 5: Design modifications to optimize integration efficiency
[0327] To optimize integration efficiency, various design modifications were tested. In one embodiment, two guide RNAs targeting upstream and downstream of the pseudo-site are introduced to increase dimer formation at the genome binding site. A model of a tetrameric complex, with two dCas9s bound closely to the pseudo-site and two dCas9 monomers unbound, is shown in Figure 19. Using this design, we show that the introduction of two target-binding gRNAs appears to have an additive effect on integration, increasing attH3 integration from approximately 5-8% with a single guide to approximately 10-13% with two guide RNAs (Figure 20). Similarly, we show that multiple guides increase integration efficiency for attH1 (Figure 21).
[0328] Another design modification to improve efficiency is the inclusion of a second gRNA targeting the donor plasmid. This guide can assist in nuclear recruitment of the donor plasmid and / or promote dimerization on the donor plasmid. A model of this tetrameric complex is shown in Figure 22. Full-length (20 bp) and truncated (16 bp) spacers were designed to target upstream and downstream of attD on the donor plasmid. The truncated spacers reduce binding affinity, potentially reducing the donor plasmid's ability to function as a protein "sink" (i.e., the binding affinity of the [donor target] far exceeds that of the [genomic target]).
[0329] Figure 23 shows that donor-targeting guides slightly increase integration efficiency.
[0330] A modification of this approach is shown in Figure 31. The donor plasmid and gRNA are designed so that the target sequence of the gRNA spacer is adjacent to attH1 and attD, resulting in a single guide targeting LSR-dCas9 to bind to both the target and the genome. A possible orientation of this strategy is shown in the schematic (top). The target sequence on the donor plasmid can be either full-length (20 bp) or truncated (16 bp). The bottom diagram shows the increased efficiency that this single-guide dual-targeting design offers. In this design, the full-length target sequence located adjacent to attD on the donor plasmid provides up to a 1.5-fold increase in efficiency compared to a control donor without a target sequence.
[0331] In summary, multiple guides targeting the genomic pseudosite or donor significantly increase integration efficiency. The pseudosite, the best candidate for multiple guides, has functional guides both upstream and downstream. Guides targeting the donor plasmid have a slightly positive effect on integration, and a preferred design includes the genomic target sequence of the gRNA on the donor, so that a single gRNA targets two sites, the genome and the donor. Targeting the donor with a full-length gRNA is preferable to a truncated guide with a mismatch in the last four bases.
[0332] Example 6: Measuring the effect of dCas9 fusions on specificity
[0333] Targeting Dn29 monomers to a single pseudosite improves efficiency and, consequently, specificity (on-target / off-target). To measure integration specificity, we developed an integration site mapping assay. HEK293FT cells are transfected with an effector (LSR or LSR-dCas9 fusion) plasmid, a donor plasmid containing a UMI, and, in the case of dCas9 fusions, a guide RNA plasmid. Genomic integration events are mapped by NGS as described in this method, so that each unique integration event is counted based on its UMI. The proportion of UMIs at each locus measures the relative integration preference of that locus over all other loci and indicates the specificity profile of each effector.
[0334] Figure 24 shows the specificity profiles of Dn29 (left) and Dn29-dCas9 (right). Using the fusion system, the specificity of attH1 increases from 30% to nearly 80%. For attH3, the specificity increases from less than 10% to more than 50% (Figure 25). To examine the relationship between efficiency and specificity, we plotted the ddPCR results of each guide, attH1 and attH3, against the specificity of the on-target site. The results show that guides with higher integration efficiency have fewer off-target integrations (Figure 26), indicating that the increased efficiency of a single pseudo-site is important for increasing specificity.
[0335] Example 7: Dn29-dCas9-mediated integration of the plasmid donor attH1 in H1 human embryonic stem cell and HepG2 hepatocellular carcinoma cell lines
[0336] Figure 41 shows Dn29-dCas9-mediated integration of a plasmid donor at attH1 in H1 human embryonic stem cells. Cells were transfected with a puromycin-expressing donor plasmid and an effector plasmid expressing both the Dn29-dCas9 effector and guide3 using FuGENE Transfection Reagent at the indicated donor-to-effector molar ratios (total mass 140 or 280 ng / well). As controls, WT Dn29 and mismatched LSR were transfected with the same Dn29 donor plasmid as the effector. On day 1, cells were split and half underwent puromycin selection. Three days after transfection, attH1 integration was measured by ddPCR from plates without selection. Three and eight days after transfection, the attH1 integration rate was measured by ddPCR from plates with selection. The results show that selection increases integration. In this example, the LSR-DBD fusion (Dn29-dCas9) and guide RNA were expressed from the same plasmid, with effector expression driven by Ef-1a and guide expression driven by U6.
[0337] Figure 42 shows Dn29-dCas9-mediated integration of a donor plasmid at attH1 in the HepG2 hepatocellular carcinoma cell line. Cells were transfected with a puromycin-expressing donor plasmid and an effector plasmid expressing both the Dn29-dCas9 effector and guide 3 using XtremeGene-9 Transfection Reagent at the specified molar ratios for cells seeded at 8–20 kJ / well, as indicated in the figure legend. After 3 days, attH1 integration was measured by ddPCR. In this example, the LSR-DBD fusion (Dn29-dCas9) and guide RNA were expressed from the same plasmid, with effector expression driven by Ef-1a and guide expression driven by U6.
Claims
1. A nucleic acid comprising a sequence encoding a fusion polypeptide, said fusion polypeptide comprising a large serine recombinase (LSR) portion and a DNA binding domain (DBD) portion.
2. The nucleic acid of claim 1 , wherein the nucleic acid sequence encodes a fusion polypeptide in which the LSR portion is fused to the N-terminal side of the DBD portion.
3. The nucleic acid of claim 1, wherein the nucleic acid sequence encoding the fusion polypeptide further comprises a nucleic acid sequence encoding a peptide linker located between the nucleic acid sequence encoding the LSR portion and the nucleic acid sequence encoding the DBD portion.
4. The nucleic acid of claim 3 , wherein the nucleic acid sequence encodes a fusion polypeptide in which the LSR portion is fused to the N-terminal side of the DBD portion via the peptide linker.
5. The nucleic acid according to claims 3 to 4, wherein the peptide linker encoded by the nucleic acid comprises at least one amino acid.
6. The nucleic acid of claim 5, wherein the peptide linker encoded by the nucleic acid comprises 2 to 100 amino acids.
7. The nucleic acid of claim 6, wherein the peptide linker encoded by the nucleic acid comprises 15 to 70 amino acids.
8. The nucleic acid according to claims 5 to 7, wherein the peptide linker encoded by the nucleic acid comprises a glycine residue and a serine residue.
9. 9. The nucleic acid of claims 5 to 8, wherein the peptide linker encoded by the nucleic acid comprises repeats of GGS, repeats of GGSS (SEQ ID NO: 584), repeats of GGGS (SEQ ID NO: 572), or repeats of GGGGS (SEQ ID NO: 596).
10. 10. The nucleic acid of claims 5 to 9, wherein the peptide linker encoded by the nucleic acid comprises one or more XTEN16 repeats.
11. 11. The nucleic acid of claim 10, wherein the polypeptide linker encoded by the nucleic acid comprises one XTEN16 repeat, two XTEN16 repeats, or three XTEN16 repeats.
12. The nucleic acid of claim 5, wherein the polypeptide linker encoded by the nucleic acid comprises the amino acid sequence of SEQ ID NOs: 11 to 15.
13. The nucleic acid of claim 5, wherein the nucleic acid sequence encoding the polypeptide linker comprises SEQ ID NOs: 20-24.
14. The nucleic acid of claims 1 to 13, wherein the LSR portion encoded by the nucleic acid comprises an amino acid sequence having at least 90% identity to SEQ ID NOs: 1-5, 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, or 291.
15. The nucleic acid of claim 14, wherein the LSR portion encoded by the nucleic acid comprises the amino acid sequence of SEQ ID NOs: 1-5, 432-443, 445-446, 448-467, 469-476, 478-492, 494-501, 276, 279, 282, 285, 288, or 291.
16. The nucleic acid of claim 15, wherein the LSR portion encoded by the nucleic acid comprises Dn29 (SEQ ID NO: 1), Pf80 (SEQ ID NO: 2), Cp36 (SEQ ID NO: 3), Nm60 (SEQ ID NO: 4), or Si74 (SEQ ID NO: 5).
17. The nucleic acid according to claim 14 or claim 16, wherein the nucleic acid sequence encoding the LSR portion comprises a nucleic acid sequence having at least 90% identity with SEQ ID NOs: 6 to 10.
18. The nucleic acid of claim 16, wherein the nucleic acid sequence encoding the LSR portion comprises the nucleic acid sequence of SEQ ID NOs: 6 to 10.
19. The nucleic acid of claims 1 to 18, wherein the fusion polypeptide encoded by the nucleic acid further comprises one or more nuclear localization signals (NLS).
20. The nucleic acid of claims 1 to 19, wherein the DBD portion encoded by the nucleic acid comprises Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12h, Cas12i, or Cas12g.
21. The nucleic acid of claim 20, wherein the Cas9, Cpf1, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12h, Cas12i, or Cas12g does not have nuclease activity and / or nickase activity.
22. 22. The nucleic acid of claim 21, wherein the DBD portion encoded by the nucleic acid comprises dCas9.
23. The nucleic acid of any one of claims 1 to 19, wherein the DBD portion encoded by the nucleic acid comprises an amino acid sequence having at least 90% identity to dCas9 (SEQ ID NO: 29), dCas9-HF1 (SEQ ID NO: 30), dCas9-SpG (SEQ ID NO: 31), or dCas9-SpG-HF1 (SEQ ID NO: 32).
24. 24. The nucleic acid of claim 23, wherein the DBD portion encoded by the nucleic acid comprises the amino acid sequence of dCas9 (SEQ ID NO: 29), dCas9-HF1 (SEQ ID NO: 30), dCas9-SpG (SEQ ID NO: 31), or dCas9-SpG-HF1 (SEQ ID NO: 32).
25. The nucleic acid according to claims 23 to 24, wherein the nucleic acid sequence encoding the DBD portion comprises a nucleic acid sequence having at least 90% identity with SEQ ID NOs: 33 to 36.
26. The nucleic acid of claim 24, wherein the nucleic acid sequence encoding the DBD portion comprises the nucleic acid sequence of SEQ ID NOs: 33 to 36.
27. 2. The nucleic acid of claim 1, wherein the fusion polypeptide encoded by the nucleic acid comprises Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), Pf80 (SEQ ID NO: 2) and dCas9 (SEQ ID NO: 29), Cp36 (SEQ ID NO: 3) and dCas9 (SEQ ID NO: 29), Nm60 (SEQ ID NO: 4) and dCas9 (SEQ ID NO: 29), or Si74 (SEQ ID NO: 5) and dCas9 (SEQ ID NO: 29).
28. The fusion polypeptide encoded by the nucleic acid further comprises a peptide linker located between the nucleic acid sequence encoding the LSR portion and the nucleic acid sequence encoding the DBD portion, wherein the LSR portion is fused to the N-terminal side of the DBD portion by the peptide linker, and the peptide linker encoded by the nucleic acid is 8 (SEQ ID NO: 11), (GGGGS) 6 (SEQ ID NO: 598), S(GGGGS) 6 S (SEQ ID NO: 12), XTEN16 (SEQ ID NO: 13), XTEN32-(GGSS) 2 (SEQ ID NO: 14), or XTEN48-(GGSS) 2 28. The nucleic acid of claim 27, comprising: (SEQ ID NO: 15).
29. The nucleic acid of claim 1, wherein the fusion polypeptide encoded by the nucleic acid comprises an amino acid sequence having at least 90% identity to SEQ ID NOs: 37-42.
30. The nucleic acid of claim 1, wherein the fusion polypeptide encoded by the nucleic acid comprises the amino acid sequence of SEQ ID NOs: 37 to 42.
31. 31. The nucleic acid of claims 1 to 30, wherein the DBD portion of the fusion polypeptide encoded by the nucleic acid binds to a guide RNA (gRNA).
32. A vector comprising any one of the nucleic acids according to claims 1 to 31.
33. A host cell comprising the vector of claim 32.
34. A nucleic acid editing system comprising a first nucleic acid according to any one of claims 1 to 30 and a second nucleic acid encoding a gRNA.
35. 35. The nucleic acid editing system of claim 34, wherein the gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracr RNA portion, the nucleic acid sequence of the spacer sequence portion is identical to the target nucleic acid sequence except that T in the target nucleic acid sequence is U in the spacer sequence portion, and the target nucleic acid sequence is located within 80 nucleotides upstream or downstream of a dinucleotide core in a binding site for the LSR portion of the fusion polypeptide on the target DNA.
36. The nucleic acid editing system of claim 35, wherein the spacer sequence portion is 16 to 20 nucleotides in length.
37. The nucleic acid editing system of claims 35 to 36, wherein the gRNA encoded by the nucleic acid is an sgRNA.
38. A nucleic acid editing system described in claims 35 to 37, in which a PAM sequence is located immediately 3' to the target nucleic acid sequence on the target DNA.
39. A nucleic acid editing system as described in claims 35 to 38, wherein the target nucleic acid sequence is located within 80 nucleotides upstream or downstream of the dinucleotide core at the attA site for the LSR portion of the fusion polypeptide on the target DNA of interest.
40. 40. The nucleic acid editing system of claim 39, wherein the attA site is a pseudo site in the mammalian target DNA of interest.
41. 41. The nucleic acid editing system of claim 40, wherein the attA site is a pseudo site (attH) in the human genome.
42. The fusion polypeptide encoded by the nucleic acid comprises Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), and the attH sites are chr10:21130404-21130406:-, chr11:77367459-77367461:-, chr1:230490334-230490336:+, chr2:14280297-14280299:+, chr9: 116464427-116464429: +, chr20: 38982599-38982601: +, chr5: 3553012-3553014: -, chr7: 134676315-134676317: -, chr10: 58514255-58514257: +, or chr4: 92338934-92338936: +. The nucleic acid editing system of claim 41.
43. The nucleic acid editing system of claim 41, wherein the fusion polypeptide encoded by the nucleic acid comprises Pf80 (SEQ ID NO: 2) and dCas9 (SEQ ID NO: 29), and the attH site is chr11:64243293-64243295.
44. The nucleic acid editing system of claims 35 to 43, wherein the tracr RNA portion comprises SEQ ID NO:
153.
45. A nucleic acid editing system as described in claims 35 to 43, wherein the target nucleic acid sequence is located on the target DNA within 80 nucleotides upstream of a dinucleotide core at the binding site for the LSR portion of the fusion polypeptide.
46. A nucleic acid editing system as described in claims 35 to 43, wherein the target nucleic acid sequence is located within 80 nucleotides downstream of a dinucleotide core in the binding site for the LSR portion of the fusion polypeptide on the target DNA.
47. 46. The nucleic acid editing system of Claim 45, further comprising a third nucleic acid encoding a second gRNA.
48. 48. The nucleic acid editing system of claim 47, wherein the second gRNA encoded by the nucleic acid comprises a spacer sequence portion and a tracr RNA portion, the nucleic acid sequence of the spacer sequence portion is identical to the target nucleic acid sequence except that T in the target nucleic acid sequence is U in the spacer sequence portion, and the target nucleic acid sequence is located within 80 nucleotides downstream of a dinucleotide core in a binding site for the LSR portion of the fusion polypeptide on the target DNA.
49. The nucleic acid editing system of claim 48, wherein the spacer sequence portion of the second gRNA is 16 to 20 nucleotides in length.
50. A nucleic acid editing system according to claims 47 to 49, wherein the second gRNA encoded by the nucleic acid is an sgRNA.
51. A nucleic acid editing system described in claims 48 to 50, wherein a PAM sequence is present immediately 3' to the target nucleic acid sequence on the target DNA.
52. A nucleic acid editing system according to claims 34 to 46, further comprising a third nucleic acid comprising a donor DNA sequence comprising an attD binding site for the LSR portion of the fusion polypeptide and a nucleic acid sequence for insertion into the target DNA of interest.
53. 53. The nucleic acid editing system of Claim 52, wherein the third nucleic acid further comprises a portion having a target nucleic acid sequence for the gRNA that is identical to the target DNA of interest.
54. the fusion polypeptide encoded by the nucleic acid is Dn29 (SEQ ID NO: 1) and dCas9 (SEQ ID NO: 29), wherein the attH site on the target DNA of interest is located at chromosomal locus chr10:21130404-21130406:-, chr11:77367459-77367461:-, chr1:230490334-230490336:+, chr2:14280297-14280299:+, chr9:116464427-116464429:+, chr20:38982 599-38982601:+, chr5:3553012-3553014:-, chr7:134676315-134676317:-, chr10:58514255-58514257:+, or chr4:92338934-92338936:+, or comprises the attH sequence found at the chromosomal locus, and wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO: 154 or a sequence having 90% identity to SEQ ID NO:
154. Pf80 (SEQ ID NO: 2) and dCas9 (SEQ ID NO: 29), wherein the attH site on the target DNA of interest is located at chromosomal locus chr11:64243293-64243295:+, chr1:162878224-162878226:+, chr11:92763120-92763122:-, chr9:103309977-103309979:-, chr13:91145766-91145768:+, chr2:1024673 61-102467363:+, chr13:99865454-99865456:+, chr9:113640780-113640782:-, chr9:123986548-123986550:-, chr15:53565450-53565452:-, or the attH sequence found at the chromosomal locus, wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO:265 or a sequence having 90% identity to SEQ ID NO:
265. Cp36 (SEQ ID NO: 3) and dCas9 (SEQ ID NO: 29), wherein the attH site on the target DNA of interest is located at chromosomal loci chr16:2789124-2789126:+, chr22:43958465-43958467:-, chr10:117762740-117762742:+, chr7:157294532-157294534:-, chr13:20558930-20558932:-, chr6:15112034 or the attH sequence found at the chromosomal locus is or comprises the attH sequence found at the chromosomal locus, and the attD binding site of the donor DNA sequence comprises SEQ ID NO:267 or a sequence having 90% identity to SEQ ID NO:
267. The attH site on the target DNA of interest comprises Nm60 (SEQ ID NO: 4) and dCas9 (SEQ ID NO: 29), and the attH site on the target DNA of interest is at the chromosomal locus chr9:83308042-83308044:-, chr13:79497139-79497141:-, chr9:131409759-131409761:+, chr4:55980785-55980787:+, chr5:96968267-96968269:+, chr6:37700280- 37700282:-, chr19:17495840-17495842:-, chr5:126546219-126546221:+, chr10:15703649-15703651:-, chr10:395348-395350:+, or comprises the attH sequence found at the chromosomal locus, and wherein the attD binding site of the donor DNA sequence comprises SEQ ID NO:234 or a sequence with 90% identity to SEQ ID NO:234; or Si74 (SEQ ID NO:5) and dCas9 (SEQ ID NO:29), wherein the attH site on the desired target DNA is chromosomal locus chr7:155557356-155557358:+, chr9:77155112-77155114:-, or comprises the attH sequence found at said chromosomal locus; and the attD binding site of the donor DNA sequence comprises SEQ ID NO:266, or a sequence having 90% identity to SEQ ID NO:
266. A nucleic acid editing system according to claims 52 to 53.
55. A nucleic acid editing system according to claims 52 to 54, wherein the third nucleic acid is a plasmid.
56. The nucleic acid editing system of claims 52 to 54, wherein the third nucleic acid is a linear amplicon.
57. A vector comprising any of the nucleic acids of the nucleic acid editing system described in claims 34 to 56.
58. 58. A host cell comprising any of the vectors of claim 57.
59. A nucleic acid editing system as described in claims 34 to 56, wherein the nucleic acid encoding the fusion polypeptide, the nucleic acid encoding the gRNA, or both, and / or, if present, a third nucleic acid encoding the second gRNA, is expressed from an inducible promoter.
60. A method for integrating a donor DNA sequence into a desired target DNA in a cell, the method comprising introducing into the cell a nucleic acid editing system described in claims 34 to 59.
61. 61. The method of claim 60, wherein the cell is a mammalian cell.
62. 62. The method of claim 61, wherein the cell is a human cell.
63. 63. The method of claim 62, wherein the cell is a human embryonic stem cell.
64. 63. The method of claim 62, wherein the cells are hepatocellular carcinoma cells.
65. 65. The method of claims 60 to 64, wherein the target DNA of interest of the cell is modified to contain an attA binding site prior to introduction of the nucleic acid editing system.
66. 66. The method of claims 60-65, wherein the donor DNA comprises an LSR attD binding site that is integrated into the intended target DNA.
67. 67. The method of claims 60 to 66, wherein the target DNA of interest of the cell is the genome of the cell.
68. 67. The method of claims 60 to 66, wherein the target DNA of interest in the cell is a plasmid.
69. A method for inverting a DNA sequence of a target DNA of interest, comprising introducing into a cell a nucleic acid editing system described in claims 34 to 51, wherein the attD binding site and the attA binding site for the LSR portion of the fusion polypeptide are located in opposite orientations on the same target DNA molecule of interest.
70. 70. The method of Claim 69, wherein the target DNA of interest of the cell has been modified prior to introduction of the nucleic acid editing system to contain an attA binding site.
71. 71. The method of claims 69-70, wherein the target DNA of interest of the cell is modified prior to introduction of the nucleic acid editing system to contain an attD binding site.
72. 72. The method of claims 69 to 71, wherein the target DNA of interest of the cell is the genome of the cell.
73. A method for excising a DNA sequence in a target DNA of interest, comprising introducing into a cell a nucleic acid editing system described in claims 34 to 51, wherein the attD binding site and the attA binding site for the LSR portion of the fusion polypeptide are located in the same orientation on the same target DNA molecule of interest.
74. 74. The method of Claim 73, wherein the target DNA of interest of the cell is modified prior to introduction of the nucleic acid editing system to contain an attA binding site.
75. 75. The method of claims 73-74, wherein the target DNA of interest of the cell is modified prior to introduction of the nucleic acid editing system to contain an attD binding site.
76. 76. The method of claims 73 to 75, wherein the target DNA of interest of the cell is the genome of the cell.
77. A method for translocating a DNA sequence between two desired linear target DNA molecules, comprising introducing into a cell a nucleic acid editing system described in claims 34 to 51, wherein an attD binding site for the LSR portion of the fusion polypeptide is located on a first linear target DNA molecule and an attA binding site for the LSR portion of the fusion polypeptide is located on a second linear target DNA molecule.
78. 78. The method of Claim 77, wherein the first target DNA molecule of interest of the cell has been modified prior to introduction of the nucleic acid editing system to contain an attA binding site.
79. 79. The method of claims 77-78, wherein the second target DNA molecule of interest in the cell is modified prior to introduction of the nucleic acid editing system to contain an attD binding site.
80. 76. The method of claims 69 to 75, wherein the linear target DNA molecule of interest of the cell is a chromosome of the cell.