Revision of genetic material using direct replacement editing
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2026-03-18
AI Technical Summary
Current genome editing methods, such as CRISPR-Cas9, face limitations in precision and efficiency due to reliance on endogenous cellular machinery, particularly in nondividing cells, and generate unwanted insertions and deletions, limiting their applicability for therapeutic and research purposes.
A self-contained gene editing system using RNA-guided endonucleases and DNA ligases, which introduce targeted nucleic acid replacements without generating double-stranded breaks, utilizing guide nucleic acids and splints to facilitate precise ligation of donor strands into genomic loci, independent of host cell repair mechanisms.
Enables precise and efficient gene editing in both dividing and nondividing cells, minimizing unwanted mutations and bypassing reliance on endogenous cellular machinery, thereby enhancing the utility of genome editing for therapeutic and research applications.
Smart Images

Figure US2024028906_14112024_PF_FP_ABST
Abstract
Description
REVISION OF GENETIC MATERIAL USING DIRECT REPLACEMENT EDITINGRELATED APPLICATIONS AND INCORPORATION BY REFERENCE
[0001] This application claims priority to US provisional application Serial No. 63 / 465,799, filed May 11, 2023, and US provisional application Serial No. 63 / 637,313, filed April 22, 2024, each incorporated by reference herein in its entirety.
[0002] The foregoing applications, and all documents cited therein or during their prosecution (“appln cited documents”) and all documents cited or referenced in the appln cited documents, and all documents cited or referenced herein (“herein cited documents”), and all documents cited or referenced in herein cited documents, together with any manufacturer’s instructions, descriptions, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated herein by reference, and may be employed in the practice of the invention. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.FIELD OF THE INVENTION
[0003] The invention provides compositions, systems, and methods for revising nucleic acids. The revision may be accomplished using a fusion protein comprising a nuclease, a ligase, an integrase, or a combination thereof. The nucleic acid revision may include ligation of a donor nucleic acid to a target nucleic acid. The nucleic acid editing may include replacement of a portion of the target nucleic acid with a revising nucleic acid.SEQUENCE LISTINGThe instant application contains a Sequence Listing which has been submitted via Patent Center and is hereby incorporated by reference in its entirety. Said .xml copy, created on May 10, 2024 is named Y9506-99016, and is 1,437,114 bytes in size.BACKGROUND OF THE INVENTION
[0004] Improved genome editing methods are needed for replacement or revision of nucleic acid sequences in the genome.
[0005] Citation or identification of any document in this application is not an admission that such document is available as prior art to the present invention.SUMMARY OF THE INVENTION
[0006] Disclosed herein, in some aspects, are systems inside a cell comprising: a fusion protein. In some aspects, are systems or compositions disclosed herein comprise: a DNA-binding protein coupled to a DNA ligase. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the coupling is covalent. Some aspects include a fusion protein comprising the DNA-binding protein (e.g., endonuclease such as an RNA-guided endonuclease) and the DNA ligase. Some aspects include a composition comprising: a cell containing a DNA-binding protein (e.g., endonuclease such as an RNA-guided endonuclease) and a DNA ligase, both of which are heterologous to the cell. Some aspects include a composition comprising: a cell containing a DNA-binding protein that is heterologous to the cell (e.g., endonuclease such as an RNA-guided endonuclease) and a DNA ligase that is endogenous to the cell. Some aspects include a composition comprising: a cell containing a DNA-binding protein (e.g., endonuclease such as an RNA-guided endonuclease) and a DNA ligase, in which the DNA ligase is endogenous to the cell. In some aspects, the DNA- binding protein is amino (N)-terminal relative to the DNA ligase within the fusion protein. In some aspects, the DNA-binding protein is carboxy (C)-terminal relative to the DNA ligase within the fusion protein. In some aspects, the connection comprises a linker comprising 1-100 amino acids. In some aspects, the coupling is non-covalent. In some aspects, the composition comprises a first polypeptide comprising at least part of the DNA-binding protein, and a second polypeptide comprising at least part of the DNA ligase, wherein the first and second polypeptides are non- covalently coupled. In some aspects, the first polypeptide comprises a first heterodimerization domain that binds a second heterodimerization domain, and wherein the second polypeptide comprises the second heterodimerization domain. In some aspects, the heterodimer domains comprise a leucine zipper, PDZ domain, streptavidin, streptavidin binding protein, foldon domain, hydrophobic moiety, or a functional binding fragment thereof. In some aspects, the first polypeptide comprises a first intein that binds a second intein, and wherein the second polypeptide comprises the second intein. In some aspects, the ligase comprises a hairpin binding motif, and wherein the DNA-binding protein and the DNA ligase are coupled with a nucleic acid comprising a scaffold that binds to the DNA-binding protein and a hairpin that binds to the hairpin binding motif. In some aspects, the hairpin binding motif comprises an MS2 coat protein (MCP) peptide, and wherein the hairpin comprises an MS2 hairpin. In some aspects, the DNA-binding protein andthe DNA ligase are coupled with a heterobifunctional molecule comprising an endonuclease binding domain and a DNA ligase binding domain. In some aspects, the heterobifunctional molecule comprises a small molecule. In some aspects, the DNA-binding protein comprises a class II CRISPR / Cas endonuclease. In some aspects, the DNA-binding protein comprises a Cas9 endonuclease. In some aspects, the DNA-binding protein comprises a nickase. In some aspects, the DNA-binding protein comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 1-13, or a functional fragment thereof. In some aspects, the DNA ligase ligates DNA strands base paired to a DNA splint. In some aspects, the DNA ligase ligates DNA strands base paired to an RNA splint. In some aspects, the DNA ligase comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 55-96, or a functional fragment thereof. In some aspects, the DNA-binding protein or the DNA ligase comprises a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, or tag polypeptide. Some aspects include a guide RNA and an integrating nucleic acid. In some aspects, a strand of the integrating nucleic acid acts as the splint for the ligase. In some aspects, the guide nucleic acid acts as a splint for the ligase. In some aspects, the splint comprises a modification. In some aspects, the modification comprises a streptavidin operatively coupled to the splint. In some aspects, the splint comprising the streptavidin is operatively coupled with the DNA ligase, and the DNA ligase is operatively coupled to a biotin. In some aspects, the modification to the splint comprises a biotin operatively linked to the splint. In some aspects, the splint comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a streptavidin. Some aspects include one or more nucleic acids encoding the composition. Some aspects include a cell comprising the composition or comprising the one or more nucleic acids.
[0007] Disclosed herein, in some aspects, are methods of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising: contacting a target nucleic acid in a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a strand break at the predetermined locus of the target nucleic acid. In some aspects, the strand break is a double-strand break.
[0008] Disclosed herein, in some aspects, are methods of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising: contacting a target nucleic acid in a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing astrand break at the predetermined locus of the target nucleic acid; introducing a pre-synthesized integrating nucleic acid to the cell; and ligating a 5' end of the pre-synthesized integrating nucleic acid to a 3' end of the nick at the predetermined locus of the target nucleic acid. In some aspects the strand break is a double-strand break. In some aspects, the strand break is a nick. In some aspects, the endonuclease comprises an RNA-guided endonuclease. In some aspects, the endonuclease comprises a class II CRISPR / Cas endonuclease. In some aspects, the endonuclease comprises a Cas9 endonuclease. In some aspects, the endonuclease comprises a nickase. In some aspects, the endonuclease comprises Cas9 nickase. Some aspects include contacting the endonuclease and the predetermined locus of the target nucleic acid with a guide nucleic acid. In some aspects, said ligating is performed by a ligase coupled to the endonuclease. In some aspects, the endonuclease comprises a fusion partner. In some aspects, the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone Hl central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof. In some aspects, the pre-synthesized integrating nucleic acid comprises a mutation in relation to the target nucleic acid. In some aspects, the nick comprises a single phosphodi ester strand break in the otherwise double stranded target nucleic acid. In some aspects, the nick comprises a non-sticky, non-blunt end of a strand of the target nucleic acid. In some aspects, the target nucleic acid comprises a chromosome of the cell. In some aspects, the cell is eukaryotic.
[0009] Disclosed herein, in some aspects, are systems inside a cell, comprising: a ligase (e.g., a heterologous or an endogenous ligase); an endonuclease that introduces a nick at a predetermined locus of a target nucleic acid; and a pre-synthesized integrating nucleic acid comprising a 5’ end that is ligated by the ligase to a 3' end of the nick at the predetermined locus of the target nucleic acid. In some aspects, the endonuclease comprises a class II CRISPR / Cas endonuclease. In some aspects, the endonuclease comprises Cas9 nickase. Some aspects include a guide nucleic acid that brings the endonuclease into proximity with the predetermined locus of the target nucleic acid. In some aspects, the ligase is coupled to the endonuclease. In some aspects, the pre-synthesized integrating nucleic acid comprises a mutation in relation to the target nucleic acid. In some aspects, the nick comprises a single phosphodiester strand break in the otherwise double stranded target nucleic acid. In some aspects, the nick comprises a non-sticky, non-blunt end of a strand of thetarget nucleic acid. In some aspects, the target nucleic acid comprises a chromosome of a cell. In some aspects, the cell is eukaryotic.
[0010] Disclosed herein, in some aspects, are systems inside a cell, comprising: a ligase; an endonuclease; and an integrating nucleic acid comprising an integrating recombination sequence. In some aspects, the integrating nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid. In some aspects, the strand break is a double-strand break. In some aspects, the strand break is a nick.
[0011] Disclosed herein, in some aspects, are systems inside a cell, comprising: a ligase; an endonuclease; an integrating nucleic acid comprising an integrating recombination sequence; an integrase; and a second integrating nucleic acid, comprising: (i) at least one cognate recombination sequence and (ii) an exogenous nucleic acid. In some aspects, the integrating nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid. In some aspects, the cognate recombination sequence of the second integrating nucleic acid couples with the at least one integrating recombination sequence of the integrating nucleic acid. In some aspects, the integrase integrates the second integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence.
[0012] Disclosed herein, in some aspects, are integrating nucleic acids configured to be ligated by a ligase to a strand break generated by an endonuclease in a target nucleic acid. In some aspects, the integrating nucleic acid comprises at least one attachment site (att) integrating recombination sequence. In some aspects, the integrating nucleic acid comprises an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the integrating nucleic acid comprises two attB integrating recombination sequence. In some aspects, the integrating nucleic acid comprises an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the integrating nucleic acid comprises two attP integrating recombination sequences. In some aspects, the integrating nucleic acid comprises the integrating nucleic acid comprises one attB integrating recombination sequence and one attP integrating recombination sequence. In some aspects, the integrating nucleic acid comprises at least one locus ofX(cross)-overinPl (LoxP) integrating recombination sequence. In some aspects, the integrating nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence.
[0013] Disclosed herein, in some aspects, are second integrating nucleic acids configured to be integrated by an integrase, in whole or in part, into a target nucleic acid at an integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one attachment site (att) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attB integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attP integrating recombination sequences. In some aspects, the second integrating nucleic acid comprises the integrating nucleic acid comprises one attB integrating recombination sequence and one attP integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one locus of X(cross)-over in Pl (LoxP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises a regulatory sequence. In some aspects, the regulatory sequence is a promoter.
[0014] Disclosed herein, in some aspects, are integrating nucleic acids or second integrating nucleic acids comprising a modified nucleotide. In some aspects, the modified nucleotide comprises a methylated nucleotide. In some aspects, the modified nucleotide comprises methylated cytosine (e.g. 5-mC), 5-hydroxymethylcytosine (5-hmC), 5 -formylcytosine (5-fC), 5- carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof In some aspects, the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
[0015] Disclosed herein are integrases that integrate a second integrating nucleic acid, in whole or in part, into a target nucleic acid at an integrating recombination sequence. In some aspects, the integrase is coupled to an endonuclease. In some aspects, the integrase is coupled to a ligase. In some aspects, the integrase is coupled to a recombination directionality factor (RDF). In some aspects, the integrase is a serine integrase. The serine integrase can be a PhiC31 bacteriophage integrase, a Bxbl mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (PaOl), a Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74). In someaspects, the integrase is a tyrosine integrase. The tyrosine integrase can be a Cre recombinase or a flippase (Flp).
[0016] Disclosed herein, in some aspects, are methods of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising: a guide nucleic acid comprising: (a) a spacer complementary to a region of a genomic locus of a genomic strand, (b) a scaffold for complexing with a DNA-binding protein, (c) an optional donor binding site that is at least partially complementary to an integrating nucleic acid, and (d) a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; and a first integrating nucleic acid comprising (i) at least one nucleic acid sequence encoding at least one integrating recombination sequence and (ii) a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by a DNA-binding protein. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. Disclosed herein, in some aspects, are systems of nucleic acids comprising: a guide nucleic acid comprising: (a) a spacer complementary to a region of a genomic locus of a genomic strand, (b) a scaffold for complexing with a DNA-binding protein, and (c) an optional donor binding site that is at least partially complementary to a splinting nucleic acid; an integrating nucleic acid comprising (i) at least one nucleic acid sequence encoding at least one integrating recombination sequence and (ii) a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by a DNA-binding protein; and a splinting nucleic acid comprising a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, and comprising an optional guide binding site that is at least partially complementary to a guide nucleic acid. In some aspects, the genomic strand is in a cell. In some aspects, the splinting nucleic acid further comprises a donor binding site that is at least partially identical or complementary to a portion of the integrating nucleic acid. In some aspects, the guide nucleic acid comprises a sequence of linking nucleic acids between the scaffold and the donor binding site. In some aspects, the guide nucleic acid comprises MS2 binding loops within the scaffold. In some aspects, the guide nucleic acid comprises MS2 binding loops between the scaffold and the donor binding site. In some aspects, the guide nucleic acid, the first integrating nucleic acid, or the splinting nucleic acid comprises a modified intemucleoside linkage. In some aspects, the modified intemucleoside linkage comprises a phosphorothioate linkage. In some aspects, the modified intemucleoside linkage comprises a phosphoacetate linkage. In some aspects, the first integrating nucleic acid, the guide nucleic acid,or the splinting nucleic acid comprises a modified nucleoside. In some aspects, the modified intemucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the guide nucleic acid or the integrating nucleic acid. In some aspects, the guide nucleic acid or the integrating nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, or a combination thereof. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end of the guide nucleic acid or the integrating nucleic acid. The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. In some aspects, the endonuclease comprises an RNA-guided endonuclease. In some aspects, the endonuclease comprises a class II CRISPR / Cas endonuclease. In some aspects, the endonuclease comprises a Cas9 endonuclease. In some aspects, the endonuclease comprises a nickase. In some aspects, the endonuclease comprises a Cas9 nickase.
[0017] In some aspects, a method disclosed herein further comprises a second integrating nucleic acid, wherein the second integrating nucleic acid comprises at least one nucleic acid sequence encoding a cognate recombination sequence; and wherein the at least one cognate recombination sequence of the second integrating nucleic acid couples with the at least one integrating recombination sequence of the first integrating nucleic acid, and wherein an integrase integrates the second integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attB integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attP integrating recombination sequences. In some aspects, the second integrating nucleic acid comprises the integrating nucleic acid comprises one attB integrating recombination sequence and one attP integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one locus of X(cross)-over in Pl (LoxP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises a regulatory sequence. In some aspects, the regulatory sequence is a promoter. In some aspects, the integrase is coupled to arecombination directionality factor (RDF). In some aspects, the integrase is a serine integrase. The serine integrase can be a PhiC31 bacteriophage integrase, a Bxbl mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (PaOl), a Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74). In some aspects, the integrase is a tyrosine integrase. The tyrosine integrase can be a Cre recombinase or a flippase (Flp). In some aspects, the first integrating nucleic acid or the second integrating nucleic acid comprises a modified nucleotide. In some aspects, the modified nucleotide comprises a methylated nucleotide. In some aspects, the modified nucleotide comprises methylated cytosine (e.g. 5-mC), 5 -hydroxymethyl cytosine (5- hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof. In some aspects, the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof. In some aspects, the splinting nucleic acid comprises a modification. In some aspects, the modification comprises a streptavidin operatively coupled to the splint. In some aspects, the splint comprising the streptavidin is operatively coupled with the DNA ligase, and the DNA ligase is operatively coupled to a biotin. In some aspects, the modification to the splint comprises a biotin operatively linked to the splint. In some aspects, the splint comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a streptavidin. In some aspects, the endonuclease comprises a fusion partner. In some aspects, the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone Hl central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof.
[0018] Disclosed herein, in some aspects, are fusion proteins, comprising: a DNA-binding protein connected to a DNA ligase. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the connection between the DNA-binding protein and the DNA ligase is covalent. Some aspects include a fusion protein comprising the DNA-binding protein upstream of the DNA ligase. Some aspects include a fusion protein comprising the DNA-binding protein downstream of the DNA ligase. In some aspects, the connection comprises a linker comprising 1-100 amino acids. In some aspects, the composition comprises a first polypeptide comprising at least part of the DNA-binding protein, and a second polypeptide comprising at least part of the DNA ligase, wherein the first and second polypeptides are bound together covalently or non-covalently. In some aspects, the first polypeptide comprisesa first heterodimerization domain that binds a second heterodimerization domain, and wherein the second polypeptide comprises the second heterodimerization domain. In some aspects, the heterodimer domains comprise a leucine zipper, PDZ domain, streptavidin, streptavidin binding protein, foldon domain, hydrophobic moiety, or a functional binding fragment thereof. In some aspects, the first polypeptide comprises a first intein that binds a second intein, and wherein the second polypeptide comprises the second intein. In some aspects, the DNA-binding protein and the DNA ligase are bound together by a small molecule. In some aspects, the DNA-binding protein comprises a class II CRISPR / Cas endonuclease. In some aspects, the DNA-binding protein comprises a Cas9 endonuclease. In some aspects, the DNA-binding protein comprises a nickase. In some aspects, the DNA-binding protein comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 1-13, or a functional fragment thereof. In some aspects, the DNA ligase ligates DNA strands base paired to a DNA splint. In some aspects, the DNA ligase ligates DNA strands base paired to an RNA splint. In some aspects, the DNA ligase comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 55-96, or a functional fragment thereof. In some aspects, the DNA-binding protein or the DNA ligase comprises a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or streptavidin. Some aspects include a guide RNA and an integrating nucleic acid. Some aspects relate to a cell comprising the composition. Some aspects include a nucleic acid encoding the composition. Some aspects include one or more nucleic acids encoding the first or second polypeptides. Some aspects include an editing method (e.g., nucleic acid) which uses the composition. Some aspects include a method of treatment using the composition. Some aspects include administering the composition to a subject.
[0019] Disclosed herein, in some aspects, are editing methods, comprising ligating an integrating recombination sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease.
[0020] Disclosed herein, in some aspects, are fusion proteins, comprising: a DNA-binding protein fused to a DNA ligase. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. Disclosed herein, in some aspects, are fusion proteins, comprising: an integrase, and a ligase or an endonuclease. In some aspects, a fusion protein disclosed herein comprises an integrase and a ligase. In some aspects, a fusion protein disclosed herein comprises an integrase, a ligase, and an endonuclease. In some aspects, afusion protein disclosed herein comprises an integrase and an endonuclease. Disclosed herein, in some aspects, are protein complexes, comprising: a DNA-binding protein bound to a DNA ligase. In some aspects, the endonuclease and the DNA ligase are bound together through heterodimerization domains. In some aspects, the heterodimerization domains comprise leucine zippers, PDZ domains, streptavidin, and streptavidin binding protein, foldon domains, hydrophobic polypeptides, an antibody that binds the Cas nickase, or an antibody that binds the DNA ligase, or one or more binding fragments thereof. Disclosed herein, in some aspects, are cells comprising the fusion protein or the protein complex. Disclosed herein, in some aspects, are cells comprising a heterologous DNA-binding protein and a DNA ligase that was introduced into the cell. Some aspects include a nuclease that is different from the DNA-binding protein. Disclosed herein, in some aspects, are guide nucleic acids, comprising: a spacer at least partially reverse complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; and a flap binding site at least partially reverse complementary to a nucleic acid flap, and an integrating nucleic acid binding site. Disclosed herein, in some aspects, are integrating nucleic acids, comprising: a single or double-stranded DNA region to be inserted into a target nucleic acid, wherein the single or double-stranded DNA region is flanked by at least one additional single-stranded region comprising a guide binding site. Disclosed herein, in some aspects, are editing systems, comprising a DNA-binding protein, the guide nucleic acid, and the integrating nucleic acid. Disclosed herein, in some aspects, are editing methods, comprising: contacting a target nucleic acid with the editing system and a DNA ligase.
[0021] Disclosed herein, in some aspects, are methods comprising ligating an integrating sequence to a nick in a target nucleic acid in a cell, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating sequence comprises a modified nucleotide. In some aspects, the integrating nucleic acid comprises a modified nucleotide. In some aspects, the modified nucleotide comprises a methylated nucleotide. In some aspects, the modified nucleotide comprises methylated cytosine (e.g. 5-mC), 5 -hydroxymethyl cytosine (5-hmC), 5-formylcytosine (5-fC), 5 -carboxyl cytosine (5-caC), N6- methyladenine (6-mA), or a combination thereof. In some aspects, the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
[0022] Disclosed herein, in some aspects, are methods comprising ligating an integrating sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating of the integrating sequence introduces methylated nucleosides into the genome.
[0023] Disclosed herein, in some aspects, are methods comprising ligating an integrating sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating of the integrating sequence removes methylated nucleosides from the genome.
[0024] Disclosed herein, in some aspects, are methods, comprising: (a) contacting a target nucleic acid in a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a nick at the predetermined locus of the target nucleic acid; (b) introducing a pre-synthesized, methylated integrating nucleic acid; and (c) ligating an end of the integrating nucleic acid to an end of the nick at the predetermined locus of the target nucleic acid.
[0025] Disclosed herein, in some aspects, are editing systems comprising: a ligase; an endonuclease that introduces a nick at a predetermined locus of a target nucleic acid; and a presynthesized, methylated integrating nucleic acid comprising an end that is ligated by the ligase to an end of the nick at the predetermined locus of the target nucleic acid.
[0026] Disclosed herein, in some aspects, are methods comprising: contacting an endonuclease with a splinting nucleic acid, the endonuclease comprising a first heterodimerization moiety, and the splinting nucleic acid comprising a second heterodimerization moiety. In some aspects, a method disclosed herein further comprises, introducing a nick or a strand break at a predetermined locus of a target nucleic acid by the endonuclease, and ligating an integrating nucleic acid to the nick or to the strand break, wherein the splinting nucleic acid binds to the integrating nucleic acid and to the target nucleic acid. In some aspects, the first heterodimerization moiety comprises biotin, and the second heterodimerization moiety comprises streptavidin or avidin, or wherein the second heterodimerization moiety comprises biotin, and the first heterodimerization moiety comprises streptavidin or avidin.
[0027] Disclosed herein, in some aspects, are systems, comprising: an endonuclease, wherein the endonuclease is coupled to a first heterodimerization moiety; and a splinting nucleic acid comprising a second heterodimerization moiety. In some aspects, a system disclosed herein further comprises a ligase. In some aspects, the endonuclease introduces a nick or a strand break at apredetermined locus of a target nucleic acid, and the ligase ligates an integrating nucleic acid to the nick or to the strand break, wherein the splinting nucleic acid binds to the integrating nucleic acid and to the target nucleic acid. In some aspects, the first heterodimerization moiety comprises biotin, and the second heterodimerization moiety comprises streptavidin or avidin, or wherein the second heterodimerization moiety comprises biotin, and the first heterodimerization moiety comprises streptavidin or avidin.
[0028] Disclosed herein, in some aspects, are systems inside a cell: at least one DNA-binding protein; at least one guide nucleic acid comprising: a spacer at least partially complementary to a genomic locus in a cell; a scaffold for complexing with the at least one DNA-binding protein; and an optional donor binding site that is at least partially complementary to an integrating nucleic acid; and at least one DNA ligase; and the integrating nucleic acid, comprising a flap binding site at least partially reverse complementary to a nucleic acid flap and optionally comprising a guide binding site that is at least partially complementary to the at least one guide nucleic acid, wherein the at least one DNA-binding protein cleaves or nicks at least one strand of the genomic locus, and wherein the at least one DNA ligase ligates an end of the integrating nucleic acid to the genomic flap site, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a singlestranded DNA. In some aspects, the integrating nucleic acid comprises a double-stranded DNA.
[0029] Disclosed herein, in some aspects, are systems inside a cell comprising: at least one DNA-binding protein comprising a first DNA-binding protein and an optional second DNA- binding protein; at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid comprising: a first spacer complementary to a first region of a genomic locus in a cell; a first scaffold for complexing with the first DNA-binding protein; and an optional first donor binding site that at least partially complementary to an integrating nucleic acid; and a first flap binding site that is at least partially identical or complementary to a first genomic flap at or adjacent to the genomic locus; and the second guide nucleic acid comprising: a second spacer complementary to a second region of the genomic locus in the cell; a second scaffold for complexing with the first or second DNA-binding protein; an optional second donor binding site that at least partially complementary to the integrating nucleic acid; and a second flap binding site that is at least partially identical or complementary to a secondgenomic flap at or adjacent to the genomic locus; at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; and at least one integrating nucleic acid comprising a first strand and a second strand: wherein the first strand comprises an optional first guide binding site that is at least partially complementary to the first guide nucleic acid; and wherein the second strand comprises an optional second guide binding site that is at least partially complementary to the second guide nucleic acid, wherein the first DNA-binding protein and / or the second DNA- binding protein each cleaves or nicks at least one strand of the genomic locus in the cell; and wherein the first DNA ligase ligates an end of the first strand of the integrating nucleic acid to the first genomic flap; and the first or second DNA ligase ligates an end of the second strand of the integrating nucleic acid to the second genomic flap, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. In some aspects, the integrating nucleic acid comprises a double-stranded DNA duplex region. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a 5’ overhang optionally comprising the first guide binding site. In some aspects, the integrating nucleic acid comprises a 5’ overhang optionally comprising the second guide binding site.
[0030] Disclosed herein, in some aspects, are systems inside a cell comprising: at least one DNA-binding protein; at least one guide nucleic acid comprising: a spacer complementary to a genomic locus in a cell; a scaffold for complexing with the at least one DNA-binding protein; and an optional donor binding site that is at least partially complementary to an integrating nucleic acid; at least one DNA ligase; and the integrating nucleic acid that: comprises an optional guide binding site that is at least partially complementary to the at least one guide nucleic acid; and comprises a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, wherein the at least one DNA-binding protein cleaves or nicks at least one strand of the genomic locus; and wherein the at least one DNA ligase ligates an end of the integrating nucleic acid to the genomic flap, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a DNA comprising a 3’ overhang. In some aspects, the 3’ overhang comprises the guide binding site. In some aspects, the 3’ overhang comprises the flapbinding site. In some aspects, the at least one DNA ligase ligates a strand of the integrating nucleic acid to the genomic nucleic acid sequence.
[0031] Disclosed herein, in some aspects, are systems inside a cell comprising: at least one DNA-binding protein comprising a first DNA-binding protein and an optional second DNA- binding protein; at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid comprising: a first spacer complementary to a first region of a genomic locus in a cell; a first scaffold for complexing with the first DNA-binding protein; and an optional first donor binding site that at least partially complementary to an integrating nucleic acid; and the second guide nucleic acid comprising: a second spacer complementary to a second region of the genomic locus in the cell; a second scaffold for complexing with the first or second DNA-binding protein; and an optional second donor binding site that at least partially complementary to the integrating nucleic acid; and at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; and the integrating nucleic acid comprising a first strand and a second strand: wherein the first strand comprises an optional first guide binding site that is at least partially complementary to the first guide nucleic acid; wherein the second strand comprises an optional second guide binding site that is at least partially complementary to the second guide nucleic acid; wherein the first strand comprises a first flap binding site that is at least partially identical or complementary to a first genomic flap at or adjacent to the genomic locus; and wherein the second strand comprises a second flap binding site that is at least partially identical or complementary to a second genomic flap at or adjacent to the genomic locus; wherein the first DNA-binding protein and / or the second DNA-binding protein each cleaves or nicks at least one strand of the genomic locus in the cell; and wherein the first DNA ligase ligates an end of the first strand of the integrating nucleic acid to the first genomic flap; and the first or second DNA ligase ligates an end of the second strand of the integrating nucleic acid to the second genomic flap, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a double-stranded DNA duplex region. In some aspects, the double-stranded DNA comprises a 3 ’ overhang optionally comprising the first guide binding site and comprising the first flap binding site. In some aspects, the double stranded DNA comprises a 3’ overhang optionally comprising the second guide binding site and comprising the second flap binding site.
[0032] The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the at least one DNA-binding protein comprises a Cas protein or a functional fragment thereof. In some aspects, the Cas protein or the functional fragment thereof comprises nickase activity. In some aspects, the at least one DNA- binding protein comprises a Cas9 nickase or a functional fragment thereof. In some aspects, the at least one DNA ligase ligates nucleic acids bound to DNA. In some aspects, the at least one DNA ligase ligates nucleic acids bound to RNA. In some aspects, the at least one DNA ligase comprises a PBCV-1 DNA ligase. In some aspects, the at least one DNA ligase is operatively coupled to the at least one DNA-binding protein. In some aspects, the at least one DNA ligase is fused to the at least one DNA-binding protein as a fusion polypeptide. In some aspects, the at least one DNA- binding protein and the at least one DNA ligase each comprises a heterodimer domain. In some aspects, the at least one DNA-binding protein and the at least one DNA ligase forms a heterodimer via the heterodimer domain. In some aspects, the at least one DNA-binding protein comprises a linker. In some aspects, the linker connects the Cas protein or a functional fragment thereof to the heterodimer domain. In some aspects, the at least one DNA-binding protein comprises a localization signal sequence. In some aspects, the at least one DNA ligase comprises a localization signal sequence. In some aspects, the localization signal sequence comprises a nuclear localization sequence (NLS). In some aspects, the at least one DNA-binding protein or the at least one DNA ligase are directed to nucleus of the cell by the NLS. In some aspects, the at least one integrating nucleic acid, such as a donor nucleic acid or a revising nucleic acid, corrects at least one genetic mutation in the at least one genomic locus. In some aspects, the at least one integrating nucleic acid inserts a coding sequence. In some aspects, the coding sequence encodes a full-length protein. In some aspects, the at least one integrating nucleic acid inserts a non-coding sequence. In some aspects, the non-coding sequence comprises a recombination sequence. In some aspects, the noncoding sequence knocks out an endogenous gene. In some aspects, the non-coding sequence comprises a regulatory element. Some aspects further include a nuclease. In some aspects, the nuclease comprises an exonuclease for digesting the genomic flap. In some aspects, the nuclease comprises a human flap endonuclease 1 (hFENl), a human exonuclease 5 (hEXO5), a T5 exonuclease, a T7 exonuclease, an exonuclease VIII, a flap endonuclease domain of E. coli Poll, a RecJF, a Lambda exonuclease, a Xni (ExoIXI), a SaFEN (Staphylococcus aureus FEN), a nuclease BAL-31, or a fragment thereof. In some aspects, the heterologous nuclease comprises anendonuclease for digesting the genomic flap, and the endonuclease is different from the at least one DNA-binding protein. In some aspects, the at least one DNA-binding protein comprises at least one additional functional domain. In some aspects, the at least one additional functional domain comprises a chromatin modifying domain. In some aspects, the at least one additional functional domain comprises a cell penetrating peptide. In some aspects, the at least one guide nucleic acid comprises at least one nucleic acid modification. In some aspects, the at least one nucleic acid modification comprises a modification to a backbone, a sugar, a base, or a combination thereof. In some aspects, the at least one DNA-binding protein is complexed with the at least one guide nucleic acid. In some aspects, the at least one guide nucleic acid is complexed with the integrating nucleic acid. In some aspects, the at least one DNA-binding protein, the at least one guide nucleic acid, the at least one at least one DNA ligase, the integrating nucleic acid, or a combination thereof is encoded by a polynucleotide. In some aspects, the polynucleotide comprises mRNA. In some aspects, the polynucleotide comprises a vector. In some aspects, the vector comprises a viral vector. In some aspects, the at least one DNA-binding protein, the at least one guide nucleic acid, the at least one at least one DNA ligase, the integrating nucleic acid, or a combination thereof is encapsulated by at least one lipid nanoparticle. In some aspects, the cell comprises a bacterial cell, a eukaryotic cell, or a plant cell. In some aspects, the eukaryotic cell comprises a mammalian cell. Some aspects include a composition comprising the system. Some aspects include a cell comprising the system. Some aspects include a cell line comprising the cell. Some aspects include a pharmaceutical composition comprising the system. Some aspects include a pharmaceutical composition comprising the composition. Some aspects include a pharmaceutical composition comprising the cell. Some aspects include a pharmaceutically acceptable: excipient, carrier, or diluent. In some aspects, the pharmaceutical composition is formulated for administering intrathecally, intraocularly, intravitreally, retinally, intravenously, intramuscularly, intraventricularly, intracerebrally, intracerebellarly, intracerebroventricularly, intraperenchymally, subcutaneously, intratumorally, pulmonarily, endotracheally, intraperitoneally, intravesically, intravaginally, intrarectally, orally, sublingually, transdermally, by inhalation, by inhaled nebulized form, by intraluminal -GI route, or a combination thereof to a subject in need thereof. Some aspects include a kit comprising: the system, the composition, or the pharmaceutical composition and a container. In some aspects, include method for modifying a cell comprising contacting a cell with the system. In some aspects, include method for modifying a cellcomprising contacting a cell with the composition. In some aspects, include method for modifying a cell comprising contacting a cell with the pharmaceutical composition. In some aspects, the cell is not a dividing cell. In some aspects, the integrating nucleic acid is inserted into the genomic locus of the cell independent of endogenous non-homologous end joining (NHEJ) and independent of endogenous homology-directed repair (HDR). Some aspects include a method for treating a disease or condition in subject in need thereof comprising: contacting the cell or the subject with the system, the composition, or the pharmaceutical composition; replacing a genomic locus in a cell with an integrating nucleic acid, thereby treating the disease or condition in the subject. In some aspects, the cell is not a dividing cell. In some aspects, the integrating nucleic acid is inserted into the genomic locus of the cell independent of endogenous non-homologous endjoining (NHEJ) and independent of endogenous homology-directed repair (HDR).
[0033] Disclosed herein, in some aspects, are guide nucleic acids comprising: a spacer that is at least partially complementary to a genomic locus in a cell; a scaffold for complexing with a DNA-binding protein; and a donor binding site that is at least partially complementary to an integrating nucleic acid. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the guide nucleic acid comprises a flap binding site that is at least partially complementary to a genomic sequence of the genomic locus. In some aspects, the guide nucleic acid comprises at least one nucleic acid modification. In some aspects, the at least one nucleic acid modification comprises a modification to a backbone, a sugar, a base, or a combination thereof. In some aspects, the guide nucleic acid comprises RNA sequence.
[0034] Accordingly, it is an object of the invention not to encompass within the invention any previously known product, process of making the product, or method of using the product such that Applicants reserve the right and hereby disclose a disclaimer of any previously known product, process, or method. It is further noted that the invention does not intend to encompass within the scope of the invention any product, process, or making of the product or method of using the product, which does not meet the written description and enablement requirements of the USPTO (35 U.S.C. §112, first paragraph) or the EPO (Article 83 of the EPC), such that Applicants reserve the right and hereby disclose a disclaimer of any previously described product, process of making the product, or method of using the product. It may be advantageous in the practice of the invention to be in compliance with Art. 53(c) EPC and Rule 28(b) and (c) EPC. All rights to explicitlydisclaim any embodiments that are the subject of any granted patent(s) of applicant in the lineage of this application or in any other lineage or in any prior filed application of any third party is explicitly reserved. Nothing herein is to be construed as a promise.
[0035] It is noted that in this disclosure and particularly in the claims and / or paragraphs, terms such as "comprises", "comprised", "comprising" and the like can have the meaning attributed to it in U.S. Patent law; e.g., they can mean "includes", "included", "including", and the like; and that terms such as "consisting essentially of' and "consists essentially of' have the meaning ascribed to them in U.S. Patent law, e.g., they allow for elements not explicitly recited, but exclude elements that are found in the prior art or that affect a basic or novel characteristic of the invention.
[0036] These and other embodiments are disclosed or are obvious from and encompassed by, the following Detailed Description.BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The patent or application fde contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0038] The following detailed description, given by way of example, but not intended to limit the invention solely to the specific embodiments described, may best be understood in conjunction with the accompanying drawings.
[0039] Fig. 1A illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus.
[0040] Fig. IB follows sequentially from Fig. 1A and illustrates a donor strand incorporated into one side of a genomic locus, the donor strand having displaced a genomic flap.
[0041] Fig. 1C follows sequentially from Fig. IB and illustrates a donor strand incorporated into one side of a genomic locus, and a nick appearing where a genomic flap has been removed.
[0042] Fig. 2A illustrates 2 guide nucleic acids, 2 endonucleases, 2 ligases, and a donor strand at a genomic locus.
[0043] Fig. 2B follows sequentially from Fig. 2A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced 2 genomic flaps.
[0044] Fig. 2C follows sequentially from Fig. 2B and illustrates a donor strand incorporated into a genomic locus, and 2 nicks appearing where genomic flaps have been removed.
[0045] Fig. 3A illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus.
[0046] Fig. 3B follows sequentially from Fig. 3A and illustrates a donor strand incorporated into one side of a genomic locus, the donor strand having displaced a genomic flap.
[0047] Fig. 3C follows sequentially from Fig. 3B and illustrates a donor strand incorporated into one side of a genomic locus, and a nick appearing where a genomic flap has been removed.
[0048] Fig. 4A illustrates 2 guide nucleic acids, 2 endonucleases, 2 ligases, and a donor strand at a genomic locus.
[0049] Fig. 4B follows sequentially from Fig. 4A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced 2 genomic flaps.
[0050] Fig. 4C follows sequentially from Fig. 4B and illustrates a donor strand incorporated into a genomic locus, and 2 nicks appearing where genomic flaps have been removed.
[0051] Fig. 5A illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus.
[0052] Fig. 5B follows sequentially from Fig. 5A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced a genomic flap.
[0053] Fig. 5C follows sequentially from Fig. 5B and illustrates a donor strand incorporated into one side of a genomic locus, and a nick appearing where a genomic flap has been removed.
[0054] Fig. 6A illustrates 2 guide nucleic acids, 2 endonucleases, 2 ligases, and a donor strand at a genomic locus.
[0055] Fig. 6B follows sequentially from Fig. 6A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced 2 genomic flaps.
[0056] Fig. 6C follows sequentially from Fig. 6B and illustrates a donor strand incorporated into a genomic locus, and 2 nicks appearing where genomic flaps have been removed.
[0057] Fig. 7 illustrates some examples of fusion protein arrangements.
[0058] Fig. 8A illustrates an exemplary nicking and ligation pattern of an exogenous first integrating nucleic acid.
[0059] Fig. 8B illustrates a DNA gel showing a pattern associated with 1 -Sided Replacer 2 performed in vitro using 30nt GBS / DBS and thermostable T4 ligase. Using a 30nt GBS / DBS combination, a donor containing a protospacer adj acent motif (PAM) mutation, and a thermostable T4 ligase (Hi-T4, NEB), we were able to produce a final Replacer product (Lane 3) correspondingto the size of our control product (Lane 1). Replacer products were not detected in the absence of nicking Cas9 (Cas9n) (Lane 2), or in the absence of the bottom donor which serves as the splint (Lanes 4 & 5).
[0060] Fig. 8C illustrates an exemplary nucleic acid gel showing a pattern associated with in vitro 1-Sided Replacer 2 using variable length GBS / DBS combinations and T4 ligase. Using regular T4 ligase (NEB), a final Replacer product corresponding to the size of the control when using multiple GBS / DBS combinations was produced, including no GBS / DBS, 20nt GBS / DBS, and 30nt GBS / DBS. Additionally, in this experiment, recoded dsDNA donors containing PAM mutation were more efficient at producing final Replacer products compared to PAM mutant dsDNA donors that were not recoded.
[0061] Fig. 9 illustrates measurement of a percentage of cells expressing green fluorescent protein (GFP), indicating gene editing from BFP to GFP by a 1 -sided Replacer 2 with nicking Cas9 and DNA ligase.
[0062] Fig. 10 illustrates sequencing reads merged and aligned to an amplicon of interest and a percentage of total reads that matched an intended edit via a 1 -sided replacer 2 with a nicking Cas9 and a T4 DNA ligase.
[0063] Fig. 11 illustrates sequencing reads merged and aligned to an amplicon of interest and a percentage of total reads that matched an intended edit via a 2-sided replacer 2 with a nicking Cas9 and a T4 DNA ligase.
[0064] Fig. 12 illustrates measurement of a percentage of cells expressing green fluorescent protein (GFP), indicating gene editing from BFP to GFP via a 1 -Sided Replacer 2 with a nicking Cas9 and a T4 DNA Ligase.
[0065] Fig. 13A and Fig. 13B depict example fusion proteins, which may be useful for a genome revising system. In some embodiments, an endonuclease, DNA ligase, and integrase can be delivered as three individual proteins (not illustrated). An endonuclease, DNA ligase, and integrase can be delivered as a combination of one double fusion protein and a third protein delivered in trans (Fig. 13A). An endonuclease, DNA ligase, and integrase can be delivered as one triple fusion protein (Fig. 13B). A linker such as a peptide linker may be included between any two of the components in each fusion protein, or additional components may be added.
[0066] Fig. 14A-14D depict vectors that can be used to deliver the second integrating nucleic acid (e.g., revising nucleic acid) to the host cell. The vector can contain a single integraseatachment site (Fig. 14A and Fig. 14B) or a plurality of integrase attachment sites (Fig. 14C and Fig. 14D). The vector can be a minicircle (Fig. 14A), a plasmid (Fig. 14B and Fig. 14C), or linearized DNA (Fig. 14D).
[0067] Fig. 15 shows an exemplary nucleic acid gel having a pattern associated with the replacement of either a 222-base pair (bp) sequence (lanes 1 and 2) or a 131 bp sequence (lanes 4 and 5) with a 38 bp attB recombination sequence. Lanes 3 and 6 represent the deletion, without replacement, of the 222 bp sequence and the 131 bp sequence, respectively. The lane on the far left represents the non-edited sequence.
[0068] Fig. 16 depicts amplicon sequencing data that quantifies the percentage of reads in a pool of cells that contain the expected precise edit of replacement of a 222 bp or a 131 bp sequence with an attB recombination sequence. The bars labeled numbers 1-6 represent the same reactions as those in Fig. 15.
[0069] Fig. 17 illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus. Illustrated on the donor strand is a 3' chemical modifications (e.g., a C3 spacer or an inverted dT), and illustrated on the splint are additional chemical modifications (e.g., 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer or 3' inverted dT on the splint.)
[0070] Fig. 18 provides sequencing data that quantifies the percentage of reads in a pool of cells, showing the editing efficiency of the ATPase copper transporting beta (ATP7B) gene using genome revising systems comprising splints having different end blocking modifications.
[0071] Fig. 19 provides sequencing data that quantifies percentages of reads in a pool of cells, showing the editing efficiency of a cystic fibrosis transmembrane conductance regulator (CFTR) gene using genome revising systems comprising splints having different end blocking modifications.
[0072] Fig. 20 illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus. In this illustration, a biotinylated splinting nucleic acid is attached to a monomeric streptavidin fused to nicking Cas9. The monomeric streptavidin can also be fused to the ligase (not illustrated).
[0073] Fig. 21 provides example fusion proteins, which may be useful for a genome revising system. In some embodiments, a DNA-binding domain of the Rad51 DNA repair protein (rad51DBD), and / or a high-mobility group nucleosome binding domain 1 (HN1) and a histone Hl central globular domain (H1G) are fused to the fusion protein.
[0074] Fig. 22 provides sequencing data that quantifies the percentage of reads in a pool of cells, showing the editing efficiency of the CFTR gene using genome revising systems comprising fusion proteins comprising, rad51DBD and / or HN1 and H1G.
[0075] Fig. 23 provides sequencing data that quantifies the percentage of reads in a pool of cells, showing the editing efficiency of the CFTR gene using genome revising systems comprising bicistronic mRNAs encoding for phosphomimetic peptide from IGFl(IGFlpml) and N-terminal peptide from NFATC2IP (NFATC2IPpl) peptides (IN peptides).
[0076] Fig. 24 illustrates a model of a G to T point mutation mediated by a 1 -sided Replacer mechanism comprising a donor nucleic acid and a splint nucleic acid. The figure depicts displacement of a 5’ flap of the target DNA, ligation of a donor nucleic acid comprising a singlebase substitution to a 3’ flap of the target DNA flap, and resolution of the mismatch by a mismatch repair (MMR) pateway.
[0077] Fig. 25 provides a model of ligase- mediated programmable gene integration (PGI) or L-PGI. The model depicts replacement or deletion mediated by a 2 x 1 -sided Replacer. (See also, Fig. 1A and Fig. 3 A) Target modifications (replacements or deletions) are determined by nick location and the sequences of the overlapping flaps formed by donor nucleic acids. Target modifications are made with single base pair resolution.
[0078] Fig. 26 illustrates the effects of single-stranded overhangs on stability of a donor-splint nucleic acid duplex at a physiological temperature (left panel). Affinity of complementary sequences can be modulated by incorporation of locked nucleic acids (LNAs) (right panel). Particularly in cases where the donor nucleic acid, but not the splinting nucleic acid is integrated at the target (see, e.g. Fig. 1 A, 2A, 3A, 5A, 6A), the splint can be freely modified.
[0079] Fig. 27 illustrates a Replacer programmed to mediate three point mutations to mutate a target coding sequence to encode GFP and disrupt the PAM.
[0080] Fig. 28 illustrates optimization of guide / splint duplex region formed by the splint biding site (SBS) of the guide and guide binding site (GBS) of the splint (top panel). The system can be optimized by adjusting the GC content and length of the SBS / GBS region (bottom left). The optimal length of the SBS / GBS region in the HEK293T GFP example is about 19 bp (bottom right). In this case, there appears to be no substantial benefit from inclusion of linker nucleotides between the flap binding site (FBS) and GBS (bottom right).
[0081] Fig 29 illustrates optimization of the flap binding site (FBS) (top panel) and donor / splint DNA dose relative to the amount of gRNA. In the HEK293T GFP example, the optimal FBS is about 9 to 13 bases (bottom left). In the HEK293T GFP example, for a 13 base FBS, the optimal molar amount of splint / donor is about 18% of the gRNA molar amount (bottom right).
[0082] Fig. 30 illustrates optimization of the donor and DBS length (top panel). In the HEK293T GFP example, there is substantial activity over a broad range of donor lengths and an optimal length of about 22-26 nt (bottom left). Overhangs of donor or splint nucleic acids were observed to reduce efficiency (bottom right).
[0083] Fig. 31 illustrates optimization of locked nucleic acids (LNAs). Fig. 31A depicts effect of varying the number of LNAs in the splint within the first 20 nt of the donor binding site (DBS) alternating from the 5’ end of the splint and within the 19 nt of the guide binding site (GBS) at the 3’ end of the splint. Fig. 31 B depicts effect of inclusion of additional LNAs in the splint near the target nick site or at the 5’ end of the splint. Fig. 31C depicts the effect of varying the number of LNAs in the donor binding site (DBS) of the splint for a 32 nt DBS. The comparison is among 12 LNAs distributed towards the target nick site, 12 LNAs distributed towards the 5’ end of the splint, and 16 alternating LNAs. Fig. 31D(SEQ ID Nos:815,556 and 816) depicts the locations and exemplary compositions of the nucleic acid binding sites for a splint with a 20 nt DBS and 19 nt GBS. (Donor: SEQ ID NO:556, / 5Phos / cgtaTgtcagggtggtcacGACgg; Splint: SEQ ID NO:557, +C*C*+GT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGgcgtgcagtgcttACGCCA+CA+AT+A C+CG+CA+G*C*+T; Guide: 3’ end of guide, e g. SEQ ID NO: 166, SEQ ID NO: 170, SEQ ID NO: 173, or SEQ ID NO:222).
[0084] Fig. 32 illustrates functional aspects of the system and components, including homology-directed repair (HDR)-independence. The high efficiency of Replacer editing is not primarily due to HDR (Fig. 32A). The efficiency of Replacer editing depends on all components, splint, donor, Cas9, and ligase. The exogenous ligase substantially increases ligase activity that may be present in target cells (Fig. 32B).
[0085] Fig. 33. Replacer efficiency can be enhanced by incorporation of nucleotide analogs such as pseudo-UTP enhance (Fig. 33A). A variety of ligases are effective. SplintR and T4 ligase demonstrated particularly high efficiency (Fig. 33B, C). Expressing nCas9 and T4 ligase from separate mRNAs resulted in higher efficiency (left panel, split nCas9 & T4, 2 mRNAs). This maybe due to the size of the mRNAs and not the protein itself, as a T4-P2A-nCas9 mRNA expressing a self-cleaving fusion (Fig. 33C) had the same efficiency to a T4-nCas9 fusion.
[0086] Fig. 34 illustrates gene editing of several endogenous target genes in HEK293T cells prior to optimization. Point mutations were introduced at reportedly high efficiency targets (e.g., AAVS1 (a safe harbor site; Anzalone et. al., Nature biotechnology 2022), and VEGFA (Wang et. al., Nature methods 2022) and at specific disease-relevant locations (HBB E6V for sickle cell disease; CFTR R553X & G551D for cystic fibrosis; ATP7B H1069Q for Wilson’s disease).
[0087] Fig. 35 illustrates effects of nCas9 modifications that enhance editing efficiency. Inclusion of chromatin-modifying peptides (CMPs) HN1 and Rad51 DNA binding domain (DBD) in fusions with nCas9 increased efficiencies.
[0088] Fig. 36 illustrates effects of donor DNA methylation and 3’ end protection. Fig. 36A: Donor methylation. Fig. 36B: Donor methylation and 3’ end protection with and without splint LNA modification. Fig. 36C: Donor methylation with O-Me splint modification or LNA splint modification.
[0089] Fig. 37 illustrates effects of splint 3’ end protection on editing efficiency. Fig. 37A: Chemically blocking the 3’ end of splints increases efficiency. The effect is shown for 3’ AltR and a 3’ C3 spacer. Fig. 37B: Splints linked at the 3’ end to biotin were tested with T4 ligase linked to streptavidin. Fig. 37C(SEQ ID Nos:815,558 and 816) depicts the locations and exemplary compositions of the nucleic acid binding sites for a splint with a 20 nt DBS and 19 nt GBS. (Donor: SEQ ID NO:558, / 5Phos / mCgtaTgtmCagggtggtmCamCGAmC*g*g-C3 spacer; Splint: SEQ ID NO:559, +C*C*+GT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGgcgtgcagtgcttACGCCA+CA+AT+A C+CG+CA+G*C*+T-C3 spacer; Guide: 3’ end of guide, e.g. SEQ ID NO: 166, SEQ ID NO: 170, SEQ ID NO: 173, or SEQ ID NO: 222).
[0090] Fig. 38 illustrates effects of nicking guides (Fig. 38A, B) and dead guides (Fig. 38B). Nicking guides nick the opposite strand to promote incorporation of desired edits upon DNA mismatch repair (MMR). Nicking guides can be screened to achieve improved efficiency with minimum indels. Dead guides include a spacer (e.g. a 15-nt spacer) that allows nCas9 to engage a target without nicking, opening chromatin in the region of the target, and are used with additional target strand spacers (Park et. al., Genome Biology 2021).
[0091] Fig. 39 illustrates high efficiency editing of disease-relevant mutations. Replacer optimization included one or more of fusing additional peptides to nCas9; chemically blocking ends of donors and splints; methylating donor DNA; including nicking guides; screening different FBS lengths; removing the splint’s 3’ LNA (HBB only); and changing the gRNA scaffold.
[0092] Fig. 40 illustrates Replacer high efficiency editing and precision compared to prime editing PE2, including without use of a nicking guide. Replacer demonstrates substantially reduced indel formation due from RNA synthesis errors, reverse transcriptase errors, and scaffold integration. A: Edited and modified reads (e.g., indels) as a proportion of all edited reads at five loci. B: Editing efficiency at the five loci. C: Indels by location, Replacer vs. PE. Replacer produces substantially fewer indels over the range of bases.
[0093] Fig. 41 illustrates 1 -sided and 2-sided Replacer deletion of disease-associated repeat regions. Top Left: HTT in-frame 96bp deletion of repeat region by 1 -sided Replacer showing efficiency of deletion and low frequency of indels. Top Right: 2-sided Replacer vs Prime on deletion of 131bp region of C9orf72 and 38bp Bxbl attB replacement. Bottom: Efficient replacement (lanes 1 and 2) or deletion (lane 3) of C9orf72 by 2-sided Replacer.
[0094] Fig. 42 illustrates effects of overlap length (A-C)) and base modification (C) in a 2 x 1-sided Replacer (e.g., as diagrammed in Fig. 25). The direction and length of the attB sequence, as well as the overlap length between the 2 Replacer donors, affects edit efficiency (B). Changing some LNAs in the splint to 2’-0Me bases increases efficiency for longer donor overlaps (e.g. 38 bp donors) (C).
[0095] Fig. 43 illustrates high Replacer efficiency in cells, including in non-dividing cells. Fig. 43A depicts 2-sided Replacer compared to prime editing in human primary hepatocytes (PHH), induced pluripotent stem cells (iPSCs) and HEK293T cells. Replacer efficiency can be optimized by LNP formulation. Fig. 43B illustrates VEGFA 175 bp replacement with a 2 bp “GT” in three cell lines using three different LNP formulations.
[0096] Fig. 44 illustrates Replacer editing efficiency in human hepatocytes. A: ATP7B mutations in PHH primary human hepatocytes and HEK293T immortalized human embryonic kidney cells. B. Replacement of NOLC 79 bp sequence with PaOl attB 33 bp sequence in PHH and HEK293T cells.
[0097] Fig. 45 illustrates effects of ligase and nCas9 order in fusion protein and split expression.
[0098] Fig. 46 depicts an overview of the Replacer system to edit human cells with a stably integrated BFP reporter that converts to GFP upon successful editing. Fig. 46A depicts an example of the Replacer editing mechanism for an A to C transversion edit. The donor DNA is homologous to the genome besides the edited C base. After the donor is ligated onto the 3’ flap, it anneals to the opposite strand of the genome and induces editing through the endogenous mismatch repair pathway. Fig. 46B illustrates a schematic of the Replacer editing workflow: nucleic acids are transfected to HEK293T cells that express a virally integrated BFP gene. Replacer edits 3 nucleotides in the BFP gene to convert it to GFP. Editing efficiency is equal to the percentage of GFP+ cells as assessed with flow cytometry. Fig. 46C illustrates an example architecture of a replacer nucleic acid and chemical modification layout. The splint includes alternating LNAs on in the DBS and GBS and the ligRNA includes three 2’-OMe nucleotides on the 5’ and 3’ ends. The Replacer splint composed of a donor binding site (DBS) that is complementary to the donor DNA, flap binding site (FBS) that hybridizes to the nicked 3’ flap, and guide binding site (GBS) that connects to the ligRNA. The ligRNA contains a typical Cas9 spacer and scaffold as well as a splint binding site (SBS) on the 3’ end that is bound to the splint. The splint and ligRNA may incorporate modifications, not limited to phosphorothioate backbone modifications (PS), locked nucleic acids (LNAs), and methylated nucleotides.
[0099] Fig. 47 displays a comparison of the editing efficiencies of nucleic acids with different number of alternating LNAs in the DBS and GBS regions (Fig. 47A), different FBS length and DNA dose using 2.1 pmol ligRNA for transfection (Fig. 47B), and different donor and DBS length (Fig. 47C). Fig. 47A(SEQ ID Nos:208, 560-563) depicts effects of LNAs incorporated into the DBS and GBS portion of the splint. Splints from top to bottom: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+A C+CG+CA+G*C*+T (LNAs in DBS / GBS - 10 / 7) (SEQ ID NO:208);+C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGC+CA+CA+AT+A C+CG+CA+G*C*+T (LNAs in DBS / GBS - 9 / 8) (SEQ ID NO:560);+C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGCCA+CA+AT+AC +CG+CA+G*C*+T (LNAs in DBS / GBS - 9 / 7) (SEQ ID NO:561);+C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGCCACA+AT+AC+ CG+CA+G*C*+T (LNAs in DBS / GBS - 9 / 6) (SEQ ID NO:562);+C*G*+TG+AC+CA+CC+CT+GA+CATACGGCGTGCAGTGCTTACGC+CA+CA+AT+AC+CG+CA+G*C*+T (LNAs in DBS / GBS - 8 / 8) (SEQ ID NO:563). The right portion of Fig. 52A depicts % GFP fluorescent cells in the transfected cells. Fig. 47B depicts effects of splint FBS length and DNA amount on editing efficiency of the HEK293T system. Fig. 47C depicts the BFP to GFP conversion efficiency based on donor and DBS length.
[0100] Fig. 48 displays the comparative BFP to GFP editing efficiencies using either a 100-nt ssODN donor for HDR or the splint and donor DNA used for Replacer, with different mRNAs. The system optimally comprises splint, donor DNA, Cas9n, T4 ligase, and the 5’phosphate. Fig. 48A illustrates the effects of subtracting system components. 5’ phosphorylation of donor DNA is not critical in the HEK293T system. Also, while HEK293T cells display endogenous ligase activity, the level is substantially increased by T4 ligase. Fig. 48B illustrates comparative efficiency of the Replacer system using Cas9 (Cas9 nuclease), Cas9n (H840A Cas9 nickase) or Cas9n + T4 DNA ligase relative to a 100 nt single-stranded oligodeoxynucleotide (ssODN) donor template with Cas9. Fig. 48C illustrates comparative efficiency of separate Replacer mRNAs on BFP to GFP conversion, using pseudouridine (ml ) in either an nCas9-T4 fusion or dual (nCas9 and T4) mRNA system.
[0101] Fig. 49 (SEQ ID Nos: 564,817 and 818)illustrates replacer editing of point mutations at endogenous genomic loci in human HEK293T cell line (Fig. 49A) using splints with alternating locked nucleic acids (LNAs) in the GBS and DBS regions (Fig. 49B). An example chemical modification layout of an efficient system: Splint:+A*G*+GC+CA+GC+AG+TG+AA+CA+AC+CA+TT+GGGCGTGGCAGTACGCCA+CA+A T+AC+CG+CA+G*C*+T-c3 SPACER (SEQ ID NO:564); Donor DNA: Z5Phos / / IME- DC / AATGGTTGTT / IME-DC / A / IME-DC / TG / IME-DC / TGG / IME-DC / / IME-DC / *T-c3 SPACER (SEQ ID NO:565); SBS portion of ligRNA: AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO: 566).
[0102] Fig. 50 displays the editing efficiencies of splints with alternating locked nucleic acids (LNAs) in the GBS and DBS regions in HBB (Fig. 50A), in CFTR (Fig. 50B), and in ATP7B (Fig. 50C).
[0103] Fig. 51 illustrates the effect of adding additional nicking guides on the BFP to GFP editing efficiencies and indels made using the splits in Fig. 50 in HBB (Fig. 51A), in CFTR (Fig. 51B), and in ATP7B (Fig. 51C). Location represents the direction and number of nucelotides between the 2 nicks
[0104] Fig. 52 illustrates the editing efficiencies and indels at different doses of splint and donor DNA as compared to ligRNA for different FBS lengths in HBB (Fig. 52A), in CFTR (Fig. 52B), and ATP7B (Fig. 52D), and in AAVS1 (Fig. 52D). Every nCas9 mRNA is used with a T4 ligase mRNA.
[0105] Fig. 53 illustrates the editing efficiencies of splints with or without a C3 spacer on the 3’ end of the splint in HBB (Fig. 53A), in CFTR (Fig. 53B), and in ATP7B (Fig. 53C).
[0106] Fig. 54 illustrates the editing efficiencies of splints with several DNA modifications: with or without a C3 spacer on the 3 ’ end, phosphorothioate (PS) bonds and a methylated cytosine (meC) in HBB (Fig. 54A), in CFTR (Fig. 54B), and in ATP7B (Fig. 54C).
[0107] Fig. 55 illustrates a schematic of a 1 sided Replacer editing system (Fig. 55A). In a 1 sided Replacer editing system, the splint links together ligase-mediated guide RNA (ImgRNA) and Donor DNA, nicks Cas9 complexes with ImgRNA and nicks the genomic DNA to open the R- loop, making the opposite stand’s flap accessible. The splint then binds to the genomic flap and positions the Donor at the nick, the Ligase installs the Donor to the genomic DNA flap, the L-PGI components dissociate, leaving only the donor DNA as a permanent edit. Fig 60B illustrates a 2 sided Replacer system editing, where there are 2 full Replacer editing complexes targeting protospacers on opposite strands of DNA. Fig. 55C illustrates an example outline of an editing process: nucleic acids transfections include 2 sets of splint, donor DNA, and ligRNA. 72 hours after transfection, genomic DNA is extracted for PCR amplification and amplicons can be visualized on a gel or sequenced to assess editing efficiency.
[0108] Fig. 56 illustrates design strategies for a 2-sided Replacer, where donor DNAs can be homologous to each other to make a replacement edit or homologous to the genome to make a deletion edit: a replacement with full overlap where the edited flaps bind each other completely (Fig. 56A), partial overlap where the 3’ segments of the edited flaps bind each other (Fig 61B), and a deletion (Fig. 56C).
[0109] Fig. 57 illustrates 1-sided and 2-sided Replacer edit of disease-associated repeat regions. Fig. 57A illustrates an agarose gel image of PCR amplicons for C9ORF72 after Replacer deletion of 222bp and 13 Ibp regions of C9orf72. The WT fragment is non-edited and the fragments on the bottom are the desired replacement or deletion. attB and attB rev comp are replacements of the nick-nick region with the Bxbl attB site in forward or reverse orientation, del F + R is a deletion, del F is the deletion without the Rev splint and donor, and del R is the deletion withoutthe Fwd splint and donor Fig. 57B illustrates an agarose gel image of PCR amplicons for the Bxb 1 attB replacement edit of a 13 Ibp region of C9ORF72, comparing donor DNAs with different chemical modifications. The fragment of the desired edit is on the bottom.
[0110] Fig. 58 illustrates the editing efficiencies for attB replacement edits at 2 targets as quantified with NGS of amplicons. T4 ligase mRNA is used alongside the nCas9 mRNAs that are shown. Fig. 58A shows the replacement of NOLC1 79 bp sequence with PaOl attB 33 bp sequence, Fig. 58B shows the VEGFA replacement of 38 bp Bxbl attB site, and Fig. 58C (SEQ ID Nos: 567-576)illustrates the effect on editing efficiencies of of splints that have different overlap length between forward and reverse splints and variation in the number and chemical modifications (LNAs and 2’-OMe) in 2-sided Replacer system showing precise editing vs. indel production, 10 being partial overlap and 38 being full overlap. Each splint is used with a donor DNA of the same length as its DBS, and all splints include the same FBS and GBS that are not shown. Efficiencies are assessed with NGS of amplicons. Top to bottom: 10 bp splint overlap with LNAs (12 of 24), forward splint DBS = +G*G*+AG+AC+CG+CC+GT+CG+TC+GA+CA+AG+CC (SEQ ID NO:567), Reverse splint DBS = +G*G*+CG+GT+CT+CC+GT+CG+TC+AG+GA+TC+AT (SEQ ID NO:568); 10 bp splint overlap with LNAs (7 of 24) and 2’0-Me nucleotides (5 of 24), Forward splint DBS = +G*G*+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+CC (SEQ ID NO: 569), Reverse splint DBS = +G*G*+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+AT (SEQ ID NO:570); 38 bp splint overlap with LNAs (19 of 38), Forward splint DBS = +A*T*+GA+TC+CT+GA+CG+AC+GG+AG+AC+CG+CC+GT+CG+TC+GA+CA+AG+CC (SEQ ID NO:571), Reverse splint DBS+G*G*+CT+TG+TC+GA+CG+AC+GG+CG+GT+CT+CC+GT+CG+TC+AG+GA+TC+AT (SEQ ID NO:572); 38 bp splint overlap with LNAs (10 of 38) and 2’-OMe nucleotides (9 of 38), Forward splint DBS: =+A*T*mGA+TCmCT+GAmCG+ACmGG+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+ CC (SEQ ID NO:573), Reverse splint DBS +G*G*mCT+TGmUC+GAmCG+ACmGG+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+A T (SEQ ID NO:574); 38 bp splint overlap with LNAs (13 of 38) and 2’-OMe nucleotides (6 of 38), Forward splint DBS: =+A*T*+GA+TC+CT+GA+CG+ACmGG+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+CC (SEQ ID NO: 575), Reverse splint DBS:+G*G*+CT+TG+TC+GA+CG+ACmGG+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+AT (SEQ ID NO: 576).
[0111] Fig. 59 illustrates replacer editing in primary human hepatocytes. Comparison of Replacer and the PE2 or PEMax prime editing systems in HEK293T cells and primary human hepatocytes (PHH) for the point mutations in ATP7B (Fig. 59A) and for the replacement of NOLC1 79 bp sequence with PaOl attB 33 bp sequence (Fig. 59B). The PBS and RTT sequences for prime editing match Replacer’s splint FBS and DBS sequences, respectively. Editing efficiency and indel generation is calculated by analyzing NGS of amplicons
[0112] Fig. 60 illustrates optimization of chemical modifications in the splint and donor DNA for Replacer editing at BFP. Fig. 60A (SEQ ID Nos:577-583) depicts variation in the number of LNAs in a 21nt DBS region of a splint. The control splint has alternating LNAs in the DBS, which is typically optimal for a 21 -nt DBS. Top to bottom: +T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT+ AC+CG+CA+G*C*+T (SEQ ID NO:577);+T*+C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:578);+T*+C*+G+T+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:579 );+T*C*+GT+GA+CC+AC+CC+TG+AC+AT+A+C+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO: 580);+T*C*+GT+GA+CC+AC+CC+TG+AC+A+T+A+C+GGCGTGCAGTGCTTACGCCA+CA+A T+AC+CG+CA+G*C*+T (SEQ ID NO:581);+T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+G+GCGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (SEQ ID NO:582);+T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GG+CGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (SEQ ID NO:583). Fig. 60B (SEQ ID Nos:208, 584-588) depicts a comparison of different chemical modifications (LNA, 2’-0Me, 2F) in splint DBA and GBS regions. Top to bottom:+C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+AC +CG+CA+G*C*+T (LNA / LNA) (SEQ ID NO:208);+C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA / i2FC / A / i2FA / T / i2FA / C / i2FC / G / i2FC / A / i2FG / *C* / 32FU / (LNA / 2’-F) (SEQ ID NO:584); / 52FC / *G* / i2FU / G / i2FA / C / i2FC / A / i2FC / C / i2FC / T / i2FG / A / i2FC / A / i2FU / A / i2FC / GGCGTGCA GTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (2’-F / LNA) (SEQ ID NO:585);+C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCAmCAmATmA CmCGmCAmG*C*mT (LNA / 2’-0Me) (SEQ ID NO:586); mC*G*mTGmACmCAmCCmCTmGAmCAmTAmCGGCGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (2’-0Me / LNA) (SEQ ID NO: 587);C*G*TGACCACCCTGACATACGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G* C*+T (DNA / LNA) (SEQ ID NO:588). Fig. 60C(SEQ ID Nos: 589, 590 and 591) illustrates the effects of varying the number and the location of LNAs in a 32nt DBS of a splint. Top to bottom: 16 LNAs of 32(+C*T*+GG+CC+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTG CTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:589)); 12 LNAs of 32, reduced near nick(+C*T*+GG+CC+CA+CC+CTC+GTG+ACC+ACC+CTG+ACA+TA+CGGCGTGCAGTGCT TACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:590)); 12 LNAs of 32, reduced near 5’ end(+C*T*+GGC+CCA+CCC+TCG+TGA+CCA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCT TACGCCA+CA+AT+AC+CG+CAG*C*T (SEQ ID NO:591)). Fig. 60D illustrates the effect of methylated DNA donors on BFP to GFP conversion efficiency, dependent on chemical modifications in the splint DBS.
[0113] Fig. 61A(SEQ ID Nos:592-594) illustrates a comparison of different pairs of splint GBS and ligRNA SBS sequences with 3’ end modifications of ligRNA. SBS folding AG is the strength of secondary structure in the ligRNA’ s 20-nt 3’ extension (SBS). All splints include the same DBS and FBS (not shown) and ligRNAs include identical scaffolds and spacers (not shown). Top to bottom: AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO:592); GUGGUUCCGGGCUGCAmU*mG*mA (SEQ ID NO: 593);CGAUUCCUGAUACUGCmU*mG*Mc (SEQ ID NO: 594).
[0114] Fig. 61B (SEQ ID Nos: 595, 818, 597, 822, 595, 822, 600, 822, 601, 823, 601, 824, 601, 825, 600, 825, 600, 826, 595 and 818) illustrates Replacer optimization, including splintoptimization by GBS / SBS length, splint composition, and FBS linker, and ligRNA optimization. Nucleotide sequences are shown as they are bound to each other and in some cases a part of the GBS or SBS functions as a single stranded linker. All splints include the same DBS and FBS (not shown) and ligRNAs include identical scaffolds and spacers (not shown). Top to bottom. Length of GBS / SBS (nt) = 24 / 19: ACGTCACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:595) and AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO:596); Length of GBS / SBS (nt) = 15 / 15: +CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:597) andAGCUGCGGUAUUmG*mU*mG (SEQ ID NO:598); Length of GBS / SBS (nt) = 19 / 15: ACGC+CA+CA+AT+AC+CG+CA+G*C*T (SEQ ID NO:599) and AGCUGCGGUAUUmG*mU*mG (SEQ ID NO:598); Length of GBS / SBS (nt) = 20 / 15; GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:600) and AGCUGCGGUAUUmG*mU*mG (SEQ ID NO:598); Length of GBS / SBS (nt) = 28 / 40: GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (SEQ ID NO:601) and CGATTTCCTGATAGTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:602); Length of GBS / SBS (nt) = 28 / 32:GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (SEQ ID NO:601) and GATAGTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:603); Length of GBS / SBS (nt) = 28 / 28; GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (SEQ ID NO:601) and GTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:604); Length of GBS / SBS (nt) = 20 / 28: GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NQ:600) and GTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:604); Length of GBS / SBS (nt) = 20 / 20: GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:600) and AGCUGCGGUAUUGUGGCmG*mU*mC (SEQ ID NO:605); Length of GBS / SBS (nt) = 19 / 19: ACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:595) and AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO:606).
[0115] Fig. 61C(SEQ ID Nos:607, 810, 609, 620, 609, 819, 607, 821, 607, 820) illustrates Replacer optimization, including splint optimization of DBS length and DBS / donor overhang. When the lengths of the splint DBS and donor DNA are not equal, there is a ssDNA overhang. All splints include the same DBS and FBS (not shown). Top to bottom. Length of DBS / donor (nt) = 20 / 28 nt: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:607) and / 5Phos / CGTATGTCAGGGTGGTCACGAGGGTGGG (SEQ ID NO 608); Length of DBS / donor= 28 / 20 nt: +C*C*+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:609) and / 5Phos / CGTATGTCAGGGTGGTCACG (SEQ ID NO:610); Length of DBS / donor = 28 / 28 nt: +C*C*+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:609) and / 5Phos / CGTATGTCAGGGTGGTCACGAGGGTGGG (SEQ ID NO:608); Length of DBS / donor = 20 / 21 nt: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:607) and / 5Phos / CGTATGTCAGGGTGGTCACGA (SEQ ID NO:611); Length of DBS / donor = 20 / 20 nt +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:607) and / 5Phos / CGTATGTCAGGGTGGTCACG (SEQ ID NO:610).
[0116] Fig. 62 depicts the evaluation of ligases and protein architecture for Replacer editing at BFP. Fig. 62A depicts the comparison of 3 different ligases as either fusion proteins with nCas9 encoded in a single mRNA, or as separate mRNAs used with nCas9 mRNA. When 2 mRNAs are used, nCas9 and the ligase are fused to leucine zippers (LZs) on either the N-terminal or C-terminal to promote colocalization. A 24-nt donor DNA & splint DBS were used to convert BFP to GFP. Fig. 62B depicts the comparison of the 3 ligases to no ligase, a T4-nCas9 fusion, and T4-P2A- nCas9 bicistronic mRNA for conversion of BFP to GFP with a 50-nt donor & DBS. Fig. 62B depicts the effect of LZs on BFP to GFP conversion efficiency. Each ligase mRNA is delivered with nCas9 mRNA without LZs on either protein (left bar), with LZs on the ligase C-terminus and nCas9 N-terminus (middle bar), or with LZs on the ligase N-terminus and nCas9 C-terminus (right bar).
[0117] Fig. 63 illustrates deletion, replacement and conversion with L-PGI in HEK293T cells (Figs. 63A-C) of a 2-sided deletion (Fig. 63), a 2-sided beacon replacement (Fig. 63 B), and a 1- sided correction / conversion (Fig. 63 C).DETAILED DESCRIPTION OF THE INVENTIONIntroduction
[0118] Advances in genome editing tools have enabled precision editing of genomes for therapeutic, agricultural, industrial, and research purposes. Some editing tools may include integration of nucleic acid sequences into the genome at a targeted location. However, the length of integrating sequences that can be used may be constrained in some editing tools.
[0119] Some nuclease-based tools such as CRISPR-Cas9 use a guide RNA to target the Cas9 protein to a specific DNA sequence specified by the spacer sequence in the guide RNA. Cas9nuclease activity then cleaves the DNA resulting in a double-stranded break (DSB). DSBs are typically repaired through endogenous DNA repair mechanisms including non-homologous end joining (NHEJ) or homology-directed repair (HDR). However, NHEJ results in a spectrum of nucleotide insertions and deletions (indels) that hinder its utility for precision editing. HDR efficiency is very low in nondividing cells and may require DNA replication. Even when HDR editing is detectable, DSB-induced indels are often prevalent, meaning that HDR may not be feasible when precision editing is desired.
[0120] Homology-independent targeted insertion (HITI) utilizes NHEJ DNA repair mechanisms active in nondividing cells for CRISPR-guided transgene integration in nondividing cells such as primary neurons, retinal pigment epithelial cells, and HSPCs. However, due to the generation of DSBs from Cas9, HITI generates high frequencies of indels, resulting in unintended mutations in addition to DSB associated toxicity.
[0121] Other methods for gene editing have additional limitations. Tools employing fusions of nicking Cas nucleases with nucleotide deaminases (e.g., base editors) can perform certain nucleotide mutations, e.g., cytosine base editors can convert C to T. While some base editors can perform precision editing at high efficiency, they are inherently limited to specific edits determined by the deaminase variant so they are only applicable to specific substitution mutations and further cannot perform precise insertion or deletion edits. Moreover, base editors are generally limited to a small editing window within a subset of the protospacer region and are therefore significantly limited by protospacer adjacent motif (PAM) availability. Finally, base editors can exhibit bystander mutations within the editing region (e.g., if two C’s are present) and have demonstrated DNA and RNA off-target deaminase activity.
[0122] Existing precision editing technologies have limitations that hamper their practical applicability in a variety of ways. In particular, they may rely on endogenous cellular machinery for editing, for example HDR machinery for nuclease-based editing and mismatch repair for base editing. No system has been reported that is independent of all endogenous factors. Reliance on endogenous factors is problematic because different cell types have different activity levels of these endogenous factors, and in many cases the activity is not sufficient to provide useful levels of editing. An example where this reliance is particularly problematic is nondividing cells, which comprise the majority of cells in adults and therefore are not amenable to many existing precision editing tools.
[0123] Accordingly, there remains a need for a system or a method for effective gene editing (e.g., revising) or for modifying gene expression by gene editing. Particularly, there remains a need for the system or method that minimally constrains the length of the revising nucleotide sequence. Additionally, there remains a need for the system or method for gene editing or modifying gene expression, where the system or the method do not rely on the endogenous components or mechanism of a cell. There also remains a need for a system or a method for correcting genetic mutations in a cell. In some cases, the correction of genetic mutation can treat a disease or condition in subject in need thereof. As will be seen below, the systems, methods, and compositions disclosed herein may be useful for addressing these needs or limitations.Overview
[0124] Described herein are self-contained gene editing systems. In some such self-contained systems, every aspect of gene editing may be controlled. Some such systems do not rely on host cell machinery to perform an editing function, or to replace or repair any aspect of a target nucleic acid such as a genomic locus. Some such systems are unaffected by a cell’s nucleotide triphosphate (dNTP) concentration because the editing may be performed without use of a polymerase. For example, an exogenous first integrating nucleic acid may be delivered and inserted into a genetic locus without transcribing a template. The editing may exclude a need to rely on a cell repair system such as HDR or NHEJ. The editing may be performed without cell cycling. The gene editing may take place in a cell or may even be performed in vitro. For example, the gene editing may even be performed in a test tube or outside of a cell.
[0125] Described herein are systems and methods for editing DNA with a donor strand without generating a double-stranded break in the genome using CRISPR-guided DNA ligases and guide nucleic acids targeting the genomic region of interest. DNA ligases are enzymes which chemically join two DNA molecules via a phosphodiester bond. DNA ligases may or may not require hybridization of the DNA molecules to a DNA or RNA backbone or “splint” which is reverse complementary to the DNA sequences that are to be ligated. Targeting of ligases to genomic nicks generated by CRISPR nucleases enables precise replacement of genomic DNA with donor strands optionally recruited by guide nucleic acids into targeted loci. The CRISPR-guided DNA ligases can be composed of DNA ligases that are fused, recruited, or unfused to the RNA-guided endonuclease by utilizing peptide linkers, heterodimerization domains, or two separate peptides, respectively.
[0126] Some aspects include a cell containing or comprising an RNA-guided endonuclease and a DNA ligase, both of which are introduced into the cell. The endonuclease or ligase may be heterologous to the cell. The endonuclease and ligase may be heterologous to the cell. The ligase may be endogenous to the cell. In some aspects, a cell comprises an RNA-guided endonuclease and a DNA ligase, both of which are heterologous to the cell. The cell may include a composition or system described herein. The cell may be used or included in a system, composition, or method described herein.
[0127] A system described herein may include a heterologous endonuclease comprising an RNA-guided endonuclease such as nicking Cas9 as well as a heterologous ligase (e.g., a DNA ligase) that can utilize an RNA splint. The guide nucleic acid optionally recruits a donor strand to the site targeted by the endonuclease (e.g., a targeted genomic locus) and also generates a splint across from the donor strand (donor strand) and genomic flap generated by the nicking Cas9, resulting in ligation of the donor strand and the genomic flap by the DNA ligase. In some embodiments, the ligase is or comprises an endogenous ligase. The system can utilize one or more guide nucleic acids that together can comprise the following components, optionally in the following order: 5 ’ spacer - scaffold - donor binding site (optional) - flap binding site 3 ’ . The donor strand (donor strand) can comprise the following sequence components: 5’ guide binding site - donor strand 3’. The guide binding site of the donor strand is at least partially reverse complementary to the donor binding site of the guide nucleic acid such that the donor hybridizes to the guide and is localized to the target site of the RNA guided endonuclease. The 5’ end of the donor sequence and the 3’ end of the genomic flap generated by nuclease nicking activity are ligated by the DNA ligase, splinted by the donor binding site and a flap binding site of the guide nucleic acid(s).
[0128] Fig. 1A-1C illustrate a non-limiting example of a system (1-sided Replacer 1). The example includes a guide nucleic acid comprising: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus. The guide nucleic acid is shown complexed with an endonuclease (e.g., a Cas9 nickase, nCas9) operatively coupled to a ligase. The guide nucleic acid may direct the endonuclease to a genomic locus that is bound by the spacer of the guide nucleic acid. The guide nucleic acid is also shown as partially complementary to a donor strand (complexing between thedonor binding site of the guide nucleic acid and guide binding site of the donor strand). The endonuclease, when directed by the guide nucleic acid, can cleave or nick at least one strand of the genomic locus, and the ligase can ligate one end of the donor strand with the cleaved or nicked end of the genomic locus, thus incorporating the donor strand into the genomic locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by a nuclease.
[0129] Fig. 2A-2C illustrate a non-limiting example of a system (2-sided Replacer 1). The guide nucleic acid in the example, similar to the guide nucleic acid of Fig 1A, comprises: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus. In Fig. 2A, a first guide nucleic acid is shown complexed with a first endonuclease operatively coupled with a first ligase and a second guide nucleic acid is complexed with a second endonuclease operatively coupled with a second ligase. The first endonuclease and the second nuclease may each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus can then be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. The insertion of the donor strand at the genomic locus may generate two genomic flaps that can be digested and removed by a nuclease.
[0130] Fig. 3A-3C illustrate a non-limiting example of a system (1-sided Replacer 2). In the example, a guide nucleic acid comprises: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; and a donor binding site for complexing with a donor strand. Also shown in Fig. 3A is a donor strand comprising at least one overhang, where the overhang comprises: a flap binding site for complexing with a genomic flap of the genomic locus; and a guide binding site for complexing with the guide nucleic acid (via the donor binding site of the guide nucleic acid). The guide nucleic acid can be complexed with an endonuclease (e.g., nCas9) operatively coupled to a ligase. The guide nucleic acid in the example directs the endonuclease and the ligase to a genomic locus that is bound by the spacer of the guide nucleic acid. The guide nucleic acid in the example is also partially complementary to a donor strand (complexing between the donor binding site of the guide nucleic acid and guide binding site of the donor strand). The endonuclease, when directed by the guide nucleic acid, can cleave at least one strand of the genomic locus, and the ligase can ligate one end of the donor strand withthe cleaved end of the genomic locus, thus incorporating the donor strand into the genomic locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by a nuclease.
[0131] Fig. 4A-4C illustrates a non-limiting example of a system (2-sided Replacer 2). In the example, where the guide nucleic acid, similar to the guide nucleic acid of Fig 3A, comprises a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; and a donor binding site for complexing with a donor strand. Also shown in Fig. 4A is a donor strand comprising two overhangs, where the overhangs each comprise a flap binding site for complexing with a genomic flap of the genomic locus; and a guide binding site for complexing with a guide nucleic acid (via a donor binding site of the guide nucleic acid). The flap binding site of the donor strand can bring the donor strand in close proximity with the genomic locus after a genomic flap is generated after the endonuclease cleaves at least one strand of the genomic locus. In Fig. 4A, a first guide nucleic acid is shown complexed with a first endonuclease operatively coupled with a first ligase and a second guide nucleic acid is complexed with a second endonuclease operatively coupled with a second ligase. In the example, the first endonuclease and the second nuclease each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus can then be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. In the example, the insertion of the donor strand at the genomic locus generates two genomic flaps that can be digested and removed by a nuclease.
[0132] A system described herein (Replacer 3) may include a heterologous endonuclease comprising an RNA-guided endonuclease such as nicking Cas9 as well as a ligase (e.g., a DNA ligase) that can utilize a DNA splint. The guide nucleic acid optionally recruits a donor strand to the site targeted by the endonuclease (e.g., a targeted genomic locus) and also generates a splint across from the donor strand (donor strand) and genomic flap generated by the nicking Cas9, resulting in ligation of the donor strand and the genomic flap by the DNA ligase. At least part of the flap binding site and donor binding site on the guide nucleic acid are DNA such that ligases that utilize DNA splints are able to catalyze the intended reaction. The system can utilize one or more guide nucleic acids that together can comprise the following components, optionally in the following order: 5’ spacer - scaffold - donor binding site (optional) - flap binding site 3’. The donor strand (donor strand) can comprise the following sequence components: 5’ guide binding site - donor strand 3’. The guide binding site of the donor strand is at least partially reversecomplementary to the donor binding site of the guide nucleic acid such that the donor hybridizes to the guide and is localized to the target site of the RNA guided endonuclease. The 5’ end of the donor sequence and the 3’ end of the genomic flap generated by nuclease nicking activity are ligated by the DNA ligase, splinted by the donor binding site and a flap binding site of the guide nucleic acid(s).
[0133] Fig. 5A-5C illustrate a non-limiting example of a system (1-sided Replacer 3). The example includes a guide nucleic acid comprising: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus, wherein at least part of the flap binding site and donor binding site are comprised of DNA. The guide nucleic acid is shown complexed with an endonuclease (e.g., a Cas9 nickase, nCas9) operatively coupled to a ligase (e.g., an endogenous ligase or an exogenous ligase). The guide nucleic acid may direct the endonuclease to a genomic locus that is bound by the spacer of the guide nucleic acid. The guide nucleic acid is also shown as partially complementary to a donor strand (complexing between the donor binding site of the guide nucleic acid and guide binding site of the donor strand). The endonuclease, when directed by the guide nucleic acid, can cleave at least one strand of the genomic locus, and the ligase can ligate one end of the donor strand with the cleaved end of the genomic locus, thus incorporating the donor strand into the genomic locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by a nuclease.
[0134] Fig. 6A-6C illustrate a non-limiting example of a system (2-sided Replacer 3). The guide nucleic acid in the example, similar to the guide nucleic acid of Fig. 5A, comprises: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus, wherein at least part of the flap binding site and donor binding site are comprised of DNA. In Fig. 6A, a first guide nucleic acid is shown complexed with a first endonuclease operatively coupled with a first ligase and a second guide nucleic acid is complexed with a second endonuclease operatively coupled with a second ligase. The first endonuclease and the second nuclease may each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus can then be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. The insertion of the donorstrand at the genomic locus may generate two genomic flaps that can be digested and removed by a nuclease.
[0135] Ligation may be performed using a DNA ligase that can utilize an RNA splint such as SplintR ligase - also known as PBCV-1 DNA Ligase - from Chlorella virus. In some aspects, the system utilizes two guide nucleic acids targeting the CRISPR-guided ligase to target sites on opposite strands flanking the genomic region of interest. In some aspects, each guide nucleic acid interacts with a corresponding donor strand in the manner described above, resulting in ligation of both donor strands which are reverse complementary with each other in the donor strand regions.
[0136] A ligase that is fused or recruited to an endonuclease, or supplied in trans, can utilize DNA as a splint, and a donor strand acts as the splint for the genomic flap generated by the endonuclease and another donor strand. In some aspects, the donor strand comprises: 5’ donor strand - flap binding site - guide binding site (optional) 3’. The flap binding site on one donor strand (Donor2) can be reverse complementary to the genomic flap, while the optional guide binding site on Donor2 is reverse complementary to the optional donor binding site of a guide nucleic acid (Guide 1), and the donor strand can be at least partially reverse complementary to a different donor strand (Donorl). The 5’ end of this Donorl and the 3’ end of the genomic flap can be ligated using the flap binding site and donor strand of the Donor2 as a splint. Such 2-sided approach utilizing dual guide nucleic acids with different spacer sequences can be adopted with Donor2, which provides the splint at the first genomic site and can be ligated on its 5’ end to a 3’ end of a different genomic flap at a nick created using a second Replacer2 guide nucleic acid (Guide2) with a spacer sequence that targets a second site. The donor binding site on the second guide nucleic acid system can optionally recruit Donorl via hybridization with its optional guide binding site, and the Donorl acts as the DNA splint for ligation of Donor2 to the 3’ end of the genomic flap at the target site of the second guide nucleic acid.
[0137] Following ligation, the remaining flaps of native genomic DNA can be excised via exogenously delivered or endogenous flap endonucleases or exonucleases. Examples of exogenous nucleases that can be introduced into the cell include human flap endonuclease 1 (hFENl), human exonuclease 5 (hEXO5), T5 exonuclease, T7 exonuclease, exonuclease VIII, the flap endonuclease domain of E. coli Poll, RecJF, Lambda exonuclease, Xni (ExoIXI) from Escherichia coli, SaFEN (Staphylococcus aureus FEN), nuclease BAL-31, or fragments thereof. The endonucleases or exonucleases can optionally be fused, recruited, or unfused to the RNA-guided endonuclease or DNA ligase by utilizing peptide linkers, heterodimerization domains, or two separate peptides, respectively.
[0138] In some aspects, the system, composition, or method described herein utilizes additional protein that binds to the cleaved or nicked site. For example, the system, composition, or method described herein can include Ku protein or Gam protein from bacteriophage Mu, where the binding of the Ku protein or Gam protein can increase ligation efficiency of the integration nucleic acid at the cleaved or nicked site.
[0139] A system or method described herein may use a nicking endonuclease and, therefore, does not generate double stranded breaks. Furthermore, the system described herein addresses the issue of poor editing efficiencies in nondividing cells through a mechanism of action which only depends on the exogenous components delivered to the cells using mRNA, viral vectors, guide nucleic acids, DNA, or peptides, or any other modalities. Therefore, the system does not require the presence of cell cycle-dependent endogenous cell processes or components such as HDR or dNTPS. As such, the system described herein allows efficiency that is not hindered in nondividing cells. Furthermore, the system enables replacement of both strands of a targeted region of the genome, which can increase editing efficiency.
[0140] A donor strand may contain a high degree of homology with the replaced genomic DNA. These donors may contain mutations to the genomic DNA such as pathogenic mutation correction, disabling of CRISPR protospacer adjacent motif (PAM) sites, disruption of the guide’s spacer sequences, other substitution mutations, or a combination thereof. Additional substitution mutations may be included to increase donor-donor homology versus donor-genome homology to promote hybridization of donor strands and incorporation into the genome. Donor strands may also encode deletions or insertions of nucleotides or may encode a complex combination of the above which then replaces the target genomic DNA. Optionally, guide and donor strands may be chemically modified using nucleic acid chemistries such as phosphorothioate bonds or 2’-O- methylation. Optionally, guide nucleic acids may include hairpin sequences. Optionally, any combination of guide nucleic acids, donor strands, and proteins can be complexed, using an annealing reaction (gradual reduction in temperature) for example, prior to delivering the editing components to the cell.
[0141] Protein components (e.g., nicking Cas9, ligase) may be modified using nuclear localization signals, cell penetrating peptides, or chromatin disrupting peptides in order to improve delivery efficiency to genomic targets.
[0142] The predominant cellular DNA repair pathway for resolving small (< 13nt) mismatches between genomic DNA strands is mismatch repair (MMR). For single stranded donor ligation, the ligated donor strand forms a DNA heteroduplex with the reverse complementary genomic DNA strand. This may also occur with competitive hybridization between ligated donor strand strands and genomic DNA strands. In these cases, MMR activity can excise and revert mismatches in the donor strand using the genomic strand as a template, resulting in reduced editing. Expression of dominant negative versions of MMR proteins has been shown to inhibit the MMR pathway and improve editing outcome in cases where similar DNA heteroduplexes are generated. In some aspects, dominant negative MMR peptides such as MSH2 (G674A) and MLH1 (del754-756) may be delivered as part of the system described herein to improve genomic editing capability, particularly in cells which overexpress the MMR pathway. In some aspects, these dominant negative MMR peptides can be delivered as a fusion (e.g., fused with any component of the system described herein), recruited, or as separate peptides.Endonucleases
[0143] Disclosed herein are endonucleases. The endonuclease may be included in a composition, system or method disclosed herein. The endonuclease may be recombinant. The endonuclease may be coupled to a ligase. The endonuclease may be coupled directly or indirectly to the ligase. The coupling may be covalent or non-covalent. The endonuclease may be bound or connected to a ligase. The endonuclease may be recruited to, be part of a fusion protein with, or be used in conjunction with the ligase. The endonuclease may be coupled to an integrase. The endonuclease may be coupled directly or indirectly to the integrase. The coupling may be covalent or non-covalent. The endonuclease may be bound or connected to an integrase. The endonuclease may be recruited to, be part of a fusion protein with, or be used in conjunction with the integrase. The endonuclease may be heterologous. Heterologous may indicate a source from without a cell. Where a heterologous endonuclease is described, a non-heterologous (e.g., endogenous) endonuclease may be used in some instances. The endonuclease may be encoded in a cell. The endonuclease may be delivered to the cell in trans. The endonuclease may catalyze cleavage of a phosphate bond within an exogenous first integrating nucleic acid. The endonuclease may beguided by a guide nucleic acid to cleave or nick a target nucleic acid for ligation of an exogenous first integrating nucleic acid at the cleavage or nick site. The endonuclease may include any aspect included in Fig. 1A-6C.
[0144] The endonuclease may be non-naturally occurring. The endonuclease may be engineered. The endonuclease may be synthetic. The endonuclease may be pre-synthetized. The endonuclease may be added to a subject or a cell. The endonuclease may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0145] At least part of the endonuclease may be included in a first polypeptide. At least part of the endonuclease may be included in a second polypeptide. The endonuclease may be split into two or more polypeptides bound together. The first polypeptide may include an N-terminal portion of the endonuclease. The first polypeptide may include a C-terminal portion of the endonuclease. The second polypeptide may include the N-terminal portion of the endonuclease. The second polypeptide may include the C-terminal portion of the endonuclease. The first or second polypeptide comprising a part of the endonuclease may be fused with at least part, or the whole, of the ligase. The first or second polypeptide comprising a part of the endonuclease may be fused with at least part, or the whole, of the integrase.
[0146] Described herein, in some aspects, is a system comprising at least one endonuclease. In some aspects, the endonuclease is a programmable endonuclease, where the endonuclease can be complexed with and directed by a guide nucleic acid described herein to a genomic locus. The endonuclease may bind DNA. In some aspects, the endonuclease is an RNA-guided endonuclease. In some aspects, the endonuclease can introduce a single-stranded break. Examples of RNA- guided endonucleases can include CRISPR / Cas endonucleases (e.g., class 2 CRISPR / Cas endonucleases such as a type II, type V, or type VI CRISPR / Cas endonucleases). A CRISPR / Cas endonuclease is also referred to as a CRISPR / Cas effector polypeptide. A suitable endonuclease is a CRISPR / Cas endonuclease (e.g., a class 2 CRISPR / Cas endonuclease such as a type II, type V, or type VI CRISPR / Cas endonuclease). In some cases, a suitable RNA-guided endonuclease is a class 2 CRISPR / Cas endonuclease. In some cases, a suitable RNA-guided endonuclease is a class 2 type II CRISPR / Cas endonuclease (e.g., a Cas9 protein). In some cases, an endonuclease includes a class 2 type V CRISPR / Cas endonuclease (e.g., a Cpfl protein, a C2cl protein, or a C2c3 protein). In some cases, a suitable RNA-guided endonuclease is a class 2 type VI CRISPR / Cas endonuclease (e.g., a C2c2 protein; also referred to as a “Cast 3a” protein). Also suitable for useis a CasX protein. Also suitable for use is a CasY protein. In some aspects, the endonuclease can include any one of the Cas described herein complexed with a guide nucleic acid (e.g., a gRNA) as an RNP complex.
[0147] In some cases, the endonuclease is a Type II CRISPR / Cas endonuclease. In some cases, the endonuclease is a Cas9. Cas9 functions as an RNA-guided endonuclease that uses a dual -guide RNA having a crRNA and trans-activating crRNA (tracrRNA) for target recognition and cleavage by a mechanism involving two nuclease active sites in Cas9 that together generate double-stranded DNA breaks (DSBs) or can individually generate single-stranded DNA breaks (SSBs). The Type II CRISPR endonuclease Cas9 and engineered dual- (dgRNA) or single guide RNA (sgRNA) form a ribonucleoprotein (RNP) complex that can be targeted to a desired DNA sequence. Guided by a dual-RNA complex or a chimeric single-guide RNA, Cas9 generates site-specific DSBs or SSBs within double-stranded DNA (dsDNA) target nucleic acids, which are repaired either by non- homologous end joining (NHEI) or homology-directed recombination (HDR). The Cas9 can be guided to a target site (e.g., stabilized at a target site) within a target nucleic acid sequence by virtue of its association with the RNA-binding segment of the Cas9 to guide RNA. A Cas9 protein can bind and / or modify (e.g., cleave, nick, methylate, demethylate, etc.) a target nucleic acid and / or a polypeptide associated with target nucleic acid (e.g., methylation or acetylation of a histone tail; e.g., when the Cas9 protein includes a fusion partner with an activity). In some cases, the Cas9 protein is a naturally-occurring protein (e.g., naturally occurs in bacterial and / or archaeal cells). In other cases, the Cas9 protein is not a naturally-occurring polypeptide (e.g., the Cas9 protein is a variant Cas9 protein, a chimeric protein, and the like).
[0148] Naturally occurring Cas9 proteins may bind a Cas9 guide RNA, are thereby directed to a specific sequence within a target nucleic acid (a target site), and cleave the target nucleic acid (e.g., cleave dsDNA to generate a double strand break, cleave ssDNA, cleave ssRNA, etc.). A chimeric Cas9 protein may include a fusion protein comprising a Cas9 polypeptide fused to a heterologous protein (referred to as a fusion partner), where the heterologous protein provides an activity (e.g., one that is not provided by the Cas9 protein). The fusion partner can provide an activity, e.g., enzymatic activity (e.g., nuclease activity, activity for DNA and / or RNA methylation, activity for DNA and / or RNA cleavage, activity for histone acetylation, activity for histone methylation, activity for RNA modification, activity for RNA-binding, activity for RNA splicing etc.). In some cases, a portion of the Cas9 protein (e.g., the RuvC domain and / or the HNHdomain) exhibits reduced nuclease activity relative to the corresponding portion of a wild type Cas9 protein (e.g., in some cases the Cas9 protein is a nickase). In some cases, the Cas9 protein is enzymatically inactive, or has reduced enzymatic activity relative to a wild-type Cas9 protein (e.g., relative to Streptococcus pyogenes Cas9). In some cases, the Cas9 is a Cas9 nickase. The Cas9 nickase can be generated by mutating a Cas9 nuclease domain. Non-limiting example of the Cas9 nickase can include SpCas9, SaCas9, CjCas9, GeoCas9, HpaCas9, andNmeCas9. In some aspects, the endonuclease described herein comprises any one of the Cas9 in Table 1. In some aspects, the endonuclease described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the Cas9 in Table 1.
[0149] Some aspects include an endonuclease such as an RNA-guided endonuclease. The RNA-guided endonuclease may comprise a class II CRISPR / Cas endonuclease. The RNA-guided endonuclease may comprise a Cas9 endonuclease. The RNA-guided endonuclease may comprise a nickase. The RNA-guided endonuclease may comprise an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 1-13, or a functional fragment thereof.
[0150] The endonuclease may introduce a single-strand break in a target nucleic acid. The endonuclease may introduce a single-strand break in a target nucleic acid without cleaving a strand opposite the single strand break. The endonuclease may include a nickase. In some instances, the endonuclease may exclude an endonuclease that introduces a double strand break. The endonuclease may exclude a restriction enzyme.
[0151] The endonuclease may be included as part of a fusion protein. In some cases, an endonuclease is a fusion protein that is fused to a heterologous polypeptide such as the heterologous ligase described herein. The heterologous polypeptide may include a fusion partner. The fusion protein may include a fusion partner such as a DNA ligase, a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, or tag polypeptide. The fusion protein may include one or more fusion partner. The fusion protein may include a ligase. The fusion protein may include a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, or tag polypeptide.
[0152] The fusion partner may be connected to the N-terminus of the endonuclease. The fusion partner may be connected to the C-terminus of the endonuclease. The endonuclease may be connected at an N-terminus or a C-terminus to a linker. The fusion partner may be connected by the fusion partner’s N-terminus or C-terminus. The fusion partner may be connected by the fusion partner’s N-terminus to the endonuclease. The fusion partner may be connected by the fusion partner’s C-terminus to the endonuclease. The fusion partner may be connected at an N-terminus or a C-terminus to a linker.
[0153] In some cases, the endonuclease comprises a linker, where the linker covalently connects the endonuclease to the heterologous polypeptide. The linker may connect the endonuclease to any fusion partner. A linker may also connect any fusion partner to another fusion partner. The linker polypeptide may have any of a variety of amino acid sequences. Proteins can be joined by a spacer peptide, generally of a flexible nature, although other chemical linkages are not excluded. Suitable linkers include polypeptides of between 4 amino acids and 40 amino acids in length, or between 4 amino acids and 25 amino acids in length. These linkers can be produced by using synthetic, linker-encoding oligonucleotides to couple the proteins, or can be encoded by a nucleic acid sequence encoding the fusion protein. Peptide linkers with a degree of flexibility can be used. The linking peptides may have virtually any amino acid sequence, bearing in mind that the preferred linkers will have a sequence that results in a generally flexible peptide. The use of small amino acids, such as glycine and alanine, are of use in creating a flexible peptide. The creation of such sequences is routine to those of skill in the art. A variety of different linkers are commercially available and are considered suitable for use. Examples of linker polypeptides include glycine polymers (G)n, glycine-serine polymers (including, for example, (GS)n(SEQ ID NO:960), (GSGGS)n(SEQ ID NO:950), (GGSGGS)n(SEQ ID NO:951), and (GGGS)n(SEQ IDNO:952), where n is an integer of at least one); glycine-alanine polymers; and alanine-serine polymers. Exemplary linkers can comprise amino acid sequences including, but not limited to, GGSG(SEQ ID NO:954), GGSGG(SEQ ID NO:955), GSGSG(SEQ ID NO:956), GSGGG(SEQ ID NO:957), GGGSG(SEQ ID NO:958), GSSSG(SEQ ID NO:959), and the like. Also suitable is a linker having the sequence (GGGGS)n(SEQ ID NO:953), where n is an integer of from 1 to 10 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10). The ordinarily skilled artisan will recognize that design of a peptide conjugated to any desired element can include linkers that are all or partially flexible, such that the linker can include a flexible linker as well as one or more portions that confer less flexible structure.[00154J One or more linkers may be included in a fusion protein. 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 linkers, or a range of linkers defined by any two of the aforementioned integers, may be included in the fusion protein. A linker may connect to an N-terminal end of at least part of the endonuclease. A linker may connect to an N-terminal end of at least part of a fusion partner. A linker may connect to an N-terminal end of at least part of a fusion ligase. A linker may connect to an N-terminal end of a nuclear localization signal. A linker may connect to an N-terminal end of a chromatin modifying domain. A linker may connect to an N-terminal end of a cell penetrating peptide. A linker may connect to an N-terminal end of a tag polypeptide. A linker may connect to a C-terminal end of at least part of the endonuclease. A linker may connect to a C-terminal end of at least part of a fusion partner. A linker may connect to a C-terminal end of at least part of a fusion ligase. A linker may connect to a C-terminal end of a nuclear localization signal. A linker may connect to a C-terminal end of a chromatin modifying domain. A linker may connect to a C- terminal end of a cell penetrating peptide. A linker may connect to a C-terminal end of a tag polypeptide.
[0155] A linker may comprise a number or range of amino acids or residues. The linker may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 amino acid residues. The linker may, in some aspects, include no more than 1, no more than 2, no more than 3, no more than 4, no more than 5, no more than 6, no more than 7, no more than 8, no more than 9, no more than 10, no more than 12, no more than 13, no more than 14, no more than 15, no more than 20, no more than25, no more than 30, no more than 35, no more than 40, no more than 45, no more than 50, no more than 55, no more than 60, no more than 65, no more than 70, no more than 75, no more than 80, no more than 85, no more than 90, no more than 95, or no more than 100 amino acid residues. A linker may include 1-10 amino acids, 1-25 amino acids, or 1-100 amino acids.
[0156] Linkers may be included anywhere in a polypeptide chain or protein described herein. For example, a linker may separate an endonuclease from a ligase. A linker may separate an endonuclease from a nuclear localization signal, a chromatin modifying domain, a cell penetrating peptide, or a tag polypeptide.
[0157] In some cases, the endonuclease comprises a nuclear localization sequence (e.g., one or more nuclear localization signals or NLSs for targeting to the nucleus). In some aspects, the NLS described herein comprises any one of the NLS in Table 2. In some aspects, the NLS described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of NLS in Table 2.
[0158] A polynucleotide encoding an NLS polypeptide may be used. An example of such a polynucleotide may be SGGSx2-bpNLS-SGGSx2:TCCGGCGGAAGCTCTGGTGGCAGCAAGCGGACCGCCGACGGCTCTGAATTCGAGAG CCCTAAGAAGAAAAGAAAGGTGAGCGGAGGCTCTAGCGGCGGAAGC (SEQ ID NO:25).
[0159] In some aspects, the endonuclease comprises a dimerization domain. The dimerization domain can be located at the N-terminus or C-terminus of the endonuclease. In some aspects, the dimerization domain allows the endonuclease to form a heterodimer with another polypeptide (e.g., the heterologous ligase). In some aspects, the dimerization domain allows the endonuclease to be functionally coupled with another polypeptide. Non-limiting examples of the dimerization domains can include a leucine zipper, an FKBP, an FRB, a Calcineurin A, a CyP-Fas, a GyrB, a GAI, a GID1, a SNAP tag, a Halo tag, a Bcl-xL, a Fab, a LOV domain, or SpyTag / SpyCatcher. Other example of dimerization domain can include an antibody such as anyone of heavy chain domain 2 (CH2) of IgM (MHD2) or IgE (EHD2), immunoglobulin Fc region, heavy chain domain 3 (CH3) of IgG or IgA, heavy chain domain 4 (CH4) of IgM or IgE, Fab, Fab2, leucine zipper motifs, barnase-barstar dimers, miniantibodies, or ZIP miniantibodies. In some aspects, the dimerization domain described herein comprises any one of the dimerization domain in Table 3. In some aspects, the dimerization domain described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of dimerization domain in Table 3.
[0160] In some aspects, the endonuclease comprises at least one additional domain. In some aspects, the at least one additional domain is a functional domain. For example, the functional domain can comprise a chromatin modifying domain or a cell penetrating peptide. In some aspects, the chromatin modifying domain described herein comprises any one of the chromatin modifying domain in Table 4. In some aspects, the chromatin modifying domain described herein comprisesa polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of chromatin modifying domain in Table 4.
[0161] In some aspects, the cell penetrating peptide described herein comprises any one of the cell penetrating peptide in Table 5. In some aspects, the cell penetrating peptide described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of cell penetrating peptide in Table 5.
[0162] In some aspects, the endonuclease comprises a tag, where the tag can be used for increasing expression, identifying, or purifying the endonuclease. In some aspects, the tag described herein comprises any one of the tag sequence in Table 6. In some aspects, the tag described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the tag sequence in Table 6.[00163J In some embodiments, the endonuclease can be expressed as split construct as one or more exteins fused to one or more inteins. Intein technology may be used to deliver large proteins into a cell by expressing the protein as two or more shorter peptide segments (exteins). Each extein may be expressed as a fusion with an intein peptide (e.g., an Npu C intein or an Npu N intein). An intein may autocatalyze fusion of two or more exteins and may autocatalyze excision of the intein from its corresponding extein. The result may be a protein complex comprising a first extein fused to a second extein and lacking inteins. An intein may be positioned N-terminal of the extein, or an intein may be positioned C-terminal of the extein. An extein may comprise a cysteine residue positioned adjacent to the intein (e.g., at the C-terminal end of an extein with an intein fused to the C-terminal end of the extein). The Cas nickase may be expressed as two or more segments. A first of the Cas nickase segment may comprise an N-terminal portion of the Cas nickase. A first segment of the Cas nickase may comprise a first intein. A second segment of the Cas nickase may comprise a C-terminal portion of the Cas nickase. A second segment of the Cas nickase may comprise a second intein. An intein may be fused to a C-terminus of an N-terminal portion of the Cas nickase. An intein may be fused to an N-terminus of a C-terminal portion of the Cas nickase. A nucleic acid sequence encoding an extein-intein fusion may fit into a delivery vector (e.g., an adeno- associated virus (AAV) vector).DNA Ligases
[0164] Disclosed herein are ligases. The ligase may be or include a DNA ligase. The ligase may be included in a composition, system or method disclosed herein. The ligase may be recombinant. The ligase may be coupled to the endonuclease. The ligase may be coupled directly or indirectly to the endonuclease. The coupling may be covalent or non-covalent. The ligase may be bound or connected to the endonuclease. The ligase may be recruited to, be part of a fusion protein with, or be used in conjunction with an endonuclease. The ligase may be coupled to an integrase. The ligase may be coupled directly or indirectly to the integrase. The coupling may be covalent or non-covalent. The ligase may be bound or connected to the integrase. The ligase may be recruited to, be part of a fusion protein with, or be used in conjunction with an integrase. The ligase may be heterologous. The ligase may be endogenous. Where a heterologous ligase is described, a non-heterologous (e.g., endogenous) ligase may be used in some cases. The ligase may be encoded in a cell. The ligase may be delivered to the cell in trans. The ligase may form a phosphodiester bond by joining two nucleic acid ends together. The ligase may join an end (e.g., 5’ or 3’ end) of a target nucleic acid to an exogenous first integrating nucleic acid (e.g., a 3’ or 5’ end of the exogenous first integrating nucleic acid). The ligase ligates an exogenous first integrating nucleic acid (e.g., a donor nucleic acid) to a cleaved or nicked end of a target nucleic acid where the cleaved or nicked end has been generated by an endonuclease such as an RNA- guided endonuclease. The ligase may include any aspect included in Fig. 1A-6C.
[0165] The ligase may be non-naturally occurring. The ligase may be engineered. The ligase may be synthetic. The ligase may be pre-synthetized. The ligase may be added to a subject or a cell. The ligase may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0166] At least part of the ligase may be included in a first polypeptide. At least part of the ligase may be included in a second polypeptide. The ligase may be split into two polypeptides bound together. The first polypeptide may include an N-terminal portion of the ligase. The first polypeptide may include a C-terminal portion of the ligase. The second polypeptide may include the N-terminal portion of the ligase. The second polypeptide may include the C-terminal portion of the ligase. The first or second polypeptide comprising a part of the ligase may be fused with at least part, or the whole, of the endonuclease. The first or second polypeptide comprising a part of the ligase may be fused with at least part, or the whole, of the integrase.
[0167] Examples of DNA ligases are hLIGl, T4 ligase, T7 ligase, and ligases from Aquifex aeolicus VF5, Neisseria meningitidis serogroup A strain Z2491, Neisseria meningitidis serogroup B strain MC58, Pseudomonas aeruginosa PA01, Vibrio cholerae El Tor N 1696, Vaccinia virus, and Emiliania huxleyi virus.
[0168] The ligase may comprise a ligase that can ligate a substrate comprising DNA. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising a DNA splint. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized to another DNA strand. The splinting DNA strand may include an RNA portion. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized across from a DNA portion of an RNA / DNA hybrid strand. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising a DNA / RNA. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising an RNA splint. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized to an RNA strand. The RNA strand may include a DNA portion. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized across from an RNA portion of an RNA / DNA hybrid strand.
[0169] In some aspects, the ligase described herein comprises any one of the ligases in Table 7. In some aspects, the ligase described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the ligase in Table 7.
[0170] Some aspects include a DNA ligase that ligates DNA strands base paired to a DNA splint. In some embodiments, the DNA ligase ligates DNA strands base paired to an RNA splint. In some embodiments, the DNA ligase comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 55-96, or a functional fragment thereof.
[0171] In some aspects, the ligases comprises at least one NLS (e.g., any one of the NLS in Table 2). In some aspects, the ligase comprises at least one additional domain. In some aspects, the at least one additional domain is a dimerization domain (e.g., any one of the dimerization domain in Table 3). In some aspects, the ligase comprising a dimerization domain can be dimerized with an endonuclease to form a heterodimer. In some aspects, the at least one additional domain is a functional domain. For example, the functional domain can comprise a chromatin modifying domain (e.g., any one of the chromatin modifying domain in Table 4) or a cellpenetrating peptide (e.g., any one of the cell penetrating peptide in Table 5). In some aspects, the ligase comprises a linker, where the linker can covalently connect the ligase with another polypeptide (e.g., the endonuclease). In some aspects, the linker covalently connects the ligase to the at least one additional domain. In some aspects, the ligase comprises a tag (e.g., any one of the tag in Table 6, where the tag can be used for increasing expression, identifying, or purifying the ligase. A linker may separate the ligase from a nuclear localization signal, a chromatin modifying domain, a cell penetrating peptide, or a tag polypeptide. Any linker described herein may be included.
[0172] The ligase may comprise a binding motif for binding to a nucleic acid motif (e.g., a hairpin motif). In some aspects, the ligase (e.g., DNA ligase) comprises an MS2 coat protein (MCP) peptide. The ligase may include a hairpin binding motif such as an MCP peptide. The MCP peptide may be useful for recruiting the ligase to a guide nucleic acid comprising an MS2 hairpin. A benefit of using a MCP peptide and MS2 hairpin is to separate the ligase and endonuclease such as a Cas nickase (or a portion of them) and allow fitting within separate vectors such as AAV vectors. In some aspects, the ligase comprises a loop region. In some aspects, the loop region is a 2a loop or a 3a loop. The loop region may comprise a 2a loop. The loop region may comprise a 3a loop.Integrases
[0173] Disclosed herein are integrases. An integrase may be an example of a recombinase, and where an integrase is described, a recombinase may be contemplated. The integrase may be or include a phage integrase. The integrase may be or include a site-specific recombinase. The integrase may be or include a serine integrase. The integrase may be or include a resolvase. The integrase may be or include a DNA invertase. The integrase may be or include a tyrosine integrase. The integrase may be or include a retrotransposase. The integrase may be included in a composition, system or method disclosed herein. The integrase may be recombinant. Any of these integrases may be modified or mutated. For example, an integrase may include an insertion, a deletion, or may be an active or functional fragment. The integrase may include a mutation of an integrase described herein. The integrase may be coupled to the endonuclease. The integrase may be coupled directly or indirectly to the endonuclease. The coupling may be covalent or non- covalent. The integrase may be bound or connected to the endonuclease. The integrase may be recruited to, be part of a fusion protein with, or be used in conjunction with an endonuclease. Theintegrase may be coupled to a ligase. The integrase may be coupled directly or indirectly to the ligase. The coupling may be covalent or non-covalent. The integrase may be bound or connected to the ligase. The integrase may be recruited to, be part of a fusion protein with, or be used in conjunction with the ligase. The integrase may be heterologous. The integrase may be endogenous. Where a heterologous integrase is described, a non-heterologous (e g., endogenous) integrase may be used in some cases. The integrase may be encoded in a cell. The integrase may be delivered to the cell in trans. The integrase introduces a second integrating nucleic acid into a target nucleic acid by recognizing and binding to an exogenous first integrating nucleic acid (e.g., a recombination sequence).[00174J The integrase may be non-naturally occurring. The integrase may be engineered. The integrase may be synthetic. The integrase may be pre-synthetized. The integrase may be added to a subject or a cell. The integrase may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0175] At least part of the integrase may be included in a first polypeptide. At least part of the integrase may be included in a second polypeptide. The integrase may be split into two polypeptides bound together. The first polypeptide may include an N-terminal portion of the integrase. The first polypeptide may include a C-terminal portion of the integrase. The second polypeptide may include the N-terminal portion of the integrase. The second polypeptide may include the C-terminal portion of the integrase. The first or second polypeptide comprising a part of the integrase may be fused with at least part, or the whole, of the endonuclease. The first or second polypeptide comprising a part of the integrase may be fused with at least part, or the whole, of the ligase.
[0176] In some aspects, the integrase described herein may be a serine integrase. A serine integrase may be referred to as a “serine recombinase.” Examples of species that contain serine recombinases are: Bacillus cereus,' Bacillus safe n si s; Bacillus tropicus, Burkholderia multivorans, Burkholderia ubonensis,' Cellulosimicrobium cellulans, Clostrid im botulinum., Clostridioides difficile,- Clostridium perfringens,' Clostridium thermobutyricunr, Cronobacter sakazakii, Desulfotomaculum nigrificans, Enterocloster clostridioformis, Enterococcus faecalis,' Enterococcus faecium, Escherichia coir, Eubacterium maltosivorans; Faecalibacterium prausnitzii,' Fusobacterium mortiferum, Klebsiella pneumoniae, Mycobacteroides abscessus,' Mycobacterium phage Bxbl,' Mycolicibacterium elephantis. Neobacillus mesonae, Nocardiaolitidiscaviarum, Paenibacillus campinasensis, Paeniclostridium sordellii, Parageobacillus caldoxylosilyticus,' Pseudomonas aeruginosa., Pseudomonas fluorescens,' Pseudomonas fulva.' Pseudomonas putida,' Pseudomonas syringae, Prochlorothrix hollandica, Rhizobiales bacterium, Rhodococcus hoagie. Ruminococcus lactaris,' Salinispora pacifica, Staphylococcus arlettae,' Staphylococcus aureus,' Staphylococcus hominis,' Streptococcus agalactiac, Streptococcus equinus,' Streptococcus mitis, Streptomyces ipomoeae, Streptomyces phage PhiC31, Tenacibaculum dicentrarchi,' Treponema denticola, Vibrio harveyi,' Vibrio hyugaensis,' Vibrio parahaemolyticus .
[0177] In some aspects, the integrase described herein may be from or may be derived from the species Bacillus cereus. In some aspects, the integrase described herein may be from or may be derived from the species Bacillus safensis. In some aspects, the integrase described herein may be from or may be derived from the species Bacillus tropicus. In some aspects, the integrase described herein may be from or may be derived from the species Burkholderia multivorans. In some aspects, the integrase described herein may be from or may be derived from the species Burkholderia ubonensis. In some aspects, the integrase described herein may be from or may be derived from the species Cellulosimicrobium cellulans. In some aspects, the integrase described herein may be from or may be derived from the species Clostridium botulinum. In some aspects, the integrase described herein may be from or may be derived from the species Clostridioides difficile. In some aspects, the integrase described herein may be from or may be derived from the species Clostridium perfringens. In some aspects, the integrase described herein may be from or may be derived from the species Clostridium thermobutyricum . In some aspects, the integrase described herein may be from or may be derived from the species Cronobacter sakazakii. In some aspects, the integrase described herein may be from or may be derived from the species Desulfotomaculum nigrificans . In some aspects, the integrase described herein may be from or may be derived from the species Enterocloster clostridioformis . In some aspects, the integrase described herein may be from or may be derived from the species Enterococcus faecalis. In some aspects, the integrase described herein may be from or may be derived from the species Enterococcus faecium. In some aspects, the integrase described herein may be from or may be derived from the species Escherichia coli. In some aspects, the integrase described herein may be from or may be derived from the species Eubacterium maltosivorans . In some aspects, the integrase described herein may be from or may be derived from the species Faecalibacteriumprausnitzii. In some aspects, the integrase described herein may be from or may be derived from the species Fusobacterium mortiferum. In some aspects, the integrase described herein may be from or may be derived from the species Klebsiella pneumoniae. In some aspects, the integrase described herein may be from or may be derived from the species Mycobacteroides abscessus. In some aspects, the integrase described herein may be from or may be derived from the species Mycobacterium phage Bxbl. In some aspects, the integrase described herein may be from or may be derived from the QCiQsMycolici bacterium elephantis. In some aspects, the integrase described herein may be from or may be derived from the species Neobacillus mesonae. In some aspects, the integrase described herein may be from or may be derived from the species Nocardia otitidiscaviarum. In some aspects, the integrase described herein may be from or may be derived from the species Paenibacillus campinasensis . In some aspects, the integrase described herein may be from or may be derived from the species Paeniclostridium sordellii. In some aspects, the integrase described herein may be from or may be derived from the species Parageobacillus caldoxylosilyticus . In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas aeruginosa. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas fluorescens. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas fulva. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas putida. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas syringae. In some aspects, the integrase described herein may be from or may be derived from the species Prochlorothrix hollandica. In some aspects, the integrase described herein may be from or may be derived from the species Rhizobiales bacterium. In some aspects, the integrase described herein may be from or may be derived from the species Rhodococcus hoagie. In some aspects, the integrase described herein may be from or may be derived from the species Ruminococcus lactaris. In some aspects, the integrase described herein may be from or may be derived from the species Salinispora pacifica. In some aspects, the integrase described herein may be from or may be derived from the species Staphylococcus arlettae. In some aspects, the integrase described herein may be from or may be derived from the species Staphylococcus aureus. In some aspects, the integrase described herein may be from or may be derived from the species Staphylococcus hominis. In some aspects, the integrase described herein may be from or may be derived from the species Streptococcus agalactiae. In some aspects,the integrase described herein may be from or may be derived from the species Streptococcus equinus. In some aspects, the integrase described herein may be from or may be derived from the species Streptococcus mitis. Streptomyces ipomoeae. In some aspects, the integrase described herein may be from or may be derived from the species Streptomyces phage PhiC31. In some aspects, the integrase described herein may be from or may be derived from the species Tenacibaculum dicentrarchi. In some aspects, the integrase described herein may be from or may be derived from the species. In some aspects, the integrase described herein may be from or may be derived from the species Treponema denticola. In some aspects, the integrase described herein may be from or may be derived from the species Vibrio harveyi. In some aspects, the integrase described herein may be from or may be derived from the species Vibrio hyugaensis. In some aspects, the integrase described herein may be from or may be derived from the species Vibrio parahaemolyticus .
[0178] In some aspects, the integrase described herein comprises any one of the integrases in Table 8. In some aspects, the integrase described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8.
[0179] In some aspects, the integrase comprises any integrase that recognizes and binds to a recombination sequence in Table 13. In some aspects, the integrase comprises any integrase that recognizes and binds to a recombination sequence in Table 13. In some aspects, the recombination sequence comprises a sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to a recombination sequence in Table 13. In some aspects, the integrase described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to that of an integrase that recognizes and binds to a polypeptide sequence of any one of the recombination sequences in Table 13.
[0180] In some aspects, the integrase described herein may be a tyrosine integrase. A tyrosine integrase may be referred to as a “tyrosine recombinase.” Examples of tyrosine integrases may include: BS codV; BS ripX; BS ydcL; CB tnpA; CollD; CP4; Cre; D29; DLP12; DN int; EC FimB; EC FimE; EC orf; EC xerC; EC xerD; Oil; 013; 080; Oadh; OCTX; OLC3; FLP; OR73; HI orf; HI rci; HI xerC; HI xerD; HK22; HP1; L2; L5; L54; X; LL orf; LL xerC; LO L5; MJ orf; MP int; MT int; MT orf; MV4; P186; P2; P21; P22; P4; P434; PA sss; PM fimB; pAEl; pCLl;pKDl; pMEA; pSAM2; pSB2; pSB3; pSDL2; pSElOl; pSE211; pSMl; pSRl; pWS58; R721; Rci; sF6; SLP1; SM orf; SsrA; SSV1; T12; Tn21; Tn4430; Tn554a; Tn554b; Tn7; Tn916; Tuc; WZ int; Xis A; or XisC.
[0181] Examples of species that contain tyrosine recombinases may include : Bacillus subtilis; Clostridium butyricum; Escherichia coli, Mycobacterium smegmatis, Dichelobacter nodosus; Staphylococcus aureus; E. coli phage; Lactobacillus gasseri; Pseudomonas aeruginosa; Lactococcis lactis; Saccharomyces cerevisiae; Haemophilus influenzae; Mycoplasma sp.; Mycobacterium tuberculosis; Lactobacillus leichmannii; Leuconostoc oenos; Methanococcus jannaschi; Mycobacterium leprae; Mycobacterium paratuberculosis; Lactobacillus delbrueckii; Salmonella typhimurium; Pseudomonas aeruginosa; Proteus mirabilis; Alcaligenes eutrophus; Chlorobium limicola; Kluyveromyces lactis; Amycolatopsis methanolica; Streptomyces ambofaciencs; Zygosaccharomyces bailii; Zygosaccharomyces bisporus; Salmonella dublin; Saccharopolyspora erythraea; Zygosaccharomyces fermentati; Zygosaccharomyces rouxii; Shigella flexneri; Streptomyces coelicolor; Serratia marcescens; Methanosarcina acetivorans; Sultolobus sp.; Streptococcus pyogenes; Bacillus thurinigiensis; Enteroccus faecalis; Lactobacillus lactis; Weeksella zoohelcum; or Anabaena sp.
[0182] In some aspects, the integrase described herein may be a gamma-delta resolvase from the Tn 1000 transposon. In some aspects, the integrase described herein may be a Hin recombinase. In some aspects, the integrase described herein may be a Tn3 resolvase from the Tn3 transposon. In some aspects, the integrase described herein may be a Tre recombinase. In some aspects, the integrase described herein may be a Dre recombinase. In some aspects, the integrase described herein may be a Cre recombinase. In some aspects, the integrase described herein may be a flippase (Flp). In some aspects, the integrase described herein may be a KD recombinase. In some aspects, the integrase described herein may be a B2 B3 recombinase. In some aspects, the integrase described herein may be an HK022 integrase. In some aspects, the integrase described herein may be a ParA integrase. In some aspects, the integrase described herein may be a Gin integrase. In some aspects, the integrase described herein the integrase may be an R4 recombinase.
[0183] In some embodiments, the integrase described herein may be a Vika recombinase. In some embodiments, the integrase described herein may be an RDF recombinase. In some embodiments, the integrase described herein may be a (pBTl recombinase. In some embodiments, the integrase described herein may be an R1 recombinase. In some embodiments, the integrasedescribed herein may be an R2 recombinase. In some embodiments, the integrase described herein may be an R3 recombinase. In some embodiments, the integrase described herein may be an R4 integrase. In some embodiments, the integrase described herein may be an R5 integrase. In some embodiments, the integrase described herein may be a TP901-1 recombinase. In some embodiments, the integrase described herein may be a Al 18 recombinase. In some embodiments, the integrase described herein may be a cpFCl recombinase. In some embodiments, the integrase described herein may be a cpCl recombinase. In some embodiments, the integrase described herein may be a MR11 recombinase. In some embodiments, the integrase described herein may be a TGI recombinase. In some embodiments, the integrase described herein may be a cp370.1 recombinase. In some embodiments, the integrase described herein may be a W0 recombinase. In some embodiments, the integrase described herein may be a BL3 recombinase. In some embodiments, the integrase described herein may be a SPBc recombinase. In some embodiments, the integrase described herein may be aK38 recombinase. In some embodiments, the integrase described herein may be a Peaches recombinase. In some embodiments, the integrase described herein may be a Veracruz recombinase. In some embodiments, the integrase described herein may be a Rebeuca recombinase. In some embodiments, the integrase described herein may be a Theia recombinase. In some embodiments, the integrase described herein may be a Benedict recombinase. In some embodiments, the integrase described herein may be a KSSJEB recombinase. In some embodiments, the integrase described herein may be a PattyP recombinase. In some embodiments, the integrase described herein may be a Doom recombinase. In some embodiments, the integrase described herein may be a Scowl recombinase. In some embodiments, the integrase described herein may be a Lockley recombinase. In some embodiments, the integrase described herein may be a Switzer recombinase. In some embodiments, the integrase described herein may be a Bob3 recombinase. In some embodiments, the integrase described herein may be a Troube recombinase. In some embodiments, the integrase described herein may be an Abrogate recombinase. In some embodiments, the integrase described herein may be an Anglerfish recombinase. In some embodiments, the integrase described herein may be a Sarfire recombinase. In some embodiments, the integrase described herein may be a SkiPole recombinase. In some embodiments, the integrase described herein may be a Conceptll recombinase. In some embodiments, the integrase described herein may be a Museum recombinase. In some embodiments, the integrase described herein may be a Severus recombinase. In some embodiments, the integrase described herein may be an Airmidrecombinase. In some embodiments, the integrase described herein may be a Benedict recombinase. In some embodiments, the integrase described herein may be a Hinder recombinase. In some embodiments, the integrase described herein may be a ICleared recombinase. In some embodiments, the integrase described herein may be a Sheen recombinase. In some embodiments, the integrase described herein may be a Mundrea recombinase. In some embodiments, the integrase described herein may be a BxZ2 recombinase. In some embodiments, the integrase described herein may be a cpRV recombinase.
[0184] In some aspects, the integrase described herein may be a retrotransposase encoded by R2. In some aspects, the integrase described herein may be a retrotransposase encoded by LI. In some aspects, the integrase described herein may be a retrotransposase encoded by Tol2. In some aspects, the integrase described herein may be a retrotransposase encoded by Tel . In some aspects, the integrase described herein may be a retrotransposase encoded by Tc3. In some aspects, the integrase described herein may be a retrotransposase encoded by Mariner (Himar 1). In some aspects, the integrase described herein may be a retrotransposase encoded by Mariner (mos 1). In some aspects, the integrase described herein may be a retrotransposase encoded by Minos.
[0185] As can be used herein, Xu et al describes methods for evaluating integrase activity in E. coll and mammalian cells and confirmed at least R4, q>C31, (pBTl, Bxbl, SPBc, TP901-1 and WP integrases to be active on substrates integrated into the genome of HT1080 cells (Xu et al., 2013, Accuracy and efficiency define Bxbl integrase as the best of fifteen candidate serine recombinases for the integration of DNA into the human genome. BMC Biotechnol. 2013 Oct 20;13:87. Doi: 10.1186 / 1472-6750-13-87). Durrant describes new large serine recombinases (LSRs) divided into three classes distinguished from one another by efficiency and specificity, including landing pad LSRs which outperform wild-type Bxbl in episomal and chromosomal integration efficiency, LSRs that achieve both efficient and site-specific integration without a landing pad, and multi-targeting LSRs with minimal site-specificity. Additionally, embodiments can include any serine recombinase such as BceINT, SSCINT, SACINT, and INT10 (s lonnidi et al., 2021; Drag-and-drop genome insertion without DNA cleavage with CRISPR directed integrases. bioRxiv 2021.11.01.466786, doi.org / 10.1101 / 2021.11.01.466786). In some embodiments, the integration site can be selected from an attB site, an attP site, an attL site, an attR site, a lox71 site a Vox site, or a FRT site. In instances in this disclosure that refer to a Cre- lox system, the Cre-lox system is referred to either as a control for programmable gene insertionor as a tool for a recombinase-mediated event separate and distinct from insertion of the donor polynucleotide template (or exogenous nucleic acid) into the integrated recognition site.
[0186] In some aspects, the integrase described herein may be coupled to a recombination directionality factor (RDF). In some aspects, the integrase described herein may be fused to an RDF. In some aspects, the integrase described herein may be linked to an RDF. The RDF may comprise a gp3 RDF. The RDF may comprise a gp47 RDF.Fusion Proteins
[0187] Disclosed herein are fusion proteins. Some aspects include a nucleic acid (e.g., an expression vector) encoding a fusion protein. The fusion protein may include an endonuclease. The fusion protein may include a ligase. The fusion protein may include an integrase. The fusion protein may include a linker. The fusion protein may include two linkers. The fusion protein may include a plurality of linkers. The endonuclease and ligase may be connected through a linker. The endonuclease and the integrase may be connected through a linker. The ligase and the integrase may be connected through a linker. The fusion protein may be an example of a covalently coupled endonuclease and DNA ligase. The fusion protein may be an example of a covalently coupled endonuclease and an integrase. The fusion protein may be an example of a covalently coupled DNA ligase and an integrase. The fusion protein may comprise an endonuclease such as an RNA- guided endonuclease fused to a DNA ligase. The fusion protein may comprise an endonuclease such as an RNA-guided endonuclease fused to an integrase such as a serine integrase or a tyrosine integrase. The fusion protein may comprise a DNA ligase fused to an integrase.
[0188] The fusion protein may be non-naturally occurring. The fusion protein may be engineered. The fusion protein may be synthetic. The fusion protein may be pre-synthetized. The fusion protein may be added to a subject or a cell. The fusion protein may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0189] The fusion protein may be a double fusion protein. The double fusion protein may include an endonuclease such as an RNA-guided endonuclease and a ligase. The double fusion protein may include an endonuclease such as an RNA-guided endonuclease and an integrase. The double fusion protein may include a ligase and an integrase.
[0190] The double fusion protein including an RNA-guided endonuclease and a DNA ligase may include one of various orientations. For example, the double fusion protein may include an RNA-guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g.,C-terminal or in the C-direction) relative to the DNA ligase. The double fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the DNA ligase. The double fusion protein may include an RNA-guided endonuclease carboxy (C)-terminal to the DNA ligase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the ligase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the ligase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The ligase may be N-terminal. The ligase may be C-terminal.
[0191] The double fusion protein including an RNA-guided endonuclease and an integrase may include one of various orientations. For example, the double fusion protein may include an RNA-guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the integrase. The double fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the integrase. The double fusion protein may include an RNA-guided endonuclease carboxy (C)-terminal to the integrase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the integrase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the integrase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The integrase may be N-terminal. The integrase may be C-terminal.
[0192] The double fusion protein including a ligase and an integrase may include one of various orientations. For example, the double fusion protein may include an integrase upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the ligase. The double fusion protein may include an integrase (N)-terminal to the ligase. The double fusion protein may include an integrase carboxy (C)-terminal to the ligase. The integrase may be in the amino direction within the fusion polypeptide relative to the ligase. The integrase may be in the carboxy direction within the fusion polypeptide relative to the ligase. The ligase may be N-terminal. The ligase may be C-terminal. The integrase may be N-terminal. The integrase may be C-terminal.
[0193] The fusion protein may be a triple fusion protein. The triple fusion protein may include an endonuclease such as an RNA-guided endonuclease, a ligase, and an integrase.
[0194] The triple fusion protein including an endonuclease, a ligase, and an integrase may include one of various orientations. For example, the triple fusion protein may include an RNA- guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the DNA ligase. The triple fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the DNA ligase. The triple fusion protein may include an RNA-guided endonuclease carboxy (C)-terminal to the DNA ligase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the ligase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the ligase. The triple fusion protein may include an RNA-guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the integrase. The triple fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the integrase. The triple fusion protein may include an RNA-guided endonuclease carboxy (C)- terminal to the integrase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the integrase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the integrase. The triple fusion protein may include an integrase upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C- direction) relative to the ligase. The triple fusion protein may include an integrase (N)-terminal to the ligase. The triple fusion protein may include an integrase carboxy (C)-terminal to the ligase. The integrase may be in the amino direction within the fusion polypeptide relative to the ligase. The integrase may be in the carboxy direction within the fusion polypeptide relative to the ligase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The ligase may be N-terminal. The ligase may be C-terminal. The integrase may be N-terminal. The integrase may be C-terminal.
[0195] The fusion protein may include a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or exonuclease. The fusion protein may include a nuclear localization signal. The fusion protein may include a chromatin modifying domain. The fusion protein may include a cell penetrating peptide. The fusion protein may include a tag polypeptide. The fusion protein may include an exonuclease. Any of the nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or exonuclease, endonuclease, ligase, or integrase may be directly connected to another or to the endonuclease, ligase, or integrase. Any of the nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or exonuclease, endonuclease, ligase, or integrase may be connected by a linker to another or to the endonuclease, ligase, or integrase. Multiple linkers may be included in the fusion protein. The fusion protein may exclude a polymerase.
[0196] A linker may include an amino acid linker. The amino acid linker may include a length of residues. The length may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 residues, or a range of residues defined by any two of the aforementioned integers. The length may include at least 1 residue, at least 2 residues, at least 3 residues, at least 4 residues, at least 5 residues, at least 6 residues, at least 7 residues, at least 8 residues, at least 9 residues, at least 10 residues, at least 15 residues, at least 20 residues, at least 25 residues, at least 30 residues, at least 40 residues, at least 50 residues, at least 60 residues, at least 70 residues, at least 80 residues, at least 90 residues, or at least 100 residues. In some aspects, the length may include less than 2 residues, less than 3 residues, less than 4 residues, less than 5 residues, less than 6 residues, less than 7 residues, less than 8 residues, less than 9 residues, less than 10 residues, less than 15 residues, less than 20 residues, less than 25 residues, less than 30 residues, less than 40 residues, less than 50 residues, less than 60 residues, less than 70 residues, less than 80 residues, less than 90 residues, or less than 100 residues. Examples of residues may include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine, or any combination thereof. The linker may be non-enzymatic or may lack any enzymatic activity.
[0197] A connection may be covalent. A covalent connection may include a peptide bond. The peptide bond may include amide bond. A connection may be between an N-terminus and another N-terminus. A connection may be between a C-terminus and another C-terminus. A connection may be between an N-terminus and a C-terminus. A connection may be between a C-terminus and an N-terminus.
[0198] The fusion protein may include connections in various orientations. The endonuclease may be connected at its C-terminus. The endonuclease may be connected at its N-terminus. The ligase may be connected at its C-terminus. The ligase may be connected at its N-terminus. The integrase may be connected at its C-terminus. The integrase may be connected at its N-terminus.
[0199] Fig. 7 illustrates some examples of fusion proteins including an endonuclease and a ligase. The figure includes examples of arrangements and orientations of the endonuclease, linker, ligase, or nuclear localization signal. Other aspects may be incorporated into the examples shown.
[0200] Fig. 13A and 13B illustrate some examples of fusion proteins including at least two of an endonuclease, a ligase, and an integrase. The figures includes examples of arrangements and orientations of the endonuclease, ligase, or integrase. Fig. 13A includes examples of double fusionproteins. Fig. 13B includes examples of triple fusion proteins. Other aspects may be incorporated into the examples shown.
[0201] In some embodiments, a fusion protein contemplated herein can comprise a DNA- binding domain of the Rad51 DNA repair protein (rad51DBD). In some embodiments, a fusion protein can comprise a high-mobility group nucleosome binding domain 1 (HN1) and a histone Hl central globular domain (H1G). In some embodiments, a fusion protein can comprise a Rad51DBD, and an HN1 and an H1G.
[0202] In some embodiments, a fusion protein contemplated herein comprises Brex27. In some embodiments, Brex27 can be fused to an endonuclease. In some embodiments, Brex27 can be fused to nCas9.
[0203] Table 9 provides example fusion proteins, which may be useful for a genome revising system. In some embodiments, a rad51DBD, and / or an HN1 and a H1G are fused to the fusion protein. In some embodiments, Brex27 is fused to the fusion protein. A fusion protein contemplated herein can comprise an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to an amino acid sequence provided in Table 9.
[0204] In some embodiments, a fusion protein contemplated herein can comprise a bicistronic mRNAs encoding for phosphomimetic peptide from IGFl(IGFlpml) and N-terminal peptide from NFATC2IP (NFATC2IPpl) peptides (IN peptides).
[0205] In some embodiments, a fusion protein comprises an mRNA encoding an IN peptide as described in Table 27 are fused to the fusion protein. A fusion protein contemplated herein can comprise an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to an amino acid sequence provided inTable 27Non-Covalently Coupled Proteins
[0206] Disclosed herein are non-covalently coupled proteins. Some aspects relate to a nucleic acid (e.g., an expression vector) encoding a protein, or encoding at least part of a protein. The proteins may include an endonuclease such as an RNA-guided endonuclease. A protein of the non- covalently coupled proteins may include a portion of an endonuclease. A protein of the non- covalently coupled proteins may include a portion of a ligase. The proteins may include a ligase such as a DNA ligase. A protein of the non-covalently coupled proteins may include an integrase. A protein of the non-covalently coupled proteins may include a portion of an integrase (e.g. a functional integrase fragment). The proteins may include an integrase such as a serine integrase. The proteins may include an integrase such as a tyrosine integrase. A protein of the non-covalently coupled proteins may include a fusion protein.
[0207] The non-covalently coupled proteins may be bound together through heterodimerization domains. Examples of heterodimerization domains may include a leucine zipper, PDZ domain, streptavidin, streptavidin binding protein, foldon domain, hydrophobic moiety, or a functional binding fragment thereof. A heterodimerization domain may include a leucine zipper. A heterodimerization domain may include a PDZ domain. A heterodimerization domain may include a streptavidin. A heterodimerization domain may include a streptavidin binding protein. A heterodimerization domain may include a foldon domain. A heterodimerization domain may include a hydrophobic moiety. A heterodimerization domain may include an antibody or antibody fragment. The non-covalently coupled proteins may be bound together through inteins.
[0208] The endonuclease and ligase may be coupled together by a separate molecule. The endonuclease and the integrase may be coupled together by a separate molecule. The ligase and the integrase may be coupled together by a separate molecule. The separate molecule maycomprise a nucleic acid (e g., a guide nucleic acid). The ligase may include a hairpin binding motif, where the RNA-guided endonuclease and the DNA ligase are coupled with the nucleic acid. The ligase may include a hairpin binding motif, where the integrase and the DNA ligase are coupled with the nucleic acid. The nucleic acid may include a scaffold that binds the RNA-guided endonuclease and a hairpin that binds to the hairpin binding motif. The hairpin binding motif may include an MS2 coat protein (MCP) peptide. The hairpin may include an MS2 hairpin.
[0209] The endonuclease and ligase may be coupled together by a heterobifunctional molecule. The endonuclease and integrase may be coupled together by a heterobifunctional molecule. The ligase and integrase may be coupled together by a heterobifunctional molecule. The heterobifunctional molecule may include an endonuclease binding domain and a DNA ligase binding domain. The heterobifunctional molecule may include an endonuclease binding domain and an integrase binding domain. The heterobifunctional molecule may include a ligase binding domain and an integrase binding domain. The heterobifunctional molecule may include an endonuclease binding domain. The endonuclease binding domain may include a heterodimerization domain. The endonuclease binding domain may include an antibody or antibody binding fragment. The heterobifunctional molecule may include a ligase binding domain such as a DNA ligase binding domain. The DNA ligase binding domain may include a heterodimerization domain. The DNA ligase binding domain may include an antibody or antibody binding fragment. The heterobifunctional molecule may include an integrase binding domain such as a serine integrase or a tyrosine integrase binding domain. The integrase binding domain may include a heterodimerization domain. The integrase binding domain may include an antibody or antibody binding fragment. The heterobifunctional molecule may include a small molecule. The small molecule may comprise a proteolysis targeting chimera (PROTAC), or a related heterobifunctional molecule.
[0210] Some aspects include a protein complex, comprising: an RNA-guided endonuclease bound to a DNA ligase. The endonuclease and the DNA ligase may be bound together through heterodimerization domains. The protein complex of embodiment 75, wherein the heterodimerization domains may comprise leucine zippers, PDZ domains, streptavidin, and streptavidin binding protein, foldon domains, hydrophobic polypeptides, an antibody that binds the Cas nickase, or an antibody that binds the DNA ligase, or one or more binding fragments thereof. The protein complex may be included in a cell. The cell may further include a heterologousRNA-guided endonuclease and a DNA ligase that that was introduced into the cell. The cell may further include a nuclease that is different from the RNA-guided endonuclease.
[0211] In some aspects, a protein complex contemplated herein can comprise a monomeric streptavidin (mSA). In some embodiments, the monomeric streptavidin can be fused to an endonuclease. In some embodiments, the monomeric streptavidin can be fused to a nicking Cas9. In some embodiments, the monomeric streptavidin can be fused to a ligase. In some embodiments, the monomeric streptavidin can be attached to a biotinylated splinting nucleic acid. In some embodiments, the biotin modification is on the / 5Biosg / end of the splinting nucleic acid. In some embodiments, the biotin modification is on the / 3Bio / end of the splinting nucleic acid.[00212J Fig. 20 illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus. In this illustration, a biotinylated splinting nucleic acid is attached to a monomeric streptavidin fused to nicking Cas9. The monomeric streptavidin can also be fused to the ligase (not illustrated).
[0213] Table 10 provides non-limiting examples of fusion protein comprising mSA contemplated herein. A fusion protein can be any fusion protein comprising an amino acid sequence provided in Table 10. A fusion protein contemplated herein can comprise an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to an amino acid sequence provided in Table 10.Guide Nucleic Acids
[0214] Disclosed herein are guide nucleic acids. The guide nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate to a nucleic acid (e.g., DNA or an expression vector) that encodes a guide nucleic acid such as a guide RNA. Provided herein are guide nucleic acids (e.g., gRNAs) that direct a programmable endonuclease (e.g., a nCas9) to a target nucleic acid (e.g., a genomic locus). The guide nucleic acid may guide an RNA-guided endonuclease to a target nucleic acid locus for nucleic acid replacement or gene editing at thelocus. A guide nucleic acid of the present disclosure may facilitate a donor strand to be inserted into a target site of the target nucleic acid. A guide nucleic acid of the present disclosure may facilitate editing of a nucleic acid sequence at a target site of the target nucleic acid. The guide nucleic acid may, in some instances, also act as a splint for a DNA ligase described herein, such as for ligating two nucleic acid strands base paired to a portion of the guide nucleic acid. The guide nucleic acid may be single stranded. The guide nucleic acid may include RNA. The guide nucleic acid may be RNA. The guide nucleic acid may include a guide RNA (gRNA). In some cases, a guide nucleic acid may include DNA.
[0215] The guide nucleic acid may be non-naturally occurring. The guide nucleic acid may be engineered. The guide nucleic acid may be synthetic. The guide nucleic acid may be pre- synthetized. The guide nucleic acid may be added to a subject or a cell. In some aspects, the guide nucleic acid does not include a template for a polymerase.
[0216] The guide nucleic acid may include an exogenous first integrating nucleic acid binding site. The exogenous first integrating nucleic acid binding site may be referred to as a “donor binding site” or vice versa.
[0217] Disclosed herein are guide nucleic acids, comprising: a spacer reverse complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; and a donor nucleic acid binding site and optionally a flap binding site reverse complementary to a nucleic acid flap.
[0218] In some aspects, the guide nucleic acid comprises a spacer complementary to a genomic locus in a cell; a scaffold for complexing with the at least one endonuclease; a donor binding site that is at least partially complementary to a donor strand; a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; or a combination thereof. In some aspects, the guide nucleic acid can direct the at least one endonuclease to cleave at least one strand of the genomic locus. In some aspects, the guide nucleic acid can be at least partially complementary to the donor strand or at least partially complementary to a genomic flap (e.g., a genomic nucleic acid sequence that is displaced and become single-stranded when the guide nucleic acid recruits the endonuclease to the genomic locus). In some aspects, the guide nucleic acid, being at least partially complementary to the donor strand or at least partially complementary to a genomic flap, brings the donor strand to close proximity of the cleaving of the genomic locus.
[0219] Disclosed herein, in some embodiments, are guide nucleic acids comprising a scaffold. The scaffold may bind a nuclease. The scaffold may bind a Cas nuclease. The scaffold may bind a nickase. The scaffold may bind a Cas nickase. The scaffold may bind an S. Pyogenes Cas9 nuclease. The scaffold may bind an S. Pyogenes Cas9 nickase. The scaffold may include a scaffold nucleic acid sequence. A system described herein may include a first guide nucleic acid. The system can include a second guide nucleic acid. The first guide nucleic acid may bind to a first Cas nickase. The second guide nucleic acid may bind to a second Cas nickase.
[0220] A guide nucleic acid may include any aspect of (i)-(iv): (i) a spacer complementary to a region of a genomic locus of a genomic strand, (ii) a scaffold for complexing with an RNA- guided endonuclease, (iii) a donor binding site that is at least partially complementary to an exogenous first integrating nucleic acid, or (iv) a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus. A guide nucleic acid may include any aspect of (i)-(iii): (i) a spacer complementary to a region of a genomic locus of a genomic strand, (ii) a scaffold for complexing with an RNA-guided endonuclease, or (iii) a donor binding site that is at least partially complementary to a splinting nucleic acid. A component of (i), (ii), or (iii) may be included in a single guide nucleic acid or may be split between or collectively included among multiple guide nucleic acids.
[0221] In some aspects, the guide nucleic acid comprises a modified intemucleoside linkage. In some aspects, the modified intemucleoside linkage comprises a phosphorothioate linkage. In some aspects, the modified intemucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the guide nucleic acid. The guide nucleic acid may include multiple modified intemucleoside linkages. For example, the guide nucleic acid may include modified intemucleoside linkages at nucleic acids of the 5’ and 3’ ends of the guide nucleic acid, such as between the last 4 nucleic acids at the 5’ end and between the last 4 nucleic acids at the 3’ end. In some aspects, the guide nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, or a combination thereof. The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. The modified nucleoside may include an LNA. The modified nucleoside may include a 2’fluoro. The modified nucleoside may include a 2’ O-alkyl. The modified nucleoside may include a methylated cytosine. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end ofthe guide nucleic acid. The guide nucleic acid may include multiple modified nucleosides. For example, the guide nucleic acid may include modified nucleosides at nucleic acids of the 5’ and 3’ ends of the guide nucleic acid, such as the last 3 nucleic acids at the 5’ end and the last 3 nucleic acids at the 3’ end.
[0222] In some aspects, the guide nucleic acid comprises at least one nucleic acid modification. In some aspect, the at least nucleic acid modification comprises modifying a backbone, a sugar, a base, or a combination thereof of the guide nucleic acid. In some aspects, the at least one nucleic acid modification can increase resistance of the guide nucleic acid to degradation (e.g., against nuclease degradation or hydrolysis). In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the at least one endonuclease. In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the donor strand. In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the genomic locus via by being complementary to the genomic flap.
[0223] In some aspects, the guide nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleic acid modifications. In some aspects, nucleic acid modification can occur at 3’OH group, 5’OH group, at the backbone, at the sugar component, or at the nucleotide base. Nucleic acid modification can include non-naturally occurring linker molecules of interstrand or intrastrand cross links. In one aspect, the modified nucleic acid comprises modification of one or more of the 3’OH or 5’OH group, the backbone, the sugar component, or the nucleotide base, or addition of non-naturally occurring linker molecules. In some aspects, modified backbone comprises a backbone other than a phosphodiester backbone. In some aspects, a modified sugar comprises a sugar other than deoxyribose (in modified DNA) or other than ribose (modified RNA). In some aspects, a modified base comprises a base other than adenine, guanine, cytosine, thymine, or uracil. In some aspects, the guide nucleic acid comprises at least one modified base. In some instances, the guide nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 15, 20, or more modified bases. In some cases, the nucleic acid modifications to the base moiety include natural and synthetic modifications of adenine, guanine, cytosine, thymine, or uracil, and purine or pyrimidine bases.
[0224] In some aspects, the at least one nucleic acid modification of the guide nucleic acid comprises a modification of any one of or any combination of: 2' modified nucleotide comprising 2'-O-methyl, 2'-O-methoxyethyl (2'-0-M0E), 2'-O-aminopropyl, 2'-deoxy, 2'-deoxy-2'-fluoro, 2'- O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-O-N-methylacetamido (2'-0-NMA); modification of one or both of the non-linking phosphate oxygens in the phosphodiester backbone linkage; modification of one or more of the linking phosphate oxygens in the phosphodiester backbone linkage; modification of a constituent of the ribose sugar; replacement of the phosphate moiety with “dephospho” linkers; modification or replacement of a naturally occurring nucleobase; modification of the ribose-phosphate backbone; modification of 5’ end of polynucleotide; modification of 3’ end of polynucleotide; modification of the deoxyribose phosphate backbone; substitution of the phosphate group; modification of the ribophosphate backbone; modifications to the sugar of a nucleotide; modifications to the base of a nucleotide; or stereopure of nucleotide. Non limiting examples of nucleic acid modification to the guide nucleic acid can include: modification of one or both of non-linking or linking phosphate oxygens in the phosphodi ester backbone linkage (e.g., sulfur (S), selenium (Se), BR3 (wherein R can be, e.g., hydrogen, alkyl, or aryl), C (e.g., an alkyl group, an aryl group, and the like), H, NR2, wherein R can be, e.g., hydrogen, alkyl, or aryl, or wherein R can be, e.g., alkyl or aryl); replacement of the phosphate moiety with “dephospho” linkers (e.g., replacement with methyl phosphonate, hydroxylamino, siloxane, carbonate, carb oxy methyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino); modification or replacement of a naturally occurring nucleobase with nucleic acid analog; modification of deoxyribose-phosphate or ribose-phosphate backbone (e.g., modifying the ribose-phosphate backbone to incorporate phosphorothioate, phosphonothioacetate, phosphoroselenates, boranophosphates, borano phosphate esters, hydrogen phosphonates, phosphonocarboxylate, phosphoroamidates, alkyl or aryl phosphonates, phosphonoacetate, or phosphotriesters; modification of 5’ end (e.g., 5’ cap or modification of 5’ cap -OH) or 3’ end of the nucleic acid sequence (3’ tail or modification of 3’ end -OH); substitution of the phosphate group with methyl phosphonate, hydroxyl ami no, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime,methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino; modification of the ribophosphate backbone to incorporate morpholino (phosphorodi ami date morpholino oligomer PMO), cyclobutyl, pyrrolidine, or peptide nucleic acid (PNA) nucleoside surrogates; modifications to the sugar of a nucleotide to incorporate locked nucleic acid (LNA), unlocked nucleic acid (UNA), ethylene nucleic acid (ENA), constrained ethyl (cEt) sugar, or bridged nucleic acid (BNA); modification of a constituent of the ribose sugar (e.g., 2’-O-methyl, 2’-O-methoxy-ethyl (2’-M0E), 2’-fluoro, 2’-aminoethyl, 2’- deoxy-2’-fuloarabinou-cleic acid, 2'-deoxy, 2'-O-methyl, 3'-phosphorothioate, 3'- phosphonoacetate (PACE), or 3 '-phosphonothioacetate (thioPACE)); modification to the base of a nucleotide (of A, T, C, G, or U); and stereopure of nucleotide (e.g., S conformation of phosphorothioate or R conformation of phosphorothioate).
[0225] The following codes may be used in sequences set forth herein.
[0226] In some aspects, the nucleic acid modification comprises at least one substitution of one or both of non-linking phosphate oxygen atoms in a phosphodiester backbone linkage of theguide nucleic acid. In some aspects, the at least one nucleic acid modification of the guide nucleic acid comprises a substitution of one or more of linking phosphate oxygen atoms in a phosphodiester backbone linkage of the guide nucleic acid. A non-limiting example of a nucleic acid modification of a phosphate oxygen atom is a sulfur atom. In some aspects, the nucleic acid modification comprises at least one modification to a sugar. In some aspects, the nucleic acid modification comprises at least one nucleic acid modification to the sugar comprising a modification of a constituent of the sugar, where the sugar is a ribose sugar. In some aspects, the nucleic acid modification of the guide nucleic acid comprises at least one modification to the constituent of the ribose sugar of the nucleotide of the guide nucleic acid comprising a 2’ -O-Methyl group. In some aspects, the nucleic acid modification comprises at least one modification comprising replacement of a phosphate moiety of the guide nucleic acid with a dephospho linker. In some aspects, the nucleic acid modification of comprises at least one modification of a phosphate backbone. In some aspects, the modification comprises a phosphorothioate group. In some aspects, the nucleic acid modifications comprises at least one modification comprising a modification to a base of a nucleotide of the guide nucleic acid. In some aspects, the nucleic acid modifications comprises at least one modification comprising an unnatural base of a nucleotide. In some aspects, the nucleic acid modifications comprises at least one modification comprising at least one stereopure nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 5’ end of the guide nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 3’ end of the guide nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to both 5’ and 3’ ends of the guide nucleic acid.
[0227] In some aspects, the guide nucleic acid described herein comprises a backbone comprising a plurality of sugar and phosphate moi eties covalently linked together. In some cases, a backbone of the guide nucleic acid comprises a phosphodiester bond linkage between a first hydroxyl group in a phosphate group on a 5’ carbon of a deoxyribose in DNA or ribose in RNA and a second hydroxyl group on a 3’ carbon of a deoxyribose in DNA or ribose in RNA. In some aspects, a backbone of the guide nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to a solvent. In some aspects, a backbone of the guide nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to nucleases. In some aspects, a backbone of the guide nucleic acid can lack a 5’ reducinghydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to hydrolytic enzymes. In some instances, a backbone of the guide nucleic acid can be represented as a polynucleotide sequence in a circular 2-dimensional format with one nucleotide after the other. In some instances, a backbone of the guide nucleic acid can be represented as a polynucleotide sequence in a looped 2-dimensional format with one nucleotide after the other. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are joined through a phosphorus-oxygen bond. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are modified into a phosphoester with a phosphorus-containing moiety. In some aspects, the guide nucleic acid comprises at least one nucleic acid modification comprising any one of: 5' adenylate, 5' guanosine-triphosphate cap, 5 'N7-Methyl guanosine-triphosphate cap, 5 'triphosphate cap, 3 'phosphate, 3 'thiophosphate, 5 'phosphate, 5 'thiophosphate, Cis-Syn thymidine dimer, trimers, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer 18, Spacer 9,3 '-3' modifications, 5'-5' modifications, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOT A, dT-Biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3 'DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linkers, 2'deoxyribonucleoside analog purine, 2 'deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methyl ribonucleoside analog, sugar modified analogs, wobble / universal bases, fluorescent dye label, 2'fluoro RNA, 2'0-methyl RNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphothioate DNA, phosphorothioate RNA, UNA, LNA, cEt, pseudouridine-5 '-triphosphate, 5-methylcytidine-5'- triphosphate, 2-O-methyl -phosphorothioate or any combinations thereof.
[0228] A nucleic acid modification can also be a phosphorothioate substitute. In some cases, a natural phosphodiester bond can be susceptible to rapid degradation by cellular nucleases and; a modification of intemucleotide linkage using phosphorothioate (PS) bond substitutes can be more stable towards hydrolysis by cellular degradation. A modification can increase stability in a polynucleic acid. A modification can also enhance biological activity. In some cases, a phosphorothioate enhanced RNA polynucleic acid can inhibit RNase A, RNase Tl, calf serum nucleases, or any combinations thereof. These properties can allow the use of PS-RNA polynucleic acids to be used in applications where exposure to nucleases is of high probability in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced between the last 3-5 nucleotides at the 5 '-or 3 '-end of a polynucleic acid which can inhibit exonuclease degradation. Insome cases, phosphorothioate bonds can be added throughout an entire polynucleic acid to reduce attack by endonucleases. In some aspects, the guide nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 100, or more intemucleotide linkage comprising PS bond. In some aspects, the guide nucleic acid comprises only PS bond as the intemucleotide linkage modification. In some aspects, all intemucleotide linkages of the guide nucleic acid herein are fully PS-modified or include phosphorothioate intemucleotide linkages.
[0229] The guide nucleic acid may include a hairpin. The hairpin may bind to a hairpin binding motif such as a hairpin binding motif on a DNA ligase. The hairpin may include an MS2 hairpin A hairpin such as an MS2 hairpin may be useful for recruiting a DNA ligase that includes an MCP peptide.
[0230] The guide nucleic acid may include any aspect included in Fig. 1A-6C. Table 11 illustrates non-limiting examples of some of the guide nucleic acids described herein. Some of the guide nucleic acids in the table include nucleic acid modifications.
[0231] The guide nucleic acid may include a sequence of linking nucleic acids (e.g., linking RNA or DNA nucleotides) between components of the guide nucleic acid. For example, the guidenucleic acid may include a sequence of linking nucleic acids between any of the following components: a spacer, a scaffold, a donor binding site, or a flap binding site. The guide nucleic acid may include a sequence of linking nucleic acids between a spacer, a scaffold, or a donor binding site. The guide nucleic acid include a sequence of linking nucleic acids between the scaffold and the donor binding site The guide nucleic acid may include a sequence of linking nucleic acids between a spacer and a scaffold. The guide nucleic acid may include multiple sequences of linking nucleic acids between components.
[0232] The sequence of linking nucleic acids may include any base, such as A, U, T, G, or C, or a combination thereof. The sequence of linking nucleic acids may include A, T, G, or C, or a combination thereof. The sequence of linking nucleic acids may include A, U, G, or C, or a combination thereof. The sequence of linking nucleic acids may include a series of As. The sequence of linking nucleic acids may include a series of Ts. The sequence of linking nucleic acids may include a series of Us. The sequence of linking nucleic acids may include a series of Cs. The sequence of linking nucleic acids may include a series of Gs.
[0233] The sequence of linking nucleic acids may include a length, such as a number of nucleotides. The length may include 1, 2, 3, 4, 5, 6, 7, 8, 9 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides, or a range defined by any two of the aforementioned numbers of nucleotides. The length may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides. In some aspects, the length may be less than 2, less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9 10, less than 11, less than 12, less than 13, less than 14, less than 15, less than 16, less than 17, less than 18, less than 19, less than 20, less than 21, less than 22, less than 23, less than 24, less than 25, less than 30, less than 35, less than 40, less than 45, less than 50, less than 55, less than 60, less than 65, less than 70, less than 75, less than 80, less than 85, less than 90, less than 95, or less than 100 nucleotides.
[0234] Some aspects relate to a guide nucleic acid comprising: a spacer that is at least partially complementary to a genomic locus in a cell; a scaffold for complexing with an RNA-guidedendonuclease; and a donor binding site that is at least partially complementary to an exogenous first integrating nucleic acid. The guide nucleic acid may further comprise a flap binding site that is at least partially complementary to a genomic sequence of the genomic locus. The guide nucleic acid may further comprise at least one nucleic acid modification. The at least one nucleic acid modification may comprise a modification to a backbone, a sugar, a base, or a combination thereof. The guide nucleic acid may comprise RNA.
[0235] Some aspects include a guide nucleic acid, comprising: a spacer at least partially reverse complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; and a flap binding site at least partially reverse complementary to a nucleic acid flap, and an exogenous first integrating nucleic acid binding site.Splint Nucleic Acids
[0236] Disclosed herein are splinting nucleic acids. The splinting nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate to a nucleic acid (e.g., DNA or RNA). The splinting nucleic acid can comprise a DNA or RNA backbone or “splint” which is reverse complementary to the DNA sequences that are to be ligated. Non-limiting examples include the Replacer guide nucleic acids depicted in Fig. 1A, 2A, 5 A, and 6A, which comprise a flap binding site (FBS) adjacent to a donor binding site (DBS). In some embodiments, a splinting nucleic acid can comprise a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, and can optionally comprise a guide binding site that is at least partially complementary to a guide nucleic acid. Non-limiting examples include Donor 2 nucleic acids depicted in Replacer 2 models of Fig. 3A and 4A which comprise a guide binding site (GBS) adjacent to a flap binding site (FBS) adjacent to a donor binding site DBS). In some aspects, the genomic strand is in a cell. In some aspects, the splinting nucleic acid further comprises a donor binding site that is at least partially identical or complementary to a portion of the integrating nucleic acid. In some aspects, the donor nucleic acid comprises a splinting nucleic acid. In some aspects, the guide nucleic acid comprises a splinting nucleic acid.
[0237] In some aspects, the splint nucleic acid comprises a modified intemucleoside linkage. In some aspects, the modified intemucleoside linkage comprises a phosphor othioate linkage. In some aspects, the modified intemucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the splint nucleic acid. The splint nucleic acid may include multiplemodified internucleoside linkages. For example, the splint nucleic acid may include modified intemucleoside linkages at nucleic acids of the 5’ and 3’ ends of the splint nucleic acid, such as between the last 4 nucleic acids at the 5’ end and between the last 4 nucleic acids at the 3’ end. In some aspects, the splint nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, or a combination thereof The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. The modified nucleoside may include an LNA. The modified nucleoside may include a 2’fluoro. The modified nucleoside may include a 2’ O-alkyl. The modified nucleoside may include a methylated cytosine. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end of the splint nucleic acid. The splint nucleic acid may include multiple modified nucleosides. For example, the splint nucleic acid may include modified nucleosides at nucleic acids of the 5’ and 3’ ends of the splint nucleic acid, such as the last 3 nucleic acids at the 5’ end and the last 3 nucleic acids at the 3’ end.
[0238] In some aspects, the splint nucleic acid comprises at least one nucleic acid modification. In some aspect, the at least nucleic acid modification comprises modifying a backbone, a sugar, a base, or a combination thereof of the splint nucleic acid. In some aspects, the at least one nucleic acid modification can increase resistance of the splint nucleic acid to degradation (e.g., against nuclease degradation or hydrolysis). In some aspects, the at least one nucleic acid modification can increase the complexing of the splint nucleic acid to the at least one endonuclease. In some aspects, the at least one nucleic acid modification can increase the complexing of the splint nucleic acid to the donor strand. In some aspects, the at least one nucleic acid modification can increase the complexing of the splint nucleic acid to the genomic locus via by being complementary to the genomic flap.
[0239] In some aspects, the splint nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleic acid modifications. In some aspects, nucleic acid modification can occur at 3’OH group, 5’OH group, at the backbone, at the sugar component, or at the nucleotide base. Nucleic acid modification can include non-naturally occurring linker molecules of interstrand or intrastrand cross links. In one aspect, the modified nucleic acid comprises modification of one or more of the 3’OH or 5’OHgroup, the backbone, the sugar component, or the nucleotide base, or addition of non-naturally occurring linker molecules. In some aspects, modified backbone comprises a backbone other than a phosphodiester backbone. In some aspects, a modified sugar comprises a sugar other than deoxyribose (in modified DNA) or other than ribose (modified RNA). In some aspects, a modified base comprises a base other than adenine, guanine, cytosine, thymine, or uracil. In some aspects, the splint nucleic acid comprises at least one modified base. In some instances, the splint nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 15, 20, or more modified bases. In some cases, the nucleic acid modifications to the base moiety include natural and synthetic modifications of adenine, guanine, cytosine, thymine, or uracil, and purine or pyrimidine bases.
[0240] In some aspects, the at least one nucleic acid modification of the splint nucleic acid comprises a modification of any one of or any combination of: 2' modified nucleotide comprising 2'-O-methyl, 2'-O-methoxyethyl (2'-0-M0E), 2'-O-aminopropyl, 2'-deoxy, 2'-deoxy-2'-fluoro, 2'- O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-0-DMA0E), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-O-N-methylacetamido (2'-0-NMA); modification of one or both of the non-linking phosphate oxygens in the phosphodiester backbone linkage; modification of one or more of the linking phosphate oxygens in the phosphodiester backbone linkage; modification of a constituent of the ribose sugar; replacement of the phosphate moiety with “dephospho” linkers; modification or replacement of a naturally occurring nucleobase; modification of the ribose-phosphate backbone; modification of 5’ end of polynucleotide; modification of 3’ end of polynucleotide; modification of the deoxyribose phosphate backbone; substitution of the phosphate group; modification of the ribophosphate backbone; modifications to the sugar of a nucleotide; modifications to the base of a nucleotide; or stereopure of nucleotide. Non limiting examples of nucleic acid modification to the splint nucleic acid can include: modification of one or both of non-linking or linking phosphate oxygens in the phosphodiester backbone linkage (e.g., sulfur (S), selenium (Se), BR3 (wherein R can be, e.g., hydrogen, alkyl, or aryl), C (e.g., an alkyl group, an aryl group, and the like), H, NR2, wherein R can be, e.g., hydrogen, alkyl, or aryl, or wherein R can be, e.g., alkyl or aryl); replacement of the phosphate moiety with “dephospho” linkers (e.g., replacement with methyl phosphonate, hydroxylamino, siloxane, carbonate, carb oxym ethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino,methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino); modification or replacement of a naturally occurring nucleobase with nucleic acid analog; modification of deoxyribose-phosphate or ribose-phosphate backbone (e.g., modifying the ribose-phosphate backbone to incorporate phosphorothioate, phosphonothioacetate, phosphoroselenates, boranophosphates, borano phosphate esters, hydrogen phosphonates, phosphonocarboxylate, phosphoroamidates, alkyl or aryl phosphonates, phosphonoacetate, or phosphotriesters; modification of 5’ end (e.g., 5’ cap or modification of 5’ cap -OH) or 3’ end of the nucleic acid sequence (3’ tail or modification of 3’ end -OH); substitution of the phosphate group with methyl phosphonate, hydroxyl amino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino; modification of the ribophosphate backbone to incorporate morpholino (phosphorodi ami date morpholino oligomer PMO), cyclobutyl, pyrrolidine, or peptide nucleic acid (PNA) nucleoside surrogates; modifications to the sugar of a nucleotide to incorporate locked nucleic acid (LNA), unlocked nucleic acid (UNA), ethylene nucleic acid (ENA), constrained ethyl (cEt) sugar, or bridged nucleic acid (BNA); modification of a constituent of the ribose sugar (e.g., 2’-O-methyl, 2’-O-methoxy-ethyl (2’-M0E), 2’-fluoro, 2’-aminoethyl, 2’- deoxy-2’-fuloarabinou-cleic acid, 2'-deoxy, 2'-O-methyl, 3 '-phosphorothioate, 3'- phosphonoacetate (PACE), or 3 '-phosphonothioacetate (thioPACE)); modification to the base of a nucleotide (of A, T, C, G, or U); and stereopure of nucleotide (e.g., S conformation of phosphorothioate or R conformation of phosphorothioate).
[0241] In some aspects, the nucleic acid modification comprises at least one substitution of one or both of non-linking phosphate oxygen atoms in a phosphodiester backbone linkage of the splint nucleic acid. In some aspects, the at least one nucleic acid modification of the splint nucleic acid comprises a substitution of one or more of linking phosphate oxygen atoms in a phosphodiester backbone linkage of the splint nucleic acid. A non-limiting example of a nucleic acid modification of a phosphate oxygen atom is a sulfur atom. In some aspects, the nucleic acid modification comprises at least one modification to a sugar. In some aspects, the nucleic acid modification comprises at least one nucleic acid modification to the sugar comprising a modification of a constituent of the sugar, where the sugar is a ribose sugar. In some aspects, the nucleic acid modification of the splint nucleic acid comprises at least one modification to theconstituent of the ribose sugar of the nucleotide of the splint nucleic acid comprising a 2’ -O-Methyl group. In some aspects, the nucleic acid modification comprises at least one modification comprising replacement of a phosphate moiety of the splint nucleic acid with a dephospho linker. In some aspects, the nucleic acid modification of comprises at least one modification of a phosphate backbone. In some aspects, the modification comprises a phosphorothioate group. In some aspects, the nucleic acid modifications comprises at least one modification comprising a modification to a base of a nucleotide of the splint nucleic acid. In some aspects, the nucleic acid modifications comprises at least one modification comprising an unnatural base of a nucleotide. In some aspects, the nucleic acid modifications comprises at least one modification comprising at least one stereopure nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 5’ end of the splint nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 3’ end of the splint nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to both 5’ and 3’ ends of the splint nucleic acid.
[0242] In some aspects, the splint nucleic acid described herein comprises a backbone comprising a plurality of sugar and phosphate moi eties covalently linked together. In some cases, a backbone of the splint nucleic acid comprises a phosphodiester bond linkage between a first hydroxyl group in a phosphate group on a 5’ carbon of a deoxyribose in DNA or ribose in RNA and a second hydroxyl group on a 3’ carbon of a deoxyribose in DNA or ribose in RNA. In some aspects, a backbone of the splint nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to a solvent. In some aspects, a backbone of the splint nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to nucleases. In some aspects, a backbone of the splint nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to hydrolytic enzymes. In some instances, a backbone of the splint nucleic acid can be represented as a polynucleotide sequence in a circular 2-dimensional format with one nucleotide after the other. In some instances, a backbone of the splint nucleic acid can be represented as a polynucleotide sequence in a looped 2-dimensional format with one nucleotide after the other. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are joined through a phosphorus-oxygen bond. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are modified into a phosphoester with a phosphorus-containing moiety. In some aspects, the splint nucleic acid comprises at least one nucleic acid modification comprisingany one of: 5' adenylate, 5' guanosine-triphosphate cap, 5 'N7-Methyl guanosine-triphosphate cap, 5 'triphosphate cap, 3 'phosphate, 3 'thiophosphate, 5 'phosphate, 5 'thiophosphate, Cis-Syn thymidine dimer, trimers, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer 18, Spacer 9,3 '-3' modifications, 5'-5' modifications, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-Biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3 'DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linkers, 2'deoxyribonucleoside analog purine, 2 'deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methyl ribonucleoside analog, sugar modified analogs, wobble / universal bases, fluorescent dye label, 2'fluoro RNA, 2'0-methyl RNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphothioate DNA, phosphorothioate RNA, UNA, LNA, cEt, pseudouridine-5 '-triphosphate, 5-methylcytidine-5'- triphosphate, 2-O-methyl -phosphorothioate or any combinations thereof.
[0243] A nucleic acid modification can also be a phosphorothioate substitute. In some cases, a natural phosphodiester bond can be susceptible to rapid degradation by cellular nucleases and; a modification of intemucleotide linkage using phosphorothioate (PS) bond substitutes can be more stable towards hydrolysis by cellular degradation. A modification can increase stability in a polynucleic acid. A modification can also enhance biological activity. In some cases, a phosphorothioate enhanced RNA polynucleic acid can inhibit RNase A, RNase Tl, calf serum nucleases, or any combinations thereof. These properties can allow the use of PS-RNA polynucleic acids to be used in applications where exposure to nucleases is of high probability in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced between the last 3-5 nucleotides at the 5 '-or 3 '-end of a polynucleic acid which can inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added throughout an entire polynucleic acid to reduce attack by endonucleases. In some aspects, the splint nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 100, or more intemucleotide linkage comprising PS bond. In some aspects, the splint nucleic acid comprises only PS bond as the intemucleotide linkage modification. In some aspects, all intemucleotide linkages of the splint nucleic acid herein are fully PS-modified or include phosphorothioate intemucleotide linkages.
[0244] The splint nucleic acid may include a hairpin. The hairpin may bind to a hairpin binding motif such as a hairpin binding motif on a DNA ligase. The hairpin may include an MS2 hairpin A hairpin such as an MS2 hairpin may be useful for recruiting a DNA ligase that includes an MCP peptide.
[0245] In some embodiments, the donor binding site (DBS) of the splint is from 12 nt to 50 nt in length, including 13, 14, 15,16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 36, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 nt length. In some embodiements, the DBS of the splint comprises one or more modified nucleotides. In some embodiments from 10% to 80% of the DBS nucleotides are modified. In some embodiments, alternating nucleotides are modified. In some embodiments, every third nucleotide is modified. In some embodiments, the DBS comprises alternating LNAs or comprises a region where every third nucleotide is an LNA. In some embodiments, the DBS comprises anternating 2’F nucleotides or comprises a region where every third nucleotide is 2’F modified. In some embodiments, the DBS comprises alternating 2’OMe nucleotides or comprises a region where every third nucleotide is 2’OMe modified.
[0246] In some embodiments, the guide binding site (GBS) of the splint comprises one or more modified nucleotides. In some embodiments from 10% to 80% of the GBS nucleotides are modified. In some embodiments, every third nucleotide is modified. In some embodiments, the GBS comprises alternating LNAs or comprises a region where every third nucleotide is an LNA. In some embodiments, the GBS comprises anternating 2’F nucleotides or comprises a region where every third nucleotide is 2’F modified. In some embodiments, the GBS comprises alternating 2’OMe nucleotides or comprises a region where every third nucleotide is 2’OMe modified.
[0247] Non-limiting examples of chemical modifications of the splinting nucleic acid and the nucleic acid sequences corresponding to these modifications are provided in Table 12, Table 25 and illustrated in Fig. 17.
[0248] In some embodiments, the chemical modification of the splint can be 3' C3 spacer. In some embodiments, the chemical modification of the splint can be 3' phosphorylation. In some embodiments, the chemical modification of the splint can be 3' phosphorylation. In some embodiments, the chemical modification of the splint can be 3' inverted T overhang. In someembodiments, the chemical modification of the splint can be 5' inverted T overhang. In some embodiments, the chemical modification of the splint can be 3' inverted T part of GBS.
[0249] In some embodiments, editing efficiency may be optimized. In some embodiments, nucleic acid duplex lengths and compositions may be adjusted. In some embodiments, a splint binding site (SBS) of a guide nucleic acid is optimized. In some embodiments, a guide binding site (GBS) of a splint is optimized. In some embodiments, a splint binding site of a donor nucleic acid is optimized. In some embodiments, the a donor binding site (DBS) of a splint nucleic acid is optimized. In some embodiments, it is preferred to avoid certain modifications of nucleic acids to be integrated at the target. In some embodiments, it is preferred to modify duplex forming nucleic acids that will not be integrated at the target. In some embodiments, it is preferred to modify a splinting nucleic acid that will not be integrated at the target. In some embodiments, it is preferred to modify a duplex -forming region of a nucleic acid that will not be integrated, such as duplexforming region of a splint nucleic acid or guide nucleic acid.
[0250] In some embodiments, editing efficiency can be optimized by adjusting the length and composition of a duplex formed by a gRNA SBS and a splint GBS. In some embodiements, editing efficiency can be optimized by a adjusting the length and composition of a duplex formed by a donor SBS and a splint DBS. In some embodiements, editing efficiency can be optimized by a adjusting the length and composition of a duplex formed by a gRNA SBS and a splint GBS.
[0251] In some embodiments, optimizing editing efficiency comprises modifying a splint GBS. In some embodiments, the length of the splint GBS is adjusted. In some embodiments, the length of the splint GBS-gRNA duplex is from 10 bp to 30 bp, or from 12 bp to 28 bp, or from 14 bp to 26 bp, or from 16 bp to 24 bp, or from 18 bp to 22 bp, or from 17 bp to 19 bp, or from 18 bp to 20 bp, or from 19 bp to 21 bp, or 17 bp, or 18 bp, or 19 bp, or 20 bo, or 21 bp, or 22 bp, or 23 bp. In some embodiments, the splint GBS comprises LNAs and the number or proportion of LNAs is adjusted. In some embodiments, the LNA nucleotides alternate with DNA nucleotides.
[0252] In some embodiments, optimizing editing efficiency comprises modifying a splint DBS. In some embodiments, the length of the splint DBS is adjusted. In some embodiments, the length of the splint GBS-donor duplex is from 18 bp to 40 bp, or from 19 bp to 36 bp, or from20 bp to 32 bp, or from 21 bp to 30 bp, or from 22 bp to 28 bp, or from 23 bp to 27 bp, or from24 bp to 26 bp, or from 18 bp to 22 bp, or from 20 bp to 24 bp, or from 22 bp to 28 bp, or from24 bp to 30 bp, or from 26 bp to 32 bp, or from 28 bp to 34 bp, or 18 bp, or 19 bp, or 20 bo, or21 bp, or 22 bp, or 23 bp, or 24 bp, or 25 bp, or 26 bo, or 27 bp, or 28 bp, or 29 bp, or 30 bp, or 31 bo, or 32 bp, or 33 bp, or 34 bp, or 35 bp, or 36 bp, or 37 bp, or 38 bp, or 39 bp, or 40 bp. In some embodiments, the splint DBS comprises LNAs and the number or proportion of LNAs is adjusted. In some embodiments, the LNA nucleotides alternate with DNA nucleotides.
[0253] In some embodiments, optimizing editing efficiency comprises modifying a splint FBS. In some embodiments, the length of the splint FBS is adjusted. In some embodiments, the length of the splint FBS-flap duplex is from 10 bp to 30 bp, or from 12 bp to 28 bp, or from 14 bp to 26 bp, or from 16 bp to 24 bp, or from 18 bp to 22 bp, or from 17 bp to 19 bp, or from 18 bp to 20 bp, or from 19 bp to 21 bp, or 17 bp, or 18 bp, or 19 bp, or 20 bo, or 21 bp, or 22 bp, or 23 bp. In some embodiments, the splint DBS comprises LNAs and the number or proportion of LNAs is adjusted. In some embodiments, the LNA nucleotides alternate with DNA nucleotides.
[0254] The following Table 12 provides non-limiting examples of splints employed herein.Target nucleic acids
[0255] Disclosed herein are target nucleic acids. The target nucleic acid may include DNA. The target nucleic acid may be DNA. The target nucleic acid may include RNA. The target nucleic acid may be in a cell. The target nucleic acid may be methylated. The target nucleic acid may be unmethylated. The target nucleic acid may comprise a genome. The target nucleic acid may comprise genomic DNA. The target nucleic acid may comprise a chromosome. The target nucleic acid may comprise a gene.
[0256] The target nucleic acid may be in a subject. The target nucleic acid may be in a cell. The target nucleic acid may be in a test tube.
[0257] The target nucleic acid may be edited. The target nucleic acid may be edited in vitro. The target nucleic acid may be edited in vivo.Integrating nucleic acidsExogenous First Integrating Nucleic Acids
[0258] Disclosed herein are exogenous first integrating nucleic acids. The exogenous first integrating nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate to a nucleic acid that encodes an exogenous first integrating nucleic acid. Provided herein are exogenous first integrating nucleic acids that are inserted into a target nucleic acid such as a host genome at a genetic locus. For example, the exogenous first integrating nucleic acid may replace a nucleic acid in the target nucleic acid. The exogenous first integrating nucleic acid may be referred to as a “donor nucleic acid,” “donor” or “donor strand.” Where a genomic locus is described, a genetic locus may be included, or vice versa. For example, the locus may be part of a host genome or may be a part of a non-genome nucleic acid. The donor may include DNA. Likewise, the target nucleic acid may include DNA. In some cases, the donor may include RNA, for example when a target nucleic acid includes RNA. The donor may include any insert, such as a gene or a regulatory element, to be inserted at a genomic locus of a target nucleic acid. The donor strand may include a sequence that is at least partially homologous to the genomic locus.The donor may, in some instances, also act as a splint for a DNA ligase described herein, such as for ligating two nucleic acid strands base paired to a portion of the splinting exogenous first integrating nucleic acid. In some cases, the splint includes one strand of the donor, and the portion being ligated may be another strand of the donor. In some cases, the splint includes a strand of the donor, and the portion being ligated may be an upstream or downstream portion of the same strand of the donor. The donor may be single stranded. The donor may be double stranded. The donor may be delivered as two strands. The donor may be delivered as multiple strands, e.g., 2 strands.
[0259] The exogenous first integrating nucleic acid may be non-naturally occurring. The donor may be engineered. The donor may be synthetic. The donor may be pre-synthetized. The donor may be added to a subject or a cell. In some aspects, the donor does not include a template for a polymerase.
[0260] Disclosed herein are exogenous first integrating nucleic acids, comprising: a doublestranded DNA region to be inserted into a target nucleic acid, wherein the double-stranded DNA region is flanked by at least one overhang comprising a flap binding site and / or guide binding site.
[0261] The exogenous first integrating nucleic acid may be ligated into a target nucleic acid such as a genomic strand. The exogenous first integrating nucleic acid may include a 5’ end that may be ligated to a 3’ terminus of a genomic strand generated by an RNA-guided endonuclease.
[0262] The donor may include any aspect included in Fig. 1A-6C. For example, the donor may include an aspect such as a guide binding site, a flap binding site, or an overhang. The donor may include a guide binding site. The donor may include 2 guide binding sites. The donor may include a flap binding site. The donor may include 2 flap binding sites. The donor may include an overhang. The donor may include 2 overhangs. The aspects may be included at a 5’ end or a 3’ end of the donor, or at both ends. A guide binding site or a flap binding site may be in an internal region of the donor.
[0263] Some aspects include an exogenous first integrating nucleic acid, comprising: a doublestranded DNA region to be inserted into a target nucleic acid, wherein the double-stranded DNA region is flanked by at least one overhang comprising a flap binding site or guide binding site.
[0264] In some aspects, the exogenous first integrating nucleic acid comprises a modified intemucleoside linkage. In some aspects, the modified intemucleoside linkage comprises a phosphorothioate linkage. In some aspects, the modified intemucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the donor nucleic acid. The exogenous firstintegrating nucleic acid may include multiple modified internucleoside linkages. For example, the exogenous first integrating nucleic acid may include modified internucleoside linkages at nucleic acids of the 5’ and 3’ ends of the exogenous first integrating nucleic acid, such as between the last 4 nucleic acids at the 5’ end and between the last 4 nucleic acids at the 3’ end. In some aspects, the exogenous first integrating nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, a 5’ O- methyl, a 2’-O-methyl, or a combination thereof. The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. The modified nucleoside may include an LNA. The modified nucleoside may include a 2’fluoro. The modified nucleoside may include a 2’ O-alkyl. The modified nucleoside may include a methylated cytosine. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end of the exogenous first integrating nucleic acid. The exogenous first integrating nucleic acid may include multiple modified nucleosides. For example, the exogenous first integrating nucleic acid may include modified nucleosides at nucleic acids of the 5’ and 3’ ends of the exogenous first integrating nucleic acid, such as the last 3 nucleic acids at the 5’ end and the last 3 nucleic acids at the 3’ end. The exogenous first integrating nucleic acid may include any modification such as a modified nucleoside or modified internucleoside linkage described in relation to guide nucleic acids, insofar as it does not interfere with the function of the exogenous first integrating nucleic acid after it is ligated into a target nucleic acid such as a host genome. The exogenous first integrating nucleic acid may include any number or combination of modifications such as a number or combination described in relation to guide nucleic acids, insofar as it does not interfere with a function of the exogenous first integrating nucleic acid. Table 11 and Table 13 include some examples of exogenous first integrating nucleic acid sequences.
[0265] The donor nucleic acid may include a methylated nucleotide. The donor nucleic acid may include an unmethylated nucleotide. An example of a methylated nucleotide may include a nucleotide including methylated cytosine. The cytosine may be methylated at a C-5 position of the cytosine ring. An example of an unmethylated nucleotide may include an unmethylated cytosine. The unmethylated nucleotide may include a cytosine that is not methylated at a C-5 position of the cytosine ring.
[0266] In some embodiments, the donor nucleic acid can comprise a modified nucleotide. In some embodiments, the donor nucleic acid can comprise a modified DNA base. In someembodiments, the modified nucleotide comprises methylation. In some embodiments, the donor nucleic acid comprises a methylated nucleoside. Non-limiting example of the methylated nucleoside can include methylated cytosine (5-mC). In some embodiments, the donor nucleic acid can comprise N6-methyladenine (6-mA). In some embodiments, the modified DNA base can comprise 5 -hydroxymethyl cytosine (5-hmC). In some embodiments, the modified DNA base can comprise 5 -formyl cytosine (5-fC). In some embodiments, the modified DNA base can comprise 5-carboxylcytosine (5-caC).
[0267] In some embodiments, the donor can introduce an epigenetic modification into the target nucleic acid. In some embodiments, the donor can introduce an epigenetic modification of a DNA base (e.g., a methylated nucleoside) into the target nucleic acid. In some embodiments, the donor can introduce a methylated DNA base into the target nucleic acid (e.g., a methylated nucleoside). In some embodiments, the donor can introduce an unmethylated DNA base into the target nucleic acid (e.g., a methylated nucleoside). In some embodiments, the integrating of the donor can introduce a methoylated nucleoside into the genome. In some embodiments, the integrating of the donor can remove a methylated nucleoside from the genome. In some embodiments, the donor can introduce a methylated cytosine (e.g. 5-mC) into the target nucleic acid. In some embodiments, the donor can introduce an N6-methyladenine (6-mA) into the target nucleic acid. In some embodiments, the donor can introduce a 5 -hydroxymethyl cytosine (5-hmC) into the target nucleic acid. In some embodiments, the donor can introduce a 5-formylcytosine (5- fC) into the target nucleic acid. In some embodiments, the donor can introduce a 5- carboxylcytosine (5-caC) into the target nucleic acid.Recombination Sequences
[0268] Disclosed herein are exogenous first integrating nucleic acids. In some aspects, the exogenous first integrating nucleic acid may comprise a recombination sequence. A recombination sequence may also be referred to as an “integrating recombination sequence,” an “integration sequence,” a “recombination site,” an “integration site,” a “site-specific recombination sequence,” or a “site- specific integration sequence.”
[0269] In some aspects, the exogenous first integrating nucleic acid may comprise a plurality of recombination sequences. In some aspects, the exogenous first integrating nucleic acid may comprise at least 1 recombination sequence. In some aspects, the donor nucleic acid may comprise at least 2 recombination sequences. In some aspects, the donor nucleic acid may comprise at least3 recombination sequences. In some aspects, the donor nucleic acid may comprise at least 4 recombination sequences. In some aspects, the donor nucleic acid may comprise at least 5 recombination sequences. In some aspects, the donor nucleic acid may comprise at least 10 recombination sequences.
[0270] In some aspects, the exogenous first integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by an integrase. In some aspects, the exogenous first nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a recombinase. In some aspects, the exogenous first nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by any integrase described herein. In some aspects, the exogenous first nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by any integrase described in Table 8. The donor nucleic acid described herein may comprise a nucleic acid sequence that is recognized or bound by an integrase that comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8.
[0271] In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a serine recombinase. The serine recombinase may be or comprise a PhiC31 (OC31) bacteriophage integrase, a Bxbl mycobacteriophage integrase, a Pseudomonas aeruginosa (PaOl) integrase, a Nocctrdia otitidiscaviarum (No67) integrase, or a Streptomyces ipomoeae (Si74) integrase. In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a OC31 bacteriophage integrase. In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Bxbl mycobacteriophage integrase. The donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Pseudomonas aeruginosa (PaOl) integrase. The donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Nocardia otitidiscaviarum (No67) integrase. The donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Streptomyces ipomoeae (Si74) integrase.
[0272] In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a tyrosine recombinase. In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Cre recombinase.In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a flippase (Flp).
[0273] In some aspects, the exogenous first integrating nucleic acid described herein comprises any one of the sequences in Table 13. In some aspects, the donor nucleic acid described herein comprises a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the nucleotide sequence of any of the recombination sequences in Table 13.
[0274] In some aspects, the exogenous first integrating nucleic acid described herein comprises any one of the sequences in Table 14. In some aspects, the donor nucleic acid described herein comprises a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the nucleotide sequence of any of the recombination sequences in Table 14.
[0275] In some aspects, the exogenous first integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a serine recombinase. In some aspects, the donor nucleic acid may comprise an attachment site (att). In some aspects, the donor nucleic acid may comprise a plurality of attachment sites. In some aspects, the donor nucleic acid may comprise an attachment site on the bacterial part (allB) integrating recombination sequence. In some aspects, the donor nucleic acid may comprise two attachment site on the bacterial part (attB) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise a plurality of attachment site on the bacterial part (attB) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the donor nucleic acid may comprise two attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise a plurality of attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise an attachment site on the bacterial part (attB) integrating recombination sequence and an attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise a plurality of attachment site on the bacterial part (attB) integrating recombination sequences and a plurality of attachment site on the phage part (attP) integrating recombination sequences.
[0276] In some aspects, the exogenous first integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a tyrosine recombinase. In some aspects,the donor nucleic acid may comprise a locus of X(cross)-over in Pl (LoxP) sequence. In some aspects, the donor nucleic acid may comprise two locus of X(cross)-over in Pl (LoxP) sequences. In some aspects, the donor nucleic acid may comprise a plurality of locus of X(cross)-over in Pl (LoxP) sequences. In some aspects, the donor nucleic acid may comprise a flippase recognition target (FRT) sequence. In some aspects, the donor nucleic acid may comprise two flippase recognition target (FRT) sequences. In some aspects, the donor nucleic acid may comprise a plurality of flippase recognition target (FRT sequences.
[0277] In some embodiments, the donor nucleic acid comprises a chemical modification. In some embodiments, the donor nucleic acid can comprise a 3' C3 spacer. In some embodiments, the donor nucleic acid can comprise a 3' inverted dT. In some embodiments, the donor nucleic acid can comprise a 3' phosphorylation. In some embodiments, the donor nucleic acid can comprise a 3' phosphorothioate bond. In some embodiments, a chemical modification on the donor nucleic acid can protect the donor DNA against nucleases. In some embodiments, the chemical modification is not integrated into the genome. In some embodiments, the chemical modification can be integrated into the genome.Second Integrating Nucleic Acids
[0278] Disclosed herein are second integrating nucleic acids. The second integrating nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate to a nucleic acid that encodes a second integrating nucleic acid. Provided herein are second integrating nucleic acids that are inserted into a target nucleic acid such as a host genome at a genetic locus such as an integrating recombination sequence. For example, the second integrating nucleic acid may replace in whole or in part an integrating recombination sequence in the target nucleic acid. The second integrating nucleic acid may be referred to as a “revising nucleic acid.” Where a genomic locus is described, a genetic locus may be included, or vice versa. For example, the locus may be part of a host genome or may be a part of a non-genome nucleic acid. The revising nucleic acid may include DNA. Likewise, the target nucleic acid may include DNA. In some cases, the revising nucleic acid may include RNA, for example when a target nucleic acid includes RNA. The revising nucleic acid may include any insert, such as a gene or a regulatory element, to be integrated at a genomic locus of a target nucleic acid such as a recombination sequence. The revising nucleic acid may be integrated, in whole or in part, into a genomic locus of a target nucleic acid such as a recombination sequence by an integrase. The revising nucleic acid may include asequence that is at least partially homologous to the genomic locus. The revising nucleic acid may be single stranded. The revising nucleic acid may be double stranded. The revising nucleic acid may be delivered as two strands. The revising nucleic acid may be delivered as multiple strands, e.g., 2 strands.
[0279] The second integrating nucleic acid may be non-naturally occurring. The revising nucleic acid may be engineered. The revising nucleic acid may be synthetic. The revising nucleic acid may be pre-synthetized. The revising nucleic acid may be added to a subject or a cell. In some aspects, the revising nucleic acid does not include a template for a polymerase.
[0280] In some aspects, the revising nucleic acid comprises a recombination sequence. The recombination sequence may be a cognate of any recombination sequence of the donor nucleic acid. In some aspects, the cognate recombination sequence of the revising nucleic acid couples with a recombination sequence of the donor nucleic acid.
[0281] In some aspects, the revising nucleic acid comprises a plurality of recombination sequences. In some aspects, the revising nucleic acid comprises a plurality of cognate recombination sequences. In some aspects, the revising integrating nucleic acid may comprise at least 1 recombination sequence. In some aspects, the revising nucleic acid may comprise at least 2 recombination sequences. In some aspects, the revising nucleic acid may comprise at least 3 recombination sequences. In some aspects, the revising nucleic acid may comprise at least 4 recombination sequences. In some aspects, the revising nucleic acid may comprise at least 5 recombination sequences. In some aspects, the revising nucleic acid may comprise at least 10 recombination sequences.
[0282] In some aspects, the second integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by an integrase. In some aspects, the second nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a recombinase. In some aspects, the revising nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by any integrase described herein. In some aspects, the revising nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by any integrase described in Table 8. The revising nucleic acid described herein may comprise a nucleic acid sequence that is recognized or bound by an integrase that comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8.
[0283] In some aspects, the revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a serine recombinase. The serine recombinase may be or comprise a PhiC31 ( C31) bacteriophage integrase, a Bxbl mycobacteriophage integrase, a Pseudomonas aeruginosa (PaOl) integrase, a Nocardia otitidiscaviarum (No67) integrase, or a Streptomyces ipomoeae (Si74) integrase. In some aspects, the revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a C31 bacteriophage integrase. In some aspects, the revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Bxbl mycobacteriophage integrase. The revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Pseudomonas aeruginosa (PaOl) integrase. The revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Nocardia otitidiscaviarum (No67) integrase. The revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Streptomyces ipomoeae (Si 74) integrase.
[0284] In some aspects, the revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a tyrosine recombinase. In some aspects, the revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Cre recombinase. In some aspects, the revising nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a flippase (Flp).
[0285] In some aspects, the revising nucleic acid described herein comprises any one of the sequences in Table 13. In some aspects, the revising nucleic acid described herein comprises a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the nucleotide sequence of any of the recombination sequences in Table 13.
[0286] In some aspects, the second integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a serine recombinase. In some aspects, the revising nucleic acid may comprise an attachment site (att). In some aspects, the revising nucleic acid may comprise a plurality of attachment sites. In some aspects, the revising nucleic acid may comprise an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the revising nucleic acid may comprise two attachment site on the bacterial part (az / B) integrating recombination sequences. In some aspects, the revising nucleic acid may comprise a plurality of attachment site on the bacterial part (a / zB) integrating recombination sequences. In some aspects,the revising nucleic acid may comprise an attachment site on the phage part (czZZP) integrating recombination sequence. In some aspects, the revising nucleic acid may comprise two attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the revising nucleic acid may comprise a plurality of attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the revising nucleic acid may comprise an attachment site on the bacterial part (at / B) integrating recombination sequence and an attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the revising nucleic acid may comprise a plurality of attachment site on the bacterial part (attB) integrating recombination sequences and a plurality of attachment site on the phage part attP) integrating recombination sequences.
[0287] In some aspects, the second integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a tyrosine recombinase. In some aspects, the revising nucleic acid may comprise a locus of X(cross)-over in Pl (LoxP) sequence. In some aspects, the revising nucleic acid may comprise two locus of X(cross)-over in Pl (LoxP) sequences. In some aspects, the revising nucleic acid may comprise a plurality of locus of X(cross)- over in Pl (LoxP) sequences. In some aspects, the revising nucleic acid may comprise a flippase recognition target (FRT) sequence. In some aspects, the revising nucleic acid may comprise two flippase recognition target (FRT) sequences. In some aspects, the revising nucleic acid may comprise a plurality of flippase recognition target (FRT) sequences.
[0288] In some aspects, the revising nucleic acid may comprise an exogenous nucleic acid sequence. In some aspects the exogenous nucleic acid sequence may be integrated in whole or in part into a genome at a target location such as an integration sequence by an integrase. In some aspects, the revising nucleic acid may further comprise regulatory elements. The regulatory elements may be promoters or enhancers. In some aspects, the revising nucleic acid may not comprise a promoter or an enhancer.
[0289] In some aspects, the revising nucleic acid may comprise a nucleotide sequence that encodes a gene. In some aspects, the revising nucleic acid may comprise a nucleotide sequence that encodes a plurality of genes. The at revising nucleic acid may comprise a coding sequence. The coding sequence may encode a full-length protein. The revising nucleic acid may comprise a non-coding sequence The non-coding sequence may be a recombination sequence. The non-coding sequence may knock out an endogenous gene. The non-coding sequence may comprise aregulatory element. The revising nucleic acid may be up to about 10 kb in length. The revising nucleic acid may be up to about 20 kb in length. The revising nucleic acid may be up to about 30 kb in length. The revising nucleic acid may be up to about 40 kb in length. The revising nucleic acid may be up to about 50 kb in length. The revising nucleic acid may be at least about 50 kb in length.
[0290] In some aspects, the revising nucleic acid may be delivered as a minicircle (Fig. 15A). In some aspects, the revising nucleic acid may be delivered as a plasmid (Fig. 15B and Fig. 15C), In some aspects, the revising nucleic acid may be delivered as linearized double strand DNA (Fig.15D)[00291J The revising nucleic acid may include a methylated nucleotide. The revising nucleic acid may include an unmethylated nucleotide. An example of a methylated nucleotide may include a nucleotide including methylated cytosine. The cytosine may be methylated at a C-5 position of the cytosine ring. An example of an unmethylated nucleotide may include an unmethylated cytosine. The unmethylated nucleotide may include a cytosine that is not methylated at a C-5 position of the cytosine ring.
[0292] In some embodiments, the revising nucleic acid can comprise a modified nucleotide. In some embodiments, the revising nucleic acid can comprise a modified DNA base. In some embodiments, the modified nucleotide comprises methylation. In some embodiments, the revising nucleic acid comprises a methylated nucleotide. Non-limiting example of the methylated nucleotide can include methylated cytosine (e.g. 5-mC). In some embodiments, the revising nucleic acid can comprise N6-methyladenine (6-mA). In some embodiments, the modified DNA base can comprise 5 -hydroxymethyl cytosine (5-hmC). In some embodiments, the modified DNA base can comprise 5-formylcytosine (5-fC). In some embodiments, the modified DNA base can comprise 5 -carboxyl cytosine (5-caC).
[0293] In some embodiments, the revising nucleic acid can introduce an epigenetic modification into the target nucleic acid. In some embodiments, the revising nucleic acid can introduce an epigenetic modification of a DNA base into the target nucleic acid. In some embodiments, the revising nucleic acid can introduce a methylated DNA base into the target nucleic acid. In some embodiments, the revising nucleic acid can introduce an unmethylated DNA base into the target nucleic acid. In some embodiments, the revising nucleic acid can introduce a methylated cytosine (e.g. 5-mC) into the target nucleic acid. In some embodiments, the revisingnucleic acid can introduce an N6-methyladenine (6-mA) into the target nucleic acid. In some embodiments, the revising nucleic acid can introduce a 5-hydroxymethylcytosine (5-hmC) into the target nucleic acid. In some embodiments, the revising nucleic acid can introduce a 5- formylcytosine (5-fC) into the target nucleic acid. In some embodiments, the revising nucleic acid can introduce a 5-carboxyl cytosine (5-caC) into the target nucleic acid.Systems
[0294] Described herein are systems for nucleic acid editing (also known as gene editing). The editing system may include an endonuclease such as an RNA-guided endonuclease, a guide nucleic acid, and a donor nucleic acid. Where gene editing is described, it is contemplated that the editing may be of a gene, regulatory element, or any sequence of a nucleic acid. It is contemplated that editing may be an integration of a nucleic acid, such as a nucleic acid sequence, into a gene, regulatory element, or any sequence of a nucleic acid. Also, where genome editing is described, such as genome editing at a genetic locus, it is contemplated that editing of a nucleic acid not comprising a genome may also be performed. For example, genome editing may refer to editing of a genome of an organism or may include editing of a nucleic acid that is not part of a genome. The systems described herein may be used in gene editing methods. Where nucleic acid integration is described, such as integration of a nucleic acid sequence at a genetic locus, it is contemplated that a nucleic acid sequence may be integrated into a second nucleic acid sequence, the second nucleic acid not comprising a genome. For example, integration may refer to integration of a nucleic acid sequence into the genome of an organism or may refer to integration of a nucleic acid sequence into a second nucleic acid sequence that is not part of a genome.
[0295] Described herein, in some aspects, is a system comprising at least one endonuclease; at least one guide nucleic acid; at least one ligase; at least one donor strand; at least one integrase; at least one exogenous first integrating nucleic acid; or a combination thereof. In some aspects, the guide nucleic acid directs the endonuclease to the genomic locus for cleaving at least one strand of the genomic locus, where, after cleavage, the donor strand is ligated and thus incorporated into the genomic locus by the ligase. In some aspects, the system comprises: a first endonuclease to be complexed with a first guide nucleic acid, where the first endonuclease can be operatively coupled to a first ligase; and a second endonuclease to be complexed with a second guide nucleic acid, where the second endonuclease can be operatively coupled to a second ligase. In such system eachof the first endonuclease and the second endonuclease can each cleave at least one strand of the genomic locus for incorporation of the donor strand.
[0296] In some aspects, the system comprises one, two, three, or more endonucleases. In some aspects, the system comprises one endonucleases. In some aspects, the two endonucleases can each be complexed with a different guide nucleic acid. In some aspects, the two endonucleases can each be operatively coupled to a ligase. In some aspects, the two endonucleases can each be operatively coupled to an integrase. In some aspects, the endonuclease is a programmable endonuclease. In some aspects, the endonuclease comprises an RNA-guided endonuclease, where the guide nucleic acid comprises a guide RNA. In some aspects, the endonuclease comprises a nickase, where the endonuclease only cleaves one strand (as opposed to making a double-stranded break). In some aspects, the endonuclease comprises a localization signal sequence to increase the accumulation of the endonuclease in the proximity of the genomic locus (e.g., in the nucleus). In some aspects, the endonuclease comprises at least one additional domain. In some aspects, the at least one additional domain is a dimerization domain. In some aspects, the endonuclease comprising a dimerization domain can be dimerized with a ligase to form a heterodimer. In some aspects, the at least one additional domain is a functional domain. For example, the functional domain can comprise a chromatin modifying domain or a cell penetrating peptide. In some aspects, the endonuclease comprises a linker, where the linker can covalently connect the endonuclease with another polypeptide (e.g., the ligase or the integrase). In some aspects, the linker covalently connects the endonuclease to the at least one additional domain. In some aspects, the endonuclease comprises a tag, where the tag can be used for increasing expression, identifying, or purifying the endonuclease.
[0297] In some aspects, the system comprises one, two, three, or more guide nucleic acids. In some aspects, the system comprises one guide nucleic acid, where the one guide nucleic acid can be complexed with at least one endonuclease. In some aspects, the system comprises two guide nucleic acids, where the two guide nucleic acids can each be complexed with the at least one endonuclease. In some aspects, the guide nucleic acid comprises a spacer complementary to a genomic locus in a cell; a scaffold for complexing with the at least one endonuclease; a donor binding site that is at least partially complementary to a donor strand; a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; or a combination thereof. In some aspects, the guide nucleic acid can direct the at least oneendonuclease to cleave at least one strand of the genomic locus. In some aspects, the guide nucleic acid can be at least partially complementary to the donor strand or at least partially complementary to a genomic flap (e.g., a genomic nucleic acid sequence that is displaced and becomes singlestranded when the guide nucleic acid recruits the endonuclease to the genomic locus). In some aspects, the guide nucleic acid, being at least partially complementary to the donor strand or at least partially complementary to a genomic flap, brings the donor strand to close proximity of the cleaving of the genomic locus. In some aspects, the guide nucleic acid comprises a splint for ligation. In some aspects, the guide nucleic acid comprises at least one nucleic acid modification. In some aspects, the at least nucleic acid modification comprises modifying a backbone, a sugar, a base, or a combination thereof of the guide nucleic acid. In some aspects, the at least one nucleic acid modification can increase resistance of the guide nucleic acid to degradation (e.g., against nuclease degradation or hydrolysis). In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the at least one endonuclease. In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the donor strand. In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the genomic locus via by being complementary to the genomic flap.
[0298] In some aspects, the system comprises one, two, three, or more ligases. In some aspects, the system comprises one ligase. In some aspects, the one ligase is operatively coupled with at least one endonuclease, at least one integrase, or both, where the ligase can ligate at least one end of the donor strand to the cleaved genomic locus, thus incorporating the donor strand into the genomic locus. In some aspects, the system comprises two ligases. In some aspects, the two ligases can each be operatively coupled to a different endonuclease, a different integrase, or each can be operatively coupled to a different endonuclease and to a different integrase, where the genomic locus is cleaved at two or more locations. In such scenario, the two ligases can each ligate one end of the donor strand to the cleaved genomic locus, thus incorporating the donor strand into the genomic locus. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising DNA. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising a DNA splint. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising a DNA / RNA. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising an RNA splint. In some aspects, the ligase comprises a ligase that can ligate a substrate comprisinga gRNA splint. In some aspects, the ligase comprises at least one additional domain. In some aspects, the at least one additional domain is a dimerization domain. In some aspects, the ligase comprising a dimerization domain can be dimerized with an endonuclease to form a heterodimer. In some aspects, the at least one additional domain is a functional domain. For example, the functional domain can comprise a chromatin modifying domain or a cell penetrating peptide. In some aspects, the ligase comprises a linker, where the linker can covalently connect the ligase with another polypeptide (e.g., the endonuclease). In some aspects, the linker covalently connects the ligase to the at least one additional domain. In some aspects, the ligase comprises a tag, where the tag can be used for increasing expression, identifying, or purifying the ligase.[00299J Disclosed herein are fusion proteins comprising: an RNA-guided endonuclease fused to a ligase. Table 15 illustrates non-limiting examples of polypeptide and nucleic acid sequences encoding a fusion polypeptide comprising components (e.g., an endonuclease fused to a ligase) of a system described herein. SEQ ID NO: 125 illustrates a nucleic acid sequence encoding the polypeptide sequence of SEQ ID NO: 126, where SEQ ID NO: 126 illustrates a fusion protein (NLS-nCas9-linker-hLIGl(119-919)-bpNLS) comprising a N-terminus NLS followed by an endonuclease (nCas9) covalently connected to a ligase (hLIGl, 119-919 fragment) via a linker followed by a C-terminus NLS. SEQ ID NO: 127 illustrates a nucleic acid sequence encoding the polypeptide sequence of SEQ ID NO: 128, where SEQ ID NO: 128 illustrates a fusion protein (NLS-nCas9-linker-hLIGl(233-919)-bpNLS) comprising a N-terminus NLS followed by an endonuclease (nCas9) covalently connected to a ligase (hLIGl, 233-919 fragment) via a linker followed by a C-terminus NLS. SEQ ID NO: 129 illustrates a nucleic acid sequence encoding the polypeptide sequence of SEQ ID: 130, where SEQ ID NO: 130 illustrates a fusion protein (NLS- nCas9-linker-SplintR-bpNLS) comprising a N-terminus NLS followed by an endonuclease (nCas9) covalently connected to a ligase (SplintR) via a linker followed by a C-terminus NLS. SEQ ID NO: 131 illustrates a nucleic acid sequence encoding the polypeptide sequence of SEQ ID NO: 132, where SEQ ID NO: 132 illustrates a fusion protein (NLS-nCas9-linker-T4LIG-bpNLS) comprising a N-terminus NLS followed by an endonuclease (nCas9) covalently connected to a ligase (T4LIG) via a linker followed by a C-terminus NLS. SEQ ID NO: 133 illustrates a nucleic acid sequence encoding an endonuclease (nCas9) comprising a N-terminus NLS and a leucine zipper (LZ) dimerization domain. SEQ ID NO: 134 illustrates a fusion protein (NLSl-hFENl- linkerl-nCas9-linker2-T4LIG-NLS2) comprising first NLS (NLS1) at N-terminus followed by anexonuclease (hFENl) covalently connected to an endonuclease (nCas9) via linkerl and further covalently connected to a ligase (T4LIG) via linker 2 followed by a second NLS (NLS2) at C- terminus. SEQ ID NO: 135 illustrates a fusion protein (NLSl-hFENl-linkerl-T4LIG-linker2- nCas9-NLS2) comprising a N-terminus NLS1 followed by an exonuclease (hFENl) covalently connected to a ligase (T4LIG) via linker 1 and further covalently connected to an endonuclease (nCas9) via linker 2 followed by a C-terminus NLS2. SEQ ID NO: 136 illustrates a fusion protein (NLS l-nCas9-linkerl -hFENl -Iinker2-T4LIG-NLS2) comprising a N-terminus NLS 1 followed by an endonuclease (nCas9) covalently connected to an exonuclease (hFENl) via linker 1 and further covalently connected to a ligase (T4LIG) via linker 2 followed by a C-terminus NLS2. SEQ ID NO: 137 illustrates a fusion protein (NLSl-T4LIG-linkerl-nCas9-linker2-hFENl-NLS2) comprising a N-terminus NLS1 followed by a ligase (T4LIG) covalently connected to an endonuclease (nCas9) via linker 1 and further covalently connected to an exonuclease (hFENl) via linker 2 followed by a C-terminus NLS2. SEQ ID NO: 138 illustrates a fusion protein (NLS1 - nCas9-linkerl-T4LIG-linker2-hFENl-NLS2) comprising a N-terminus NLS1 followed by an endonuclease (nCas9) covalently connected to a ligase (T4LIG) via linker 1 and further covalently connected to an exonuclease (hFENl) via linker 2 followed by a C-terminus NLS2. SEQ ID NO:139 illustrates a fusion protein (NLS l-T4LIG-linkerl -hFENl -linker2-nCas9-NLS2) comprising a N-terminus NLS1 followed by a ligase (T4LIG) covalently connected to an exonuclease (hFENl) via linker 1 and further covalently connected to an endonuclease (nCas9) via linker 2 followed by a C-terminus NLS2. SEQ ID NO: 140 illustrates a fusion protein (NLS1- T5 EXO-linkerl-nCas9-linker2-T4LIG-NLS2) comprising a N-terminus NLS1 followed by an exonuclease (EXO) covalently connected to an endonuclease (nCas9) via linker 1 and further covalently connected to a ligase (T4LIG) via linker 2 followed by a C-terminus NLS2. SEQ ID NO: 141 illustrates a nucleic acid sequence encoding a fusion protein (LZ-SplintR-bpNLS) comprising a ligase (SplintR) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 142 illustrates a nucleic acid sequence encoding a fusion protein (LZ-T4LIG-bpNLS) comprising a ligase (T4LIG) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 143 illustrates a nucleic acid sequence encoding a fusion protein (LZ-hLIG 233-919 polypeptide fragment-bpNLS) comprising a ligase (hLIG) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 144 illustrates a nucleic acid sequence encoding a fusion protein (LZ-hLIGl 119-919 polypeptide fragment-bpNLS) comprising a ligase (hLIG) fused to a dimerization domain (LZ) and an NLS.SEQ ID NO: 145 illustrates a nucleic acid sequence encoding a fusion protein (T4-LZ) comprising a ligase (T4) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 146 illustrates a nucleic acid sequence encoding a fusion protein (LZ-hLIG4( 1-620)) comprising a ligase polypeptide fragment (hLIG4( 1-620)) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 147 illustrates a nucleic acid sequence encoding a fusion protein (LZ-nCas9) comprising an endonuclease (nCas9) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 148 illustrates a nucleic acid sequence encoding a fusion protein (SplintR-LZ) comprising a ligase (SplintR) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 149 illustrates a nucleic acid sequence encoding a fusion protein (hLIG4(l-620)-LZ) comprising a ligase polypeptide fragment (hL!G4( 1-620)) fused to a dimerization domain (LZ) and an NLS. SEQ ID NO: 150 illustrates a nucleic acid sequence encoding a fusion protein (nCas9-hLIG4(l-620)) comprising a ligase polypeptide fragment (hLIG4( 1-620)) fused to an endonuclease (nCas9) and an NLS. SEQ ID NO: 151 illustrates a nucleic acid sequence encoding a fusion protein (T4-nCas9) comprising a ligase (T4) fused to an endonuclease (nCas9) and an NLS. SEQ ID NO: 152 illustrates a nucleic acid sequence encoding a fusion protein (SplintR-nCas9) comprising a ligase (SplintR) fused to an endonuclease (nCas9) and an NLS. SEQ ID NO: 153 illustrates a nucleic acid sequence encoding a fusion protein (hLIG4(l-620)-nCas9) comprising a ligase polypeptide fragment (hLIG4( 1-620)) fused to an endonuclease (nCas9) and an NLS.ATTCAAGTACAAGTCTGCCTTTATGCGTTTGACCTGATCTATCTTAATGGAGAGAGTTTGGTGAGAGAACCCTTGAGCAGACGACGGCAGCTCTTGAGAGAAAATTTCGTAGAAACTGAGGGGGAGTTCGTCTTTGCGACTAGTCTCGACACCAAAGACATTGAGCAAATCGCGGAATTCCTCGAACAGTCAGTTAAAGACTCCTGCGAAGGTCTGATGGTTAAGACTCTTGACGTGGATGCTACCTACGAGATAGCTAAGCGGTCACACAATTGGCTGAAACTGAAAAAGGACTATCTGGATGGAGTTGGGGACACGCTGGATTTGGTCGTTATCGGGGCCTATCTGGGACGCGGTAAGCGGGCAGGGAGATATGGTGGATTCCTCCTCGCTTCATACGATGAGGACTCTGAAGAGCTGCAGGCTATATGCAAACTTGGGACGGGTTTTTCCGATGAAGAATTGGAGGAACATCATCAGTCACTGAAGGCCCTTGTATTGCCAAGTCCACGCCCATACGTACGAATCGATGGAGCAGTAATCCCTGACCACTGGCTTGACCCGTCCGCCGTCTGGGAAGTAAAGTGCGCGGATCTCTCTCTCAGTCCGATCTACCCAGCCGCACGGGGGCTGGTTGACAGTGACAAGGGTATCAGCCTGCGATTTCCTCGATTCATACGCGTCCGGGAAGACAAGCAACCGGAACAGGCTACGACCTCTGCACAGGTCGCATGTTTGTATAGAAAACAGAGCCAAATTCAGAATCAACAAGGCGAAGACAGTGGGTCCGATCCTGAAGATACCTACTCAGGCGGCAGTAAACGGACAGCTGATAGCCAACACTCAACTCCTCCGAAGACTAAAAGGAAGGTAGAGTTCGAACCAAAAAAGAAAAGGAAAGTGTAALZ- ATGCTCGAGATCGAGGCGGCGTTCCTTGAACGCGAGAACACTGC hLIGl(l GCTGGAAACGAGGGTCGCGGAACTCCGCCAGAGGGTTCAACGG 19-919)- TTGAGGAATCGAGTGAGTCAGTACCGAACCCGATATGGACCACT bpNLS GGGTGGCGGGAAATCAGGGGGCTCATCCGGCGGCTCCAGCGGG AGCGAAACCCCGGGTACCTCAGAATCTGCGACGCCAGAAAGCTC AGGCGGATCTAGCGGCGGTAGTTCACCGAAGCGCCGGACTGCAC GAAAGCAACTGCCAAAACGGACTATACAAGAAGTCCTGGAAGA ACAAAGCGAAGATGAGGATCGCGAAGCCAAGCGCAAGAAAGAG GAAGAGGAAGAAGAGACTCCAAAGGAGTCCTTGACCGAAGCAG AAGTCGCAACGGAGAAGGAAGGTGAGGATGGGGATCAGCCAAC AACCCCGCCTAAACCTCTGAAAACCTCTAAGGCGGAGACACCAA CTGAGAGTGTCAGCGAACCGGAGGTAGCCACGAAACAAGAGCT TCAGGAGGAAGAAGAACAGACAAAGCCACCTCGGCGGGCTCCC AAAACCCTTAGCTCCTTCTTCACGCCTCGAAAGCCAGCAGTGAA GAAAGAAGTGAAGGAGGAGGAACCTGGCGCCCCTGGAAAGGAG GGCGCAGCCGAGGGCCCGCTGGACCCTTCAGGGTATAACCCGGC AAAAAATAATTACCACCCGGTCGAGGACGCTTGTTGGAAACCAGGCCAAAAGGTACCTTACCTCGCCGTCGCTAGGACCTTTGAGAAG ATAGAGGAAGTTAGTGCTAGGTTGAGAATGGTCGAAACCCTTAG TAACCTTCTCAGGTCCGTAGTCGCCCTTAGTCCCCCAGACCTGCT TCCGGTGCTGTACCTGTCCCTGAACCATCTCGGTCCCCCCCAACA GGGACTGGAGTTGGGCGTCGGTGACGGCGTTCTCCTGAAAGCGG TTGCACAAGCTACAGGAAGGCAACTGGAATCTGTCCGGGCTGAG GCTGCAGAGAAAGGTGACGTGGGGCTTGTGGCAGAGAATAGTC GGTCAACACAGCGGCTGATGCTGCCACCGCCCCCGCTTACGGCT
[0300] Disclosed herein are protein complexes comprising: an RNA-guided endonuclease bound to a ligase. The endonuclease and the ligase may be bound together through heterodimerization domains. The heterodimerization domains may include one or more of leucine zippers, PDZ domains, streptavidin, and streptavidin binding protein, foldon domains, hydrophobic polypeptides, an antibody that binds the Cas nickase, or an antibody that binds the ligase, or one or more binding fragments thereof.
[0301] Disclosed herein are protein complexes comprising: an RNA-guided endonuclease bound to an integrase. The endonuclease and the integrase may be bound together through heterodimerization domains. The heterodimerization domains may include one or more of leucine zippers, PDZ domains, streptavidin, and streptavidin binding protein, foldon domains,hydrophobic polypeptides, an antibody that binds the Cas nickase, or an antibody that binds the integrase, or one or more binding fragments thereof.
[0302] Disclosed herein are protein complexes comprising: a ligase bound to an integrase. The endonuclease and the integrase may be bound together through heterodimerization domains. The heterodimerization domains may include one or more of leucine zippers, PDZ domains, streptavidin, and streptavidin binding protein, foldon domains, hydrophobic polypeptides, an antibody that binds the ligase, or an antibody that binds the integrase, or one or more binding fragments thereof.
[0303] Disclosed herein are protein complexes comprising: an RNA-guided endonuclease bound to a ligase and to an integrase; a ligase bound to an RNA-guided endonuclease and an integrase; and an integrase bound to an RNA-guided endonuclease and a ligase.
[0304] In some aspects, the system comprises at least one donor strand. In some aspects, the donor strand comprises a nucleic acid sequence that is at least partially homologous to the genomic locus targeted by the at least one guide nucleic acid. In some aspects, the donor strand comprises a nucleic acid sequence that is not homologous to the genomic locus targeted by the at least one guide nucleic acid. In some aspects, the donor strand is a single-stranded or a double-stranded nucleic acid. In some aspects, the donor strand comprising double-stranded nucleic acid comprises at least one overhang. In some aspects, the overhang comprises a guide binding site that is at least partially complementary to a guide nucleic acid. In some aspects, the overhang comprises a genomic flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus. In some aspects, the donor strand comprises two overhangs, where the first overhang: comprises a first guide binding site that is at least partially complementary to a first guide nucleic acid; or a first genomic flap binding site that is at least partially identical or complementary to a first genomic flap at or adjacent to the genomic locus; and the second overhang: comprises a second guide binding site that is at least partially complementary to a second guide nucleic acid; or a second genomic flap binding site that is at least partially identical or complementary to a second genomic flap at or adjacent to the genomic locus. In some aspects, the donor strand corrects at least one genetic mutation in the at least one genomic locus. In some aspects, the donor strand comprises a coding sequence. In some aspects, the coding sequence encodes a full-length protein or a fragment thereof. In some aspects, the donor strand comprises a non-coding sequence. In some aspects, the non-coding sequence comprises arecombination sequence. In some aspects, the non-coding sequence knocks out an endogenous gene. In some aspects, the non-coding sequence comprises a regulatory element.
[0305] In some aspects, the system comprises a nuclease. The nuclease may be heterologous. In some aspects, the nuclease comprises an exonuclease for digesting the genomic flap. In some aspects, the exonuclease is a 5’ exonuclease. Non-limiting example of the exonuclease can include a human flap endonuclease 1 (hFENl), a human exonuclease 5 (hEXO5), a T5 exonuclease, a T7 exonuclease, an exonuclease VIII, a flap endonuclease domain of E. coli Poll, a RecJF, a Lambda exonuclease, a Xni (ExoIXI), a SaFEN (Staphylococcus aureus FEN), a nuclease BAL-31, or a fragment thereof. In some aspects, the exonuclease comprises an exonuclease in Table 16. In some aspects, the exonuclease comprises a polypeptide sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the exonuclease in Table 16.
[0306] In some aspects, the system comprises at least one additional endonuclease that is different from the at least one programmable endonuclease described herein. In some aspects, the at least one additional endonuclease can digest the genomic flap.[00307J In some aspects, the system comprises a dominant negative MMR peptide to improve genomic editing capability, particularly in cells which overexpress the MMR pathway. In some aspects, the dominant negative MMR peptide can be delivered as a fusion (e g., fused with any component of the system described herein), recruited, or as separate peptide. Table 17 lists nonlimiting examples of the MMR peptide sequences.
[0308] The system may relate to a 1 -sided Replacer 1. Some aspects include a system comprising: (a) at least one RNA-guided endonuclease; (b) at least one guide nucleic acid comprising: (i) a spacer complementary to a genomic locus in a cell, (ii) a scaffold for complexing with the at least one RNA-guided endonuclease, (iii) an optional donor binding site that is at least partially complementary to an exogenous first integrating nucleic acid, and (iv) a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; and (c) at least one DNA ligase; and (d) the exogenous first integrating nucleic acid, optionally comprising a guide binding site that is at least partially complementary to the at least one guide nucleic acid, wherein the at least one RNA-guided endonuclease cleaves at least one strand of the genomic locus, and wherein the at least one DNA ligase ligates an end of the exogenous first integrating nucleic acid to the genomic flap site, thereby replacing a region of thegenomic locus with the exogenous first integrating nucleic acid in the cell. The exogenous first integrating nucleic acid may comprise a single- stranded DNA.
[0309] The system may relate to a 2-sided Replacer 1. Some aspects include a system comprising: (a) at least one RNA-guided endonuclease comprising a first RNA-guided endonuclease and an optional second RNA-guided endonuclease; (b) at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid comprising: (i) a first spacer complementary to a first region of a genomic locus in a cell, (ii) a first scaffold for complexing with the first RNA-guided endonuclease, and (iii) an optional first donor binding site that at least partially complementary to an exogenous first integrating nucleic acid, and (iv) a first flap binding site that is at least partially identical or complementary to a first genomic flap at or adjacent to the genomic locus; and the second guide nucleic acid comprising: (i) a second spacer complementary to a second region of the genomic locus in the cell, (ii) a second scaffold for complexing with the first or second RNA-guided endonuclease, (iii) an optional second donor binding site that at least partially complementary to the exogenous first integrating nucleic acid, and (iv) a second flap binding site that is at least partially identical or complementary to a second genomic flap at or adjacent to the genomic locus; (c) at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; and (d) at least one exogenous first integrating nucleic acid comprising a first strand and a second strand: (i) wherein the first strand comprises an optional first guide binding site that is at least partially complementary to the first guide nucleic acid, and (ii) wherein the second strand comprises an optional second guide binding site that is at least partially complementary to the second guide nucleic acid, wherein the first RNA-guided endonuclease and / or the second RNA-guided endonuclease each cleaves at least one strand of the genomic locus in the cell; and wherein the first DNA ligase ligates an end of the first strand of the exogenous first integrating nucleic acid to the first genomic flap; and the first or second DNA ligase ligates an end of the second strand of the exogenous first integrating nucleic acid to the second genomic flap, thereby replacing a region of the genomic locus with the exogenous first integrating nucleic acid in the cell. The exogenous first integrating nucleic acid may comprise a double-stranded DNA duplex region. The exogenous first integrating nucleic acid may comprise a 5’ overhang optionally comprising the first guide binding site. The exogenous first integrating nucleic acid may comprise a 5’ overhang optionally comprising the second guide binding site.
[0310] The system may relate to 1 -sided Replacer 2. Some aspects include a system comprising: (a) at least one RNA-guided endonuclease; (b) at least one guide nucleic acid comprising: (i) a spacer complementary to a genomic locus in a cell, (ii) a scaffold for complexing with the at least one RNA-guided endonuclease, and (iii) an optional donor binding site that is at least partially complementary to an exogenous first integrating nucleic acid; (c) at least one DNA ligase; and (d) the exogenous first integrating nucleic acid that: (i) comprises an optional guide binding site that is at least partially complementary to the at least one guide nucleic acid, and (ii) comprises a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, wherein the at least one RNA-guided endonuclease cleaves at least one strand of the genomic locus; and wherein the at least one DNA ligase ligates an end of the exogenous first integrating nucleic acid to the genomic flap, thereby replacing a region of the genomic locus with the exogenous first integrating nucleic acid in the cell. The exogenous first integrating nucleic acid may comprise a DNA comprising a 3’ overhang. The 3’ overhang may comprise the guide binding site. The 3’ overhang may comprise the flap binding site. The at least one DNA ligase may ligate a strand of the exogenous first integrating nucleic acid to the genomic nucleic acid sequence.
[0311] The system may relate to 2-sided Replacer 2. Some aspects include a system comprising: (a) at least one RNA-guided endonuclease comprising a first RNA-guided endonuclease and an optional second RNA-guided endonuclease; (b) at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid comprising: (i) a first spacer complementary to a first region of a genomic locus in a cell, (ii) a first scaffold for complexing with the first RNA-guided endonuclease, and (iii) an optional first donor binding site that at least partially complementary to an exogenous first integrating nucleic acid; and the second guide nucleic acid comprising: (i) a second spacer complementary to a second region of the genomic locus in the cell, (ii) a second scaffold for complexing with the first or second RNA-guided endonuclease, and (iii) an optional second donor binding site that at least partially complementary to the exogenous first integrating nucleic acid; and at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; and the exogenous first integrating nucleic acid comprising a first strand and a second strand: wherein the first strand comprises an optional first guide binding site that is at least partially complementary to the first guide nucleic acid; wherein the second strand comprises an optional second binding site that is atleast partially complementary to the second guide nucleic acid; wherein the first strand comprises a first flap binding site that is at least partially identical or complementary to a first genomic flap at or adjacent to the genomic locus; and wherein the second strand comprises a second flap binding site that is at least partially identical or complementary to a second genomic flap at or adjacent to the genomic locus; wherein the first RNA-guided endonuclease and / or the second RNA-guided endonuclease each cleaves at least one strand of the genomic locus in the cell; and wherein the first DNA ligase ligates an end of the first strand of the exogenous first integrating nucleic acid to the first genomic flap; and the first or second DNA ligase ligates an end of the second strand of the exogenous first integrating nucleic acid to the second genomic flap, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. The exogenous first integrating nucleic acid may comprise a double-stranded DNA duplex region. The double-stranded DNA may comprise a 3’ overhang optionally comprising the first guide binding site and comprising the first flap binding site. The double stranded DNA may comprise a 3’ overhang optionally comprising the second guide binding site and comprising the second flap binding site.
[0312] In the system, the at least one RNA-guided endonuclease may comprise a Cas protein or a functional fragment thereof. The Cas protein or the functional fragment thereof may comprise nickase activity The at least one RNA-guided endonuclease may comprise a Cas9 nickase or a functional fragment thereof. The at least one DNA ligase may ligate nucleic acids bound to DNA. The at least one DNA ligase may ligate nucleic acids bound to RNA. The at least one DNA ligase may comprise a PBCV-1 DNA ligase. The at least one DNA ligase may be operatively coupled to the at least one RNA-guided endonuclease. The at least one DNA ligase may be fused to the at least one RNA-guided endonuclease as a fusion polypeptide. The at least one RNA-guided endonuclease and the at least one DNA ligase may comprise a heterodimer domain. The at least one RNA-guided endonuclease and the at least one DNA ligase may form a heterodimer via the heterodimer domain. The at least one RNA-guided endonuclease may comprise a linker. The linker may connect the Cas protein or a functional fragment thereof to the heterodimer domain. The at least one RNA-guided endonuclease may comprise a localization signal sequence. The at least one DNA ligase may comprise a localization signal sequence. The localization signal sequence may comprise a nuclear localization sequence (NLS). The at least one RNA-guided endonuclease or the at least one DNA ligase may be directed to nucleus of the cell by the NLS. The at least one integrating nucleic acid, such as a donor nucleic acid or a revising nucleic acid, may correct atleast one genetic mutation in the at least one genomic locus. The at least one integrating nucleic acid may insert a coding sequence. The coding sequence may encode a full-length protein. The at least one integrating nucleic acid may insert a non-coding sequence The non-coding sequence may comprise a recombination sequence. The non-coding sequence may knock out an endogenous gene. The non-coding sequence may comprise a regulatory element. The system may further include a nuclease. The nuclease may comprise an exonuclease for digesting the genomic flap. The nuclease may comprise a human flap endonuclease 1 (hFENl), a human exonuclease 5 (hEXO5), a T5 exonuclease, a T7 exonuclease, an exonuclease VIII, a flap endonuclease domain of E. coli Poll, a RecJF, a Lambda exonuclease, a Xni (ExoIXI), a SaFEN (Staphylococcus aureus FEN), a nuclease BAL-31, or a fragment thereof. The heterologous nuclease may comprise an endonuclease for digesting the genomic flap, and the endonuclease may be different from the at least one RNA-guided endonuclease. The at least one RNA-guided endonuclease may comprise at least one additional functional domain. The at least one additional functional domain may comprise a chromatin modifying domain. The at least one additional functional domain may comprise a cell penetrating peptide. The at least one guide nucleic acid may comprise at least one nucleic acid modification. The at least one nucleic acid modification may comprise a modification to a backbone, a sugar, a base, or a combination thereof. The at least one RNA-guided endonuclease may be complexed with the at least one guide nucleic acid. The at least one guide nucleic acid may be complexed with the exogenous first integrating nucleic acid. The at least one RNA-guided endonuclease, the at least one guide nucleic acid, the at least one at least one DNA ligase, the exogenous first integrating nucleic acid, or a combination thereof may be encoded by a polynucleotide. The polynucleotide may comprise mRNA. The polynucleotide may comprise a vector. The vector may comprise a viral vector. The at least one RNA-guided endonuclease, the at least one guide nucleic acid, the at least one at least one DNA ligase, the exogenous first integrating nucleic acid, or a combination thereof may be encapsulated by at least one lipid nanoparticle. The cell may comprise a bacterial cell or a prokaryotic cell. The cell may include a prokaryotic cell. The prokaryotic cell may include a bacterial cell. The editing may be performed in a cytoplasm of the bacterial cell. The cell may include a eukaryotic cell. The eukaryotic cell may include an animal cell or a plant cell. The eukaryotic cell may include a plant cell. The eukaryotic cell may include an animal cell. The eukaryotic cell may comprise a mammalian cell. The editing may be performed in a cytoplasm of the eukaryotic cell. The editing may be performed in a nucleus of the eukaryoticcell. The system, or any aspect of the system, may be included in a composition, or in a cell such as a cell line.
[0313] Some aspects relate to a system that includes nucleic acids. The system may include guide nucleic acids, integrating nucleic acids, or a combination thereof. Some aspects relate to a system of nucleic acids. The system may include a system of guide nucleic acids. The system may include a system of integrating nucleic acids. The system of nucleic acids may further include other aspects such as additional nucleic acids or non-nucleic acid components.
[0314] The system of nucleic acids may include a guide nucleic acid. The guide nucleic acid may include a spacer. The spacer may be complementary to a region of a locus (e.g., genomic locus) of a target nucleic acid such as a genomic strand. The target nucleic acid may be in a cell. The genomic strand may be in a cell. The target nucleic acid may be in vitro. The guide nucleic acid may include a scaffold. The scaffold may complex with an endonuclease such as an RNA- guided endonuclease. The guide nucleic acid may include a flap binding site. The flap binding site may be complementary or at least partially complementary to a flap such as a genomic flap. The flap binding site may be identical or at least partially identical to a flap such as a genomic flap. The flap may be at the locus. The flap may be adjacent to the locus. The guide nucleic acid may include a donor binding site. The donor binding site may be complementary to an integrating nucleic acid. The donor binding site may be partially complementary to an integrating nucleic acid. The donor binding site may be complementary to a splinting nucleic acid. The donor binding site may be partially complementary to a splinting nucleic acid. Components of the guide nucleic acid may be included in 1 guide nucleic acid. More than one guide nucleic acid may be used. Components of the guide nucleic acid may collectively be included among multiple guide nucleic acids. Components of the guide nucleic acid may split between multiple guide nucleic acids.
[0315] The system of nucleic acids may include an exogenous first integrating nucleic acid. The exogenous first integrating nucleic acid may include a 5’ end to be ligated. The 5’ end may be ligated. The 5’ end may be ligated to a 3’ terminus. The 3’ terminus may be of a target nucleic acid strand (e.g., a genomic strand). The 3’ terminus may be generated by an endonuclease such as an RNA-guided endonuclease. The exogenous first integrating nucleic acid may include a 5’ end to be ligated to a 3’ terminus of a genomic strand generated by an RNA-guided endonuclease. Components of the exogenous first integrating nucleic acid may be included in 1 or 2 complementary strands. Components of the exogenous first integrating nucleic acid may beincluded in 1 donor nucleic acid. More than one integrating nucleic acid may be used. Components of the donor nucleic acid may collectively be included among multiple donor nucleic acids. Components of the exogenous first integrating nucleic acid may split between multiple integrating nucleic acids.
[0316] The system of nucleic acids may include a splinting nucleic acid (also referred to as a “splinting strand”). The splinting strand may hybridize to two nucleic acids comprising ends to be ligated. The splinting nucleic acid may include a flap binding site. The flap binding site may be complementary to a flap. The flap binding site may be partially complementary to a flap. The flap binding site may be identical to a flap. The flap binding site may be partially identical to a flap. The flap may be at a locus of a target nucleic acid. The flap may be adjacent to a locus of a target nucleic acid. The flap may be a genomic flap. The locus may be a genomic locus. The flap binding site may be at least partially identical or complementary to a genomic flap at or adjacent to a genomic locus. The splinting nucleic acid may include a guide binding site. The guide binding site may be complementary to a guide nucleic acid. The guide binding site may be partially complementary to a guide nucleic acid. Components of the splinting nucleic acid may be included in 1 splinting nucleic acid. More than one splinting nucleic acid may be used. The splinting nucleic acid may include a donor binding site. The donor binding site may be complementary to a donor nucleic acid. The donor binding site may be partially complementary to a donor nucleic acid.
[0317] The splinting strand may be or include DNA. The splinting strand may be or include RNA. The splinting nucleic acid may be included as part of an exogenous first integrating nucleic acid. The splinting nucleic acid may be included as a strand of a double stranded exogenous first integrating nucleic acid. The splinting nucleic acid may be included as part of a guide nucleic acid.
[0318] The system of nucleic acids may include: (a) a guide nucleic acid comprising: (i) a spacer complementary to a region of a genomic locus of a genomic strand, (ii) a scaffold for complexing with RNA-guided endonuclease, (iii) an optional donor binding site that is at least partially complementary to an exogenous first integrating nucleic acid, and (iv) a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; and (b) an exogenous first integrating nucleic acid comprising a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by an RNA-guided endonuclease. A component of (i), (ii), (iii), or (iv) may be included in a single guide nucleic acid or may be split between or collectively included among multiple guide nucleic acids.
[0319] The system of nucleic acids may include: (a) a guide nucleic acid comprising (i) a spacer complementary to a region of a genomic locus of a genomic strand, (ii) a scaffold for complexing with an RNA-guided endonuclease, and (iii) an optional donor binding site that is at least partially complementary to a splinting nucleic acid; (b) an exogenous first integrating nucleic acid comprising a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by an RNA- guided endonuclease; and (c) a splinting nucleic acid comprising a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, and comprising an optional guide binding site that is at least partially complementary to a guide nucleic acid. A component of (i), (ii), or (iii) may be included in a single guide nucleic acid or may be split between or collectively included among multiple guide nucleic acids.
[0320] In some aspects, the system described herein can be delivered into a cell, where one or more of the components of the system can be delivered into the cell together. In some aspects, each component of the system can be delivered into the cell separately. In some aspects, the system can be encoded by a polynucleotide such as a heterologous polynucleotide, where the polynucleotide is delivered into a cell and where the polynucleotide is expressed by the cell to generate the components of the cell. In some aspects, the system can be encoded and delivered into the cell via a polynucleotide comprising mRNA. In some aspects, the system can be encoded and delivered into the cell via a polynucleotide comprising a vector. In some aspects, the vector comprises a viral vector. The system can be encapsulated in a lipid or nanoparticle, or multiple lipids or nanoparticles. In some aspects, the system can be encapsulated in at least one lipid nanoparticle. In some aspects, the system comprises a ribonucleoprotein (RNP). For example, at least one RNA-guided endonuclease described herein (e.g., a Cas9) can be complexed with at least one guide nucleic acid described herein (e.g., forming a CRISPR ribonucleoprotein) for delivery. In some aspects, the system comprises at least one RNP comprising an RNA-guided endonuclease complexed with at least one first guide nucleic acid or with at least one second guide n...
Claims
WHAT IS CLAIMED:
1. A system, comprising: an endonuclease; a DNA ligase; an integrating nucleic acid, wherein the integrating nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid; and a splint nucleic acid which is configured to locate the 5’ end of the integrating nucleic acid to the 3’ end of the strand break in the target nucleic acid generated by the endonuclease.
2. The system of claim 1, wherein the system comprises a guide RNA and the endonuclease comprises an RNA-guided endonuclease.
3. The system of claim 1, wherein the endonuclease is coupled to the DNA ligase.
4. The system of claim 3, wherein the coupling is non-covalent.
5. The system of claim 4, wherein the endonuclease coupled to the DNA ligase comprises a first polypeptide comprising at least part of the endonuclease, and a second polypeptide comprising at least part of the DNA ligase, wherein the first and second polypeptides are non-covalently coupled.
6. The system of claim 5, wherein the first polypeptide comprises a first heterodimerization domain that binds a second heterodimerization domain, and wherein the second polypeptide comprises the second heterodimerization domain.
7. The system of claim 6 wherein the heterodimerization domains comprise a leucine zipper, PDZ domain, streptavidin, streptavidin binding protein, foldon domain, hydrophobic moiety, or a functional binding fragment thereof.
8. The system of claim 3, wherein the coupling is covalent.
9. The system of claim 4, comprising a fusion protein comprising the endonuclease and the DNA ligase.
10. The system of claim 9, wherein the endonuclease is amino (N)-terminal relative to the DNA ligase within the fusion protein.
11. The system of claim 9, wherein the endonuclease is carboxy (C)-terminal relative to the DNA ligase within the fusion protein.
12. The system of claim 9, wherein the fusion protein comprises a linker comprising 1-100 amino acids connecting the endonuclease and the DNA ligase.
13. The system of claim 7, wherein the first polypeptide comprises a first intein that binds a second intein, and wherein the second polypeptide comprises the second intein.
14. The system of claim 3, wherein the ligase comprises a hairpin binding motif, and wherein the endonuclease and the DNA ligase are coupled with a nucleic acid comprising a scaffold that binds to the endonuclease and a hairpin that binds to the hairpin binding motif.
15. The system of claim 14, wherein the hairpin binding motif comprises an MS2 coat protein (MCP) peptide, and wherein the hairpin comprises an MS2 hairpin.
16. The system of claim 3, wherein the endonuclease and the DNA ligase are coupled with a heterobifunctional molecule comprising an endonuclease binding domain and a DNA ligase binding domain.
17. The system of claim 16, wherein the heterobifunctional molecule comprises a small molecule.
18. The system of claim 1, wherein the splint is part of a guide nucleic acid or wherein a strand of the integrating nucleic acid comprises the splint.
19. The system of claim 1, wherein the splint nucleic acid comprises a donor nucleic acid binding site (DBS) that is at least partially complementary to the integrating nucleic acid, a flap binding site (FBS) that is at least partially complementary to the target nucleic acid, and a guide binding site (GBS) that is at least partially complementary to the guide nucleic acid.
20. The system of claim 1, wherein the system comprises: a second integrating nucleic acid, wherein the second integrating nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a second target nucleic acid strand; anda second splint nucleic acid which is configured to locate the 5’ end of the second integrating nucleic acid to the 3’ end of the strand break in the second strand of the target nucleic acid generated by the endonuclease.
21. The system of claim 20, wherein the system comprises a guide RNA and the endonuclease comprises an RNA-guided endonuclease.
22. The system of claim 20, wherein the second integrating nucleic is at least partially complementary to the first integrating nucleic acid.
23. The system of claim 20, wherein the DBS of the first splint and the DBS of the second splint overlap by at least 10 bp or more, or at least 15 bp or more, or at least 20 bp or more, or at least 30 bp or more, or at least 40 bp or more.
24. A system comprising a cell containing a heterologous RNA-guided endonuclease, a DNA ligase, and an integrating nucleic acid comprising an integrating recombination sequence and configured to be ligated by the DNA ligase to a strand break generated by the endonuclease in a target nucleic acid.
25. The system of claim 24, wherein the DNA ligase is endogenous to the cell.
26. The system of claim 24, wherein the DNA ligase is a heterologous DNA ligase.
27. The system of any one of claims 1-26, wherein the endonuclease comprises anRNA-guided endonuclease.
28. The system of claim 27, wherein the endonuclease comprises a class II CRISPR / Cas endonuclease.
29. The system of claim 28, wherein the endonuclease comprises a Cas9 endonuclease.
30. The system of claim 27, wherein the endonuclease comprises a nickase.
31. The system of claim 28, wherein the endonuclease comprises a Cas9 nickase.
32. The system of any one of claims 1-31, wherein the endonuclease or the DNA ligase comprises a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, nucleic acid binding domain, or streptavidin.
33. The system of any one of claims 1-32, further comprising a guide nucleic acid.
34. The system of any one of claims 1-33, wherein the integrating nucleic acid comprises at least one attachment site (att) integrating recombination sequence.
35. The system of claim 34, wherein the integrating nucleic acid comprises an attachment site on the bacterial part (attB) integrating recombination sequence.
36. The system of claim 34, wherein the integrating nucleic acid comprises two attachment sites on the bacterial part (attB) integrating recombination sequences.
37. The system of claim 34, wherein the integrating nucleic acid comprises an attachment site on the phage part (attP) integrating recombination sequence.
38. The system of claim 34, wherein the integrating nucleic acid comprises two attachment sites on the phage part (attP) integrating recombination sequences.
39. The system of claim 34, wherein the integrating nucleic acid comprises one attachment site on the bacterial part (attB) integrating recombination sequence and one attachment site on the phage part (attP) integrating recombination sequence.
40. The system of any one of claims 1-39, further comprising: an integrase; and a second integrating nucleic acid comprising: at least one cognate recombination sequence; and an exogenous nucleic acid, wherein the at least one cognate recombination sequence of the second integrating nucleic acid couples with the at least one integrating recombination sequence of the integrating nucleic acid, and wherein the integrase integrates the second integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence.
41. The system of claim 40, wherein the second integrating nucleic acid comprises at least one attachment site (att) cognate recombination sequence.
42. The system of claim 41, wherein the second integrating nucleic acid comprises an attachment site on the bacterial part (attB) cognate recombination sequence.
43. The system of claim 41, wherein the second integrating nucleic acid comprises two attachment sites on the bacterial part (attB) cognate recombination sequences.
44. The system of claim 41, wherein the second integrating nucleic acid comprises an attachment site on the phage part (attP) cognate recombination sequence.
45. The system of claim 41, wherein the second integrating nucleic acid comprises two attachment sites on the phage part (attP) cognate recombination sequences.
46. The system of claim 41, wherein the second integrating nucleic acid comprises one attachment site on the bacterial part (attB) integrating recombination sequence and one attachment site on the phage part (attP) integrating recombination sequence.
47. The system of claim 40, wherein the integrase is coupled to the endonuclease or the ligase.
48. The system of claim 40, wherein the integrase comprises a serine integrase.
49. The system of claim 48, wherein the serine integrase comprises a PhiC31 bacteriophage integrase, a Bxbl mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (PaOl), a Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74).
50. The system of claim 40, wherein the integrase is coupled to a recombination directionality factor (RDF).
51. The system of any one of paragraphs 1-50, wherein the integrating nucleic acid or the second integrating nucleic acid comprises a modified nucleotide.
52. The system of claim 51, wherein the modified nucleotide comprises a methylated nucleotide.
53. The system of claim 51, wherein the modified nucleotide comprises methylated cytosine (e.g. 5-mC), 5-hydroxymethylcytosine (5-hmC), 5 -formyl cytosine (5-fC), 5- carboxyl cytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
54. The system of claim 53, wherein the modified nucleotide comprises methylated cytosine 5-mC.
55. The system of claim 51 wherein the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
56. The system of any one of claims 1-55, wherein the splint comprises a modification.
57. The system of claim 56, wherein the modification comprises a streptavidin operatively coupled to the splint.
58. The system of claim 57, wherein the splint comprising the streptavidin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a biotin.
59. The system of claim 56, wherein the modification comprises a biotin operatively coupled to the splint.
60. The system of claim 59, wherein the splint comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a streptavidin.
61. The system of any one of claims 1-60, wherein the endonuclease comprises a fusion partner.
62. The system of claim 61, wherein the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone Hl central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof.
63. A method of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, the method comprising introducing into the host cell a system of nucleic acids comprising: a) a guide nucleic acid comprising: i. a spacer complementary to a region of a genomic locus of a genomic strand, ii. a scaffold for complexing with an endonuclease, iii. an optional donor binding site that is at least partially complementary to an integrating nucleic acid, and iv. a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; andb) a first integrating nucleic acid comprising at least one nucleic acid sequence encoding at least one integrating recombination sequence, wherein the first integrating nucleic acid further comprises a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by the endonuclease.
64. A method of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, the method comprising introducing into the host cell a system of nucleic acids comprising: a) a guide nucleic acid comprising: a spacer complementary to a region of a genomic locus of a genomic strand, a scaffold for complexing with an endonuclease, and an optional splint binding site that is at least partially complementary to a splinting nucleic acid; b) a first integrating nucleic acid comprising at least one nucleic acid sequence encoding at least one integrating recombination sequence, wherein the first integrating nuclei acid further comprises a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by the endonuclease; and c) a splinting nucleic acid comprising a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, and comprising an optional guide binding site (GBS) that is at least partially complementary to a guide nucleic acid.
65. The method of any one of claims 63 or 64, wherein the splinting nucleic acid further comprises a donor binding site (DBS) that is at least partially identical or complementary to a portion of the first integrating nucleic acid.
66. The method of claim 64, wherein the GBS comprises a modified nucleotide.
67. The method of claim 66, wherein the GBS comprises a region wherein alternating nucleotides or every third nucleotide comprises an LNA.
68. The method of claim 65, wherein the DBS comprises a modified nucleotide.
69. The method of claim 68, wherein the DBS comprises a region wherein alternating nucleotides or every third nucleotide comprises an LNA.
70. The method of any one of claims 63 or 64, wherein the splinting nucleic acid comprises a binding sited having modified nucleotides in the configuration of a splinting nucleic acid of Table 12.
71. The method of any one of claims 63 or 64, wherein the guide nucleic acid comprises a sequence of linking nucleic acids between the scaffold and the donor binding site.
72. The method of any one of claims 63 or 64, wherein the guide nucleic acid comprises MS2 binding loops within the scaffold.
73. The method of claim 64, wherein the guide nucleic acid comprises MS2 binding loops between the scaffold and the donor binding site.
74. The method of any one of claims 63 or 64, wherein the guide nucleic acid, the first integrating nucleic acid, or the splinting nucleic acid comprises a modified internucleoside linkage.
75. The method of claim 74, wherein the modified internucleoside linkage comprises a phosphorothioate linkage.
76. The method of claim 75, wherein the modified internucleoside linkage comprises a phosphonoacetate linkage.
77. The method of claim 63 or 64, wherein the guide nucleic acid, the first integrating nucleic acid, or the splinting nucleic acid comprises a modified nucleoside.
78. The method of claim 77, wherein the modified nucleoside comprises a locked nucleic acid (LNA), a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof.
79. The method of claim 63 or 64, wherein the endonuclease comprises an RNA- guided endonuclease.
80. The method of claim 79, wherein the endonuclease comprises a class II CRISPR / Cas endonuclease.
81. The method of claim 80, wherein the endonuclease comprises a Cas9 endonuclease.
82. The method of claim 81, wherein the endonuclease comprises a nickase.
83. The method of claim 82, wherein the endonuclease comprises a Cas9 nickase.
84. The method of any one of claims 63-83, wherein the method further comprises a second integrating nucleic acid, wherein the second integrating nucleic acid comprises: at least one nucleic acid sequence encoding a cognate recombination sequence; and wherein the at least one cognate recombination sequence of the second integrating nucleic acid couples with the at least one integrating recombination sequence of the first integrating nucleic acid, and wherein an integrase integrates the second integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence.
85. The method of claim 84, wherein the second integrating nucleic acid comprises at least one attachment site (att) cognate recombination sequence.
86. The method of claim 85, wherein the second integrating nucleic acid comprises an attachment site on the bacterial part (attB) cognate recombination sequence.
87. The method of claim 85, wherein the second integrating nucleic acid comprises two attachment sites on the bacterial part (attB) cognate recombination sequences.
88. The method of claim 85, wherein the second integrating nucleic acid comprises an attachment site on the phage part (attP) cognate recombination sequence.
89. The method of claim 85, wherein the second integrating nucleic acid comprises two attachment sites on the phage part (attP) cognate recombination sequences.
90. The method of claim 85, wherein the second integrating nucleic acid comprises one attachment site on the bacterial part (attB) integrating recombination sequence and one attachment site on the phage part (attP) integrating recombination sequence.
91. The method of claim 84, wherein the second integrating nucleic acid comprises at least one locus of X(cross)-over in Pl (LoxP) cognate recombination sequence.
92. The method of claim 84, wherein the first integrating nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence.
93. The method of any one of claims 84-92, wherein the second integrating nucleic acid comprises a regulatory sequence.
94. The method of claim 93, wherein the regulatory sequence is a promoter.
95. The method of any one of claims 84-94, wherein the integrase is a serine integrase.
96. The method of claim 95, wherein the serine integrase is a PhiC31 bacteriophage integrase, a Bxbl mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (PaOl), a Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74).
97. The method of any one of claims 95 or 96, wherein the integrase is coupled to a recombination directionality factor (RDF).
98. The method of any one of claims 84-94, wherein the integrase is a tyrosine integrase.
99. The method of claim 98, wherein the tyrosine integrase is a Cre recombinase.
100. The method of claim 98, wherein the tyrosine integrase is a flippase (Flp).
101. The method of any one of claims 63-100, wherein the first integrating nucleic acid or the second integrating nucleic acid comprises a modified nucleotide.
102. The method of claim 101, wherein the modified nucleotide comprises a methylated nucleotide.
103. The method of claim 102, wherein the methylated nucleotide comprises methylated cytosine (e.g. 5-mC), 5 -hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5- carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
104. The method of claim 103, wherein the methylated nucleotide comprises methylated cytosine 5-mC.
105. The method of claim 101, wherein the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
106. The method of any one of claims 63-105, wherein the splinting nucleic acid comprises a modification.
107. The method of claim 106, wherein the splinting nucleic acid comprises a sequence of Table 12.
108. The method of claim 106, wherein the modification comprises a streptavidin operatively coupled to the splinting nucleic acid.
109. The method of claim 108, wherein the splinting nucleic acid comprising the streptavidin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a biotin.
110. The method of claim 106, wherein the modification comprises a biotin operatively coupled to the splinting nucleic acid.
111. The method of claim 110, wherein the splinting nucleic acid comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a streptavidin.
112. The method of any one of claims 63-111, wherein the endonuclease comprises a fusion partner.
113. The method of claim 112, wherein the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone Hl central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof.
114. An editing method, comprising: ligating an integrating recombination sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease.
115. A fusion protein, the fusion protein comprising:(i) an integrase; and(ii) a ligase or an endonuclease.
116. The fusion protein of claim 115, wherein the ligase or endonuclease comprises the ligase.
117. The fusion protein of claim 115, wherein the ligase or endonuclease comprises the endonuclease.
118. The fusion protein of claim 115, wherein the ligase or endonuclease comprises the ligase and the endonuclease.
119. A method, comprising: ligating an integrating sequence to a nick in a target nucleic acid in a cell, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating sequence comprises a modified nucleotide.
120. The method of claim 119, wherein the integrating nucleic acid comprises a modified nucleotide.
121. The method of claim 120, wherein the modified nucleotide comprises a methylated nucleotide.
122. The method of claim 120, wherein the modified nucleotide comprises methylated cytosine (e.g. 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5- carboxyl cytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
123. The method of claim 121, wherein the methylated nucleotide comprises methylated cytosine 5-mC.
124. The method of claim 119, wherein the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
125. An editing method, comprising: ligating an integrating sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating of the integrating sequence introduces methylated nucleosides into the genome.
126. An editing method, comprising: ligating an integrating sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, andwherein the integrating of the integrating sequence removes methylated nucleosides from the genome.
127. An editing method, comprising: contacting a target nucleic acid in a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a nick at the predetermined locus of the target nucleic acid; introducing a pre-synthesized, methylated integrating nucleic acid; and ligating an end of the integrating nucleic acid to an end of the nick at the predetermined locus of the target nucleic acid.
128. An editing system, comprising: a ligase; an endonuclease that introduces a nick at a predetermined locus of a target nucleic acid; and a pre-synthesized, methylated integrating nucleic acid comprising an end that is ligated by the ligase to an end of the nick at the predetermined locus of the target nucleic acid.
129. A method, comprising: contacting an endonuclease with a splinting nucleic acid, the endonuclease comprising a first heterodimerization moiety, and the splinting nucleic acid comprising a second heterodimerization moiety.
130. The method of claim 129, further comprising introducing a nick or a strand break at a predetermined locus of a target nucleic acid by the endonuclease, and ligating an integrating nucleic acid to the nick or to the strand break, wherein the splinting nucleic acid binds to the integrating nucleic acid and to the target nucleic acid.
131. The method of claim 129 or 130, wherein the first heterodimerization moiety comprises biotin, and the second heterodimerization moiety comprises streptavidin or avidin, or wherein the second heterodimerization moiety comprises biotin, and the first heterodimerization moiety comprises streptavidin or avidin.
132. A system, comprising:an endonuclease, wherein the endonuclease is coupled to a first heterodimerization moiety; and a splinting nucleic acid comprising a second heterodimerization moiety.
133. The system of claim 132, further comprising the ligase.
134. The method of claim 133, wherein the endonuclease introduces a nick or a strand break at a predetermined locus of a target nucleic acid, and the ligase ligates an integrating nucleic acid to the nick or to the strand break, wherein the splinting nucleic acid binds to the integrating nucleic acid and to the target nucleic acid.
135. The system of any one of claims 132-134, wherein the first heterodimerization moiety comprises biotin, and the second heterodimerization moiety comprises streptavidin or avidin, or wherein the second heterodimerization moiety comprises biotin, and the first heterodimerization moiety comprises streptavidin or avidin.