Revision of genetic material using direct replacement editing
Fusion proteins with DNA-binding proteins and ligases enable precise integration of exogenous nucleic acids into target nucleic acids, addressing the inefficiencies of existing genome editing methods by facilitating accurate sequence replacement and modification.
Patent Information
- Application Number
- PCT/US2025/025847
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-10
- Filing Date
- 2025-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
Existing genome editing methods are inadequate for efficient replacement or revision of nucleic acid sequences in the genome.
The use of fusion proteins comprising a DNA-binding protein, such as an RNA-guided endonuclease, coupled with a DNA ligase, to introduce targeted strand breaks and ligate exogenous nucleic acids into a target nucleic acid, facilitated by guide RNAs and integrating nucleic acids with specific recombination sequences.
This approach enables precise and efficient integration of exogenous nucleic acids into target nucleic acids, allowing for accurate replacement or revision of genomic sequences, including the introduction of modified nucleotides.
Smart Images

Figure US2025025847_30102025_PF_FP_ABST
Abstract
Description
PATENT Docket No. J4040-99018 REVISION OF GENETIC MATERIAL USING DIRECT REPLACEMENT EDITING RELATED APPLICATIONS AND INCORPORATION BY REFERENCE
[0001] This application claims priority to US provisional application Serial No. 63 / 637,319, filed April 22, 2024, US provisional application Serial No.63 / 643,320, filed May 6, 2024, and US provisional application Serial No.63 / 645,681, filed May 10, 2024, each incorporated by reference herein in its entirety.
[0002] The foregoing applications, and all documents cited therein or during their prosecution (“appln cited documents”) and all documents cited or referenced in the appln cited documents, and all documents cited or referenced herein (“herein cited documents”), and all documents cited or referenced in herein cited documents, together with any manufacturer’s instructions, descriptions, product specifications, and product sheets for any products mentioned herein or in any document incorporated by reference herein, are hereby incorporated herein by reference, and may be employed in the practice of the invention. More specifically, all referenced documents are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference. SEQUENCE LISTING
[0003] The instant application contains a Sequence Listing which has been submitted via Patent Center and is hereby incorporated by reference in its entirety. Said .xml copy, created on April 22, 2025 is named J4040-99018, and is 2,164,977 bytes in size. FIELD OF THE INVENTION
[0004] The invention provides compositions, systems, and methods for revising nucleic acids. The revision may be accomplished using a fusion protein comprising a nuclease, a ligase, an integrase, or a combination thereof. The nucleic acid revision may include ligation of a donor nucleic acid to a target nucleic acid. The nucleic acid editing may include replacement of a portion of the target nucleic acid with a revising nucleic acid. BACKGROUND OF THE INVENTION
[0005] Improved genome editing methods are needed for replacement or revision of nucleic acid sequences in the genome. DM2\21326271.11PATENT Docket No. J4040-99018
[0006] Citation or identification of any document in this application is not an admission that such document is available as prior art to the present invention. SUMMARY OF THE INVENTION
[0007] Disclosed herein, in some aspects, are systems inside a cell comprising: a fusion protein. In some aspects, are systems or compositions disclosed herein comprise: a DNA-binding protein coupled to a DNA ligase. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the coupling is covalent. Some aspects include a fusion protein comprising the DNA-binding protein (e.g., endonuclease such as an RNA-guided endonuclease) and the DNA ligase. Some aspects include a composition comprising: a cell containing a DNA-binding protein (e.g., endonuclease such as an RNA-guided endonuclease) and a DNA ligase, both of which are heterologous to the cell. Some aspects include a composition comprising: a cell containing a DNA-binding protein that is heterologous to the cell (e.g., endonuclease such as an RNA-guided endonuclease) and a DNA ligase that is endogenous to the cell. Some aspects include a composition comprising: a cell containing a DNA-binding protein (e.g., endonuclease such as an RNA-guided endonuclease) and a DNA ligase, in which the DNA ligase is endogenous to the cell. In some aspects, the DNA- binding protein is amino (N)-terminal relative to the DNA ligase within the fusion protein. In some aspects, the DNA-binding protein is carboxy (C)-terminal relative to the DNA ligase within the fusion protein. In some aspects, the connection comprises a linker comprising 1-100 amino acids. In some aspects, the coupling is non-covalent. In some aspects, the composition comprises a first polypeptide comprising at least part of the DNA-binding protein, and a second polypeptide comprising at least part of the DNA ligase, wherein the first and second polypeptides are non- covalently coupled. In some aspects, the first polypeptide comprises a first heterodimerization domain that binds a second heterodimerization domain, and wherein the second polypeptide comprises the second heterodimerization domain. In some aspects, the heterodimer domains comprise a leucine zipper, PDZ domain, streptavidin, streptavidin binding protein, foldon domain, hydrophobic moiety, or a functional binding fragment thereof. In some aspects, the first polypeptide comprises a first intein that binds a second intein, and wherein the second polypeptide comprises the second intein. In some aspects, the ligase comprises a hairpin binding motif, and wherein the DNA-binding protein and the DNA ligase are coupled with a nucleic acid comprising a scaffold that binds to the DNA-binding protein and a hairpin that binds to the hairpin binding DM2\21326271.12PATENT Docket No. J4040-99018 motif. In some aspects, the hairpin binding motif comprises an MS2 coat protein (MCP) peptide, and wherein the hairpin comprises an MS2 hairpin. In some aspects, the DNA-binding protein and the DNA ligase are coupled with a heterobifunctional molecule comprising an endonuclease binding domain and a DNA ligase binding domain. In some aspects, the heterobifunctional molecule comprises a small molecule. In some aspects, the DNA-binding protein comprises a class II CRISPR / Cas endonuclease. In some aspects, the DNA-binding protein comprises a Cas9 endonuclease. In some aspects, the DNA-binding protein comprises a nickase. In some aspects, the DNA-binding protein comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 1-13, or a functional fragment thereof. In some aspects, the DNA ligase ligates DNA strands base paired to a DNA splint. In some aspects, the DNA ligase ligates DNA strands base paired to an RNA splint. In some aspects, the DNA ligase comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 55-96, or a functional fragment thereof. In some aspects, the DNA-binding protein or the DNA ligase comprises a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, or tag polypeptide. Some aspects include a guide RNA and an integrating nucleic acid. In some aspects, a strand of the integrating nucleic acid acts as the splint for the ligase. In some aspects, the guide nucleic acid acts as a splint for the ligase. In some aspects, the splint comprises a modification. In some aspects, the modification comprises a streptavidin operatively coupled to the splint. In some aspects, the splint comprising the streptavidin is operatively coupled with the DNA ligase, and the DNA ligase is operatively coupled to a biotin. In some aspects, the modification to the splint comprises a biotin operatively linked to the splint. In some aspects, the splint comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a streptavidin. Some aspects include one or more nucleic acids encoding the composition. Some aspects include a cell comprising the composition or comprising the one or more nucleic acids.
[0008] Disclosed herein, in some aspects, are methods of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising: contacting a target nucleic acid in a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a strand break at the predetermined locus of the target nucleic acid. In some aspects, the strand break is a double-strand break. DM2\21326271.13PATENT Docket No. J4040-99018
[0009] Disclosed herein, in some aspects, are methods of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising: contacting a target nucleic acid in a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a strand break at the predetermined locus of the target nucleic acid; introducing a pre-synthesized integrating nucleic acid to the cell; and ligating a 5' end of the pre-synthesized integrating nucleic acid to a 3' end of the nick at the predetermined locus of the target nucleic acid. In some aspects the strand break is a double-strand break. In some aspects, the strand break is a nick. In some aspects, the endonuclease comprises an RNA-guided endonuclease. In some aspects, the endonuclease comprises a class II CRISPR / Cas endonuclease. In some aspects, the endonuclease comprises a Cas9 endonuclease. In some aspects, the endonuclease comprises a nickase. In some aspects, the endonuclease comprises Cas9 nickase. Some aspects include contacting the endonuclease and the predetermined locus of the target nucleic acid with a guide nucleic acid. In some aspects, said ligating is performed by a ligase coupled to the endonuclease. In some aspects, the endonuclease comprises a fusion partner. In some aspects, the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone H1 central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof. In some aspects, the pre-synthesized integrating nucleic acid comprises a mutation in relation to the target nucleic acid. In some aspects, the nick comprises a single phosphodiester strand break in the otherwise double stranded target nucleic acid. In some aspects, the nick comprises a non-sticky, non-blunt end of a strand of the target nucleic acid. In some aspects, the target nucleic acid comprises a chromosome of the cell. In some aspects, the cell is eukaryotic.
[0010] Disclosed herein, in some aspects, are systems inside a cell, comprising: a ligase (e.g., a heterologous or an endogenous ligase); an endonuclease that introduces a nick at a predetermined locus of a target nucleic acid; and a pre-synthesized integrating nucleic acid comprising a 5’ end that is ligated by the ligase to a 3' end of the nick at the predetermined locus of the target nucleic acid. In some aspects, the endonuclease comprises a class II CRISPR / Cas endonuclease. In some aspects, the endonuclease comprises Cas9 nickase. Some aspects include a guide nucleic acid that brings the endonuclease into proximity with the predetermined locus of the target nucleic acid. In some aspects, the ligase is coupled to the endonuclease. In some aspects, the pre-synthesized integrating nucleic acid comprises a mutation in relation to the target nucleic acid. In some aspects, DM2\21326271.14PATENT Docket No. J4040-99018 the nick comprises a single phosphodiester strand break in the otherwise double stranded target nucleic acid. In some aspects, the nick comprises a non-sticky, non-blunt end of a strand of the target nucleic acid. In some aspects, the target nucleic acid comprises a chromosome of a cell. In some aspects, the cell is eukaryotic.
[0011] Disclosed herein, in some aspects, are systems inside a cell, comprising: a ligase; an endonuclease; and an integrating nucleic acid comprising an integrating recombination sequence. In some aspects, the integrating nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid. In some aspects, the strand break is a double-strand break. In some aspects, the strand break is a nick.
[0012] Disclosed herein, in some aspects, are systems inside a cell, comprising: a ligase; an endonuclease; an integrating nucleic acid comprising an integrating recombination sequence; an integrase; and a second integrating nucleic acid, comprising: (i) at least one cognate recombination sequence and (ii) an exogenous nucleic acid. In some aspects, the integrating nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid. In some aspects, the cognate recombination sequence of the second integrating nucleic acid couples with the at least one integrating recombination sequence of the integrating nucleic acid. In some aspects, the integrase integrates the second integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence.
[0013] Disclosed herein, in some aspects, are integrating nucleic acids configured to be ligated by a ligase to a strand break generated by an endonuclease in a target nucleic acid. In some aspects, the integrating nucleic acid comprises at least one attachment site (att) integrating recombination sequence. In some aspects, the integrating nucleic acid comprises an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the integrating nucleic acid comprises two attB integrating recombination sequence. In some aspects, the integrating nucleic acid comprises an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the integrating nucleic acid comprises two attP integrating recombination sequences. In some aspects, the integrating nucleic acid comprises the integrating nucleic acid comprises one attB integrating recombination sequence and one attP integrating recombination sequence. In some aspects, the integrating nucleic acid comprises at least one locus of X(cross)-over in P1 (LoxP) integrating recombination sequence. In some aspects, the integrating DM2\21326271.15PATENT Docket No. J4040-99018 nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence.
[0014] Disclosed herein, in some aspects, are second integrating nucleic acids configured to be integrated by an integrase, in whole or in part, into a target nucleic acid at an integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one attachment site (att) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attB integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attP integrating recombination sequences. In some aspects, the second integrating nucleic acid comprises the integrating nucleic acid comprises one attB integrating recombination sequence and one attP integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one locus of X(cross)-over in P1 (LoxP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises a regulatory sequence. In some aspects, the regulatory sequence is a promoter.
[0015] Disclosed herein, in some aspects, are integrating nucleic acids or second integrating nucleic acids comprising a modified nucleotide. In some aspects, the modified nucleotide comprises a methylated nucleotide. In some aspects, the modified nucleotide comprises methylated cytosine (e.g. 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5- carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof. In some aspects, the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
[0016] Disclosed herein are integrases that integrate a second integrating nucleic acid, in whole or in part, into a target nucleic acid at an integrating recombination sequence. In some aspects, the integrase is coupled to an endonuclease. In some aspects, the integrase is coupled to a ligase. In some aspects, the integrase is coupled to a recombination directionality factor (RDF). In some aspects, the integrase is a serine integrase. The serine integrase can be a PhiC31 bacteriophage integrase, a Bxb1 mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (Pa01), a DM2\21326271.16PATENT Docket No. J4040-99018 Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74). In some aspects, the integrase is a tyrosine integrase. The tyrosine integrase can be a Cre recombinase or a flippase (Flp).
[0017] In some embodiments, disclosed herein are methods for ligase-mediated programmable genomic integration (L-PGI) using donor DNA, splint DNA, guide RNAs, and fusion proteins. The donor includes modular binding domains for the guide RNA and target genomic overhang. Splint DNA can carry a guide-binding site (GBS), donor-binding site (DBS), and flap-binding site (FBS), which together stabilize the integration complex. Guide RNAs can be configured with linker-modified domains (lmgRNAs) that facilitate strand-specific binding.
[0018] Disclosed herein, in some aspects, are methods of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising: a guide nucleic acid comprising: (a) a spacer complementary to a region of a genomic locus of a genomic strand, (b) a scaffold for complexing with a DNA-binding protein, (c) an optional donor binding site that is at least partially complementary to an integrating nucleic acid, and (d) a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; and a first integrating nucleic acid comprising (i) at least one nucleic acid sequence encoding at least one integrating recombination sequence and (ii) a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by a DNA-binding protein. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. Disclosed herein, in some aspects, are systems of nucleic acids comprising: a guide nucleic acid comprising: (a) a spacer complementary to a region of a genomic locus of a genomic strand, (b) a scaffold for complexing with a DNA-binding protein, and (c) an optional donor binding site that is at least partially complementary to a splinting nucleic acid; an integrating nucleic acid comprising (i) at least one nucleic acid sequence encoding at least one integrating recombination sequence and (ii) a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by a DNA-binding protein; and a splinting nucleic acid comprising a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, and comprising an optional guide binding site that is at least partially complementary to a guide nucleic acid. In some aspects, the genomic strand is in a cell. In some aspects, the splinting nucleic acid further comprises a donor binding site that is at least partially identical or complementary to a portion of the integrating nucleic acid. In some aspects, the guide nucleic acid comprises a sequence of linking nucleic acids DM2\21326271.17PATENT Docket No. J4040-99018 between the scaffold and the donor binding site. In some aspects, the guide nucleic acid comprises MS2 binding loops within the scaffold. In some aspects, the guide nucleic acid comprises MS2 binding loops between the scaffold and the donor binding site. In some aspects, the guide nucleic acid, the first integrating nucleic acid, or the splinting nucleic acid comprises a modified internucleoside linkage. In some aspects, the modified internucleoside linkage comprises a phosphorothioate linkage. In some aspects, the modified internucleoside linkage comprises a phosphoacetate linkage. In some aspects, the first integrating nucleic acid, the guide nucleic acid, or the splinting nucleic acid comprises a modified nucleoside. In some aspects, the modified internucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the guide nucleic acid or the integrating nucleic acid. In some aspects, the guide nucleic acid or the integrating nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, or a combination thereof. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end of the guide nucleic acid or the integrating nucleic acid. The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. In some aspects, the endonuclease comprises an RNA-guided endonuclease. In some aspects, the endonuclease comprises a class II CRISPR / Cas endonuclease. In some aspects, the endonuclease comprises a Cas9 endonuclease. In some aspects, the endonuclease comprises a nickase. In some aspects, the endonuclease comprises a Cas9 nickase.
[0019] In some aspects, a method disclosed herein further comprises a second integrating nucleic acid, wherein the second integrating nucleic acid comprises at least one nucleic acid sequence encoding a cognate recombination sequence; and wherein the at least one cognate recombination sequence of the second integrating nucleic acid couples with the at least one integrating recombination sequence of the first integrating nucleic acid, and wherein an integrase integrates the second integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attB integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises two attP integrating recombination sequences. In some aspects, the second DM2\21326271.18PATENT Docket No. J4040-99018 integrating nucleic acid comprises the integrating nucleic acid comprises one attB integrating recombination sequence and one attP integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one locus of X(cross)-over in P1 (LoxP) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence. In some aspects, the second integrating nucleic acid comprises a regulatory sequence. In some aspects, the regulatory sequence is a promoter. In some aspects, the integrase is coupled to a recombination directionality factor (RDF). In some aspects, the integrase is a serine integrase. The serine integrase can be a PhiC31 bacteriophage integrase, a Bxb1 mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (Pa01), a Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74). In some aspects, the integrase is a tyrosine integrase. The tyrosine integrase can be a Cre recombinase or a flippase (Flp). In some aspects, the first integrating nucleic acid or the second integrating nucleic acid comprises a modified nucleotide. In some aspects, the modified nucleotide comprises a methylated nucleotide. In some aspects, the modified nucleotide comprises methylated cytosine (e.g. 5-mC), 5-hydroxymethylcytosine (5- hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof. In some aspects, the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof. In some aspects, the splinting nucleic acid comprises a modification. In some aspects, the modification comprises a streptavidin operatively coupled to the splint. In some aspects, the splint comprising the streptavidin is operatively coupled with the DNA ligase, and the DNA ligase is operatively coupled to a biotin. In some aspects, the modification to the splint comprises a biotin operatively linked to the splint. In some aspects, the splint comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a streptavidin. In some aspects, the endonuclease comprises a fusion partner. In some aspects, the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone H1 central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof.
[0020] Disclosed herein, in some aspects, are fusion proteins, comprising: a DNA-binding protein connected to a DNA ligase. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the connection between DM2\21326271.19PATENT Docket No. J4040-99018 the DNA-binding protein and the DNA ligase is covalent. Some aspects include a fusion protein comprising the DNA-binding protein upstream of the DNA ligase. Some aspects include a fusion protein comprising the DNA-binding protein downstream of the DNA ligase. In some aspects, the connection comprises a linker comprising 1-100 amino acids. In some aspects, the composition comprises a first polypeptide comprising at least part of the DNA-binding protein, and a second polypeptide comprising at least part of the DNA ligase, wherein the first and second polypeptides are bound together covalently or non-covalently. In some aspects, the first polypeptide comprises a first heterodimerization domain that binds a second heterodimerization domain, and wherein the second polypeptide comprises the second heterodimerization domain. In some aspects, the heterodimer domains comprise a leucine zipper, PDZ domain, streptavidin, streptavidin binding protein, foldon domain, hydrophobic moiety, or a functional binding fragment thereof. In some aspects, the first polypeptide comprises a first intein that binds a second intein, and wherein the second polypeptide comprises the second intein. In some aspects, the DNA-binding protein and the DNA ligase are bound together by a small molecule. In some aspects, the DNA-binding protein comprises a class II CRISPR / Cas endonuclease. In some aspects, the DNA-binding protein comprises a Cas9 endonuclease. In some aspects, the DNA-binding protein comprises a nickase. In some aspects, the DNA-binding protein comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 1-13, or a functional fragment thereof. In some aspects, the DNA ligase ligates DNA strands base paired to a DNA splint. In some aspects, the DNA ligase ligates DNA strands base paired to an RNA splint. In some aspects, the DNA ligase comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 55-96, or a functional fragment thereof. In some aspects, the DNA-binding protein or the DNA ligase comprises a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or streptavidin. Some aspects include a guide RNA and an integrating nucleic acid. Some aspects relate to a cell comprising the composition. Some aspects include a nucleic acid encoding the composition. Some aspects include one or more nucleic acids encoding the first or second polypeptides. Some aspects include an editing method (e.g., nucleic acid) which uses the composition. Some aspects include a method of treatment using the composition. Some aspects include administering the composition to a subject. DM2\21326271.110PATENT Docket No. J4040-99018
[0021] Disclosed herein, in some aspects, are editing methods, comprising ligating an integrating recombination sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease.
[0022] Disclosed herein, in some aspects, are fusion proteins, comprising: a DNA-binding protein fused to a DNA ligase. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. Disclosed herein, in some aspects, are fusion proteins, comprising: an integrase, and a ligase or an endonuclease. In some aspects, a fusion protein disclosed herein comprises an integrase and a ligase. In some aspects, a fusion protein disclosed herein comprises an integrase, a ligase, and an endonuclease. In some aspects, a fusion protein disclosed herein comprises an integrase and an endonuclease. Disclosed herein, in some aspects, are protein complexes, comprising: a DNA-binding protein bound to a DNA ligase. In some aspects, the endonuclease and the DNA ligase are bound together through heterodimerization domains. In some aspects, the heterodimerization domains comprise leucine zippers, PDZ domains, streptavidin, and streptavidin binding protein, foldon domains, hydrophobic polypeptides, an antibody that binds the Cas nickase, or an antibody that binds the DNA ligase, or one or more binding fragments thereof. Disclosed herein, in some aspects, are cells comprising the fusion protein or the protein complex. Disclosed herein, in some aspects, are cells comprising a heterologous DNA-binding protein and a DNA ligase that was introduced into the cell. Some aspects include a nuclease that is different from the DNA-binding protein. Disclosed herein, in some aspects, are guide nucleic acids, comprising: a spacer at least partially reverse complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; and a flap binding site at least partially reverse complementary to a nucleic acid flap, and an integrating nucleic acid binding site. Disclosed herein, in some aspects, are integrating nucleic acids, comprising: a single or double-stranded DNA region to be inserted into a target nucleic acid, wherein the single or double-stranded DNA region is flanked by at least one additional single-stranded region comprising a guide binding site. Disclosed herein, in some aspects, are editing systems, comprising a DNA-binding protein, the guide nucleic acid, and the integrating nucleic acid. Disclosed herein, in some aspects, are editing methods, comprising: contacting a target nucleic acid with the editing system and a DNA ligase.
[0023] Disclosed herein, in some aspects, are methods comprising ligating an integrating sequence to a nick in a target nucleic acid in a cell, wherein the nick has been generated by DM2\21326271.111PATENT Docket No. J4040-99018 contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating sequence comprises a modified nucleotide. In some aspects, the integrating nucleic acid comprises a modified nucleotide. In some aspects, the modified nucleotide comprises a methylated nucleotide. In some aspects, the modified nucleotide comprises methylated cytosine (e.g.5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6- methyladenine (6-mA), or a combination thereof. In some aspects, the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
[0024] Disclosed herein, in some aspects, are methods comprising ligating an integrating sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating of the integrating sequence introduces methylated nucleosides into the genome.
[0025] Disclosed herein, in some aspects, are methods comprising ligating an integrating sequence to a nick in a target nucleic acid, wherein the nick has been generated by contacting the target nucleic acid with an RNA-guided endonuclease, and wherein the integrating of the integrating sequence removes methylated nucleosides from the genome.
[0026] Disclosed herein, in some aspects, are methods, comprising: (a) contacting a target nucleic acid in a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a nick at the predetermined locus of the target nucleic acid; (b) introducing a pre-synthesized, methylated integrating nucleic acid; and (c) ligating an end of the integrating nucleic acid to an end of the nick at the predetermined locus of the target nucleic acid.
[0027] Disclosed herein, in some aspects, are editing systems comprising: a ligase; an endonuclease that introduces a nick at a predetermined locus of a target nucleic acid; and a pre- synthesized, methylated integrating nucleic acid comprising an end that is ligated by the ligase to an end of the nick at the predetermined locus of the target nucleic acid.
[0028] Disclosed herein, in some aspects, are methods comprising: contacting an endonuclease with a splinting nucleic acid, the endonuclease comprising a first heterodimerization moiety, and the splinting nucleic acid comprising a second heterodimerization moiety. In some aspects, a method disclosed herein further comprises, introducing a nick or a strand break at a predetermined locus of a target nucleic acid by the endonuclease, and ligating an integrating nucleic acid to the nick or to the strand break, wherein the splinting nucleic acid binds to the integrating nucleic acid DM2\21326271.112PATENT Docket No. J4040-99018 and to the target nucleic acid. In some aspects, the first heterodimerization moiety comprises biotin, and the second heterodimerization moiety comprises streptavidin or avidin, or wherein the second heterodimerization moiety comprises biotin, and the first heterodimerization moiety comprises streptavidin or avidin.
[0029] Disclosed herein, in some aspects, are systems, comprising: an endonuclease, wherein the endonuclease is coupled to a first heterodimerization moiety; and a splinting nucleic acid comprising a second heterodimerization moiety. In some aspects, a system disclosed herein further comprises a ligase. In some aspects, the endonuclease introduces a nick or a strand break at a predetermined locus of a target nucleic acid, and the ligase ligates an integrating nucleic acid to the nick or to the strand break, wherein the splinting nucleic acid binds to the integrating nucleic acid and to the target nucleic acid. In some aspects, the first heterodimerization moiety comprises biotin, and the second heterodimerization moiety comprises streptavidin or avidin, or wherein the second heterodimerization moiety comprises biotin, and the first heterodimerization moiety comprises streptavidin or avidin.
[0030] Disclosed herein, in some aspects, are systems inside a cell: at least one DNA-binding protein; at least one guide nucleic acid comprising: a spacer at least partially complementary to a genomic locus in a cell; a scaffold for complexing with the at least one DNA-binding protein; and an optional donor binding site that is at least partially complementary to an integrating nucleic acid; and at least one DNA ligase; and the integrating nucleic acid, comprising a flap binding site at least partially reverse complementary to a nucleic acid flap and optionally comprising a guide binding site that is at least partially complementary to the at least one guide nucleic acid, wherein the at least one DNA-binding protein cleaves or nicks at least one strand of the genomic locus, and wherein the at least one DNA ligase ligates an end of the integrating nucleic acid to the genomic flap site, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a single- stranded DNA. In some aspects, the integrating nucleic acid comprises a double-stranded DNA.
[0031] Disclosed herein, in some aspects, are systems inside a cell comprising: at least one DNA-binding protein comprising a first DNA-binding protein and an optional second DNA- binding protein; at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid comprising: a first spacer complementary to a first DM2\21326271.113PATENT Docket No. J4040-99018 region of a genomic locus in a cell; a first scaffold for complexing with the first DNA-binding protein; and an optional first donor binding site that at least partially complementary to an integrating nucleic acid; and a first flap binding site that is at least partially identical or complementary to a first genomic flap at or adjacent to the genomic locus; and the second guide nucleic acid comprising: a second spacer complementary to a second region of the genomic locus in the cell; a second scaffold for complexing with the first or second DNA-binding protein; an optional second donor binding site that at least partially complementary to the integrating nucleic acid; and a second flap binding site that is at least partially identical or complementary to a second genomic flap at or adjacent to the genomic locus; at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; and at least one integrating nucleic acid comprising a first strand and a second strand: wherein the first strand comprises an optional first guide binding site that is at least partially complementary to the first guide nucleic acid; and wherein the second strand comprises an optional second guide binding site that is at least partially complementary to the second guide nucleic acid, wherein the first DNA-binding protein and / or the second DNA- binding protein each cleaves or nicks at least one strand of the genomic locus in the cell; and wherein the first DNA ligase ligates an end of the first strand of the integrating nucleic acid to the first genomic flap; and the first or second DNA ligase ligates an end of the second strand of the integrating nucleic acid to the second genomic flap, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. In some aspects, the integrating nucleic acid comprises a double-stranded DNA duplex region. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a 5’ overhang optionally comprising the first guide binding site. In some aspects, the integrating nucleic acid comprises a 5’ overhang optionally comprising the second guide binding site.
[0032] Disclosed herein, in some aspects, are systems inside a cell comprising: at least one DNA-binding protein; at least one guide nucleic acid comprising: a spacer complementary to a genomic locus in a cell; a scaffold for complexing with the at least one DNA-binding protein; and an optional donor binding site that is at least partially complementary to an integrating nucleic acid; at least one DNA ligase; and the integrating nucleic acid that: comprises an optional guide binding site that is at least partially complementary to the at least one guide nucleic acid; and comprises a flap binding site that is at least partially identical or complementary to a genomic flap DM2\21326271.114PATENT Docket No. J4040-99018 at or adjacent to the genomic locus, wherein the at least one DNA-binding protein cleaves or nicks at least one strand of the genomic locus; and wherein the at least one DNA ligase ligates an end of the integrating nucleic acid to the genomic flap, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a DNA comprising a 3’ overhang. In some aspects, the 3’ overhang comprises the guide binding site. In some aspects, the 3’ overhang comprises the flap binding site. In some aspects, the at least one DNA ligase ligates a strand of the integrating nucleic acid to the genomic nucleic acid sequence.
[0033] Disclosed herein, in some aspects, are systems inside a cell comprising: at least one DNA-binding protein comprising a first DNA-binding protein and an optional second DNA- binding protein; at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, the first guide nucleic acid comprising: a first spacer complementary to a first region of a genomic locus in a cell; a first scaffold for complexing with the first DNA-binding protein; and an optional first donor binding site that at least partially complementary to an integrating nucleic acid; and the second guide nucleic acid comprising: a second spacer complementary to a second region of the genomic locus in the cell; a second scaffold for complexing with the first or second DNA-binding protein; and an optional second donor binding site that at least partially complementary to the integrating nucleic acid; and at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; and the integrating nucleic acid comprising a first strand and a second strand: wherein the first strand comprises an optional first guide binding site that is at least partially complementary to the first guide nucleic acid; wherein the second strand comprises an optional second guide binding site that is at least partially complementary to the second guide nucleic acid; wherein the first strand comprises a first flap binding site that is at least partially identical or complementary to a first genomic flap at or adjacent to the genomic locus; and wherein the second strand comprises a second flap binding site that is at least partially identical or complementary to a second genomic flap at or adjacent to the genomic locus; wherein the first DNA-binding protein and / or the second DNA-binding protein each cleaves or nicks at least one strand of the genomic locus in the cell; and wherein the first DNA ligase ligates an end of the first strand of the integrating nucleic acid to the first genomic flap; and the first or second DNA ligase ligates an end of the second strand of the integrating DM2\21326271.115PATENT Docket No. J4040-99018 nucleic acid to the second genomic flap, thereby replacing a region of the genomic locus with the integrating nucleic acid in the cell. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the integrating nucleic acid comprises a double-stranded DNA duplex region. In some aspects, the double-stranded DNA comprises a 3’ overhang optionally comprising the first guide binding site and comprising the first flap binding site. In some aspects, the double stranded DNA comprises a 3’ overhang optionally comprising the second guide binding site and comprising the second flap binding site.
[0034] The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the at least one DNA-binding protein comprises a Cas protein or a functional fragment thereof. In some aspects, the Cas protein or the functional fragment thereof comprises nickase activity. In some aspects, the at least one DNA- binding protein comprises a Cas9 nickase or a functional fragment thereof. In some aspects, the at least one DNA ligase ligates nucleic acids bound to DNA. In some aspects, the at least one DNA ligase ligates nucleic acids bound to RNA. In some aspects, the at least one DNA ligase comprises a PBCV-1 DNA ligase. In some aspects, the at least one DNA ligase is operatively coupled to the at least one DNA-binding protein. In some aspects, the at least one DNA ligase is fused to the at least one DNA-binding protein as a fusion polypeptide. In some aspects, the at least one DNA- binding protein and the at least one DNA ligase each comprises a heterodimer domain. In some aspects, the at least one DNA-binding protein and the at least one DNA ligase forms a heterodimer via the heterodimer domain. In some aspects, the at least one DNA-binding protein comprises a linker. In some aspects, the linker connects the Cas protein or a functional fragment thereof to the heterodimer domain. In some aspects, the at least one DNA-binding protein comprises a localization signal sequence. In some aspects, the at least one DNA ligase comprises a localization signal sequence. In some aspects, the localization signal sequence comprises a nuclear localization sequence (NLS). In some aspects, the at least one DNA-binding protein or the at least one DNA ligase are directed to nucleus of the cell by the NLS. In some aspects, the at least one integrating nucleic acid, such as a donor nucleic acid or a revising nucleic acid, corrects at least one genetic mutation in the at least one genomic locus. In some aspects, the at least one integrating nucleic acid inserts a coding sequence. In some aspects, the coding sequence encodes a full-length protein. In some aspects, the at least one integrating nucleic acid inserts a non-coding sequence. In some aspects, the non-coding sequence comprises a recombination sequence. In some aspects, the non- DM2\21326271.116PATENT Docket No. J4040-99018 coding sequence knocks out an endogenous gene. In some aspects, the non-coding sequence comprises a regulatory element. Some aspects further include a nuclease. In some aspects, the nuclease comprises an exonuclease for digesting the genomic flap. In some aspects, the nuclease comprises a human flap endonuclease 1 (hFEN1), a human exonuclease 5 (hEXO5), a T5 exonuclease, a T7 exonuclease, an exonuclease VIII, a flap endonuclease domain of E. coli PolI, a RecJF, a Lambda exonuclease, a Xni (ExoIXI), a SaFEN (Staphylococcus aureus FEN), a nuclease BAL-31, or a fragment thereof. In some aspects, the heterologous nuclease comprises an endonuclease for digesting the genomic flap, and the endonuclease is different from the at least one DNA-binding protein. In some aspects, the at least one DNA-binding protein comprises at least one additional functional domain. In some aspects, the at least one additional functional domain comprises a chromatin modifying domain. In some aspects, the at least one additional functional domain comprises a cell penetrating peptide. In some aspects, the at least one guide nucleic acid comprises at least one nucleic acid modification. In some aspects, the at least one nucleic acid modification comprises a modification to a backbone, a sugar, a base, or a combination thereof. In some aspects, the at least one DNA-binding protein is complexed with the at least one guide nucleic acid. In some aspects, the at least one guide nucleic acid is complexed with the integrating nucleic acid. In some aspects, the at least one DNA-binding protein, the at least one guide nucleic acid, the at least one at least one DNA ligase, the integrating nucleic acid, or a combination thereof is encoded by a polynucleotide. In some aspects, the polynucleotide comprises mRNA. In some aspects, the polynucleotide comprises a vector. In some aspects, the vector comprises a viral vector. In some aspects, the at least one DNA-binding protein, the at least one guide nucleic acid, the at least one at least one DNA ligase, the integrating nucleic acid, or a combination thereof is encapsulated by at least one lipid nanoparticle. In some aspects, the cell comprises a bacterial cell, a eukaryotic cell, or a plant cell. In some aspects, the eukaryotic cell comprises a mammalian cell. Some aspects include a composition comprising the system. Some aspects include a cell comprising the system. Some aspects include a cell line comprising the cell. Some aspects include a pharmaceutical composition comprising the system. Some aspects include a pharmaceutical composition comprising the composition. Some aspects include a pharmaceutical composition comprising the cell. Some aspects include a pharmaceutically acceptable: excipient, carrier, or diluent. In some aspects, the pharmaceutical composition is formulated for administering intrathecally, intraocularly, intravitreally, retinally, intravenously, intramuscularly, DM2\21326271.117PATENT Docket No. J4040-99018 intraventricularly, intracerebrally, intracerebellarly, intracerebroventricularly, intraperenchymally, subcutaneously, intratumorally, pulmonarily, endotracheally, intraperitoneally, intravesically, intravaginally, intrarectally, orally, sublingually, transdermally, by inhalation, by inhaled nebulized form, by intraluminal-GI route, or a combination thereof to a subject in need thereof. Some aspects include a kit comprising: the system, the composition, or the pharmaceutical composition and a container. In some aspects, include method for modifying a cell comprising contacting a cell with the system. In some aspects, include method for modifying a cell comprising contacting a cell with the composition. In some aspects, include method for modifying a cell comprising contacting a cell with the pharmaceutical composition. In some aspects, the cell is not a dividing cell. In some aspects, the integrating nucleic acid is inserted into the genomic locus of the cell independent of endogenous non-homologous end joining (NHEJ) and independent of endogenous homology-directed repair (HDR). Some aspects include a method for treating a disease or condition in subject in need thereof comprising: contacting the cell or the subject with the system, the composition, or the pharmaceutical composition; replacing a genomic locus in a cell with an integrating nucleic acid, thereby treating the disease or condition in the subject. In some aspects, the cell is not a dividing cell. In some aspects, the integrating nucleic acid is inserted into the genomic locus of the cell independent of endogenous non-homologous end joining (NHEJ) and independent of endogenous homology-directed repair (HDR).
[0035] Disclosed herein, in some aspects, are guide nucleic acids comprising: a spacer that is at least partially complementary to a genomic locus in a cell; a scaffold for complexing with a DNA-binding protein; and a donor binding site that is at least partially complementary to an integrating nucleic acid. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some aspects, the guide nucleic acid comprises a flap binding site that is at least partially complementary to a genomic sequence of the genomic locus. In some aspects, the guide nucleic acid comprises at least one nucleic acid modification. In some aspects, the at least one nucleic acid modification comprises a modification to a backbone, a sugar, a base, or a combination thereof. In some aspects, the guide nucleic acid comprises RNA sequence.
[0036] Disclosed herein, in some aspects, is a system comprising a splinting nucleic acid comprising a guide binding site (GBS), a flap binding site (FBS), and a donor binding site (DBS). In certain embodiments, the guide binding site (GBS) is 15-25 nucleotides in length, the flap DM2\21326271.118PATENT Docket No. J4040-99018 binding site (FBS) is 10-20 nucleotides in length, and the donor binding site (DBS) is 20-30 nucleotides in length. In a specific embodiment, the guide binding site (GBS) is 19 nucleotides in length, the flap binding site (FBS) is 13 nucleotides in length, and the donor binding site (DBS) is 24 nucleotides in length. In further embodiments, the splinting nucleic acid comprises locked nucleic acids (LNAs) in at least one of the GBS, FBS, or DBS. For instance, the splinting nucleic acid may comprise alternating locked nucleic acids (LNAs) throughout the DBS and in the 3' end of the GBS. In another aspect, a system comprises a first complex comprising a first endonuclease, a first DNA ligase, and a first integrating nucleic acid; and a second complex comprising a second endonuclease, a second DNA ligase, and a second integrating nucleic acid. The first and second complexes are configured to generate strand breaks on opposite strands of a target nucleic acid and ligate the first and second integrating nucleic acids to the strand breaks. In some configurations, the first and second integrating nucleic acids comprise complementary overlapping sequences, which may be 5-25 base pairs in length, for example, 10 base pairs in length. Such systems may be configured to delete a region of the target nucleic acid between the two strand breaks or to replace a region of the target nucleic acid between the two strand breaks with an exogenous nucleic acid sequence. In some embodiments of the systems described herein, a nucleic acid encoding the endonuclease, the DNA ligase, or both, comprises a modified nucleotide, such as N1- methylpseudouridine. In other embodiments, the DNA ligase is linked to a first heterodimerization domain and the endonuclease is linked to a second heterodimerization domain that interacts with the first heterodimerization domain. For example, the first heterodimerization domain may be at the C-terminus of the DNA ligase and the second heterodimerization domain may be at the N- terminus of the endonuclease, and these domains may comprise leucine zipper domains. In certain embodiments, the integrating nucleic acid lacks methylation and comprises one or more phosphorothioate bonds, for instance, two phosphorothioate bonds. In another aspect, a method of introducing modifications into a target nucleic acid in a host cell comprises introducing into the host cell a first complex (comprising a first endonuclease, a first DNA ligase, and a first integrating nucleic acid) and a second complex (comprising a second endonuclease, a second DNA ligase, and a second integrating nucleic acid), and generating strand breaks on opposite strands of the target nucleic acid and ligating the first and second integrating nucleic acids to the strand breaks. The host cell may be a primary human hepatocyte, a primary cynomolgus hepatocyte, a primary mouse hepatocyte, a hematopoietic stem cell, or an induced pluripotent stem cell. The complexes DM2\21326271.119PATENT Docket No. J4040-99018 may be delivered to the host cell in a lipid nanoparticle (LNP), which may comprise ionizable lipids, phospholipids, cholesterol, and PEG-lipids. In a further aspect, a method of determining editing quality comprises introducing a gene editing system into a host cell, amplifying a region of a target nucleic acid comprising an edit site, sequencing the amplified region, and determining a fidelity metric, defined as the proportion of sequence reads containing an intended edit without insertions or deletions relative to the total number of sequence reads containing the intended edit. In yet another aspect, a method of integrated gene editing comprises introducing into a host cell a first complex (comprising a first endonuclease, a first DNA ligase, and a first integrating nucleic acid) and a second complex (comprising a second endonuclease, a second DNA ligase, and a second integrating nucleic acid); integrating an integrase recognition sequence into a target nucleic acid; introducing an integrase into the host cell; and integrating an exogenous nucleic acid at the integrase recognition sequence. The integrase recognition sequence may be a Bxb1 attachment site, such as a 38 base pair attB site or a 52 base pair attP site. The exogenous nucleic acid may be delivered by a helper-dependent adenovirus (HDAd) or adeno-associated virus (AAV).
[0037] Accordingly, it is an object of the invention not to encompass within the invention any previously known product, process of making the product, or method of using the product such that Applicants reserve the right and hereby disclose a disclaimer of any previously known product, process, or method. It is further noted that the invention does not intend to encompass within the scope of the invention any product, process, or making of the product or method of using the product, which does not meet the written description and enablement requirements of the USPTO (35 U.S.C. §112, first paragraph) or the EPO (Article 83 of the EPC), such that Applicants reserve the right and hereby disclose a disclaimer of any previously described product, process of making the product, or method of using the product. It may be advantageous in the practice of the invention to be in compliance with Art. 53(c) EPC and Rule 28(b) and (c) EPC. All rights to explicitly disclaim any embodiments that are the subject of any granted patent(s) of applicant in the lineage of this application or in any other lineage or in any prior filed application of any third party is explicitly reserved. Nothing herein is to be construed as a promise.
[0038] It is noted that in this disclosure and particularly in the claims and / or paragraphs, terms such as "comprises", "comprised", "comprising" and the like can have the meaning attributed to it in U.S. Patent law; e.g., they can mean "includes", "included", "including", and the like; and that terms such as "consisting essentially of" and "consists essentially of" have the meaning ascribed DM2\21326271.120PATENT Docket No. J4040-99018 to them in U.S. Patent law, e.g., they allow for elements not explicitly recited, but exclude elements that are found in the prior art or that affect a basic or novel characteristic of the invention.
[0039] These and other embodiments are disclosed or are obvious from and encompassed by, the following Detailed Description. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0041] The following detailed description, given by way of example, but not intended to limit the invention solely to the specific embodiments described, may best be understood in conjunction with the accompanying drawings.
[0042] Fig.1A illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus.
[0043] Fig.1B follows sequentially from Fig.1A and illustrates a donor strand incorporated into one side of a genomic locus, the donor strand having displaced a genomic flap.
[0044] Fig.1C follows sequentially from Fig.1B and illustrates a donor strand incorporated into one side of a genomic locus, and a nick appearing where a genomic flap has been removed.
[0045] Fig.2A illustrates 2 guide nucleic acids, 2 endonucleases, 2 ligases, and a donor strand at a genomic locus.
[0046] Fig.2B follows sequentially from Fig.2A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced 2 genomic flaps.
[0047] Fig.2C follows sequentially from Fig.2B and illustrates a donor strand incorporated into a genomic locus, and 2 nicks appearing where genomic flaps have been removed.
[0048] Fig.3A illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus.
[0049] Fig.3B follows sequentially from Fig.3A and illustrates a donor strand incorporated into one side of a genomic locus, the donor strand having displaced a genomic flap.
[0050] Fig.3C follows sequentially from Fig.3B and illustrates a donor strand incorporated into one side of a genomic locus, and a nick appearing where a genomic flap has been removed.
[0051] Fig.4A illustrates 2 guide nucleic acids, 2 endonucleases, 2 ligases, and a donor strand at a genomic locus. DM2\21326271.121PATENT Docket No. J4040-99018
[0052] Fig.4B follows sequentially from Fig.4A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced 2 genomic flaps.
[0053] Fig.4C follows sequentially from Fig.4B and illustrates a donor strand incorporated into a genomic locus, and 2 nicks appearing where genomic flaps have been removed.
[0054] Fig.5A illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus.
[0055] Fig.5B follows sequentially from Fig.5A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced a genomic flap.
[0056] Fig.5C follows sequentially from Fig.5B and illustrates a donor strand incorporated into one side of a genomic locus, and a nick appearing where a genomic flap has been removed.
[0057] Fig.6A illustrates 2 guide nucleic acids, 2 endonucleases, 2 ligases, and a donor strand at a genomic locus.
[0058] Fig.6B follows sequentially from Fig.6A and illustrates a donor strand incorporated into a genomic locus, the donor strand having displaced 2 genomic flaps.
[0059] Fig.6C follows sequentially from Fig.6B and illustrates a donor strand incorporated into a genomic locus, and 2 nicks appearing where genomic flaps have been removed.
[0060] Fig.7 illustrates some examples of fusion protein arrangements.
[0061] Fig. 8A illustrates an exemplary nicking and ligation pattern of an exogenous first integrating nucleic acid.
[0062] Fig. 8B illustrates a DNA gel showing a pattern associated with 1-Sided Replacer 2 performed in vitro using 30nt GBS / DBS and thermostable T4 ligase. Using a 30nt GBS / DBS combination, a donor containing a protospacer adjacent motif (PAM) mutation, and a thermostable T4 ligase (Hi-T4, NEB), we were able to produce a final Replacer product (Lane 3) corresponding to the size of our control product (Lane 1). Replacer products were not detected in the absence of nicking Cas9 (Cas9n) (Lane 2), or in the absence of the bottom donor which serves as the splint (Lanes 4 & 5).
[0063] Fig.8C illustrates an exemplary nucleic acid gel showing a pattern associated with in vitro 1-Sided Replacer 2 using variable length GBS / DBS combinations and T4 ligase. Using regular T4 ligase (NEB), a final Replacer product corresponding to the size of the control when using multiple GBS / DBS combinations was produced, including no GBS / DBS, 20nt GBS / DBS, and 30nt GBS / DBS. Additionally, in this experiment, recoded dsDNA donors containing PAM DM2\21326271.122PATENT Docket No. J4040-99018 mutation were more efficient at producing final Replacer products compared to PAM mutant dsDNA donors that were not recoded.
[0064] Fig. 9 illustrates measurement of a percentage of cells expressing green fluorescent protein (GFP), indicating gene editing from BFP to GFP by a 1-sided Replacer 2 with nicking Cas9 and DNA ligase.
[0065] Fig.10 illustrates sequencing reads merged and aligned to an amplicon of interest and a percentage of total reads that matched an intended edit via a 1-sided replacer 2 with a nicking Cas9 and a T4 DNA ligase.
[0066] Fig.11 illustrates sequencing reads merged and aligned to an amplicon of interest and a percentage of total reads that matched an intended edit via a 2-sided replacer 2 with a nicking Cas9 and a T4 DNA ligase.
[0067] Fig. 12 illustrates measurement of a percentage of cells expressing green fluorescent protein (GFP), indicating gene editing from BFP to GFP via a 1-Sided Replacer 2 with a nicking Cas9 and a T4 DNA Ligase.
[0068] Fig. 13A and Fig. 13B depict example fusion proteins, which may be useful for a genome revising system. In some embodiments, an endonuclease, DNA ligase, and integrase can be delivered as three individual proteins (not illustrated). An endonuclease, DNA ligase, and integrase can be delivered as a combination of one double fusion protein and a third protein delivered in trans (Fig.13A). An endonuclease, DNA ligase, and integrase can be delivered as one triple fusion protein (Fig. 13B). A linker such as a peptide linker may be included between any two of the components in each fusion protein, or additional components may be added.
[0069] Fig.14A-14D depict vectors that can be used to deliver the second integrating nucleic acid (e.g., revising nucleic acid) to the host cell. The vector can contain a single integrase attachment site (Fig.14A and Fig.14B) or a plurality of integrase attachment sites (Fig.14C and Fig. 14D). The vector can be a minicircle (Fig. 14A), a plasmid (Fig. 14B and Fig. 14C), or linearized DNA (Fig.14D).
[0070] Fig. 15 shows an exemplary nucleic acid gel having a pattern associated with the replacement of either a 222-base pair (bp) sequence (lanes 1 and 2) or a 131 bp sequence (lanes 4 and 5) with a 38 bp attB recombination sequence. Lanes 3 and 6 represent the deletion, without replacement, of the 222 bp sequence and the 131 bp sequence, respectively. The lane on the far left represents the non-edited sequence. DM2\21326271.123PATENT Docket No. J4040-99018
[0071] Fig. 16 depicts amplicon sequencing data that quantifies the percentage of reads in a pool of cells that contain the expected precise edit of replacement of a 222 bp or a 131 bp sequence with an attB recombination sequence. The bars labeled numbers 1-6 represent the same reactions as those in Fig.15.
[0072] Fig.17 illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus. Illustrated on the donor strand is a 3' chemical modifications (e.g., a C3 spacer or an inverted dT), and illustrated on the splint are additional chemical modifications (e.g., 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer or 3' inverted dT on the splint.)
[0073] Fig. 18 provides sequencing data that quantifies the percentage of reads in a pool of cells, showing the editing efficiency of the ATPase copper transporting beta (ATP7B) gene using genome revising systems comprising splints having different end blocking modifications.
[0074] Fig.19 provides sequencing data that quantifies percentages of reads in a pool of cells, showing the editing efficiency of a cystic fibrosis transmembrane conductance regulator (CFTR) gene using genome revising systems comprising splints having different end blocking modifications.
[0075] Fig.20 illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus. In this illustration, a biotinylated splinting nucleic acid is attached to a monomeric streptavidin fused to nicking Cas9. The monomeric streptavidin can also be fused to the ligase (not illustrated).
[0076] Fig.21 provides example fusion proteins, which may be useful for a genome revising system. In some embodiments, a DNA-binding domain of the Rad51 DNA repair protein (rad51DBD), and / or a high-mobility group nucleosome binding domain 1 (HN1) and a histone H1 central globular domain (H1G) are fused to the fusion protein.
[0077] Fig. 22 provides sequencing data that quantifies the percentage of reads in a pool of cells, showing the editing efficiency of the CFTR gene using genome revising systems comprising fusion proteins comprising, rad51DBD and / or HN1 and H1G.
[0078] Fig. 23 provides sequencing data that quantifies the percentage of reads in a pool of cells, showing the editing efficiency of the CFTR gene using genome revising systems comprising bicistronic mRNAs encoding for phosphomimetic peptide from IGF1(IGF1pm1) and N-terminal peptide from NFATC2IP (NFATC2IPp1) peptides (IN peptides). DM2\21326271.124PATENT Docket No. J4040-99018
[0079] Fig.24 illustrates a model of a G to T point mutation mediated by a 1-sided Replacer mechanism comprising a donor nucleic acid and a splint nucleic acid. The figure depicts displacement of a 5’ flap of the target DNA, ligation of a donor nucleic acid comprising a single- base substitution to a 3’ flap of the target DNA flap, and resolution of the mismatch by a mismatch repair (MMR) pateway.
[0080] Fig.25 provides a model of ligase- mediated programmable gene integration (PGI) or L-PGI. The model depicts replacement or deletion mediated by a 2 x 1-sided Replacer. (See also, Fig. 1A and Fig. 3A) Target modifications (replacements or deletions) are determined by nick location and the sequences of the overlapping flaps formed by donor nucleic acids. Target modifications are made with single base pair resolution.
[0081] Fig.26 illustrates the effects of single-stranded overhangs on stability of a donor-splint nucleic acid duplex at a physiological temperature (left panel). Affinity of complementary sequences can be modulated by incorporation of locked nucleic acids (LNAs) (right panel). Particularly in cases where the donor nucleic acid, but not the splinting nucleic acid is integrated at the target (see, e.g. Fig.1A, 2A, 3A, 5A, 6A), the splint can be freely modified.
[0082] Fig.27 illustrates a Replacer programmed to mediate three point mutations to mutate a target coding sequence to encode GFP and disrupt the PAM.
[0083] Fig. 28 illustrates optimization of guide / splint duplex region formed by the splint biding site (SBS) of the guide and guide binding site (GBS) of the splint (top panel). The system can be optimized by adjusting the GC content and length of the SBS / GBS region (bottom left). The optimal length of the SBS / GBS region in the HEK293T GFP example is about 19 bp (bottom right). In this case, there appears to be no substantial benefit from inclusion of linker nucleotides between the flap binding site (FBS) and GBS (bottom right).
[0084] Fig 29 illustrates optimization of the flap binding site (FBS) (top panel) and donor / splint DNA dose relative to the amount of gRNA. In the HEK293T GFP example, the optimal FBS is about 9 to 13 bases (bottom left). In the HEK293T GFP example, for a 13 base FBS, the optimal molar amount of splint / donor is about 18% of the gRNA molar amount (bottom right).
[0085] Fig. 30 illustrates optimization of the donor and DBS length (top panel). In the HEK293T GFP example, there is substantial activity over a broad range of donor lengths and an DM2\21326271.125PATENT Docket No. J4040-99018 optimal length of about 22-26 nt (bottom left). Overhangs of donor or splint nucleic acids were observed to reduce efficiency (bottom right).
[0086] Fig.31 illustrates optimization of locked nucleic acids (LNAs). Fig.31A depicts effect of varying the number of LNAs in the splint within the first 20 nt of the donor binding site (DBS) alternating from the 5’ end of the splint and within the 19 nt of the guide binding site (GBS) at the 3’ end of the splint. Fig.31B depicts effect of inclusion of additional LNAs in the splint near the target nick site or at the 5’ end of the splint. Fig.31C depicts the effect of varying the number of LNAs in the donor binding site (DBS) of the splint for a 32 nt DBS. The comparison is among 12 LNAs distributed towards the target nick site, 12 LNAs distributed towards the 5’ end of the splint, and 16 alternating LNAs. Fig. 31D (SEQ ID Nos:815,556 and 816) depicts the locations and exemplary compositions of the nucleic acid binding sites for a splint with a 20 nt DBS and 19 nt GBS. (Donor: SEQ ID NO:556, / 5Phos / cgtaTgtcagggtggtcacGACgg; Splint: SEQ ID NO:557, +C*C*+GT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGgcgtgcagtgcttACGCCA+CA+AT+A C+CG+CA+G*C*+T; Guide: 3’ end of guide, e.g. SEQ ID NO:166, SEQ ID NO:170, SEQ ID NO:173, or SEQ ID NO:222).
[0087] Fig. 33. Replacer efficiency can be enhanced by incorporation of nucleotide analogs such as pseudo-UTP enhance (Fig.33A). A variety of ligases are effective. SplintR and T4 ligase demonstrated particularly high efficiency (Fig. 33B, C). Expressing nCas9 and T4 ligase from separate mRNAs resulted in higher efficiency (left panel, split nCas9 & T4, 2 mRNAs). This may be due to the size of the mRNAs and not the protein itself, as a T4-P2A-nCas9 mRNA expressing a self-cleaving fusion (Fig.33C) had the same efficiency to a T4-nCas9 fusion.
[0088] Fig. 32 illustrates functional aspects of the system and components, including homology-directed repair (HDR)-independence. The high efficiency of Replacer editing is not primarily due to HDR (Fig.32A). The efficiency of Replacer editing depends on all components, splint, donor, Cas9, and ligase. The exogenous ligase substantially increases ligase activity that may be present in target cells (Fig.32B).
[0089] Fig. 34 illustrates gene editing of several endogenous target genes in HEK293T cells prior to optimization. Point mutations were introduced at reportedly high efficiency targets (e.g., AAVS1 (a safe harbor site; Anzalone et. al., Nature biotechnology 2022), and VEGFA (Wang et. al., Nature methods 2022) and at specific disease-relevant locations (HBB E6V for sickle cell disease; CFTR R553X & G551D for cystic fibrosis; ATP7B H1069Q for Wilson’s disease). DM2\21326271.126PATENT Docket No. J4040-99018
[0090] Fig. 35 illustrates effects of nCas9 modifications that enhance editing efficiency. Inclusion of chromatin-modifying peptides (CMPs) HN1 and Rad51 DNA binding domain (DBD) in fusions with nCas9 increased efficiencies.
[0091] Fig.36 illustrates effects of donor DNA methylation and 3’ end protection. Fig.36A: Donor methylation. Fig. 36B: Donor methylation and 3’ end protection with and without splint LNA modification. Fig. 36C: Donor methylation with O-Me splint modification or LNA splint modification.
[0092] Fig. 37 illustrates effects of splint 3’ end protection on editing efficiency. Fig. 37A: Chemically blocking the 3’ end of splints increases efficiency. The effect is shown for 3’ AltR and a 3’ C3 spacer. Fig.37B: Splints linked at the 3’ end to biotin were tested with T4 ligase linked to streptavidin. Fig. 37C (SEQ ID Nos:815,558 and 816) depicts the locations and exemplary compositions of the nucleic acid binding sites for a splint with a 20 nt DBS and 19 nt GBS. (Donor: SEQ ID NO:558, / 5Phos / mCgtaTgtmCagggtggtmCamCGAmC*g*g-C3 spacer; Splint: SEQ ID NO:559, +C*C*+GT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGgcgtgcagtgcttACGCCA+CA+AT+A C+CG+CA+G*C*+T-C3 spacer; Guide: 3’ end of guide, e.g. SEQ ID NO:166, SEQ ID NO:170, SEQ ID NO:173, or SEQ ID NO:222).
[0093] Fig.38 illustrates effects of nicking guides (Fig.38A, B) and dead guides (Fig.38B). Nicking guides nick the opposite strand to promote incorporation of desired edits upon DNA mismatch repair (MMR). Nicking guides can be screened to achieve improved efficiency with minimum indels. Dead guides include a spacer (e.g. a 15-nt spacer) that allows nCas9 to engage a target without nicking, opening chromatin in the region of the target, and are used with additional target strand spacers (Park et. al., Genome Biology 2021).
[0094] Fig. 39 illustrates high efficiency editing of disease-relevant mutations. Replacer optimization included one or more of: fusing additional peptides to nCas9; chemically blocking ends of donors and splints; methylating donor DNA; including nicking guides; screening different FBS lengths; removing the splint’s 3’ LNA (HBB only); and changing the gRNA scaffold.
[0095] Fig. 40 illustrates Replacer high efficiency editing and precision compared to prime editing PE2, including without use of a nicking guide. Replacer demonstrates substantially reduced indel formation due from RNA synthesis errors, reverse transcriptase errors, and scaffold integration. A: Edited and modified reads (e.g., indels) as a proportion of all edited reads at five DM2\21326271.127PATENT Docket No. J4040-99018 loci. B: Editing efficiency at the five loci. C: Indels by location, Replacer vs. PE. Replacer produces substantially fewer indels over the range of bases.
[0096] Fig. 41 illustrates 1-sided and 2-sided Replacer deletion of disease-associated repeat regions. Top Left: HTT in-frame 96bp deletion of repeat region by 1-sided Replacer showing efficiency of deletion and low frequency of indels. Top Right: 2-sided Replacer vs Prime on deletion of 131bp region of C9orf72 and 38bp Bxb1 attB replacement. Bottom: Efficient replacement (lanes 1 and 2) or deletion (lane 3) of C9orf72 by 2-sided Replacer.
[0097] Fig. 42 illustrates effects of overlap length (A-C)) and base modification (C) in a 2 x 1-sided Replacer (e.g., as diagrammed in Fig. 25). The direction and length of the attB sequence, as well as the overlap length between the 2 Replacer donors, affects edit efficiency (B). Changing some LNAs in the splint to 2’-OMe bases increases efficiency for longer donor overlaps (e.g.38 bp donors) (C).
[0098] Fig.43 illustrates high Replacer efficiency in cells, including in non-dividing cells. Fig. 43A depicts 2-sided Replacer compared to prime editing in human primary hepatocytes (PHH), induced pluripotent stem cells (iPSCs) and HEK293T cells. Replacer efficiency can be optimized by LNP formulation. Fig.43B illustrates VEGFA 175 bp replacement with a 2 bp “GT” in three cell lines using three different LNP formulations.
[0099] Fig. 44 illustrates Replacer editing efficiency in human hepatocytes. A: ATP7B mutations in PHH primary human hepatocytes and HEK293T immortalized human embryonic kidney cells. B. Replacement of NOLC 79 bp sequence with Pa01 attB 33 bp sequence in PHH and HEK293T cells.
[0100] Fig. 45 illustrates effects of ligase and nCas9 order in fusion protein and split expression.
[0101] Fig.46 illustrates models of a ligase-mediated programmable genomic integration (L- PGI) system (Fig.46A) and a 2-Sided L-PGI system (Fig.46B). As depicted, the components of an L-PGI system comprise mRNA encoding nCas9 and DNA ligase, an ssDNA donor template, an ssDNA splint, and a guide RNA (lmgRNA) suitable for ligase mediated programable genomic integration. Fig.46C exemplifies relative efficiency of beacon placement for 1-Sided and 2-Sided L-PGI in human hepatocytes. DM2\21326271.128PATENT Docket No. J4040-99018
[0102] Fig.47 illustrates optimization of ligase-mediated programmable genomic integration (L-PGI) at F9 and PAH loci in PHH (Fig.47A), PCH (Fig.47B) and PMH (Fig.47C) by variation of dose of donor:splint (used in 1:1 molar ratio).
[0103] Fig.48 PGI (L-PGI, right) compared with integrase-mediated PGI (I-PGI, left) in PHH and PMH. L-PGI editing can be used for integrative gene therapies for diseases where a single mutation affects most patients or for exon replacement. L-PGI edits are efficient in primary cells as they can perform scarless editing with an edit size of ~1-100 bps, but cannot be delivered virally. I-PGI editing can be used for Integrative gene therapies for diseases with heterogenous causal mutations and for advanced cell engineering (e.g. logic gates, multi-CARs, etc.). L-PGI edits are efficient in primary cells as they can be used with any writing enzyme, they can perform edits of 1,000 bps+, and can be delivered virally and non-virally (SEQ ID Nos: 1316-1351).
[0104] Fig.49 illustrates efficient beacon placement by L-PGI at hF9, hALB and hPAH loci in PHH (Fig.49A) and PMH (Fig.49B).
[0105] Fig.50 illustrates beacon placement by L-PGI and LNP delivery. All-in-one (AIO) or split LNP delivery of L-PGI components improved beacon placement in PHH. AIO L-PGI: nCas9 / DNA ligase mRNA + forward-lmgRNA + forward-donor + forward-splint + reverse-lmgRNA + reverse-donor + reverse-splint; Split L-PGI: LNP1: nCas9 / DNA ligase mRNA + forward-lmgRNA + forward-donor + forward-splint; LNP2: nCas9 / DNA ligase mRNA + reverse-lmgRNA + reverse-donor + reverse-splint.
[0106] Fig. 51 depicts an overview of the Replacer system to edit human cells with a stably integrated BFP reporter that converts to GFP upon successful editing. Fig.51A depicts an example of the Replacer editing mechanism for an A to C transversion edit. The donor DNA is homologous to the genome besides the edited C base. After the donor is ligated onto the 3’ flap, it anneals to the opposite strand of the genome and induces editing through the endogenous mismatch repair pathway. Fig. 51B illustrates a schematic of the Replacer editing workflow: nucleic acids are transfected to HEK293T cells that express a virally integrated BFP gene. Replacer edits 3 nucleotides in the BFP gene to convert it to GFP. Editing efficiency is equal to the percentage of GFP+ cells as assessed with flow cytometry. Fig. 51C illustrates an example architecture of a replacer nucleic acid and chemical modification layout. The splint includes alternating LNAs on in the DBS and GBS and the ligRNA includes three 2’-OMe nucleotides on the 5’ and 3’ ends. The Replacer splint composed of a donor binding site (DBS) that is complementary to the donor DM2\21326271.129PATENT Docket No. J4040-99018 DNA, flap binding site (FBS) that hybridizes to the nicked 3’ flap, and guide binding site (GBS) that connects to the ligRNA. The ligRNA contains a typical Cas9 spacer and scaffold as well as a splint binding site (SBS) on the 3’ end that is bound to the splint. The splint and ligRNA may incorporate modifications, not limited to phosphorothioate backbone modifications (PS), locked nucleic acids (LNAs), and methylated nucleotides.
[0107] Fig.52 displays a comparison of the editing efficiencies of nucleic acids with different number of alternating LNAs in the DBS and GBS regions (Fig. 52A), different FBS length and DNA dose using 2.1 pmol ligRNA for transfection (Fig.52B), and different donor and DBS length (Fig. 52C). Fig.52A(SEQ ID Nos:208,560-563) depicts effects of LNAs incorporated into the DBS and GBS portion of the splint. Splints from top to bottom: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+A C+CG+CA+G*C*+T (LNAs in DBS / GBS - 10 / 7) (SEQ ID NO:208); +C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGC+CA+CA+AT+A C+CG+CA+G*C*+T (LNAs in DBS / GBS - 9 / 8) (SEQ ID NO:560); +C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGCCA+CA+AT+AC +CG+CA+G*C*+T (LNAs in DBS / GBS - 9 / 7) (SEQ ID NO:561); +C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGCCACA+AT+AC+ CG+CA+G*C*+T (LNAs in DBS / GBS - 9 / 6) (SEQ ID NO:562); +C*G*+TG+AC+CA+CC+CT+GA+CATACGGCGTGCAGTGCTTACGC+CA+CA+AT+AC +CG+CA+G*C*+T (LNAs in DBS / GBS - 8 / 8) (SEQ ID NO:563). The right portion of Fig.52A depicts % GFP fluorescent cells in the transfected cells. Fig. 52B depicts effects of splint FBS length and DNA amount on editing efficiency of the HEK293T system. Fig.52C depicts the BFP to GFP conversion efficiency based on donor and DBS length.
[0108] Fig.53 displays the comparative BFP to GFP editing efficiencies using either a 100-nt ssODN donor for HDR or the splint and donor DNA used for Replacer, with different mRNAs. The system optimally comprises splint, donor DNA, Cas9n, T4 ligase, and the 5’phosphate. Fig. 53A illustrates the effects of subtracting system components.5’ phosphorylation of donor DNA is not critical in the HEK293T system. Also, while HEK293T cells display endogenous ligase activity, the level is substantially increased by T4 ligase. Fig. 53B illustrates comparative efficiency of the Replacer system using Cas9 (Cas9 nuclease), Cas9n (H840A Cas9 nickase) or Cas9n + T4 DNA ligase relative to a 100 nt single-stranded oligodeoxynucleotide (ssODN) donor DM2\21326271.130PATENT Docket No. J4040-99018 template with Cas9. Fig.53C illustrates comparative efficiency of separate Replacer mRNAs on BFP to GFP conversion, using pseudouridine (m1Ψ) in either an nCas9-T4 fusion or dual (nCas9 and T4) mRNA system.
[0109] Fig.54(SEQ ID Nos:564,817 and 818) illustrates replacer editing of point mutations at endogenous genomic loci in human HEK293T cell line (Fig.54A) using splints with alternating locked nucleic acids (LNAs) in the GBS and DBS regions (Fig. 54B). An example chemical modification layout of an efficient system: Splint: +A*G*+GC+CA+GC+AG+TG+AA+CA+AC+CA+TT+GGGCGTGGCAGTACGCCA+CA+A T+AC+CG+CA+G*C*+T-c3 SPACER (SEQ ID NO:564); Donor DNA: / 5Phos / / IME- DC / AATGGTTGTT / IME-DC / A / IME-DC / TG / IME-DC / TGG / IME-DC / / IME-DC / *T-c3 SPACER (SEQ ID NO:565); SBS portion of ligRNA: AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO:566).
[0110] Fig.55 displays the editing efficiencies of splints with alternating locked nucleic acids (LNAs) in the GBS and DBS regions in HBB (Fig.55A), in CFTR (Fig.55B), and in ATP7B (Fig. 55C).
[0111] Fig. 56 illustrates the effect of adding additional nicking guides on the BFP to GFP editing efficiencies and indels made using the splits in Fig.55 in HBB (Fig.56A), in CFTR (Fig. 56B), and in ATP7B (Fig. 56C). Location represents the direction and number of nucelotides between the 2 nicks
[0112] Fig. 57 illustrates the editing efficiencies and indels at different doses of splint and donor DNA as compared to ligRNA for different FBS lengths in HBB (Fig.57A), in CFTR (Fig. 57B), and ATP7B (Fig.56D), and in AAVS1 (Fig.57D). Every nCas9 mRNA is used with a T4 ligase mRNA.
[0113] Fig.58 illustrates the editing efficiencies of splints with or without a C3 spacer on the 3’ end of the splint in HBB (Fig.58A), in CFTR (Fig.58B), and in ATP7B (Fig.58C).
[0114] Fig. 59 illustrates the editing efficiencies of splints with several DNA modifications: with or without a C3 spacer on the 3’ end, phosphorothioate (PS) bonds and a methylated cytosine (meC) in HBB (Fig.59A), in CFTR (Fig.59B), and in ATP7B (Fig.59C).
[0115] Fig.60 illustrates a schematic of a 1 sided Replacer editing system (Fig.60A). In a 1 sided Replacer editing system, the splint links together ligase-mediated guide RNA (lmgRNA) and Donor DNA, nicks Cas9 complexes with lmgRNA and nicks the genomic DNA to open the R- DM2\21326271.131PATENT Docket No. J4040-99018 loop, making the opposite stand’s flap accessible. The splint then binds to the genomic flap and positions the Donor at the nick, the Ligase installs the Donor to the genomic DNA flap, the L-PGI components dissociate, leaving only the donor DNA as a permanent edit. Fig 60B illustrates a 2 sided Replacer system editing, where there are 2 full Replacer editing complexes targeting protospacers on opposite strands of DNA. Fig. 60C illustrates an example outline of an editing process: nucleic acids transfections include 2 sets of splint, donor DNA, and ligRNA. 72 hours after transfection, genomic DNA is extracted for PCR amplification and amplicons can be visualized on a gel or sequenced to assess editing efficiency.
[0116] Fig.61 illustrates design strategies for a 2-sided Replacer, where donor DNAs can be homologous to each other to make a replacement edit or homologous to the genome to make a deletion edit: a replacement with full overlap where the edited flaps bind each other completely (Fig. 61A), partial overlap where the 3’ segments of the edited flaps bind each other (Fig 61B), and a deletion (Fig.61C).
[0117] Fig. 62 illustrates 1-sided and 2-sided Replacer edit of disease-associated repeat regions. Fig.62A illustrates an agarose gel image of PCR amplicons for C9ORF72 after Replacer deletion of 222bp and 131bp regions of C9orf72. The WT fragment is non-edited and the fragments on the bottom are the desired replacement or deletion. attB and attB rev comp are replacements of the nick-nick region with the Bxb1 attB site in forward or reverse orientation. del F + R is a deletion, del F is the deletion without the Rev splint and donor, and del R is the deletion without the Fwd splint and donor Fig.62B illustrates an agarose gel image of PCR amplicons for the Bxb1 attB replacement edit of a 131bp region of C9ORF72, comparing donor DNAs with different chemical modifications. The fragment of the desired edit is on th bottom.
[0118] Fig. 63 illustrates the editing efficiencies for attB replacement edits at 2 targets as quantified with NGS of amplicons. T4 ligase mRNA is used alongside the nCas9 mRNAs that are shown. Fig. 63A shows the replacement of NOLC1 79 bp sequence with Pa01 attB 33 bp sequence, Fig.63B shows the VEGFA replacement of 38 bp Bxb1 attB site, and Fig.63C (SEQ ID Nos:567-576)illustrates the effect on editing efficiencies of of splints that have different overlap length between forward and reverse splints and variation in the number and chemical modifications (LNAs and 2’-OMe) in 2-sided Replacer system showing precise editing vs. indel production, 10 being partial overlap and 38 being full overlap. Each splint is used with a donor DNA of the same length as its DBS, and all splints include the same FBS and GBS that are not DM2\21326271.132PATENT Docket No. J4040-99018 shown. Efficiencies are assessed with NGS of amplicons. Top to bottom: 10 bp splint overlap with LNAs (12 of 24), forward splint DBS = +G*G*+AG+AC+CG+CC+GT+CG+TC+GA+CA+AG+CC (SEQ ID NO:567), Reverse splint DBS = +G*G*+CG+GT+CT+CC+GT+CG+TC+AG+GA+TC+AT (SEQ ID NO:568); 10 bp splint overlap with LNAs (7 of 24) and 2’O-Me nucleotides (5 of 24), Forward splint DBS = +G*G*+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+CC (SEQ ID NO:569), Reverse splint DBS = +G*G*+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+AT (SEQ ID NO:570); 38 bp splint overlap with LNAs (19 of 38), Forward splint DBS = +A*T*+GA+TC+CT+GA+CG+AC+GG+AG+AC+CG+CC+GT+CG+TC+GA+CA+AG+CC (SEQ ID NO:571), Reverse splint DBS = +G*G*+CT+TG+TC+GA+CG+AC+GG+CG+GT+CT+CC+GT+CG+TC+AG+GA+TC+AT (SEQ ID NO:572); 38 bp splint overlap with LNAs (10 of 38) and 2’-OMe nucleotides (9 of 38), Forward splint DBS: = +A*T*mGA+TCmCT+GAmCG+ACmGG+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+ CC (SEQ ID NO:573), Reverse splint DBS = +G*G*mCT+TGmUC+GAmCG+ACmGG+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+A T (SEQ ID NO:574); 38 bp splint overlap with LNAs (13 of 38) and 2’-OMe nucleotides (6 of 38), Forward splint DBS: = +A*T*+GA+TC+CT+GA+CG+ACmGG+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+C C (SEQ ID NO:575), Reverse splint DBS: = +G*G*+CT+TG+TC+GA+CG+ACmGG+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+AT (SEQ ID NO:576).
[0119] Fig. 64 illustrates replacer editing in primary human hepatocytes. Comparison of Replacer and the PE2 or PEMax prime editing systems in HEK293T cells and primary human hepatocytes (PHH) for the point mutations in ATP7B (Fig. 64A) and for the replacement of NOLC1 79 bp sequence with Pa01 attB 33 bp sequence (Fig.64B). The PBS and RTT sequences for prime editing match Replacer’s splint FBS and DBS sequences, respectively. Editing efficiency and indel generation is calculated by analyzing NGS of amplicons
[0120] Fig.65 illustrates optimization of chemical modifications in the splint and donor DNA for Replacer editing at BFP. Fig.65A (SEQ ID Nos:577-583) depicts variation in the number of LNAs in a 21nt DBS region of a splint. The control splint has alternating LNAs in the DBS, which DM2\21326271.133PATENT Docket No. J4040-99018 is typically optimal for a 21-nt DBS. Top to bottom: +T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT+ AC+CG+CA+G*C*+T (SEQ ID NO:577); +T*+C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (SEQ ID NO:578); +T*+C*+G+T+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+A T+AC+CG+CA+G*C*+T (SEQ ID NO:579 ); +T*C*+GT+GA+CC+AC+CC+TG+AC+AT+A+C+GGCGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (SEQ ID NO:580); +T*C*+GT+GA+CC+AC+CC+TG+AC+A+T+A+C+GGCGTGCAGTGCTTACGCCA+CA+A T+AC+CG+CA+G*C*+T (SEQ ID NO:581); +T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+G+GCGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (SEQ ID NO:582); +T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GG+CGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (SEQ ID NO:583). Fig.65B depicts a comparison of different chemical modifications(SEQ ID Nos:208,584-588) (LNA, 2’-OMe, 2F) in splint DBA and GBS regions. Top to bottom: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+AC +CG+CA+G*C*+T (LNA / LNA) (SEQ ID NO:208); +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA / i2FC / A / i2FA / T / i2FA / C / i2FC / G / i2FC / A / i2FG / *C* / 32FU / (LNA / 2’-F) (SEQ ID NO:584); / 52FC / *G* / i2FU / G / i2FA / C / i2FC / A / i2FC / C / i2FC / T / i2FG / A / i2FC / A / i2FU / A / i2FC / GGCGTGCA GTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (2’-F / LNA) (SEQ ID NO:585); +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCAmCAmATmA CmCGmCAmG*C*mT (LNA / 2’-OMe) (SEQ ID NO:586); mC*G*mTGmACmCAmCCmCTmGAmCAmTAmCGGCGTGCAGTGCTTACGCCA+CA+AT +AC+CG+CA+G*C*+T (2’-OMe / LNA) (SEQ ID NO:587); C*G*TGACCACCCTGACATACGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G* C*+T (DNA / LNA) (SEQ ID NO:588). Fig.65C(SEQ ID Nos:589,590 and 591) illustrates the effects of varying the number and the location of LNAs in a 32nt DBS of a splint. Top to bottom: 16 LNAs of 32 DM2\21326271.134PATENT Docket No. J4040-99018 (+C*T*+GG+CC+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTG CTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:589)); 12 LNAs of 32, reduced near nick (+C*T*+GG+CC+CA+CC+CTC+GTG+ACC+ACC+CTG+ACA+TA+CGGCGTGCAGTGCT TACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:590)); 12 LNAs of 32, reduced near 5’ end (+C*T*+GGC+CCA+CCC+TCG+TGA+CCA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCT TACGCCA+CA+AT+AC+CG+CAG*C*T (SEQ ID NO:591)). Fig. 65D illustrates the effect of methylated DNA donors on BFP to GFP conversion efficiency, dependent on chemical modifications in the splint DBS.
[0121] Fig. 66A(SEQ ID Nos:592-594) illustrates a comparison of different pairs of splint GBS and ligRNA SBS sequences with 3’ end modifications of ligRNA. SBS folding ΔG is the strength of secondary structure in the ligRNA’s 20-nt 3’ extension (SBS). All splints include the same DBS and FBS (not shown) and ligRNAs include identical scaffolds and spacers (not shown). Top to bottom: AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO:592); GUGGUUCCGGGCUGCAmU*mG*mA (SEQ ID NO:593); CGAUUCCUGAUACUGCmU*mG*Mc (SEQ ID NO:594).
[0122] Fig. 66B (SEQ ID Nos:595,818,597,822,595,822,600,822,601,823,601,824,601,825,600,825,600,826,595 and 818) illustrates Replacer optimization, including splint optimization by GBS / SBS length, splint composition, and FBS linker, and ligRNA optimization. Nucleotide sequences are shown as they are bound to each other and in some cases a part of the GBS or SBS functions as a single stranded linker. All splints include the same DBS and FBS (not shown) and ligRNAs include identical scaffolds and spacers (not shown). Top to bottom. Length of GBS / SBS (nt) = 24 / 19: ACGTCACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:595) and AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO:596); Length of GBS / SBS (nt) = 15 / 15: +CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:597) and AGCUGCGGUAUUmG*mU*mG (SEQ ID NO:598); Length of GBS / SBS (nt) = 19 / 15: ACGC+CA+CA+AT+AC+CG+CA+G*C*T (SEQ ID NO:599) and AGCUGCGGUAUUmG*mU*mG (SEQ ID NO:598); Length of GBS / SBS (nt) = 20 / 15; GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:600) and AGCUGCGGUAUUmG*mU*mG (SEQ ID NO:598); Length of GBS / SBS (nt) = 28 / 40: DM2\21326271.135PATENT Docket No. J4040-99018 GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (SEQ ID NO:601) and CGATTTCCTGATAGTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:602); Length of GBS / SBS (nt) = 28 / 32: GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (SEQ ID NO:601) and GATAGTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:603); Length of GBS / SBS (nt) = 28 / 28; GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (SEQ ID NO:601) and GTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:604); Length of GBS / SBS (nt) = 20 / 28: GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:600) and GTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (SEQ ID NO:604); Length of GBS / SBS (nt) = 20 / 20: GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:600) and AGCUGCGGUAUUGUGGCmG*mU*mC (SEQ ID NO:605); Length of GBS / SBS (nt) = 19 / 19: ACGC+CA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:595) and AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO:606).
[0123] Fig. 66C (SEQ ID Nos:607,810,609,620,609,819,607,821,607,820) illustrates Replacer optimization, including splint optimization of DBS length and DBS / donor overhang. When the lengths of the splint DBS and donor DNA are not equal, there is a ssDNA overhang. All splints include the same DBS and FBS (not shown). Top to bottom. Length of DBS / donor (nt) = 20 / 28 nt: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:607) and / 5Phos / CGTATGTCAGGGTGGTCACGAGGGTGGG (SEQ ID NO:608); Length of DBS / donor = 28 / 20 nt: +C*C*+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:609) and / 5Phos / CGTATGTCAGGGTGGTCACG (SEQ ID NO:610); Length of DBS / donor = 28 / 28 nt: +C*C*+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:609) and / 5Phos / CGTATGTCAGGGTGGTCACGAGGGTGGG (SEQ ID NO:608); Length of DBS / donor = 20 / 21 nt: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:607) and / 5Phos / CGTATGTCAGGGTGGTCACGA (SEQ ID NO:611); Length of DBS / donor = 20 / 20 nt +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (SEQ ID NO:607) and / 5Phos / CGTATGTCAGGGTGGTCACG (SEQ ID NO:610).
[0124] Fig. 67 depicts the evaluation of ligases and protein architecture for Replacer editing at BFP. Fig.67A depicts the comparison of 3 different ligases as either fusion proteins with nCas9 encoded in a single mRNA, or as separate mRNAs used with nCas9 mRNA. When 2 mRNAs are used, nCas9 and the ligase are fused to leucine zippers (LZs) on either the N-terminal or C-terminal DM2\21326271.136PATENT Docket No. J4040-99018 to promote colocalization. A 24-nt donor DNA & splint DBS were used to convert BFP to GFP. Fig. 67B depicts the comparison of the 3 ligases to no ligase, a T4-nCas9 fusion, and T4-P2A- nCas9 bicistronic mRNA for conversion of BFP to GFP with a 50-nt donor & DBS. Fig. 67B depicts the effect of LZs on BFP to GFP conversion efficiency. Each ligase mRNA is delivered with nCas9 mRNA without LZs on either protein (left bar), with LZs on the ligase C-terminus and nCas9 N-terminus (middle bar), or with LZs on the ligase N-terminus and nCas9 C-terminus (right bar).
[0125] Fig.68 illustrates a screen of different DNA ligases using ddPCR. Ligases were fused to N- or C-terminus of fusion construct with Cas9 nickase, a ddPCR assay was performed, and data was normalized by mRNA purity readout using fragment analyzer (%BP / %mRNA purity).
[0126] Fig.69 illustrates ligase-mediated programmable genomic integration (L-PGI) results in a Mouse model. Fig.69A displays total editing, Fig.69B displays the perfect beacon placement, and Fig.69C displays beacon fidelity results.
[0127] Fig.70 shows partial edit fidelity of I-PGI vs L-PGI from the same study as depicted in Fig.69. Partial edit is defined as the insertion of complete donor sequence or reverse transcribed sequence in either forward or reverse orientation. Fig.70 illustrates the donor insertion fidelity of L-PGI at different ratios in forward and reverse reads.
[0128] Fig. 71 illustrates the precise editing and indels of L-PGI across different species: in primary mouse hepatocyte (Fig. 71A), in primary Cyno hepatocyte (Fig. 71B), primary human hepatocyte (Fig. 71C), and the fidelity of L-PGI across different species: primary mouse hepatocyte (Fig.71D), in primary Cyno hepatocyte (Fig.71E), primary human hepatocyte (Fig. 71F).
[0129] Fig. 72 illustrates nucleotide sequences of splints in L-PGI systems shown in Fig.71 in primary human hepatocyte (PHH, Fig.72A), in primary cyno hepatocytes (PCH, Fig.72B), and primary mouse hepatocyte (PMH, Fig.72C) (SEQ ID Nos:1352-1449).
[0130] Fig.73 illustrates deletion, replacement and conversion with L-PGI in HEK293T cells (Figs.73A-C) and Primary hepatocytes (Figs.73D-F) of a 2-sided deletion (Figs.73 A and D), a 2-sided beacon replacement (Figs.73 B and E), and a 1-sided correction / conversion (Figs.73 C and F).
[0131] Fig.74 illustrates an overview of the L-PGI editing mechanism, biochemical proof of concept, and optimizations in a HEK293T GFP reporter cell line. Fig. 74A depicts a schematic DM2\21326271.137PATENT Docket No. J4040-99018 diagram illustrating the L-PGI editing complex, including nCas9 nickase, lmgRNA, splint DNA, donor DNA, and DNA ligase interacting at the target genomic site. Fig. 74B illustrates the proposed mechanism following ligation, including donor strand flap displacement, equilibration, and resolution via mismatch repair. Fig. 74C displays a denaturing PAGE gel validating the L- PGI strategy in vitro using recombinant proteins and a 5'-Cy5 labeled DNA substrate, showing the expected 200-nt ligation product only when all components (nCas9, lmgRNA, splint, donor, T4 ligase) are present (Lane 3). Fig.74D depicts optimization of the Donor Binding Site (DBS) length on the splint, showing peak %GFP+ conversion between 24-30 nt. Fig.74E depicts optimization of the Flap Binding Site (FBS) length on the splint, showing peak %GFP+ conversion between 11-13 nt. Fig. 74F (SEQ ID Nos: 1208-1222) depicts optimization of the Guide Binding Site (GBS) on the splint and the Splint Binding Site (SBS) on the lmgRNA, showing efficiencies for various length combinations and sequences, including modified nucleotides (N=LNA, N=2'-OMe, N=RNA, N=DNA, *=PS bond). Exemplary interacting sequences shown include: (GBS / SBS length 24 / 19): 5’-Splint GBS-3’: ACGTCA C GCCACAATACCGCAG*C*T (SEQ ID NO:1208); 3’-lmgRNA SBS-5’: U*G*CGGUGUUAUGGCGUC G A… (SEQ ID NO: 1209); (GBS / SBS length 15 / 15): 5’-Splint GBS-3’: …C A CAATACCGCAG*C*TT (SEQ ID NO:1210) ; 3’-lmgRNA SBS-5’: G*U*GUUAUGGCGUC G A… (SEQ ID NO:1211); (GBS / SBS length 19 / 19): 5’-Splint GBS-3’: …ACGCC A CAATACCGCAG*C*T (SEQ ID NO:1212); 3’- lmgRNA SBS-5’: G*U*GUUAUGGCGUC G A… (SEQ ID NO:1211); (GBS / SBS length 20 / 20): 5’-Splint GBS-3’: GACGCC A CAATACCGCAG*C*T (SEQ ID NO:1213); 3’-lmgRNA SBS- 5’: G*U*GUUAUGGCGUC G A… (SEQ ID NO:1211); (GBS / SBS length 28 / 28): 5’-Splint GBS-3’: G A CGCCACAATACCGCAG*C*T (SEQ ID NO:1214); 3’-lmgRNA SBS-5’: C*U*GCGGUGUUAUGGCGUC G A… (SEQ ID NO:1215); A C GCCACAATACCGCAG*C*T (SEQ ID NO:1216); U*G*CGGUGUUAUGGCGUC G A…(SEQ IDNO; 1209); …G A CGCCACAATACCGCAGCTGGCAGC*A*C (SEQ IDNO; 1217); C*U*GCGGUGUUAUGGCGUCGACCGTCG T GATAGTCCTTAGC…(SEQ IDNO; 1218); G A CGCCACAATACCGCAGCTGGCAGC*A*C (SEQ IDNO; 1219); C*U*GCGGUGUUAUGGCGUCGACCGTCG T GATAG…(SEQ IDNO; 1220); G A CGCCACAATACCGCAGCTGGCAGC*A*C (SEQ IDNO; 1219); C*U*GCGGUGUUAUGGCGUCGACCGTCG T G…(SEQ IDNO; 1221); G A CGCCACAATACCGCAG*C*T (SEQ IDNO; 1214); C*U*GCGGUGUUAUGGCGUC G DM2\21326271.138PATENT Docket No. J4040-99018 ACCGTCGTG…(SEQ IDNO; 1222); (Note: '+' indicates LNA, '*' indicates phosphorothioate bond, 'm' indicates 2'-OMe. Fig.74G compares editing efficiencies (%GFP+) for various enzyme configurations, including nCas9 fused or co-localized (via Leucine Zippers, LZ; or P2A self- cleaving peptide) with different ligases (T4, SplintR, hLig4 truncations). Fig.74H summarizes the optimal parameters found from the optimization experiments.
[0132] Fig.75 illustrates a comparison of prime editing (PE) versus L-PGI for installing point mutations in 3 disease-relevant loci in primary human hepatocytes (PHH). Fig.75A compares the efficiency (% Edit by NGS) and fidelity (%) for installing a 2 nt H1069Q mutation in ATP7B using L-PGI versus PE2, PE3, and PEMax. Fig. 75B compares engineered nCas9-RT variants versus PE2 mRNA for placing a 38 bp attB site in F9 intron 1 via a twin PE mechanism. Fig.75C compares the efficiency and fidelity for installing a 2 nt C282Y mutation in HFE using L-PGI versus PE-TB2 (an engineered nCas9-RT fusion). Fig. 75D provides a schematic map showing the target window for guides within the APOA1 promoter region. Fig.75E presents an estimation plot comparing the editing efficiency (% Edit by ICE) of 1-3 nt substitutions in the APOA1 promoter region using either PE-TB3 or L-PGI with identical spacer sequences. Fig.75F compares the editing fidelity (%) for a subset of the APOA1 spacers tested in Fig.75E.
[0133] Fig. 76 illustrates a comparison of prime editing (PE) versus L-PGI for larger small corrections (14 nt substitution or insertion) in PHH. Fig.76A provides schematics illustrating the 14 nt edit strategy for sequence substitution versus insertion relative to the nick site. Fig. 76B compares the highest total editing efficiencies (% Edit by NGS) and maximum fidelities (%) observed for 14 nt substitution or insertion in PAH intron 1 using either PE-TB2 or L-PGI after component ratio optimization. Fig. 76C compares the overall lowest indel generation rates observed between PE-TB2 and L-PGI across different edits (HFE 2nt, PAH 14nt substitution, PAH 14nt insertion). Fig.76D shows representative NGS alignment results for the top 15 reads for 14 nt insertions in PAH intron 1 using L-PGI (top panel) and PE-TB2 (bottom panel), highlighting differences in error profiles (SEQ ID Nos: 1223-1251). Fig.76E compares the efficiency (% Edit by NGS) of placing a 38 nt Bxb1 attB site in PAH intron 1 using L-PGI versus PE-TB2.
[0134] Fig.77 illustrates the Paired L-PGI (pL-PGI) design and efficiencies across cell types for attB placement and large excisions. Fig. 77A depicts a schematic diagram of two L-PGI complexes targeting opposite strands simultaneously, using forward (Fwd) and reverse (Rev) components (nCas9, lmgRNA, splint, donor, ligase). Fig.77B shows schematics for two pL-PGI DM2\21326271.139PATENT Docket No. J4040-99018 strategies: replacement using partially overlapping reverse complementary donors (left) and precise deletion using donors homologous to sequences flanking the excision site (right). Fig.77C compares the perfect edit and indel frequencies (% Edit by NGS) for placing a 33 bp Pa01 attB site in NOLC1 of HEK293T cells using paired PE2, paired PEMax, or pL-PGI. Fig.77D shows pL-PGI efficiencies (% Edit by ddPCR) for placing Pa01 attB in NOLC1 of hematopoietic stem cells (HSCs), Bxb1 attB in B2M of induced pluripotent stem cells (iPSCs), and Bxb1 attB plus stop codon in CIITA of iPSCs. Fig.77E compares the efficiencies (% Edit by NGS) of achieving a 175 bp excision at the VEGFA locus in HEK293T, iPSCs, and PHH using either dual PE2 or pL-PGI.
[0135] Fig.78 illustrates pL-PGI attB placement optimizations across therapeutic loci and its implementation for Bxb1-mediated integration in PHH. Fig.78A shows the efficiency (% Edit by NGS) and fidelity (%) of placing a 38 bp Bxb1 attB site in F9 intron 1 using pL-PGI with varying symmetric overlaps (10, 20, 38 nt) between the forward and reverse donors. Fig. 78B compares the efficiency and fidelity of placing a 38 bp Bxb1 attB site using optimized pL-PGI (10 nt overlap) versus dual PE-TB2 (20 nt overlap) at three therapeutic loci in PHH: F9, Albumin (ALB), and PAH. Fig. 78C provides a schematic illustration of the two-step integration process: Step 1 involves pL-PGI-mediated placement of attB in an intron, and Step 2 involves Bxb1 integrase- mediated insertion of a cargo (containing attP) delivered via a vector (e.g., HDAd). Fig.78D shows the integration efficiency (% Edit by ddPCR) measured at the attL and attR junctions and the residual unconverted attB for a 30 kb HDAd cargo, comparing attB sites placed by either dual PE or pL-PGI. Fig.78E compares the attB conversion rate (integration / (integration + residual attB)) between dual PE and pL-PGI strategies across both integration junctions.
[0136] Fig.79 illustrates the translation of pL-PGI across primary cell species in vitro, delivery with lipid nanoparticles (LNP), and efficiency in mice. Fig.79A compares perfect edit and indel efficiencies (% Edit by NGS) for placing a 38 bp Bxb1 attB site in PAH intron 1 using pL-PGI versus dual PE-TB2 in primary human (PHH), cynomolgus monkey (PCH), and mouse (PMH) hepatocytes. Fig.79B compares the fidelity (%) for the same edits shown in Fig.79A. Fig.79C depicts schematics for two LNP formulation strategies for pL-PGI: all-in-one (AIO) with all 8 components in one particle, and split delivery with forward and reverse components in separate particles. Fig.79D shows dose-response curves comparing the potency (% Edit by NGS vs RNA dose) of dual PE-TB2 AIO LNP versus pL-PGI AIO and split LNPs for attB placement in F9 in DM2\21326271.140PATENT Docket No. J4040-99018 PHH, along with calculated EC50 values. Fig. 79E provides a schematic overview of the LNP formulation and intravenous injection process for in vivo mouse studies. Fig. 79F compares in vivo editing efficiency (% Edit by ddPCR) for 38 bp attB placement in mouse PAH using LNPs delivering either dual PE-TB3 or pL-PGI. Fig.79G compares in vivo editing efficiency (% Edit by ddPCR) for a 14 nt insertion in mouse PAH using LNPs delivering either PE-TB3 or L-PGI. Fig.79H shows serum levels of liver toxicity markers Alanine transaminase (ALT) and Aspartate transaminase (AST) in mice 24 hours and 7 days after LNP injection for saline control, dual PE- TB3, and pL-PGI groups.
[0137] Fig.80 illustrates additional optimization experiments in the HEK293T BFP-to-GFP reporter cell line. Fig.80A shows the target BFP sequence and the 3 nt edit required to convert it to GFP. Fig.80B depicts the nucleic acid architecture showing interactions between the lmgRNA, splint, and donor DNA, highlighting chemical modification patterns used, including 5' phosphate and 3' phosphorothioates on the donor, alternating DNA / LNA and terminal phosphorothioates on the splint, and 2'-OMe and terminal phosphorothioates on the lmgRNA. Fig.80C shows %GFP+ efficiency using splints with varying numbers of LNAs in the DBS or GBS regions (SEQ ID NO: 1252). Fig.80D shows the effect of placing additional LNAs in the splint DBS or FBS compared to the optimized alternating LNA pattern (SEQ ID NO: 1253). Fig. 80E evaluates additional enzyme formats, testing different ligases (T4, SplintR, hLIG4) fused or co-localized (via LZs) with nCas9. Fig. 80F shows the effect of using pseudouridine (m1Ψ) versus uridine in the nCas9 / T4 mRNAs for both fusion and split constructs. Fig. 80G compares editing using an HDR ssODN template versus L-PGI oligonucleotides with Cas9, nCas9, or nCas9+T4 ligase. Fig.80H presents a component dropout experiment, assessing %GFP+ efficiency when omitting the splint, donor, nCas9, T4 ligase, or donor 5' phosphorylation.
[0138] Fig. 81 illustrates L-PGI optimizations for placing point mutations at 3 endogenous loci (HBB, CFTR, ATP7B) in HEK293T cells. All experiments aimed to install specific disease- relevant mutations. Fig.81A shows optimization of the splint FBS length for each of the three loci. Fig. 81B shows the effect on precise editing efficiency and indel generation when fusing DNA binding domains (Rad51, HN1, H1G) to nCas9, delivered with separate T4 ligase mRNA. Fig. 81C compares editing efficiencies with or without a 3' C3 spacer modification on the splint for each locus. Fig. 81D shows the effect of adding modifications to the donor DNA, specifically DM2\21326271.141PATENT Docket No. J4040-99018 cytosine methylation (meC), 2x 3' phosphorothioate bonds (2x 3' PS), and a 3' C3 spacer, on precise editing and indel rates.
[0139] Fig. 82 illustrates transfection optimization for 14 nt edits (substitution or insertion) using L-PGI and PE-TB2 in PHH. Fig.82A shows perfect edit and indel efficiencies (% Edit by NGS) for L-PGI at titrated doses of splint and donor oligonucleotides. Fig.82B shows perfect edit and indel rates for PE-TB2 using titrated doses of pegRNA. Fig.82C shows the fidelity (%) of L- PGI edits across the varying splint / donor doses from Fig.82A. Fig.82D shows the fidelity (%) of PE-TB2 edits across the varying pegRNA doses from Fig.82B.
[0140] Fig.83 illustrates the schematic of pL-PGI transfection and detection methods, along with initial efficacy data. Fig. 83A diagrams the experimental workflow in HEK293T cells involving co-transfection of mRNAs and synthetic oligonucleotides, followed by gDNA isolation and analysis by gel electrophoresis, NGS, or ddPCR. Fig. 83B shows an agarose gel of PCR amplicons following pL-PGI editing at the C9ORF72 locus, indicating bands corresponding to WT, desired replacements (attB or attB rev comp), or deletions (del F+R, del F, del R) based on nick-nick distance. Fig. 83C shows an agarose gel comparing pL-PGI replacement edits at C9ORF72 using donors with different chemical modifications (None, meC, 2x 3'PS, meC + 2x 3'PS). Fig.83D shows ddPCR results indicating detectable insertion (% Edit) of a 345 nt sequence using pL-PGI with varying splint:donor doses in PHH.
[0141] Fig. 84 illustrates pL-PGI attB placement optimizations and subsequent Bxb1- mediated integration in PHH. Fig. 84A shows results from a cytotoxicity assay (Cell Titer Glo) assessing PHH viability after transfection with varying doses of PE-TB1 components versus pL- PGI components. Fig.84B shows the dose-dependent efficiency and fidelity (% Edit by NGS) for 38 bp Bxb1 attB placement in F9 intron 1 from the same experiment as Fig. 84A. Fig. 84C measures IFN-beta concentration in PHH supernatant 6 hours post-transfection with individual L- PGI components, showing no significant innate immune response compared to controls. Fig.84D shows efficiency and fidelity results for an attB overlap walk in PAH, varying the overlap length (10, 20, 38 nt). Fig.84E shows efficiency and fidelity results for a similar symmetric overlap walk (20, 38, 52 nt) for placing a 52 bp Bxb1 attP site. Fig. 84F (SEQ ID Nos: 1254-1257) displays results from a donor chemical modification screen for pL-PGI attB placement, assessing perfect edit vs indels (% Edit by NGS) and fidelity (%) for donors with varying combinations of 5' phosphorylation ( / Phos / ), 3' C3 spacer ( / C3 / ), methylated cytosine (C̲ or / iMe-dC / ), and DM2\21326271.142PATENT Docket No. J4040-99018 phosphorothioate bonds (*).Fig.84G shows the corresponding attB conversion rates (%) observed using the donors from Fig.84F in a two-step integration experiment involving subsequent AAV cargo and Bxb1 mRNA delivery. Fig.84H shows a schematic map of the 30 kb helper-dependent adenovirus (HDAd) cargo used for integration, containing attP, F9 CDS, GFP reporter, and stuffer sequence. Fig. 84I provides schematic diagrams illustrating the ddPCR primer and probe placements used for specific detection of residual attB, and integration across the attL or attR junctions.
[0142] Fig.85 (SEQ ID Nos:1258-1315) illustrates comparisons of pL-PGI with PE across species in vitro and provides additional data for in vivo translation. Fig. 85A presents representative NGS alignment results for dual PE-TB2 mediated attB placement in PAH intron 1 in PHH, PCH, and PMH. Fig. 85B presents corresponding NGS alignment results for pL-PGI mediated attB placement in the same loci and cell types, highlighting differences in edit outcomes and error profiles compared to PE-TB2. Fig. 85C compares attB placement efficiencies (% Edit by ddPCR) in F9 in PHH using pL-PGI with unpurified versus HPLC-purified lmgRNAs, compared to dual PE-TB1 with purified atgRNAs. Fig.85D shows attB placement efficiency and fidelity in PAH in PHH achieved using pL-PGI with HPLC-purified donors. Fig. 85E shows in vitro potency curves (% Edit by ddPCR vs RNA dose) comparing AIO LNPs delivering either PE- TB2 or pL-PGI for attB placement in F9 intron 1 of PMH. DETAILED DESCRIPTION OF THE INVENTION Introduction
[0143] Advances in genome editing tools have enabled precision editing of genomes for therapeutic, agricultural, industrial, and research purposes. Some editing tools may include integration of nucleic acid sequences into the genome at a targeted location. However, the length of integrating sequences that can be used may be constrained in some editing tools.
[0144] Some nuclease-based tools such as CRISPR-Cas9 use a guide RNA to target the Cas9 protein to a specific DNA sequence specified by the spacer sequence in the guide RNA. Cas9 nuclease activity then cleaves the DNA resulting in a double-stranded break (DSB). DSBs are typically repaired through endogenous DNA repair mechanisms including non-homologous end joining (NHEJ) or homology-directed repair (HDR). However, NHEJ results in a spectrum of nucleotide insertions and deletions (indels) that hinder its utility for precision editing. HDR DM2\21326271.143PATENT Docket No. J4040-99018 efficiency is very low in nondividing cells and may require DNA replication. Even when HDR editing is detectable, DSB-induced indels are often prevalent, meaning that HDR may not be feasible when precision editing is desired.
[0145] Homology-independent targeted insertion (HITI) utilizes NHEJ DNA repair mechanisms active in nondividing cells for CRISPR-guided transgene integration in nondividing cells such as primary neurons, retinal pigment epithelial cells, and HSPCs. However, due to the generation of DSBs from Cas9, HITI generates high frequencies of indels, resulting in unintended mutations in addition to DSB associated toxicity.
[0146] Other methods for gene editing have additional limitations. Tools employing fusions of nicking Cas nucleases with nucleotide deaminases (e.g., base editors) can perform certain nucleotide mutations, e.g., cytosine base editors can convert C to T. While some base editors can perform precision editing at high efficiency, they are inherently limited to specific edits determined by the deaminase variant so they are only applicable to specific substitution mutations and further cannot perform precise insertion or deletion edits. Moreover, base editors are generally limited to a small editing window within a subset of the protospacer region and are therefore significantly limited by protospacer adjacent motif (PAM) availability. Finally, base editors can exhibit bystander mutations within the editing region (e.g., if two C’s are present) and have demonstrated DNA and RNA off-target deaminase activity.
[0147] Existing precision editing technologies have limitations that hamper their practical applicability in a variety of ways. In particular, they may rely on endogenous cellular machinery for editing, for example HDR machinery for nuclease-based editing and mismatch repair for base editing. No system has been reported that is independent of all endogenous factors. Reliance on endogenous factors is problematic because different cell types have different activity levels of these endogenous factors, and in many cases the activity is not sufficient to provide useful levels of editing. An example where this reliance is particularly problematic is nondividing cells, which comprise the majority of cells in adults and therefore are not amenable to many existing precision editing tools.
[0148] Accordingly, there remains a need for a system or a method for effective gene editing (e.g., revising) or for modifying gene expression by gene editing. Particularly, there remains a need for the system or method that minimally constrains the length of the revising nucleotide sequence. Additionally, there remains a need for the system or method for gene editing or modifying gene DM2\21326271.144PATENT Docket No. J4040-99018 expression, where the system or the method do not rely on the endogenous components or mechanism of a cell. There also remains a need for a system or a method for correcting genetic mutations in a cell. In some cases, the correction of genetic mutation can treat a disease or condition in subject in need thereof. As will be seen below, the systems, methods, and compositions disclosed herein may be useful for addressing these needs or limitations. Overview
[0149] Described herein are self-contained gene editing systems. In some such self-contained systems, every aspect of gene editing may be controlled. Some such systems do not rely on host cell machinery to perform an editing function, or to replace or repair any aspect of a target nucleic acid such as a genomic locus. Some such systems are unaffected by a cell’s nucleotide triphosphate (dNTP) concentration because the editing may be performed without use of a polymerase. For example, an exogenous first integrating nucleic acid may be delivered and inserted into a genetic locus without transcribing a template. The editing may exclude a need to rely on a cell repair system such as HDR or NHEJ. The editing may be performed without cell cycling. The gene editing may take place in a cell or may even be performed in vitro. For example, the gene editing may even be performed in a test tube or outside of a cell.
[0150] Described herein are systems and methods for editing DNA with a donor strand without generating a double-stranded break in the genome using CRISPR-guided DNA ligases and guide nucleic acids targeting the genomic region of interest. DNA ligases are enzymes which chemically join two DNA molecules via a phosphodiester bond. DNA ligases may or may not require hybridization of the DNA molecules to a DNA or RNA backbone or “splint” which is reverse complementary to the DNA sequences that are to be ligated. Targeting of ligases to genomic nicks generated by CRISPR nucleases enables precise replacement of genomic DNA with donor strands optionally recruited by guide nucleic acids into targeted loci. The CRISPR-guided DNA ligases can be composed of DNA ligases that are fused, recruited, or unfused to the RNA-guided endonuclease by utilizing peptide linkers, heterodimerization domains, or two separate peptides, respectively.
[0151] In certain embodiments, these systems and methods may be referred to as Ligase- mediated Programmable Genomic Integration (L-PGI). L-PGI utilizes specific configurations of guide RNAs, splint nucleic acids, and donor nucleic acids, often incorporating chemical modifications, to achieve high-fidelity editing. Furthermore, paired L-PGI (pL-PGI) approaches, DM2\21326271.145PATENT Docket No. J4040-99018 employing two L-PGI complexes simultaneously targeting opposite strands, may be used for efficient deletion or replacement of larger genomic segments or for insertion of sequences like site-specific recombination sites.
[0152] Some aspects include a cell containing or comprising an RNA-guided endonuclease and a DNA ligase, both of which are introduced into the cell. The endonuclease or ligase may be heterologous to the cell. The endonuclease and ligase may be heterologous to the cell. The ligase may be endogenous to the cell. In some aspects, a cell comprises an RNA-guided endonuclease and a DNA ligase, both of which are heterologous to the cell. The cell may include a composition or system described herein. The cell may be used or included in a system, composition, or method described herein.
[0153] A system described herein may include a heterologous endonuclease comprising an RNA-guided endonuclease such as nicking Cas9 as well as a heterologous ligase (e.g., a DNA ligase) that can utilize an RNA splint. The guide nucleic acid optionally recruits a donor strand to the site targeted by the endonuclease (e.g., a targeted genomic locus) and also generates a splint across from the donor strand (donor strand) and genomic flap generated by the nicking Cas9, resulting in ligation of the donor strand and the genomic flap by the DNA ligase. In some embodiments, the ligase is or comprises an endogenous ligase. The system can utilize one or more guide nucleic acids that together can comprise the following components, optionally in the following order: 5’ spacer - scaffold - donor binding site (optional) - flap binding site 3’. The donor strand (donor strand) can comprise the following sequence components: 5’ guide binding site - donor strand 3’. The guide binding site of the donor strand is at least partially reverse complementary to the donor binding site of the guide nucleic acid such that the donor hybridizes to the guide and is localized to the target site of the RNA guided endonuclease. The 5’ end of the donor sequence and the 3’ end of the genomic flap generated by nuclease nicking activity are ligated by the DNA ligase, splinted by the donor binding site and a flap binding site of the guide nucleic acid(s).
[0154] Fig. 1A-1C illustrate a non-limiting example of a system (1-sided Replacer 1). The example includes a guide nucleic acid comprising: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus. The guide nucleic acid is shown complexed with an endonuclease (e.g., a Cas9 DM2\21326271.146PATENT Docket No. J4040-99018 nickase, nCas9) operatively coupled to a ligase. The guide nucleic acid may direct the endonuclease to a genomic locus that is bound by the spacer of the guide nucleic acid. The guide nucleic acid is also shown as partially complementary to a donor strand (complexing between the donor binding site of the guide nucleic acid and guide binding site of the donor strand). The endonuclease, when directed by the guide nucleic acid, can cleave or nick at least one strand of the genomic locus, and the ligase can ligate one end of the donor strand with the cleaved or nicked end of the genomic locus, thus incorporating the donor strand into the genomic locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by a nuclease.
[0155] Fig. 2A-2C illustrate a non-limiting example of a system (2-sided Replacer 1). The guide nucleic acid in the example, similar to the guide nucleic acid of Fig 1A, comprises: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus. In Fig. 2A, a first guide nucleic acid is shown complexed with a first endonuclease operatively coupled with a first ligase and a second guide nucleic acid is complexed with a second endonuclease operatively coupled with a second ligase. The first endonuclease and the second nuclease may each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus can then be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. The insertion of the donor strand at the genomic locus may generate two genomic flaps that can be digested and removed by a nuclease.
[0156] Fig.3A-3C illustrate a non-limiting example of a system (1-sided Replacer 2). In the example, a guide nucleic acid comprises: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; and a donor binding site for complexing with a donor strand. Also shown in Fig.3A is a donor strand comprising at least one overhang, where the overhang comprises: a flap binding site for complexing with a genomic flap of the genomic locus; and a guide binding site for complexing with the guide nucleic acid (via the donor binding site of the guide nucleic acid). The guide nucleic acid can be complexed with an endonuclease (e.g., nCas9) operatively coupled to a ligase. The guide nucleic acid in the example directs the endonuclease and the ligase to a genomic locus that is bound by the spacer of the guide nucleic acid. The guide nucleic acid in the example is also partially complementary to a donor DM2\21326271.147PATENT Docket No. J4040-99018 strand (complexing between the donor binding site of the guide nucleic acid and guide binding site of the donor strand). The endonuclease, when directed by the guide nucleic acid, can cleave at least one strand of the genomic locus, and the ligase can ligate one end of the donor strand with the cleaved end of the genomic locus, thus incorporating the donor strand into the genomic locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by a nuclease.
[0157] Fig.4A-4C illustrates a non-limiting example of a system (2-sided Replacer 2). In the example, where the guide nucleic acid, similar to the guide nucleic acid of Fig 3A, comprises a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; and a donor binding site for complexing with a donor strand. Also shown in Fig. 4A is a donor strand comprising two overhangs, where the overhangs each comprise a flap binding site for complexing with a genomic flap of the genomic locus; and a guide binding site for complexing with a guide nucleic acid (via a donor binding site of the guide nucleic acid). The flap binding site of the donor strand can bring the donor strand in close proximity with the genomic locus after a genomic flap is generated after the endonuclease cleaves at least one strand of the genomic locus. In Fig.4A, a first guide nucleic acid is shown complexed with a first endonuclease operatively coupled with a first ligase and a second guide nucleic acid is complexed with a second endonuclease operatively coupled with a second ligase. In the example, the first endonuclease and the second nuclease each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus can then be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. In the example, the insertion of the donor strand at the genomic locus generates two genomic flaps that can be digested and removed by a nuclease.
[0158] A system described herein (Replacer 3) may include a heterologous endonuclease comprising an RNA-guided endonuclease such as nicking Cas9 as well as a ligase (e.g., a DNA ligase) that can utilize a DNA splint. The guide nucleic acid optionally recruits a donor strand to the site targeted by the endonuclease (e.g., a targeted genomic locus) and also generates a splint across from the donor strand (donor strand) and genomic flap generated by the nicking Cas9, resulting in ligation of the donor strand and the genomic flap by the DNA ligase. At least part of the flap binding site and donor binding site on the guide nucleic acid are DNA such that ligases that utilize DNA splints are able to catalyze the intended reaction. The system can utilize one or more guide nucleic acids that together can comprise the following components, optionally in the DM2\21326271.148PATENT Docket No. J4040-99018 following order: 5’ spacer - scaffold - donor binding site (optional) - flap binding site 3’. The donor strand (donor strand) can comprise the following sequence components: 5’ guide binding site - donor strand 3’. The guide binding site of the donor strand is at least partially reverse complementary to the donor binding site of the guide nucleic acid such that the donor hybridizes to the guide and is localized to the target site of the RNA guided endonuclease. The 5’ end of the donor sequence and the 3’ end of the genomic flap generated by nuclease nicking activity are ligated by the DNA ligase, splinted by the donor binding site and a flap binding site of the guide nucleic acid(s).
[0159] Fig. 5A-5C illustrate a non-limiting example of a system (1-sided Replacer 3). The example includes a guide nucleic acid comprising: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus, wherein at least part of the flap binding site and donor binding site are comprised of DNA. The guide nucleic acid is shown complexed with an endonuclease (e.g., a Cas9 nickase, nCas9) operatively coupled to a ligase (e.g., an endogenous ligase or an exogenous ligase). The guide nucleic acid may direct the endonuclease to a genomic locus that is bound by the spacer of the guide nucleic acid. The guide nucleic acid is also shown as partially complementary to a donor strand (complexing between the donor binding site of the guide nucleic acid and guide binding site of the donor strand). The endonuclease, when directed by the guide nucleic acid, can cleave at least one strand of the genomic locus, and the ligase can ligate one end of the donor strand with the cleaved end of the genomic locus, thus incorporating the donor strand into the genomic locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by a nuclease.
[0160] Fig. 6A-6C illustrate a non-limiting example of a system (2-sided Replacer 3). The guide nucleic acid in the example, similar to the guide nucleic acid of Fig.5A, comprises: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting an endonuclease described herein; a donor binding site for complexing with a donor strand; and a flap binding site for complexing with a genomic flap of the genomic locus, wherein at least part of the flap binding site and donor binding site are comprised of DNA. In Fig. 6A, a first guide nucleic acid is shown complexed with a first endonuclease operatively coupled with a first ligase and a second guide nucleic acid is complexed with a second endonuclease operatively coupled with a second ligase. DM2\21326271.149PATENT Docket No. J4040-99018 The first endonuclease and the second nuclease may each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus can then be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. The insertion of the donor strand at the genomic locus may generate two genomic flaps that can be digested and removed by a nuclease.
[0161] Ligation may be performed using a DNA ligase that can utilize an RNA splint such as SplintR ligase – also known as PBCV-1 DNA Ligase – from Chlorella virus. In some aspects, the system utilizes two guide nucleic acids targeting the CRISPR-guided ligase to target sites on opposite strands flanking the genomic region of interest. In some aspects, each guide nucleic acid interacts with a corresponding donor strand in the manner described above, resulting in ligation of both donor strands which are reverse complementary with each other in the donor strand regions.
[0162] A ligase that is fused or recruited to an endonuclease, or supplied in trans, can utilize DNA as a splint, and a donor strand acts as the splint for the genomic flap generated by the endonuclease and another donor strand. In some aspects, the donor strand comprises: 5’ donor strand - flap binding site - guide binding site (optional) 3’. The flap binding site on one donor strand (Donor2) can be reverse complementary to the genomic flap, while the optional guide binding site on Donor2 is reverse complementary to the optional donor binding site of a guide nucleic acid (Guide1), and the donor strand can be at least partially reverse complementary to a different donor strand (Donor1). The 5’ end of this Donor1 and the 3’ end of the genomic flap can be ligated using the flap binding site and donor strand of the Donor2 as a splint. Such 2-sided approach utilizing dual guide nucleic acids with different spacer sequences can be adopted with Donor2, which provides the splint at the first genomic site and can be ligated on its 5’ end to a 3’ end of a different genomic flap at a nick created using a second Replacer2 guide nucleic acid (Guide2) with a spacer sequence that targets a second site. The donor binding site on the second guide nucleic acid system can optionally recruit Donor1 via hybridization with its optional guide binding site, and the Donor1 acts as the DNA splint for ligation of Donor2 to the 3’ end of the genomic flap at the target site of the second guide nucleic acid.
[0163] Following ligation, the remaining flaps of native genomic DNA can be excised via exogenously delivered or endogenous flap endonucleases or exonucleases. Examples of exogenous nucleases that can be introduced into the cell include human flap endonuclease 1 (hFEN1), human exonuclease 5 (hEXO5), T5 exonuclease, T7 exonuclease, exonuclease VIII, the DM2\21326271.150PATENT Docket No. J4040-99018 flap endonuclease domain of E. coli PolI, RecJF, Lambda exonuclease, Xni (ExoIXI) from Escherichia coli, SaFEN (Staphylococcus aureus FEN), nuclease BAL-31, or fragments thereof. The endonucleases or exonucleases can optionally be fused, recruited, or unfused to the RNA- guided endonuclease or DNA ligase by utilizing peptide linkers, heterodimerization domains, or two separate peptides, respectively.
[0164] In some aspects, the system, composition, or method described herein utilizes additional protein that binds to the cleaved or nicked site. For example, the system, composition, or method described herein can include Ku protein or Gam protein from bacteriophage Mu, where the binding of the Ku protein or Gam protein can increase ligation efficiency of the integration nucleic acid at the cleaved or nicked site.
[0165] A system or method described herein may use a nicking endonuclease and, therefore, does not generate double stranded breaks. Furthermore, the system described herein addresses the issue of poor editing efficiencies in nondividing cells through a mechanism of action which only depends on the exogenous components delivered to the cells using mRNA, viral vectors, guide nucleic acids, DNA, or peptides, or any other modalities. Therefore, the system does not require the presence of cell cycle-dependent endogenous cell processes or components such as HDR or dNTPS. As such, the system described herein allows efficiency that is not hindered in nondividing cells. Furthermore, the system enables replacement of both strands of a targeted region of the genome, which can increase editing efficiency.
[0166] A donor strand may contain a high degree of homology with the replaced genomic DNA. These donors may contain mutations to the genomic DNA such as pathogenic mutation correction, disabling of CRISPR protospacer adjacent motif (PAM) sites, disruption of the guide’s spacer sequences, other substitution mutations, or a combination thereof. Additional substitution mutations may be included to increase donor-donor homology versus donor-genome homology to promote hybridization of donor strands and incorporation into the genome. Donor strands may also encode deletions or insertions of nucleotides or may encode a complex combination of the above which then replaces the target genomic DNA. Optionally, guide and donor strands may be chemically modified using nucleic acid chemistries such as phosphorothioate bonds or 2’-O- methylation. Optionally, guide nucleic acids may include hairpin sequences. Optionally, any combination of guide nucleic acids, donor strands, and proteins can be complexed, using an DM2\21326271.151PATENT Docket No. J4040-99018 annealing reaction (gradual reduction in temperature) for example, prior to delivering the editing components to the cell.
[0167] Protein components (e.g., nicking Cas9, ligase) may be modified using nuclear localization signals, cell penetrating peptides, or chromatin disrupting peptides in order to improve delivery efficiency to genomic targets.
[0168] The predominant cellular DNA repair pathway for resolving small (< 13nt) mismatches between genomic DNA strands is mismatch repair (MMR). For single stranded donor ligation, the ligated donor strand forms a DNA heteroduplex with the reverse complementary genomic DNA strand. This may also occur with competitive hybridization between ligated donor strand strands and genomic DNA strands. In these cases, MMR activity can excise and revert mismatches in the donor strand using the genomic strand as a template, resulting in reduced editing. Expression of dominant negative versions of MMR proteins has been shown to inhibit the MMR pathway and improve editing outcome in cases where similar DNA heteroduplexes are generated. In some aspects, dominant negative MMR peptides such as MSH2 (G674A) and MLH1 (del754-756) may be delivered as part of the system described herein to improve genomic editing capability, particularly in cells which overexpress the MMR pathway. In some aspects, these dominant negative MMR peptides can be delivered as a fusion (e.g., fused with any component of the system described herein), recruited, or as separate peptides. Endonucleases
[0169] Disclosed herein are endonucleases. The endonuclease may be included in a composition, system or method disclosed herein. The endonuclease may be recombinant. The endonuclease may be coupled to a ligase. The endonuclease may be coupled directly or indirectly to the ligase. The coupling may be covalent or non-covalent. The endonuclease may be bound or connected to a ligase. The endonuclease may be recruited to, be part of a fusion protein with, or be used in conjunction with the ligase. The endonuclease may be coupled to an integrase. The endonuclease may be coupled directly or indirectly to the integrase. The coupling may be covalent or non-covalent. The endonuclease may be bound or connected to an integrase. The endonuclease may be recruited to, be part of a fusion protein with, or be used in conjunction with the integrase. The endonuclease may be heterologous. Heterologous may indicate a source from without a cell. Where a heterologous endonuclease is described, a non-heterologous (e.g., endogenous) endonuclease may be used in some instances. The endonuclease may be encoded in a cell. The DM2\21326271.152PATENT Docket No. J4040-99018 endonuclease may be delivered to the cell in trans. The endonuclease may catalyze cleavage of a phosphate bond within an exogenous first integrating nucleic acid. The endonuclease may be guided by a guide nucleic acid to cleave or nick a target nucleic acid for ligation of an exogenous first integrating nucleic acid at the cleavage or nick site. The endonuclease may include any aspect included in Fig.1A-6C.
[0170] The endonuclease may be non-naturally occurring. The endonuclease may be engineered. The endonuclease may be synthetic. The endonuclease may be pre-synthetized. The endonuclease may be added to a subject or a cell. The endonuclease may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0171] At least part of the endonuclease may be included in a first polypeptide. At least part of the endonuclease may be included in a second polypeptide. The endonuclease may be split into two or more polypeptides bound together. The first polypeptide may include an N-terminal portion of the endonuclease. The first polypeptide may include a C-terminal portion of the endonuclease. The second polypeptide may include the N-terminal portion of the endonuclease. The second polypeptide may include the C-terminal portion of the endonuclease. The first or second polypeptide comprising a part of the endonuclease may be fused with at least part, or the whole, of the ligase. The first or second polypeptide comprising a part of the endonuclease may be fused with at least part, or the whole, of the integrase.
[0172] Described herein, in some aspects, is a system comprising at least one endonuclease. In some aspects, the endonuclease is a programmable endonuclease, where the endonuclease can be complexed with and directed by a guide nucleic acid described herein to a genomic locus. The endonuclease may bind DNA. In some aspects, the endonuclease is an RNA-guided endonuclease. In some aspects, the endonuclease can introduce a single-stranded break. Examples of RNA- guided endonucleases can include CRISPR / Cas endonucleases (e.g., class 2 CRISPR / Cas endonucleases such as a type II, type V, or type VI CRISPR / Cas endonucleases). A CRISPR / Cas endonuclease is also referred to as a CRISPR / Cas effector polypeptide. A suitable endonuclease is a CRISPR / Cas endonuclease (e.g., a class 2 CRISPR / Cas endonuclease such as a type II, type V, or type VI CRISPR / Cas endonuclease). In some cases, a suitable RNA-guided endonuclease is a class 2 CRISPR / Cas endonuclease. In some cases, a suitable RNA-guided endonuclease is a class 2 type II CRISPR / Cas endonuclease (e.g., a Cas9 protein). In some cases, an endonuclease includes a class 2 type V CRISPR / Cas endonuclease (e.g., a Cpf1 protein, a C2c1 protein, or a C2c3 DM2\21326271.153PATENT Docket No. J4040-99018 protein). In some cases, a suitable RNA-guided endonuclease is a class 2 type VI CRISPR / Cas endonuclease (e.g., a C2c2 protein; also referred to as a “Cas13a” protein). Also suitable for use is a CasX protein. Also suitable for use is a CasY protein. In some aspects, the endonuclease can include any one of the Cas described herein complexed with a guide nucleic acid (e.g., a gRNA) as an RNP complex.
[0173] In some cases, the endonuclease is a Type II CRISPR / Cas endonuclease. In some cases, the endonuclease is a Cas9. Cas9 functions as an RNA-guided endonuclease that uses a dual-guide RNA having a crRNA and trans-activating crRNA (tracrRNA) for target recognition and cleavage by a mechanism involving two nuclease active sites in Cas9 that together generate double-stranded DNA breaks (DSBs) or can individually generate single-stranded DNA breaks (SSBs). The Type II CRISPR endonuclease Cas9 and engineered dual- (dgRNA) or single guide RNA (sgRNA) form a ribonucleoprotein (RNP) complex that can be targeted to a desired DNA sequence. Guided by a dual-RNA complex or a chimeric single-guide RNA, Cas9 generates site-specific DSBs or SSBs within double-stranded DNA (dsDNA) target nucleic acids, which are repaired either by non- homologous end joining (NHEJ) or homology-directed recombination (HDR). The Cas9 can be guided to a target site (e.g., stabilized at a target site) within a target nucleic acid sequence by virtue of its association with the RNA-binding segment of the Cas9 to guide RNA. A Cas9 protein can bind and / or modify (e.g., cleave, nick, methylate, demethylate, etc.) a target nucleic acid and / or a polypeptide associated with target nucleic acid (e.g., methylation or acetylation of a histone tail; e.g., when the Cas9 protein includes a fusion partner with an activity). In some cases, the Cas9 protein is a naturally-occurring protein (e.g., naturally occurs in bacterial and / or archaeal cells). In other cases, the Cas9 protein is not a naturally-occurring polypeptide (e.g., the Cas9 protein is a variant Cas9 protein, a chimeric protein, and the like).
[0174] Naturally occurring Cas9 proteins may bind a Cas9 guide RNA, are thereby directed to a specific sequence within a target nucleic acid (a target site), and cleave the target nucleic acid (e.g., cleave dsDNA to generate a double strand break, cleave ssDNA, cleave ssRNA, etc.). A chimeric Cas9 protein may include a fusion protein comprising a Cas9 polypeptide fused to a heterologous protein (referred to as a fusion partner), where the heterologous protein provides an activity (e.g., one that is not provided by the Cas9 protein). The fusion partner can provide an activity, e.g., enzymatic activity (e.g., nuclease activity, activity for DNA and / or RNA methylation, activity for DNA and / or RNA cleavage, activity for histone acetylation, activity for DM2\21326271.154PATENT Docket No. J4040-99018 histone methylation, activity for RNA modification, activity for RNA-binding, activity for RNA splicing etc.). In some cases, a portion of the Cas9 protein (e.g., the RuvC domain and / or the HNH domain) exhibits reduced nuclease activity relative to the corresponding portion of a wild type Cas9 protein (e.g., in some cases the Cas9 protein is a nickase). In some cases, the Cas9 protein is enzymatically inactive, or has reduced enzymatic activity relative to a wild-type Cas9 protein (e.g., relative to Streptococcus pyogenes Cas9). In some cases, the Cas9 is a Cas9 nickase. The Cas9 nickase can be generated by mutating a Cas9 nuclease domain. Non-limiting example of the Cas9 nickase can include SpCas9, SaCas9, CjCas9, GeoCas9, HpaCas9, and NmeCas9. In some aspects, the endonuclease described herein comprises any one of the Cas9 in Table 1. In some aspects, the endonuclease described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the Cas9 in Table 1. Table 1. Non-limiting examples of Cas9 polypeptide sequence QDM2\21326271.155PATENT Docket No. J4040-99018 VYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLAN GEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVDM2\21326271.156PATENT Docket No. J4040-99018 QEEFYKFIKPILEKMDGTEELLVKLKREDLLRKQRTFDNGSIPHQIH LGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFDM2\21326271.157PATENT Docket No. J4040-99018 QTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLV VAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKDM2\21326271.158PATENT Docket No. J4040-99018 STRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHHDP QTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKDM2\21326271.159PATENT Docket No. J4040-99018 LYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMYHH DPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIDM2\21326271.160PATENT Docket No. J4040-99018 RSGKIQTVVKTKLSEIKLDASGHFPMYGKESDPRTYEAIRQRLLEH NNDPKKAFQEPLYKPKKNGEPGPVIRTVKIIDTKNQVIPLNDGKTVDM2\21326271.161PATENT Docket No. J4040-99018 VVACSTASMQQKITKAFQRHESIEYVDTETGEVKFRIPQPWDFFRQ EVMIRVFSDQPCEDLVEKLSARPEALHDNVTPLFVSRAPNRKMSGp g . he RNA-guided endonuclease may comprise a class II CRISPR / Cas endonuclease. The RNA-guided endonuclease may comprise a Cas9 endonuclease. The RNA-guided endonuclease may comprise a nickase. The RNA-guided endonuclease may comprise an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 1-13, or a functional fragment thereof.
[0176] The endonuclease may introduce a single-strand break in a target nucleic acid. The endonuclease may introduce a single-strand break in a target nucleic acid without cleaving a strand opposite the single strand break. The endonuclease may include a nickase. In some instances, the DM2\21326271.162PATENT Docket No. J4040-99018 endonuclease may exclude an endonuclease that introduces a double strand break. The endonuclease may exclude a restriction enzyme.
[0177] The endonuclease may be included as part of a fusion protein. In some cases, an endonuclease is a fusion protein that is fused to a heterologous polypeptide such as the heterologous ligase described herein. The heterologous polypeptide may include a fusion partner. The fusion protein may include a fusion partner such as a DNA ligase, a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, or tag polypeptide. The fusion protein may include one or more fusion partner. The fusion protein may include a ligase. The fusion protein may include a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, or tag polypeptide.
[0178] The fusion partner may be connected to the N-terminus of the endonuclease. The fusion partner may be connected to the C-terminus of the endonuclease. The endonuclease may be connected at an N-terminus or a C-terminus to a linker. The fusion partner may be connected by the fusion partner’s N-terminus or C-terminus. The fusion partner may be connected by the fusion partner’s N-terminus to the endonuclease. The fusion partner may be connected by the fusion partner’s C-terminus to the endonuclease. The fusion partner may be connected at an N-terminus or a C-terminus to a linker.
[0179] In some cases, the endonuclease comprises a linker, where the linker covalently connects the endonuclease to the heterologous polypeptide. The linker may connect the endonuclease to any fusion partner. A linker may also connect any fusion partner to another fusion partner. The linker polypeptide may have any of a variety of amino acid sequences. Proteins can be joined by a spacer peptide, generally of a flexible nature, although other chemical linkages are not excluded. Suitable linkers include polypeptides of between 4 amino acids and 40 amino acids in length, or between 4 amino acids and 25 amino acids in length. These linkers can be produced by using synthetic, linker-encoding oligonucleotides to couple the proteins, or can be encoded by a nucleic acid sequence encoding the fusion protein. Peptide linkers with a degree of flexibility can be used. The linking peptides may have virtually any amino acid sequence, bearing in mind that the preferred linkers will have a sequence that results in a generally flexible peptide. The use of small amino acids, such as glycine and alanine, are of use in creating a flexible peptide. The creation of such sequences is routine to those of skill in the art. A variety of different linkers are commercially available and are considered suitable for use. Examples of linker polypeptides DM2\21326271.163PATENT Docket No. J4040-99018 include glycine polymers (G)n, glycine-serine polymers (including, for example, (GS)n(SEQ ID NO:960), (GSGGS)n(SEQ ID NO:950), (GGSGGS)n(SEQ ID NO:951), and (GGGS)n(SEQ ID NO:952), where n is an integer of at least one); glycine-alanine polymers; and alanine-serine polymers. Exemplary linkers can comprise amino acid sequences including, but not limited to, GGSG(SEQ ID NO:954), GGSGG(SEQ ID NO:955), GSGSG(SEQ ID NO:956), GSGGG(SEQ ID NO:957), GGGSG(SEQ ID NO:958), GSSSG(SEQ ID NO:959), and the like. Also suitable is a linker having the sequence (GGGGS)n(SEQ ID NO:953), where n is an integer of from 1 to 10 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10). The ordinarily skilled artisan will recognize that design of a peptide conjugated to any desired element can include linkers that are all or partially flexible, such that the linker can include a flexible linker as well as one or more portions that confer less flexible structure.
[0180] One or more linkers may be included in a fusion protein.1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 linkers, or a range of linkers defined by any two of the aforementioned integers, may be included in the fusion protein. A linker may connect to an N-terminal end of at least part of the endonuclease. A linker may connect to an N-terminal end of at least part of a fusion partner. A linker may connect to an N-terminal end of at least part of a fusion ligase. A linker may connect to an N-terminal end of a nuclear localization signal. A linker may connect to an N-terminal end of a chromatin modifying domain. A linker may connect to an N-terminal end of a cell penetrating peptide. A linker may connect to an N-terminal end of a tag polypeptide. A linker may connect to a C-terminal end of at least part of the endonuclease. A linker may connect to a C-terminal end of at least part of a fusion partner. A linker may connect to a C-terminal end of at least part of a fusion ligase. A linker may connect to a C-terminal end of a nuclear localization signal. A linker may connect to a C-terminal end of a chromatin modifying domain. A linker may connect to a C- terminal end of a cell penetrating peptide. A linker may connect to a C-terminal end of a tag polypeptide.
[0181] A linker may comprise a number or range of amino acids or residues. The linker may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 amino acid residues. The linker may, in some aspects, include no more than 1, no more than 2, no more than 3, no more than 4, no DM2\21326271.164PATENT Docket No. J4040-99018 more than 5, no more than 6, no more than 7, no more than 8, no more than 9, no more than 10, no more than 12, no more than 13, no more than 14, no more than 15, no more than 20, no more than 25, no more than 30, no more than 35, no more than 40, no more than 45, no more than 50, no more than 55, no more than 60, no more than 65, no more than 70, no more than 75, no more than 80, no more than 85, no more than 90, no more than 95, or no more than 100 amino acid residues. A linker may include 1-10 amino acids, 1-25 amino acids, or 1-100 amino acids.
[0182] Linkers may be included anywhere in a polypeptide chain or protein described herein. For example, a linker may separate an endonuclease from a ligase. A linker may separate an endonuclease from a nuclear localization signal, a chromatin modifying domain, a cell penetrating peptide, or a tag polypeptide.
[0183] In some cases, the endonuclease comprises a nuclear localization sequence (e.g., one or more nuclear localization signals or NLSs for targeting to the nucleus). In some aspects, the NLS described herein comprises any one of the NLS in Table 2. In some aspects, the NLS described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of NLS in Table 2. Table 2. Non-limiting examples of NLS polypeptide sequence :DM2\21326271.165PATENT Docket No. J4040-99018
[0184] A polynucleotide encoding an NLS polypeptide may be used. An example of such a polynucleotide may be SGGSx2-bpNLS-SGGSx2: TCCGGCGGAAGCTCTGGTGGCAGCAAGCGGACCGCCGACGGCTCTGAATTCGAGAG CCCTAAGAAGAAAAGAAAGGTGAGCGGAGGCTCTAGCGGCGGAAGC (SEQ ID NO:25).
[0185] In some aspects, the endonuclease comprises a dimerization domain. The dimerization domain can be located at the N-terminus or C-terminus of the endonuclease. In some aspects, the dimerization domain allows the endonuclease to form a heterodimer with another polypeptide (e.g., the heterologous ligase). In some aspects, the dimerization domain allows the endonuclease to be functionally coupled with another polypeptide. Non-limiting examples of the dimerization domains can include a leucine zipper, an FKBP, an FRB, a Calcineurin A, a CyP-Fas, a GyrB, a GAI, a GID1, a SNAP tag, a Halo tag, a Bcl-xL, a Fab, a LOV domain, or SpyTag / SpyCatcher. Other example of dimerization domain can include an antibody such as anyone of heavy chain domain 2 (CH2) of IgM (MHD2) or IgE (EHD2), immunoglobulin Fc region, heavy chain domain 3 (CH3) of IgG or IgA, heavy chain domain 4 (CH4) of IgM or IgE, Fab, Fab2, leucine zipper motifs, barnase-barstar dimers, miniantibodies, or ZIP miniantibodies. In some aspects, the dimerization domain described herein comprises any one of the dimerization domain in Table 3. In some aspects, the dimerization domain described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of dimerization domain in Table 3. Table 3. Non-limiting examples of dimerization domain sequence O:
[00186] In some aspects, the endonuclease comprises at least one additional domain. In some aspects, the at least one additional domain is a functional domain. For example, the functional domain can comprise a chromatin modifying domain or a cell penetrating peptide. In some aspects, the chromatin modifying domain described herein comprises any one of the chromatin modifying DM2\21326271.166PATENT Docket No. J4040-99018 domain in Table 4. In some aspects, the chromatin modifying domain described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of chromatin modifying domain in Table 4. Table 4. Non-limiting examples of chromatin modifying domain polypeptide sequence SE ID
[0087] n some aspects, t e ce penetrat ng pept de descrbed ere n compr ses any one o the cell penetrating peptide in Table 5. In some aspects, the cell penetrating peptide described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of cell penetrating peptide in Table 5. DM2\21326271.167PATENT Docket No. J4040-99018 Table 5. Non-limiting examples of cell penetrating peptide polypeptide sequence N C ll t ti tid SE ID NO:for increasing expression, identifying, or purifying the endonuclease. In some aspects, the tag described herein comprises any one of the tag sequence in Table 6. In some aspects, the tag described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the tag sequence in Table 6. Table 6. Non-limiting examples of tag polypeptide sequence DDM2\21326271.168PATENT Docket No. J4040-99018 NKGETAMTINGPWAWSNIDTSKVNYGVTVLPTFKGQPSKP FVGVLSAGINAASPNKELAKEFLENYLLTDEGLEAVNKDKor more exteins fused to one or more inteins. Intein technology may be used to deliver large proteins into a cell by expressing the protein as two or more shorter peptide segments (exteins). Each extein may be expressed as a fusion with an intein peptide (e.g., an Npu C intein or an Npu N intein). An intein may autocatalyze fusion of two or more exteins and may autocatalyze excision of the intein from its corresponding extein. The result may be a protein complex comprising a first extein fused to a second extein and lacking inteins. An intein may be positioned N-terminal of the extein, or an intein may be positioned C-terminal of the extein. An extein may comprise a cysteine residue positioned adjacent to the intein (e.g., at the C-terminal end of an extein with an intein fused to the C-terminal end of the extein). The Cas nickase may be expressed as two or more segments. A first of the Cas nickase segment may comprise an N-terminal portion of the Cas nickase. A first segment of the Cas nickase may comprise a first intein. A second segment of the Cas nickase may comprise a C-terminal portion of the Cas nickase. A second segment of the Cas nickase may comprise a second intein. An intein may be fused to a C-terminus of an N-terminal portion of the Cas nickase. An intein may be fused to an N-terminus of a C-terminal portion of the Cas nickase. A nucleic acid sequence encoding an extein-intein fusion may fit into a delivery vector (e.g., an adeno- associated virus (AAV) vector). DM2\21326271.169PATENT Docket No. J4040-99018 DNA Ligases
[0190] Disclosed herein are ligases. The ligase may be or include a DNA ligase. The ligase may be included in a composition, system or method disclosed herein. The ligase may be recombinant. The ligase may be coupled to the endonuclease. The ligase may be coupled directly or indirectly to the endonuclease. The coupling may be covalent or non-covalent. The ligase may be bound or connected to the endonuclease. The ligase may be recruited to, be part of a fusion protein with, or be used in conjunction with an endonuclease. The ligase may be coupled to an integrase. The ligase may be coupled directly or indirectly to the integrase. The coupling may be covalent or non-covalent. The ligase may be bound or connected to the integrase. The ligase may be recruited to, be part of a fusion protein with, or be used in conjunction with an integrase. The ligase may be heterologous. The ligase may be endogenous. Where a heterologous ligase is described, a non-heterologous (e.g., endogenous) ligase may be used in some cases. The ligase may be encoded in a cell. The ligase may be delivered to the cell in trans. The ligase may form a phosphodiester bond by joining two nucleic acid ends together. The ligase may join an end (e.g., 5’ or 3’ end) of a target nucleic acid to an exogenous first integrating nucleic acid (e.g., a 3’ or 5’ end of the exogenous first integrating nucleic acid). The ligase ligates an exogenous first integrating nucleic acid (e.g., a donor nucleic acid) to a cleaved or nicked end of a target nucleic acid where the cleaved or nicked end has been generated by an endonuclease such as an RNA- guided endonuclease. The ligase may include any aspect included in Fig.1A-6C.
[0191] The ligase may be non-naturally occurring. The ligase may be engineered. The ligase may be synthetic. The ligase may be pre-synthetized. The ligase may be added to a subject or a cell. The ligase may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0192] At least part of the ligase may be included in a first polypeptide. At least part of the ligase may be included in a second polypeptide. The ligase may be split into two polypeptides bound together. The first polypeptide may include an N-terminal portion of the ligase. The first polypeptide may include a C-terminal portion of the ligase. The second polypeptide may include the N-terminal portion of the ligase. The second polypeptide may include the C-terminal portion of the ligase. The first or second polypeptide comprising a part of the ligase may be fused with at least part, or the whole, of the endonuclease. The first or second polypeptide comprising a part of the ligase may be fused with at least part, or the whole, of the integrase. DM2\21326271.170PATENT Docket No. J4040-99018
[0193] Examples of DNA ligases are hLIG1, T4 ligase, T7 ligase, and ligases from Aquifex aeolicus VF5, Neisseria meningitidis serogroup A strain Z2491, Neisseria meningitidis serogroup B strain MC58, Pseudomonas aeruginosa PA01, Vibrio cholerae El Tor N1696, Vaccinia virus, and Emiliania huxleyi virus.
[0194] The ligase may comprise a ligase that can ligate a substrate comprising DNA. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising a DNA splint. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized to another DNA strand. The splinting DNA strand may include an RNA portion. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized across from a DNA portion of an RNA / DNA hybrid strand. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising a DNA / RNA. In some aspects, the ligase comprises a ligase that can ligate a substrate comprising an RNA splint. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized to an RNA strand. The RNA strand may include a DNA portion. For example, a DNA ligase may ligate a 5’ phosphate to a 3’ hydroxyl of two DNA strands that are hybridized across from an RNA portion of an RNA / DNA hybrid strand.
[0195] In some aspects, the ligase described herein comprises any one of the ligases in Table 7. In some aspects, the ligase described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the ligase in Table 7. Table 7. Non-limiting examples of ligase polypeptide sequence DDM2\21326271.171PATENT Docket No. J4040-99018 KESLTEAEVATEKEGEDGDQPTTPPKPLKTSKAETPTESVS EPEVATKQELQEEEEQTKPPRRAPKTLSSFFTPRKPAVKKEDM2\21326271.172PATENT Docket No. J4040-99018 AGKGKTAEARKTWLEEQGMILKQTFCEVPDLDRIIPVLLE HGLERLPEHCKLSPGIPLKPMLAHPTRGISEVLKRFEEAAFDM2\21326271.173PATENT Docket No. J4040-99018 QSKSFPPAAKSLLTIQEVDEFLLRLSKLTKEDEQQQALQDI ASRCTANDLKCIIRLIKHDLKMNSGAKHVLDALDPNAYEADM2\21326271.174PATENT Docket No. J4040-99018 AKSLLTIQEVDEFLLRLSKLTKEDEQQQALQDIASRCTAND LKCIIRLIKHDLKMNSGAKHVLDALDPNAYEAFKASRNLQDM2\21326271.175PATENT Docket No. J4040-99018 ENQVVNNLDEAKVIYKKYIDQGLEGIILKNIDGLWENARS KNLYKFKEVIDVDLKIVGIYPHRKDPTKAGGFILESECGKIDM2\21326271.176PATENT Docket No. J4040-99018 NAD- MESIEQQLTELRTTLRHHEYLYHVMDAPEIPDAEYDRLMR 68 dependent E ELRELETKHPELITPDSPTQRVGAAPLAAFSQIRHEVPMLSDM2\21326271.177PATENT Docket No. J4040-99018 Vaccinia virus MTSLREFRKLCCDIYHASGYKEKSKLIRDFITDRDDKYLII 71 DNA ligase KLLLPGLDDRIYNMNDKQIIKLYSIIFKQSQEDMLQDLGYGDM2\21326271.178PATENT Docket No. J4040-99018 Alteromonas MQFFLTVFCLLLITAVTHVNAEDKLDIVDGLQLAKQYSHS 74 mediterranea RQDINIAEYWVSEKLDGIRARWDGTELRTRNNNKIAAPADM2\21326271.179PATENT Docket No. J4040-99018 IGDVRLISCKTTTECKALIDRGYDILHPNWVLDCIAYKRLIL IEPNYCFNVSQKMRAVAEKRVDCLGDSFENDISETKLSSLDM2\21326271.180PATENT Docket No. J4040-99018 Mouse DNA MASSQTSQTVAAHVPFADLCSTLERIQKGKDRAEKIRHFK 79 ligase IV EFLDSWRKFHDALHKNRKDVTDSFYPAMRLILPQLERERDM2\21326271.181PATENT Docket No. J4040-99018 Arabidopsis MTEEIKFSVLVSLFNWIQKSKTSSQKRSKFRKFLDTYCKPS 81 DNA ligase IV DYFVAVRLIIPSLDRERGSYGLKESVLATCLIDALGISRDAPDM2\21326271.182PATENT Docket No. J4040-99018 KFFNKVLDGGSNSVSVGSETEECNTDKKMVHIDASEAYK EVTDQFIDIVNGSESLRDYAASIIDEAKGDISRALNIYYSKPDM2\21326271.183PATENT Docket No. J4040-99018 KDIRIGDAVIIKKAGDIIPEVVGVVVDRRDGDETPFAMPTH CPECESELVRLEGEVALRCLNPNCPAQLRERLIHFASRAADM2\21326271.184PATENT Docket No. J4040-99018 LYRLKLEDLLKLEGFAETRARNLLRAIEASKQRPLSRLLFG LGIRHVGKTTAELLVQRFASIDELAAATIDELAALEGVGPIDM2\21326271.185PATENT Docket No. J4040-99018 Thermus MTLEEARRRVNELRDLIRYHNYLYYVLDAPEISDAEYDRL 90 species AK16D LRELKELEERFPELKSPDSPTEQVGARPLEATFRPVRHPTRDM2\21326271.186PATENT Docket No. J4040-99018 KSHIESFFADKLIETPADIFRLFQKRQLLIEREGWGELSVDN LISAIDKRRKVPFDRFLFALGIRHVGAVTARDLAKSYQTWDM2\21326271.187PATENT Docket No. J4040-99018 SSKYDPGSRDKSWIKLKPDFVDGMGDTLDLLILGGYYGE GRRRSGAVSTFLMGVRAPPEAAKRVGGAAHPLFYPFCKVNA splint. In some embodiments, the DNA ligase ligates DNA strands base paired to an RNA splint. In some embodiments, the DNA ligase comprises an amino acid sequence at least 80% identical to the amino acid sequence of any one of SEQ ID NOS: 55-96, or a functional fragment thereof.
[0197] In some aspects, the ligases comprises at least one NLS (e.g., any one of the NLS in Table 2). In some aspects, the ligase comprises at least one additional domain. In some aspects, the at least one additional domain is a dimerization domain (e.g., any one of the dimerization domain in Table 3). In some aspects, the ligase comprising a dimerization domain can be dimerized with an endonuclease to form a heterodimer. In some aspects, the at least one additional domain is a functional domain. For example, the functional domain can comprise a chromatin modifying domain (e.g., any one of the chromatin modifying domain in Table 4) or a cell DM2\21326271.188PATENT Docket No. J4040-99018 penetrating peptide (e.g., any one of the cell penetrating peptide in Table 5). In some aspects, the ligase comprises a linker, where the linker can covalently connect the ligase with another polypeptide (e.g., the endonuclease). In some aspects, the linker covalently connects the ligase to the at least one additional domain. In some aspects, the ligase comprises a tag (e.g., any one of the tag in Table 6, where the tag can be used for increasing expression, identifying, or purifying the ligase. A linker may separate the ligase from a nuclear localization signal, a chromatin modifying domain, a cell penetrating peptide, or a tag polypeptide. Any linker described herein may be included.
[0198] The ligase may comprise a binding motif for binding to a nucleic acid motif (e.g., a hairpin motif). In some aspects, the ligase (e.g., DNA ligase) comprises an MS2 coat protein (MCP) peptide. The ligase may include a hairpin binding motif such as an MCP peptide. The MCP peptide may be useful for recruiting the ligase to a guide nucleic acid comprising an MS2 hairpin. A benefit of using a MCP peptide and MS2 hairpin is to separate the ligase and endonuclease such as a Cas nickase (or a portion of them) and allow fitting within separate vectors such as AAV vectors. In some aspects, the ligase comprises a loop region. In some aspects, the loop region is a 2a loop or a 3a loop. The loop region may comprise a 2a loop. The loop region may comprise a 3a loop. Integrases
[0199] Disclosed herein are integrases. An integrase may be an example of a recombinase, and where an integrase is described, a recombinase may be contemplated. The integrase may be or include a phage integrase. The integrase may be or include a site-specific recombinase. The integrase may be or include a serine integrase. The integrase may be or include a resolvase. The integrase may be or include a DNA invertase. The integrase may be or include a tyrosine integrase. The integrase may be or include a retrotransposase. The integrase may be included in a composition, system or method disclosed herein. The integrase may be recombinant. Any of these integrases may be modified or mutated. For example, an integrase may include an insertion, a deletion, or may be an active or functional fragment. The integrase may include a mutation of an integrase described herein. The integrase may be coupled to the endonuclease. The integrase may be coupled directly or indirectly to the endonuclease. The coupling may be covalent or non- covalent. The integrase may be bound or connected to the endonuclease. The integrase may be recruited to, be part of a fusion protein with, or be used in conjunction with an endonuclease. The DM2\21326271.189PATENT Docket No. J4040-99018 integrase may be coupled to a ligase. The integrase may be coupled directly or indirectly to the ligase. The coupling may be covalent or non-covalent. The integrase may be bound or connected to the ligase. The integrase may be recruited to, be part of a fusion protein with, or be used in conjunction with the ligase. The integrase may be heterologous. The integrase may be endogenous. Where a heterologous integrase is described, a non-heterologous (e.g., endogenous) integrase may be used in some cases. The integrase may be encoded in a cell. The integrase may be delivered to the cell in trans. The integrase introduces a second integrating nucleic acid into a target nucleic acid by recognizing and binding to an exogenous first integrating nucleic acid (e.g., a recombination sequence).
[0200] The integrase may be non-naturally occurring. The integrase may be engineered. The integrase may be synthetic. The integrase may be pre-synthetized. The integrase may be added to a subject or a cell. The integrase may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0201] At least part of the integrase may be included in a first polypeptide. At least part of the integrase may be included in a second polypeptide. The integrase may be split into two polypeptides bound together. The first polypeptide may include an N-terminal portion of the integrase. The first polypeptide may include a C-terminal portion of the integrase. The second polypeptide may include the N-terminal portion of the integrase. The second polypeptide may include the C-terminal portion of the integrase. The first or second polypeptide comprising a part of the integrase may be fused with at least part, or the whole, of the endonuclease. The first or second polypeptide comprising a part of the integrase may be fused with at least part, or the whole, of the ligase.
[0202] In some aspects, the integrase described herein may be a serine integrase. A serine integrase may be referred to as a “serine recombinase.” Examples of species that contain serine recombinases are: Bacillus cereus; Bacillus safensis; Bacillus tropicus; Burkholderia multivorans; Burkholderia ubonensis; Cellulosimicrobium cellulans; Clostridium botulinum; Clostridioides difficile; Clostridium perfringens; Clostridium thermobutyricum; Cronobacter sakazakii; Desulfotomaculum nigrificans; Enterocloster clostridioformis; Enterococcus faecalis; Enterococcus faecium; Escherichia coli; Eubacterium maltosivorans; Faecalibacterium prausnitzii; Fusobacterium mortiferum; Klebsiella pneumoniae; Mycobacteroides abscessus; Mycobacterium phage Bxb1; Mycolicibacterium elephantis; Neobacillus mesonae; Nocardia DM2\21326271.190PATENT Docket No. J4040-99018 otitidiscaviarum; Paenibacillus campinasensis; Paeniclostridium sordellii; Parageobacillus caldoxylosilyticus; Pseudomonas aeruginosa; Pseudomonas fluorescens; Pseudomonas fulva; Pseudomonas putida; Pseudomonas syringae; Prochlorothrix hollandica; Rhizobiales bacterium; Rhodococcus hoagie; Ruminococcus lactaris; Salinispora pacifica; Staphylococcus arlettae; Staphylococcus aureus; Staphylococcus hominis; Streptococcus agalactiae; Streptococcus equinus; Streptococcus mitis; Streptomyces ipomoeae; Streptomyces phage PhiC31; Tenacibaculum dicentrarchi; Treponema denticola; Vibrio harveyi; Vibrio hyugaensis; Vibrio parahaemolyticus.
[0203] In some aspects, the integrase described herein may be from or may be derived from the species Bacillus cereus. In some aspects, the integrase described herein may be from or may be derived from the species Bacillus safensis. In some aspects, the integrase described herein may be from or may be derived from the species Bacillus tropicus. In some aspects, the integrase described herein may be from or may be derived from the species Burkholderia multivorans. In some aspects, the integrase described herein may be from or may be derived from the species Burkholderia ubonensis. In some aspects, the integrase described herein may be from or may be derived from the species Cellulosimicrobium cellulans. In some aspects, the integrase described herein may be from or may be derived from the species Clostridium botulinum. In some aspects, the integrase described herein may be from or may be derived from the species Clostridioides difficile. In some aspects, the integrase described herein may be from or may be derived from the species Clostridium perfringens. In some aspects, the integrase described herein may be from or may be derived from the species Clostridium thermobutyricum. In some aspects, the integrase described herein may be from or may be derived from the species Cronobacter sakazakii. In some aspects, the integrase described herein may be from or may be derived from the species Desulfotomaculum nigrificans. In some aspects, the integrase described herein may be from or may be derived from the species Enterocloster clostridioformis. In some aspects, the integrase described herein may be from or may be derived from the species Enterococcus faecalis. In some aspects, the integrase described herein may be from or may be derived from the species Enterococcus faecium. In some aspects, the integrase described herein may be from or may be derived from the species Escherichia coli. In some aspects, the integrase described herein may be from or may be derived from the species Eubacterium maltosivorans. In some aspects, the integrase described herein may be from or may be derived from the species Faecalibacterium DM2\21326271.191PATENT Docket No. J4040-99018 prausnitzii. In some aspects, the integrase described herein may be from or may be derived from the species Fusobacterium mortiferum. In some aspects, the integrase described herein may be from or may be derived from the species Klebsiella pneumoniae. In some aspects, the integrase described herein may be from or may be derived from the species Mycobacteroides abscessus. In some aspects, the integrase described herein may be from or may be derived from the species Mycobacterium phage Bxb1. In some aspects, the integrase described herein may be from or may be derived from the species Mycolicibacterium elephantis. In some aspects, the integrase described herein may be from or may be derived from the species Neobacillus mesonae. In some aspects, the integrase described herein may be from or may be derived from the species Nocardia otitidiscaviarum. In some aspects, the integrase described herein may be from or may be derived from the species Paenibacillus campinasensis. In some aspects, the integrase described herein may be from or may be derived from the species Paeniclostridium sordellii. In some aspects, the integrase described herein may be from or may be derived from the species Parageobacillus caldoxylosilyticus. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas aeruginosa. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas fluorescens. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas fulva. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas putida. In some aspects, the integrase described herein may be from or may be derived from the species Pseudomonas syringae. In some aspects, the integrase described herein may be from or may be derived from the species Prochlorothrix hollandica. In some aspects, the integrase described herein may be from or may be derived from the species Rhizobiales bacterium. In some aspects, the integrase described herein may be from or may be derived from the species Rhodococcus hoagie. In some aspects, the integrase described herein may be from or may be derived from the species Ruminococcus lactaris. In some aspects, the integrase described herein may be from or may be derived from the species Salinispora pacifica. In some aspects, the integrase described herein may be from or may be derived from the species Staphylococcus arlettae. In some aspects, the integrase described herein may be from or may be derived from the species Staphylococcus aureus. In some aspects, the integrase described herein may be from or may be derived from the species Staphylococcus hominis. In some aspects, the integrase described herein may be from or may be derived from the species Streptococcus agalactiae. In some aspects, DM2\21326271.192PATENT Docket No. J4040-99018 the integrase described herein may be from or may be derived from the species Streptococcus equinus. In some aspects, the integrase described herein may be from or may be derived from the species Streptococcus mitis. Streptomyces ipomoeae. In some aspects, the integrase described herein may be from or may be derived from the species Streptomyces phage PhiC31. In some aspects, the integrase described herein may be from or may be derived from the species Tenacibaculum dicentrarchi. In some aspects, the integrase described herein may be from or may be derived from the species. In some aspects, the integrase described herein may be from or may be derived from the species Treponema denticola. In some aspects, the integrase described herein may be from or may be derived from the species Vibrio harveyi. In some aspects, the integrase described herein may be from or may be derived from the species Vibrio hyugaensis. In some aspects, the integrase described herein may be from or may be derived from the species Vibrio parahaemolyticus.
[0204] In some aspects, the integrase described herein comprises any one of the integrases in Table 8. In some aspects, the integrase described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8. Table 8. Non-limiting examples of large serine recombinase (LSR) sequences :DM2\21326271.193PATENT Docket No. J4040-99018 IPIVSELLALGVTIVSTQEGVFRQGNVMDLIHLIMR LDASHKESSLKSAKILDTKNLQRELGGYVGGKAPDM2\21326271.194PATENT Docket No. J4040-99018 Bm99 Burkholderia MAKKPKAKVYSYLRFSDPKQAAGSSADRQMEYA 255 multivorans ARWATEHDMQLDATLTLRDEGLSAFHQRHIKQGDM2\21326271.195PATENT Docket No. J4040-99018 Nm60 Neobacillus MSRPTGLTIDIYLRKSRKDLEEEKKASESGETYDT 258 mesonae LERHRRTLFAVAKKERHNIANIYEEVVSGESVSERDM2\21326271.196PATENT Docket No. J4040-99018 Cs56 Cronobacter MEKRKLYSYIRWSSDKQAKGSSLQRQLETARRV 261 sakazakii AHENDLELVEIIDAGLSAFRSKHLEKGSLGAFIEADM2\21326271.197PATENT Docket No. J4040-99018 AERQAQLSSEAVVPAGWKLLPTGELFGDWWSRQ DLTARNVWLRSMGVRARFKRDDKTLYIDLGNLNDM2\21326271.198PATENT Docket No. J4040-99018 MREVGDKNVLERYFVPAENHQIELDEAIRATEEL TALLGTMTSATMRSSLTAQLAALDSRIASLEKLPTDM2\21326271.199PATENT Docket No. J4040-99018 NHLVLAPHEGIISSDLWLKCRVKCLEAQQIKPYQK AKNTWLAGKIKCGACGYALVDKHYSTTRSRYLLDM2\21326271.1100PATENT Docket No. J4040-99018 DHKQNKPGYKSEHNAFAGLLKHECGGALVRKFH VASGKTYQYHVCANARDGKCNVTKNFKNIEVALDM2\21326271.1101PATENT Docket No. J4040-99018 GRTCEQPHIRVDELEQAVMEQVKRLPLKHKVKK RAFDFKPVENKIATIDKQKERLLDLYLNEHLDNEDM2\21326271.1102PATENT Docket No. J4040-99018 VKATYQGVYRCNNVPDGRCNVPTIKRKPFDKWM LDNIVGFLERDDGNNTDKRKAEIEYQISLVTSKLKDM2\21326271.1103PATENT Docket No. J4040-99018 WLTTDGQLVPHVAPHIAQAFQDYADGLGERRICR KLRESGLEEFSKTNATTVRRWLKNRTAIGYWNDIDM2\21326271.1104PATENT Docket No. J4040-99018 RFIVDKLLSGKSANEVVRLLESKKKPPGITKWNR KTVLGWMRNPILRGHTKHGDLLIKNTHEPIISEDEDM2\21326271.1105PATENT Docket No. J4040-99018 MTVNQSEAIIVKEVFSSYLNGRSITKLRDDLNEKY PKTPAWSYRTIRQMLDNPVYCGYNKYKGQVYPGDM2\21326271.1106PATENT Docket No. J4040-99018 RKMYKLKVDEDTIEIVKLIFDKYLELRSLSKLYKY MYENGIKGTRGGNLDPSALSLILKNPAYVKADKSDM2\21326271.1107PATENT Docket No. J4040-99018 SMTEPNLEDDEMSLYIDAMQGATNEIYVRKLSKS VKRGHNDRALRGDLPGDVQFGLKRLKDGSIVLDDM2\21326271.1108PATENT Docket No. J4040-99018 ERMQLGKLGRAKSGKSMMWAKTSYGYNYHKET GTVTINPAQALAVKFIFKSYLAGRSITKLRDDLNEDM2\21326271.1109PATENT Docket No. J4040-99018 ERENLGERVKMGQNEKARQGQFSAPAPFGFIKEG KSLVKNHEQGEILLEIIDKVKKGYSTRQIANYLDDDM2\21326271.1110PATENT Docket No. J4040-99018 NVKMSMNAKARSGEAITGRVLGYKLSLNPLTQK NDLVIDENEANIVREIFGLYLNHNKGLKAITTILNQDM2\21326271.1111PATENT Docket No. J4040-99018 LNFTRYGTFAHIFEGKPESKRYEWTIAHVKAILKS EVYIGNSVHNRQSTVSFKSKKKVRKPESEWFRVEDM2\21326271.1112PATENT Docket No. J4040-99018 HNHLIVDDYAADIVREIYKLYLQGIGKGRIGRILS DRGVLIPSLYKRNVQGINYHNANAKAETHLWSYDM2\21326271.1113PATENT Docket No. J4040-99018 APKRMTPVWLNSDGTIREDVAPWIKTAFELYVSG VGKSTIAKRLRESGVERLAKASGPGVEGWLRNKADM2\21326271.1114PATENT Docket No. J4040-99018 QLWDITDQDVHVVTLVDGRIYTKDMDFEDIMLA GLIMQRAHEESETKSKRLQEKWQERRTLGKFIHKDM2\21326271.1115PATENT Docket No. J4040-99018 VMIRANEESETKQRRSNAFLKSALNQYQANGKIR RLGSDPSWLDFNKDNDTYSFNERVEVIRRILNLYNDM2\21326271.1116PATENT Docket No. J4040-99018 IAINDGVDSAKGDNDFTPFCNLFNDFYAKDTSKK VRAIKRAQGQAGEHLTKPPYGYMVSPTDKKQWIDM2\21326271.1117PATENT Docket No. J4040-99018 FQKLLSDVRANLIDLIIFTRLDRWFRSLRHYLNTQ EVLDKHNVSWTAIRQPFFDTSTAQGRTFVNTSMADM2\21326271.1118PATENT Docket No. J4040-99018 TSERVTDGMLALGTAGKWTGGICPSGMKSVRMQ CGDKYHSYLQIDPDTIWRPKMLYELLLDGNPITRIDM2\21326271.1119PATENT Docket No. J4040-99018 RKKCGVTIKSVSEPIMEGMFGRLVEMIIEWSDEFY SVNLSGEVLRGMTQKALEHGYQLTPCLGYDAVGDM2\21326271.1120PATENT Docket No. J4040-99018 KAAYANGTDTLEEYAANKKKISAEIARLEAELQQ ESNVKPINKKAFAKRVSEIIKYISDPHNSEAAKNQDM2\21326271.1121PATENT Docket No. J4040-99018 Rsa2IN human gut MADIQPVKNGALYIRVSTHLQEELSPDAQKRLLM 476 T metagenome EYAEAHNIIVLKEHIYIDSGISGRSARQRPQFNNMIDM2\21326271.1122PATENT Docket No. J4040-99018 Bt1INT Streptomyces MSPFIAPDVPEHLLDTVRVFLYARQSKGRSDGSD 479 virus phiBT1 VSTEAQLAAGRALVASRNAQGGARWVVAGEFVDM2\21326271.1123PATENT Docket No. J4040-99018 EDPIKIDETYLKEIVYMFHQTFNDLESEKQKEFISK FIRTIRYTVKEQQPIRPDKSKTGKGKQKVIITEVEF *o a recombination sequence in Table 13. In some aspects, the integrase comprises any integrase that recognizes and binds to a recombination sequence in Table 13. In some aspects, the recombination sequence comprises a sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to a recombination sequence in Table 13. In some aspects, the integrase described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to that of an integrase that recognizes and binds to a polypeptide sequence of any one of the recombination sequences in Table 13.
[0206] In some aspects, the integrase described herein may be a tyrosine integrase. A tyrosine integrase may be referred to as a “tyrosine recombinase.” Examples of tyrosine integrases may include: BS codV; BS ripX; BS ydcL; CB tnpA; Col1D; CP4; Cre; D29; DLP12; DN int; EC FimB; EC FimE; EC orf; EC xerC; EC xerD; Φ11; Φ13; Φ80; Φadh; ΦCTX; ΦLC3; FLP; ΦR73; HI orf; HI rci; HI xerC; HI xerD; HK22; HP1; L2; L5; L54; λ; LL orf; LL xerC; LO L5; MJ orf; MP int; MT int; MT orf; MV4; P186; P2; P21; P22; P4; P434; PA sss; PM fimB; pAE1; pCL1; pKD1; pMEA; pSAM2; pSB2; pSB3; pSDL2; pSE101; pSE211; pSM1; pSR1; pWS58; R721; Rci; sF6; SLP1; SM orf; SsrA; SSV1; T12; Tn21; Tn4430; Tn554a; Tn554b; Tn7; Tn916; Tuc; WZ int; XisA; or XisC. DM2\21326271.1124PATENT Docket No. J4040-99018
[0207] Examples of species that contain tyrosine recombinases may include : Bacillus subtilis; Clostridium butyricum; Escherichia coli, Mycobacterium smegmatis, Dichelobacter nodosus; Staphylococcus aureus; E. coli phage; Lactobacillus gasseri; Pseudomonas aeruginosa; Lactococcis lactis; Saccharomyces cerevisiae; Haemophilus influenzae; Mycoplasma sp.; Mycobacterium tuberculosis; Lactobacillus leichmannii; Leuconostoc oenos; Methanococcus jannaschi; Mycobacterium leprae; Mycobacterium paratuberculosis; Lactobacillus delbrueckii; Salmonella typhimurium; Pseudomonas aeruginosa; Proteus mirabilis; Alcaligenes eutrophus; Chlorobium limicola; Kluyveromyces lactis; Amycolatopsis methanolica; Streptomyces ambofaciencs; Zygosaccharomyces bailii; Zygosaccharomyces bisporus; Salmonella dublin; Saccharopolyspora erythraea; Zygosaccharomyces fermentati; Zygosaccharomyces rouxii; Shigella flexneri; Streptomyces coelicolor; Serratia marcescens; Methanosarcina acetivorans; Sultolobus sp.; Streptococcus pyogenes; Bacillus thurinigiensis; Enteroccus faecalis; Lactobacillus lactis; Weeksella zoohelcum; or Anabaena sp.
[0208] In some aspects, the integrase described herein may be a gamma-delta resolvase from the Tn1000 transposon. In some aspects, the integrase described herein may be a Hin recombinase. In some aspects, the integrase described herein may be a Tn3 resolvase from the Tn3 transposon. In some aspects, the integrase described herein may be a Tre recombinase. In some aspects, the integrase described herein may be a Dre recombinase. In some aspects, the integrase described herein may be a Cre recombinase. In some aspects, the integrase described herein may be a flippase (Flp). In some aspects, the integrase described herein may be a KD recombinase. In some aspects, the integrase described herein may be a B2 B3 recombinase. In some aspects, the integrase described herein may be an HK022 integrase. In some aspects, the integrase described herein may be a ParA integrase. In some aspects, the integrase described herein may be a Gin integrase. In some aspects, the integrase described herein the integrase may be an R4 recombinase.
[0209] In some embodiments, the integrase described herein may be a Vika recombinase. In some embodiments, the integrase described herein may be an RDF recombinase. In some embodiments, the integrase described herein may be a φBT1 recombinase. In some embodiments, the integrase described herein may be an R1 recombinase. In some embodiments, the integrase described herein may be an R2 recombinase. In some embodiments, the integrase described herein may be an R3 recombinase. In some embodiments, the integrase described herein may be an R4 integrase. In some embodiments, the integrase described herein may be an R5 integrase. In some DM2\21326271.1125PATENT Docket No. J4040-99018 embodiments, the integrase described herein may be a TP901-1 recombinase. In some embodiments, the integrase described herein may be a A118 recombinase. In some embodiments, the integrase described herein may be a φFC1 recombinase. In some embodiments, the integrase described herein may be a φC1 recombinase. In some embodiments, the integrase described herein may be a MR11 recombinase. In some embodiments, the integrase described herein may be a TG1 recombinase. In some embodiments, the integrase described herein may be a φ370.1 recombinase. In some embodiments, the integrase described herein may be a Wβ recombinase. In some embodiments, the integrase described herein may be a BL3 recombinase. In some embodiments, the integrase described herein may be a SPBc recombinase. In some embodiments, the integrase described herein may be a K38 recombinase. In some embodiments, the integrase described herein may be a Peaches recombinase. In some embodiments, the integrase described herein may be a Veracruz recombinase. In some embodiments, the integrase described herein may be a Rebeuca recombinase. In some embodiments, the integrase described herein may be a Theia recombinase. In some embodiments, the integrase described herein may be a Benedict recombinase. In some embodiments, the integrase described herein may be a KSSJEB recombinase. In some embodiments, the integrase described herein may be a PattyP recombinase. In some embodiments, the integrase described herein may be a Doom recombinase. In some embodiments, the integrase described herein may be a Scowl recombinase. In some embodiments, the integrase described herein may be a Lockley recombinase. In some embodiments, the integrase described herein may be a Switzer recombinase. In some embodiments, the integrase described herein may be a Bob3 recombinase. In some embodiments, the integrase described herein may be a Troube recombinase. In some embodiments, the integrase described herein may be an Abrogate recombinase. In some embodiments, the integrase described herein may be an Anglerfish recombinase. In some embodiments, the integrase described herein may be a Sarfire recombinase. In some embodiments, the integrase described herein may be a SkiPole recombinase. In some embodiments, the integrase described herein may be a ConceptII recombinase. In some embodiments, the integrase described herein may be a Museum recombinase. In some embodiments, the integrase described herein may be a Severus recombinase. In some embodiments, the integrase described herein may be an Airmid recombinase. In some embodiments, the integrase described herein may be a Benedict recombinase. In some embodiments, the integrase described herein may be a Hinder recombinase. In some embodiments, the integrase described herein may be a ICleared recombinase. In some DM2\21326271.1126PATENT Docket No. J4040-99018 embodiments, the integrase described herein may be a Sheen recombinase. In some embodiments, the integrase described herein may be a Mundrea recombinase. In some embodiments, the integrase described herein may be a BxZ2 recombinase. In some embodiments, the integrase described herein may be a φRV recombinase.
[0210] In some aspects, the integrase described herein may be a retrotransposase encoded by R2. In some aspects, the integrase described herein may be a retrotransposase encoded by L1. In some aspects, the integrase described herein may be a retrotransposase encoded by Tol2. In some aspects, the integrase described herein may be a retrotransposase encoded by Tc1. In some aspects, the integrase described herein may be a retrotransposase encoded by Tc3. In some aspects, the integrase described herein may be a retrotransposase encoded by Mariner (Himar 1). In some aspects, the integrase described herein may be a retrotransposase encoded by Mariner (mos 1). In some aspects, the integrase described herein may be a retrotransposase encoded by Minos.
[0211] As can be used herein, Xu et al describes methods for evaluating integrase activity in E. coli and mammalian cells and confirmed at least R4, φC31, φBT1, Bxb1, SPBc, TP901-1 and Wβ integrases to be active on substrates integrated into the genome of HT1080 cells (Xu et al., 2013, Accuracy and efficiency define Bxb1 integrase as the best of fifteen candidate serine recombinases for the integration of DNA into the human genome. BMC Biotechnol. 2013 Oct 20;13:87. Doi: 10.1186 / 1472-6750-13-87). Durrant describes new large serine recombinases (LSRs) divided into three classes distinguished from one another by efficiency and specificity, including landing pad LSRs which outperform wild-type Bxb1 in episomal and chromosomal integration efficiency, LSRs that achieve both efficient and site-specific integration without a landing pad, and multi-targeting LSRs with minimal site-specificity. Additionally, embodiments can include any serine recombinase such as BceINT, SSCINT, SACINT, and INT10 (see Ionnidi et al., 2021; Drag-and-drop genome insertion without DNA cleavage with CRISPR directed integrases. bioRxiv 2021.11.01.466786, doi.org / 10.1101 / 2021.11.01.466786). In some embodiments, the integration site can be selected from an attB site, an attP site, an attL site, an attR site, a lox71 site a Vox site, or a FRT site. In instances in this disclosure that refer to a Cre- lox system, the Cre-lox system is referred to either as a control for programmable gene insertion or as a tool for a recombinase-mediated event separate and distinct from insertion of the donor polynucleotide template (or exogenous nucleic acid) into the integrated recognition site. DM2\21326271.1127PATENT Docket No. J4040-99018
[0212] In some aspects, the integrase described herein may be coupled to a recombination directionality factor (RDF). In some aspects, the integrase described herein may be fused to an RDF. In some aspects, the integrase described herein may be linked to an RDF. The RDF may comprise a gp3 RDF. The RDF may comprise a gp47 RDF. Fusion Proteins
[0213] Disclosed herein are fusion proteins. Some aspects include a nucleic acid (e.g., an expression vector) encoding a fusion protein. The fusion protein may include an endonuclease. The fusion protein may include a ligase. The fusion protein may include an integrase. The fusion protein may include a linker. The fusion protein may include two linkers. The fusion protein may include a plurality of linkers. The endonuclease and ligase may be connected through a linker. The endonuclease and the integrase may be connected through a linker. The ligase and the integrase may be connected through a linker. The fusion protein may be an example of a covalently coupled endonuclease and DNA ligase. The fusion protein may be an example of a covalently coupled endonuclease and an integrase. The fusion protein may be an example of a covalently coupled DNA ligase and an integrase. The fusion protein may comprise an endonuclease such as an RNA- guided endonuclease fused to a DNA ligase. The fusion protein may comprise an endonuclease such as an RNA-guided endonuclease fused to an integrase such as a serine integrase or a tyrosine integrase. The fusion protein may comprise a DNA ligase fused to an integrase.
[0214] The fusion protein may be non-naturally occurring. The fusion protein may be engineered. The fusion protein may be synthetic. The fusion protein may be pre-synthetized. The fusion protein may be added to a subject or a cell. The fusion protein may be encoded by a nucleic acid. The encoding nucleic acid may be engineered, synthetic, or added to a subject or a cell.
[0215] The fusion protein may be a double fusion protein. The double fusion protein may include an endonuclease such as an RNA-guided endonuclease and a ligase. The double fusion protein may include an endonuclease such as an RNA-guided endonuclease and an integrase. The double fusion protein may include a ligase and an integrase.
[0216] The double fusion protein including an RNA-guided endonuclease and a DNA ligase may include one of various orientations. For example, the double fusion protein may include an RNA-guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the DNA ligase. The double fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the DNA ligase. The double fusion protein DM2\21326271.1128PATENT Docket No. J4040-99018 may include an RNA-guided endonuclease carboxy (C)-terminal to the DNA ligase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the ligase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the ligase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The ligase may be N-terminal. The ligase may be C-terminal.
[0217] The double fusion protein including an RNA-guided endonuclease and an integrase may include one of various orientations. For example, the double fusion protein may include an RNA-guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the integrase. The double fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the integrase. The double fusion protein may include an RNA-guided endonuclease carboxy (C)-terminal to the integrase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the integrase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the integrase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The integrase may be N-terminal. The integrase may be C-terminal.
[0218] The double fusion protein including a ligase and an integrase may include one of various orientations. For example, the double fusion protein may include an integrase upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the ligase. The double fusion protein may include an integrase (N)-terminal to the ligase. The double fusion protein may include an integrase carboxy (C)-terminal to the ligase. The integrase may be in the amino direction within the fusion polypeptide relative to the ligase. The integrase may be in the carboxy direction within the fusion polypeptide relative to the ligase. The ligase may be N-terminal. The ligase may be C-terminal. The integrase may be N-terminal. The integrase may be C-terminal.
[0219] The fusion protein may be a triple fusion protein. The triple fusion protein may include an endonuclease such as an RNA-guided endonuclease, a ligase, and an integrase.
[0220] The triple fusion protein including an endonuclease, a ligase, and an integrase may include one of various orientations. For example, the triple fusion protein may include an RNA- guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C- terminal or in the C-direction) relative to the DNA ligase. The triple fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the DNA ligase. The triple fusion protein may DM2\21326271.1129PATENT Docket No. J4040-99018 include an RNA-guided endonuclease carboxy (C)-terminal to the DNA ligase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the ligase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the ligase. The triple fusion protein may include an RNA-guided endonuclease upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C-direction) relative to the integrase. The triple fusion protein may include an RNA-guided endonuclease amino (N)-terminal to the integrase. The triple fusion protein may include an RNA-guided endonuclease carboxy (C)- terminal to the integrase. The endonuclease may be in the amino direction within the fusion polypeptide relative to the integrase. The endonuclease may be in the carboxy direction within the fusion polypeptide relative to the integrase. The triple fusion protein may include an integrase upstream (e.g., N-terminal or in the N-direction) or downstream (e.g., C-terminal or in the C- direction) relative to the ligase. The triple fusion protein may include an integrase (N)-terminal to the ligase. The triple fusion protein may include an integrase carboxy (C)-terminal to the ligase. The integrase may be in the amino direction within the fusion polypeptide relative to the ligase. The integrase may be in the carboxy direction within the fusion polypeptide relative to the ligase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The ligase may be N-terminal. The ligase may be C-terminal. The integrase may be N-terminal. The integrase may be C-terminal.
[0221] The fusion protein may include a nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or exonuclease. The fusion protein may include a nuclear localization signal. The fusion protein may include a chromatin modifying domain. The fusion protein may include a cell penetrating peptide. The fusion protein may include a tag polypeptide. The fusion protein may include an exonuclease. Any of the nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or exonuclease, endonuclease, ligase, or integrase may be directly connected to another or to the endonuclease, ligase, or integrase. Any of the nuclear localization signal, chromatin modifying domain, cell penetrating peptide, tag polypeptide, or exonuclease, endonuclease, ligase, or integrase may be connected by a linker to another or to the endonuclease, ligase, or integrase. Multiple linkers may be included in the fusion protein. The fusion protein may exclude a polymerase.
[0222] A linker may include an amino acid linker. The amino acid linker may include a length of residues. The length may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, DM2\21326271.1130PATENT Docket No. J4040-99018 90, or 100 residues, or a range of residues defined by any two of the aforementioned integers. The length may include at least 1 residue, at least 2 residues, at least 3 residues, at least 4 residues, at least 5 residues, at least 6 residues, at least 7 residues, at least 8 residues, at least 9 residues, at least 10 residues, at least 15 residues, at least 20 residues, at least 25 residues, at least 30 residues, at least 40 residues, at least 50 residues, at least 60 residues, at least 70 residues, at least 80 residues, at least 90 residues, or at least 100 residues. In some aspects, the length may include less than 2 residues, less than 3 residues, less than 4 residues, less than 5 residues, less than 6 residues, less than 7 residues, less than 8 residues, less than 9 residues, less than 10 residues, less than 15 residues, less than 20 residues, less than 25 residues, less than 30 residues, less than 40 residues, less than 50 residues, less than 60 residues, less than 70 residues, less than 80 residues, less than 90 residues, or less than 100 residues. Examples of residues may include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine, or any combination thereof. The linker may be non-enzymatic or may lack any enzymatic activity.
[0223] For L-PGI systems, certain ligases may exhibit higher efficiency. For example, T4 DNA ligase has been shown to be very effective. Regarding protein architecture, delivering the nuclease (e.g., nCas9) and the DNA ligase (e.g., T4 ligase) as separate proteins, encoded by separate mRNAs, may result in higher editing efficiency compared to delivering them as a single fusion protein, potentially due to factors related to mRNA size, stability, or translation. Co- localization of the separate nuclease and ligase proteins may be facilitated using interacting heterodimerization domains, such as leucine zippers (e.g., LZ1 / LZ2 pairs or EE / RR pairs), fused to each protein. Furthermore, the nuclease component (e.g., nCas9) may optionally be fused to additional domains to enhance performance, such as a Rad51 DNA binding domain (Rad51DBD) to potentially stabilize the DNA flap, or chromatin-modifying domains like high-mobility group nucleosome binding domain 1 (HN1) and / or histone H1 central globular domain (H1G) to potentially increase accessibility to the target locus, particularly in challenging chromatin environments.
[0224] A connection may be covalent. A covalent connection may include a peptide bond. The peptide bond may include amide bond. A connection may be between an N-terminus and another N-terminus. A connection may be between a C-terminus and another C-terminus. A connection DM2\21326271.1131PATENT Docket No. J4040-99018 may be between an N-terminus and a C-terminus. A connection may be between a C-terminus and an N-terminus.
[0225] The fusion protein may include connections in various orientations. The endonuclease may be connected at its C-terminus. The endonuclease may be connected at its N-terminus. The ligase may be connected at its C-terminus. The ligase may be connected at its N-terminus. The integrase may be connected at its C-terminus. The integrase may be connected at its N-terminus.
[0226] Fig. 7 illustrates some examples of fusion proteins including an endonuclease and a ligase. The figure includes examples of arrangements and orientations of the endonuclease, linker, ligase, or nuclear localization signal. Other aspects may be incorporated into the examples shown.
[0227] Fig.13A and 13B illustrate some examples of fusion proteins including at least two of an endonuclease, a ligase, and an integrase. The figures includes examples of arrangements and orientations of the endonuclease, ligase, or integrase. Fig.13A includes examples of double fusion proteins. Fig.13B includes examples of triple fusion proteins. Other aspects may be incorporated into the examples shown.
[0228] In some embodiments, a fusion protein contemplated herein can comprise a DNA- binding domain of the Rad51 DNA repair protein (rad51DBD). In some embodiments, a fusion protein can comprise a high-mobility group nucleosome binding domain 1 (HN1) and a histone H1 central globular domain (H1G). In some embodiments, a fusion protein can comprise a Rad51DBD, and an HN1 and an H1G.
[0229] In some embodiments, a fusion protein contemplated herein comprises Brex27. In some embodiments, Brex27 can be fused to an endonuclease. In some embodiments, Brex27 can be fused to nCas9.
[0230] Table 9 provides example fusion proteins, which may be useful for a genome revising system. In some embodiments, a rad51DBD, and / or an HN1 and a H1G are fused to the fusion protein. In some embodiments, Brex27 is fused to the fusion protein. A fusion protein contemplated herein can comprise an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to an amino acid sequence provided in Table 9. Table 9. Non-limiting examples of DNA binding domain from Rad51 or chromatinDM2\21326271.1132PATENT Docket No. J4040-99018 Name Sequence SEQID NO:DM2\21326271.1133PATENT Docket No. J4040-99018 GTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFD SGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFDM2\21326271.1134PATENT Docket No. J4040-99018 PLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNG YAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKDM2\21326271.1135PATENT Docket No. J4040-99018 GKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGD SLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEnic mRNAs encoding for phosphomimetic peptide from IGF1(IGF1pm1) and N-terminal peptide from NFATC2IP (NFATC2IPp1) peptides (IN peptides).
[0232] In some embodiments, a fusion protein comprises an mRNA encoding an IN peptide as described in Table 30 are fused to the fusion protein. A fusion protein contemplated herein can comprise an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to an amino acid sequence provided in Table 30. Non-Covalently Coupled Proteins
[0233] Disclosed herein are non-covalently coupled proteins. Some aspects relate to a nucleic acid (e.g., an expression vector) encoding a protein, or encoding at least part of a protein. The proteins may include an endonuclease such as an RNA-guided endonuclease. A protein of the non- covalently coupled proteins may include a portion of an endonuclease. A protein of the non- covalently coupled proteins may include a portion of a ligase. The proteins may include a ligase such as a DNA ligase. A protein of the non-covalently coupled proteins may include an integrase. A protein of the non-covalently coupled proteins may include a portion of an integrase (e.g. a functional integrase fragment). The proteins may include an integrase such as a serine integrase. DM2\21326271.1136PATENT Docket No. J4040-99018 The proteins may include an integrase such as a tyrosine integrase. A protein of the non-covalently coupled proteins may include a fusion protein.
[0234] The non-covalently coupled proteins may be bound together through heterodimerization domains. Examples of heterodimerization domains may include a leucine zipper, PDZ domain, streptavidin, streptavidin binding protein, foldon domain, hydrophobic moiety, or a functional binding fragment thereof. A heterodimerization domain may include a leucine zipper. A heterodimerization domain may include a PDZ domain. A heterodimerization domain may include a streptavidin. A heterodimerization domain may include a streptavidin binding protein. A heterodimerization domain may include a foldon domain. A heterodimerization domain may include a hydrophobic moiety. A heterodimerization domain may include an antibody or antibody fragment. The non-covalently coupled proteins may be bound together through inteins.
[0235] The endonuclease and ligase may be coupled together by a separate molecule. The endonuclease and the integrase may be coupled together by a separate molecule. The ligase and the integrase may be coupled together by a separate molecule. The separate molecule may comprise a nucleic acid (e.g., a guide nucleic acid). The ligase may include a hairpin binding motif, where the RNA-guided endonuclease and the DNA ligase are coupled with the nucleic acid. The ligase may include a hairpin binding motif, where the integrase and the DNA ligase are coupled with the nucleic acid. The nucleic acid may include a scaffold that binds the RNA-guided endonuclease and a hairpin that binds to the hairpin binding motif. The hairpin binding motif may include an MS2 coat protein (MCP) peptide. The hairpin may include an MS2 hairpin.
[0236] The endonuclease and ligase may be coupled together by a heterobifunctional molecule. The endonuclease and integrase may be coupled together by a heterobifunctional molecule. The ligase and integrase may be coupled together by a heterobifunctional molecule. The heterobifunctional molecule may include an endonuclease binding domain and a DNA ligase binding domain. The heterobifunctional molecule may include an endonuclease binding domain and an integrase binding domain. The heterobifunctional molecule may include a ligase binding domain and an integrase binding domain. The heterobifunctional molecule may include an endonuclease binding domain. The endonuclease binding domain may include a heterodimerization domain. The endonuclease binding domain may include an antibody or antibody binding fragment. The heterobifunctional molecule may include a ligase binding domain such as a DNA ligase binding domain. The DNA ligase binding domain may include a DM2\21326271.1137PATENT Docket No. J4040-99018 heterodimerization domain. The DNA ligase binding domain may include an antibody or antibody binding fragment. The heterobifunctional molecule may include an integrase binding domain such as a serine integrase or a tyrosine integrase binding domain. The integrase binding domain may include a heterodimerization domain. The integrase binding domain may include an antibody or antibody binding fragment. The heterobifunctional molecule may include a small molecule. The small molecule may comprise a proteolysis targeting chimera (PROTAC), or a related heterobifunctional molecule.
[0237] Some aspects include a protein complex, comprising: an RNA-guided endonuclease bound to a DNA ligase. The endonuclease and the DNA ligase may be bound together through heterodimerization domains. The protein complex of embodiment 75, wherein the heterodimerization domains may comprise leucine zippers, PDZ domains, streptavidin, and streptavidin binding protein, foldon domains, hydrophobic polypeptides, an antibody that binds the Cas nickase, or an antibody that binds the DNA ligase, or one or more binding fragments thereof. The protein complex may be included in a cell. The cell may further include a heterologous RNA-guided endonuclease and a DNA ligase that that was introduced into the cell. The cell may further include a nuclease that is different from the RNA-guided endonuclease.
[0238] In some aspects, a protein complex contemplated herein can comprise a monomeric streptavidin (mSA). In some embodiments, the monomeric streptavidin can be fused to an endonuclease. In some embodiments, the monomeric streptavidin can be fused to a nicking Cas9. In some embodiments, the monomeric streptavidin can be fused to a ligase. In some embodiments, the monomeric streptavidin can be attached to a biotinylated splinting nucleic acid. In some embodiments, the biotin modification is on the / 5Biosg / end of the splinting nucleic acid. In some embodiments, the biotin modification is on the / 3Bio / end of the splinting nucleic acid.
[0239] Fig.20 illustrates a guide nucleic acid, an endonuclease, a ligase, and a donor strand at a genomic locus. In this illustration, a biotinylated splinting nucleic acid is attached to a monomeric streptavidin fused to nicking Cas9. The monomeric streptavidin can also be fused to the ligase (not illustrated).
[0240] Table 10 provides non-limiting examples of fusion protein comprising mSA contemplated herein. A fusion protein can be any fusion protein comprising an amino acid sequence provided in Table 10. A fusion protein contemplated herein can comprise an amino acid DM2\21326271.1138PATENT Docket No. J4040-99018 sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to an amino acid sequence provided in Table 10. Table 10. Non-limiting examples of fusion proteins comprising monomeric streptavidin (mSA) fused to nCas or T4 DDM2\21326271.1139PATENT Docket No. J4040-99018 ADGNLTGQYENRAQGTGCQNSPYTLTGRYNGTKLEWRVEW NNSTENCHSRTEWRGQYQGGAEARINTQWNLTYEGGSGPADM2\21326271.1140PATENT Docket No. J4040-99018 ESRTASNGIANKSLKGTISEKEAQCMKFQVWDYVPLVEIYSL PAFRLKYDVRFSKLEQMTSGYDKVILIENQVVNNLDEAKVIY
[0241] Disclosed herein are guide nucleic acids. The guide nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate to a nucleic acid (e.g., DNA or an expression vector) that encodes a guide nucleic acid such as a guide RNA. Provided herein are guide nucleic acids (e.g., gRNAs) that direct a programmable endonuclease (e.g., a nCas9) to a target nucleic acid (e.g., a genomic locus). The guide nucleic acid may guide an RNA-guided endonuclease to a target nucleic acid locus for nucleic acid replacement or gene editing at the locus. A guide nucleic acid of the present disclosure may facilitate a donor strand to be inserted into a target site of the target nucleic acid. A guide nucleic acid of the present disclosure may facilitate editing of a nucleic acid sequence at a target site of the target nucleic acid. The guide nucleic acid may, in some instances, also act as a splint for a DNA ligase described herein, such as for ligating two nucleic acid strands base paired to a portion of the guide nucleic acid. The guide nucleic acid may be single stranded. The guide nucleic acid may include RNA. The guide nucleic acid may be RNA. The guide nucleic acid may include a guide RNA (gRNA). In some cases, a guide nucleic acid may include DNA.
[0242] The guide nucleic acid may be non-naturally occurring. The guide nucleic acid may be engineered. The guide nucleic acid may be synthetic. The guide nucleic acid may be pre- synthetized. The guide nucleic acid may be added to a subject or a cell. In some aspects, the guide nucleic acid does not include a template for a polymerase.
[0243] The guide nucleic acid may include an exogenous first integrating nucleic acid binding site. The exogenous first integrating nucleic acid binding site may be referred to as a “donor binding site” or vice versa. DM2\21326271.1141PATENT Docket No. J4040-99018
[0244] Disclosed herein are guide nucleic acids, comprising: a spacer reverse complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; and a donor nucleic acid binding site and optionally a flap binding site reverse complementary to a nucleic acid flap.
[0245] In some aspects, the guide nucleic acid comprises a spacer complementary to a genomic locus in a cell; a scaffold for complexing with the at least one endonuclease; a donor binding site that is at least partially complementary to a donor strand; a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; or a combination thereof. In some aspects, the guide nucleic acid can direct the at least one endonuclease to cleave at least one strand of the genomic locus. In some aspects, the guide nucleic acid can be at least partially complementary to the donor strand or at least partially complementary to a genomic flap (e.g., a genomic nucleic acid sequence that is displaced and become single-stranded when the guide nucleic acid recruits the endonuclease to the genomic locus). In some aspects, the guide nucleic acid, being at least partially complementary to the donor strand or at least partially complementary to a genomic flap, brings the donor strand to close proximity of the cleaving of the genomic locus.
[0246] In some embodiments, in an L-PGI system, the guide nucleic acid may be referred to as a ligation-initiating guide RNA (lmgRNA). An lmgRNA may comprise, in addition to the spacer and scaffold, a 3' extension sequence termed a Splint Binding Site (SBS), configured to hybridize with a corresponding Guide Binding Site (GBS) on a splint nucleic acid. The stability and function of the lmgRNA may be enhanced through chemical modifications, such as incorporating 2'-O- methyl (2'OMe) nucleotides and / or phosphorothioate (PS) linkages at or near the 5' and 3' termini. The length and sequence of the SBS may be optimized, for instance, an SBS providing a hybridization region of about 19-20 base pairs with the splint GBS, potentially including sequences that form secondary structures, may enhance performance in some systems.
[0247] Disclosed herein, in some embodiments, are guide nucleic acids comprising a scaffold. The scaffold may bind a nuclease. The scaffold may bind a Cas nuclease. The scaffold may bind a nickase. The scaffold may bind a Cas nickase. The scaffold may bind an S. Pyogenes Cas9 nuclease. The scaffold may bind an S. Pyogenes Cas9 nickase. The scaffold may include a scaffold nucleic acid sequence. A system described herein may include a first guide nucleic acid. The system can include a second guide nucleic acid. The first guide nucleic acid may bind to a first Cas nickase. The second guide nucleic acid may bind to a second Cas nickase. DM2\21326271.1142PATENT Docket No. J4040-99018
[0248] A guide nucleic acid may include any aspect of (i)-(iv): (i) a spacer complementary to a region of a genomic locus of a genomic strand, (ii) a scaffold for complexing with an RNA- guided endonuclease, (iii) a donor binding site that is at least partially complementary to an exogenous first integrating nucleic acid, or (iv) a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus. A guide nucleic acid may include any aspect of (i)-(iii): (i) a spacer complementary to a region of a genomic locus of a genomic strand, (ii) a scaffold for complexing with an RNA-guided endonuclease, or (iii) a donor binding site that is at least partially complementary to a splinting nucleic acid. A component of (i), (ii), or (iii) may be included in a single guide nucleic acid or may be split between or collectively included among multiple guide nucleic acids.
[0249] In some aspects, the guide nucleic acid comprises a modified internucleoside linkage. In some aspects, the modified internucleoside linkage comprises a phosphorothioate linkage. In some aspects, the modified internucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the guide nucleic acid. The guide nucleic acid may include multiple modified internucleoside linkages. For example, the guide nucleic acid may include modified internucleoside linkages at nucleic acids of the 5’ and 3’ ends of the guide nucleic acid, such as between the last 4 nucleic acids at the 5’ end and between the last 4 nucleic acids at the 3’ end. In some aspects, the guide nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, or a combination thereof. The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. The modified nucleoside may include an LNA. The modified nucleoside may include a 2’fluoro. The modified nucleoside may include a 2’ O-alkyl. The modified nucleoside may include a methylated cytosine. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end of the guide nucleic acid. The guide nucleic acid may include multiple modified nucleosides. For example, the guide nucleic acid may include modified nucleosides at nucleic acids of the 5’ and 3’ ends of the guide nucleic acid, such as the last 3 nucleic acids at the 5’ end and the last 3 nucleic acids at the 3’ end.
[0250] In some aspects, the guide nucleic acid comprises at least one nucleic acid modification. In some aspect, the at least nucleic acid modification comprises modifying a backbone, a sugar, a base, or a combination thereof of the guide nucleic acid. In some aspects, the at least one nucleic DM2\21326271.1143PATENT Docket No. J4040-99018 acid modification can increase resistance of the guide nucleic acid to degradation (e.g., against nuclease degradation or hydrolysis). In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the at least one endonuclease. In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the donor strand. In some aspects, the at least one nucleic acid modification can increase the complexing of the guide nucleic acid to the genomic locus via by being complementary to the genomic flap.
[0251] In some aspects, the guide nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleic acid modifications. In some aspects, nucleic acid modification can occur at 3’OΗ group, 5’OΗ group, at the backbone, at the sugar component, or at the nucleotide base. Nucleic acid modification can include non-naturally occurring linker molecules of interstrand or intrastrand cross links. In one aspect, the modified nucleic acid comprises modification of one or more of the 3’OΗ or 5’OΗ group, the backbone, the sugar component, or the nucleotide base, or addition of non-naturally occurring linker molecules. In some aspects, modified backbone comprises a backbone other than a phosphodiester backbone. In some aspects, a modified sugar comprises a sugar other than deoxyribose (in modified DNA) or other than ribose (modified RNA). In some aspects, a modified base comprises a base other than adenine, guanine, cytosine, thymine, or uracil. In some aspects, the guide nucleic acid comprises at least one modified base. In some instances, the guide nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 15, 20, or more modified bases. In some cases, the nucleic acid modifications to the base moiety include natural and synthetic modifications of adenine, guanine, cytosine, thymine, or uracil, and purine or pyrimidine bases.
[0252] In some aspects, the at least one nucleic acid modification of the guide nucleic acid comprises a modification of any one of or any combination of: 2' modified nucleotide comprising 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl, 2'-deoxy, 2'-deoxy-2'-fluoro, 2'- O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-O-N-methylacetamido (2'-O-NMA); modification of one or both of the non-linking phosphate oxygens in the phosphodiester backbone linkage; modification of one or more of the linking phosphate oxygens DM2\21326271.1144PATENT Docket No. J4040-99018 in the phosphodiester backbone linkage; modification of a constituent of the ribose sugar; replacement of the phosphate moiety with “dephospho” linkers; modification or replacement of a naturally occurring nucleobase; modification of the ribose-phosphate backbone; modification of 5’ end of polynucleotide; modification of 3’ end of polynucleotide; modification of the deoxyribose phosphate backbone; substitution of the phosphate group; modification of the ribophosphate backbone; modifications to the sugar of a nucleotide; modifications to the base of a nucleotide; or stereopure of nucleotide. Non limiting examples of nucleic acid modification to the guide nucleic acid can include: modification of one or both of non-linking or linking phosphate oxygens in the phosphodiester backbone linkage (e.g., sulfur (S), selenium (Se), BR3 (wherein R can be, e.g., hydrogen, alkyl, or aryl), C (e.g., an alkyl group, an aryl group, and the like), H, NR2, wherein R can be, e.g., hydrogen, alkyl, or aryl, or wherein R can be, e.g., alkyl or aryl); replacement of the phosphate moiety with “dephospho” linkers (e.g., replacement with methyl phosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino); modification or replacement of a naturally occurring nucleobase with nucleic acid analog; modification of deoxyribose-phosphate or ribose-phosphate backbone (e.g., modifying the ribose-phosphate backbone to incorporate phosphorothioate, phosphonothioacetate, phosphoroselenates, boranophosphates, borano phosphate esters, hydrogen phosphonates, phosphonocarboxylate, phosphoroamidates, alkyl or aryl phosphonates, phosphonoacetate, or phosphotriesters; modification of 5’ end (e.g., 5’ cap or modification of 5’ cap -OH) or 3’ end of the nucleic acid sequence (3’ tail or modification of 3’ end -OH); substitution of the phosphate group with methyl phosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino; modification of the ribophosphate backbone to incorporate morpholino (phosphorodiamidate morpholino oligomer PMO), cyclobutyl, pyrrolidine, or peptide nucleic acid (PNA) nucleoside surrogates; modifications to the sugar of a nucleotide to incorporate locked nucleic acid (LNA), unlocked nucleic acid (UNA), ethylene nucleic acid (ENA), constrained ethyl (cEt) sugar, or bridged nucleic acid (BNA); modification of a constituent of the ribose sugar (e.g., 2’-O-methyl, 2’-O-methoxy-ethyl (2’-MOE), 2’-fluoro, 2’-aminoethyl, 2’- DM2\21326271.1145PATENT Docket No. J4040-99018 deoxy-2’-fuloarabinou-cleic acid, 2′-deoxy, 2′-O-methyl, 3′-phosphorothioate, 3′- phosphonoacetate (PACE), or 3′-phosphonothioacetate (thioPACE)); modification to the base of a nucleotide (of A, T, C, G, or U); and stereopure of nucleotide (e.g., S conformation of phosphorothioate or R conformation of phosphorothioate).
[0253] The following codes may be used in sequences set forth herein. Nucleic Acid Modification Codes * Ph h hi b d
[0254] In some aspects, the nucleic acid modification comprises at least one substitution of one or both of non-linking phosphate oxygen atoms in a phosphodiester backbone linkage of the guide nucleic acid. In some aspects, the at least one nucleic acid modification of the guide nucleic acid comprises a substitution of one or more of linking phosphate oxygen atoms in a phosphodiester backbone linkage of the guide nucleic acid. A non-limiting example of a nucleic acid modification of a phosphate oxygen atom is a sulfur atom. In some aspects, the nucleic acid modification comprises at least one modification to a sugar. In some aspects, the nucleic acid modification comprises at least one nucleic acid modification to the sugar comprising a modification of a constituent of the sugar, where the sugar is a ribose sugar. In some aspects, the DM2\21326271.1146PATENT Docket No. J4040-99018 nucleic acid modification of the guide nucleic acid comprises at least one modification to the constituent of the ribose sugar of the nucleotide of the guide nucleic acid comprising a 2’-O-Methyl group. In some aspects, the nucleic acid modification comprises at least one modification comprising replacement of a phosphate moiety of the guide nucleic acid with a dephospho linker. In some aspects, the nucleic acid modification of comprises at least one modification of a phosphate backbone. In some aspects, the modification comprises a phosphorothioate group. In some aspects, the nucleic acid modifications comprises at least one modification comprising a modification to a base of a nucleotide of the guide nucleic acid. In some aspects, the nucleic acid modifications comprises at least one modification comprising an unnatural base of a nucleotide. In some aspects, the nucleic acid modifications comprises at least one modification comprising at least one stereopure nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 5’ end of the guide nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 3’ end of the guide nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to both 5’ and 3’ ends of the guide nucleic acid.
[0255] In some aspects, the guide nucleic acid described herein comprises a backbone comprising a plurality of sugar and phosphate moieties covalently linked together. In some cases, a backbone of the guide nucleic acid comprises a phosphodiester bond linkage between a first hydroxyl group in a phosphate group on a 5’ carbon of a deoxyribose in DNA or ribose in RNA and a second hydroxyl group on a 3’ carbon of a deoxyribose in DNA or ribose in RNA. In some aspects, a backbone of the guide nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to a solvent. In some aspects, a backbone of the guide nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to nucleases. In some aspects, a backbone of the guide nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to hydrolytic enzymes. In some instances, a backbone of the guide nucleic acid can be represented as a polynucleotide sequence in a circular 2-dimensional format with one nucleotide after the other. In some instances, a backbone of the guide nucleic acid can be represented as a polynucleotide sequence in a looped 2-dimensional format with one nucleotide after the other. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are joined through a phosphorus-oxygen bond. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are modified into a phosphoester with a phosphorus-containing moiety. In DM2\21326271.1147PATENT Docket No. J4040-99018 some aspects, the guide nucleic acid comprises at least one nucleic acid modification comprising any one of: 5′ adenylate, 5′ guanosine-triphosphate cap, 5′N7-Methylguanosine-triphosphate cap, 5′triphosphate cap, 3′phosphate, 3′thiophosphate, 5′phosphate, 5′thiophosphate, Cis-Syn thymidine dimer, trimers, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer 18, Spacer 9,3′-3′ modifications, 5′-5′ modifications, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-Biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3′DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linkers, 2′deoxyribonucleoside analog purine, 2′deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2′-O-methyl ribonucleoside analog, sugar modified analogs, wobble / universal bases, fluorescent dye label, 2′fluoro RNA, 2′O-methyl RNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphothioate DNA, phosphorothioate RNA, UNA, LNA, cEt, pseudouridine-5′-triphosphate, 5-methylcytidine-5′- triphosphate, 2-O-methyl -phosphorothioate or any combinations thereof.
[0256] A nucleic acid modification can also be a phosphorothioate substitute. In some cases, a natural phosphodiester bond can be susceptible to rapid degradation by cellular nucleases and; a modification of internucleotide linkage using phosphorothioate (PS) bond substitutes can be more stable towards hydrolysis by cellular degradation. A modification can increase stability in a polynucleic acid. A modification can also enhance biological activity. In some cases, a phosphorothioate enhanced RNA polynucleic acid can inhibit RNase A, RNase T1, calf serum nucleases, or any combinations thereof. These properties can allow the use of PS-RNA polynucleic acids to be used in applications where exposure to nucleases is of high probability in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced between the last 3-5 nucleotides at the 5′-or 3′-end of a polynucleic acid which can inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added throughout an entire polynucleic acid to reduce attack by endonucleases. In some aspects, the guide nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 100, or more internucleotide linkage comprising PS bond. In some aspects, the guide nucleic acid comprises only PS bond as the internucleotide linkage modification. In some aspects, all internucleotide linkages of the guide nucleic acid herein are fully PS-modified or include phosphorothioate internucleotide linkages. DM2\21326271.1148PATENT Docket No. J4040-99018
[0257] The guide nucleic acid may include a hairpin. The hairpin may bind to a hairpin binding motif such as a hairpin binding motif on a DNA ligase. The hairpin may include an MS2 hairpin A hairpin such as an MS2 hairpin may be useful for recruiting a DNA ligase that includes an MCP peptide.
[0258] The guide nucleic acid may include any aspect included in Fig. 1A-6C. Table 11 illustrates non-limiting examples of some of the guide nucleic acids described herein. Some of the guide nucleic acids in the table include nucleic acid modifications. Table 11. Examples of nucleic acid sequences Q :DM2\21326271.1149PATENT Docket No. J4040-99018 Rep1.BFP.Rev mG*mA*mC*GUAGCCUUCGGGCAUGGGUUUAAGAGCUAUG 101 Guide.SpPAM CUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUDM2\21326271.1150PATENT Docket No. J4040-99018 Rep2.BFP2GF / 5Phos / AaccggcaagctgccGgtgccctggcccaccCTCgtgaccaccctgaccTACg 112 P.TopDonor.S gcgtgcagtgcttcagccgctaccccgaccacatgaagcagcacgacttcttcaagtcAgccatgDM2\21326271.1151PATENT Docket No. J4040-99018 Rep2.mGL- ATGgtgagcaagggcgaggagctgttcaccggggtggtgcccatcctggtcgagctggacgg 122 CBX1.TopDon cgacgtaaacggccacaagttcagcgtccgcggcgagggcgagggcgatgccaccaacggca[ ] e gu e nuc e c ac may nc u e a sequence o n ng nuc e c ac s (e.g., n ng RNA or DNA nucleotides) between components of the guide nucleic acid. For example, the guide nucleic acid may include a sequence of linking nucleic acids between any of the following components: a spacer, a scaffold, a donor binding site, or a flap binding site. The guide nucleic acid may include a sequence of linking nucleic acids between a spacer, a scaffold, or a donor binding site. The guide nucleic acid include a sequence of linking nucleic acids between the scaffold and the donor binding site The guide nucleic acid may include a sequence of linking DM2\21326271.1152PATENT Docket No. J4040-99018 nucleic acids between a spacer and a scaffold. The guide nucleic acid may include multiple sequences of linking nucleic acids between components.
[0260] The sequence of linking nucleic acids may include any base, such as A, U, T, G, or C, or a combination thereof. The sequence of linking nucleic acids may include A, T, G, or C, or a combination thereof. The sequence of linking nucleic acids may include A, U, G, or C, or a combination thereof. The sequence of linking nucleic acids may include a series of As. The sequence of linking nucleic acids may include a series of Ts. The sequence of linking nucleic acids may include a series of Us. The sequence of linking nucleic acids may include a series of Cs. The sequence of linking nucleic acids may include a series of Gs.
[0261] The sequence of linking nucleic acids may include a length, such as a number of nucleotides. The length may include 1, 2, 3, 4, 5, 6, 7, 8, 910, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides, or a range defined by any two of the aforementioned numbers of nucleotides. The length may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 910, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides. In some aspects, the length may be less than 2, less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9 10, less than 11, less than 12, less than 13, less than 14, less than 15, less than 16, less than 17, less than 18, less than 19, less than 20, less than 21, less than 22, less than 23, less than 24, less than 25, less than 30, less than 35, less than 40, less than 45, less than 50, less than 55, less than 60, less than 65, less than 70, less than 75, less than 80, less than 85, less than 90, less than 95, or less than 100 nucleotides.
[0262] Some aspects relate to a guide nucleic acid comprising: a spacer that is at least partially complementary to a genomic locus in a cell; a scaffold for complexing with an RNA-guided endonuclease; and a donor binding site that is at least partially complementary to an exogenous first integrating nucleic acid. The guide nucleic acid may further comprise a flap binding site that is at least partially complementary to a genomic sequence of the genomic locus. The guide nucleic acid may further comprise at least one nucleic acid modification. The at least one nucleic acid DM2\21326271.1153PATENT Docket No. J4040-99018 modification may comprise a modification to a backbone, a sugar, a base, or a combination thereof. The guide nucleic acid may comprise RNA.
[0263] Some aspects include a guide nucleic acid, comprising: a spacer at least partially reverse complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; and a flap binding site at least partially reverse complementary to a nucleic acid flap, and an exogenous first integrating nucleic acid binding site. Splint Nucleic Acids
[0264] Disclosed herein are splinting nucleic acids. The splinting nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate to a nucleic acid (e.g., DNA or RNA). The splinting nucleic acid can comprise a DNA or RNA backbone or “splint” which is reverse complementary to the DNA sequences that are to be ligated. Non-limiting examples include the Replacer guide nucleic acids depicted in Fig. 1A, 2A, 5A, and 6A, which comprise a flap binding site (FBS) adjacent to a donor binding site (DBS). In some embodiments, a splinting nucleic acid can comprise a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus, and can optionally comprise a guide binding site that is at least partially complementary to a guide nucleic acid. Non-limiting examples include Donor 2 nucleic acids depicted in Replacer 2 models of Fig.3A and 4A which comprise a guide binding site (GBS) adjacent to a flap binding site (FBS) adjacent to a donor binding site DBS). In some aspects, the genomic strand is in a cell. In some aspects, the splinting nucleic acid further comprises a donor binding site that is at least partially identical or complementary to a portion of the integrating nucleic acid. In some aspects, the donor nucleic acid comprises a splinting nucleic acid. In some aspects, the guide nucleic acid comprises a splinting nucleic acid.
[0265] In some aspects, the splint nucleic acid plays a crucial role in coordinating the interaction between the lmgRNA, the donor DNA, and the target genomic site. It typically comprises three functional domains: a Guide Binding Site (GBS) at its 3' end that hybridizes to the Splint Binding Site (SBS) of the lmgRNA; a central Flap Binding Site (FBS) that hybridizes to the displaced 3' genomic flap generated by the nickase; and a Donor Binding Site (DBS) at its 5' end that hybridizes to the donor DNA. The lengths of these domains are key parameters for optimization, with typical ranges being, for example, about 15-25 nucleotides (e.g., 19-20 nt) for the GBS, about 10-15 nucleotides (e.g., 11-13 nt) for the FBS, and about 20-30 nucleotides (e.g., DM2\21326271.1154PATENT Docket No. J4040-99018 24-26 nt) for the DBS. To achieve the necessary stability and affinity for efficient L-PGI, the splint nucleic acid often incorporates chemical modifications, particularly within the GBS and DBS regions. Locked Nucleic Acids (LNAs) and / or 2'-O-Methyl (2'OMe) modified bases may be used, often in patterns such as alternating with unmodified DNA bases, to precisely tune hybridization properties. The choice and pattern of modifications (e.g., LNA vs. 2'OMe) may depend on the specific application, such as the length of the donor overlap in pL-PGI systems. Additionally, the 3' end of the splint nucleic acid may be protected from degradation, for example, using a C3 spacer or an inverted dT modification, which can further enhance editing efficiency.
[0266] In some aspects, the splint nucleic acid comprises a modified internucleoside linkage. In some aspects, the modified internucleoside linkage comprises a phosphorothioate linkage. In some aspects, the modified internucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the splint nucleic acid. The splint nucleic acid may include multiple modified internucleoside linkages. For example, the splint nucleic acid may include modified internucleoside linkages at nucleic acids of the 5’ and 3’ ends of the splint nucleic acid, such as between the last 4 nucleic acids at the 5’ end and between the last 4 nucleic acids at the 3’ end. In some aspects, the splint nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, or a combination thereof. The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. The modified nucleoside may include an LNA. The modified nucleoside may include a 2’fluoro. The modified nucleoside may include a 2’ O-alkyl. The modified nucleoside may include a methylated cytosine. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end of the splint nucleic acid. The splint nucleic acid may include multiple modified nucleosides. For example, the splint nucleic acid may include modified nucleosides at nucleic acids of the 5’ and 3’ ends of the splint nucleic acid, such as the last 3 nucleic acids at the 5’ end and the last 3 nucleic acids at the 3’ end.
[0267] In some aspects, the splint nucleic acid comprises at least one nucleic acid modification. In some aspect, the at least nucleic acid modification comprises modifying a backbone, a sugar, a base, or a combination thereof of the splint nucleic acid. In some aspects, the at least one nucleic acid modification can increase resistance of the splint nucleic acid to degradation (e.g., against nuclease degradation or hydrolysis). In some aspects, the at least one nucleic acid modification DM2\21326271.1155PATENT Docket No. J4040-99018 can increase the complexing of the splint nucleic acid to the at least one endonuclease. In some aspects, the at least one nucleic acid modification can increase the complexing of the splint nucleic acid to the donor strand. In some aspects, the at least one nucleic acid modification can increase the complexing of the splint nucleic acid to the genomic locus via by being complementary to the genomic flap.
[0268] In some aspects, the splint nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleic acid modifications. In some aspects, nucleic acid modification can occur at 3’OΗ group, 5’OΗ group, at the backbone, at the sugar component, or at the nucleotide base. Nucleic acid modification can include non-naturally occurring linker molecules of interstrand or intrastrand cross links. In one aspect, the modified nucleic acid comprises modification of one or more of the 3’OΗ or 5’OΗ group, the backbone, the sugar component, or the nucleotide base, or addition of non-naturally occurring linker molecules. In some aspects, modified backbone comprises a backbone other than a phosphodiester backbone. In some aspects, a modified sugar comprises a sugar other than deoxyribose (in modified DNA) or other than ribose (modified RNA). In some aspects, a modified base comprises a base other than adenine, guanine, cytosine, thymine, or uracil. In some aspects, the splint nucleic acid comprises at least one modified base. In some instances, the splint nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 15, 20, or more modified bases. In some cases, the nucleic acid modifications to the base moiety include natural and synthetic modifications of adenine, guanine, cytosine, thymine, or uracil, and purine or pyrimidine bases.
[0269] In some aspects, the at least one nucleic acid modification of the splint nucleic acid comprises a modification of any one of or any combination of: 2' modified nucleotide comprising 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl, 2'-deoxy, 2'-deoxy-2'-fluoro, 2'- O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-O-N-methylacetamido (2'-O-NMA); modification of one or both of the non-linking phosphate oxygens in the phosphodiester backbone linkage; modification of one or more of the linking phosphate oxygens in the phosphodiester backbone linkage; modification of a constituent of the ribose sugar; replacement of the phosphate moiety with “dephospho” linkers; modification or replacement of a DM2\21326271.1156PATENT Docket No. J4040-99018 naturally occurring nucleobase; modification of the ribose-phosphate backbone; modification of 5’ end of polynucleotide; modification of 3’ end of polynucleotide; modification of the deoxyribose phosphate backbone; substitution of the phosphate group; modification of the ribophosphate backbone; modifications to the sugar of a nucleotide; modifications to the base of a nucleotide; or stereopure of nucleotide. Non limiting examples of nucleic acid modification to the splint nucleic acid can include: modification of one or both of non-linking or linking phosphate oxygens in the phosphodiester backbone linkage (e.g., sulfur (S), selenium (Se), BR3 (wherein R can be, e.g., hydrogen, alkyl, or aryl), C (e.g., an alkyl group, an aryl group, and the like), H, NR2, wherein R can be, e.g., hydrogen, alkyl, or aryl, or wherein R can be, e.g., alkyl or aryl); replacement of the phosphate moiety with “dephospho” linkers (e.g., replacement with methyl phosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino); modification or replacement of a naturally occurring nucleobase with nucleic acid analog; modification of deoxyribose-phosphate or ribose-phosphate backbone (e.g., modifying the ribose-phosphate backbone to incorporate phosphorothioate, phosphonothioacetate, phosphoroselenates, boranophosphates, borano phosphate esters, hydrogen phosphonates, phosphonocarboxylate, phosphoroamidates, alkyl or aryl phosphonates, phosphonoacetate, or phosphotriesters; modification of 5’ end (e.g., 5’ cap or modification of 5’ cap -OH) or 3’ end of the nucleic acid sequence (3’ tail or modification of 3’ end -OH); substitution of the phosphate group with methyl phosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino; modification of the ribophosphate backbone to incorporate morpholino (phosphorodiamidate morpholino oligomer PMO), cyclobutyl, pyrrolidine, or peptide nucleic acid (PNA) nucleoside surrogates; modifications to the sugar of a nucleotide to incorporate locked nucleic acid (LNA), unlocked nucleic acid (UNA), ethylene nucleic acid (ENA), constrained ethyl (cEt) sugar, or bridged nucleic acid (BNA); modification of a constituent of the ribose sugar (e.g., 2’-O-methyl, 2’-O-methoxy-ethyl (2’-MOE), 2’-fluoro, 2’-aminoethyl, 2’- deoxy-2’-fuloarabinou-cleic acid, 2′-deoxy, 2′-O-methyl, 3′-phosphorothioate, 3′- phosphonoacetate (PACE), or 3′-phosphonothioacetate (thioPACE)); modification to the base of DM2\21326271.1157PATENT Docket No. J4040-99018 a nucleotide (of A, T, C, G, or U); and stereopure of nucleotide (e.g., S conformation of phosphorothioate or R conformation of phosphorothioate).
[0270] In some aspects, the nucleic acid modification comprises at least one substitution of one or both of non-linking phosphate oxygen atoms in a phosphodiester backbone linkage of the splint nucleic acid. In some aspects, the at least one nucleic acid modification of the splint nucleic acid comprises a substitution of one or more of linking phosphate oxygen atoms in a phosphodiester backbone linkage of the splint nucleic acid. A non-limiting example of a nucleic acid modification of a phosphate oxygen atom is a sulfur atom. In some aspects, the nucleic acid modification comprises at least one modification to a sugar. In some aspects, the nucleic acid modification comprises at least one nucleic acid modification to the sugar comprising a modification of a constituent of the sugar, where the sugar is a ribose sugar. In some aspects, the nucleic acid modification of the splint nucleic acid comprises at least one modification to the constituent of the ribose sugar of the nucleotide of the splint nucleic acid comprising a 2’-O-Methyl group. In some aspects, the nucleic acid modification comprises at least one modification comprising replacement of a phosphate moiety of the splint nucleic acid with a dephospho linker. In some aspects, the nucleic acid modification of comprises at least one modification of a phosphate backbone. In some aspects, the modification comprises a phosphorothioate group. In some aspects, the nucleic acid modifications comprises at least one modification comprising a modification to a base of a nucleotide of the splint nucleic acid. In some aspects, the nucleic acid modifications comprises at least one modification comprising an unnatural base of a nucleotide. In some aspects, the nucleic acid modifications comprises at least one modification comprising at least one stereopure nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 5’ end of the splint nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to a 3’ end of the splint nucleic acid. In some aspects, the at least one nucleic acid modification can be positioned proximal to both 5’ and 3’ ends of the splint nucleic acid.
[0271] In some aspects, the splint nucleic acid described herein comprises a backbone comprising a plurality of sugar and phosphate moieties covalently linked together. In some cases, a backbone of the splint nucleic acid comprises a phosphodiester bond linkage between a first hydroxyl group in a phosphate group on a 5’ carbon of a deoxyribose in DNA or ribose in RNA and a second hydroxyl group on a 3’ carbon of a deoxyribose in DNA or ribose in RNA. In some DM2\21326271.1158PATENT Docket No. J4040-99018 aspects, a backbone of the splint nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to a solvent. In some aspects, a backbone of the splint nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to nucleases. In some aspects, a backbone of the splint nucleic acid can lack a 5’ reducing hydroxyl, a 3’ reducing hydroxyl, or both, capable of being exposed to hydrolytic enzymes. In some instances, a backbone of the splint nucleic acid can be represented as a polynucleotide sequence in a circular 2-dimensional format with one nucleotide after the other. In some instances, a backbone of the splint nucleic acid can be represented as a polynucleotide sequence in a looped 2-dimensional format with one nucleotide after the other. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are joined through a phosphorus-oxygen bond. In some cases, a 5’ hydroxyl, a 3’ hydroxyl, or both, are modified into a phosphoester with a phosphorus-containing moiety. In some aspects, the splint nucleic acid comprises at least one nucleic acid modification comprising any one of: 5′ adenylate, 5′ guanosine-triphosphate cap, 5′N7-Methylguanosine-triphosphate cap, 5′triphosphate cap, 3′phosphate, 3′thiophosphate, 5′phosphate, 5′thiophosphate, Cis-Syn thymidine dimer, trimers, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer 18, Spacer 9,3′-3′ modifications, 5′-5′ modifications, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-Biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3′DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linkers, 2′deoxyribonucleoside analog purine, 2′deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2′-O-methyl ribonucleoside analog, sugar modified analogs, wobble / universal bases, fluorescent dye label, 2′fluoro RNA, 2′O-methyl RNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphothioate DNA, phosphorothioate RNA, UNA, LNA, cEt, pseudouridine-5′-triphosphate, 5-methylcytidine-5′- triphosphate, 2-O-methyl -phosphorothioate or any combinations thereof.
[0272] A nucleic acid modification can also be a phosphorothioate substitute. In some cases, a natural phosphodiester bond can be susceptible to rapid degradation by cellular nucleases and; a modification of internucleotide linkage using phosphorothioate (PS) bond substitutes can be more stable towards hydrolysis by cellular degradation. A modification can increase stability in a polynucleic acid. A modification can also enhance biological activity. In some cases, a phosphorothioate enhanced RNA polynucleic acid can inhibit RNase A, RNase T1, calf serum DM2\21326271.1159PATENT Docket No. J4040-99018 nucleases, or any combinations thereof. These properties can allow the use of PS-RNA polynucleic acids to be used in applications where exposure to nucleases is of high probability in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced between the last 3-5 nucleotides at the 5′-or 3′-end of a polynucleic acid which can inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added throughout an entire polynucleic acid to reduce attack by endonucleases. In some aspects, the splint nucleic acid comprises at least one, two, three, four, five, six, seven, eight, nine, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 100, or more internucleotide linkage comprising PS bond. In some aspects, the splint nucleic acid comprises only PS bond as the internucleotide linkage modification. In some aspects, all internucleotide linkages of the splint nucleic acid herein are fully PS-modified or include phosphorothioate internucleotide linkages.
[0273] The splint nucleic acid may include a hairpin. The hairpin may bind to a hairpin binding motif such as a hairpin binding motif on a DNA ligase. The hairpin may include an MS2 hairpin A hairpin such as an MS2 hairpin may be useful for recruiting a DNA ligase that includes an MCP peptide.
[0274] In some embodiments, the donor binding site (DBS) of the splint is from 12 nt to 50 nt in length, including 13, 14, 15,16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 36, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 nt length. In some embodiments, every third nucleotide is modified. In some embodiments, the DBS comprises alternating LNAs or comprises a region where every third nucleotide is an LNA. In some embodiments, the DBS comprises anternating 2’F nucleotides, or comprises a region where every third nucleotide is 2’F modified. In some embodiments, the DBS comprises alternating 2’OMe nucleotides or comprises a region where every third nucleotide is 2’OMe modified.
[0275] In some embodiments, the guide binding site (GBS) of the splint comprises one or more modified nucleotides. In some embodiments from 10% to 80% of the GBS nucleotides are modified. In some embodiments, every third nucleotide is modified. In some embodiments, the GBS comprises alternating LNAs or comprises a region where every third nucleotide is an LNA. In some embodiments, the GBS comprises anternating 2’F nucleotides, or comprises a region where every third nucleotide is 2’F modified. In some embodiments, the GBS comprises alternating 2’OMe nucleotides or comprises a region where every third nucleotide is 2’OMe modified. DM2\21326271.1160PATENT Docket No. J4040-99018
[0276] Non-limiting examples of chemical modifications of the splinting nucleic acid and the nucleic acid sequences corresponding to these modifications are provided in Table 12, Table 28 and illustrated in Fig.17.
[0277] In some embodiments, the chemical modification of the splint can be 3' C3 spacer. In some embodiments, the chemical modification of the splint can be 3' phosphorylation. In some embodiments, the chemical modification of the splint can be 3' phosphorylation. In some embodiments, the chemical modification of the splint can be 3' inverted T overhang. In some embodiments, the chemical modification of the splint can be 5' inverted T overhang. In some embodiments, the chemical modification of the splint can be 3' inverted T part of GBS.
[0278] In some embodiments, editing efficiency may be optimized. In some embodiments, nucleic acid duplex lengths and compositions may be adjusted. In some embodiments, a splint binding site (SBS) of a guide nucleic acid is optimized. In some embodiments, a guide binding site (GBS) of a splint is optimized. In some embodiments, a splint binding site of a donor nucleic acid is optimized. In some embodiments, the donor binding site (DBS) of a splint nucleic acid is optimized. In some embodiments, it is preferred to avoid certain modifications of nucleic acids to be integrated at the target. In some embodiments, it is preferred to modify duplex forming nucleic acids that will not be integrated at the target. In some embodiments, it is preferred to modify a splinting nucleic acid that will not be integrated at the target. In some embodiments, it is preferred to modify a duplex-forming region of a nucleic acid that will not be integrated, such as duplex- forming region of a splint nucleic acid or guide nucleic acid.
[0279] In some embodiments, editing efficiency can be optimized by adjusting the length and composition of a duplex formed by a gRNA SBS and a splint GBS. In some embodiements, editing efficiency can be optimized by a adjusting the length and composition of a duplex formed by a donor SBS and a splint DBS. In some embodiements, editing efficiency can be optimized by a adjusting the length and composition of a duplex formed by a gRNA SBS and a splint GBS.
[0280] In some embodiments, optimizing editing efficiency comprises modifying a splint GBS. In some embodiments, the length of the splint GBS is adjusted. In some embodiments, the length of the splint GBS-gRNA duplex is from 10 bp to 30 bp, or from 12 bp to 28 bp, or from 14 bp to 26 bp, or from 16 bp to 24 bp, or from 18 bp to 22 bp, or from 17 bp to 19 bp, or from 18 bp to 20 bp, or from 19 bp to 21 bp, or 17 bp, or 18 bp, or 19 bp, or 20 bo, or 21 bp, or 22 bp, DM2\21326271.1161PATENT Docket No. J4040-99018 or 23 bp. In some embodiments, the splint GBS comprises LNAs and the number or proportion of LNAs is adjusted. In some embodiments, the LNA nucleotides alternate with DNA nucleotides.
[0281] In some embodiments, optimizing editing efficiency comprises modifying a splint DBS. In some embodiments, the length of the splint DBS is adjusted. In some embodiments, the length of the splint GBS-donor duplex is from 18 bp to 40 bp, or from 19 bp to 36 bp, or from 20 bp to 32 bp, or from 21 bp to 30 bp, or from 22 bp to 28 bp, or from 23 bp to 27 bp, or from 24 bp to 26 bp, or from 18 bp to 22 bp, or from 20 bp to 24 bp, or from 22 bp to 28 bp, or from 24 bp to 30 bp, or from 26 bp to 32 bp, or from 28 bp to 34 bp, or 18 bp, or 19 bp, or 20 bo, or 21 bp, or 22 bp, or 23 bp, or 24 bp, or 25 bp, or 26 bo, or 27 bp, or 28 bp, or 29 bp, or 30 bp, or 31 bo, or 32 bp, or 33 bp, or 34 bp, or 35 bp, or 36 bp, or 37 bp, or 38 bp, or 39 bp, or 40 bp. In some embodiments, the splint DBS comprises LNAs and the number or proportion of LNAs is adjusted. In some embodiments, the LNA nucleotides alternate with DNA nucleotides.
[0282] In some embodiments, optimizing editing efficiency comprises modifying a splint FBS. In some embodiments, the length of the splint FBS is adjusted. In some embodiments, the length of the splint FBS-flap duplex is from 10 bp to 30 bp, or from 12 bp to 28 bp, or from 14 bp to 26 bp, or from 16 bp to 24 bp, or from 18 bp to 22 bp, or from 17 bp to 19 bp, or from 18 bp to 20 bp, or from 19 bp to 21 bp, or 17 bp, or 18 bp, or 19 bp, or 20 bp, or 21 bp, or 22 bp, or 23 bp. In some embodiments, the splint DBS comprises LNAs and the number or proportion of LNAs is adjusted. In some embodiments, the LNA nucleotides alternate with DNA nucleotides.
[0283] Also provided herein in are figures and tables showing a variety of splint nucleic acids, guide nucleic acids and donor nucleic acids. It will be apparent from the sequences and descriptions of DBS, FPS, and GBS sites of splints herein and sequences and descriptions of splint biding sites of donor nucleic acids herein, and sequences and descriptions of SBS sites of guides herein, which donors, splints, and guides are suitable to be used together and how to make suitable donors, splints, and guides.
[0284] The following Table 12 provides non-limiting examples of splints employed herein. Table 12. Non-limiting examples of splint nucleic acid sequences * *DM2\21326271.1162PATENT Docket No. J4040-99018 s003 +C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C *+T (SEQ ID NO:561) * C * * C Q + A G C A A + + A C A CDM2\21326271.1163PATENT Docket No. J4040-99018 s024 +T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+G+GCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA +G*C*+T (SEQ ID NO:582) A F F * * A A C * C + + C G + CDM2\21326271.1164PATENT Docket No. J4040-99018 s046 +C*A*+GT+AA+CG+GC+AG+AC+TT+CT+CT+AC+AGGAGTCAGGTGCACGCCA+CA+AT+AC+CG+C A+G*C*T / 3SpC3 / (SEQ ID NO:644) A A G + C G + G C G G A A G T T +GC+DM2\21326271.1165PATENT Docket No. J4040-99018 s067+A*G*+GC+CA+GC+AG+TG+AA+CA+AC+CA+TT+GGGCGTGGCAGTACGCCA+CA+AT+AC+CG+CA+G*C*+T (SEQ ID NO:540) A+CC ++AG+ C + CG+CG +CTDM2\21326271.1166PATENT Docket No. J4040-99018 s087+G*G*+AG+AC+CG+CC+GT+CG+TC+GA+CA+AG+CCTCTGGCCTGCAGATCATGC+AG+CC+CG+GA+AC+C*A*+C (SEQ ID NO:193) TCTG CGCGC C T
[0285] The L-PGI system operates through a distinct mechanism involving direct ligation. Upon recruitment to the target site by the lmgRNA, the nuclease (e.g., nCas9) introduces a nick on the non-target strand, creating a 3' genomic flap. The splint nucleic acid, bound to both the lmgRNA (via GBS-SBS interaction) and the donor DNA (via DBS hybridization), then hybridizes its FBS domain to this genomic flap. This positions the 5' phosphate end of the donor DNA adjacent to the 3' hydroxyl end of the genomic flap, templated by the splint. The DNA ligase (delivered either fused to, co-localized with, or separately from the nuclease) catalyzes the formation of a phosphodiester bond, ligating the donor DNA directly onto the genomic flap. For point mutations or small edits, the ligated donor strand, containing the desired edit, may then hybridize to the complementary target strand, creating a heteroduplex. This heteroduplex can be DM2\21326271.1167PATENT Docket No. J4040-99018 resolved by the cell's endogenous mismatch repair (MMR) pathway, potentially leading to permanent incorporation of the edit.
[0286] For larger deletions or replacements, a paired L-PGI (pL-PGI) approach may be employed. This involves using two separate L-PGI complexes targeting opposite strands flanking the region of interest. Each complex generates a nick and ligates its respective donor DNA onto the corresponding genomic flap. If the donors are designed for deletion, they contain homology arms that allow the newly ligated strands to hybridize to the genomic sequence outside the opposing nick site after the intervening genomic segment is excised (e.g., by endogenous exonucleases). If designed for replacement, the donors contain complementary overlapping sequences corresponding to the desired insertion; after ligation and excision of the intervening genomic segment, these donors can anneal to each other, resulting in the replacement of the original sequence with the sequence encoded by the donors.
[0287] Notably, these L-PGI and pL-PGI mechanisms primarily rely on the activity of the delivered nuclease and ligase, and the hybridization properties of the provided oligonucleotides. They do not necessitate homology-directed repair (HDR) pathways and may have reduced reliance on endogenous kinases or ligases compared to other editing strategies, potentially contributing to their efficacy in non-dividing or slowly dividing cells.
[0288] L-PGI System Optimization
[0289] In some aspects, the efficiency and fidelity of L-PGI and pL-PGI systems may be enhanced through optimization of various components and parameters. Key parameters include the lengths of the functional domains within the splint nucleic acid (DBS, FBS, GBS) and the donor nucleic acid, as well as the length of overlap between donors in pL-PGI systems. Chemical modifications play a significant role; incorporating LNAs or 2'OMe modifications in specific patterns within the splint DBS and GBS can tune affinity and stability, while 3' end protection (e.g., C3 spacer, PS bonds) on donors and / or splints can increase resistance to degradation. Methylation of donor DNA bases (e.g., 5-methylcytosine) may also enhance stability or enable epigenetic editing. The ratio of delivered components, such as the molar ratio of DNA oligonucleotides (splint and donor) to lmgRNA, can impact efficiency. Furthermore, the use of additional guide RNAs, such as nicking guides (ngRNAs) targeting the opposite strand to promote MMR resolution or dead guides to potentially enhance chromatin accessibility, may improve outcomes for specific targets. Enzyme architecture, including the choice of ligase (e.g., T4 ligase) DM2\21326271.1168PATENT Docket No. J4040-99018 and whether the nuclease and ligase are delivered as a fusion protein or as separate components (potentially co-localized via heterodimerization domains like leucine zippers), also represents an avenue for optimization. Optimization of these parameters may be performed empirically for specific target sites, cell types, or desired edits to maximize performance.
[0290] Advantages and Applications of L-PGI Systems
[0291] In some aspects, the L-PGI and pL-PGI systems offer several potential advantages for genome editing. Notably, these systems have demonstrated high fidelity, generating significantly fewer unintended insertions or deletions (indels) at the target site compared to some other editing methods like prime editing or DSB-based approaches. This high precision is particularly valuable for therapeutic applications. Furthermore, L-PGI systems have shown robust efficiency in various cell types, including primary cells and potentially non-dividing or slowly dividing cells such as hepatocytes, hematopoietic stem cells, and induced pluripotent stem cells, where RT-based or HDR-dependent methods may be less efficient. The L-PGI approach is versatile, capable of mediating point mutations, insertions, and deletions. The pL-PGI variant is particularly suited for larger edits, such as the precise deletion of genomic segments or the insertion / replacement of sequences like transcription factor binding sites or site-specific recombination sites (e.g., attB or attP sites). The ability to efficiently place recombination sites enables a two-step strategy for integrating very large DNA cargos (tens of kilobases) using site-specific integrases like Bxb1. L- PGI systems have demonstrated activity across different species (e.g., human, monkey, mouse) and are compatible with clinically relevant delivery modalities such as lipid nanoparticles (LNPs), facilitating potential in vivo therapeutic applications. Potential applications include, but are not limited to, correcting point mutations associated with genetic diseases (e.g., in HBB, CFTR, ATP7B, HFE), inserting therapeutic genes or regulatory elements, potentially via landing sites integrated using pL-PGI, installing or modifying transcription factor binding sites to modulate gene expression, and deleting pathogenic repeat expansions. Target nucleic acids
[0292] Disclosed herein are target nucleic acids. The target nucleic acid may include DNA. The target nucleic acid may be DNA. The target nucleic acid may include RNA. The target nucleic acid may be in a cell. The target nucleic acid may be methylated. The target nucleic acid may be unmethylated. The target nucleic acid may comprise a genome. The target nucleic acid may DM2\21326271.1169PATENT Docket No. J4040-99018 comprise genomic DNA. The target nucleic acid may comprise a chromosome. The target nucleic acid may comprise a gene.
[0293] The target nucleic acid may be in a subject. The target nucleic acid may be in a cell. The target nucleic acid may be in a test tube.
[0294] The target nucleic acid may be edited. The target nucleic acid may be edited in vitro. The target nucleic acid may be edited in vivo. Integrating nucleic acids Exogenous First Integrating Nucleic Acids
[0295] Disclosed herein are exogenous first integrating nucleic acids. The exogenous first integrating nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate to a nucleic acid that encodes an exogenous first integrating nucleic acid. Provided herein are exogenous first integrating nucleic acids that are inserted into a target nucleic acid such as a host genome at a genetic locus. For example, the exogenous first integrating nucleic acid may replace a nucleic acid in the target nucleic acid. The exogenous first integrating nucleic acid may be referred to as a “donor nucleic acid,” “donor” or “donor strand.” Where a genomic locus is described, a genetic locus may be included, or vice versa. For example, the locus may be part of a host genome or may be a part of a non-genome nucleic acid. The donor may include DNA. Likewise, the target nucleic acid may include DNA. In some cases, the donor may include RNA, for example when a target nucleic acid includes RNA. The donor may include any insert, such as a gene or a regulatory element, to be inserted at a genomic locus of a target nucleic acid. The donor strand may include a sequence that is at least partially homologous to the genomic locus. The donor may, in some instances, also act as a splint for a DNA ligase described herein, such as for ligating two nucleic acid strands base paired to a portion of the splinting exogenous first integrating nucleic acid. In some cases, the splint includes one strand of the donor, and the portion being ligated may be another strand of the donor. In some cases, the splint includes a strand of the donor, and the portion being ligated may be an upstream or downstream portion of the same strand of the donor. The donor may be single stranded. The donor may be double stranded. The donor may be delivered as two strands. The donor may be delivered as multiple strands, e.g., 2 strands.
[0296] The exogenous first integrating nucleic acid may be non-naturally occurring. The donor may be engineered. The donor may be synthetic. The donor may be pre-synthetized. The donor DM2\21326271.1170PATENT Docket No. J4040-99018 may be added to a subject or a cell. In some aspects, the donor does not include a template for a polymerase.
[0297] Disclosed herein are exogenous first integrating nucleic acids, comprising: a double- stranded DNA region to be inserted into a target nucleic acid, wherein the double-stranded DNA region is flanked by at least one overhang comprising a flap binding site and / or guide binding site.
[0298] The exogenous first integrating nucleic acid may be ligated into a target nucleic acid such as a genomic strand. The exogenous first integrating nucleic acid may include a 5’ end that may be ligated to a 3’ terminus of a genomic strand generated by an RNA-guided endonuclease.
[0299] The donor may include any aspect included in Fig.1A-6C. For example, the donor may include an aspect such as a guide binding site, a flap binding site, or an overhang. The donor may include a guide binding site. The donor may include 2 guide binding sites. The donor may include a flap binding site. The donor may include 2 flap binding sites. The donor may include an overhang. The donor may include 2 overhangs. The aspects may be included at a 5’ end or a 3’ end of the donor, or at both ends. A guide binding site or a flap binding site may be in an internal region of the donor.
[0300] Some aspects include an exogenous first integrating nucleic acid, comprising: a double- stranded DNA region to be inserted into a target nucleic acid, wherein the double-stranded DNA region is flanked by at least one overhang comprising a flap binding site or guide binding site.
[0301] In some aspects, the exogenous first integrating nucleic acid comprises a modified internucleoside linkage. In some aspects, the modified internucleoside linkage comprises a phosphorothioate linkage. In some aspects, the modified internucleoside linkage is between any of the 4 terminal nucleosides at a 5’ end or at a 3’ end of the donor nucleic acid. The exogenous first integrating nucleic acid may include multiple modified internucleoside linkages. For example, the exogenous first integrating nucleic acid may include modified internucleoside linkages at nucleic acids of the 5’ and 3’ ends of the exogenous first integrating nucleic acid, such as between the last 4 nucleic acids at the 5’ end and between the last 4 nucleic acids at the 3’ end. In some aspects, the exogenous first integrating nucleic acid comprises a modified nucleoside. In some aspects, the modified nucleoside comprises a locked nucleic acid (LNA), a 2’ fluoro, a 2’ O-alkyl, a 5’ O- methyl, a 2’-O-methyl, or a combination thereof. The modified nucleoside may include an LNA, a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof. The modified nucleoside may include an LNA. The modified nucleoside may include a 2’fluoro. DM2\21326271.1171PATENT Docket No. J4040-99018 The modified nucleoside may include a 2’ O-alkyl. The modified nucleoside may include a methylated cytosine. In some aspects, the modified nucleoside is any of the 3 terminal nucleosides at a 5’ end or at a 3’ end of the exogenous first integrating nucleic acid. The exogenous first integrating nucleic acid may include multiple modified nucleosides. For example, the exogenous first integrating nucleic acid may include modified nucleosides at nucleic acids of the 5’ and 3’ ends of the exogenous first integrating nucleic acid, such as the last 3 nucleic acids at the 5’ end and the last 3 nucleic acids at the 3’ end. The exogenous first integrating nucleic acid may include any modification such as a modified nucleoside or modified internucleoside linkage described in relation to guide nucleic acids, insofar as it does not interfere with the function of the exogenous first integrating nucleic acid after it is ligated into a target nucleic acid such as a host genome. The exogenous first integrating nucleic acid may include any number or combination of modifications such as a number or combination described in relation to guide nucleic acids, insofar as it does not interfere with a function of the exogenous first integrating nucleic acid. Table 11 and Table 13 include some examples of exogenous first integrating nucleic acid sequences.
[0302] In some apsects, for L-PGI applications, the donor nucleic acid (or integrating nucleic acid) is typically a pre-synthesized DNA oligonucleotide. Its length may be optimized depending on the desired edit; for example, lengths of about 22-26 nucleotides may be effective for introducing point mutations or small changes, while longer donors are used for insertions or replacements. A 5' phosphate group is generally included to enable ligation. The donor nucleic acid may incorporate chemical modifications to enhance stability or function. For instance, the 3' end may be protected against exonuclease degradation using modifications such as a C3 spacer or phosphorothioate linkages. Furthermore, bases within the donor, such as cytosines, may be methylated (e.g., 5-methylcytosine). Such methylation may enhance stability through improved hybridization with modified splints (e.g., containing LNAs) or may be used to introduce specific epigenetic marks into the target locus. For paired L-PGI (pL-PGI) applications involving two donor nucleic acids (one for each strand break), the donors may be designed with complementary overlapping regions to facilitate replacement edits, or with sequences homologous to the regions flanking the intended deletion site for deletion edits. The length of the overlap in replacement donors may be optimized, with overlaps of about 10 base pairs sometimes proving effective, for example, when inserting attB recombination sites. DM2\21326271.1172PATENT Docket No. J4040-99018
[0303] The donor nucleic acid may include a methylated nucleotide. The donor nucleic acid may include an unmethylated nucleotide. An example of a methylated nucleotide may include a nucleotide including methylated cytosine. The cytosine may be methylated at a C-5 position of the cytosine ring. An example of an unmethylated nucleotide may include an unmethylated cytosine. The unmethylated nucleotide may include a cytosine that is not methylated at a C-5 position of the cytosine ring.
[0304] In some embodiments, the donor nucleic acid can comprise a modified nucleotide. In some embodiments, the donor nucleic acid can comprise a modified DNA base. In some embodiments, the modified nucleotide comprises methylation. In some embodiments, the donor nucleic acid comprises a methylated nucleoside. Non-limiting example of the methylated nucleoside can include methylated cytosine (5-mC). In some embodiments, the donor nucleic acid can comprise N6-methyladenine (6-mA). In some embodiments, the modified DNA base can comprise 5-hydroxymethylcytosine (5-hmC). In some embodiments, the modified DNA base can comprise 5-formylcytosine (5-fC). In some embodiments, the modified DNA base can comprise 5-carboxylcytosine (5-caC).
[0305] In certain embodiments, the donor nucleic acid comprises one or more modifications to the 3' end. For example, a 3’ C3 spacer, which is a short 3 carbon chain (C3) attached to the terminal 3’ hydroxyl group of the oligonucleotide, can be used. Such modifications may, for example, block extension by a polymerase or exonuclease degradation.
[0306] In some embodiments, the donor can introduce an epigenetic modification into the target nucleic acid. In some embodiments, the donor can introduce an epigenetic modification of a DNA base (e.g., a methylated nucleoside) into the target nucleic acid. In some embodiments, the donor can introduce a methylated DNA base into the target nucleic acid (e.g., a methylated nucleoside). In some embodiments, the donor can introduce an unmethylated DNA base into the target nucleic acid (e.g., a methylated nucleoside). In some embodiments, the integrating of the donor can introduce a methoylated nucleoside into the genome. In some embodiments, the integrating of the donor can remove a methylated nucleoside from the genome. In some embodiments, the donor can introduce a methylated cytosine (e.g. 5-mC) into the target nucleic acid. In some embodiments, the donor can introduce an N6-methyladenine (6-mA) into the target nucleic acid. In some embodiments, the donor can introduce a 5-hydroxymethylcytosine (5-hmC) into the target nucleic acid. In some embodiments, the donor can introduce a 5-formylcytosine (5- DM2\21326271.1173PATENT Docket No. J4040-99018 fC) into the target nucleic acid. In some embodiments, the donor can introduce a 5- carboxylcytosine (5-caC) into the target nucleic acid. Recombination Sequences
[0307] Disclosed herein are exogenous first integrating nucleic acids. In some aspects, the exogenous first integrating nucleic acid may comprise a recombination sequence. A recombination sequence may also be referred to as an “integrating recombination sequence,” an “integration sequence,” a “recombination site,” an “integration site,” a “site-specific recombination sequence,” or a “site-specific integration sequence.”
[0308] In some aspects, the exogenous first integrating nucleic acid may comprise a plurality of recombination sequences. In some aspects, the exogenous first integrating nucleic acid may comprise at least 1 recombination sequence. In some aspects, the donor nucleic acid may comprise at least 2 recombination sequences. In some aspects, the donor nucleic acid may comprise at least 3 recombination sequences. In some aspects, the donor nucleic acid may comprise at least 4 recombination sequences. In some aspects, the donor nucleic acid may comprise at least 5 recombination sequences. In some aspects, the donor nucleic acid may comprise at least 10 recombination sequences.
[0309] In some aspects, the exogenous first integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by an integrase. In some aspects, the exogenous first nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a recombinase. In some aspects, the exogenous first nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by any integrase described herein. In some aspects, the exogenous first nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by any integrase described in Table 8. The donor nucleic acid described herein may comprise a nucleic acid sequence that is recognized or bound by an integrase that comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8.
[0310] In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a serine recombinase. The serine recombinase may be or comprise a PhiC31 (ΦC31) bacteriophage integrase, a Bxb1 mycobacteriophage integrase, a Pseudomonas aeruginosa (Pa01) integrase, a Nocardia otitidiscaviarum (No67) integrase, or a DM2\21326271.1174PATENT Docket No. J4040-99018 Streptomyces ipomoeae (Si74) integrase. In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a ΦC31 bacteriophage integrase. In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Bxb1 mycobacteriophage integrase. The donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Pseudomonas aeruginosa (Pa01) integrase. The donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Nocardia otitidiscaviarum (No67) integrase. The donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Streptomyces ipomoeae (Si74) integrase.
[0311] In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a tyrosine recombinase. In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a Cre recombinase. In some aspects, the donor nucleic acid described herein may comprise a nucleotide sequence recognized or bound by a flippase (Flp).
[0312] In some aspects, the exogenous first integrating nucleic acid described herein comprises any one of the sequences in Table 13. In some aspects, the donor nucleic acid described herein comprises a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the nucleotide sequence of any of the recombination sequences in Table 13. Table 13. Non-limiting examples of large serine recombinase (LSR) recombination Q :DM2\21326271.1175PATENT Docket No. J4040-99018 Si74 TAGTGACGTCTGTCCGCGCAG 327 TTCCGACGCAGTTTCCGACGA 328 TGATCGAGGGAGTGTGTGCTT GTACGAGGACGAGGACAGACDM2\21326271.1176PATENT Docket No. J4040-99018 AAAACCCTCTCCAGTTGTTTTT CTCCTGCGCCGCCATCCCTGG TCTTGCCCC GTCTTTTCCGDM2\21326271.1177PATENT Docket No. J4040-99018 CTTCGTTTCGCCGAAGCGTAC CGCAGGCCTTCGCCGAAGCGC AAGATGGAAATCATAAACCTT TGGAAGTGTACGATCTCGCGC CAAATGCGCCATTCGATCTTG GCACGCAGGAACTTGATGACGDM2\21326271.1178PATENT Docket No. J4040-99018 CATAGTATACGAAAGGACGC GCTAAAAAACGAAAGGACGG ATACCATAATCAATGGGAACA CATGCCATGAATATATTCGAT TTTGAATTTTCGTAAAAAAAA CACTATCGCCAGCGCTACGAADM2\21326271.1179PATENT Docket No. J4040-99018 ATCTGAATGTTACACTGCCAC CTGATGGGGATGCAAGTGGGC CTTTTCGACGAAGGTGTGTGG GGCAAGCGCAAACTGCTGGTG TTTC CCGGDM2\21326271.1180PATENT Docket No. J4040-99018 ACTGAATGTGCTTCCAGTACA AAATCCCTTTGTTATCAAATC GGTTGTTACACCGTTAGGCAA CCTGCCCTTATGATGGATGGA AAAAATAAATATCCATAG ATCACGCTGGTCATTTCCCDM2\21326271.1181PATENT Docket No. J4040-99018 GTGTCGGTGGTTCAAATCCGC AGGCGGATTTGAACCACCGAC CTCCGGGCACCATTAACTTGC ACGCGGATTTTCAATCCGCTG TGAAAATGCTAAGTTTTTGGA CTCTACCAACTGAGCTATCCGDM2\21326271.1182PATENT Docket No. J4040-99018 Efs2 AGTTTAATACTAAAAGAGGTA 437 TGAAAAGCTATTTTATACAAC 438 TATTAATTTTATTTAAAAATTG GGGGGCATAGCTCAGTTGGTTDM2\21326271.1183PATENT Docket No. J4040-99018 Vp82 CGTACATGGCTCCTAGTGTGT 453 TGTCTTGTAAGCACCCATCCC 454 ACGATGGGAAATAACCAAAA TGTAGTTATGCACTGTACATTDM2\21326271.1184PATENT Docket No. J4040-99018 Sss2INT ATGGAGCCGTTCTCCGCGGAC 497 GGCCGCGAGGTCGTGTTCGTC 498 GTCATGGACTACGGCGTGATC GTCATGTTGAGGTTCACGACCDM2\21326271.1185PATENT Docket No. J4040-99018 SluINT TTTTTGTATGTTAGTTGTGTCA 525 TGGGTGGTACAGGTGCCACAT 526 CTGGGTAGACCTAAATAGTGA TAGTTGTACCATTTATG eincomprises any one of the sequences in Table 14. In some aspects, the donor nucleic acid described herein comprises a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the nucleotide sequence of any of the recombination sequences in Table 14. Table 14. D ri ti n F r rd S n (5’-3’) SEQ IDR S 5’ 3’ SE ID:DM2\21326271.1186PATENT Docket No. J4040-99018 Bxb1_AttP_T GTGGTTTGTCTGGTCAACCAC 686 TGGGTTTGTACCGTACACCACTGAGGAC 757C_site CGCGtcCTCAGTGGTGTACGG GCGGTGGTTGACCAGACAAACCACTACAAACCCADM2\21326271.1187PATENT Docket No. J4040-99018 Bxb1_AttB_4 GGCCGGCTTGTCGACGACGGC 702 CCGGATGATCCTGACGACGGAGGCCGCC 7736_GC_site GgcCTCCGTCGTCAGGATCAT GTCGTCGACAAGCCGGCCCCGGDM2\21326271.1188PATENT Docket No. J4040-99018 Bxb1_AttB_3 GGCTTGTCGACGACGGCGtcC 719 ATGATCCTGACGACGGAGGACGCCGTCG 7918_TC_site TCCGTCGTCAGGATCAT TCGACAAGCCDM2\21326271.1189PATENT Docket No. J4040-99018 87_Int30_Att_28DM2\21326271.1190PATENT Docket No. J4040-99018 N680429_5 CATTATATGTTTTTACAATCC 738 cattatatgttcttacagtatggcggcccggattgtaaa 80460_31_50bp GGGCCGCCATACTGTAAGAAC aacatataatgATATAATG
[0314] In some aspects, the exogenous first integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a serine recombinase. In some aspects, the donor nucleic acid may comprise an attachment site (att). In some aspects, the donor nucleic acid may comprise a plurality of attachment sites. In some aspects, the donor nucleic acid may comprise an attachment site on the bacterial part (attB) integrating recombination sequence. In some aspects, the donor nucleic acid may comprise two attachment site on the bacterial part (attB) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise a plurality of attachment site on the bacterial part (attB) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise an attachment site on the phage part (attP) integrating recombination sequence. In some aspects, the donor nucleic acid may comprise two attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the DM2\21326271.1191PATENT Docket No. J4040-99018 donor nucleic acid may comprise a plurality of attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise an attachment site on the bacterial part (attB) integrating recombination sequence and an attachment site on the phage part (attP) integrating recombination sequences. In some aspects, the donor nucleic acid may comprise a plurality of attachment site on the bacterial part (attB) integrating recombination sequences and a plurality of attachment site on the phage part (attP) integrating recombination sequences.
[0315] In some aspects, the exogenous first integrating nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a tyrosine recombinase. In some aspects, the donor nucleic acid may comprise a locus of X(cross)-over in P1 (LoxP) sequence. In some aspects, the donor nucleic acid may comprise two locus of X(cross)-over in P1 (LoxP) sequences. In some aspects, the donor nucleic acid may comprise a plurality of locus of X(cross)-over in P1 (LoxP) sequences. In some aspects, the donor nucleic acid may comprise a flippase recognition target (FRT) sequence. In some aspects, the donor nucleic acid may comprise two flippase recognition target (FRT) sequences. In some aspects, the donor nucleic acid may comprise a plurality of flippase recognition target (FRT) sequences.
[0316] In some embodiments, the donor nucleic acid comprises a chemical modification. In some embodiments, the donor nucleic acid can comprise a 3' C3 spacer. In certain embodiments, the donor nucleic acid comprises one or more modifications to the 3' end. For example, a 3’ C3 spacer, which is a short 3 carbon chain (C3) attached to the terminal 3’ hydroxyl group of the oligonucleotide, can be used. Such modifications may, for example, block extension by a polymerase or exonuclease degradation. In some embodiments, the donor nucleic acid can comprise a 3' inverted dT. In some embodiments, the donor nucleic acid can comprise a 3' phosphorylation. In some embodiments, the donor nucleic acid can comprise a 3' phosphorothioate bond. In some embodiments, a chemical modification on the donor nucleic acid can protect the donor DNA against nucleases. In some embodiments, the chemical modification is not integrated into the genome. In some embodiments, the chemical modification can be integrated into the genome. Second Integrating Nucleic Acids
[0317] Disclosed herein are second integrating nucleic acids. The second integrating nucleic acid may be included in a composition, system, or method disclosed herein. Some aspects relate DM2\21326271.1192PATENT Docket No. J4040-99018 to a nucleic acid that encodes a second integrating nucleic acid. Provided herein are second integrating nucleic acids that are inserted into a target nucleic acid such as a host genome at a genetic locus such as an integrating recombination sequence. For example, the second integrating nucleic acid may replace in whole or in part an integrating recombination sequence in the target nucleic acid. The second integrating nucleic acid may be referred to as a “revising nucleic acid.” Where a genomic locus is described, a genetic locus may be included, or vice versa. For example, the locus may be part of a host genome or may be a part of a non-genome nucleic acid. The revising nucleic acid may include DNA. Likewise, the target nucleic acid may include DNA. In some cases, the revising nucleic acid may include RNA, for example when a target nucleic acid includes RNA. The revising nucleic acid may include any insert, such as a gene or a regulatory element, to be integrated at a genomic locus of a target nucleic acid such as a recombination sequence. The revising nucleic acid may be integrated, in whole or in part, into a genomic locus of a target nucleic acid such as a recombination sequence by an integrase. The revising nucleic acid may include a sequence that is at least partially homologous to the genomic locus. The revising nucleic acid may be single stranded. The revising nucleic acid may be double stranded. The revising nucleic acid may be delivered as two strands. The revising nucleic acid may be delivered as multiple strands, e.g., 2 strands.
[0318] The second integrating nucleic acid may be non-naturally occurring. The revising nucleic acid may be engineered. The revising nucleic acid may be synthetic. The revising nucleic acid may be pre-synthetized. The revising nucleic acid may be added to a subject or a cell. In some aspects, the revising nucleic acid does not include a template for a polymerase.
[0319] In some aspects, the revising nucleic acid comprises a recombination sequence. The recombination sequence may be a cognate of any recombination sequence of the donor nucleic acid. In some aspects, the cognate recombination sequence of the revising nucleic acid couples with a recombination sequence of the donor nucleic acid.
[0320] In some aspects, the revising nucleic acid comprises a plurality of recombination sequences. In some aspects, the re...
Claims
PATENT Docket No. J4040-99018 WHAT IS CLAIMED:
1. A system in a cell, comprising: a) an endonuclease; b) a DNA ligase; c) a first guide nucleic acid comprising: i. a spacer complementary to a first region of a first genomic strand of a genomic locus, ii. a scaffold for complexing with the endonuclease, and iii. a splint binding site that is at least partially complementary to a first splint nucleic acid; d) a first donor nucleic acid comprising an integrating recombination sequence, wherein the first donor nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid; e) a first splint nucleic acid comprising a donor binding site (DBS) that is at least partially complementary to the first donor nucleic acid, a flap binding site (FBS) that is at least partially complementary to a first genomic flap at or adjacent to the genomic locus, and a guide binding site (GBS) that is at least partially complementary to the first guide nucleic acid; f) a second guide nucleic acid comprising: i. a spacer complementary to a second region of a second strand of the genomic locus, ii. a scaffold for complexing with the endonuclease, and ii. a splint binding site that is at least partially complementary to a second splint nucleic acid; g) a second donor nucleic acid comprising a second integrating recombination sequence complementary to the first integrating recombination sequence, wherein the second donor nucleic acid is configured to be ligated by the ligase to a second strand break generated by the endonuclease in the target nucleic acid; h) a second splint nucleic acid comprising a donor binding site (DBS) that is at least partially complementary to the second donor nucleic acid, a flap binding site (FBS) that is at least partially complementary to a second genomic flap at or adjacent to the genomic locus, 433PATENT Docket No. J4040-99018 and a guide binding site (GBS) that is at least partially complementary to the second guide nucleic acid; wherein the first donor nucleic acid and the second donor nucleic acid are configured to overlap to form a duplex.
2. The system of claim 1, wherein the first splint and the second splint each comprises a donor binding site (DBS) from 22-26 nucleotides in length, a flap binding site (FBS) from 11-13 nucleotides, and a guide binding site (GBS) from 19-20 nucleotides.
3. The system of any of claims 1 or 2, wherein the first donor nucleic acid and the second donor nucleic acid form a 10 bp overlap.
4. A system in a cell, comprising: a) an endonuclease; b) a DNA ligase; c) a guide nucleic acid comprising: i. a spacer complementary to a region of a genomic locus of a genomic strand, ii. a scaffold for complexing with the endonuclease, and iii. a splint binding site that is at least partially complementary to a splint nucleic acid; d) a donor nucleic acid comprising an integrating recombination sequence, wherein the donor nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid; and e) a splint nucleic acid comprising a donor binding site (DBS) that is at least partially complementary to the donor nucleic acid, a flap binding site (FBS) that is at least partially complementary to a genomic flap at or adjacent to the genomic locus, a guide binding site (GBS) that is at least partially complementary to the guide nucleic acid.
5. The system of claim 4, wherein the splint comprises a donor binding site (DBS) from 22-26 nucleotides in length, a flap binding site (FBS) from 11-13 nucleotides, and a guide binding site (GBS) from 19-20 nucleotides.
6. The system of any one of claim 1 to 5, wherein the GBS comprises a modified nucleotide. 434PATENT Docket No. J4040-99018 7. The system of any of claim 1 - 6, wherein the GBS comprises a region wherein alternating nucleotides or every third nucleotide comprises a locked nucleic acid (LNA).
8. The system of any of claims 1 - 7, wherein the DBS comprises a modified nucleotide.
9. The system of any of claims 1 - 8, wherein the DBS comprises a region wherein alternating nucleotides or every third nucleotide comprises a locked nucleic acid (LNA).
10. The system of any of claims 1 - 9, wherein the splint comprises modified nucleotides in the configuration of a splint nucleic acid of Table 63.
11. The system of any claims 1 - 10, wherein the splint nucleic acid comprises a Guide Binding Site (GBS) and a Donor Binding Site (DBS) as set forth for a splint selected from spl001, spl002, spl003, spl004, spl005, spl006, spl007, spl008, spl009, spl010, spl011, spl012, spl013, spl014, spl015, spl016, spl017, spl018, spl019, spl020, spl021, spl022, spl023, spl049, spl050, spl051, spl052, spl053, spl054, spl055, spl056, spl057, spl058, spl059, spl060, spl061, spl062, spl063, spl064, spl065, spl066, spl067, spl068, spl069, spl070, spl073, spl074, spl075, spl076, spl077, spl078, spl089, spl090, spl091, spl092, spl093, spl094, spl095, spl096, spl097, spl098, spl099, spl100, spl101, spl102, spl103, spl104, spl105, spl106, spl107, spl108, spl109, spl110, and spl111 in Table 63.
12. The system of any claims 1 - 11, wherein the endonuclease comprises an RNA- guided endonuclease.
13. The system of claim 12, wherein the endonuclease comprises a class II CRISPR / Cas endonuclease.
14. The system of claim 12, wherein the endonuclease comprises a Cas9 endonuclease.
15. The system of claim 12, wherein the endonuclease comprises a nickase.
16. The system of claim 15, wherein the endonuclease comprises a Cas9 nickase.
17. The system of any of claims 1 - 16, further comprising: an integrase; and an integrating nucleic acid comprising: 435PATENT Docket No. J4040-99018 at least one cognate recombination sequence; and an exogenous nucleic acid, wherein the at least one cognate recombination sequence of the integrating nucleic acid couples with the at least one integrating recombination sequence of the donor nucleic acid, and wherein the integrase integrates the integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence.
18. The system of claim 12, wherein the integrating nucleic acid comprises at least one attachment site (att) cognate recombination sequence.
19. The system of claim 18, wherein the integrating nucleic acid comprises an attachment site on the bacterial part (attB) cognate recombination sequence.
20. The system of claim 18, wherein the integrating nucleic acid comprises two attachment sites on the bacterial part (attB) cognate recombination sequences.
21. The system of claim 18, wherein the integrating nucleic acid comprises an attachment site on the phage part (attP) cognate recombination sequence.
22. The system of claim 18, wherein the integrating nucleic acid comprises two attachment sites on the phage part (attP) cognate recombination sequences.
23. The system of claim 18, wherein the integrating nucleic acid comprises one attachment site on the bacterial part (attB) cognate recombination sequence and one attachment site on the phage part (attP) cognate recombination sequence.
24. The system of claim 12, wherein the integrating nucleic acid comprises at least one locus of X(cross)-over in P1 (LoxP) cognate recombination sequence.
25. The system of claim 12, wherein the donor nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence.
26. The system of any one of claims 12-25, wherein the second integrating nucleic acid comprises a regulatory sequence.
27. The system of claim 26, wherein the regulatory sequence is a promoter.
28. The system of anyone of claims 12-27, wherein the integrase is coupled to the endonuclease or the ligase. 436PATENT Docket No. J4040-99018 29. The system of any one of claims 12-28, wherein the integrase is a serine integrase.
30. The system of claim 29, wherein the serine integrase comprises a PhiC31 bacteriophage integrase, a Bxb1 mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (Pa01), a Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74).
31. The system of any one of claims 29 or 30, wherein the integrase is coupled to a recombination directionality factor (RDF).
32. The system of any one of claims 12-28 wherein the integrase is a tyrosine integrase.
33. The system of claim 32, wherein the tyrosine integrase is a Cre recombinase.
34. The system of claim 32, wherein the tyrosine integrase is a flippase (Flp).
35. The system of any one of claims 1-34, wherein the donor nucleic acid or the integrating nucleic acid comprises a modified nucleotide.
36. The system of claim 35, wherein the modified nucleotide comprises a methylated nucleotide.
37. The system of claim 35, wherein the modified nucleotide comprises methylated cytosine (e.g.5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5- carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
38. The system of claim 37, wherein the modified nucleotide comprises methylated cytosine 5-mC.
39. The system of claim 35, wherein the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
40. The system of any one of claims 1-39, wherein the splint comprises a streptavidin or biotin operatively coupled to the splint.
41. The system of claim 40, wherein the splint comprising the streptavidin is operatively coupled with the DNA ligase, said DNA ligase operatively coupled to a biotin, or 437PATENT Docket No. J4040-99018 wherein the splint comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase operatively coupled to a streptavidin.
42. The system of any one of claims 1-41, wherein the endonuclease comprises a fusion partner.
43. The system of claim 42, wherein the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone H1 central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof.
44. A method of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, the method comprising introducing into the host cell: a) an endonuclease; b) a DNA ligase; c) a first guide nucleic acid comprising: i. a spacer complementary to a first region of a first genomic strand of a genomic locus, ii. a scaffold for complexing with the endonuclease, and iii. a splint binding site that is at least partially complementary to a first splint nucleic acid; d) a first donor nucleic acid comprising an integrating recombination sequence, wherein the first donor nucleic acid is configured to be ligated by the ligase to a strand break generated by the endonuclease in a target nucleic acid; e) a first splint nucleic acid comprising a donor binding site (DBS) that is at least partially complementary to the first donor nucleic acid, a flap binding site (FBS) that is at least partially complementary to a first genomic flap at or adjacent to the genomic locus, and a guide binding site (GBS) that is at least partially complementary to the first guide nucleic acid; f) a second guide nucleic acid comprising: i. a spacer complementary to a second region of a second strand of the genomic locus, ii. a scaffold for complexing with the endonuclease, and ii. a splint binding site that is at least partially complementary to a second splint nucleic acid; 438PATENT Docket No. J4040-99018 g) a second donor nucleic acid comprising a second integrating recombination sequence complementary to the first integrating recombination sequence, wherein the second donor nucleic acid is configured to be ligated by the ligase to a second strand break generated by the endonuclease in the target nucleic acid; h) a second splint nucleic acid comprising a donor binding site (DBS) that is at least partially complementary to the second donor nucleic acid, a flap binding site (FBS) that is at least partially complementary to a second genomic flap at or adjacent to the genomic locus, and a guide binding site (GBS) that is at least partially complementary to the second guide nucleic acid; wherein the first donor nucleic acid and the second donor nucleic acid are configured to overlap to form a duplex.
45. The method of claim 44, wherein the first splint and the second splint each comprises a donor binding site (DBS) from 22-26 nucleotides in length, a flap binding site (FBS) from 11-13 nucleotides, and a guide binding site (GBS) from 19-20 nucleotides.
46. The method of claim 44 or 45, wherein the first guide nucleic acid, the first donor nucleic acid, and the first splint nucleic acid are delivered in a first LNP and the second guide nucleic acid, the second donor nucleic acid, and the second splint are delivered in a second LNP.
47. The method of claim 44 or 45, wherein the first guide nucleic acid, the first donor nucleic acid, and the first splint nucleic acid the second guide nucleic acid, the second donor nucleic acid, and the second splint are delivered together in the same LNP.
48. The method of any of claims 44 to 47, wherein the exogenous integrating nucleic acid is delivered by a viral vector.
49. A method of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, the method comprising introducing into the host cell a system of nucleic acids comprising: a) a guide nucleic acid comprising: i. a spacer complementary to a region of a genomic locus of a genomic strand, ii. a scaffold for complexing with an endonuclease, 439PATENT Docket No. J4040-99018 iii. an optional donor binding site that is at least partially complementary to a donor nucleic acid, and iv. a flap binding site that is at least partially identical or complementary to a genomic flap at or adjacent to the genomic locus; and b) a donor nucleic acid comprising at least one nucleic acid sequence encoding at least one integrating recombination sequence, wherein the donor nucleic acid further comprises a 5’ end to be ligated to a 3’ terminus of the genomic strand generated by the endonuclease.
50. A method of integrating an exogenous nucleic acid into a target nucleic acid in a host cell, the method comprising introducing into the host cell a system of nucleic acids comprising: a) a guide nucleic acid comprising: i. a spacer complementary to a region of a genomic locus of a genomic strand, ii. a scaffold for complexing with the endonuclease, and ii. a splint binding site that is at least partially complementary to a splint nucleic acid; b) a donor nucleic acid comprising an integrating recombination sequence, wherein the donor nucleic acid is configured to be ligated by a ligase to a strand break generated by a endonuclease in a target nucleic acid; and c) a splint nucleic acid comprising a donor binding site (DBS) that is at least partially complementary to the donor nucleic acid, a flap binding site (FBS) that is at least partially complementary to a genomic flap at or adjacent to the genomic locus, a guide binding site (GBS) that is at least partially complementary to the guide nucleic acid.
51. The method of claim 49 or 50, wherein the splinting nucleic acid further comprises a donor binding site (DBS) that is at least partially identical or complementary to a portion of the donor nucleic acid.
52. The method of claim 49 or 50, wherein the GBS comprises a modified nucleotide.
53. The method of claim 52, wherein the GBS comprises a region wherein alternating nucleotides or every third nucleotide comprises an LNA.
54. The method of claim 51, wherein the DBS comprises a modified nucleotide. 440PATENT Docket No. J4040-99018 55. The method of claim 54, wherein the DBS comprises a region wherein alternating nucleotides or every third nucleotide comprises an LNA.
56. The method of any one of claims 49 or 50, wherein the splinting nucleic acid comprises a binding site having modified nucleotides in the configuration of a splinting nucleic acid of Table 12.
57. The method of any one of claims 49 or 50, wherein the guide nucleic acid comprises a sequence of linking nucleic acids between the scaffold and the donor binding site.
58. The method of any one of claims 49 or 50, wherein the guide nucleic acid comprises MS2 binding loops within the scaffold.
59. The method of claim 50, wherein the guide nucleic acid comprises MS2 binding loops between the scaffold and the donor binding site.
60. The method of any one of claims 49 or 50, wherein the guide nucleic acid, the donor nucleic acid, or the splinting nucleic acid comprises a modified internucleoside linkage.
61. The method of claim 60, wherein the modified internucleoside linkage comprises a phosphorothioate linkage.
62. The method of claim 61, wherein the modified internucleoside linkage comprises a phosphonoacetate linkage.
63. The method of claim 49 or 50, wherein the guide nucleic acid, the donor nucleic acid, or the splinting nucleic acid comprises a modified nucleoside.
64. The method of claim 63, wherein the modified nucleoside comprises a locked nucleic acid (LNA), a 2’fluoro, a 2’ O-alkyl, a methylated cytosine, an inverted thymidine, or a combination thereof.
65. The method of claim 49 or 50, wherein the endonuclease comprises an RNA- guided endonuclease.
66. The method of claim 65, wherein the endonuclease comprises a class II CRISPR / Cas endonuclease. 441PATENT Docket No. J4040-99018 67. The method of claim 66, wherein the endonuclease comprises a Cas9 endonuclease.
68. The method of claim 67, wherein the endonuclease comprises a nickase.
69. The method of claim 68, wherein the endonuclease comprises a Cas9 nickase.
70. The method of any one of claims 49-69, wherein the method further comprises an integrating nucleic acid, wherein the integrating nucleic acid comprises: at least one nucleic acid sequence encoding a cognate recombination sequence; and wherein the at least one cognate recombination sequence of the integrating nucleic acid couples with the at least one integrating recombination sequence of the donor nucleic acid, and wherein an integrase integrates the integrating nucleic acid, in whole or in part, into the target nucleic acid at the integrating recombination sequence.
71. The method of claim 70, wherein the integrating nucleic acid comprises at least one attachment site (att) cognate recombination sequence.
72. The method of claim 71, wherein the integrating nucleic acid comprises an attachment site on the bacterial part (attB) cognate recombination sequence.
73. The method of claim 71, wherein the integrating nucleic acid comprises two attachment sites on the bacterial part (attB) cognate recombination sequences.
74. The method of claim 71, wherein the second integrating nucleic acid comprises an attachment site on the phage part (attP) cognate recombination sequence.
75. The method of claim 71, wherein the integrating nucleic acid comprises two attachment sites on the phage part (attP) cognate recombination sequences.
76. The method of claim 71, wherein the integrating nucleic acid comprises one attachment site on the bacterial part (attB) integrating recombination sequence and one attachment site on the phage part (attP) integrating recombination sequence.
77. The method of claim 70, wherein the integrating nucleic acid comprises at least one locus of X(cross)-over in P1 (LoxP) cognate recombination sequence.
78. The method of claim 70, wherein the donor nucleic acid comprises at least one flippase recognition target (FRT) integrating recombination sequence. 442PATENT Docket No. J4040-99018 79. The method of any one of claims 70-78, wherein the integrating nucleic acid comprises a regulatory sequence.
80. The method of claim 79, wherein the regulatory sequence is a promoter.
81. The method of any one of claims 70-80, wherein the integrase is a serine integrase.
82. The method of claim 81, wherein the serine integrase is a PhiC31 bacteriophage integrase, a Bxb1 mycobacteriophage integrase, a Pseudomonas aeruginosa integrase (Pa01), a Nocardia otitidiscaviarum integrase (No67), or a Streptomyces ipomoeae integrase (Si74).
83. The method of any one of claims 81 or 82, wherein the integrase is coupled to a recombination directionality factor (RDF).
84. The method of any one of claims 70-80, wherein the integrase is a tyrosine integrase.
85. The method of claim 84, wherein the tyrosine integrase is a Cre recombinase.
86. The method of claim 84, wherein the tyrosine integrase is a flippase (Flp).
87. The method of any one of claims 43-86, wherein the donor nucleic acid or the integrating nucleic acid comprises a modified nucleotide.
88. The method of claim 87, wherein the modified nucleotide comprises a methylated nucleotide.
89. The method of claim 88, wherein the methylated nucleotide comprises methylated cytosine (e.g.5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5- carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
90. The method of claim 89, wherein the methylated nucleotide comprises methylated cytosine 5-mC.
91. The method of claim 87, wherein the modified nucleotide comprises 5’ Inverted Dideoxy-T, 3' phosphorylation, 3' C3 spacer, 3' inverted dT, or a combination thereof.
92. The method of any one of claims 49-91, wherein the splinting nucleic acid comprises a modification. 443PATENT Docket No. J4040-99018 93. The method of claim 92, wherein the splinting nucleic acid comprises a sequence of Table 12.
94. The method of claim 92, wherein the modification comprises a streptavidin operatively coupled to the splinting nucleic acid.
95. The method of claim 94, wherein the splinting nucleic acid comprising the streptavidin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a biotin.
96. The method of claim 92, wherein the modification comprises a biotin operatively coupled to the splinting nucleic acid.
97. The method of claim 96, wherein the splinting nucleic acid comprising the biotin is operatively coupled with the DNA ligase, said DNA ligase is operatively coupled to a streptavidin.
98. The method of any one of claims 49-97, wherein the endonuclease comprises a fusion partner.
99. The method of claim 98, wherein the fusion partner comprises: a streptavidin; a Rad51 DNA repair protein (rad51DBD) or fragment thereof; a high-mobility group nucleosome binding domain 1 (HN1) or fragment thereof; a histone H1 central globular domain (H1G) or fragment thereof; Brex27 or fragment thereof; or a combination thereof. 444
Citation Information
Patent Citations
Systems, methods, and compositions for site-specific genetic engineering using programmable addition via site-specific targeting elements (PASTE)
US20220145293A1
Direct replacement genome editing
US20230151353A1
Methods and compositions for prime editing nucleotide sequences
US20230340467A1
Methods and compositions for directed genome editing
WO2021188840A1
Revision of genetic material using direct replacement editing
WO2024233949A1