Modification of genetic material using direct substitution editing
Fusion proteins with DNA-binding proteins and ligases enable precise nucleic acid modification by introducing targeted breaks and integrating exogenous nucleic acids, addressing the limitations of existing genome-editing methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- トランジット セラピューティクス インコーポレイテッド
- Filing Date
- 2024-05-10
- Publication Date
- 2026-05-26
AI Technical Summary
Existing genome-editing methods are inadequate for efficient substitution or modification of nucleic acid sequences in the genome.
The use of fusion proteins comprising a DNA-binding protein, such as an RNA-guided endonuclease, coupled with a DNA ligase, to introduce targeted strand breaks and facilitate the ligation of exogenous nucleic acids into a target nucleic acid, utilizing systems and methods that include intracellular components like ligases, endonucleases, and integrases to modify nucleic acids through ligation and integration.
This approach enables precise and efficient modification of nucleic acid sequences by introducing targeted strand breaks and integrating exogenous nucleic acids, enhancing the capability to replace or modify genomic regions with high specificity and accuracy.
Smart Images

Figure 2026516883000001_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 465,799, filed May 11, 2023, and U.S. Provisional Application No. 63 / 637,313, filed April 22, 2024, each of which is hereby incorporated by reference in its entirety.
[0002] The foregoing applications, and all documents cited in those applications or during their examination (the "application - cited documents"), and all documents cited or referenced in the application - cited documents, and all documents cited or referenced in this specification (the "specification - cited documents"), and all documents cited or referenced in this specification, together with the manufacturer's instructions, descriptions, product specifications, and product sheets for any products mentioned in this specification or for any documents incorporated by reference into this specification, are hereby incorporated by reference and may be used in the practice of the invention. More specifically, all reference documents are incorporated by reference as if each individual document was specifically and individually indicated to be incorporated by reference.
[0003] The present invention provides compositions, systems, and methods for modifying nucleic acids. The modification can be achieved using a fusion protein comprising a nuclease, ligase, integrase, or a combination thereof. Nucleic acid modification can include ligation of a donor nucleic acid to a target nucleic acid. Nucleic acid editing can include replacing a portion of a target nucleic acid with a modified nucleic acid.
[0004] Sequence Listing This application includes a sequence listing submitted via Patent Center, which is hereby incorporated by reference in its entirety. The.xml copy created on May 10, 2024, is named Y9506 - 99016 and is 1,437,114 bytes in size.
Background Art
[0005] Improved genome - editing methods are necessary for the substitution or modification of nucleic acid sequences in the genome.
[0006] Any reference or specification of any document in this application does not constitute an admission that such document is available as prior art for the present invention. [Overview of the project] [Means for solving the problem]
[0007] In some embodiments, intracellular systems comprising fusion proteins are disclosed herein. In some embodiments, the systems or compositions disclosed herein comprise a DNA-binding protein coupled to a DNA ligase. The DNA-binding protein may comprise an endonuclease. The endonuclease may comprise an RNA-guided endonuclease. In some embodiments, the coupling is covalent. In some embodiments, the fusion protein comprises a DNA-binding protein (e.g., an endonuclease such as an RNA-guided endonuclease) and a DNA ligase. In some embodiments, the cell comprises a cell containing a DNA-binding protein (e.g., an endonuclease such as an RNA-guided endonuclease) and a DNA ligase, both of which are heterogeneous to the cell. In some embodiments, the composition comprises a cell containing a DNA-binding protein (e.g., an endonuclease such as an RNA-guided endonuclease) that is heterogeneous to the cell and a DNA ligase that is endogenous to the cell. Some embodiments include a composition comprising a cell containing a DNA-binding protein (e.g., an endonuclease such as an RNA-guided endonuclease) and a DNA ligase, wherein the DNA ligase is endogenous to the cell. In some embodiments, the DNA-binding protein is amino(N)-terminated relative to the DNA ligase in the fusion protein. In some embodiments, the DNA-binding protein is carboxy(C)-terminated relative to the DNA ligase in the fusion protein. In some embodiments, the linkage includes a linker comprising 1 to 100 amino acids. In some embodiments, the coupling is non-covalent. In some embodiments, the composition comprises a first polypeptide comprising at least a portion of the DNA-binding protein and a second polypeptide comprising at least a portion of the DNA ligase, wherein the first and second polypeptides are non-covalently coupled. In some embodiments, the first polypeptide comprises a first heterodimerizing domain that binds to a second heterodimerizing domain, and the second polypeptide comprises a second heterodimerizing domain.In some embodiments, the heterodimer domain comprises a leucine zipper, a PDZ domain, streptavidin, a streptavidin-binding protein, a Foldon domain, a hydrophobic moiety, or a functional binding fragment thereof. In some embodiments, the first polypeptide comprises a first intein that binds to a second intein, and the second polypeptide comprises a second intein. In some embodiments, the ligase comprises a hairpin binding motif, and the DNA-binding protein and DNA ligase are coupled with a nucleic acid comprising a scaffold that binds to the DNA-binding protein and a hairpin that binds to the hairpin binding motif. In some embodiments, the hairpin binding motif comprises an MS2 coat protein (MCP) peptide, and the hairpin comprises an MS2 hairpin. In some embodiments, the DNA-binding protein and DNA ligase are coupled with a heterobifunctional molecule comprising an endonuclease-binding domain and a DNA ligase-binding domain. In some embodiments, the heterobifunctional molecule comprises a small molecule. In some embodiments, the DNA-binding protein comprises a class II CRISPR / Cas endonuclease. In some embodiments, the DNA-binding protein includes a Cas9 endonuclease. In some embodiments, the DNA-binding protein includes a nickase. In some embodiments, the DNA-binding protein includes an amino acid sequence that is at least 80% identical to one of the amino acid sequences of SEQ ID NOs: 1-13, or a functional fragment thereof. In some embodiments, the DNA ligase ligates a DNA strand that forms a base pair with a DNA sprint. In some embodiments, the DNA ligase ligates a DNA strand that forms a base pair with an RNA sprint. In some embodiments, the DNA ligase includes an amino acid sequence that is at least 80% identical to one of the amino acid sequences of SEQ ID NOs: 55-96, or a functional fragment thereof. In some embodiments, the DNA-binding protein or DNA ligase includes a nuclear localization signal, a chromatin modification domain, a cell-permeable peptide, or a tag polypeptide. In some embodiments, it includes a guide RNA and an embedded nucleic acid. In some embodiments, the strand of the embedded nucleic acid acts as a ligase sprint. In some embodiments, the guide nucleic acid acts as a ligase sprint.In some embodiments, the sprint includes a modification. In some embodiments, the modification includes streptavidin operably coupled to the sprint. In some embodiments, the streptavidin-containing sprint is operably coupled to a DNA ligase, and the DNA ligase is operably coupled to biotin. In some embodiments, the modification to the sprint includes biotin operably coupled to the sprint. In some embodiments, the biotin-containing sprint is operably coupled to a DNA ligase, and the DNA ligase is operably coupled to streptavidin. In some embodiments, the composition includes one or more nucleic acids encoding the composition. In some embodiments, the composition includes a cell containing one or more nucleic acids.
[0008] In some embodiments, methods for incorporating an exogenous nucleic acid into a target nucleic acid in a host cell are disclosed herein, comprising contacting the intracellular target nucleic acid with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a strand break at the predetermined locus of the target nucleic acid. In some embodiments, the strand break is a double-strand break.
[0009] In some embodiments, a method for incorporating an exogenous nucleic acid into a target nucleic acid in a host cell is disclosed herein, comprising: contacting the intracellular target nucleic acid with an endonuclease at a predetermined locus of the target nucleic acid to thereby introduce a strand break at the predetermined locus of the target nucleic acid; introducing a pre-synthesized incorporated nucleic acid into the cell; and ligating the 5' end of the pre-synthesized incorporated nucleic acid to the 3' end of the nick at the predetermined locus of the target nucleic acid. In some embodiments, the strand break is a double-strand break. In some embodiments, the strand break is a nick. In some embodiments, the endonuclease includes an RNA-guided endonuclease. In some embodiments, the endonuclease includes a class II CRISPR / Cas endonuclease. In some embodiments, the endonuclease includes a Cas9 endonuclease. In some embodiments, the endonuclease includes a nickase. In some embodiments, the endonuclease includes a Cas9 nickase. In some embodiments, the ligation involves contacting a predetermined locus of an endonuclease and a target nucleic acid with a guide nucleic acid. In some embodiments, the ligation is carried out by a ligase bound to the endonuclease. In some embodiments, the endonuclease includes a fusion partner. In some embodiments, the fusion partner includes streptavidin; Rad51 DNA repair protein (rad51DBD) or a fragment thereof; high-mobility group nucleosome-binding domain 1 (HN1) or a fragment thereof; histone H1 central globular domain (H1G) or a fragment thereof; Brex27 or a fragment thereof; or a combination thereof. In some embodiments, the pre-synthesized integrated nucleic acid includes mutations related to the target nucleic acid. In some embodiments, the nick includes a single phosphodiester chain break in the otherwise double-stranded target nucleic acid. In some embodiments, the nick includes a non-sticky, non-blunt end of the strand of the target nucleic acid. In some embodiments, the target nucleic acid includes a chromosome of a cell. In some embodiments, the cell is a eukaryote.
[0010] In some embodiments, a system is disclosed herein that is intracellular, comprising: a ligase (e.g., a heterologous ligase or an endogenous ligase); an endonuclease that introduces a nick to a predetermined locus of a target nucleic acid; and a pre-synthesized integrated nucleic acid having a 5' end ligated by the ligase to the 3' end of the nick at the predetermined locus of the target nucleic acid. In some embodiments, the endonuclease comprises a class II CRISPR / Cas endonuclease. In some embodiments, the endonuclease comprises a Cas9 nickase. In some embodiments, a guide nucleic acid is included to bring the endonuclease close to a predetermined locus of the target nucleic acid. In some embodiments, the ligase is coupled to the endonuclease. In some embodiments, the pre-synthesized integrated nucleic acid comprises a mutation related to the target nucleic acid. In some embodiments, the nick comprises a single phosphodiester chain break in an otherwise double-stranded target nucleic acid. In some embodiments, the nick comprises a non-sticky, non-blunt end of a strand of the target nucleic acid. In some embodiments, the target nucleic acid comprises a cellular chromosome. In some aspects, cells are eukaryotes.
[0011] In some embodiments, a system within a cell comprising a ligase, an endonuclease, and an embedded nucleic acid containing an embedded recombinant sequence is disclosed herein. In some embodiments, the embedded nucleic acid is configured to be ligated by the ligase to strand breaks generated by the endonuclease in the target nucleic acid. In some embodiments, the strand breaks are double-strand breaks. In some embodiments, the strand breaks are nicks.
[0012] In some embodiments, a system within a cell is disclosed herein, comprising a ligase; an endonuclease; an embedded nucleic acid containing an embedded recombinant sequence; an integrase; and a second embedded nucleic acid, comprising (i) at least one homogeneous recombinant sequence and (ii) an exogenous nucleic acid. In some embodiments, the embedded nucleic acid is configured to be ligated by the ligase to strand breaks generated by the endonuclease in the target nucleic acid. In some embodiments, the homogeneous recombinant sequence of the second embedded nucleic acid is coupled with at least one embedded recombinant sequence of the embedded nucleic acid. In some embodiments, the integrase integrates the second embedded nucleic acid, whole or partially, into the target nucleic acid at the embedded recombinant sequence.
[0013] In some embodiments, it is disclosed herein that a nucleic acid configured to be ligated by a ligase to a chain break generated by an endonuclease in a target nucleic acid is incorporated. In some embodiments, the incorporated nucleic acid includes at least one attachment site (att) for incorporating the recombinant sequence. In some embodiments, the incorporated nucleic acid includes an attachment site on a bacterial portion (attB) for incorporating the recombinant sequence. In some embodiments, the incorporated nucleic acid includes two attB-incorporating recombinant sequences. In some embodiments, the incorporated nucleic acid includes a binding site on a phage portion (attP) for incorporating the recombinant sequence. In some embodiments, the incorporated nucleic acid includes two attP-incorporating recombinant sequences. In some embodiments, the incorporated nucleic acid includes one attB-incorporating recombinant sequence and one attP-incorporating recombinant sequence. In some embodiments, the incorporated nucleic acid includes at least one X (crossover) locus in P1 (LoxP) for incorporating the recombinant sequence. In some embodiments, the incorporated nucleic acid includes at least one flippase-recognition target (FRT) for incorporating the recombinant sequence.
[0014] In some embodiments, a second embedding nucleic acid is disclosed herein, configured to be integrated whole or partially into a target nucleic acid with an integrated recombinant sequence by an integrase. In some embodiments, the second embedding nucleic acid includes at least one attachment site (att) for incorporating the recombinant sequence. In some embodiments, the second embedding nucleic acid includes a binding site on a bacterial portion (attB) incorporating the recombinant sequence. In some embodiments, the second embedding nucleic acid includes two attB embedding recombinant sequences. In some embodiments, the second embedding nucleic acid includes a binding site on a phage portion (attP) incorporating the recombinant sequence. In some embodiments, the second embedding nucleic acid includes two attP embedding recombinant sequences. In some embodiments, the second embedding nucleic acid includes an embedding nucleic acid comprising one attB embedding recombinant sequence and one attP embedding recombinant sequence. In some embodiments, the second embedding nucleic acid includes at least one X (crossover) locus in P1 (LoxP) incorporating the recombinant sequence. In some embodiments, the second embedded nucleic acid comprises at least one flippase recognition target (FRT) that incorporates the recombinant sequence. In some embodiments, the second embedded nucleic acid comprises a regulatory sequence. In some embodiments, the regulatory sequence is a promoter.
[0015] In some embodiments, a second embedded nucleic acid comprising an embedded nucleic acid or a modified nucleotide is disclosed herein. In some embodiments, the modified nucleotide comprises a methylated nucleotide. In some embodiments, the modified nucleotide comprises methylated cytosine (e.g., 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof. In some embodiments, the modified nucleotide comprises 5' reverse dideoxy-T, 3' phosphorylation, 3'C3 spacer, 3' reverse dT, or a combination thereof.
[0016] This specification discloses integrases for incorporating a second integrated nucleic acid into a target nucleic acid, either whole or partially, via an integrated recombinant sequence. In some embodiments, the integrase is coupled to an endonuclease. In some embodiments, the integrase is coupled to a ligase. In some embodiments, the integrase is coupled to a recombination direction factor (RDF). In some embodiments, the integrase is a serine integrase. The serine integrase may be PhiC31 bacteriophage integrase, Bxb1 mycobacteriophage integrase, Pseudomonas aeruginosa integrase (Pa01), Nocardia otitidiscaviarum integrase (No67), or Streptomyces ipomoeae integrase (Si74). In some embodiments, the integrase is a tyrosine integrase. The tyrosine integrase may be Cre recombinase or flippase (Flp).
[0017] In some embodiments, a method for incorporating an exogenous nucleic acid into a target nucleic acid in a host cell is disclosed herein, comprising a guide nucleic acid comprising (a) a spacer complementary to a genomic locus region of a genome strand, (b) a scaffold for forming a complex with a DNA-binding protein, (c) an optional donor-binding site at least partially complementary to the integration nucleic acid, and (d) a flap-binding site at or adjacent to a genomic locus that is at least partially identical or complementary to a genomic flap; and a first integration nucleic acid comprising (i) at least one nucleic acid sequence encoding at least one integration recombinant sequence, and (ii) a 5' end ligated to the 3' end of a genome strand produced by a DNA-binding protein. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some embodiments, a nucleic acid system is disclosed herein, comprising: a guide nucleic acid comprising (a) a spacer complementary to the region of a genomic locus of the genome strand, (b) a scaffold for forming a complex with a DNA-binding protein, and (c) an optional donor-binding site at least partially complementary to the sprint nucleic acid; an integrated nucleic acid comprising (i) at least one nucleic acid sequence encoding at least one integrated recombinant sequence, and (ii) a 5' end ligated to the 3' end of the genome strand produced by the DNA-binding protein; and a sprint nucleic acid comprising a flap-binding site at or adjacent to the genomic locus that is at least partially identical or complementary to the genomic flap, and an optional guide-binding site at least partially complementary to the guide nucleic acid. In some embodiments, the genome strand is intracellular. In some embodiments, the sprint nucleic acid further comprises a donor-binding site at least partially identical or complementary to a portion of the integrated nucleic acid. In some embodiments, the guide nucleic acid comprises a sequence of linked nucleic acid between the scaffold and the donor-binding site. In some embodiments, the guide nucleic acid includes an MS2 binding loop within the scaffold. In some embodiments, the guide nucleic acid includes an MS2 binding loop between the scaffold and the donor binding site.In some embodiments, the guide nucleic acid, the first integrated nucleic acid, or the sprint nucleic acid includes a modified nucleoside bond. In some embodiments, the modified nucleoside bond includes a phosphorothioate bond. In some embodiments, the modified nucleoside bond includes a phosphoacetate bond. In some embodiments, the first integrated nucleic acid, the guide nucleic acid, or the sprint nucleic acid includes a modified nucleoside. In some embodiments, the modified nucleoside bond is located between any of the four terminal nucleosides at the 5' or 3' end of the guide nucleic acid or the integrated nucleic acid. In some embodiments, the guide nucleic acid or the integrated nucleic acid includes a modified nucleoside. In some embodiments, the modified nucleoside includes a locked nucleic acid (LNA), a 2' fluoro, a 2' O-alkyl, or a combination thereof. In some embodiments, the modified nucleoside is any of the three terminal nucleosides at the 5' or 3' end of the guide nucleic acid or the integrated nucleic acid. The modified nucleoside may include LNA, 2'-fluoro, 2'O-alkyl, methylated cytosine, reverse thymidine, or a combination thereof. In some embodiments, the endonuclease includes an RNA-guided endonuclease. In some embodiments, the endonuclease includes a class II CRISPR / Cas endonuclease. In some embodiments, the endonuclease includes a Cas9 endonuclease. In some embodiments, the endonuclease includes a nickasase. In some embodiments, the endonuclease includes a Cas9 nickasase.
[0018] In some embodiments, the methods disclosed herein further include a second embedded nucleic acid, wherein the second embedded nucleic acid includes at least one nucleic acid sequence encoding a homogeneous recombinant sequence, wherein at least one homogeneous recombinant sequence of the second embedded nucleic acid is coupled with at least one embedded recombinant sequence of the first embedded nucleic acid, and an integrase incorporates all or part of the second embedded nucleic acid into the target nucleic acid at the embedded recombinant sequence. In some embodiments, the second embedded nucleic acid includes a binding site on a bacterial portion (attB) incorporating the recombinant sequence. In some embodiments, the second embedded nucleic acid includes two attB embedded recombinant sequences. In some embodiments, the second embedded nucleic acid includes a binding site on a phage portion (attP) incorporating the recombinant sequence. In some embodiments, the second embedded nucleic acid includes two attP embedded recombinant sequences. In some embodiments, the second embedded nucleic acid includes an embedded nucleic acid comprising one attB embedded recombinant sequence and one attP embedded recombinant sequence. In some embodiments, the second integrated nucleic acid includes at least one X (crossover) locus in P1(LoxP) that incorporates the recombinant sequence. In some embodiments, the second integrated nucleic acid includes at least one flippase recognition target (FRT) that incorporates the recombinant sequence. In some embodiments, the second integrated nucleic acid includes a regulatory sequence. In some embodiments, the regulatory sequence is a promoter. In some embodiments, the integrase is coupled to a recombination direction factor (RDF). In some embodiments, the integrase is a serine integrase. The serine integrase may be PhiC31 bacteriophage integrase, Bxb1 mycobacteriophage integrase, Pseudomonas aeruginosa integrase (Pa01), Nocardia otitidiscaviarum integrase (No67), or Streptomyces ipomoeae integrase (Si74). In some embodiments, the integrase is a tyrosine integrase. Tyrosine integrase may be Cre recombinase or flippase (Flp).In some embodiments, the first or second embedded nucleic acid comprises a modified nucleotide. In some embodiments, the modified nucleotide comprises a methylated nucleotide. In some embodiments, the modified nucleotide comprises methylated cytosine (e.g., 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof. In some embodiments, the modified nucleotide comprises 5' reverse dideoxy-T, 3' phosphorylation, 3'C3 spacer, 3' reverse dT, or a combination thereof. In some embodiments, the sprint nucleic acid comprises a modification. In some embodiments, the modification comprises streptavidin operably coupled to the sprint. In some embodiments, the streptavidin-containing sprint is operably coupled to a DNA ligase, and the DNA ligase is operably coupled to biotin. In some embodiments, the modification to the sprint includes biotin operably coupled to the sprint. In some embodiments, the biotin-containing sprint is operably coupled to a DNA ligase, the DNA ligase is operably coupled to streptavidin. In some embodiments, the endonuclease includes a fusion partner. In some embodiments, the fusion partner includes streptavidin; Rad51 DNA repair protein (rad51DBD) or a fragment thereof; high-mobility nucleosome-binding domain 1 (HN1) or a fragment thereof; histone H1 central globular domain (H1G) or a fragment thereof; Brex27 or a fragment thereof; or a combination thereof.
[0019] In some embodiments, fusion proteins comprising: a DNA-binding protein coupled to a DNA ligase; the DNA-binding protein may include an endonuclease; the endonuclease may include an RNA-guided endonuclease; in some embodiments, the linkage between the DNA-binding protein and the DNA ligase is covalent; in some embodiments, the fusion protein comprises a DNA-binding protein upstream of the DNA ligase; in some embodiments, the fusion protein comprises a DNA-binding protein downstream of the DNA ligase; in some embodiments, the linkage comprises a linker comprising 1 to 100 amino acids; in some embodiments, the composition comprises a first polypeptide comprising at least a portion of the DNA-binding protein and a second polypeptide comprising at least a portion of the DNA ligase, wherein the first and second polypeptides are linked together by covalent or noncovalent bonds; in some embodiments, the first polypeptide comprises a first heterodimerizing domain that binds to a second heterodimerizing domain, and the second polypeptide comprises a second heterodimerizing domain. In some embodiments, the heterodimer domain includes a leucine zipper, a PDZ domain, streptavidin, a streptavidin-binding protein, a Foldon domain, a hydrophobic moiety, or a functional binding fragment thereof. In some embodiments, the first polypeptide includes a first intein that binds to a second intein, and the second polypeptide includes a second intein. In some embodiments, the DNA-binding protein and DNA ligase are bound together by a small molecule. In some embodiments, the DNA-binding protein includes a class II CRISPR / Cas endonuclease. In some embodiments, the DNA-binding protein includes a Cas9 endonuclease. In some embodiments, the DNA-binding protein includes a nickase. In some embodiments, the DNA-binding protein includes an amino acid sequence that is at least 80% identical to any one of the amino acid sequences of SEQ ID NOs: 1-13, or a functional fragment thereof. In some embodiments, the DNA ligase ligates the DNA strand that forms base pairs with the DNA sprint. In some embodiments, DNA ligases ligate the DNA strand that forms a base pair with the RNA sprint.In some embodiments, the DNA ligase comprises an amino acid sequence that is at least 80% identical to any one of the amino acid sequences of SEQ ID NOs. 55-96, or a functional fragment thereof. In some embodiments, the DNA-binding protein or DNA ligase comprises a nuclear localization signal, a chromatin modification domain, a cell-permeable peptide, a tag polypeptide, or streptavidin. In some embodiments, it comprises guide RNA and an integrated nucleic acid. In some embodiments, it relates to a cell comprising the composition. In some embodiments, it comprises a nucleic acid encoding the composition. In some embodiments, it comprises one or more nucleic acids encoding a first or second polypeptide. In some embodiments, it comprises an editing method (e.g., nucleic acid) using the composition. In some embodiments, it comprises a therapeutic method using the composition. In some embodiments, it comprises administering the composition to a target.
[0020] In some embodiments, editing methods are disclosed herein that include ligating an embedded recombinant sequence into a nick in a target nucleic acid, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease.
[0021] In some embodiments, fusion proteins comprising: a DNA-binding protein fused to a DNA ligase; the DNA-binding protein may comprise an endonuclease; the endonuclease may comprise an RNA-guided endonuclease; in some embodiments, fusion proteins comprising an integrase, ligase, or endonuclease are disclosed herein; in some embodiments, the fusion protein disclosed herein comprises an integrase and a ligase; in some embodiments, the fusion protein disclosed herein comprises an integrase, a ligase, and an endonuclease; in some embodiments, the fusion protein disclosed herein comprises an integrase and an endonuclease; in some embodiments, a protein complex comprising: a DNA-binding protein bound to a DNA ligase; in some embodiments, the endonuclease and DNA ligase are bound together via a heterodimerizing domain. In some embodiments, the heterodimer domain comprises a leucine zipper, a PDZ domain, streptavidin, and a streptavidin-binding protein, a Foldon domain, a hydrophobic polypeptide, an antibody that binds to Cas nickase, or an antibody that binds to a DNA ligase, or one or more binding fragments thereof. In some embodiments, cells comprising a fusion protein or protein complex are disclosed herein. In some embodiments, cells comprising a heterologous DNA-binding protein and a DNA ligase introduced into the cell are disclosed herein. In some embodiments, a nuclease different from the DNA-binding protein is disclosed herein. In some embodiments, a guide nucleic acid comprising: a spacer at least partially inversely complementary to a first region of the target nucleic acid; a scaffold configured to bind to an endonuclease; a flap-binding site at least partially complementary to the nucleic acid flap, and an embedded nucleic acid-binding site. In some embodiments, an embedded nucleic acid comprising a single-stranded or double-stranded DNA region inserted into the target nucleic acid, wherein at least one additional single-stranded region comprising a guide-binding site is adjacent.In some embodiments, editing systems comprising a DNA-binding protein, a guide nucleic acid, and an embedded nucleic acid are disclosed herein. In some embodiments, editing methods comprising contacting a target nucleic acid with the editing system and a DNA ligase are disclosed herein.
[0022] In some embodiments, methods are disclosed herein that include ligating an integration sequence to a nick of a target nucleic acid in a cell, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease, and the integration sequence comprises a modified nucleotide. In some embodiments, the integration nucleic acid comprises a modified nucleotide. In some embodiments, the modified nucleotide comprises a methylated nucleotide. In some embodiments, the modified nucleotide comprises methylated cytosine (e.g., 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof. In some embodiments, the modified nucleotide comprises a 5' reverse dideoxy-T, 3' phosphorylation, 3'C3 spacer, 3' reverse dT, or a combination thereof.
[0023] In some embodiments, editing methods are disclosed herein, comprising ligating an embedded sequence to a nick in a target nucleic acid, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease, and the embedding of the embedded sequence introduces a methylated nucleoside into the genome.
[0024] In some embodiments, editing methods are disclosed herein, comprising ligating a ligated sequence to a nick in a target nucleic acid, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease, and the ligation of the ligated sequence removes a methylated nucleoside from the genome.
[0025] In some embodiments, a method is disclosed herein that includes (a) contacting a target nucleic acid within a cell with an endonuclease at a predetermined locus of the target nucleic acid, thereby introducing a nick at the predetermined locus of the target nucleic acid; (b) introducing a pre-synthesized methylated integration nucleic acid; and (c) ligating the ends of the nucleic acid during integration to the ends of the nick at the predetermined locus of the target nucleic acid.
[0026] In some embodiments, an editing system is disclosed herein that includes a ligase, an endonuclease that introduces a nick at a predetermined locus of a target nucleic acid, and a pre-synthesized methylated integration nucleic acid that includes ends that are ligated by the ligase to the ends of the nick at the predetermined locus of the target nucleic acid.
[0027] In some embodiments, a method is disclosed herein that includes contacting an endonuclease that includes a first heterodimerization moiety with a sprint nucleic acid that includes a second heterodimerization moiety. In some embodiments, the method further includes introducing a nick or a strand break at a predetermined locus of a target nucleic acid by the endonuclease and ligating an integration nucleic acid to the nick or the strand break, wherein the sprint nucleic acid binds to the integration nucleic acid and the target nucleic acid. In some embodiments, the first heterodimerization moiety includes biotin and the second heterodimerization moiety includes streptavidin or avidin, or the second heterodimerization moiety includes biotin and the first heterodimerization moiety includes streptavidin or avidin.
[0028] In some embodiments, a system is disclosed herein that includes an endonuclease bound to a first heterodimerization moiety and a split nucleic acid that includes a second heterodimerization moiety. In some embodiments, the system disclosed herein further includes a ligase. In some embodiments, the endonuclease introduces a nick or strand break at a predetermined locus of a target nucleic acid, the ligase ligates an integration nucleic acid to the nick or strand break, and the split nucleic acid binds to the integration nucleic acid and the target nucleic acid. In some embodiments, the first heterodimerization moiety includes biotin, the second heterodimerization moiety includes streptavidin or avidin, or the second heterodimerization moiety includes biotin and the first heterodimerization moiety includes streptavidin or avidin.
[0029] In some embodiments, at least one DNA binding protein and at least one guide nucleic acid, a spacer that is at least partially complementary to an intracellular genomic locus, a scaffold for forming a complex with at least one DNA binding protein, an optional donor binding site that is at least partially complementary to an integration nucleic acid, at least one DNA ligase, and a flap binding site that is at least partially reverse complementary to a nucleic acid flap, optionally including a guide binding site that is at least partially complementary to at least one guide nucleic acid, wherein at least one DNA binding protein cleaves or nicks at least one strand of the genomic locus, and at least one DNA ligase ligates the ends of the integration nucleic acid to the genomic flap site, thereby replacing the region of the genomic locus with an intracellular integration nucleic acid, an intracellular system comprising at least one guide nucleic acid and an integration nucleic acid is disclosed herein. The DNA binding protein can include an endonuclease. The endonuclease can include an RNA-guided endonuclease. In some embodiments, the integration nucleic acid includes single-stranded DNA. In some embodiments, the integration nucleic acid includes double-stranded DNA.
[0030] In some embodiments, an intracellular system comprising: a first guide nucleic acid comprising at least one DNA-binding protein including a first DNA-binding protein and an optional second DNA-binding protein; a first guide nucleic acid comprising at least one guide nucleic acid including a first guide nucleic acid comprising a first spacer complementary to a first region of a genomic locus in the cell; a first scaffold for forming a complex with the first DNA-binding protein; and an optional first donor binding site at least partially complementary to the integrated nucleic acid; and a first flap binding site at or adjacent to the genomic locus that is at least partially identical or complementary to the first genomic flap; and a second spacer complementary to a second region of a genomic locus in the cell; a second scaffold for forming a complex with the first or second DNA-binding protein; an optional second donor binding site at least partially complementary to the integrated nucleic acid; and at or adjacent to the genomic locus A second flap binding site that is at least partially identical or complementary to the second genome flap; at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; a second guide nucleic acid comprising at least one integrated nucleic acid comprising a first strand and a second strand, wherein the first strand comprises an optional first guide binding site that is at least partially complementary to the first guide nucleic acid; and the second strand comprises an optional second guide binding site that is at least partially complementary to the second guide nucleic acid. An intracellular system is disclosed herein in which a DNA-binding protein 1 and / or a second DNA-binding protein each cleave or cleave at least one strand of a genomic locus in a cell; a first DNA ligase ligates the end of the first strand of the embedded nucleic acid to a first genomic flap; and a first or second DNA ligase ligates the end of the second strand of the embedded nucleic acid to a second genomic flap, thereby replacing the region of the genomic locus with the intracellular embedded nucleic acid. In some embodiments, the embedded nucleic acid comprises a double-stranded DNA region. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease.In some embodiments, the embedded nucleic acid includes a 5' overhang that optionally includes a first guide binding site. In some embodiments, the embedded nucleic acid includes a 5' overhang that optionally includes a second guide binding site.
[0031] In some embodiments, intracellular systems are disclosed herein, comprising: at least one DNA-binding protein; at least one guide nucleic acid, a spacer complementary to an intracellular genomic locus; a scaffold for forming a complex with at least one DNA-binding protein; and an optional donor-binding site at least partially complementary to the integrated nucleic acid; at least one guide nucleic acid comprising at least one DNA ligase, and an optional guide-binding site at least partially complementary to the at least one guide nucleic acid; a flap-binding site at or adjacent to a genomic locus, at least partially identical or complementary to a genomic flap, wherein at least one DNA-binding protein cleaves or cuts at least one strand of the genomic locus; and an integrated nucleic acid, wherein at least one DNA ligase ligates the ends of the integrated nucleic acid to the genomic flap, thereby replacing the region of the genomic locus with the intracellular integrated nucleic acid. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some embodiments, the embedded nucleic acid comprises DNA including a 3' overhang. In some embodiments, the 3' overhang includes a guide binding site. In some embodiments, the 3' overhang includes a flap binding site. In some embodiments, at least one DNA ligase ligates the strands of the embedded nucleic acid to a genomic nucleic acid sequence.
[0032] In some embodiments, an intracellular system comprising at least one DNA-binding protein comprising a first DNA-binding protein and an optional second DNA-binding protein; a first guide nucleic acid and at least one guide nucleic acid comprising a second guide nucleic acid, wherein the first guide nucleic acid comprises a first spacer complementary to a first region of a genomic locus in the cell; a first scaffold for forming a complex with the first DNA-binding protein; and an optional first donor-binding site at least partially complementary to the integrated nucleic acid, and the second guide nucleic acid comprises a second spacer complementary to a second region of a genomic locus in the cell; a second scaffold for forming a complex with the first or second DNA-binding protein; and an optional second donor-binding site at least partially complementary to the integrated nucleic acid; and at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase, wherein the integrated nucleic acid comprises a first strand and a second strand, the first strand being at least partially complementary to the first guide nucleic acid The first strand includes any first guide binding site; the second strand includes any second guide binding site that is at least partially complementary to the second guide nucleic acid; the first strand includes a first flap binding site at or adjacent to a genomic locus that is at least partially identical or complementary to the first genomic flap; the second strand includes a second flap binding site at or adjacent to a genomic locus that is at least partially identical or complementary to the second genomic flap; the An intracellular system is disclosed herein in which a DNA-binding protein 1 and / or a second DNA-binding protein each cleave or cleave at least one strand of a genomic locus in a cell; a first DNA ligase ligates the end of the first strand of the embedded nucleic acid to a first genomic flap; and a first or second DNA ligase ligates the end of the second strand of the embedded nucleic acid to a second genomic flap, thereby replacing the region of the genomic locus with the intracellular embedded nucleic acid. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some embodiments, the embedded nucleic acid includes a double-stranded DNA region.In some embodiments, the double-stranded DNA optionally includes a first guide binding site and a 3' overhang including a first flap binding site. In some embodiments, the double-stranded DNA optionally includes a second guide binding site and a 3' overhang including a second flap binding site.
[0033] The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some embodiments, at least one DNA-binding protein includes a Cas protein or a functional fragment thereof. In some embodiments, the Cas protein or a functional fragment thereof includes nickase activity. In some embodiments, at least one DNA-binding protein includes Cas9 nickase or a functional fragment thereof. In some embodiments, at least one DNA ligase ligates nucleic acids bound to DNA. In some embodiments, at least one DNA ligase ligates nucleic acids bound to RNA. In some embodiments, at least one DNA ligase includes a PBCV-1 DNA ligase. In some embodiments, at least one DNA ligase is operably coupled to at least one DNA-binding protein. In some embodiments, at least one DNA ligase is fused to at least one DNA-binding protein as a fusion polypeptide. In some embodiments, at least one DNA-binding protein and at least one DNA ligase each include a heterodimer domain. In some embodiments, at least one DNA-binding protein and at least one DNA ligase form a heterodimer via a heterodimer domain. In some embodiments, at least one DNA-binding protein includes a linker. In some embodiments, the linker annexes a Cas protein or a functional fragment thereof to the heterodimer domain. In some embodiments, at least one DNA-binding protein includes a localization signal sequence. In some embodiments, at least one DNA ligase includes a localization signal sequence. In some embodiments, the localization signal sequence includes a nuclear localization sequence (NLS). In some embodiments, at least one DNA-binding protein or at least one DNA ligase is directed to the nucleus of a cell by the NLS. In some embodiments, at least one embedded nucleic acid, e.g., a donor nucleic acid or a modification nucleic acid, modifies at least one gene mutation at at least one genomic locus. In some embodiments, at least one embedded nucleic acid inserts a coding sequence.In some embodiments, the coding sequence encodes a full-length protein. In some embodiments, at least one embedded nucleic acid inserts a non-coding sequence. In some embodiments, the non-coding sequence includes a recombinant sequence. In some embodiments, the non-coding sequence knocks out an endogenous gene. In some embodiments, the non-coding sequence includes a regulatory element. In some embodiments, the nuclease further includes a nuclease. In some embodiments, the nuclease includes an exonuclease for digesting the genomic flap. In some embodiments, the nuclease includes human flap endonuclease 1 (hFEN1), human exonuclease 5 (hEXO5), T5 exonuclease, T7 exonuclease, exonuclease VIII, flap endonuclease domain of Escherichia coli (E. coli) PolI, RecJF, lambda exonuclease, Xni (ExoIXI), SaFEN (Staphylococcus aureus FEN), nuclease BAL-31, or fragments thereof. In some embodiments, the heterologous nuclease includes an endonuclease for digesting the genomic flap, and the endonuclease is distinct from at least one DNA-binding protein. In some embodiments, the at least one DNA-binding protein includes at least one further functional domain. In some embodiments, the at least one further functional domain includes a chromatin-modifying domain. In some embodiments, at least one additional functional domain comprises a cell-permeable peptide. In some embodiments, at least one guide nucleic acid comprises at least one nucleic acid modification. In some embodiments, at least one nucleic acid modification comprises a modification of the backbone, sugar, base, or a combination thereof. In some embodiments, at least one DNA-binding protein is complexed with at least one guide nucleic acid. In some embodiments, at least one guide nucleic acid is complexed with an embedded nucleic acid. In some embodiments, at least one DNA-binding protein, at least one guide nucleic acid, at least one DNA ligase, an embedded nucleic acid, or a combination thereof is encoded by a polynucleotide.In some embodiments, the polynucleotide comprises mRNA. In some embodiments, the polynucleotide comprises a vector. In some embodiments, the vector comprises a viral vector. In some embodiments, at least one DNA-binding protein, at least one guide nucleic acid, at least one DNA ligase, an embedded nucleic acid, or a combination thereof is encapsulated by at least one lipid nanoparticle. In some embodiments, the cell comprises a bacterial cell, a eukaryotic cell, or a plant cell. In some embodiments, the eukaryotic cell comprises a mammalian cell. In some embodiments, the composition comprises a system. In some embodiments, the cell comprises a system. In some embodiments, the cell comprises a cell line comprising cells. In some embodiments, the pharmaceutical composition comprises a system. In some embodiments, the pharmaceutical composition comprises a composition. In some embodiments, the pharmaceutical composition comprises cells. In some embodiments, the excipient, carrier, or diluent is pharmaceutically acceptable. In some embodiments, the pharmaceutical composition is formulated for administration to a target requiring it by intrathecal, intraocular, intravitreous, retinal, intravenous, intramuscular, intraventricular, intracerebral, intracerebral, intracerebellar, intracerebral, intraparenchymal, subcutaneous, intratumoral, intrapulmonary, intratracheal, intraperitoneal, intrabladder, vaginal, intrarectal, oral, sublingual, transdermal, inhalation, inhalation spray form, alumina-GI route, or a combination thereof. In some embodiments, the system includes a composition, or a kit comprising the pharmaceutical composition and a container. In some embodiments, the method of modifying cells includes contacting cells with the system. In some embodiments, the method of modifying cells includes contacting cells with the composition. In some embodiments, the method of modifying cells includes contacting cells with the pharmaceutical composition. In some embodiments, the cells are not dividing cells. In some embodiments, the integrated nucleic acid is inserted into the genomic locus of the cell independently of endogenous non-homologous end joining (NHEJ) and independently of endogenous homologous recombination repair (HDR).Some embodiments include a method for treating a disease or condition in a subject requiring treatment of the disease or condition, comprising: contacting a cell or subject with a system, composition, or pharmaceutical composition; and replacing a genomic locus within the cell with an embedded nucleic acid to treat the disease or condition in the subject. In some embodiments, the cell is not a dividing cell. In some embodiments, the embedded nucleic acid is inserted into the genomic locus of the cell independently of endogenous non-homologous end joining (NHEJ) and independently of endogenous homologous recombination repair (HDR).
[0034] In some embodiments, guide nucleic acids are disclosed herein, comprising: a spacer at least partially complementary to a genomic locus in a cell; a scaffold for forming a complex with a DNA-binding protein; and a donor-binding site at least partially complementary to the embedded nucleic acid. The DNA-binding protein may include an endonuclease. The endonuclease may include an RNA-guided endonuclease. In some embodiments, the guide nucleic acid comprises a flap-binding site at least partially complementary to the genomic sequence of the genomic locus. In some embodiments, the guide nucleic acid comprises at least one nucleic acid modification. In some embodiments, the at least one nucleic acid modification comprises a modification to the backbone, sugar, base, or a combination thereof. In some embodiments, the guide nucleic acid comprises an RNA sequence.
[0035] Therefore, the object of the present invention is not to include any previously known product, process for producing such product, or method for using such product, to the extent that the applicant retains rights and discloses herein the waiver of any previously known product, process, or method. Furthermore, the present invention is not intended to include within its scope any product, process, or method for using such product that does not meet the written description and validity requirements of the USPTO (35 U.S. SC § 112, first paragraph) or the EPO (EPC Section 83), and it should be noted that the applicant reserves rights and hereby discloses the waiver of any previously described product, process for producing such product, or method for using such product. In practicing the present invention, it may be advantageous to comply with EPC Section 53(c) and EPC Section 28(b) and (c). All rights expressly to waive any embodiments that are patented by the applicant in the lineage of this application or any other lineage or any previously filed application of any third party are expressly reserved. Nothing in this specification should be construed as a promise.
[0036] Please note that in this disclosure, particularly in the claims and / or paragraphs, terms such as “equipment,” “includes,” and “contains” may have meanings derived therefrom under U.S. patent law. For example, they may mean “include,” “contains,” and “contains.” Terms such as “essentially,” and “essentially,” may have meanings derived therefrom under U.S. patent law, for example, they enable elements that are not expressly enumerated but exclude elements found in the prior art or that affect the fundamental or novel features of the present invention.
[0037] These and other embodiments are disclosed in the following detailed description or are evident from the following detailed description and are included therein.
[0038] A patent or application file includes at least one drawing made in color. A copy of this patent or patent application publication, accompanied by one or more color drawings, will be provided by the Patent Office upon request and payment of the necessary fees.
[0039] The following detailed description, given as an example but not intended to limit the invention to the specific embodiments described, can be best understood in conjunction with the accompanying drawings. [Brief explanation of the drawing]
[0040] [Figure 1A] This figure shows the guide nucleic acid, endonuclease, ligase, and donor strand at a genomic locus. [Figure 1B] Figures 1A and 1A sequentially show donor strands incorporated into one side of a genomic locus, with the donor strands replacing the genomic flap. [Figure 1C] Figures 1B and 1B sequentially show the donor strand incorporated into one side of a genomic locus, and the nicks that appear where the genomic flap has been removed. [Figure 2A] This figure shows two guide nucleic acids, two endonucleases, two ligases, and a donor strand at a genomic locus. [Figure 2B] Following Figure 2A, we show the donor strands integrated into the genomic loci, which replace two genomic flaps. [Figure 2C] Figure 2B and subsequent figures show the donor strand incorporated into the genomic locus, and the two nicks that appear where the genomic flap was removed. [Figure 3A] This figure shows the guide nucleic acid, endonuclease, ligase, and donor strand at a genomic locus. [Figure 3B] Figures 3A and 3A sequentially show donor strands incorporated into one side of a genomic locus, with the donor strands replacing the genomic flap. [Figure 3C]Figures 3B and 3B sequentially show the donor strand incorporated into one side of a genomic locus, and the nicks that appear where the genomic flap has been removed. [Figure 4A] This figure shows two guide nucleic acids, two endonucleases, two ligases, and a donor strand at a genomic locus. [Figure 4B] Following Figure 4A, we show the donor strands integrated into the genomic loci, which replace two genomic flaps. [Figure 4C] Figure 4B and subsequent figures show the donor strand incorporated into the genomic locus, and the two nicks that appear where the genomic flap was removed. [Figure 5A] This figure shows the guide nucleic acid, endonuclease, ligase, and donor strand at a genomic locus. [Figure 5B] Following Figure 5A, we show the donor strands integrated into the genomic locus, where the donor strands replace the genomic flap. [Figure 5C] Figures 5B and 5B sequentially show the donor strand incorporated into one side of a genomic locus, and the nicks that appear where the genomic flap has been removed. [Figure 6A] This figure shows two guide nucleic acids, two endonucleases, two ligases, and a donor strand at a genomic locus. [Figure 6B] Following Figure 6A, we show the donor strands integrated into the genomic loci, which replace two genomic flaps. [Figure 6C] Figure 6B and subsequent figures show the donor strand incorporated into the genomic locus, and the two nicks that appear where the genomic flap was removed. [Figure 7] This figure shows some examples of fusion protein configurations. [Figure 8A] This figure shows an exemplary nicking and ligation pattern of the first exogenous incorporated nucleic acid. [Figure 8B]This figure shows a DNA gel exhibiting a pattern associated with unilateral replacer 2, performed in vitro using 30nt GBS / DBS and a heat-stable T4 ligase. Using a combination of 30nt GBS / DBS, a donor containing a protospacer adjacent motif (PAM) mutation, and a heat-stable T4 ligase (Hi-T4, NEB), we were able to produce a final replacer product (lane 3) corresponding to the size of our control product (lane 1). No replacer product was detected in the absence of nicking Cas9 (Cas9n) (lane 2) or in the absence of a lower donor functioning as a sprint (lanes 4 and 5). [Figure 8C] This figure shows an exemplary nucleic acid gel exhibiting patterns associated with variable-length GBS / DBS combinations and in vitro unilateral replacer 2 using T4 ligase. Using standard T4 ligase (NEB), final replacer products corresponding to the control size were constructed using multiple GBS / DBS combinations, including those without GBS / DBS, those containing 20nt GBS / DBS, and those containing 30nt GBS / DBS. Furthermore, in this experiment, recorded dsDNA donors containing PAM mutations were more efficient in the production of the final replacer product compared to dsDNA donors without recorded PAM mutations. [Figure 9] The results show the percentage of cells expressing green fluorescent protein (GFP), demonstrating gene editing from BFP to GFP using a one-sided replacer 2 containing Nicking Cas9 and DNA ligase. [Figure 10] This figure shows the sequencing reads merged and aligned to the target amplicon, as well as the percentage of all reads that matched the intended edit via a one-sided replacer 2 containing Nicking Cas9 and T4 DNA ligases. [Figure 11] This figure shows the sequencing reads merged and aligned to the target amplicon, as well as the percentage of all reads that matched the intended edit via a bilateral replacer 2 containing Nicking Cas9 and T4 DNA ligases. [Figure 12] The results show the measurement of the percentage of cells expressing green fluorescent protein (GFP), demonstrating gene editing from BFP to GFP via a one-sided replacer 2 containing Nicking Cas9 and T4 DNA ligase. [Figures 13A-13B] This figure shows examples of fusion proteins that may be useful in genome revision systems. In some embodiments, endonucleases, DNA ligases, and integrases may be delivered as three individual proteins (not shown). Endonucleases, DNA ligases, and integrases may be delivered as a combination of one double fusion protein and a third protein delivered in trans (Figure 13A). Endonucleases, DNA ligases, and integrases may be delivered as one triple fusion protein (Figure 13B). Linkers, such as peptide linkers, may be included between any two components in each fusion protein, or additional components may be added. [Figure 14A-14D] This figure shows a vector that can be used to deliver a second integrated nucleic acid (e.g., a modified nucleic acid) to a host cell. The vector may contain a single integrase binding site (Figures 14A and 14B) or multiple integrase binding sites (Figures 14C and 14D). The vector may be a minicircle (Figure 14A), a plasmid (Figures 14B and 14C), or linear DNA (Figure 14D). [Figure 15] This figure shows an exemplary nucleic acid gel with patterns associated with substituting either a 222 base pair (bp) sequence (lanes 1 and 2) or a 131 bp sequence (lanes 4 and 5) with a 38 bp attB recombinant sequence. Lanes 3 and 6 represent unsubstituted deletions of the 222 bp and 131 bp sequences, respectively. The leftmost lane represents the unedited sequence. [Figure 16] This figure shows amplicon sequencing data quantifying the percentage of reads in a pool of cells that contain the expected accurate edit of 222 bp or 131 bp sequences by attB recombinant sequences. The bars numbered 1-6 represent the same reactions as in Figure 15. [Figure 17] This figure shows the guide nucleic acid, endonuclease, ligase, and donor strand at a genomic locus. 3' chemical modifications (e.g., C3 spacer or reverse dT) are shown on the donor strand, and additional chemical modifications (e.g., 5' reverse dideoxy-T, 3' phosphorylation, 3' C3 spacer, or 3' reverse dT on the sprint) are shown on the sprint. [Figure 18] This provides sequencing data that quantifies the percentage of reads in a cell pool, demonstrating the editing efficiency of the ATPase copper transport beta (ATP7B) gene using a genome revision system that includes sprints with different terminal blocking modifications. [Figure 19] This study provides sequencing data that quantifies the percentage of reads in a cell pool, demonstrating the editing efficiency of cystic fibrosis membrane conductance regulator (CFTR) genes using a genome revision system that includes sprints with different terminal blocking modifications. [Figure 20] This figure shows the guide nucleic acid, endonuclease, ligase, and donor strand at a genomic locus. In this figure, the biotinylated sprint nucleic acid is bound to monomeric streptavidin fused to nickeling Cas9. Monomeric streptavidin can also be fused to a ligase (not shown). [Figure 21] This provides examples of fusion proteins that may be useful in genome revision systems. In some embodiments, the DNA-binding domain (rad51DBD) of the Rad51 DNA repair protein, and / or the high-mobility nucleosome-binding domain 1 (HN1) and the histone H1 central globular domain (H1G) are fused to the fusion protein. [Figure 22] This provides sequencing data that quantifies the percentage of reads in a cell pool, demonstrating the editing efficiency of the CFTR gene using a genome modification system that includes rad51DBD and / or a fusion protein containing HN1 and H1G. [Figure 23]We provide sequencing data that quantifies the percentage of reads in the cell pool, and demonstrate the editing efficiency of the CFTR gene using a genome modification system that includes bicistronic mRNA encoding the phosphorylation-mimicking peptide derived from IGF1 (IGF1pm1) and the N-terminal peptide derived from NFATC2IP (NFATC2IPp1) (IN peptide). [Figure 24] This figure shows a model of a G-to-T point mutation mediated by a one-sided replacer mechanism including donor nucleic acids and sprint nucleic acids. The figure illustrates the displacement of the 5' flap of the target DNA, ligation of the donor nucleic acid including a single nucleotide substitution to the 3' flap of the target DNA flap, and the resolution of the mismatch by the mismatch repair (MMR) patternway. [Figure 25] A model of ligase-mediated programmable gene integration (PGI) or L-PGI is provided. This model shows substitutions or deletions mediated by a 2× one-sided replacer. (See also Figures 1A and 3A.) The targeted modification (substitution or deletion) is determined by the location of the nick and the sequence of the duplicated flap formed by the donor nucleic acid. Targeted modifications are performed at single-base pair resolution. [Figure 26] The effect of single-strand overhang on the stability of the donor-sprint nucleic acid double helix at physiological temperatures is shown (left panel). The affinity of the complementary sequence can be modulated by the incorporation of locked nucleic acid (LNA) (right panel). In particular, when the donor nucleic acid rather than the sprint nucleic acid is incorporated into the target (see, for example, Figures 1A, 2A, 3A, 5A, and 6A), the sprint can be freely modified. [Figure 27] This figure shows a replacer programmed to mediate a triple mutation in order to mutate the target coding sequence to encode GFP and disrupt PAM. [Figure 28]This figure shows the optimization of the guide / sprint biregion formed by the guide's sprint binding site (SBS) and the sprint's (upper panel's) guide binding site (GBS). The system can be optimized by adjusting the GC content and the length of the SBS / GBS region (bottom left). In the HEK293T GFP example, the optimal length of the SBS / GBS region is approximately 19 bp (bottom right). In this case, there appears to be no substantial benefit from including linker nucleotides between the flap binding site (FBS) and the GBS (bottom right). [Figure 29] The panel above shows the optimization of flap binding sites (FBS) and donor / sprint DNA doses relative to the amount of gRNA. In the HEK293T GFP example, the optimal FBS is approximately 9–13 bases (bottom left). In the HEK293T GFP example, with a 13-base FBS, the optimal molar amount of sprint / donor is approximately 18% of the molar amount of gRNA (bottom right). [Figure 30] The optimization of donor and DBS lengths is shown (top panel). In the case of HEK293T GFP, substantial activity is observed over a wide range of donor lengths and an optimal length of approximately 22–26 nt (bottom left). Overhangs in the donor nucleic acid or sprint nucleic acid were observed to reduce efficiency (bottom right). [Figure 31]This figure shows the optimization of locked nucleic acids (LNAs). Figure 31A shows the effect of varying the number of LNAs in a sprint within the first 20 nt of alternating donor binding sites (DBS) from the 5' end of the sprint and within 19 nt of guide binding sites (GBS) at the 3' end of the sprint. Figure 31B shows the effect of including additional LNAs in the sprint near the target nick site or at the 5' end of the sprint. Figure 31C shows the effect of varying the number of LNAs in the donor binding sites (DBS) of the sprint for 32 nt DBS. The comparison is between 12 LNAs distributed toward the target nick site, 12 LNAs distributed toward the 5' end of the sprint, and 16 alternating LNAs. Figure 31D (SEQ ID NOs. 815, 556, and 816) shows the location and exemplary composition of nucleic acid binding sites for sprints with 20 nt DBS and 19 nt GBS. (Donor: SEQ ID NO: 556, / 5Phos / cgtaTgtcagggtggtcacGACgg; Sprint: SEQ ID NO: 557, +C*C*+GT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGgcgtgcagtgcttACGCCA+CA+AT+AC+CG+CA+G*C*+T; Guide: 3' end of the guide, e.g., SEQ ID NO: 166, SEQ ID NO: 170, SEQ ID NO: 173 or SEQ ID NO: 222). [Figure 32] This figure shows the functional aspects of the system and its components, including homologous recombination repair (HDR) independence. The high efficiency of replacer editing is not primarily due to HDR (Figure 32A). The efficiency of replacer editing depends on all components, including the splint, donor, Cas9, and ligase. Exogenous ligase substantially increases the ligase activity that may be present in target cells (Figure 32B). [Figure 33]Replacer efficiency can be enhanced by incorporating nucleotide analogs such as pseudo-UTP enhancement (Figure 33A). Various ligases are effective. Sprint R and T4 ligases showed particularly high efficiency (Figures 33B, C). Higher efficiency was obtained when nCas9 ligase and T4 ligase were expressed from separate mRNAs (left panel, split nCas9 and T4, two mRNAs). This may be due to mRNA size and not the protein itself, as T4-P2A-nCas9 mRNA expressing a self-cleaving fusion (Figure 33C) had the same efficiency as T4-nCas9 fusion. [Figure 34] This figure shows gene editing of several endogenous target genes in HEK293T cells before optimization. Point mutations were introduced into reported high-efficiency targets (e.g., AAVS1 (safe harbor site; Anzalone et al., Nature biotechnology 2022) and VEGFA (Wang et al., Nature method 2022)) as well as specific disease-associated sites (HBB E6V for sickle cell disease; CFTR R553X and G551D for cystic fibrosis; ATP7B H1069Q for Wilson's disease). [Figure 35] This figure shows the effect of nCas9 modification on improving editing efficiency. Including chromatin-modified peptide (CMP) HN1 and Rad51 DNA-binding domain (DBD) in the nCas9 fusion improved efficiency. [Figure 36] This figure shows the effects of donor DNA methylation and 3' end protection. Figure 36A: Donor methylation. Figure 36B: Donor methylation and 3' end protection with and without splint LNA modification. Figure 36C: Donor methylation by O-Me splint modification or LNA splint modification. [Figure 37]This figure shows the effect of sprint 3' end protection on editing efficiency. Figure 37A: Chemically blocking the 3' end of the sprint improves efficiency. This effect is shown for 3'AltR and 3'C3 spacers. Figure 37B: Splints ligated to biotin at the 3' end were tested using streptavidin-ligated T4 ligase. Figure 37C (SEQ ID NOs. 815, 558 and 816) shows the location of nucleic acid binding sites and exemplary composition for sprints with 20 nt DBS and 19 nt GBS. (Donor: SEQ ID NO: 558, / 5Phos / mCgtaTgtmCagggtggtmCamCGAmC*g*g-C3 spacer; Sprint: SEQ ID NO: 559, +C*C*+GT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGgcgtgcagtgcttACGCCA+CA+AT+AC+CG+CA+G*C*+T-C3 spacer; Guide: 3' end of the guide, e.g., SEQ ID NO: 166, SEQ ID NO: 170, SEQ ID NO: 173 or SEQ ID NO: 222). [Figure 38] This figure shows the effects of nicking guides (Figures 38A, B) and dead guides (Figure 38B). Nicking induces nicking into the opposite strand to facilitate the incorporation of the desired edit during DNA mismatch repair (MMR). Nicking guides can be screened to achieve improved efficiency with minimal indels. Dead guides allow nCas9 to engage with the target without nicking, opening the chromatin in the target region and including spacers (e.g., 15nt spacers) used in conjunction with further target strand spacers (Park et.al., Genome Biology 2021). [Figure 39] This figure shows highly efficient editing of disease-associated mutations. Replacer optimization included one or more of the following: fusing additional peptides to nCas9; chemically blocking the ends of donor and sprints; methylated donor DNA; including a nicking guide; screening for different FBS lengths; a step to remove the 3'LNA from the sprint (HBB only); and altering the gRNA backbone. [Figure 40]This figure shows the high-efficiency editing and precision of the replacer compared to prime edit PE2, including the absence of a nicking guide. The replacer demonstrates substantially reduced indel formation due to RNA synthesis errors, reverse transcriptase errors, and scaffold incorporation. A: Edited and modified reads (e.g., indels) shown as the percentage of all edited reads at five loci. B: Editing efficiency at five loci. C: Indels by location, replacer vs. PE. The replacer produces substantially fewer indels across a range of bases. [Figure 41] This figure shows unilateral and bilateral replacer deletions in disease-associated repeat regions. Top left: Unilateral replacer deletes a 96bp HTT in-frame in a repeat region, demonstrating indel deletion efficiency and low frequency. Top right: Bilateral replacer versus prime for deletions of a 131bp region and a 38bp Bxb1 attB substitution in C9orf72. Bottom: Efficient replacement (lanes 1 and 2) or deletion (lane 3) of C9orf72 by bilateral replacer. [Figure 42] This figure shows the effects of overlap length (A-C) and base modification (C) in a 2× one-sided replacer (e.g., as shown in Figure 25). The orientation and length of the attB sequence, as well as the overlap length between the two replacer donors, affect editing efficiency (B). Changing some LNAs to 2'-OMe bases during the sprint increases efficiency for longer donor overlaps (e.g., 38 bp donors) (C). [Figure 43] This figure shows high replacer efficiency in cells including non-dividing cells. Figure 43A shows bilateral replacer compared to prime editing in human primary hepatocytes (PHH), induced pluripotent stem cells (iPSCs), and HEK293T cells. Replacer efficiency can be optimized by LNP formulations. Figure 43B shows 175bp VEGFA substitution with 2bp "GT" in three cell lines using three different LNP formulations. [Figure 44]This figure shows the replacer editing efficiency in human hepatocytes. A: ATP7B mutation in primary human hepatocytes (PHH) and immortalized human fetal kidney cells (HEK293T). B: Substitution of the NOLC 79bp sequence with the Pa01 attB 33bp sequence in PHH cells and HEK293T cells. [Figure 45] This figure shows the effects of ligase and nCas9 sequencing on fusion protein and split expression. [Figure 46] This figure outlines a replacer system for editing human cells using a stably integrated BFP reporter that converts to GFP upon successful editing. Figure 46A shows an example of the replacer editing mechanism for A-to-C conversion editing. The donor DNA is homologous to the genome except for the edited C base. After the donor is ligated to the 3' flap, the donor anneals to the opposite strand of the genome, inducing editing via the endogenous mismatch repair pathway. Figure 46B is a schematic diagram of the replacer editing workflow. The nucleic acid is transfected into HEK293T cells expressing a BFP gene integrated into the virus. The replacer edits three nucleotides of the BFP gene to convert to GFP. The editing efficiency is equal to the percentage of GFP+ cells evaluated by flow cytometry. Figure 46C shows an exemplary architecture of the replacer nucleic acid and chemical modification layout. The sprint contains alternating LNAs in DBS and GBS, and the ligRNA contains three 2'-OMe nucleotides at the 5' and 3' ends. The replace suprint consists of a donor binding site (DBS) complementary to the donor DNA, a flap binding site (FBS) that hybridizes to a nicked 3' flap, and a guide binding site (GBS) that connects to the ligRNA. The ligRNA contains a typical Cas9 spacer and scaffold, as well as a 3' terminal sprint binding site (SBS) that binds to the sprint. The sprint and ligRNA can incorporate modifications that are not limited to phosphorothioate backbone modifications (PS), locked nucleic acids (LNA), and methylated nucleotides. [Figure 47]This figure shows a comparison of nucleic acid editing efficiencies with different numbers of alternating LNAs in the DBS and GBS regions (Figure 47A), different FBS lengths and DNA doses using 2.1 pmol of ligRNA for transfection (Figure 47B), and different donor and DBS lengths (Figure 47C). Figure 47A (SEQ ID NOs. 208, 560-563) shows the effect of LNAs incorporated into the DBS and GBS portions of the sprint. Sprint from top to bottom:+C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(LNAs in DBS / GBS-10 / 7)(SEQ ID NO: 208);+C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGC+CA+CA+AT+AC+CG+CA+G*C*+T(LNAs in DBS / GBS-9 / 8)(SEQ ID NO:560);+C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(LNAs in DBS / GBS-9 / 7)(SEQ ID NO: 561);+C*G*+TG+AC+CA+CC+CT+GA+CA+TACGGCGTGCAGTGCTTACGCCACA+AT+AC+CG+CA+G*C*+T(LNAs in DBS / GBS-9 / 6)(SEQ ID NO: 562);+C*G*+TG+AC+CA+CC+CT+GA+CATACGGCGTGCAGTGCTTACGC+CA+CA+AT+AC+CG+CA+G*C*+T(LNAs in DBS / GBS-8 / 8)(SEQ ID NO: 563). The right portion of Figure 52A shows %GFP fluorescent cells in transfected cells. Figure 47B shows the effect of sprint FBS length and DNA quantity on the editing efficiency of the HEK293T system. Figure 47C shows the conversion efficiency from BFP to GFP based on donor length and DBS length. [Figure 48]This figure shows a comparison of BFP vs. GFP editing efficiency using either a sprint and donor DNA with different mRNAs, either for a 100nt ssODN donor or replacer for HDR. The system optimally contains sprint, donor DNA, Cas9n, T4 ligase, and 5' phosphate. Figure 48A shows the effect of subtracting system components. 5' phosphorylation of donor DNA is not important in the HEK293T system. Also, HEK293T cells exhibit endogenous ligase activity, but its level is substantially increased by T4 ligase. Figure 48B shows the comparative efficiency of replacer systems using Cas9 (Cas9 nuclease), Cas9n (H840A Cas9 nickase), or Cas9n + T4 DNA ligase against a 100nt single-stranded oligodeoxynucleotide (ssODN) donor template with Cas9. Figure 48C shows the comparative efficiency of separate replacer mRNAs for BFP to GFP conversion using pseudouridine (m1Ψ) in either an nCas9-T4 fusion or a dual (nCas9 and T4) mRNA system. [Figure 49] Sequence IDs 564, 817, and 818 demonstrate replacer editing of point mutations at endogenous genomic loci in the human HEK293T cell line (Figure 49A) using a sprint containing nucleic acids (LNAs) alternately locked in GBS and DBS regions (Figure 49B). Exemplary chemical modification layout of an efficient system: Sprint: +A*G*+GC+CA+GC+AG+TG+AA+CA+AC+CA+TT+GGGCGTGGCAGTACGCCA+CA+AT+AC+CG+CA+G*C*+T-c3 SPACER (SEQ ID NO: 564); Donor DNA: / 5Phos / / IME-DC / AATGGTTGTT / IME-DC / A / IME-DC / TG / IME-DC / TGG / IME-DC / / IME-DC / *T-c3S PACER (SEQ ID NO: 565); SBS portion of ligRNA: AGCUGCGGUAUUGUGGmC*mG*mU (SEQ ID NO: 566). [Figure 50]This figure shows the editing efficiency of sprints with alternating locked nucleic acids (LNAs) in the GBS and DBS regions of HBB (Figure 50A), CFTR (Figure 50B), and ATP7B (Figure 50C). [Figure 51] Figure 51A shows the editing efficiency from BFP to GFP created using the splitting method in Figure 50, and the effect of adding additional nicking guides to the indels in HBB (Figure 51A), CFTR (Figure 51B), and ATP7B (Figure 51C). The positions represent the orientation and number of nucleotides between the two nicks. [Figure 52] This figure shows the editing efficiency and indels in sprint and donor DNA at different doses compared with ligRNA for different FBS lengths in HBB (Figure 52A), CFTR (Figure 52B), ATP7B (Figure 52D), and AAVS1 (Figure 52D). All nCas9 mRNAs were used with T4 ligase mRNA. [Figure 53] This figure shows the editing efficiency of splints with and without a C3 spacer at the 3' end of the splint in HBB (Figure 53A), CFTR (Figure 53B), and ATP7B (Figure 53C). [Figure 54] This figure shows the editing efficiency of sprints with several DNA modifications: phosphorothioate (PS) binding and methylated cytosine (meC) in HBB (Figure 54A), CFTR (Figure 54B), and ATP7B (Figure 54C), regardless of whether or not they have a C3 spacer at the 3' end. [Figure 55]Figure 55A shows a schematic diagram of a one-sided replacer editing system. In the one-sided replacer editing system, the sprint ligase-mediated guide RNA (lmgRNA) and donor DNA together, cleaving the lmgRNA-Cas9 complex, cleaving the genomic DNA, opening the R loop, and making the flap on the opposite side of the stand accessible. The sprint then binds to the genomic flap, placing the donor at the nick, the ligase places the donor at the genomic DNA flap, the L-PGI component dissociates, and only the donor DNA remains as a permanent edit. Figure 60B shows a two-sided replacer editing system, where two complete replacer editing complexes target protospacers on the opposite strand of DNA. Figure 55C shows an example of an overview of the editing process. Nucleic acid transfection includes two sets of sprints, donor DNA, and ligRNA. 72 hours after transfection, the genomic DNA can be extracted for PCR amplification, and the amplicons can be visualized on a gel or sequenced to evaluate the editing efficiency. [Figure 56] Design strategies for bilateral replacers are presented, where the donor DNA replacers may be homologous to each other for editing or homologous to the genome for deletion editing. These include complete overlap (Figure 56A) where the edited flaps are completely joined to each other, partial overlap (Figure 61B) where the 3' segments of the edited flaps are joined to each other, and replacement by deletion (Figure 56C). [Figure 57]This figure shows unilateral and bilateral replacer editing of disease-related repeat regions. Figure 57A shows agarose gel images of PCR amplicons for C9ORF72 after replacer deletion of the 222bp and 131bp regions of C9orf72. The WT fragment is unedited, and the fragment at the bottom is the desired substitution or deletion. attB and att Brev comp are replacements of nickel-nicked regions by forward or reverse Bxb1 attB sites. del F+R is a deletion, delF is a deletion with Rev sprint and no donor, and del R is a deletion with Fwd sprint and no donor. Figure 57B shows agarose gel images of PCR amplicons for Bxb1 attB substitution editing of the 131bp region of C9ORF72, comparing donor DNA with different chemical modifications. The fragment with the desired edit is at the bottom. [Figure 58]This figure shows the editing efficiency of attB substitution editing at two targets, quantified by Amplicon NGS. T4 ligase mRNA is used together with the shown nCas9 mRNA. Figure 58A shows the substitution of the 79bp sequence of NOLC1 with the 33bp sequence of Pa01 attB, Figure 58B shows the VEGFA substitution of the Bxb1 attB site at 38bp, and Figure 58C (SEQ ID NOs. 567-576) shows the effect on editing efficiency of sprints with different overlap lengths and variations in number and chemical modifications (LNA and 2'-OMe) between forward and reverse sprints in a bilateral replacer system showing accurate editing versus indel production, where 10 is partial overlap and 38 is full overlap. Each sprint is used with donor DNA of the same length as its DBS, and all sprints contain the same FBS and GBS, not shown. Efficiency is evaluated by Amplicon NGS.Up and down: 10bp sprint overlap with LNA (12 out of 24), forward sprint DBS=+G*G*+AG+AC+CG+CC+GT+CG+TC+GA+CA+AG+CC (SEQ ID NO: 567), reverse sprint DBS=+G*G*+CG+GT+CT+CC+GT+CG+AG+GA+TC+AT (SEQ ID NO: 568); 10bp sprint overlap with LNA (7 out of 24) and 2'O-Me nucleotides (5 out of 24), forward sprint DBS=+G*G*+AGmAC+CGmCC+GTm CG+TCmGA+CAmAG+CC (Sequence ID 569), Reverse sprint DBS=+G*G*+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+AT (Sequence ID 570); 38bp sprint overlaps with LNA (19 out of 38), Forward sprint DBS=+A*T*+GA+TC+CT+GA+CG+AC+GG+AG+AC+CG+CC+GT+CG+TC+GA+CA+AG+CC (Sequence ID 571), Reverse sprint DBS=+G*G*+CT+TG+TC+GA+CG+ AC+GG+CG+GT+CT+CC+GT+TC+AG+GA+TC+AT (SEQ ID NO: 572); 38bp sprint overlap with LNA (10 out of 38) and 2'-OMenucleotide (9 out of 38), forward sprint DBS:=+A*T*mGA+TCmCT+GAmCG+ACmGG+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+CC (SEQ ID NO: 573), reverse sprint DBS=+G*G*mCT+TGmUC+GAmCG+ACmGG+CGmGT+CTmCC+GTmCG+TC mAG+GAmTC+AT (SEQ ID NO: 574); 38bp sprint overlap with LNA (13 out of 38) and 2'-OMenucleotide (6 out of 38), forward sprint DBS:=+A*T*+GA+TC+CT+GA+CG+ACmGG+AGmAC+CGmCC+GTmCG+TCmGA+CAmAG+CC (SEQ ID NO: 575), reverse sprint DBS:=+G*G*+CT+TG+TC+GA+CG+ACmGG+CGmGT+CTmCC+GTmCG+TCmAG+GAmTC+AT (SEQ ID NO: 576). [Figure 59]This figure shows replacer editing in primary human hepatocytes. It compares the replacer and PE2 or PEMax prime editing systems in HEK293T cells and primary human hepatocytes (PHH) for point mutations in ATP7B (Figure 59A) and replacement of the 79bp sequence of NOLC1 with the 33bp sequence of Pa01 attB (Figure 59B). The PBS and RTT sequences for prime editing match the sprint FBS and DBS sequences of the replacer, respectively. Editing efficiency and indel generation are calculated by analyzing NGS of the amplicons. [Figure 60]This figure shows the optimization of chemical modifications in sprints and donor DNA for replacer editing in BFP. Figure 60A (SEQ ID NOs 577-583) shows the variation in the number of LNAs in the 21nt DBS region of the sprint. The control sprint has alternating LNAs in the DBS, which is typically optimal for 21nt DBS. Upper and lower: +T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (Sequence number 57 7);+T*+C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(Sequence number 5 78);+T*+C*+G+T+GA+CC+AC+CC+TG+AC+AT+AC+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(Sequence number No. 579);+T*C*+GT+GA+CC+AC+CC+TG+AC+AT+A+C+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(Array Number 580);+T*C*+GT+GA+CC+AC+CC+TG+AC+A+T+A+C+GGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(Sequence No. 581);+T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+G+GCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(Sequence No. 582);+T*C*+GT+GA+CC+AC+CC+TG+AC+AT+AC+GG+CGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(Sequence No. 583). Figure 60B (Sequence Nos. 208, 584-588) shows a comparison of different chemical modifications (LNA, 2'-OMe, 2F) in the sprint DBA and GBS regions.Up / down:+C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(LNA / LNA)(SEQ ID NO: 208);+C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA / i2FC / A / i 2FA / T / i2FA / C / i2FC / G / i2FC / A / i2FG / *C* / 32FU / (LNA / 2'-F)(SEQ ID NO. 584); / 52FC / *G* / i2FU / G / i2FA / C / i2FC / A / i2FC / C / i2FC / T / i2FG / A / i2FC / A / i2FU / A / i2FC / GGCGTGCAGTGCTTACGCCA+CA +AT+AC+CG+CA+G*C*+T(2'-F / LNA)(SEQ ID NO: 585);+C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCAmCAmATmACmCGmCAmG*C*mT(LNA / 2'-OMe)(SEQ ID NO: 586);mC*G*mTGmACmCAmCCmCTm GAmCAmTAmCGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(2'-OMe / LNA) (SEQ ID NO: 587); C*G*TGACCACCCTGACATACGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T(DNA / LNA) (SEQ ID NO: 588). Figure 60C (Sequence IDs 589, 590, and 591) shows the effect of changing the number and position of LNAs in the DBS at 32nt of the sprint.Up and down: 32 16LNAs (+C*T*+GG+CC+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (Sequence No. 589)); 32 12LNAs, reduced proximity nick (+C*T*+GG+CC+CA+CC+CTC+GTG+ACC+ACC+CTG+ACA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CA+G*C*+T (Sequence No. 590)); 32 12LNAs, reduced proximity 5' end (+C*T*+GGC+CCA+CCC+TCG+TGA+CCA+CC+CT+GA+CA+TA+CGGCGTGCAGTGCTTACGCCA+CA+AT+AC+CG+CAG*C*T (Sequence No. 591)). Figure 60D shows the effect of methylated DNA donors on the conversion efficiency from BFP to GFP, which depends on the chemical modification of sprint DBS. [Figure 61A] (Sequence IDs 592-594) are diagrams comparing different pairs of sprint GBS and ligRNA SBS sequences with 3' terminal modifications of ligRNA. SBS folding ΔG is the strength of secondary structure in the 20nt 3' extension (SBS) of ligRNA. All sprints contain the same DBS and FBS (not shown), and the ligRNA contains the same scaffold and spacer (not shown). Top and bottom: AGCUGCGGUAUUGUGGmC*mG*mU (Sequence ID 592); GUGGUUCCGGGCUGCAmU*mG*mA (Sequence ID 593); CGAUUCCUGAUACUGCmU*mG*Mc (Sequence ID 594). [Figure 61B](Sequences 595, 818, 597, 822, 595, 822, 600, 822, 601, 823, 601, 824, 601, 825, 600, 825, 600, 826, 595, and 818) are diagrams showing replacer optimization, including GBS / SBS length, sprint composition, and sprint optimization by FBS linker, as well as ligRNA optimization. Nucleotide sequences are shown bound together, and in some cases, a portion of the GBS or SBS functions as a single-stranded linker. All sprints contain the same DBS and FBS (not shown), and the ligRNA contains the same scaffold and spacer (not shown). From top to bottom. Length of GBS / SBS (nt) = 24 / 19: ACCGTCACGC+CA+CA+AT+AC+CG+CA+G*C*+T (Sequence ID 595) and AGCUGCGGUAUUGUGGmC*mG*mU (Sequence ID 596); Length of GBS / SBS (nt) = 15 / 15: +CA+CA+AT+AC+CG+CA+G*C*+T (Sequence ID 597) and AGCUGCGGUAUUmG*mU*mG (Sequence ID 598); Length of GBS / SBS (nt) = 19 / 15: ACGC+CA+CA+AT+AC+CG+CA+G*C*T (Sequence ID 599) and AGCUGCGGUAUUmG*mU*mG (Sequence ID 598); Length of GBS / SBS (nt) = 20 / 15: GACGC+CA+CA+AT+AC+CG+ CA+G*C*+T (Sequence ID 600) and AGCUGCGGUAUUmG*mU*mG (Sequence ID 598); GBS / SBS length (nt) = 28 / 40: GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (Sequence ID 601) and CGATTTCCTGATAGTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (Sequence ID 602); GBS / SBS length (nt) = 28 / 32: GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (Sequence ID 601) and GATAGTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (Sequence ID 603); GBS / SBS length (nt) = 28 / 28;GACGC+CA+CA+AT+AC+CG+CA+GC+TG+GC+AG+C*A*+C (Sequence ID 601) and GTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (Sequence ID 604); Length of GBS / SBS (nt) = 20 / 28: GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (Sequence ID 600) and GTGCTGCCAGCUGCGGUAUUGUGGCG*mU*mC (Sequence ID 6 04); GBS / SBS length (nt) = 20 / 20: GACGC+CA+CA+AT+AC+CG+CA+G*C*+T (Sequence ID 600) and AGCUGCGGUAUUGUGGCmG*mU*mC (Sequence ID 605); GBS / SBS length (nt) = 19 / 19: ACGC+CA+CA+AT+AC+CG+CA+G*C*+T (Sequence ID 595) and AGCUGCGGUAUUGUGGmC*mG*mU (Sequence ID 606). [Figure 61C](Sequence IDs 607, 810, 609, 620, 609, 819, 607, 821, 607, 820) are diagrams showing replacer optimizations, including sprint optimization of DBS length and DBS / donor overhang. An ssDNA overhang exists when the sprint DBS and donor DNA lengths are not equal. All sprints contain the same DBS and FBS (not shown). Top to bottom. DBS / donor length (nt) = 20 / 28nt: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (Sequence ID 607) and / 5Phos / CGTATGTCAGGGTGGTCACGAGGGTGGG (Sequence ID 608); DBS / donor length = 28 / 20nt: +C*C*+CA+CC+CT+CG+TG+AC+CA+CC+CT+GA+CA+TA+CG (Sequence ID 609) and / 5Phos / CGTATGTCAGGGTGGTCACG (Sequence ID 610); DBS / donor length = 28 / 28nt: +C*C*+CA+CC+CT+CG+TG+AC+CA+C C+CT+GA+CA+TA+CG (Sequence ID 609) and / 5Phos / CGTATGTCAGGGTGGTCACGAGGGTGGG (Sequence ID 608); DBS / donor length = 20 / 21nt: +C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (Sequence ID 607) and / 5Phos / CGTATGTCAGGGTGGTCACGA (Sequence ID 611); DBS / donor length = 20 / 20nt + C*G*+TG+AC+CA+CC+CT+GA+CA+TA+CG (Sequence ID 607) and / 5Phos / CGTATGTCAGGGTGGTCACG (Sequence ID 610). [Figure 62]This figure shows the evaluation of ligase and protein architectures for replacer editing in BFP. Figure 62A shows a comparison of three different ligases as a fusion protein with nCas9 encoded in a single mRNA, or as separate mRNAs used with nCas9 mRNA. When using two mRNAs, nCas9 and the ligase are fused to either the N-terminus or C-terminus leucine zipper (LZ) to promote co-localization. BFP was converted to GFP using 24nt donor DNA and sprint DBS. Figure 62B shows a comparison of the three ligases against no ligase, a T4-nCas9 fusion, and T4-P2A-nCas9 biscistronic mRNA for converting BFP to GFP using 50nt donor and DBS. Figure 62B shows the effect of LZ on the conversion efficiency from BFP to GFP. Each ligase mRNA is delivered with either nCas9 mRNA without an LZ on the protein (left bar), mRNA with an LZ on the ligase C-terminus and nCas9 N-terminus (center bar), or mRNA with an LZ on the ligase N-terminus and nCas9 C-terminus (right bar). [Figure 63] This figure shows deletion, substitution, and conversion by L-PGI in HEK293T cells (Figures 63A-C) with bilateral deletion (Figure 63), bilateral beacon replacement (Figure 63B), and unilateral correction / conversion (Figure 63C). [Modes for carrying out the invention]
[0041] Introduction Advances in genome editing tools have made precise genome editing possible for therapeutic, agricultural, industrial, and research purposes. Some editing tools may involve the insertion of nucleic acid sequences into the genome at a target site. However, some editing tools may have limitations on the length of the usable integrated sequence.
[0042] Several nuclease-based tools, such as CRISPR-Cas9, use guide RNA to target the Cas9 protein to a specific DNA sequence designated by a spacer sequence in the guide RNA. The Cas9 nuclease activity then cleaves the DNA, resulting in a double-strand break (DSB). DSBs are typically repaired by endogenous DNA repair mechanisms, including non-homologous end joining (NHEJ) or homologous recombination repair (HDR). However, NHEJ results in a series of nucleotide insertions and deletions (indels) that hinder its usefulness for precise editing. HDR efficiency is very low in non-dividing cells and may require DNA replication. Even when HDR editing is detectable, DSB-induced indels are often common, meaning HDR is not viable when precise editing is desired.
[0043] Homologous-independent target insertion (HITI) utilizes the NHEJ DNA repair mechanism, which is active in non-dividing cells, for CRISPR-induced transgene integration in non-dividing cells such as primary neurons, retinal pigment epithelial cells, and HSPCs. However, for the generation of double-stroke subunits (DSBs) from Cas9, HITI generates a high frequency of indels, leading to unintended mutations in addition to DSB-related toxicity.
[0044] Other methods for gene editing have further limitations. Tools using fusions of Nicking Cas nucleases and nucleotide deaminases (e.g., base editors) can perform specific nucleotide mutations; for example, a cytosine base editor can convert C to T. While some base editors can perform highly efficient and precise editing, they are inherently limited to specific edits determined by deaminase variants, and therefore are only applicable to specific substitution mutations, and cannot perform more precise insertion or deletion edits. Furthermore, base editors are generally limited to small editing windows within a subset of protospacer regions and are thus significantly limited by the availability of protospacer adjacent motifs (PAMs). Finally, base editors can exhibit bystander mutations within the editing region (e.g., the presence of two Cs) and have demonstrated DNA and RNA off-target deaminase activity.
[0045] Existing precision editing techniques have limitations that hinder their practical applicability in various ways. In particular, they may rely on endogenous cellular mechanisms for editing, such as HDR mechanisms for nuclease-based editing and mismatch repair for base editing. No systems that are completely independent of endogenous factors have been reported. Dependence on endogenous factors is problematic because different cell types have different activity levels of these endogenous factors, and often the activity is insufficient to provide a useful level of editing. An example where this dependence is particularly problematic is non-dividing cells, which make up the majority of adult cells and are therefore unsuitable for many existing precision editing tools.
[0046] Therefore, there is still a need for effective gene editing (e.g., modification) or systems or methods for modifying gene expression through gene editing. In particular, there is still a need for systems or methods that limit the length of the modified nucleotide sequence to a minimum. Furthermore, there is still a need for systems or methods for gene editing or modification of gene expression that are independent of the endogenous components or mechanisms of the cell. There is also still a need for systems or methods for correcting gene mutations in cells. In some cases, the correction of gene mutations can treat the disease or condition in which it is needed. As will be seen below, the systems, methods, and compositions disclosed herein may be useful in addressing these needs or limitations.
[0047] overview This specification describes self-contained gene editing systems. In some such self-contained systems, all aspects of gene editing can be controlled. Some such systems are independent of host cell mechanisms for performing editing functions or for replacing or repairing any aspect of a target nucleic acid, such as a genomic locus. Some such systems can perform editing without the use of polymerase and are therefore unaffected by the cellular nucleotide triphosphate (dNTP) concentration. For example, an exogenous first integrated nucleic acid can be delivered without transcribing a template and inserted into a locus. Editing can eliminate the need to rely on cell repair systems such as HDR or NHEJ. Editing may be performed without a cell cycle. Gene editing may be performed intracellularly or in vitro. For example, gene editing may be performed in vitro or extracellularly.
[0048] This specification describes a system and method for editing DNA with a donor strand without generating double-strand breaks in the genome, using a CRISPR-guided DNA ligase and a guide nucleic acid that targets a genomic region of interest. A DNA ligase is an enzyme that chemically joins two DNA molecules via a phosphodiester bond. The DNA ligase may or may not require hybridization of the DNA molecule to a DNA or RNA backbone, or to a "sprint" that is inversely complementary to the DNA sequence to be ligated. Targeting of the ligase to a genomic nick produced by a CRISPR nuclease allows for precise replacement of the genomic DNA with a donor strand, which may be recruited to the target locus by the guide nucleic acid. CRISPR-guided DNA ligases may consist of a DNA ligase that is fused, recruited, or unfused to an RNA guide endonuclease by utilizing a peptide linker, a heterodimerizing domain, or two distinct peptides, respectively.
[0049] Some embodiments include cells containing or comprising RNA-guided endonucleases and DNA ligases, both of which are introduced into the cells. The endonuclease or ligase may be heterogeneous to the cell. The endonuclease and ligase may be heterogeneous to the cell. The ligase may be endogenous to the cell. In some embodiments, the cell comprises RNA-guided endonucleases and DNA ligases, both of which are heterogeneous to the cell. The cell may comprise the compositions or systems described herein. The cell may be used in or included in the systems, compositions, or methods described herein.
[0050] The system described herein may include heterologous endonucleases, including RNA guide endonucleases such as Nicking Cas9, and heterologous ligases (e.g., DNA ligases) that can utilize RNA sprinting. The guide nucleic acid optionally recruits a donor strand to a site targeted by the endonuclease (e.g., a targeted genomic locus), and also generates a sprint from the donor strand (donor strand) and genomic flap generated by Nicking Cas9, resulting in ligation of the donor strand and genomic flap by the DNA ligase. In some embodiments, the ligase is an endogenous ligase or comprises an endogenous ligase. The system can utilize one or more guide nucleic acids that together optionally contain the following components in the following order: 5' spacer - scaffold - donor binding site (optional) - flap binding site 3'. The donor strand (donor strand) may contain the following sequence components: 5' guide binding site - donor strand 3'. The guide binding site of the donor strand is at least partially inversely complementary to the donor binding site of the guide nucleic acid, so that the donor hybridizes to the guide and localizes to the target site of the RNA guide endonuclease. The 5' end of the donor sequence and the 3' end of the genomic flap generated by nuclease-nicking activity are ligated by DNA ligase and sprinted by the donor binding site and the flap binding site of the guide nucleic acid.
[0051] Figures 1A–1C show a non-limiting example of the system (one-sided replacer 1). This example includes: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting the endonuclease described herein; a donor binding site for forming a complex with a donor strand; and a guide nucleic acid including a flap binding site for forming a complex with the genomic flap of the genomic locus. The guide nucleic acid is shown complexing with an endonuclease (e.g., Cas9 nickas, nCas9) operably bound to a ligase. The guide nucleic acid can induce the endonuclease at the genomic locus bound by the spacer of the guide nucleic acid. The guide nucleic acid is also shown as being partially complementary to the donor strand (complexing between the donor binding site of the guide nucleic acid and the guide binding site of the donor strand). An endonuclease, directed by a guide nucleic acid, can cleave or nick at least one strand of a genomic locus, and a ligase can ligate one end of the donor strand to the cleaved or nicked end of the genomic locus, thus incorporating the donor strand into the locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by the nuclease.
[0052] Figures 2A to 2C show a non-limiting example of the system (bilateral replacer 1). The guide nucleic acid in the example, similar to the guide nucleic acid in Figure 1A, includes a spacer for targeting a genomic locus; a scaffold for complexing and recruiting the endonuclease described herein; a donor binding site for forming a complex with a donor strand; and a flap binding site for forming a complex with the genomic flap of the genomic locus. In Figure 2A, the first guide nucleic acid is shown complexing with a first endonuclease functionally coupled to a first ligase, and the second guide nucleic acid is shown complexing with a second endonuclease functionally coupled to a second ligase. The first and second nucleases can each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus can then be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. Insertion of a donor strand at a genomic locus can generate two genomic flaps that can be digested and removed by nucleases.
[0053] Figures 3A–3C show a non-limiting example of the system (one-sided replacer 2). In this example, the guide nucleic acid includes a spacer for targeting a genomic locus; a scaffold for complexing and recruiting the endonuclease described herein; and a donor binding site for forming a complex with the donor strand. Figure 3A also shows a donor strand including at least one overhang, the overhang including a flap binding site for forming a complex with the genomic flap of the genomic locus; and a guide binding site for forming a complex with the guide nucleic acid (via the donor binding site of the guide nucleic acid). The guide nucleic acid can complex with an endonuclease (e.g., nCas9) operably bound to a ligase. The guide nucleic acid in the example directs the endonuclease and ligase to the genomic locus bound by the spacer of the guide nucleic acid. The guide nucleic acid in the example is also partially complementary to the donor strand (complexing between the donor binding site of the guide nucleic acid and the guide binding site of the donor strand). An endonuclease, directed by a guide nucleic acid, can cleave at least one strand of a genomic locus, and a ligase can ligate one end of the donor strand to the cleaved end of the genomic locus, thus incorporating the donor strand into the locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by the nuclease.
[0054] Figures 4A–4C show non-limiting examples of the system (bilateral replacer 2). In an example where the guide nucleic acid is similar to the guide nucleic acid in Figure 3A, and includes a spacer for targeting a genomic locus, it includes a scaffold for complexing and recruiting the endonuclease described herein, and a donor binding site for forming a complex with the donor strand. Figure 4A also shows a donor strand including two overhangs, each including a flap binding site for forming a complex with the genomic flap of the genomic locus and a guide binding site for forming a complex with the guide nucleic acid (via the donor binding site of the guide nucleic acid). The flap binding site of the donor strand allows the donor strand to be brought close to the genomic locus after the genomic flap is generated following the endonuclease cleaving at least one strand of the genomic locus. In Figure 4A, the first guide nucleic acid is shown complexing with the first endonuclease functionally coupled to the first ligase, and the second guide nucleic acid is shown complexing with the second endonuclease functionally coupled to the second ligase. In this example, the first and second nucleases each cleave at least one strand of the genomic locus. The two cleaved ends of the genomic locus are then ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. In this example, the insertion of the donor strand into the genomic locus generates two genomic flaps that can be digested and removed by the nucleases.
[0055] The system described herein (Replacer 3) may include heterologous endonucleases, including an RNA guide endonuclease such as Nicking Cas9 and a ligase that can utilize DNA sprints (e.g., DNA ligase). The guide nucleic acid optionally recruits a donor strand to a site targeted by the endonuclease (e.g., a targeted genomic locus), and also generates a sprint from the donor strand (donor strand) and genomic flap produced by Nicking Cas9, resulting in ligation of the donor strand and genomic flap by the DNA ligase. At least a portion of the flap binding site and donor binding site on the guide nucleic acid is DNA such that the ligase utilizing the DNA sprint can catalyze the intended reaction. The system may utilize one or more guide nucleic acids that together optionally contain the following components in the following order: 5' spacer - scaffold - donor binding site (optional) - flap binding site 3'. The donor strand (donor strand) may contain the following sequence components: 5' guide binding site - donor strand 3'. The guide binding site of the donor strand is at least partially inversely complementary to the donor binding site of the guide nucleic acid so that the donor hybridizes to the guide and localizes to the target site of the RNA guide endonuclease. The 5' end of the donor sequence and the 3' end of the genomic flap generated by nuclease-nicking activity are ligated by DNA ligase and sprinted by the donor binding site and the flap binding site of the guide nucleic acid.
[0056] Figures 5A to 5C show a non-limiting example of the system (one-sided replacer 3). This example includes: a spacer for targeting a genomic locus; a scaffold for complexing and recruiting the endonuclease described herein; a donor binding site for forming a complex with a donor strand; and a flap binding site for forming a complex with the genomic flap of the genomic locus, the guide nucleic acid comprising the flap binding site, wherein at least a portion of the flap binding site and the donor binding site consists of DNA. The guide nucleic acid is shown complexing with an endonuclease (e.g., Cas9 nickas, nCas9) operably bound to a ligase (e.g., endogenous ligase or exogenous ligase). The guide nucleic acid can induce the endonuclease at the genomic locus bound by the spacer of the guide nucleic acid. The guide nucleic acid is also shown as being partially complementary to the donor strand (complexing between the donor binding site of the guide nucleic acid and the guide binding site of the donor strand). An endonuclease, directed by a guide nucleic acid, can cleave at least one strand of a genomic locus, and a ligase can ligate one end of the donor strand to the cleaved end of the genomic locus, thus incorporating the donor strand into the locus. The incorporation of the donor strand into the genomic locus may generate a genomic flap that can be digested and removed by the nuclease.
[0057] Figures 6A to 6C show non-limiting examples of the system (bilateral replacer 3). The guide nucleic acid in the examples, similar to the guide nucleic acid in Figure 5A, includes a spacer for targeting a genomic locus; a scaffold for complexing and recruiting the endonuclease described herein; a donor binding site for forming a complex with a donor strand; and a flap binding site for forming a complex with the genomic flap of the genomic locus, wherein at least a portion of the flap binding site and the donor binding site consists of DNA. In Figure 6A, the first guide nucleic acid is shown complexed with a first endonuclease functionally coupled to a first ligase, and the second guide nucleic acid is shown complexed with a second endonuclease functionally coupled to a second ligase. The first endonuclease and the second nuclease can each cleave at least one strand of the genomic locus. Next, the two cleaved ends of the genomic locus can be ligated to the two ends of the donor strand, thereby incorporating the donor strand into the genomic locus. Insertion of the donor strand into the genomic locus can generate two genomic flaps that can be digested and removed by nucleases.
[0058] The ligation can be carried out using a DNA ligase that can utilize RNA sprinting, such as Chlorella virus-derived sprint R ligase (also known as PBCV-1 DNA ligase). In some embodiments, the system utilizes two guide nucleic acids that target a CRISPR guide ligase to a target site on the opposite strand adjacent to the genomic region of interest. In some embodiments, each guide nucleic acid interacts with the corresponding donor strand in the manner described above, resulting in ligation of both donor strands that are inversely complementary to each other in the donor strand region.
[0059] A ligase, either fused to or recruited to an endonuclease, or supplied trans-, can utilize DNA as a sprint, with the donor strand acting as a sprint for the genomic flap generated by the endonuclease and another donor strand. In some embodiments, the donor strand comprises a 5' donor strand - flap binding site - optional guide binding site 3'. The flap binding site on one donor strand (donor 2) may be inversely complementary to the genomic flap, while an optional guide binding site on donor 2 may be inversely complementary to an optional donor binding site on a guide nucleic acid (guide 1), and the donor strand may be at least partially inversely complementary to a different donor strand (donor 1). The 5' end of donor 1 and the 3' end of the genomic flap can be ligated using the flap binding site and donor strand of donor 2 as a sprint. Such a bilateral approach utilizing dual guide nucleic acids with different spacer sequences can be employed with Donor2, a nick constructed using a second replacer 2 guide nucleic acid (Guide2) having a spacer sequence that provides a sprint to a first genomic site and targets a second site, with its 5' end ligated to the 3' end of a different genomic flap. The donor binding site on the second guide nucleic acid system can optionally recruit Donor1 via hybridization with its optional guide binding site, with Donor1 acting as a DNA sprint for the ligation of Donor2 to the 3' end of the genomic flap at the target site of the second guide nucleic acid.
[0060] After ligation, the remaining flap of native genomic DNA can be excised by exogenously delivered or endogenous flap endonucleases or exonucleases. Examples of exogenous nucleases that can be introduced into cells include human flap endonuclease 1 (hFEN1), human exonuclease 5 (hEXO5), T5 exonuclease, T7 exonuclease, exonuclease VIII, flap endonuclease domain of Escherichia coli (E. coli) PolI, RecJF, lambda exonuclease, Xni (ExoIXI) derived from Escherichia coli (E. coli), SaFEN (Staphylococcus aureus FEN), nuclease BAL-31, or fragments thereof. The endonuclease or exonuclease may be fused to, recruited to, or unfused to an RNA guide endonuclease or DNA ligase by utilizing a peptide linker, a heterodimerizing domain, or two distinct peptides, respectively.
[0061] In some embodiments, the systems, compositions, or methods described herein utilize additional proteins that bind to cleavage or nick sites. For example, the systems, compositions, or methods described herein may include a Ku protein or a Gam protein derived from bacteriophage Mu, the binding of which can increase the ligation efficiency of the embedded nucleic acid at the cleavage or nick site.
[0062] The systems or methods described herein can utilize nickel endonucleases and therefore do not produce double-strand breaks. Furthermore, the systems described herein address the problem of low editing efficiency in non-dividing cells due to their mechanism of action, which relies solely on exogenous components delivered to cells using mRNA, viral vectors, guide nucleic acids, DNA or peptides, or any other modality. Therefore, this system does not require cell cycle-dependent endogenous cell processes or the presence of components such as HDRs or dNTPs. Thus, the systems described herein enable unhindered efficiency in non-dividing cells. Moreover, this system can enable the substitution of both strands of the target region of the genome, thereby increasing editing efficiency.
[0063] The donor strand may contain a high degree of homology to the substituted genomic DNA. These donors may contain mutations to the genomic DNA, such as pathogenic mutation correction, CRISPR protospacer adjacent motif (PAM) site deactivation, disruption of the guide spacer sequence, other substitution mutations, or combinations thereof. Further substitution mutations may be included to increase donor-donor homology versus donor-genomic homology and promote hybridization and integration of the donor strand into the genome. The donor strand may also encode nucleotide deletions or insertions, or complex combinations of the above that substitute for the target genomic DNA. Optionally, the guide strand and donor strand may be chemically modified using nucleic acid chemistry, such as phosphorothioate bonding or 2'-O-methylation. Optionally, the guide nucleic acid may contain hairpin sequences. Optionally, any combination of guide nucleic acid, donor strand, and protein can be complexed, for example, using an annealing reaction (gradual decrease in temperature) before delivering the editing components to the cell.
[0064] Protein components (e.g., Nicking Cas9, ligases) can be modified using nuclear localization signals, cell-permeable peptides, or chromatin disruption peptides to improve delivery efficiency to genomic targets.
[0065] The primary cellular DNA repair pathway for resolving small (<13nt) mismatches between genomic DNA strands is mismatch repair (MMR). In single-stranded donor ligation, the ligated donor strand forms a DNA heteroduplex with the reverse-complementary genomic DNA strand. This can also occur through competitive hybridization between the ligated donor strand and the genomic DNA strand. In these cases, MMR activity can eliminate and reverse the donor strand mismatch using the genomic strand as a template, resulting in reduced editing. Dominant-negative expression of MMR proteins has been shown to inhibit the MMR pathway and improve editing outcomes when similar DNA heteroduplexes are generated. In some embodiments, dominant-negative MMR peptides such as MSH2(G674A) and MLH1(del754-756) may be delivered as part of the systems described herein to improve genome editing capabilities, particularly in cells overexpressing the MMR pathway. In some embodiments, these dominant-negative MMR peptides may be delivered by fusion (e.g., fusion with any component of the system described herein), mobilization, or as separate peptides.
[0066] Endonuclease Endonucleases are disclosed herein. Endonucleases may be included in compositions, systems, or methods disclosed herein. Endonucleases may be recombinant. Endonucleases may be bound to ligases. Endonucleases may be bound to ligases directly or indirectly. The coupling may be covalent or non-covalent. Endonucleases may be bound to or linked to ligases. Endonucleases may be recruited to be part of a fusion protein with a ligase, or may be used together with a ligase. Endonucleases may be coupled to integrases. Endonucleases may be coupled to or linked to integrases directly or indirectly. The coupling may be covalent or non-covalent. Endonucleases may be bound to or linked to integrases. Endonucleases may be recruited to be part of a fusion protein with an integrase, or may be part of a fusion protein with an integrase, or may be used together with an integrase. Endonucleases may be heterologous. Heterogeneous may indicate a source from no cell. Where heterogeneous endonucleases are described, non-heterogeneous (e.g., endogenous) endonucleases may be used in some examples. Endonucleases may be encoded intracellularly. Endonucleases may be delivered to cells in trans. Endonucleases may catalyze the cleavage of phosphate bonds in an exogenous first integrated nucleic acid. Endonucleases may be guided by a guide nucleic acid to cleave or nick a target nucleic acid for ligation of the exogenous first integrated nucleic acid at a cleavage site or nick site. Endonucleases may include any embodiment shown in Figures 1A to 6C.
[0067] Endonucleases may not be naturally occurring. Endonucleases may be manipulated. Endonucleases may be synthesized. Endonucleases may be pre-synthesized. Endonucleases may be added to a subject or cell. Endonucleases may be encoded by nucleic acids. Encoding nucleic acids may be manipulated, synthesized, or added to a subject or cell.
[0068] At least a portion of the endonuclease may be contained in the first polypeptide. At least a portion of the endonuclease may be contained in the second polypeptide. The endonuclease may be divided into two or more polypeptides that are bound together. The first polypeptide may contain the N-terminal portion of the endonuclease. The first polypeptide may contain the C-terminal portion of the endonuclease. The second polypeptide may contain the N-terminal portion of the endonuclease. The second polypeptide may contain the C-terminal portion of the endonuclease. The first or second polypeptide containing a portion of the endonuclease may be fused with at least a portion or all of the ligase. The first or second polypeptide containing a portion of the endonuclease may be fused with at least a portion or all of the integrase.
[0069] In some embodiments, a system comprising at least one endonuclease is described herein. In some embodiments, the endonuclease is a programmable endonuclease that can form a complex with a guide nucleic acid described herein and be directed to a genomic locus by the guide nucleic acid. The endonuclease can bind to DNA. In some embodiments, the endonuclease is an RNA guide endonuclease. In some embodiments, the endonuclease can introduce single-strand breaks. Examples of RNA guide endonucleases include CRISPR / Cas endonucleases (e.g., class 2 CRISPR / Cas endonucleases, e.g., type II, type V, or type VI CRISPR / Cas endonucleases). CRISPR / Cas endonucleases are also called CRISPR / Cas effector polypeptides. The appropriate endonuclease is a CRISPR / Cas endonuclease (e.g., class 2 CRISPR / Cas endonucleases, e.g., type II, type V, or type VI CRISPR / Cas endonucleases). In some cases, the appropriate RNA guide endonuclease is a class 2 CRISPR / Cas endonuclease. In some cases, the appropriate RNA guide endonuclease is a class 2 type II CRISPR / Cas endonuclease (e.g., Cas9 protein). In some cases, the endonuclease includes a class 2 type V CRISPR / Cas endonuclease (e.g., Cpf1 protein, C2c1 protein, or C2c3 protein). In some cases, the appropriate RNA guide endonuclease is a class 2 type VI CRISPR / Cas endonuclease (e.g., C2c2 protein; also called "Cas13a" protein). The CasX protein is also suitable for use. The CasY protein is also suitable for use. In some embodiments, the endonuclease may include one of the Cass described herein, which is complexed with a guide nucleic acid (e.g., gRNA) as an RNP complex.
[0070] In some cases, the endonuclease is a type II CRISPR / Cas endonuclease. In other cases, the endonuclease is Cas9. Cas9 functions as an RNA-guided endonuclease for target recognition and cleavage via a mechanism in which the two nuclease active sites in Cas9 can combine to produce double-strand DNA breaks (DSBs) or individually produce single-strand DNA breaks (SSBs) using a dual guide RNA having crRNA and trans-activating crRNA (tracrRNA). The type II CRISPR endonuclease Cas9 and the manipulated dual guide RNA (dgRNA) or single guide RNA (sgRNA) form a ribonucleoprotein (RNP) complex that can target a desired DNA sequence. Guided by the dual RNA complex or chimeric single-strand guide RNA, Cas9 generates site-specific DSBs or SSBs within the double-strand DNA (dsDNA) target nucleic acid, which are repaired by either non-homologous end joining (NHEJ) or homologous recombination (HDR). Cas9 can induce RNA by associating with its RNA-binding segment, thereby guiding it to a target site within a target nucleic acid sequence (e.g., stabilized at the target site). The Cas9 protein can bind to and / or modify (e.g., cleavage, nicks, methylating agents, demethylating agents, etc.) target nucleic acids and / or polypeptides associated with them (e.g., histone tail methylation or acetylation; e.g., if the Cas9 protein contains an active fusion partner). In some cases, the Cas9 protein is a naturally occurring protein (e.g., naturally occurring in bacterial and / or archaeal cells). In other cases, the Cas9 protein is not a naturally occurring polypeptide (e.g., the Cas9 protein is a mutant Cas9 protein, a chimeric protein, etc.).
[0071] Naturally occurring Cas9 proteins can bind to Cas9 guide RNA, thereby being directed to specific sequences within a target nucleic acid (target site) and cleaving the target nucleic acid (e.g., cleaving dsDNA to produce double-strand breaks, cleaving ssDNA, cleaving ssRNA, etc.). Chimeric Cas9 proteins may include fusion proteins containing a Cas9 polypeptide fused to a heterologous protein (e.g., one not provided by the Cas9 protein), where the heterologous protein provides activity. The fusion partner can provide activity, such as enzymatic activity (e.g., nuclease activity, activity for DNA and / or RNA methylation, activity for DNA and / or RNA cleavage, activity for histone acetylation, activity for histone methylation, activity for RNA modification, activity for RNA binding, activity for RNA splicing, etc.). In some cases, a portion of the Cas9 protein (e.g., the RuvC domain and / or HNH domain) exhibits reduced nuclease activity compared to the corresponding portion of the wild-type Cas9 protein (e.g., in some cases, the Cas9 protein is a nickase). In some cases, the Cas9 protein is enzymatically inactive or has reduced enzymatic activity compared to the wild-type Cas9 protein (e.g., compared to Streptococcus pyogenes Cas9). In some cases, Cas9 is a Cas9 nickase. Cas9 nickases can be produced by mutating the Cas9 nuclease domain. Non-limiting examples of Cas9 nickases include SpCas9, SaCas9, CjCas9, GeoCas9, HpaCas9, and NmeCas9. In some embodiments, the endonucleases described herein include one of the Cas9s listed in Table 1. In some embodiments, the endonucleases described herein contain a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to any one of the Cas9 polypeptide sequences in Table 1.
[0072] [Table 1]
[0073] [Table 2]
[0074] [Table 3]
[0075] [Table 4]
[0076] [Table 5]
[0077] [Table 6]
[0078] [Table 7]
[0079] [Table 8]
[0080] Some embodiments include endonucleases such as RNA-guided endonucleases. RNA-guided endonucleases may include class II CRISPR / Cas endonucleases. RNA-guided endonucleases may include Cas9 endonucleases. RNA-guided endonucleases may include nickases. RNA-guided endonucleases may include an amino acid sequence that is at least 80% identical to any one of the amino acid sequences of SEQ ID NOs. 1 to 13, or a functional fragment thereof.
[0081] Endonucleases can introduce single-strand breaks into target nucleic acids. Endonucleases can introduce single-strand breaks into target nucleic acids without cleaving the opposite strand of the single-strand break. Endonucleases may include nickases. In some cases, endonucleases may exclude endonucleases that introduce double-strand breaks. Endonucleases may exclude restriction enzymes.
[0082] An endonuclease may be included as part of the fusion protein. In some cases, the endonuclease is a fusion protein fused to a heterologous polypeptide, such as a heterologous ligase as described herein. The heterologous polypeptide may include a fusion partner. The fusion protein may include a fusion partner such as a DNA ligase, a nuclear localization signal, a chromatin modification domain, a cell-permeable peptide, or a tag polypeptide. The fusion protein may include one or more fusion partners. The fusion protein may include a ligase. The fusion protein may include a nuclear localization signal, a chromatin modification domain, a cell-permeable peptide, or a tag polypeptide.
[0083] The fusion partner can be linked to the N-terminus of the endonuclease. The fusion partner can be linked to the C-terminus of the endonuclease. The endonuclease can be linked to a linker at its N-terminus or C-terminus. The fusion partner can be linked by its N-terminus or C-terminus. The fusion partner can be linked to the endonuclease by its N-terminus. The fusion partner can be linked to the endonuclease by its C-terminus. The fusion partner can be linked to a linker at its N-terminus or C-terminus.
[0084] In some cases, endonucleases contain linkers, which covalently connect the endonuclease to a heterologous polypeptide. The linker can link the endonuclease to any fusion partner. The linker can also link any fusion partner to another fusion partner. Linker polypeptides can have any of a variety of amino acid sequences. Proteins can be linked by spacer peptides of a generally flexible nature, but other chemical bonds are not excluded. Suitable linkers include polypeptides of 4 to 40 amino acids or 4 to 25 amino acids in length. These linkers can be produced by coupling proteins using oligonucleotides encoding synthetic linkers, or they can be encoded by nucleic acid sequences encoding fusion proteins. Peptide linkers with a certain degree of flexibility can be used. Note that preferred linkers generally have sequences that result in flexible peptides, but linked peptides can have substantially any amino acid sequence. The use of small amino acids such as glycine and alanine is useful for creating flexible peptides. Creating such sequences is routine for those skilled in the art. A variety of different linkers are commercially available and considered suitable for use. Examples of linker polypeptides include glycine polymers (G)n, glycine-serine polymers (e.g., (GS)n (SEQ ID NO: 960), (GSGGS)n (SEQ ID NO: 950), (GGSGGS)n (SEQ ID NO: 951), and (GGGS)n (SEQ ID NO: 952), where n is at least an integer of 1); glycine-alanine polymers; and alanine-serine polymers. Exemplary linkers may include, but are not limited to, amino acid sequences such as GGSG (SEQ ID NO: 954), GGSGG (SEQ ID NO: 955), GSGSG (SEQ ID NO: 956), GSGGG (SEQ ID NO: 957), GGGSG (SEQ ID NO: 958), GSSSG (SEQ ID NO: 959), etc. A linker having the sequence (GGGGS)n (SEQ ID NO: 953) is also suitable, where n is an integer from 1 to 10 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10).Those skilled in the art will recognize that the design of a peptide conjugated to any desired element may include a linker that is all or partially flexible, and as a result, the linker may include a flexible linker as well as one or more parts that give a less flexible structure.
[0085] One or more linkers may be included in the fusion protein. A range of linkers may be included in the fusion protein, defined by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 linkers, or any two of the aforementioned integers. A linker may be attached to the N-terminus of at least some of the endonucleases. A linker may be attached to the N-terminus of at least some of the fusion partner. A linker may be attached to the N-terminus of at least some of the fusion ligases. A linker may be attached to the N-terminus of a nuclear localization signal. A linker may be attached to the N-terminus of a chromatin modification domain. A linker may be attached to the N-terminus of a cell-permeable peptide. A linker may be attached to the N-terminus of a tag polypeptide. A linker may be attached to the C-terminus of at least some of the endonucleases. A linker may be attached to the C-terminus of at least some of the fusion partner. A linker may be attached to the C-terminus of at least some of the fusion ligases. A linker may be attached to the C-terminus of a nuclear localization signal. The linker can be attached to the C-terminus of a chromatin modification domain. The linker can be attached to the C-terminus of a cell-permeable peptide. The linker can be attached to the C-terminus of a tagged polypeptide.
[0086] The linker may contain several or a range of amino acids or residues. The linker may contain at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 amino acid residues. The linker may, in some embodiments, contain 1 or fewer, 2 or fewer, 3 or fewer, 4 or fewer, 5 or fewer, 6 or fewer, 7 or fewer, 8 or fewer, 9 or fewer, 10 or fewer, 12 or fewer, 13 or fewer, 14 or fewer, 15 or fewer, 20 or fewer, 25 or fewer, 30 or fewer, 35 or fewer, 40 or fewer, 45 or fewer, 50 or fewer, 55 or fewer, 60 or fewer, 65 or fewer, 70 or fewer, 75 or fewer, 80 or fewer, 85 or fewer, 90 or fewer, 95 or fewer, or 100 or fewer amino acid residues. The linker may contain 1 to 10 amino acids, 1 to 25 amino acids, or 1 to 100 amino acids.
[0087] Linkers may be included anywhere in the polypeptide chain or protein described herein. For example, a linker can separate an endonuclease from a ligase. A linker can separate an endonuclease from a nuclear localization signal, a chromatin modification domain, a cell-permeable peptide, or a tagged polypeptide.
[0088] In some cases, the endonuclease includes a nuclear localization sequence (e.g., one or more nuclear localization signals or NLSs for targeting the nucleus). In some embodiments, the NLS described herein includes any one of the NLSs in Table 2. In some embodiments, the NLS described herein includes a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to any one of the NLSs in Table 2.
[0089] [Table 9]
[0090] Polynucleotides encoding NLS polypeptides may be used. An example of such a polynucleotide may be SGGSx2-bpNLS-SGGSx2:TCCGGCGGAAGCTCTGGTGGCAGCAAGCGGACCGCCGACGGCTCTGAATTCGAGAGCCCTAAGAAGAAAAGAAAGGTGAGCGGAGGCTCTAGCGGCGGAAGC (SEQ ID NO: 25).
[0091] In some embodiments, the endonuclease contains a dimerizing domain. The dimerizing domain may be located at the N-terminus or C-terminus of the endonuclease. In some embodiments, the dimerizing domain allows the endonuclease to form a heterodimer with another polypeptide (e.g., a heterologous ligase). In some embodiments, the dimerizing domain allows the endonuclease to be functionally coupled with another polypeptide. Non-limiting examples of dimerizing domains include leucine zipper, FKBP, FRB, calcineurin A, CyP-Fas, GyrB, GAI, GID1, SNAP tag, Halo tag, Bcl-xL, Fab, LOV domain, or SpyTag / SpyCatcher. Other examples of dimerization domains include heavy chain domain 2 (CH2) of IgM (MHD2) or IgE (EHD2), immunoglobulin Fc region, heavy chain domain 3 (CH3) of IgG or IgA, heavy chain domain 4 (CH4) of IgM or IgE, antibodies such as Fab, Fab2, leucine zipper motif, Vernus-Buster dimer, mini-antibody, or ZIP mini-antibody. In some embodiments, the dimerization domains described herein include any one of the dimerization domains in Table 3. In some embodiments, the dimerization domains described herein include polypeptide sequences that are at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the dimerization domains in Table 3.
[0092] [Table 10]
[0093] In some embodiments, the endonuclease comprises at least one additional domain. In some embodiments, the at least one additional domain is a functional domain. For example, the functional domain may include a chromatin modification domain or a cell-permeable peptide. In some embodiments, the chromatin modification domain described herein comprises any one of the chromatin modification domains in Table 4. In some embodiments, the chromatin modification domain described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the chromatin modification domains in Table 4.
[0094] [Table 11]
[0095] In some embodiments, the cell-permeable peptides described herein comprise any one of the cell-permeable peptides in Table 5. In some embodiments, the cell-permeable peptides described herein comprise a polypeptide sequence that is identical to any one of the cell-permeable peptides in Table 5 by at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more.
[0096] [Table 12]
[0097] In some embodiments, the endonuclease comprises a tag, which can be used to increase, identify, or purify the expression of the endonuclease. In some embodiments, the tags described herein comprise one of the tag sequences in Table 6. In some embodiments, the tags described herein comprise a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to one of the tag sequences in Table 6.
[0098] [Table 13]
[0099] In some embodiments, an endonuclease may be expressed as a split construct as one or more exteins fused to one or more inteins. Intein technology can be used to deliver large proteins into cells by expressing the protein as two or more shorter peptide segments (exteins). Each extein may be expressed as a fusion with an intein peptide (e.g., NpuC intein or NpuN intein). An intein may autocatalyze the fusion of two or more exteins and autocatalyze the excision of intein from its corresponding extein. The result may be a protein complex lacking intein, containing a first extein fused to a second extein. The intein may be located at the N-terminus of the extein, or it may be located at the C-terminus of the extein. The extein may contain a cysteine residue located adjacent to the intein (e.g., an extein with an intein fused to its C-terminus at the C-terminus of the extein). Cas nickases may be expressed as two or more segments. The first Cas nickase segment may include the N-terminal portion of Cas nickase. The first segment of Cas nickase may include the first intein. The second segment of Cas nickase may include the C-terminal portion of Cas nickase. The second segment of Cas nickase may include the second intein. The intein may be fused to the C-terminus of the N-terminal portion of Cas nickase. The intein may be fused to the N-terminus of the C-terminal portion of Cas nickase. The nucleic acid sequence encoding the exstein-intein fusion may be compatible with a delivery vector (e.g., an adeno-associated virus (AAV) vector).
[0100] DNA ligase This specification discloses ligases. Ligases may be DNA ligases or may comprise DNA ligases. Ligases may be included in compositions, systems, or methods disclosed herein. Ligases may be recombinant. Ligases may be bound to endonucleases. Ligases may be directly or indirectly linked to endonucleases. The coupling may be covalent or non-covalent. Ligases may be bound to or linked to endonucleases. Ligases may be part of a fusion protein with endonucleases or may be used together with endonucleases. Ligases may be coupled to integrases. Ligases may be directly or indirectly linked to integrases. The coupling may be covalent or non-covalent. Ligases may be bound to or linked to integrases. Ligases may be recruited to be part of a fusion protein with integrases or may be used together with integrases. Ligases may be heterogeneous. Ligases may be endogenous. Where heterologous ligases are described, non-heterologous (e.g., endogenous) ligases may be used in some cases. Ligases may be encoded intracellularly. Ligases may be delivered to cells in trans. Ligases may form a phosphodiester bond by ligating two nucleic acid ends together. Ligases may ligate the end of a target nucleic acid (e.g., the 5' or 3' end) to an exogenous first integrated nucleic acid (e.g., the 3' or 5' end of an exogenous first integrated nucleic acid). Ligates the exogenous first integrated nucleic acid (e.g., donor nucleic acid) to a cleaved or nicked end of the target nucleic acid, which is produced by an endonuclease such as an RNA guide endonuclease. Ligases may include any embodiment shown in Figures 1A to 6C.
[0101] Ligases may not be naturally occurring. Ligases may be manipulated. Ligases may be synthesized. Ligases may be pre-synthesized. Ligases may be added to a subject or cell. Ligases may be encoded by nucleic acids. Encoding nucleic acids may be manipulated, synthesized, or added to a subject or cell.
[0102] At least a portion of the ligase may be contained in the first polypeptide. At least a portion of the ligase may be contained in the second polypeptide. The ligase may be split into two polypeptides that are bound together. The first polypeptide may contain the N-terminal portion of the ligase. The first polypeptide may contain the C-terminal portion of the ligase. The second polypeptide may contain the N-terminal portion of the ligase. The second polypeptide may contain the C-terminal portion of the ligase. The first or second polypeptide containing a portion of the ligase may be fused with at least a portion or all of the endonuclease. The first or second polypeptide containing a portion of the ligase may be fused with at least a portion or all of the integrase.
[0103] Examples of DNA ligases include hLIG1, T4 ligase, T7 ligase, and ligases derived from Aquifex aeolicus VF5, Neisseria meningitidis serogroup A strain Z2491, Neisseria meningitidis serogroup B strain MC58, Pseudomonas aeruginosa PA01, Vibrio cholerae (El Tor N1696), Vaccinia virus, and Emiliania huxleyi virus.
[0104] Ligase may include a ligase capable of ligating a DNA-containing substrate. In some embodiments, ligase includes a ligase capable of ligating a substrate containing a DNA sprint. For example, a DNA ligase may ligate a 5' phosphate to the 3' hydroxyl of two DNA strands that hybridize into another DNA strand. The sprint DNA strand may contain an RNA portion. For example, a DNA ligase may ligate a 5' phosphate to the 3' hydroxyl of two DNA strands that hybridize across the DNA portion of an RNA / DNA hybrid strand. In some embodiments, ligase includes a ligase capable of ligating a DNA / RNA-containing substrate. In some embodiments, ligase includes a ligase capable of ligating a substrate containing an RNA sprint. For example, a DNA ligase may ligate a 5' phosphate to the 3' hydroxyl of two DNA strands that hybridize into an RNA strand. The RNA strand may contain a DNA portion. For example, a DNA ligase may ligate a 5' phosphate group to the 3' hydroxyl groups of two DNA strands that hybridize across the RNA portion of an RNA / DNA hybrid strand.
[0105] In some embodiments, the ligase described herein comprises one of the ligases in Table 7. In some embodiments, the ligase described herein comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of one of the ligases in Table 7.
[0106] [Table 14]
[0107] [Table 15]
[0108] [Table 16]
[0109] Table 17
[0110] Table 18
[0111] Table 19
[0112] Table 20
[0113] Table 21
[0114] Table 22
[0115] Table 23
[0116] Table 24
[0117] Table 25
[0118] Table 26
[0119] [Table 27]
[0120] [Table 28]
[0121] [Table 29]
[0122] [Table 30]
[0123] [Table 31]
[0124] Some embodiments include a DNA ligase that ligates a DNA strand that forms a base pair with a DNA sprint. In some embodiments, the DNA ligase ligates a DNA strand that forms a base pair with an RNA sprint. In some embodiments, the DNA ligase includes an amino acid sequence, or a functional fragment thereof, that is at least 80% identical to any one of the amino acid sequences of SEQ ID NOs. 55 to 96.
[0125] In some embodiments, the ligase comprises at least one NLS (e.g., any one of the NLS in Table 2). In some embodiments, the ligase comprises at least one further domain. In some embodiments, the at least one further domain is a dimerization domain (e.g., any one of the dimerization domains in Table 3). In some embodiments, the ligase comprising the dimerization domain can be dimerized with an endonuclease to form a heterodimer. In some embodiments, the at least one additional domain is a functional domain. For example, the functional domain may comprise a chromatin modification domain (e.g., any one of the chromatin modification domains in Table 4) or a cell-permeable peptide (e.g., any one of the cell-permeable peptides in Table 5). In some embodiments, the ligase comprises a linker that can covalently link the ligase to another polypeptide (e.g., an endonuclease). In some embodiments, the linker covalently links the ligase to at least one further domain. In some embodiments, the ligase includes a tag (e.g., any one of the tags in Table 6), which can be used to increase, identify, or purify the ligase expression. A linker can isolate the ligase from nuclear localization signals, chromatin modification domains, cell-permeable peptides, or tag polypeptides. Any linker described herein may be included.
[0126] Ligases may contain binding motifs for binding to nucleic acid motifs (e.g., hairpin motifs). In some embodiments, ligases (e.g., DNA ligases) contain an MS2 coat protein (MCP) peptide. Ligases may contain hairpin binding motifs such as the MCP peptide. The MCP peptide may be useful for recruiting the ligase to a guide nucleic acid containing an MS2 hairpin. The advantage of using the MCP peptide and MS2 hairpin is that the ligase and endonuclease, e.g., Cas niccasase (or a part thereof), can be separated and fitted into a separate vector such as an AAV vector. In some embodiments, ligases contain a loop region. In some embodiments, the loop region is a 2a loop or a 3a loop. The loop region may contain a 2a loop. The loop region may contain a 3a loop.
[0127] Integrase Integrases are disclosed herein. An integrase may be an example of a recombinase, and where an integrase is described, a recombinase may be intended. An integrase may be or may comprise a phage integrase. An integrase may be or may comprise a site-specific recombinase. An integrase may be or may comprise a serine integrase. An integrase may be a resol-based or may comprise a resol-based. An integrase may be a DNA invertase or may comprise a DNA invertase. An integrase may be a tyrosine integrase or may comprise a tyrosine integrase. An integrase may be a retrotransposase or may comprise a retrotransposase. Integrases may be included in compositions, systems, or methods disclosed herein. Integrases may be recombinant. Any of these integrases may be modified or mutated. For example, an integrase may include insertions, deletions, or active or functional fragments. Integrases may include mutations of the integrases described herein. Integrases may be coupled to endonucleases. Integrases may be coupled directly or indirectly to endonucleases. The coupling may be covalent or non-covalent. Integrases may be bound to or ligated to endonucleases. Integrases may be part of a fusion protein with an endonuclease or may be used together with an endonuclease. Integrases may be coupled to ligases. Integrases may be linked directly or indirectly to ligases. The coupling may be covalent or non-covalent. Integrases may be bound to or ligated to ligases. Integrases may be part of a fusion protein with a ligase or may be used together with a ligase. Integrases may be heterologous. Integrases may be endogenous. Where heterologous integrases are described, non-heterologous (e.g., endogenous) integrases may be used in some cases.Integrases can be encoded within cells. Integrases can be delivered to cells in trans. Integrases introduce a second integrated nucleic acid into a target nucleic acid by recognizing and binding to an exogenous first integrated nucleic acid (e.g., a recombinant sequence).
[0128] Integrases may not be naturally occurring. Integrases may be manipulated. Integrases may be synthetic. Integrases may be pre-synthesized. Integrases may be added to a subject or cell. Integrases may be encoded by nucleic acids. Encoding nucleic acids may be manipulated, synthesized, or added to a subject or cell.
[0129] At least a portion of the integrase may be contained in the first polypeptide. At least a portion of the integrase may be contained in the second polypeptide. The integrase may be split into two polypeptides that are bound together. The first polypeptide may contain the N-terminal portion of the integrase. The first polypeptide may contain the C-terminal portion of the integrase. The second polypeptide may contain the N-terminal portion of the integrase. The second polypeptide may contain the C-terminal portion of the integrase. The first or second polypeptide containing a portion of the integrase may be fused with at least a portion or all of the endonuclease. The first or second polypeptide containing a portion of the integrase may be fused with at least a portion or all of the ligase.
[0130] In some embodiments, the integrases described herein may be serine integrases. Serine integrases may be referred to as "serine recombinases." Examples of species containing serine recombinases are as follows: Bacillus cereus; Bacillus safensis; Bacillus tropicus; Burkholderia multivorans; Burkholderia ubonensis; Cellulosimicrobium cellulans; Clostridium botulinum; Clostridioides difficile; Clostridium perfringens; Clostridium thermobutyricum; Cronobacter sakazakii; Desulfotomaculum nigrificus * Enterococcus nigrificans*; *Enterococcus clostridioformis*; *Enterococcus faecalis*; *Enterococcus faecium*; *Escherichia coli*; *Eubacterium maltosivorans*; *Faecalibacterium prausnitzii*; *Fusobacterium mortiferum*; *Klebsiella pneumoniae*; *Mycobacteroides abscessus*; *Mycobacterium phage* Bxb1;Mycolicibacterium elephantis; Neobacillus mesonae; Nocardia otitidiscaviarum; Paenibacillus campinasensis; Paeniclostridium sordellii; Parageobacillus caldoxylosilyticus; Pseudomonas aeruginosa; Pseudomonas fluorescens; Pseudomonas fulva; Pseudomonas putida; Pseudomonas syringae; Prochlorothrix hollandica hollandica); Rhizobiales bacterium; Rhodococcus hoagie; Ruminococcus lactaris; Salinispora pacifica; Staphylococcus arlettae; Streptococcus equinus; Staphylococcus hominis; Streptococcus agalactiae; Streptococcus equinus; Streptococcus mitis; Streptomyces ipomoeae; Streptomyces phage (phage)PhiC31; Tenacibaculum dicentrarchi;Treponema denticola; Vibrio harveyi; Vibrio hyugaensis; Vibrio parahaemolyticus.
[0131] In some embodiments, the integrase described herein may be from or derived from the species Bacillus cereus. In some embodiments, the integrase described herein may be from or derived from the species Bacillus safensis. In some embodiments, the integrase described herein may be from or derived from the species Bacillus tropicus. In some embodiments, the integrase described herein may be from or derived from the species Burkholderia multivorans. In some embodiments, the integrase described herein may be from or derived from the species Burkholderia ubonensis. In some embodiments, the integrase described herein may be from or derived from the species Cellulosimicrobium cellulans. In some embodiments, the integrase described herein may be from or derived from the species Clostridium botulinum. In some embodiments, the integrases described herein may be from or derived from the species Clostridioides difficile.In some embodiments, the integrase described herein may be from or derived from the species Clostridium perfringens. In some embodiments, the integrase described herein may be from or derived from the species Clostridium thermobutyricum. In some embodiments, the integrase described herein may be from or derived from the species Cronobacter sakazakii. In some embodiments, the integrase described herein may be from or derived from the species Desulfotomaculum nigrificans. In some embodiments, the integrase described herein may be from or derived from the species Enterocloster clostridioformis. In some embodiments, the integrase described herein may be from or derived from the species Enterococcus faecalis. In some embodiments, the integrases described herein may be from or derived from the species Enterococcus faecium.In some embodiments, the integrases described herein may be from or derived from the species Escherichia coli. In some embodiments, the integrases described herein may be from or derived from the species Eubacterium maltosivorans. In some embodiments, the integrases described herein may be from or derived from the species Faecalibacterium prausnitzii. In some embodiments, the integrases described herein may be from or derived from the species Fusobacterium mortiferum. In some embodiments, the integrases described herein may be from or derived from the species Klebsiella pneumoniae. In some embodiments, the integrases described herein may be from or derived from the species Mycobacteroides abscessus. In some embodiments, the integrases described herein may be from or derived from the Mycobacterium phage Bxb1 species.In some embodiments, the integrase described herein may be from or derived from the species Mycolicibacterium elephantis. In some embodiments, the integrase described herein may be from or derived from the species Neobacillus mesonae. In some embodiments, the integrase described herein may be from or derived from the species Nocardia otitidiscaviarum. In some embodiments, the integrase described herein may be from or derived from the species Paenibacillus campinasensis. In some embodiments, the integrase described herein may be from or derived from the species Paeniclostridium sordellii. In some embodiments, the integrase described herein may be from or derived from the species Parageobacillus caldoxylosilyticus. In some embodiments, the integrases described herein may be from or derived from the species Pseudomonas aeruginosa.In some embodiments, the integrase described herein may be from or derived from the species Pseudomonas fluorescens. In some embodiments, the integrase described herein may be from or derived from the species Pseudomonas fulva. In some embodiments, the integrase described herein may be from or derived from the species Pseudomonas putida. In some embodiments, the integrase described herein may be from or derived from the species Pseudomonas syringae. In some embodiments, the integrase described herein may be from or derived from the species Prochlorothrix hollandica. In some embodiments, the integrase described herein may be from or derived from the species Rhizobiales bacterium. In some embodiments, the integrase described herein may be from or derived from the species Rhodococcus hoagie. In some embodiments, the integrase described herein may be from or derived from the species Ruminococcus lactaris.In some embodiments, the integrase described herein may be from or derived from the species Salinispora pacifica. In some embodiments, the integrase described herein may be from or derived from the species Staphylococcus arlettae. In some embodiments, the integrase described herein may be from or derived from the species Streptococcus equinus. In some embodiments, the integrase described herein may be from Staphylococcus. In some embodiments, the integrase described herein may be from or derived from the species Staphylococcus hominis. In some embodiments, the integrase described herein may be from or derived from the species Streptococcus agalactiae. In some embodiments, the integrase described herein may be from or derived from the species Streptococcus equinus. In some embodiments, the integrase described herein may be from or derived from the species Streptococcus mitis. Streptomyces ipomoeae. In some embodiments, the integrase described herein may be from or derived from the species Streptomyces phage PhiC31. In some embodiments, the integrase described herein may be from or derived from the species Tenacibaculum dicentrarchi. In some embodiments, the integrase described herein may be from or derived from that species. In some embodiments, the integrase described herein may be from or derived from the species Treponema denticola.In some embodiments, the integrase described herein may be from or derived from the species Vibrio harveyi. In some embodiments, the integrase described herein may be from or derived from the species Vibrio hyugaensis. In some embodiments, the integrase described herein may be from or derived from the species Vibrio parahaemolyticus.
[0132] In some embodiments, the integrases described herein include any one of the integrases in Table 8. In some embodiments, the integrases described herein include a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8.
[0133] [Table 32]
[0134] [Table 33]
[0135] [Table 34]
[0136] [Table 35]
[0137] Table 36
[0138] Table 37
[0139] Table 38
[0140] Table 39
[0141] Table 40
[0142] Table 41
[0143] Table 42
[0144] Table 43
[0145] Table 44
[0146] Table 45
[0147] Table 46
[0148] Table 47
[0149] Table 48
[0150] Table 49
[0151] Table 50
[0152] Table 51
[0153] Table 52
[0154] Table 53
[0155] Table 54
[0156] Table 55
[0157] Table 56
[0158] Table 57
[0159] Table 58
[0160] Table 59
[0161] Table 60
[0162] Table 61
[0163] Table 62
[0164] Table 63
[0165] In some embodiments, the integrase includes any integrase that recognizes and binds to the recombinant sequences in Table 13. In some embodiments, the integrase includes any integrase that recognizes and binds to the recombinant sequences in Table 13. In some embodiments, the recombinant sequence includes a sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the recombinant sequences in Table 13. In some embodiments, the integrase described herein includes a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of an integrase that recognizes and binds to any one of the polypeptide sequences of the recombinant sequences in Table 13.
[0166] In some embodiments, the integrases described herein may be tyrosine integrases. Tyrosine integrases may be referred to as "tyrosine recombinases." Examples of tyrosine integrases include BS codV;BS ripX;BS ydcL;CB tnpA;Col1D;CP4;Cre;D29;DLP12;DN int;EC FimB;EC FimE;EC orf;EC xerC;EC xerD;Φ11;Φ13;Φ80;Φadh;ΦCTX;ΦLC3;FLP;ΦR73;HI orf;HI rci;HI xerC;HI xerD;HK22;HP1;L2;L5;L54;λ;LL orf;LL xerC;LO L5;MJ orf;MP int;MT int;MT orf;MV4;P186;P2;P21;P22;P4;P434;PA sss;PM fimB;pAE1;pCL1;pKD1;pMEA;pSAM2;pSB2;pSB3;pSDL2;pSE101;pSE211;pSM1;pSR1;pWS58;R721;Rci;sF6;SLP1;SM orf;SsrA;SSV1;T12;Tn21;Tn4430;Tn554a;Tn554b;Tn7;Tn916;Tuc;WZ int;XisA; or XisC may be listed.
[0167] Examples of species containing tyrosine recombinase include Bacillus subtilis; Clostridium butyricum; Escherichia coli, Mycobacterium smegmatis, Dichelobacter nodosus; Staphylococcus aureus; E. coli phage; Lactobacillus gasseri; Pseudomonas aeruginosa; Lactococcus lactis; Saccharomyces cerevisiae; Haemophilus influenzae; and Mycoplasma sp.); Mycobacterium tuberculosis; Lactobacillus leichmannii; Leuconostoc oenos; Methanococcus jannaschi; Mycobacterium leprae; Mycobacterium paratuberculosis; Lactobacillus delbrueckii; Salmonella typhimurium; Pseudomonas aeruginosa; Proteus mirabilis; Alcaligenes eutrophus; Chlorobium limicola; Kluyveromyces lactis lactis); Amycolatopsis methanolica; Streptomyces ambofaciencs; Zygosaccharomyces bailii; Zygosaccharomyces bisporus; Salmonella dublin; Saccharopolyspora erythraea; Zygosaccharomyces fermentati; Zygosaccharomyces rouxii; Shigella flexneri; Streptomyces coelicolor; Serratia marcescens marcescens); Methanosarcina acetivorans; Sultolobus sp.Possible examples include Streptococcus pyogenes, Bacillus thurinigiensis, Enterococcus faecalis, Lactobacillus lactis, Weeksella zoohelcum, or Anabaena sp.
[0168] In some embodiments, the integrase described herein may be a gamma delta lysol base derived from a Tn1000 transposon. In some embodiments, the integrase described herein may be a Hin recombinase. In some embodiments, the integrase described herein may be a Tn3 lysol base derived from a Tn3 transposon. In some embodiments, the integrase described herein may be a Tre recombinase. In some embodiments, the integrase described herein may be a Dre recombinase. In some embodiments, the integrase described herein may be a Cre recombinase. In some embodiments, the integrase described herein may be a flippase (Flp). In some embodiments, the integrase described herein may be a KD recombinase. In some embodiments, the integrase described herein may be a B2B3 recombinase. In some embodiments, the integrase described herein may be an HK022 integrase. In some embodiments, the integrase described herein may be a ParA integrase. In some embodiments, the integrase described herein may be a Gin integrase. In some embodiments, the integrase described herein may be an R4 recombinase.
[0169] In some embodiments, the integrase described herein may be Vika recombinase. In some embodiments, the integrase described herein may be RDF recombinase. In some embodiments, the integrase described herein may be φBT1 recombinase. In some embodiments, the integrase described herein may be R1 recombinase. In some embodiments, the integrase described herein may be R2 recombinase. In some embodiments, the integrase described herein may be R3 recombinase. In some embodiments, the integrase described herein may be R4 integrase. In some embodiments, the integrase described herein may be R5 integrase. In some embodiments, the integrase described herein may be TP901-1 recombinase. In some embodiments, the integrase described herein may be A118 recombinase. In some embodiments, the integrase described herein may be φFC1 recombinase. In some embodiments, the integrase described herein may be φC1 recombinase. In some embodiments, the integrase described herein may be MR11 recombinase. In some embodiments, the integrase described herein may be TG1 recombinase. In some embodiments, the integrase described herein may be φ370.1 recombinase. In some embodiments, the integrase described herein may be Wβ recombinase. In some embodiments, the integrase described herein may be BL3 recombinase. In some embodiments, the integrase described herein may be SPBc recombinase. In some embodiments, the integrase described herein may be K38 recombinase. In some embodiments, the integrase described herein may be peach recombinase.In some embodiments, the integrase described herein may be Veracruz recombinase. In some embodiments, the integrase described herein may be Rebeuca recombinase. In some embodiments, the integrase described herein may be Theia recombinase. In some embodiments, the integrase described herein may be Benedict recombinase. In some embodiments, the integrase described herein may be KSSJEB recombinase. In some embodiments, the integrase described herein may be PattyP recombinase. In some embodiments, the integrase described herein may be Doom recombinase. In some embodiments, the integrase described herein may be Scaul recombinase. In some embodiments, the integrase described herein may be Lockley recombinase. In some embodiments, the integrase described herein may be Switcher recombinase. In some embodiments, the integrase described herein may be Bob3 recombinase. In some embodiments, the integrase described herein may be Troube recombinase. In some embodiments, the integrase described herein may be abrogate recombinase. In some embodiments, the integrase described herein may be anglefish recombinase. In some embodiments, the integrase described herein may be Sarfire recombinase. In some embodiments, the integrase described herein may be SkiPole recombinase. In some embodiments, the integrase described herein may be Concept II recombinase. In some embodiments, the integrase described herein may be Museum recombinase.In some embodiments, the integrase described herein may be Severus recombinase. In some embodiments, the integrase described herein may be Aeramido recombinase. In some embodiments, the integrase described herein may be Benedict recombinase. In some embodiments, the integrase described herein may be Hinder recombinase. In some embodiments, the integrase described herein may be ICleared recombinase. In some embodiments, the integrase described herein may be Sheen recombinase. In some embodiments, the integrase described herein may be Mundrea recombinase. In some embodiments, the integrase described herein may be BxZ2 recombinase. In some embodiments, the integrase described herein may be φRV recombinase.
[0170] In some embodiments, the integrase described herein may be a retrotransposase encoded by R2. In some embodiments, the integrase described herein may be a retrotransposase encoded by L1. In some embodiments, the integrase described herein may be a retrotransposase encoded by Tol2. In some embodiments, the integrase described herein may be a retrotransposase encoded by Tc1. In some embodiments, the integrase described herein may be a retrotransposase encoded by Tc3. In some embodiments, the integrase described herein may be a retrotransposase encoded by Mariner (Himar 1). In some embodiments, the integrase described herein may be a retrotransposase encoded by Mariner (mos 1). In some embodiments, the integrase described herein may be a retrotransposase encoded by Minos.
[0171] As available herein, Xu et al. have described a method for evaluating integrase activity in Escherichia coli (E. coli) and mammalian cells, confirming that at least R4, φC31, φBT1, Bxb1, SPBc, TP901-1, and Wβ integrases are active on substrates integrated into the genome of HT1080 cells (Xu et al., 2013, Accuracy and efficiency define Bxb1 integrase as the best of fifteen candidate serine recombinases for the integration of DNA into the human genome. BMC Biotechnol. 2013 Oct 20;13:87. Doi:10.1186 / 1472-6750-13-87). Durrant describes novel large serine recombinases (LSRs) that are divided into three classes, distinguished by efficiency and specificity, including landing pad LSRs that are superior to wild-type Bxb1 in episomal and chromosome integration efficiency, LSRs that achieve both efficient site-specific integration and site-specific integration without a landing pad, and multi-targeting LSRs with minimal site specificity. Furthermore, embodiments may include any serine recombinase such as BceINT, SSCINT, SACINT, and INT10 (see Ionnidi et al., 2021; Drag-and-drop genome insertion without DNA cleavage with CRISPR directed integrases.bioRxiv 2021.11.01.466786, doi.org / 10.1101 / 2021.11.01.466786). In some embodiments, the integration site can be selected from the attB site, attP site, attL site, attR site, lox71 site, Vox site, or FRT site.In the examples of this disclosure referring to the Cre-lox system, the Cre-lox system is shown either as a control for programmable gene insertion or as a tool for a different recombinase-mediated event distinct from the insertion of a donor polynucleotide template (or exogenous nucleic acid) into an integrated recognition site.
[0172] In some embodiments, the integrases described herein may be coupled to a recombinant direction factor (RDF). In some embodiments, the integrases described herein may be fused to an RDF. In some embodiments, the integrases described herein may be linked to an RDF. The RDF may include gp3 RDF. The RDF may include gp47 RDF.
[0173] Fusion protein Fusion proteins are disclosed herein. Some embodiments include nucleic acids (e.g., expression vectors) encoding the fusion protein. The fusion protein may include an endonuclease. The fusion protein may include a ligase. The fusion protein may include an integrase. The fusion protein may include a linker. The fusion protein may include two linkers. The fusion protein may include multiple linkers. The endonuclease and ligase may be linked via a linker. The endonuclease and integrase may be linked via a linker. The ligase and integrase may be linked via a linker. The fusion protein may be an example of a covalently bound endonuclease and DNA ligase. The fusion protein may be an example of a covalently bound endonuclease and integrase. The fusion protein may be an example of a covalently bound DNA ligase and integrase. The fusion protein may include an endonuclease such as an RNA-guided endonuclease fused to a DNA ligase. The fusion protein may include an endonuclease, such as an RNA-guided endonuclease, fused to an integrase, such as serine integrase or tyrosine integrase. The fusion protein may also include a DNA ligase fused to an integrase.
[0174] Fusion proteins may not exist naturally. Fusion proteins may be manipulated. Fusion proteins may be synthesized. Fusion proteins may be synthesized beforehand. Fusion proteins may be added to a subject or cell. Fusion proteins may be encoded by nucleic acids. The encoding nucleic acid may be manipulated, synthesized, or added to a subject or cell.
[0175] A fusion protein may be a double fusion protein. A double fusion protein may contain endonucleases such as RNA-guided endonucleases and ligases. A double fusion protein may contain endonucleases such as RNA-guided endonucleases and integrases. A double fusion protein may contain ligases and integrases.
[0176] A double fusion protein comprising an RNA guide endonuclease and a DNA ligase may have one of several orientations. For example, the double fusion protein may include an RNA guide endonuclease upstream (e.g., N-terminal or N-direction) or downstream (e.g., C-terminal or C-direction) of the DNA ligase. The double fusion protein may include an amino (N) terminus of the RNA guide endonuclease relative to the DNA ligase. The double fusion protein may include a carboxyl (C) terminus of the RNA guide endonuclease relative to the DNA ligase. The endonuclease may be amino-oriented within the fusion polypeptide relative to the ligase. The endonuclease may be carboxyl-oriented within the fusion polypeptide relative to the ligase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The ligase may be N-terminal. The ligase may be C-terminal.
[0177] A double fusion protein containing an RNA guide endonuclease and an integrase may have one of several orientations. For example, the double fusion protein may contain an RNA guide endonuclease upstream (e.g., N-terminal or N-direction) or downstream (e.g., C-terminal or C-direction) of the integrase. The double fusion protein may contain the amino (N) terminus of the RNA guide endonuclease for integrase. The double fusion protein may contain the carboxy (C) terminus of the RNA guide endonuclease for integrase. The endonuclease may be amino-oriented within the fusion polypeptide relative to integrase. The endonuclease may be carboxyl-oriented within the fusion polypeptide relative to integrase. The endonuclease may be N-terminal. The endonuclease may be C-terminal. The integrase may be N-terminal. The integrase may be C-terminal.
[0178] A dual fusion protein containing a ligase and an integrase can include one of various orientations. For example, the dual fusion protein can include an integrase upstream (e.g., N-terminal or N-direction) or downstream (e.g., C-terminal or C-direction) relative to the ligase. The dual fusion protein can include an (N)-terminal of the integrase relative to the ligase. The dual fusion protein can include a carboxy (C)-terminal of the integrase relative to the ligase. The integrase can be in the amino direction within the fusion polypeptide relative to the ligase. The integrase can be in the carboxy direction within the fusion polypeptide relative to the ligase. The ligase can be N-terminal. The ligase can be C-terminal. The integrase can be N-terminal. The integrase can be C-terminal.
[0179] The fusion protein can be a triple fusion protein. The triple fusion protein can include an endonuclease such as an RNA-guided endonuclease, a ligase, and an integrase.
[0180] A triple fusion protein comprising an endonuclease, ligase, and integrase may have one of several orientations. For example, a triple fusion protein may include an RNA guide endonuclease upstream (e.g., N-terminal or N-direction) or downstream (e.g., C-terminal or C-direction) of a DNA ligase. A triple fusion protein may include an RNA guide endonuclease amino(N) terminus for a DNA ligase. A triple fusion protein may include a carboxyl(C) terminus for an RNA guide endonuclease for a DNA ligase. The endonuclease may be amino-oriented within the fusion polypeptide relative to the ligase. The endonuclease may be carboxyl-oriented within the fusion polypeptide relative to the ligase. A triple fusion protein may include an RNA guide endonuclease upstream (e.g., N-terminal or N-direction) or downstream (e.g., C-terminal or C-direction) of an integrase. A triple fusion protein may include an RNA guide endonuclease amino(N) terminus for an integrase. A triple fusion protein may contain the carboxyl (C) terminus of an RNA guide endonuclease relative to integrase. The endonuclease may be amino-oriented relative to integrase within the fusion polypeptide. The endonuclease may be carboxyl-oriented relative to integrase within the fusion polypeptide. A triple fusion protein may contain an integrase upstream (e.g., N-terminus or N-direction) or downstream (e.g., C-terminus or C-direction) relative to the ligase. A triple fusion protein may contain the integrase (N) terminus relative to the ligase. A triple fusion protein may contain the integrase carboxyl (C) terminus relative to the ligase. The integrase may be amino-oriented relative to the ligase within the fusion polypeptide. The integrase may be carboxyl-oriented relative to the ligase within the fusion polypeptide. The endonuclease may be N-terminus. The endonuclease may be C-terminus. The ligase may be N-terminus. The ligase may be C-terminus. The integrase may be N-terminus. The integrase may be C-terminus.
[0181] The fusion protein may contain a nuclear localization signal, a chromatin modification domain, a cell-permeable peptide, a tagged polypeptide, or an exonuclease. The fusion protein may contain a nuclear localization signal. The fusion protein may contain a chromatin modification domain. The fusion protein may contain a cell-permeable peptide. The fusion protein may contain a tagged polypeptide. The fusion protein may contain an exonuclease. Any of the nuclear localization signal, chromatin modification domain, cell-permeable peptide, tagged polypeptide, or exonuclease, endonuclease, ligase, or integrase may be linked to another, or directly to an endonuclease, ligase, or integrase. Any of the nuclear localization signal, chromatin modification domain, cell-permeable peptide, tagged polypeptide, or exonuclease, endonuclease, ligase, or integrase may be linked to another, or to an endonuclease, ligase, or integrase by a linker. Multiple linkers may be included in the fusion protein. Fusion proteins may exclude polymerase.
[0182] The linker may include an amino acid linker. The amino acid linker may include residues of a certain length. The length may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 residues, or a range of residues defined by any two of the aforementioned integers. The length may include at least 1 residue, at least 2 residues, at least 3 residues, at least 4 residues, at least 5 residues, at least 6 residues, at least 7 residues, at least 8 residues, at least 9 residues, at least 10 residues, at least 15 residues, at least 20 residues, at least 25 residues, at least 30 residues, at least 40 residues, at least 50 residues, at least 60 residues, at least 70 residues, at least 80 residues, at least 90 residues, or at least 100 residues. In some embodiments, the length may include less than 2 residues, less than 3 residues, less than 4 residues, less than 5 residues, less than 6 residues, less than 7 residues, less than 8 residues, less than 9 residues, less than 10 residues, less than 15 residues, less than 20 residues, less than 25 residues, less than 30 residues, less than 40 residues, less than 50 residues, less than 60 residues, less than 70 residues, less than 80 residues, less than 90 residues, or less than 100 residues. Examples of residues include alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine, or any combination thereof. The linker may be non-enzymatic or may lack any enzymatic activity.
[0183] The connection may be a covalent bond. The covalent bond may include a peptide bond. The peptide bond may include an amide bond. The connection may be between one N-terminus and another N-terminus. The connection may be between one C-terminus and another C-terminus. The connection may be between one N-terminus and another C-terminus. The connection may be between one C-terminus and another N-terminus.
[0184] Fusion proteins may contain junctions in various orientations. Endonucleases can be ligated at their C-terminus. Endonucleases can be ligated at their N-terminus. Ligases can be ligated at their C-terminus. Ligases can be ligated at their N-terminus. Integrases can be ligated at their C-terminus. Integrases can be ligated at their N-terminus.
[0185] Figure 7 shows several examples of fusion proteins containing endonucleases and ligases. The figure includes examples of the arrangement and orientation of endonucleases, linkers, ligases, or nuclear localization signals. Other embodiments may be incorporated into the illustrated examples.
[0186] Figures 13A and 13B show several examples of fusion proteins containing at least two endonucleases, ligases, and integrases. The figures include examples of arrangement and orientation of endonucleases, ligases, or integrases. Figure 13A includes an example of a double fusion protein. Figure 13B includes an example of a triple fusion protein. Other embodiments can be incorporated into the illustrated examples.
[0187] In some embodiments, the fusion protein envisioned herein may include the DNA-binding domain of the Rad51 DNA repair protein (rad51DBD). In some embodiments, the fusion protein may include the high-mobility group nucleosome-binding domain 1 (HN1) and the histone H1 central globular domain (H1G). In some embodiments, the fusion protein may include Rad51DBD, as well as HN1 and H1G.
[0188] In some embodiments, the fusion protein intended herein comprises Brex27. In some embodiments, Brex27 can be fused to an endonuclease. In some embodiments, Brex27 can be fused to nCas9.
[0189] Table 9 provides examples of fusion proteins that may be useful in genome revision systems. In some embodiments, rad51DBD, and / or HN1 and H1G, are fused to the fusion protein. In some embodiments, Brex27 is fused to the fusion protein. The fusion proteins contemplated herein may contain amino acid sequences that are at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the amino acid sequences provided in Table 27.
[0190] [Table 64]
[0191] [Table 65]
[0192] [Table 66]
[0193] [Table 67]
[0194] In some embodiments, the fusion protein envisioned herein may comprise a bisistronic mRNA encoding a phosphorylation-mimicking peptide derived from IGF1 (IGF1pm1) and an N-terminal peptide (IN peptide) derived from NFATC2IP (NFATC2IPp1) peptide.
[0195] In some embodiments, the fusion protein comprises mRNA encoding the IN peptide described in Table 27, which is fused to the fusion protein. The fusion proteins contemplated herein may contain amino acid sequences that are at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the amino acid sequences provided in Table 9.
[0196] Non-covalent proteins Non-covalently coupled proteins are disclosed herein. Some embodiments relate to nucleic acids (e.g., expression vectors) that encode a protein, or at least a portion of a protein. The protein may include an endonuclease such as an RNA guide endonuclease. The protein of a non-covalently coupled protein may include a portion of an endonuclease. The protein of a non-covalently coupled protein may include a portion of a ligase. The protein may include a ligase such as a DNA ligase. The protein of a non-covalently coupled protein may include an integrase. The protein of a non-covalently coupled protein may include a portion of an integrase (e.g., a functional integrase fragment). The protein may include an integrase such as a serine integrase. The protein may include an integrase such as a tyrosine integrase. The protein of a non-covalently coupled protein may include a fusion protein.
[0197] Non-covalently coupled proteins can be bound together via heterodimerization domains. Examples of heterodimerization domains include leucine zippers, PDZ domains, streptavidin, streptavidin-binding proteins, Foldon domains, hydrophobic moieties, or their functional binding fragments. Heterodimerization domains may include leucine zippers. Heterodimerization domains may include PDZ domains. Heterodimerization domains may include streptavidin. Heterodimerization domains may include streptavidin-binding proteins. Heterodimerization domains may include Foldon domains. Heterodimerization domains may include hydrophobic moieties. Heterodimerization domains may include antibodies or antibody fragments. Non-covalently coupled proteins can be bound together via inteins.
[0198] Endonucleases and ligases may be linked to each other by separate molecules. Endonucleases and integrases may be linked to each other by separate molecules. Ligases and integrases may be linked together by separate molecules. The separate molecules may include nucleic acids (e.g., guide nucleic acids). Ligases may include hairpin-binding motifs, and RNA guide endonucleases and DNA ligases are bound to nucleic acids. Ligases may include hairpin-binding motifs, and integrases and DNA ligases are bound to nucleic acids. Nucleic acids may include a scaffold that binds to the RNA guide endonuclease and a hairpin that binds to the hairpin-binding motif. The hairpin-binding motif may include an MS2 coat protein (MCP) peptide. The hairpin may include an MS2 hairpin.
[0199] Endonucleases and ligases can be bound together by heterobifunctional molecules. Endonucleases and integrases can be coupled together by heterobifunctional molecules. Ligases and integrases can be linked together by heterobifunctional molecules. Heterobifunctional molecules may include an endonuclease-binding domain and a DNA ligase-binding domain. Heterobifunctional molecules may include an endonuclease-binding domain and an integrase-binding domain. Heterobifunctional molecules may include a ligase-binding domain and an integrase-binding domain. Heterobifunctional molecules may include an endonuclease-binding domain. The endonuclease-binding domain may include a heterodimerization domain. The endonuclease-binding domain may include an antibody or antibody-binding fragment. Heterobifunctional molecules may include a ligase-binding domain such as a DNA ligase-binding domain. The DNA ligase-binding domain may include a heterodimerization domain. The DNA ligase-binding domain may include an antibody or antibody-binding fragment. Heterobifunctional molecules may contain integrase-binding domains, such as serine integrase or tyrosine integrase-binding domains. Integrase-binding domains may contain heterodimerization domains. Integrase-binding domains may contain antibodies or antibody-binding fragments. Heterobifunctional molecules may contain small molecules. Small molecules may contain proteolytically targeted chimeras (PROTACs) or related heterobifunctional molecules.
[0200] Some embodiments include a protein complex comprising an RNA guide endonuclease conjugated to a DNA ligase. The endonuclease and DNA ligase may be conjugated to each other via heterodimerization domains. The protein complex according to Embodiment 75, wherein the heterodimerization domain may comprise a leucine zipper, a PDZ domain, streptavidin, and a streptavidin-binding protein, a Foldon domain, a hydrophobic polypeptide, an antibody conjugated to Cas nickase, or an antibody conjugated to a DNA ligase, or one or more binding fragments thereof. The protein complex may be contained in a cell. The cell may further contain a heterologous RNA guide endonuclease and a DNA ligase introduced into the cell. The cell may further contain a nuclease different from the RNA guide endonuclease.
[0201] In some embodiments, the protein complex contemplated herein may include monomeric streptavidin (mSA). In some embodiments, monomeric streptavidin can be fused to an endonuclease. In some embodiments, monomeric streptavidin can be fused to Nicking Cas9. In some embodiments, monomeric streptavidin can be fused to a ligase. In some embodiments, monomeric streptavidin can be bound to a biotinylated splint nucleic acid. In some embodiments, the biotin modification is located at the / 5Biosg / terminus of the splint nucleic acid. In some embodiments, the biotin modification is located at the / 3Bio / terminus of the splint nucleic acid.
[0202] Figure 20 shows the guide nucleic acid, endonuclease, ligase, and donor strand at a genomic locus. In this figure, the biotinylated sprint nucleic acid is bound to monomeric streptavidin fused to nickeling Cas9. Monomeric streptavidin can also be fused to a ligase (not shown).
[0203] Table 10 provides non-limiting examples of fusion proteins containing mSAs as intended herein. A fusion protein may be any fusion protein containing the amino acid sequences provided in Table 10. A fusion protein as intended herein may contain amino acid sequences that are at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the amino acid sequences provided in Table 10.
[0204] [Table 68]
[0205] [Table 69]
[0206] [Table 70]
[0207] Guide nucleic acids Guide nucleic acids are disclosed herein. Guide nucleic acids may be included in the compositions, systems, or methods disclosed herein. Some embodiments relate to nucleic acids (e.g., DNA or expression vectors) encoding guide nucleic acids, such as guide RNA. Guide nucleic acids (e.g., gRNA) that direct programmable endonucleases (e.g., nCas9) to target nucleic acids (e.g., genomic loci) are provided herein. Guide nucleic acids can guide RNA guide endonucleases to target nucleic acid loci for nucleic acid substitution or gene editing at the locus. Guide nucleic acids of this disclosure may facilitate the insertion of a donor strand into a target site of the target nucleic acid. Guide nucleic acids of this disclosure may facilitate the editing of the nucleic acid sequence at a target site of the target nucleic acid. Guide nucleic acids may also act as sprints for DNA ligases described herein, such as ligating two nucleic acid strands that form a base pair with a portion of the guide nucleic acid. Guide nucleic acids may be single-stranded. Guide nucleic acids may include RNA. Guide nucleic acids may include guide RNA (gRNA). In some cases, the guide nucleic acid may contain DNA.
[0208] Guide nucleic acids may not be naturally occurring. Guide nucleic acids may be manipulated. Guide nucleic acids may be synthesized. Guide nucleic acids may be pre-synthesized. Guide nucleic acids may be added to a subject or cell. In some embodiments, guide nucleic acids do not contain a polymerase template.
[0209] The guide nucleic acid may include an exogenous first integrated nucleic acid binding site. This exogenous first integrated nucleic acid binding site may be referred to as the “donor binding site,” or vice versa.
[0210] A guide nucleic acid is disclosed herein, comprising a spacer inversely complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; and a donor nucleic acid binding site, and optionally a flap binding site inversely complementary to the nucleic acid flap.
[0211] In some embodiments, the guide nucleic acid includes a spacer complementary to a genomic locus within the cell; a scaffold for complexing with at least one endonuclease; a donor binding site at least partially complementary to the donor strand; a flap binding site at or adjacent to the genomic locus that is at least partially identical or complementary to the genomic flap; or a combination thereof. In some embodiments, the guide nucleic acid can instruct at least one endonuclease to cleave at least one strand of the genomic locus. In some embodiments, the guide nucleic acid may be at least partially complementary to the donor strand or at least partially complementary to the genomic flap (e.g., a genomic nucleic acid sequence that is replaced and becomes single-stranded when the guide nucleic acid recruits an endonuclease to the genomic locus). In some embodiments, the guide nucleic acid, which is at least partially complementary to the donor strand or at least partially complementary to the genomic flap, brings the donor strand closer to the cleavage of the genomic locus.
[0212] In some embodiments, guide nucleic acids comprising a scaffold are disclosed herein. The scaffold may bind to a nuclease. The scaffold may bind to a Cas nuclease. The scaffold may bind to a nickase. The scaffold may bind to a Cas nickase. The scaffold may bind to a S. Pyogenes (S. pyogenes) Cas9 nuclease. The scaffold may bind to a S. Pyogenes (S. pyogenes) Cas9 nickase. The scaffold may comprise a scaffold nucleic acid sequence. The system described herein may comprise a first guide nucleic acid. The system may comprise a second guide nucleic acid. The first guide nucleic acid may bind to a first Cas nickase. The second guide nucleic acid may bind to a second Cas nickase.
[0213] The guide nucleic acid may include any of the following embodiments: (i) a spacer complementary to the region of the genomic locus of the genome strand, (ii) a scaffold for forming a complex with an RNA guide endonuclease, (iii) a donor binding site at least partially complementary to the first exogenous integrated nucleic acid, or (iv) a flap binding site at or adjacent to the genomic locus that is at least partially identical or complementary to the genomic flap. The guide nucleic acid may include any of the following embodiments: (i) a spacer complementary to the region of the genomic locus of the genome strand, (ii) a scaffold for forming a complex with an RNA guide endonuclease, or (iii) a donor binding site at least partially complementary to the sprint nucleic acid. The components of (i), (ii), or (iii) may be contained in a single guide nucleic acid, or may be divided among multiple guide nucleic acids, or may be contained collectively.
[0214] In some embodiments, the guide nucleic acid includes modified nucleoside bonds. In some embodiments, the modified nucleoside bonds include phosphorothioate bonds. In some embodiments, the modified nucleoside bonds are located between either the four terminal nucleosides at the 5' or 3' end of the guide nucleic acid. The guide nucleic acid may include multiple modified nucleoside bonds. For example, the guide nucleic acid may include modified nucleoside bonds between the nucleic acids at the 5' and 3' ends of the guide nucleic acid, e.g., between the last four nucleic acids at the 5' end and between the last four nucleic acids at the 3' end. In some embodiments, the guide nucleic acid includes modified nucleosides. In some embodiments, the modified nucleosides include locked nucleic acids (LNA), 2'-fluoro, 2'O-alkyl, or combinations thereof. The modified nucleosides may include LNA, 2'-fluoro, 2'O-alkyl, methylated cytosine, reverse thymidine, or combinations thereof. The modified nucleoside may contain LNA. The modified nucleoside may contain a 2'-fluoronucleotide. The modified nucleoside may contain a 2'O-alkyl nucleotide. The modified nucleoside may contain a methylated cytosine nucleotide. In some embodiments, the modified nucleoside is one of the three terminal nucleosides at the 5' or 3' end of the guide nucleic acid. The guide nucleic acid may contain multiple modified nucleosides. For example, the guide nucleic acid may contain modified nucleosides at the 5' and 3' terminal nucleic acids of the guide nucleic acid, e.g., the last three nucleic acids at the 5' end and the last three nucleic acids at the 3' end.
[0215] In some embodiments, the guide nucleic acid includes at least one nucleic acid modification. In some embodiments, at least one nucleic acid modification includes modifying the backbone, sugars, bases, or combinations thereof of the guide nucleic acid. In some embodiments, at least one nucleic acid modification can increase the guide nucleic acid's resistance to degradation (e.g., to nuclease degradation or hydrolysis). In some embodiments, at least one nucleic acid modification can increase the complexation of the guide nucleic acid to at least one endonuclease. In some embodiments, at least one nucleic acid modification can increase the complexation of the guide nucleic acid to a donor strand. In some embodiments, at least one nucleic acid modification can increase the complexation of the guide nucleic acid to a genomic locus by being complementary to a genomic flap.
[0216] In some embodiments, the guide nucleic acid contains at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty Nucleic acid modifications may include linker molecules that do not exist naturally, either interchain or intrachain. In one embodiment, the modified nucleic acid includes modifications of one or more 3'OH or 5'OH groups, a skeleton, a sugar component, or a nucleotide base, or the addition of a linker molecule that does not exist naturally. In some embodiments, the modified skeleton includes a skeleton other than a phosphodiester skeleton. In some embodiments, the modified sugar includes a sugar other than deoxyribose (in modified DNA) or a sugar other than ribose (in modified RNA). In some embodiments, the modified base includes a base other than adenine, guanine, cytosine, thymine, or uracil. In some embodiments, the guide nucleic acid includes at least one modified base. In some examples, the guide nucleic acid includes at least one, two, three, four, five, six, seven, eight, nine, ten, fifteen, twenty, or more modified bases. In some cases, nucleic acid modifications to the base moiety include natural and synthetic modifications of adenine, guanine, cytosine, thymine, or uracil, and purine or pyrimidine bases.
[0217] In some embodiments, at least one nucleic acid modification of the guide nucleic acid is a 2'-modified nucleotide comprising 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl, 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-ON-methylacetamide (2'-O-NMA) in the phosphodiester backbone bond. Modification of one or both of the unbound phosphate oxygens; modification of one or more linked phosphate oxygens in a phosphodiester backbone; modification of components of a ribose sugar; substitution of the phosphate moiety by a "dephospho" linker; alteration or substitution of naturally occurring nucleic acid bases; modification of the ribose-phosphate backbone; modification of the 5' end of a polynucleotide; modification of the 3' end of a polynucleotide; modification of the phosphate backbone of deoxyribose; substitution of phosphate groups; modification of the ribophosphate backbone; modification of the sugar of a nucleotide; modification of the base of a nucleotide; or any one or any combination thereof of steric hindrance of a nucleotide. Non-limiting examples of nucleic acid modifications to guide nucleic acids include: modification of one or both of the unbonded or bonded phosphate oxygen atoms in the phosphodiester backbone (e.g., sulfur (S), selenium (Se), BR3 (wherein R may be, for example, hydrogen, alkyl, or aryl), C (e.g., alkyl group, aryl group, etc.), H, NR2, where R may be, for example, hydrogen, alkyl, or aryl); substitution of the phosphate moiety with a "dephospho" linker (e.g., substitution with methylphosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino); modification or substitution with nucleic acid analogs of naturally occurring nucleic acid bases;Modification of the deoxyribose-phosphate or ribose-phosphate skeleton (e.g., modification of the ribose-phosphate skeleton to incorporate phosphorothioates, phosphonothioacetates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphonocarboxylates, phosphoramides, alkyl or arylphosphonates, phosphonoacetates, or phosphotryesters); modification of the 5' end of the nucleic acid sequence (e.g., 5' cap or 5' cap); Modification of the cap-OH group) or modification of the 3' end (modification of the 3' end or 3'-OH group); methylphosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, foracetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino Substitution by; modification of the ribophosphate skeleton to incorporate morpholino (phosphodiamidate morpholino oligomer PMO), cyclobutyl, pyrrolidine, or peptide nucleic acid (PNA) nucleoside substitutes; modification of nucleotide sugars to incorporate locked nucleic acid (LNA), unlocked nucleic acid (UNA), ethylene nucleic acid (ENA), restricted ethyl (cEt) sugar, or cross-linked nucleic acid (BNA); modification of ribose sugar components (e.g., 2'-O-methyl, 2'-O -Methoxy-ethyl (2'-MOE), 2'-fluoro, 2'-aminoethyl, 2'-deoxy-2'-florabino-cleaving acid, 2'-deoxy, 2'-O-methyl, 3'-phosphorothioate, 3'-phosphonoacetate (PACE) or 3'-phosphonothioacetate (thioPACE); modifications to the bases of nucleotides (of A, T, C, G, or U); and steric hindrance of nucleotides (e.g., S conformation of phosphorothioate or R conformation of phosphorothioate).
[0218] The following codes may be used in the order specified herein.
[0219] [Table 71]
[0220] In some embodiments, the nucleic acid modification includes the substitution of at least one unbound phosphate oxygen atom in the phosphodiester backbone bond of the guide nucleic acid. In some embodiments, at least one nucleic acid modification of the guide nucleic acid includes the substitution of one or more bound phosphate oxygen atoms in the phosphodiester backbone bond of the guide nucleic acid. A non-limiting example of nucleic acid modification of a phosphate oxygen atom is a sulfur atom. In some embodiments, the nucleic acid modification includes at least one modification of a sugar. In some embodiments, the nucleic acid modification includes at least one nucleic acid modification to a sugar, which is a ribose sugar, and includes modification of a component of the sugar. In some embodiments, the nucleic acid modification of the guide nucleic acid includes at least one modification of a component of the ribose sugar of the nucleotide of the guide nucleic acid, which includes a 2'-O-methyl group. In some embodiments, the nucleic acid modification includes at least one modification, which includes substituting the phosphate portion of the guide nucleic acid with a dephospholinker. In some embodiments, the nucleic acid modification includes at least one modification of the phosphate backbone. In some embodiments, the modification includes a phosphorothioate group. In some embodiments, the nucleic acid modification includes at least one modification, which includes a modification to the bases of the nucleotides of the guide nucleic acid. In some embodiments, the nucleic acid modification includes at least one modification, which includes a non-native base of the nucleotide. In some embodiments, the nucleic acid modification includes at least one modification, which includes at least one sterically pure nucleic acid. In some embodiments, at least one nucleic acid modification may be located proximal to the 5' end of the guide nucleic acid. In some embodiments, at least one nucleic acid modification may be located proximal to the 3' end of the guide nucleic acid. In some embodiments, at least one nucleic acid modification may be located proximal to both the 5' and 3' ends of the guide nucleic acid.
[0221] In some embodiments, the guide nucleic acid described herein comprises a skeleton comprising a plurality of covalently bonded sugar and phosphate moieties. In some cases, the guide nucleic acid skeleton comprises a phosphodiester bond between a first hydroxyl group of a phosphate group on the 5' carbon of deoxyribose in DNA or ribose in RNA and a second hydroxyl group on the 3' carbon of deoxyribose in DNA or ribose in RNA. In some embodiments, the guide nucleic acid skeleton may lack a 5'-reduced hydroxyl, a 3'-reduced hydroxyl, or both that can be exposed to a solvent. In some embodiments, the guide nucleic acid skeleton may lack a 5'-reduced hydroxyl, a 3'-reduced hydroxyl, or both that can be exposed to a nuclease. In some embodiments, the guide nucleic acid skeleton may lack a 5'-reduced hydroxyl, a 3'-reduced hydroxyl, or both that can be exposed to a hydrolase. In some examples, the guide nucleic acid skeleton can be represented as a polynucleotide sequence in a cyclic two-dimensional format, where one nucleotide is present sequentially. In some examples, the guide nucleic acid backbone can be represented as a polynucleotide sequence in a loop-like two-dimensional format, with nucleotides arranged sequentially. In some cases, a 5'-hydroxyl, a 3'-hydroxyl, or both are linked via a phosphorus-oxygen bond. In some cases, the 5'-hydroxyl, a 3'-hydroxyl, or both are modified with a phosphoester having a phosphorus-containing moiety.In some embodiments, the guide nucleic acid is 5'-adenylic acid, 5'-guanosine triphosphate cap, 5'N7-methylguanosine triphosphate cap, 5'-triphosphate cap, 3'-phosphate, 3'-thiophosphate, 5'-phosphate, 5'-thiophosphate, cis-scinthymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer18, Spacer9, 3'-3' modification, 5'-5' modification, abasic, acridine, azobenzene, biotin, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3'DABCYL, black hole quencher, black hole quencher 2, DABCYL The invention comprises at least one nucleic acid modification containing any one of the following: SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methylribonoside analog, sugar-modified analog, fluctuation / universal base, fluorescent dye labeling, 2'-fluoroRNA, 2'O-methylRNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphorothioate DNA, phosphorothioate RNA, UNA, LNA, cEt, pseudouridine-5'-triphosphate, 5-methylcytidine-5'-triphosphate, 2-O-methyl-phosphorothioate, or any combination thereof.
[0222] Nucleic acid modifications may include phosphorothioate substitutions. In some cases, the natural phosphodiester bond is susceptible to rapid degradation by cellular nucleases; modification of internucleotide bonds using phosphorothioate (PS) bond substitutions may be more stable against hydrolysis by cellular degradation. Modifications can increase the stability of polynucleic acids. Modifications can also enhance biological activity. In some cases, phosphorothioate-enhanced RNA polynucleic acids can inhibit RNase A, RNase T1, bovine serum nucleases, or any combination thereof. These properties may enable the use of PS-RNA polynucleic acids in applications where exposure to nucleases is highly probable in vivo or in vitro. For example, a phosphorothioate (PS) bond can be introduced between the last 3-5 nucleotides of the 5' or 3' end of a polynucleic acid to inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added to the entire polynucleic acid to reduce attack by endonucleases. In some embodiments, the guide nucleic acid includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 50, 100, or more internucleotide bonds, including PS bonds. In some embodiments, the guide nucleic acid includes only PS bonds as internucleotide bond modifications. In some embodiments, all internucleotide bonds of the guide nucleic acid herein are either fully PS modified or include phosphorothioate internucleotide bonds.
[0223] The guide nucleic acid may contain hairpins. The hairpins may bind to hairpin-binding motifs on DNA ligases, such as hairpin-binding motifs. The hairpins may include MS2 hairpins and MS2 hairpin A hairpins, which may be useful for recruiting DNA ligases containing MCP peptides.
[0224] The guide nucleic acid may include any of the embodiments shown in Figures 1A to 6C. Table 11 illustrates some non-limiting examples of guide nucleic acids described herein. Some of the guide nucleic acids in the table include nucleic acid modifications.
[0225] [Table 72]
[0226] [Table 73]
[0227] [Table 74]
[0228] [Table 75]
[0229] The guide nucleic acid may include sequences of linking nucleic acids (e.g., linked RNA or DNA nucleotides) between the components of the guide nucleic acid. For example, the guide nucleic acid may include sequences of linking nucleic acids between any of the following components: spacer, scaffold, donor binding site, or flap binding site. The guide nucleic acid may include sequences of linking nucleic acids between the spacer, scaffold, or donor binding site. The guide nucleic acid includes sequences of linking nucleic acids between the scaffold and the donor binding site. The guide nucleic acid may include sequences of linking nucleic acids between the spacer and the scaffold. The guide nucleic acid may include multiple sequences of linking nucleic acids between components.
[0230] The sequence of linked nucleic acids may contain any base, such as A, U, T, G, or C, or combinations thereof. The sequence of linked nucleic acids may contain A, T, G, or C, or combinations thereof. The sequence of linked nucleic acids may contain A, U, G, or C, or combinations thereof. The sequence of linked nucleic acids may contain a series of As. The sequence of linked nucleic acids may contain a series of Ts. The sequence of linked nucleic acids may contain a series of U. The sequence of linked nucleic acids may contain a series of Cs. The sequence of linked nucleic acids may contain a series of G.
[0231] The sequence of linked nucleic acids may have a length, for example, several nucleotides. The length may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 nucleotides, or a range defined by any two of the aforementioned numbers of nucleotides. The length may include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides. In some embodiments, the length may be less than 2, less than 3, less than 4, less than 5, less than 6, less than 7, less than 8, less than 9, less than 10, less than 11, less than 12, less than 13, less than 14, less than 15, less than 16, less than 17, less than 18, less than 19, less than 20, less than 21, less than 22, less than 23, less than 24, less than 25, less than 30, less than 35, less than 40, less than 45, less than 50, less than 55, less than 60, less than 65, less than 70, less than 75, less than 80, less than 85, less than 90, less than 95, or less than 100 nucleotides.
[0232] Some embodiments relate to a guide nucleic acid comprising a spacer at least partially complementary to an intracellular genomic locus; a scaffold for complexing with an RNA guide endonuclease; and a donor binding site at least partially complementary to an exogenous first integrated nucleic acid. The guide nucleic acid may further comprise a flap binding site at least partially complementary to the genomic sequence of the genomic locus. The guide nucleic acid may further comprise at least one nucleic acid modification. The at least one nucleic acid modification may comprise a modification to the backbone, sugars, bases, or combinations thereof. The guide nucleic acid may comprise RNA.
[0233] Some embodiments include a spacer at least partially inversely complementary to a first region of a target nucleic acid; a scaffold configured to bind to an endonuclease; a guide nucleic acid comprising a flap binding site at least partially complementary to the nucleic acid flap; and an exogenous first embedded nucleic acid binding site.
[0234] Sprint nucleic acid This specification discloses sprint nucleic acids. Sprint nucleic acids may be included in the compositions, systems, or methods disclosed herein. Some embodiments relate to nucleic acids (e.g., DNA or RNA). Sprint nucleic acids may include a DNA or RNA backbone or “sprint” that is inversely complementary to the DNA sequence being ligated. Non-limiting examples include the replacer guide nucleic acid shown in Figures 1A, 2A, 5A, and 6A, which includes a flap binding site (FBS) adjacent to a donor binding site (DBS). In some embodiments, the sprint nucleic acid may include a flap binding site that is at least partially identical or complementary to a genomic locus or an adjacent genomic flap, and may optionally include a guide binding site that is at least partially complementary to the guide nucleic acid. Non-limiting examples include the donor 2 nucleic acid shown in the replacer 2 model of Figures 3A and 4A, which includes a guide binding site (GBS) adjacent to a flap binding site (FBS) adjacent to a donor binding site DBS. In some embodiments, the genomic strand is intracellular. In some embodiments, the sprint nucleic acid further includes a donor binding site that is at least partially identical or complementary to a portion of the integrated nucleic acid. In some embodiments, the donor nucleic acid includes the sprint nucleic acid. In some embodiments, the guide nucleic acid includes the sprint nucleic acid.
[0235] In some embodiments, the sprint nucleic acid includes modified internucleoside bonds. In some embodiments, the modified internucleoside bonds include phosphorothioate bonds. In some embodiments, the modified internucleoside bonds are located between either the four terminal nucleosides at the 5' or 3' end of the sprint nucleic acid. The sprint nucleic acid may include multiple modified internucleoside bonds. For example, the sprint nucleic acid may include modified internucleoside bonds between the nucleic acids at the 5' and 3' ends of the sprint nucleic acid, e.g., between the last four nucleic acids at the 5' end and between the last four nucleic acids at the 3' end. In some embodiments, the sprint nucleic acid includes modified nucleosides. In some embodiments, the modified nucleosides include locked nucleic acids (LNA), 2'-fluoro, 2'O-alkyl, or a combination thereof. The modified nucleosides may include LNA, 2'-fluoro, 2'O-alkyl, methylated cytosine, reverse thymidine, or a combination thereof. The modified nucleoside may contain LNA. The modified nucleoside may contain a 2'-fluoronucleotide. The modified nucleoside may contain a 2'O-alkyl nucleotide. The modified nucleoside may contain a methylated cytosine nucleotide. In some embodiments, the modified nucleoside is one of the three terminal nucleosides at the 5' or 3' end of the sprint nucleic acid. The sprint nucleic acid may contain multiple modified nucleosides. For example, the sprint nucleic acid may contain modified nucleosides at the 5' and 3' terminal nucleic acids of the sprint nucleic acid, for example, the last three nucleic acids at the 5' end and the last three nucleic acids at the 3' end.
[0236] In some embodiments, the sprint nucleic acid comprises at least one nucleic acid modification. In some embodiments, at least the nucleic acid modification comprises modifying the backbone, sugars, bases, or combinations thereof of the sprint nucleic acid. In some embodiments, at least one nucleic acid modification can increase the resistance of the sprint nucleic acid to degradation (e.g., to nuclease degradation or hydrolysis). In some embodiments, at least one nucleic acid modification can increase the complexation of the sprint nucleic acid to at least one endonuclease. In some embodiments, at least one nucleic acid modification can increase the complexation of the sprint nucleic acid to a donor strand. In some embodiments, at least one nucleic acid modification can increase the complexation of the sprint nucleic acid to a genomic locus by being complementary to a genomic flap.
[0237] In some embodiments, the sprinted nucleic acid contains at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more nucleic acid modifications. In some embodiments, the nucleic acid modifications may occur at 3'OH groups, 5'OH groups, backbone, sugar components, or nucleotide bases. The nucleic acid modifications may include linker molecules that do not exist naturally, either interchain or intrachain crosslinking. In one embodiment, the modified nucleic acid includes modifications of one or more 3'OH or 5'OH groups, a skeleton, a sugar component, or a nucleotide base, or the addition of a linker molecule that does not exist in nature. In some embodiments, the modified skeleton includes a skeleton other than a phosphodiester skeleton. In some embodiments, the modified sugar includes a sugar other than deoxyribose (in modified DNA) or a sugar other than ribose (in modified RNA). In some embodiments, the modified base includes a base other than adenine, guanine, cytosine, thymine, or uracil. In some embodiments, the sprint nucleic acid includes at least one modified base. In some examples, the sprint nucleic acid includes at least one, two, three, four, five, six, seven, eight, nine, ten, fifteen, twenty, or more modified bases. In some cases, nucleic acid modifications to the base portion include natural and synthetic modifications of adenine, guanine, cytosine, thymine, or uracil, and purine or pyrimidine bases.
[0238] In some embodiments, at least one nucleic acid modification of the sprint nucleic acid is a 2'-modified nucleotide comprising 2'-O-methyl, 2'-O-methoxyethyl (2'-O-MOE), 2'-O-aminopropyl, 2'-deoxy, 2'-deoxy-2'-fluoro, 2'-O-aminopropyl (2'-O-AP), 2'-O-dimethylaminoethyl (2'-O-DMAOE), 2'-O-dimethylaminopropyl (2'-O-DMAP), 2'-O-dimethylaminoethyloxyethyl (2'-O-DMAEOE), or 2'-ON-methylacetamide (2'-O-NMA); phosphodiester backbone linkage Modification of one or both of the unbound phosphate oxygen atoms in a nucleotide; modification of one or more linked phosphate oxygen atoms in a phosphodiester backbone; modification of components of a ribose sugar; substitution of the phosphate moiety by a "dephospho" linker; alteration or substitution of naturally occurring nucleic acid bases; modification of the ribose-phosphate skeleton; modification of the 5' end of a polynucleotide; modification of the 3' end of a polynucleotide; modification of the phosphate skeleton of a deoxyribose; substitution of a phosphate group; modification of the ribophosphate skeleton; modification of the sugar of a nucleotide; modification of the base of a nucleotide; or any one of these steric hindrance modifications of a nucleotide, or any combination thereof. Non-limiting examples of nucleic acid modifications to sprint nucleic acids may include: modifications of one or both of the unbonded or bonded phosphate oxygen atoms in the phosphodiester backbone (e.g., sulfur (S), selenium (Se), BR3 (wherein R may be, for example, hydrogen, alkyl, or aryl), C (e.g., alkyl group, aryl group, etc.), H, NR2, where R may be, for example, hydrogen, alkyl, or aryl); or modifications of the phosphate moiety. Substitution by a sulfolinker (e.g., substitution with methylphosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, formacetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino); modification or substitution with nucleic acid analogs of naturally occurring nucleic acid bases;Modification of the deoxyribose-phosphate or ribose-phosphate skeleton (e.g., modification of the ribose-phosphate skeleton to incorporate phosphorothioates, phosphonothioacetates, phosphoroselenates, boranophosphates, boranophosphate esters, hydrogen phosphonates, phosphonocarboxylates, phosphoramides, alkyl or arylphosphonates, phosphonoacetates, or phosphotryesters); 5' end of nucleic acid sequence (e.g., 5' cap or 5') Modification of the cap-OH group) or modification of the 3' end (modification of the 3' end or 3'-OH group); methylphosphonate, hydroxylamino, siloxane, carbonate, carboxymethyl, carbamate, amide, thioether, ethylene oxide linker, sulfonate, sulfonamide, thioformacetal, foracetal, oxime, methyleneimino, methylenemethylimino, methylenehydrazo, methylenedimethylhydrazo, or methyleneoxymethylimino Substitution with mino; modification of the ribophosphate skeleton to incorporate morpholino (phosphodiamidate morpholino oligomer PMO), cyclobutyl, pyrrolidine, or peptide nucleic acid (PNA) nucleoside substitutes; modification of nucleotide sugars to incorporate locked nucleic acid (LNA), unlocked nucleic acid (UNA), ethylene nucleic acid (ENA), restricted ethyl (cEt) sugar, or cross-linked nucleic acid (BNA); modification of ribose sugar components (e.g., 2'-O-methyl '-O-methoxyethyl (2'-MOE), 2'-fluoro, 2'-aminoethyl, 2'-deoxy-2'-florabinocleaving acid, 2'-deoxy, 2'-O-methyl, 3'-phosphorothioate, 3'-phosphonoacetate (PACE) or 3'-phosphonothioacetate (thioPACE); modifications to the bases of nucleotides (of A, T, C, G, or U); and steric hindrance of nucleotides (e.g., S conformation of phosphorothioate or R conformation of phosphorothioate).
[0239] In some embodiments, nucleic acid modification includes the substitution of at least one unbound phosphate oxygen atom in the phosphodiester backbone bond of the sprint nucleic acid. In some embodiments, at least one nucleic acid modification of the sprint nucleic acid includes the substitution of one or more bound phosphate oxygen atoms in the phosphodiester backbone bond of the sprint nucleic acid. A non-limiting example of nucleic acid modification of a phosphate oxygen atom is a sulfur atom. In some embodiments, nucleic acid modification includes at least one modification to a sugar. In some embodiments, nucleic acid modification includes at least one nucleic acid modification to a sugar, which is a ribose sugar, and includes modification of a component of the sugar. In some embodiments, nucleic acid modification of the sprint nucleic acid includes at least one modification to a component of the ribose sugar of the nucleotide of the sprint nucleic acid, which includes a 2'-O-methyl group. In some embodiments, nucleic acid modification includes at least one modification, which includes substituting the phosphate portion of the sprint nucleic acid with a dephospholinker. In some embodiments, nucleic acid modification includes at least one modification of the phosphate backbone. In some embodiments, the modification includes a phosphorothioate group. In some embodiments, the nucleic acid modification includes at least one modification involving a modification to the bases of the nucleotides of the sprint nucleic acid. In some embodiments, the nucleic acid modification includes at least one modification involving a non-native base of the nucleotide. In some embodiments, the nucleic acid modification includes at least one modification involving at least one sterically pure nucleic acid. In some embodiments, at least one nucleic acid modification may be located proximal to the 5' end of the sprint nucleic acid. In some embodiments, at least one nucleic acid modification may be located proximal to the 3' end of the sprint nucleic acid. In some embodiments, at least one nucleic acid modification may be located proximal to both the 5' and 3' ends of the sprint nucleic acid.
[0240] In some embodiments, the sprint nucleic acids described herein include a skeleton comprising a plurality of covalently bonded sugar and phosphate moieties. In some cases, the skeleton of the sprint nucleic acid includes a phosphodiester bond between a first hydroxyl group of a phosphate group on the 5' carbon of deoxyribose in DNA or ribose in RNA and a second hydroxyl group on the 3' carbon of deoxyribose in DNA or ribose in RNA. In some embodiments, the skeleton of the sprint nucleic acid may lack a 5'-reduced hydroxyl, a 3'-reduced hydroxyl, or both that can be exposed to a solvent. In some embodiments, the skeleton of the sprint nucleic acid may lack a 5'-reduced hydroxyl, a 3'-reduced hydroxyl, or both that can be exposed to a nuclease. In some embodiments, the skeleton of the sprint nucleic acid may lack a 5'-reduced hydroxyl, a 3'-reduced hydroxyl, or both that can be exposed to a hydrolase. In some examples, the skeleton of the sprint nucleic acid can be represented as a polynucleotide sequence in a cyclic two-dimensional format in which one nucleotide is present sequentially. In some cases, the backbone of a sprint nucleic acid can be represented as a polynucleotide sequence in a loop-like two-dimensional format, with one nucleotide following another. In some cases, a 5'-hydroxyl, a 3'-hydroxyl, or both are linked via a phosphorus-oxygen bond. In some cases, the 5'-hydroxyl, a 3'-hydroxyl, or both are modified with a phosphoester having a phosphorus-containing moiety.In some embodiments, the sprint nucleic acid is 5'-adenylic acid, 5'-guanosine triphosphate cap, 5'N7-methylguanosine triphosphate cap, 5'-triphosphate cap, 3'-phosphate, 3'-thiophosphate, 5'-phosphate, 5'-thiophosphate, cis-scinthymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, Spacer18, Spacer9, 3'-3' modification, 5'-5' modification, abasic, acridine, azobenzene, biotin, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP TEG, DNP-X, DOTA, dT-biotin, dual biotin, PC biotin, psoralen C2, psoralen C6, TINA, 3'DABCYL, black hole quencher, black hole quencher 2, DABCYL The invention comprises at least one nucleic acid modification containing any one of the following: SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methylribonoside analog, sugar-modified analog, fluctuation / universal base, fluorescent dye labeling, 2'-fluoroRNA, 2'O-methylRNA, methylphosphonate, phosphodiester DNA, phosphodiester RNA, phosphorothioate DNA, phosphorothioate RNA, UNA, LNA, cEt, pseudouridine-5'-triphosphate, 5-methylcytidine-5'-triphosphate, 2-O-methyl-phosphorothioate, or any combination thereof.
[0241] Nucleic acid modifications may include phosphorothioate substitutions. In some cases, the natural phosphodiester bond is susceptible to rapid degradation by cellular nucleases; modification of internucleotide bonds using phosphorothioate (PS) bond substitutions may be more stable against hydrolysis by cellular degradation. Modifications can increase the stability of polynucleic acids. Modifications can also enhance biological activity. In some cases, phosphorothioate-enhanced RNA polynucleic acids can inhibit RNase A, RNase T1, bovine serum nucleases, or any combination thereof. These properties may enable the use of PS-RNA polynucleic acids in applications where exposure to nucleases is highly probable in vivo or in vitro. For example, a phosphorothioate (PS) bond can be introduced between the last 3-5 nucleotides of the 5' or 3' end of a polynucleic acid to inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added to the entire polynucleic acid to reduce attack by endonucleases. In some embodiments, the sprint nucleic acid contains at least one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, fifteen, sixteen, seventeen, eighteen, nineteen, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, twenty, thirty, fifty, fifty, hundred, or more internucleotide bonds, including PS bonds. In some embodiments, the sprint nucleic acid contains only PS bonds as internucleotide bond modifications. In some embodiments, all internucleotide bonds of the sprint nucleic acid herein are either fully PS modified or contain phosphorothioate internucleotide bonds.
[0242] Sprint nucleic acids may contain hairpins. These hairpins may bind to hairpin-binding motifs on DNA ligases, such as hairpin-binding motifs. The hairpins may include MS2 hairpins and MS2 hairpin A hairpins, which may be useful for recruiting DNA ligases containing MCP peptides.
[0243] In some embodiments, the donor binding site (DBS) of the sprint is 12 nt to 50 nt in length and includes lengths of 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 36, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, and 49 nt. In some embodiments, the DBS of the sprint contains one or more modified nucleotides. In some embodiments, 10% to 80% of the DBS nucleotides are modified. In some embodiments, alternating nucleotides are modified. In some embodiments, modification occurs every third nucleotide. In some embodiments, the DBS contains alternating LNAs or contains regions where every three nucleotides are LNAs. In some embodiments, the DBS includes anternating a 2'F nucleotide or includes a region in which every third nucleotide is 2'F modified. In some embodiments, the DBS includes alternating 2'OMe nucleotides or includes a region in which every third nucleotide is 2'OMe modified.
[0244] In some embodiments, the sprint guide binding site (GBS) contains one or more modified nucleotides. In some embodiments, 10% to 80% of the GBS nucleotides are modified. In some embodiments, modification occurs every third nucleotide. In some embodiments, the GBS contains alternating LNAs or regions where every third nucleotide is an LNA. In some embodiments, the GBS contains a sequence of 2'F nucleotides or regions where every third nucleotide is 2'F modified. In some embodiments, the GBS contains alternating 2'OMe nucleotides or regions where every third nucleotide is 2'OMe modified.
[0245] Non-limiting examples of chemical modifications to sprint nucleic acids and the corresponding nucleic acid sequences are provided in Tables 12 and 25, and shown in Figure 17.
[0246] In some embodiments, the chemical modification of the sprint may be a 3'C3 spacer. In some embodiments, the chemical modification of the sprint may be a 3' phosphorylation. In some embodiments, the chemical modification of the sprint may be a 3' phosphorylation. In some embodiments, the chemical modification of the sprint may be a 3' reverse T overhang. In some embodiments, the chemical modification of the sprint may be a 5' reverse T overhang. In some embodiments, the chemical modification of the sprint may be a 3' reverse T portion of the GBS.
[0247] In some embodiments, editing efficiency can be optimized. In some embodiments, nucleic acid double-strand length and composition can be adjusted. In some embodiments, the sprint binding site (SBS) of the guide nucleic acid is optimized. In some embodiments, the guide binding site (GBS) of the sprint is optimized. In some embodiments, the sprint binding site of the donor nucleic acid is optimized. In some embodiments, the donor binding site (DBS) of the sprint nucleic acid is optimized. In some embodiments, it is preferable to avoid certain modifications of the nucleic acid that is incorporated into the target. In some embodiments, it is preferable to modify the double-strand forming nucleic acid that is not incorporated into the target. In some embodiments, it is preferable to modify the double-strand forming region of nucleic acids that are not integrated, such as the double-strand forming region of the sprint nucleic acid or guide nucleic acid.
[0248] In some embodiments, editing efficiency can be optimized by adjusting the length and composition of the double helix formed by the gRNA SBS and sprint GBS. In some embodiments, editing efficiency can be optimized by adjusting the length and composition of the double helix formed by the donor SBS and sprint DBS. In some embodiments, editing efficiency can be optimized by adjusting the length and composition of the double helix formed by the gRNA SBS and sprint GBS. In some embodiments, optimizing editing efficiency includes modifying the sprint GBS.
[0249] In some embodiments, the length of the sprint GBS is adjusted. In some embodiments, the length of the sprint GBS-gRNA double helix is 10 bp to 30 bp, or 12 bp to 28 bp, or 14 bp to 26 bp, or 16 bp to 24 bp, or 18 bp to 22 bp, or 17 bp to 19 bp, or 18 bp to 20 bp, or 19 bp to 21 bp, or 17 bp, or 18 bp, or 19 bp, or 20 bp, or 21 bp, or 22 bp, or 23 bp. In some embodiments, the sprint GBS contains LNA, and the number or proportion of LNA is adjusted. In some embodiments, LNA nucleotides are alternating with DNA nucleotides.
[0250] In some embodiments, optimizing editing efficiency involves modifying the sprint DBS. In some embodiments, the length of the sprint DBS is adjusted. In some embodiments, the length of the sprinted GBS donor double helix is 18 bp to 40 bp, or 19 bp to 36 bp, or 20 bp to 32 bp, or 21 bp to 30 bp, or 22 bp to 28 bp, or 23 bp to 27 bp, or 24 bp to 26 bp, or 18 bp to 22 bp, or 20 bp to 24 bp, or 22 bp to 28 bp, or 24 bp to 30 bp, or 26 bp to 32 bp, or 28 bp to 34 bp, or 18 bp, or 19 bp, or 20 bp, or 21 bp, or 22 bp, or 23 bp, or 24 bp, or 25 bp, or 26 bp, or 27 bp, or 28 bp, or 29 bp, or 30 bp, or 31 bp, or 32 bp, or 33 bp, or 34 bp, or 35 bp, or 36 bp, or 37 bp, or 38 bp, or 39 bp, or 40 bp. In some embodiments, the sprint DBS contains LNA, and the number or proportion of LNA is adjusted. In some embodiments, the LNA nucleotides are alternating with DNA nucleotides.
[0251] In some embodiments, optimizing editing efficiency involves modifying the sprint FBS. In some embodiments, the length of the sprint FBS is adjusted. In some embodiments, the length of the sprint FBS-flap double helix is 10 bp to 30 bp, or 12 bp to 28 bp, or 14 bp to 26 bp, or 16 bp to 24 bp, or 18 bp to 22 bp, or 17 bp to 19 bp, or 18 bp to 20 bp, or 19 bp to 21 bp, or 17 bp, or 18 bp, or 19 bp, or 20 bp, or 21 bp, or 22 bp, or 23 bp. In some embodiments, the sprint DBS contains LNA, and the number or proportion of LNA is adjusted. In some embodiments, LNA nucleotides are alternating with DNA nucleotides. Table 12 below provides non-limiting examples of sprints used herein.
[0252] [Table 76]
[0253] [Table 77]
[0254] [Table 78]
[0255] [Table 79]
[0256] [Table 80]
[0257] [Table 81]
[0258] target nucleic acid Target nucleic acids are disclosed herein. Target nucleic acids may include DNA. Target nucleic acids may also include RNA. Target nucleic acids may be present in cells. Target nucleic acids may be methylated. Target nucleic acids may be unmethylated. Target nucleic acids may include genomes. Target nucleic acids may include genomic DNA. Target nucleic acids may include chromosomes. Target nucleic acids may include genes.
[0259] The target nucleic acid may be present within the sample. The target nucleic acid may be present within cells. The target nucleic acid may be present in a test tube.
[0260] The target nucleic acid may be edited. The target nucleic acid may be edited in vitro. The target nucleic acid may be edited in vivo.
[0261] Embedded nucleic acid Exogenous first integrated nucleic acid This specification discloses an exogenous first embedded nucleic acid. The exogenous first embedded nucleic acid may be included in the compositions, systems, or methods disclosed herein. Some embodiments relate to nucleic acids encoding the exogenous first embedded nucleic acid. This specification provides an exogenous first embedded nucleic acid that is inserted into a target nucleic acid, such as a host genome, at a locus. For example, the exogenous first embedded nucleic acid may replace a nucleic acid in the target nucleic acid. The exogenous first embedded nucleic acid may be called a “donor nucleic acid,” a “donor,” or a “donor strand.” Where a genomic locus is described, the locus may be included, or vice versa. For example, the locus may be part of the host genome or part of a non-genomic nucleic acid. The donor may include DNA. Similarly, the target nucleic acid may include DNA. In some cases, for example, if the target nucleic acid includes RNA, the donor may include RNA. The donor may include any insert, such as a gene or regulatory element, that is inserted into the genomic locus of the target nucleic acid. The donor strand may contain a sequence at least partially homologous to a genomic locus. In some cases, the donor may also act as a sprint for the DNA ligase described herein, for example, to ligate two nucleic acid strands that base-pair with a portion of an exogenous first integrated nucleic acid. In some cases, the sprint may contain one strand of the donor, and the portion to be ligated may be another strand of the donor. In some cases, the sprint may contain a strand of the donor, and the portion to be ligated may be an upstream or downstream portion of the same strand of the donor. The donor may be single-stranded. The donor may be double-stranded. The donor may be delivered as a double strand. The donor may be delivered as multiple strands, for example, two strands.
[0262] The first exogenous incorporated nucleic acid may not exist in nature. The donor may be engineered. The donor may be synthetic. The donor may be pre-synthesized. The donor may be added to the subject or cells. In some embodiments, the donor does not contain a polymerase template.
[0263] Disclosed herein is a first exogenous embedded nucleic acid, comprising a double-stranded DNA region to be inserted into a target nucleic acid, the double-stranded DNA region being adjacent to at least one overhang including a flap binding site and / or a guide binding site.
[0264] The first exogenous integrated nucleic acid can be ligated to a target nucleic acid, such as a genome strand. The first exogenous integrated nucleic acid may include a 5' end that can be ligated to the 3' end of the genome strand generated by an RNA guide endonuclease.
[0265] The donor may include any of the embodiments shown in Figures 1A to 6C. For example, the donor may include embodiments such as a guide junction site, a flap junction site, or an overhang. The donor may include a guide junction site. The donor may include two guide junction sites. The donor may include a flap junction site. The donor may include two flap junction sites. The donor may include an overhang. The donor may include two overhangs. The embodiments may be included at the 5' end or 3' end of the donor, or at both ends. The guide junction site or flap junction site may be located within the internal region of the donor.
[0266] Some embodiments include an exogenous first embedded nucleic acid, the nucleic acid being a double-stranded DNA region to be inserted into a target nucleic acid, the double-stranded DNA region being adjacent to at least one overhang including a flap binding site or a guide binding site.
[0267] In some embodiments, the exogenous first integrated nucleic acid includes modified nucleoside bonds. In some embodiments, the modified nucleoside bonds include phosphorothioate bonds. In some embodiments, the modified nucleoside bonds are located between either the four terminal nucleosides at the 5' or 3' end of the donor nucleic acid. The exogenous first integrated nucleic acid may include a plurality of modified nucleoside bonds. For example, the exogenous first integrated nucleic acid may include modified nucleoside bonds between the 5' and 3' terminal nucleic acids of the exogenous first integrated nucleic acid, e.g., between the last four nucleic acids at the 5' end and between the last four nucleic acids at the 3' end. In some embodiments, the exogenous first integrated nucleic acid includes modified nucleosides. In some embodiments, the modified nucleosides include locked nucleic acids (LNA), 2'-fluoro, 2'O-alkyl, 5'O-methyl, 2'-O-methyl, or combinations thereof. The modified nucleoside may include LNA, 2'-fluoro, 2'O-alkyl, methylated cytosine, reverse thymidine, or a combination thereof. The modified nucleoside may include LNA. The modified nucleoside may include 2'-fluoro. The modified nucleoside may include 2'O-alkyl. The modified nucleoside may also include methylated cytosine. In some embodiments, the modified nucleoside is one of the three terminal nucleosides at the 5' or 3' end of the exogenous first integrated nucleic acid. The exogenous first integrated nucleic acid may contain multiple modified nucleosides. For example, the exogenous first integrated nucleic acid may contain modified nucleosides at the 5' and 3' terminal nucleic acids of the exogenous first integrated nucleic acid, e.g., the last three nucleic acids at the 5' end and the last three nucleic acids at the 3' end. The exogenous first integrated nucleic acid may include any modifications, such as modified nucleosides or modified nucleoside-linking, described in relation to the guide nucleic acid, as long as they do not interfere with the function of the exogenous first integrated nucleic acid after ligation to a target nucleic acid, such as the host genome. The exogenous first integrated nucleic acid may include any number or combination of modifications, such as any number or combination, described in relation to the guide nucleic acid, as long as they do not interfere with the function of the exogenous first integrated nucleic acid. Tables 11 and 13 include some examples of exogenous first integrated nucleic acid sequences.
[0268] Donor nucleic acids may contain methylated nucleotides. Donor nucleic acids may contain unmethylated nucleotides. An example of a methylated nucleotide is a nucleotide containing methylated cytosine. Cytosine may be methylated at the C-5 position of the cytosine ring. An example of an unmethylated nucleotide may include unmethylated cytosine. Unmethylated nucleotides may contain cytosine that is not methylated at the C-5 position of the cytosine ring.
[0269] In some embodiments, the donor nucleic acid may include modified nucleotides. In some embodiments, the donor nucleic acid may include modified DNA bases. In some embodiments, the modified nucleotides may include methylation. In some embodiments, the donor nucleic acid may include methylated nucleosides. A non-limiting example of a methylated nucleoside is methylated cytosine (5-mC). In some embodiments, the donor nucleic acid may include N6-methyladenine (6-mA). In some embodiments, the modified DNA base may include 5-hydroxymethylcytosine (5-hmC). In some embodiments, the modified DNA base may include 5-formylcytosine (5-fC). In some embodiments, the modified DNA base may include 5-carboxylcytosine (5-caC).
[0270] In some embodiments, the donor can introduce epigenetic modifications to the target nucleic acid. In some embodiments, the donor can introduce epigenetic modifications of DNA bases (e.g., methylated nucleosides) to the target nucleic acid. In some embodiments, the donor can introduce methylated DNA bases to the target nucleic acid (e.g., methylated nucleosides). In some embodiments, the donor can introduce unmethylated DNA bases to the target nucleic acid (e.g., methylated nucleosides). In some embodiments, the donor integration can introduce methoxylated nucleosides into the genome. In some embodiments, the donor integration can remove methylated nucleosides from the genome. In some embodiments, the donor can introduce methylated cytosine (e.g., 5-mC) into the target nucleic acid. In some embodiments, the donor can introduce N6-methyladenine (6-mA) into the target nucleic acid. In some embodiments, the donor can introduce 5-hydroxymethylcytosine (5-hmC) into the target nucleic acid. In some embodiments, the donor can introduce 5-formylcytosine (5-fC) into the target nucleic acid. In some embodiments, the donor can introduce 5-carboxylcytosine (5-caC) into the target nucleic acid.
[0271] Recombinant sequence This specification discloses a first exogenous integrated nucleic acid. In some embodiments, the first exogenous integrated nucleic acid may include a recombinant sequence. The recombinant sequence may also be referred to as an “integrated recombinant sequence,” “integrated sequence,” “recombination site,” “integration site,” “site-specific recombinant sequence,” or “site-specific integrated sequence.”
[0272] In some embodiments, the exogenous first integrated nucleic acid may contain a plurality of recombinant sequences. In some embodiments, the exogenous first integrated nucleic acid may contain at least one recombinant sequence. In some embodiments, the donor nucleic acid may contain at least two recombinant sequences. In some embodiments, the donor nucleic acid may contain at least three recombinant sequences. In some embodiments, the donor nucleic acid may contain at least four recombinant sequences. In some embodiments, the donor nucleic acid may contain at least five recombinant sequences. In some embodiments, the donor nucleic acid may contain at least ten recombinant sequences.
[0273] In some embodiments, the exogenous first integrated nucleic acid described herein may include a nucleic acid sequence recognized or bound by an integrase. In some embodiments, the exogenous first nucleic acid described herein may include a nucleic acid sequence recognized or bound by a recombinase. In some embodiments, the exogenous first nucleic acid described herein may include a nucleic acid sequence recognized or bound by any integrase described herein. In some embodiments, the exogenous first nucleic acid described herein may include a nucleic acid sequence recognized or bound by any integrase listed in Table 8. The donor nucleic acid described herein may include a nucleic acid sequence recognized or bound by an integrase that includes a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8.
[0274] In some embodiments, the donor nucleic acids described herein may include nucleotide sequences recognized or bound by serine recombinases. The serine recombinases may be, or include, PhiC31 (ΦC31) bacteriophage integrase, Bxb1 mycobacteriophage integrase, Pseudomonas aeruginosa (Pa01) integrase, Nocardia otitidiscaviarum (No67) integrase, or Streptomyces ipomoeae (Si74) integrase. In some embodiments, the donor nucleic acids described herein may include nucleotide sequences recognized or bound by ΦC31 bacteriophage integrase. In some embodiments, the donor nucleic acids described herein may include nucleotide sequences recognized or bound by Bxb1 mycobacterium integrase. The donor nucleic acids described herein may include nucleotide sequences recognized or bound by Pseudomonas aeruginosa (Pa01) integrase. The donor nucleic acids described herein may include nucleotide sequences recognized or bound by Nocardia otitidiscaviarum (No67) integrase. The donor nucleic acids described herein may include nucleotide sequences recognized or bound by Streptomyces ipomoeae (Si74) integrase.
[0275] In some embodiments, the donor nucleic acids described herein may comprise a nucleotide sequence recognized or bound by tyrosine recombinase. In some embodiments, the donor nucleic acids described herein may comprise a nucleotide sequence recognized or bound by Cre recombinase. In some embodiments, the donor nucleic acids described herein may comprise a nucleotide sequence recognized or bound by flippase (Flp).
[0276] In some embodiments, the exogenous first integrated nucleic acid described herein comprises one of the sequences in Table 13. In some embodiments, the donor nucleic acid described herein comprises a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to any of the recombinant sequences in Table 13.
[0277] [Table 82]
[0278] [Table 83]
[0279] [Table 84]
[0280] [Table 85]
[0281] [Table 86]
[0282] [Table 87]
[0283] [Table 88]
[0284] [Table 89]
[0285] [Table 90]
[0286] [Table 91]
[0287] [Table 92]
[0288] In some embodiments, the exogenous first integrated nucleic acid described herein comprises one of the sequences in Table 14. In some embodiments, the donor nucleic acid described herein comprises a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to any of the recombinant sequences in Table 14.
[0289] [Table 93]
[0290] [Table 94]
[0291] [Table 95]
[0292] [Table 96]
[0293] [Table 97]
[0294] [Table 98]
[0295] [Table 99]
[0296] In some embodiments, the exogenous first integrated nucleic acid described herein may include a nucleic acid sequence recognized or bound by a serine recombinase. In some embodiments, the donor nucleic acid may include a binding site (att). In some embodiments, the donor nucleic acid may include a plurality of binding sites. In some embodiments, the donor nucleic acid may include a binding site on a bacterial portion (attB) incorporating the recombinant sequence. In some embodiments, the donor nucleic acid may include two binding sites on a bacterial portion (attB) incorporating the recombinant sequence. In some embodiments, the donor nucleic acid may include a plurality of binding sites on a bacterial portion (attB) incorporating the recombinant sequence. In some embodiments, the donor nucleic acid may include a binding site on a phage portion (attP) incorporating the recombinant sequence. In some embodiments, the donor nucleic acid may include two binding sites on a phage portion (attP) incorporating the recombinant sequence. In some embodiments, the donor nucleic acid may include a plurality of binding sites on a phage portion (attP) incorporating the recombinant sequence. In some embodiments, the donor nucleic acid may include binding sites on the bacterial portion (attB) that incorporates the recombinant sequence and binding sites on the phage portion (attP) that incorporates the recombinant sequence. In some embodiments, the donor nucleic acid may include multiple binding sites on the bacterial portion (attB) that incorporates the recombinant sequence and multiple binding sites on the phage portion (attP) that incorporate the recombinant sequence.
[0297] In some embodiments, the exogenous first integrated nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a tyrosine recombinase. In some embodiments, the donor nucleic acid may comprise an X (crossover) locus in the P1(LoxP) sequence. In some embodiments, the donor nucleic acid may comprise two X (crossover) loci in the P1(LoxP) sequence. In some embodiments, the donor nucleic acid may comprise multiple X (crossover) loci in the P1(LoxP) sequence. In some embodiments, the donor nucleic acid may comprise a flippase recognition target (FRT) sequence. In some embodiments, the donor nucleic acid may comprise two flippase recognition target (FRT) sequences. In some embodiments, the donor nucleic acid may comprise multiple flippase recognition target (FRT) sequences.
[0298] In some embodiments, the donor nucleic acid includes chemical modifications. In some embodiments, the donor nucleic acid may include a 3'C3 spacer. In some embodiments, the donor nucleic acid may include a 3' inverted dT. In some embodiments, the donor nucleic acid may include 3' phosphorylation. In some embodiments, the donor nucleic acid may include a 3' phosphorothioate bond. In some embodiments, chemical modifications on the donor nucleic acid can protect the donor DNA from nucleases. In some embodiments, the chemical modifications are not incorporated into the genome. In some embodiments, the chemical modifications can be incorporated into the genome.
[0299] Second embedded nucleic acid This specification discloses a second embedded nucleic acid. The second embedded nucleic acid may be included in the compositions, systems, or methods disclosed herein. Some embodiments relate to nucleic acids encoding the second embedded nucleic acid. This specification provides a second embedded nucleic acid that is inserted into a target nucleic acid, such as a host genome, at a locus, such as an embedded recombinant sequence. For example, the second embedded nucleic acid may replace all or part of an embedded recombinant sequence in the target nucleic acid. The second embedded nucleic acid may be referred to as a “modified nucleic acid.” Where a genomic locus is described, the locus may be included, or vice versa. For example, the locus may be part of the host genome or part of a non-genomic nucleic acid. The modified nucleic acid may include DNA. Similarly, the target nucleic acid may include DNA. In some cases, the modified nucleic acid may include RNA, for example, if the target nucleic acid includes RNA. The modified nucleic acid may include any insert, such as a gene or regulatory element, which is embedded into the genomic locus of the target nucleic acid, such as a recombinant sequence. The modified nucleic acid may be embedded whole or partially into the genomic locus of the target nucleic acid, such as a recombinant sequence, by integrase. Modified nucleic acids may contain sequences that are at least partially homologous to the genomic locus. Modified nucleic acids may be single-stranded. Modified nucleic acids may be double-stranded. Modified nucleic acids may be delivered as double-stranded. Modified nucleic acids may be delivered as multiple strands, for example, two strands.
[0300] The second incorporated nucleic acid may not be naturally occurring. The modified nucleic acid may be engineered. The modified nucleic acid may be synthetic. The modified nucleic acid may be pre-synthesized. The modified nucleic acid may be added to a subject or cell. In some embodiments, the modified nucleic acid does not contain a polymerase template.
[0301] In some embodiments, the modified nucleic acid includes a recombinant sequence. The recombinant sequence may be a homolog of any recombinant sequence of the donor nucleic acid. In some embodiments, the homologous recombinant sequence of the modified nucleic acid binds to the recombinant sequence of the donor nucleic acid.
[0302] In some embodiments, the modified nucleic acid comprises multiple recombinant sequences. In some embodiments, the modified nucleic acid comprises multiple homogeneous recombinant sequences. In some embodiments, the modified embedded nucleic acid may comprise at least one recombinant sequence. In some embodiments, the modified nucleic acid may comprise at least two recombinant sequences. In some embodiments, the modified nucleic acid may comprise at least three recombinant sequences. In some embodiments, the modified nucleic acid may comprise at least four recombinant sequences. In some embodiments, the modified nucleic acid may comprise at least five recombinant sequences. In some embodiments, the modified nucleic acid may comprise at least ten recombinant sequences.
[0303] In some embodiments, the second embedded nucleic acid described herein may include a nucleic acid sequence recognized or bound by an integrase. In some embodiments, the second nucleic acid described herein may include a nucleic acid sequence recognized or bound by a recombinase. In some embodiments, the modified nucleic acid described herein may include a nucleic acid sequence recognized or bound by any integrase described herein. In some embodiments, the modified nucleic acid described herein may include a nucleic acid sequence recognized or bound by any integrase listed in Table 8. The modified nucleic acid described herein may include a nucleic acid sequence recognized or bound by an integrase that comprises a polypeptide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to the polypeptide sequence of any one of the integrases in Table 8.
[0304] In some embodiments, the modified nucleic acids described herein may include nucleotide sequences recognized or bound by serine recombinases. The serine recombinases may be, or include, PhiC31 (ΦC31) bacteriophage integrase, Bxb1 mycobacteriophage integrase, Pseudomonas aeruginosa (Pa01) integrase, Nocardia otitidiscaviarum (No67) integrase, or Streptomyces ipomoeae (Si74) integrase. In some embodiments, the modified nucleic acids described herein may include nucleotide sequences recognized or bound by ΦC31 bacteriophage integrase. In some embodiments, the modified nucleic acids described herein may include nucleotide sequences recognized or bound by Bxb1 mycobacterium integrase. The modified nucleic acids described herein may include nucleotide sequences recognized or bound by Pseudomonas aeruginosa (Pa01) integrase. The modified nucleic acids described herein may include nucleotide sequences recognized or bound by Nocardia otitidiscaviarum (No67) integrase. The modified nucleic acids described herein may include nucleotide sequences recognized or bound by Streptomyces ipomoeae (Si74) integrase.
[0305] In some embodiments, the modified nucleic acids described herein may include nucleotide sequences recognized or bound by tyrosine recombinase. In some embodiments, the modified nucleic acids described herein may include nucleotide sequences recognized or bound by Cre recombinase. In some embodiments, the modified nucleic acids described herein may include nucleotide sequences recognized or bound by flippase (Flp).
[0306] In some embodiments, the modified nucleic acids described herein include one of the sequences in Table 13. In some embodiments, the modified nucleic acids described herein include a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more identical to any of the recombinant sequences in Table 13.
[0307] In some embodiments, the second incorporated nucleic acid described herein may include a nucleic acid sequence recognized or bound by a serine recombinase. In some embodiments, the modified nucleic acid may include an attachment site (att). In some embodiments, the modified nucleic acid may include a plurality of binding sites. In some embodiments, the modified nucleic acid may include an attachment site on a bacterial portion (attB) that incorporates the recombinant sequence. In some embodiments, the modified nucleic acid may include two binding sites on a bacterial portion (attB) that incorporates the recombinant sequence. In some embodiments, the modified nucleic acid may include a plurality of binding sites on a bacterial portion (attB) that incorporates the recombinant sequence. In some embodiments, the modified nucleic acid may include a binding site on a phage portion (attP) that incorporates the recombinant sequence. In some embodiments, the modified nucleic acid may include two binding sites on a phage portion (attP) that incorporates the recombinant sequence. In some embodiments, the modified nucleic acid may include a plurality of binding sites on a phage portion (attP) that incorporates the recombinant sequence. In some embodiments, the modified nucleic acid may include an attachment site on the bacterial portion (attB) that incorporates the recombinant sequence and an attachment site on the phage portion (attP) that incorporates the recombinant sequence. In some embodiments, the modified nucleic acid may include a plurality of binding sites on the bacterial portion (attB) that incorporates the recombinant sequence and a plurality of binding sites on the phage portion (attP) that incorporates the recombinant sequence.
[0308] In some embodiments, the second integrated nucleic acid described herein may comprise a nucleic acid sequence recognized or bound by a tyrosine recombinase. In some embodiments, the modified nucleic acid may comprise an X (crossover) locus in the P1(LoxP) sequence. In some embodiments, the modified nucleic acid may comprise two X (crossover) loci in the P1(LoxP) sequence. In some embodiments, the modified nucleic acid may comprise multiple X (crossover) loci in the P1(LoxP) sequence. In some embodiments, the modified nucleic acid may comprise a flippase recognition target (FRT) sequence. In some embodiments, the modified nucleic acid may comprise two flippase recognition target (FRT) sequences. In some embodiments, the modified nucleic acid may comprise multiple flippase recognition target (FRT) sequences.
[0309] In some embodiments, the modified nucleic acid may include an exogenous nucleic acid sequence. In some embodiments, the exogenous nucleic acid sequence may be fully or partially integrated into the genome at a target site, such as an integrase-mediated integration sequence. In some embodiments, the modified nucleic acid may further include a regulatory element. The regulatory element may be a promoter or an enhancer. In some embodiments, the modified nucleic acid may not include a promoter or an enhancer.
[0310] In some embodiments, the modified nucleic acid may include a nucleotide sequence that codes for a gene. In some embodiments, the modified nucleic acid may include a nucleotide sequence that codes for multiple genes. The modified nucleic acid may include a coding sequence. The coding sequence may code for a full-length protein. The modified nucleic acid may include a non-coding sequence. The non-coding sequence may be a recombinant sequence. The non-coding sequence may knock out an endogenous gene. The non-coding sequence may include a regulatory element. The modified nucleic acid may be up to about 10 kb in length. The modified nucleic acid may be up to about 20 kb in length. The modified nucleic acid may be up to about 30 kb in length. The modified nucleic acid may be up to about 40 kb in length. The modified nucleic acid may be up to about 50 kb in length. The modified nucleic acid may be at least about 50 kb in length.
[0311] In some embodiments, the modified nucleic acid may be delivered as a minicircle (Figure 15A). In some embodiments, the modified nucleic acid may be delivered as a plasmid (Figures 15B and 15C). In some embodiments, the modified nucleic acid may be delivered as linear double-stranded DNA (Figure 15D).
[0312] Modified nucleic acids may include methylated nucleotides. Modified nucleic acids may also include unmethylated nucleotides. An example of a methylated nucleotide is a nucleotide containing methylated cytosine. Cytosine may be methylated at the C-5 position of the cytosine ring. An example of an unmethylated nucleotide may include unmethylated cytosine. Unmethylated nucleotides may contain cytosine that is not methylated at the C-5 position of the cytosine ring.
[0313] In some embodiments, the modified nucleic acid may include modified nucleotides. In some embodiments, the modified nucleic acid may include modified DNA bases. In some embodiments, the modified nucleotides include methylation. In some embodiments, the modified nucleic acid may include methylated nucleotides. Non-limiting examples of methylated nucleotides may include methylated cytosine (e.g., 5-mC). In some embodiments, the modified nucleic acid may include N6-methyladenine (6-mA). In some embodiments, the modified DNA base may include 5-hydroxymethylcytosine (5-hmC). In some embodiments, the modified DNA base may include 5-formylcytosine (5-fC). In some embodiments, the modified DNA base may include 5-carboxylcytosine (5-caC).
[0314] In some embodiments, the modified nucleic acid can introduce epigenetic modifications to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce epigenetic modifications of DNA bases to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce methylated DNA bases to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce unmethylated DNA bases to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce methylated cytosine (e.g., 5-mC) to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce N6-methyladenine (6-mA) to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce 5-hydroxymethylcytosine (5-hmC) to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce 5-formylcytosine (5-fC) to the target nucleic acid. In some embodiments, the modified nucleic acid can introduce 5-carboxylcytosine (5-caC) to the target nucleic acid.
[0315] system Systems for nucleic acid editing (also known as gene editing) are described herein. The editing system may include an endonuclease such as an RNA guide endonuclease, a guide nucleic acid, and a donor nucleic acid. Where gene editing is described, the editing is intended to be of a gene, a regulatory element, or any sequence of nucleic acid. The editing is intended to be the integration of a nucleic acid, such as a nucleic acid sequence, into a gene, a regulatory element, or any sequence of nucleic acid. Furthermore, where genome editing, such as genome editing at a gene locus, is described, the editing of nucleic acids that do not involve the genome is also intended to be possible. For example, genome editing may refer to editing the genome of an organism or may include editing of nucleic acids that are not part of the genome. The systems described herein may be used in gene editing methods. Where nucleic acid integration, such as the integration of a nucleic acid sequence at a gene locus, is described, the nucleic acid sequence may be integrated into a second nucleic acid sequence, and it is assumed that the second nucleic acid does not involve the genome. For example, integration may refer to the integration of a nucleic acid sequence into the genome of an organism or to the integration of a nucleic acid sequence into a second nucleic acid sequence that is not part of the genome.
[0316] In some embodiments, systems comprising at least one endonuclease; at least one guide nucleic acid; at least one ligase; at least one donor strand; at least one integrase; at least one exogenous first integration nucleic acid; or a combination thereof are described herein. In some embodiments, the guide nucleic acid guides the endonuclease to the genomic locus to cleave at least one strand of the genomic locus, and after cleavage, the donor strand is ligated by the ligase and thus integrated into the genomic locus. In some embodiments, the system comprises a first endonuclease that forms a complex with the first guide nucleic acid and can be operably coupled with the first ligase; and a second endonuclease that complexes with the second guide nucleic acid and can be operably coupled to the second ligase. In such a system, each of the first and second endonucleases can, respectively, cleave at least one strand of the genomic locus for the integration of the donor strand.
[0317] In some embodiments, the system comprises one, two, three or more endonucleases. In some embodiments, the system comprises one endonuclease. In some embodiments, two endonucleases can be complexed with different guide nucleic acids. In some embodiments, two endonucleases can each be operably ligated to a ligase. In some embodiments, two endonucleases can each be operably coupled to an integrase. In some embodiments, the endonucleases are programmable endonucleases. In some embodiments, the endonucleases comprise an RNA guide endonuclease, and the guide nucleic acid comprises a guide RNA. In some embodiments, the endonucleases comprise a nickase, and the endonucleases cleave only one strand (as opposed to performing a double-strand break). In some embodiments, the endonucleases comprise a localization signal sequence for increasing the accumulation of endonucleases near genomic loci (e.g., in the nucleus). In some embodiments, the endonucleases comprise at least one additional domain. In some embodiments, at least one additional domain is a dimerization domain. In some embodiments, an endonucleases containing dimerization domains can be dimerized with a ligase to form a heterodimer. In some embodiments, at least one additional domain is a functional domain. For example, the functional domain may include a chromatin modification domain or a cell-permeable peptide. In some embodiments, the endonucleases include a linker, which can covalently link the endonucleases to another polypeptide (e.g., a ligase or integrase). In some embodiments, the linker covalently connects the endonucleases to at least one additional domain. In some embodiments, the endonucleases include a tag, which can be used to increase, identify, or purify the expression of the endonucleases.
[0318] In some embodiments, the system comprises one, two, three or more guide nucleic acids. In some embodiments, the system comprises one guide nucleic acid, which can complex with at least one endonuclease. In some embodiments, the system comprises two guide nucleic acids, each of which can complex with at least one endonuclease. In some embodiments, the guide nucleic acid comprises a spacer complementary to a genomic locus in the cell; a scaffold for complexing with at least one endonuclease; a donor binding site at least partially complementary to the donor strand; a flap binding site at or adjacent to the genomic locus that is at least partially identical or complementary to the genomic flap; or a combination thereof. In some embodiments, the guide nucleic acid can direct at least one endonuclease to cleave at least one strand of the genomic locus. In some embodiments, the guide nucleic acid may be at least partially complementary to the donor strand or at least partially complementary to a genomic flap (e.g., a genomic nucleic acid sequence that is replaced and becomes single-stranded when the guide nucleic acid recruits an endonuclease to a genomic locus). In some embodiments, a guide nucleic acid that is at least partially complementary to the donor strand or at least partially complementary to a genomic flap brings the donor strand closer to the cleavage of the genomic locus. In some embodiments, the guide nucleic acid includes a sprint for ligation. In some embodiments, the guide nucleic acid includes at least one nucleic acid modification. In some embodiments, at least one nucleic acid modification includes modifying the backbone, sugars, bases, or combinations thereof of the guide nucleic acid. In some embodiments, at least one nucleic acid modification can increase the guide nucleic acid's resistance to degradation (e.g., against nuclease degradation or hydrolysis). In some embodiments, at least one nucleic acid modification can increase the complexation of the guide nucleic acid to at least one endonuclease. In some embodiments, at least one nucleic acid modification can increase the complexation of the guide nucleic acid with the donor strand.In some embodiments, at least one nucleic acid modification can increase the complexation of guide nucleic acids to genomic loci by being complementary to the genomic flap.
[0319] In some embodiments, the system comprises one, two, three, or more ligases. In some embodiments, the system comprises one ligase. In some embodiments, one ligase is operably coupled with at least one endonuclease, at least one integrase, or both, and the ligase can ligate at least one end of a donor strand to a cleaved genomic locus, thus incorporating the donor strand into the genomic locus. In some embodiments, the system comprises two ligases. In some embodiments, the two ligases can each be functionally coupled to a different endonuclease, a different integrase, or each can be functionally coupled to a different endonuclease and a different integrase, where the genomic locus is cleaved at two or more positions. In such scenarios, the two ligases can each ligate one end of a donor strand to a cleaved genomic locus, thus incorporating the donor strand into the genomic locus. In some embodiments, the ligase comprises a ligase capable of ligating a substrate containing DNA. In some embodiments, the ligase comprises a ligase capable of ligating a substrate containing a DNA splint. In some embodiments, the ligase comprises a ligase capable of ligating a substrate containing DNA / RNA. In some embodiments, the ligase comprises a ligase capable of ligating a substrate containing an RNA splint. In some embodiments, the ligase comprises a ligase capable of ligating a substrate containing a gRNA splint. In some embodiments, the ligase comprises at least one additional domain. In some embodiments, at least one additional domain is a dimerization domain. In some embodiments, the ligase comprising the dimerization domain can be dimerized with an endonuclease to form a heterodimer. In some embodiments, at least one additional domain is a functional domain. For example, the functional domain may include a chromatin modification domain or a cell-permeable peptide.In some embodiments, the ligase includes a linker that can covalently link the ligase to another polypeptide (e.g., an endonuclease). In some embodiments, the linker covalently links the ligase to at least one further domain. In some embodiments, the ligase includes a tag that can be used to increase, identify, or purify the expression of the ligase.
[0320] A fusion protein comprising an RNA-guided endonuclease fused to a ligase is disclosed herein. Table 15 shows non-limiting examples of polypeptides and nucleic acid sequences encoding a fusion polypeptide comprising components of the system described herein (e.g., an endonuclease fused to a ligase). Sequence ID No. 125 shows a nucleic acid sequence encoding the polypeptide sequence of Sequence ID No. 126, which shows a fusion protein (NLS-nCas9-linker-hLIG1(119-919)-bpNLS) comprising an N-terminal NLS followed by a linker and a subsequent C-terminal NLS, and an endonuclease (nCas9) covalently bound to a ligase (hLIG1, 119-919 fragment). Sequence ID 127 shows a nucleic acid sequence encoding the polypeptide sequence of Sequence ID 128, which represents a fusion protein (NLS-nCas9-linker-hLIG1(233-919)-bpNLS) containing an endonuclease (nCas9) covalently bound to a ligase (hLIG1, 233-919 fragment) via an N-terminal NLS followed by a linker and a C-terminal NLS. Sequence ID 129 shows a nucleic acid sequence encoding the polypeptide sequence of Sequence ID 130, which represents a fusion protein (NLS-nCas9-linker-sprintR-bpNLS) containing an endonuclease (nCas9) covalently bound to a ligase (sprintR) via an N-terminal NLS followed by a linker and a C-terminal NLS. Sequence ID 131 shows a nucleic acid sequence encoding the polypeptide sequence of Sequence ID 132, which shows a fusion protein (NLS-nCas9-linker-T4LIG-bpNLS) containing an N-terminal NLS followed by an endonuclease (nCas9) covalently bound to a ligase (T4LIG) via a linker and a subsequent C-terminal NLS. Sequence ID 133 shows a nucleic acid sequence encoding an endonuclease (nCas9) containing an N-terminal NLS and a leucine zipper (LZ) dimerization domain.Sequence ID 134 shows a fusion protein (NLS1-hFEN1-linker1-nCas9-linker2-nCas9-NLS2) containing an exonuclease (hFEN1) with a first NLS (NLS1) at its N-terminus, which is then covalently bound to an endonuclease (nCas9) via linker 1, and further covalently bound to a ligase (T4LIG) via linker 2, followed by a second NLS (NLS2) at its C-terminus. Sequence ID 135 shows a fusion protein (NLS1-hFEN1-linker1-T4LIG-linker2-nCas9-NLS2) containing an exonuclease (hFEN1) with an N-terminal NLS1, which is then covalently bound to a ligase (T4LIG) via linker 1, and further covalently bound to an endonuclease (nCas9) via linker 2 and the subsequent C-terminal NLS2. Sequence ID 136 shows a fusion protein (NLS1-nCas9-linker1-hFEN1-linker2-T4LIG-NLS2) containing an N-terminal NLS1, followed by an endonuclease (nCas9) covalently bound to an exonuclease (hFEN1) via linker1, and further covalently bound to a ligase (T4LIG) via linker2 and then C-terminal NLS2. Sequence ID 137 shows a fusion protein (NLS1-T4LIG-linker1-nCas9-linker2-hFEN1-NLS2) containing an N-terminal NLS1, followed by an endonuclease (nCas9) covalently bound to an exonuclease (hFEN1) via linker1, and further covalently bound to an exonuclease (hFEN1) via linker2 and then C-terminal NLS2. Sequence ID 138 shows a fusion protein (NLS1-nCas9-linker1-T4LIG-linker2-hFEN1-NLS2) containing an endonuclease (nCas9) that is covalently bound to a ligase (T4LIG) via N-terminal NLS1 and subsequent linker1, and further covalently bound to an exonuclease (hFEN1) via linker2 and subsequent C-terminal NLS2.Sequence ID 139 shows a fusion protein (NLS1-T4LIG-linker1-hFEN1-linker2-nCas9-NLS2) containing a ligase (T4LIG) covalently bound to an exonuclease (hFEN1) via N-terminal NLS1 and subsequent linker1, and further covalently bound to an endonuclease (nCas9) via linker2 and subsequent C-terminal NLS2. Sequence ID 140 shows a fusion protein (NLS1-T5 EXO-linker1-nCas9-linker2-T4LIG-NLS2) containing an exonuclease (EXO) covalently bound to an endonuclease (nCas9) via N-terminal NLS1 and subsequent linker1, and further covalently bound to a ligase (T4LIG) via linker2 and subsequent C-terminal NLS2. Sequence ID 141 shows the nucleic acid sequence encoding a fusion protein (LZ-Sprint R-bpNLS) containing a ligase (Sprint R) fused to a dimerizing domain (LZ) and NLS. Sequence ID 142 shows the nucleic acid sequence encoding a fusion protein (LZ-T4LIG-bpNLS) containing a ligase (T4LIG) fused to a dimerizing domain (LZ) and NLS. Sequence ID 143 shows the nucleic acid sequence encoding a fusion protein (LZ-hLIG 233-919 polypeptide fragment-bpNLS) containing a ligase (hLIG) fused to a dimerizing domain (LZ) and NLS. Sequence ID 144 shows the nucleic acid sequence encoding a fusion protein (LZ-hLIG1 119-919 polypeptide fragment-bpNLS) containing a ligase (hLIG) fused to a dimerizing domain (LZ) and NLS. Sequence ID 145 shows the nucleic acid sequence encoding a fusion protein (T4-LZ) containing a ligase (T4) fused to a dimerizing domain (LZ) and NLS. Sequence ID 146 shows the nucleic acid sequence encoding a fusion protein (LZ-hLIG4(1-620)) containing a ligase polypeptide fragment (hLIG4(1-620)) fused to a dimerizing domain (LZ) and NLS. Sequence ID 147 shows the nucleic acid sequence encoding a fusion protein (LZ-nCas9) containing an endonuclease (nCas9) fused to a dimerizing domain (LZ) and NLS.Sequence ID 148 shows the nucleic acid sequence encoding a fusion protein (Sprint R-LZ) containing a ligase (Sprint R) fused to a dimerization domain (LZ) and NLS. Sequence ID 149 shows the nucleic acid sequence encoding a fusion protein (hLIG4(1-620)-LZ) containing a ligase polypeptide fragment (hLIG4(1-620)) fused to a dimerization domain (LZ) and NLS. Sequence ID 150 shows the nucleic acid sequence encoding a fusion protein (nCas9-hLIG4(1-620)) containing a ligase polypeptide fragment (hLIG4(1-620)) fused to an endonuclease (nCas9). Sequence ID 151 shows the nucleic acid sequence encoding a fusion protein (T4-nCas9) containing a ligase (T4) fused to an endonuclease (nCas9) and NLS. Sequence ID 152 shows the nucleic acid sequence encoding a fusion protein (Sprint R-nCas9) containing a ligase (Sprint R) fused to an endonuclease (nCas9) and NLS. Sequence ID 153 shows the nucleic acid sequence encoding a fusion protein (hLIG4(1-620)-nCas9) containing a ligase polypeptide fragment (hLIG4(1-620)) fused to an endonuclease (nCas9) and NLS.
[0321] [Table 100]
[0322] [Table 101]
[0323] [Table 102]
[0324] [Table 103]
[0325] [Table 104]
[0326] Table 105
[0327] Table 106
[0328] Table 107
[0329] Table 108
[0330] Table 109
[0331] Table 110
[0332] Table 111
[0333] Table 112
[0334] Table 113
[0335] Table 114
[0336] Table 115
[0337] Table 116
[0338] Table 117
[0339] Table 118
[0340] Table 119
[0341] Table 120
[0342] Table 121
[0343] Table 122
[0344]
Table 123
[0345] Table 124
[0346] Table 125
[0347] Table 126
[0348] Table 127
[0349] Table 128
[0350] Table 129
[0351] Table 130
[0352] Table 131
[0353] Table 132
[0354] Table 133
[0355] Table 134
[0356] [Table 135]
[0357] [Table 136]
[0358] [Table 137]
[0359] [Table 138]
[0360] [Table 139]
[0361] A protein complex comprising a ligase-bound RNA guide endonuclease is disclosed herein. The endonuclease and ligase may be bound to each other via a heterodimerizing domain. The heterodimerizing domain may comprise one or more of the following: a leucine zipper, a PDZ domain, streptavidin, a streptavidin-binding protein, a foldon domain, a hydrophobic polypeptide, an antibody bound to Cas nickase, or an antibody bound to ligase, or one or more binding fragments thereof.
[0362] A protein complex comprising an RNA-guided endonuclease bound to an integrase is disclosed herein. The endonuclease and integrase may be bound to each other via a heterodimerizing domain. The heterodimerizing domain may comprise one or more of the following: a leucine zipper, a PDZ domain, streptavidin, and a streptavidin-binding protein, a foldon domain, a hydrophobic polypeptide, an antibody bound to Cas nickase, or an antibody bound to integrase, or one or more binding fragments thereof.
[0363] Protein complexes comprising ligases bound to integrases are disclosed herein. Endonucleases and integrases may be bound to each other via heterodimerization domains. The heterodimerization domain may comprise one or more of the following: a leucine zipper, a PDZ domain, streptavidin, and streptavidin-binding protein, a Foldon domain, a hydrophobic polypeptide, an antibody bound to a ligase, or an antibody bound to an integrase, or one or more binding fragments thereof.
[0364] Protein complexes comprising RNA guide endonucleases bound to ligases and integrases; ligases and integrases bound to RNA guide endonucleases; and integrases bound to RNA guide endonucleases and ligases are disclosed herein.
[0365] In some embodiments, the system includes at least one donor strand. In some embodiments, the donor strand includes a nucleic acid sequence that is at least partially homologous to a genomic locus targeted by at least one guide nucleic acid. In some embodiments, the donor strand includes a nucleic acid sequence that is not homologous to a genomic locus targeted by at least one guide nucleic acid. In some embodiments, the donor strand is single-stranded or double-stranded nucleic acid. In some embodiments, a donor strand containing a double-stranded nucleic acid includes at least one overhang. In some embodiments, the overhang includes a guide binding site that is at least partially complementary to the guide nucleic acid. In some embodiments, the overhang includes a genomic flap binding site that is at least partially identical or complementary to a genomic locus or an adjacent genomic flap. In some embodiments, the donor strand comprises two overhangs, the first overhang comprising a first guide binding site at least partially complementary to a first guide nucleic acid, or the first genomic flap binding site being at least partially identical or complementary to the first genomic flap at or adjacent to a genomic locus, the second overhang comprising a second guide binding site at least partially complementary to a second guide nucleic acid, or the second genomic flap binding site being at least partially identical or complementary to the second genomic flap at or adjacent to a genomic locus. In some embodiments, the donor strand modifies at least one gene mutation at at least one genomic locus. In some embodiments, the donor strand comprises a coding sequence. In some embodiments, the coding sequence encodes a full-length protein or a fragment thereof. In some embodiments, the donor strand comprises a non-coding sequence. In some embodiments, the non-coding sequence comprises a recombinant sequence. In some embodiments, the non-coding sequence knocks out an endogenous gene. In some embodiments, the non-coded array includes modifier elements.
[0366] In some embodiments, the system includes a nuclease. The nuclease may be heterogeneous. In some embodiments, the nuclease includes an exonuclease for digesting the genomic flap. In some embodiments, the exonuclease is a 5' exonuclease. Non-limiting examples of exonucleases may include human flap endonuclease 1 (hFEN1), human exonuclease 5 (hEXO5), T5 exonuclease, T7 exonuclease, exonuclease VIII, flap endonuclease domain of Escherichia coli PolI, RecJF, lambda exonuclease, Xni (ExoIXI), SaFEN (Staphylococcus aureus FEN), nuclease BAL-31, or fragments thereof. In some embodiments, the exonuclease includes the exonucleases listed in Table 16. In some embodiments, the exonuclease contains at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or more polypeptide sequences identical to any one of the exonucleases in Table 16.
[0367] [Table 140]
[0368] [Table 141]
[0369] [Table 142]
[0370] [Table 143]
[0371] In some embodiments, the system includes at least one additional endonuclease different from the at least one programmable endonuclease described herein. In some embodiments, the at least one additional endonuclease is capable of digesting genomic flaps.
[0372] In some embodiments, the system includes a dominant-negative MMR peptide for improving genome editing capabilities, particularly in cells overexpressing the MMR pathway. In some embodiments, the dominant-negative MMR peptide may be delivered by fusion (e.g., fusion with any component of the system described herein), recruitment, or as a separate peptide. Table 17 lists non-limiting examples of MMR peptide sequences.
[0373] [Table 144]
[0374] [Table 145]
[0375] [Table 146]
[0376] [Table 147]
[0377] The system may relate to a one-sided replacer 1. Some embodiments include (a) at least one RNA guide endonuclease; (b) at least one guide nucleic acid comprising (i) a spacer complementary to an intracellular genomic locus, (ii) a scaffold for forming a complex with the at least one RNA guide endonuclease, (iii) an arbitrary donor binding site at least partially complementary to an exogenous first integrated nucleic acid, and (iv) a flap binding site at or adjacent to a genomic locus that is at least partially identical or complementary to a genomic flap. The present invention comprises an exogenous first integrated nucleic acid, which optionally includes at least one guide nucleic acid; (c) at least one DNA ligase; and (d) a guide binding site at least partially complementary to the guide nucleic acid, wherein at least one RNA guide endonuclease cleaves at least one strand of a genomic locus, and at least one DNA ligase ligates the end of the exogenous first integrated nucleic acid to a genomic flap site, thereby replacing the region of the genomic locus with the exogenous first integrated nucleic acid within the cell. The exogenous first integrated nucleic acid may include single-stranded DNA.
[0378] The system may also be associated with bilateral replacers 1.Some embodiments include: (a) at least one RNA guide endonuclease comprising a first RNA guide endonuclease and an optional second RNA guide endonuclease; (b) at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, wherein the first guide nucleic acid comprises (i) a first spacer complementary to a first region of an intracellular genomic locus, (ii) a first scaffold for forming a complex with the first RNA guide endonuclease, and (iii) an optional first RNA guide endonuclease at least partially complementary to an exogenous first integrated nucleic acid. (iv) a donor binding site, and a first flap binding site at or adjacent to a genomic locus that is at least partially identical or complementary to the first genomic flap, wherein the second guide nucleic acid includes (i) a second spacer complementary to the second region of the genomic locus in the cell, (ii) a second scaffold for forming a complex with the first or second RNA guide endonuclease, (iii) any second donor binding site at least partially complementary to the exogenous first integrated nucleic acid, and (iv) at or adjacent to the genomic locus (c) at least one guide nucleic acid adjacent to the locus and comprising a second flap binding site at least partially identical or complementary to the second genomic flap; (d) at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase; (i) the first strand comprising an optional first guide binding site at least partially complementary to the first guide nucleic acid; and (ii) the second strand comprising an optional second guide binding site at least partially complementary to the second guide nucleic acid. The present invention comprises at least one exogenous first integrated nucleic acid, wherein a first RNA guide endonuclease and / or a second RNA guide endonuclease each cleave at least one strand of an intracellular genomic locus; a first DNA ligase ligates the end of the first strand of an exogenous first integrated nucleic acid to a first genomic flap; and the first or second DNA ligase ligates the end of the second strand of the exogenous first integrated nucleic acid to a second genomic flap, thereby replacing the region of the genomic locus with the intracellular exogenous first integrated nucleic acid.The first exogenous integrated nucleic acid may include a double-stranded DNA region. The first exogenous integrated nucleic acid may optionally include a 5' overhang containing a first guide binding site. The first exogenous integrated nucleic acid may optionally include a 5' overhang containing a second guide binding site.
[0379] This system may also relate to a one-sided replacer 2. Some embodiments include: (a) at least one RNA guide endonuclease; (b) at least one guide nucleic acid comprising (i) a spacer complementary to an intracellular genomic locus, (ii) a scaffold for forming a complex with at least one RNA guide endonuclease, and (iii) an arbitrary donor binding site at least partially complementary to an exogenous first integrated nucleic acid; (c) at least one DNA ligase; and (d) an exogenous first integrated nucleic acid comprising (i) a small portion of the guide nucleic acid. (ii) comprising any guide-binding site that is at least partially complementary to the genomic flap, and comprising a flap-binding site at or adjacent to the genomic locus, wherein at least one RNA guide endonuclease cleaves at least one strand of the genomic locus; and comprising an exogenous first integrated nucleic acid, wherein at least one DNA ligase ligates the end of the exogenous first integrated nucleic acid to the genomic flap, thereby replacing the region of the genomic locus with the intracellular exogenous first integrated nucleic acid. The exogenous first integrated nucleic acid may comprise DNA with a 3' overhang. The 3' overhang may comprise a guide-binding site. The 3' overhang may comprise a flap-binding site. At least one DNA ligase may ligate the strand of the exogenous first integrated nucleic acid to the genomic nucleic acid sequence.
[0380] The system may also relate to a bilateral replacer 2. Some embodiments include (a) at least one RNA guide endonuclease comprising a first RNA guide endonuclease and an optional second RNA guide endonuclease; (b) at least one guide nucleic acid comprising a first guide nucleic acid and a second guide nucleic acid, wherein the first guide nucleic acid comprises (i) a first spacer complementary to a first region of an intracellular genomic locus, (ii) a first scaffold for forming a complex with the first RNA guide endonuclease, and (iii) at least one exogenous first integrated nucleic acid. The second guide nucleic acid comprises (i) a second spacer complementary to a second region of a genomic locus in a cell, (ii) a second scaffold for forming a complex with a first or second RNA guide endonuclease, and (iii) at least one DNA ligase comprising a first DNA ligase and an optional second DNA ligase, the exogenous first integrated nucleic acid comprising a first strand and A genomic flap comprising two strands, wherein the first strand comprises any first guide binding site at least partially complementary to the first guide nucleic acid; the second strand comprises any second binding site at least partially complementary to the second guide nucleic acid; the first strand comprises a first flap binding site at or adjacent to a genomic locus that is at least partially identical or complementary to the first genomic flap; the second strand comprises a first flap binding site at or adjacent to a genomic locus that is at least partially identical or complementary to the second genomic flap The DNA ligase comprises a second flap binding site; the first RNA guide endonuclease and / or the second RNA guide endonuclease each cleave at least one strand of the intracellular genomic locus; the first DNA ligase ligates the end of the first strand of the exogenous first integrated nucleic acid to the first genomic flap; and the first or second DNA ligase ligates the end of the second strand of the exogenous first integrated nucleic acid to the second genomic flap, thereby replacing the region of the genomic locus with the intracellular integrated nucleic acid. The exogenous first integrated nucleic acid may include a double-stranded DNA region.The double-stranded DNA may optionally include a first guide binding site and a 3' overhang including a first flap binding site. The double-stranded DNA may optionally include a second guide binding site and a 3' overhang including a second flap binding site.
[0381] In this system, at least one RNA guide endonuclease may contain a Cas protein or a functional fragment thereof. The Cas protein or a functional fragment thereof may contain nickase activity. At least one RNA guide endonuclease may contain a Cas9 nickase or a functional fragment thereof. At least one DNA ligase may ligate nucleic acids bound to DNA. At least one DNA ligase may ligate nucleic acids bound to RNA. At least one DNA ligase may contain a PBCV-1 DNA ligase. At least one DNA ligase may be operablely coupled with at least one RNA guide endonuclease. At least one DNA ligase may be fused to at least one RNA guide endonuclease as a fusion polypeptide. At least one RNA guide endonuclease and at least one DNA ligase may contain a heterodimer domain. At least one RNA guide endonuclease and at least one DNA ligase may form a heterodimer via a heterodimer domain. At least one RNA guide endonuclease may contain a linker. The linker may ligate a Cas protein or a functional fragment thereof to a heterodimer domain. At least one RNA guide endonuclease may contain a localization signal sequence. At least one DNA ligase may contain a localization signal sequence. The localization signal sequence may contain a nuclear localization sequence (NLS). At least one RNA guide endonuclease or at least one DNA ligase may be directed to the nucleus of a cell by the NLS. At least one embedded nucleic acid, e.g., a donor nucleic acid or a modified nucleic acid, may modify at least one gene mutation at at least one genomic locus. At least one embedded nucleic acid may insert a coding sequence. The coding sequence may code for a full-length protein. At least one embedded nucleic acid may insert a non-coding sequence. The non-coding sequence may contain a recombinant sequence. The non-coding sequence may knock out an endogenous gene. The non-coding sequence may contain a regulatory element. The system may contain additional nucleases.Nucleases may include exonucleases for digesting genomic flaps. Nucleases may include human flap endonuclease 1 (hFEN1), human exonuclease 5 (hEXO5), T5 exonuclease, T7 exonuclease, exonuclease VIII, flap endonuclease domain of Escherichia coli (E. coli) PolI, RecJF, lambda exonuclease, Xni (ExoIXI), SaFEN (Staphylococcus aureus FEN), nuclease BAL-31, or fragments thereof. Heterogeneous nucleases may include endonucleases for digesting genomic flaps, and the endonucleases may differ from at least one RNA-guided endonuclease. At least one RNA-guided endonuclease may contain at least one additional functional domain. At least one further functional domain may include a chromatin modification domain. At least one additional functional domain may include a cell-permeable peptide. At least one guide nucleic acid may include at least one nucleic acid modification. At least one nucleic acid modification may include modifications to the backbone, sugars, bases, or combinations thereof. At least one RNA guide endonuclease may be complexed with at least one guide nucleic acid. At least one guide nucleic acid may be complexed with an exogenous first integrated nucleic acid. At least one RNA guide endonuclease, at least one guide nucleic acid, at least one DNA ligase, an exogenous first integrated nucleic acid, or a combination thereof may be encoded by a polynucleotide. The polynucleotide may include mRNA. The polynucleotide may include a vector. The vector may include a viral vector. At least one RNA guide endonuclease, at least one guide nucleic acid, at least one DNA ligase, an exogenous first integrated nucleic acid, or a combination thereof may be encapsulated by at least one lipid nanoparticle. Cells may include bacterial cells or prokaryotic cells. Cells may include prokaryotic cells. Prokaryotic cells may include bacterial cells. Editing may be performed in the cytoplasm of bacterial cells. Cells may include eukaryotic cells.Eukaryotic cells may include animal cells or plant cells. Eukaryotic cells may include plant cells. Eukaryotic cells may include animal cells. Eukaryotic cells may include mammalian cells. Editing may occur in the cytoplasm of eukaryotic cells. Editing may occur in the nucleus of eukaryotic cells. The system, or any aspect of the system, may be included in a composition or in cells such as a cell line.
[0382] Some embodiments relate to systems comprising nucleic acids. The system may include a guide nucleic acid, an embedded nucleic acid, or a combination thereof. Some embodiments relate to nucleic acid systems. This system may include a guide nucleic acid system. This system may include a nucleic acid embedding system. The nucleic acid system may further include other embodiments such as additional nucleic acid or non-nucleic acid components.
[0383] The nucleic acid system may include a guide nucleic acid. The guide nucleic acid may include a spacer. The spacer may be complementary to the region of the target nucleic acid's locus (e.g., a genomic locus), such as a genome strand. The target nucleic acid may be present in a cell. The genome strand may be present in a cell. The target nucleic acid may be in vitro. The guide nucleic acid may include a scaffold. The scaffold may form a complex with an endonuclease, such as an RNA guide endonuclease. The guide nucleic acid may include a flap binding site. The flap binding site may be complementary to or at least partially complementary to a flap, such as a genomic flap. The flap binding site may be identical to or at least partially identical to a flap, such as a genomic flap. The flap may be at its location. The flap may be adjacent to the trajectory. The guide nucleic acid may include a donor binding site. The donor binding site may be complementary to the integrated nucleic acid. The donor binding site may be partially complementary to the integrated nucleic acid. The donor binding site may be complementary to the sprint nucleic acid. The donor binding site may be partially complementary to the sprint nucleic acid. The components of the guide nucleic acid may be contained within a single guide nucleic acid. Two or more guide nucleic acids may be used. The components of the guide nucleic acid may be contained together within multiple guide nucleic acids. The components of the guide nucleic acid may be divided among multiple guide nucleic acids.
[0384] The nucleic acid system may include an exogenous first integrated nucleic acid. The exogenous first integrated nucleic acid may include a ligated 5' end. The 5' end may be ligated to a 3' end. The 3' end may belong to a target nucleic acid strand (e.g., a genome strand). The 3' end may be produced by an endonuclease such as an RNA guide endonuclease. The exogenous first integrated nucleic acid may include a 5' end that is ligated to the 3' end of a genome strand produced by an RNA guide endonuclease. The components of the exogenous first integrated nucleic acid may be contained in one or two complementary strands. The components of the exogenous first integrated nucleic acid may be contained in one donor nucleic acid. Two or more integrated nucleic acids may be used. The components of a donor nucleic acid may be contained together in multiple donor nucleic acids. The components of the exogenous first integrated nucleic acid may be divided among multiple integrated nucleic acids.
[0385] A nucleic acid system may include a sprint nucleic acid (also called a "sprint strand"). A sprint strand may hybridize to two nucleic acids, each containing a terminal to be ligated. A sprint nucleic acid may include a flap binding site. The flap binding site may be complementary to the flap. The flap binding site may be partially complementary to the flap. The flap binding site may be identical to the flap. The flap binding site may be partially identical to the flap. The flap may be located at the site of the target nucleic acid. The flap may be adjacent to the locus of the target nucleic acid. The flap may be a genomic flap. The locus may be a genomic locus. The flap binding site may be at least partially identical or complementary to the genomic flap, either at or adjacent to the genomic locus. A sprint nucleic acid may include a guide binding site. The guide binding site may be complementary to the guide nucleic acid. The guide binding site may be partially complementary to the guide nucleic acid. A single sprint nucleic acid may contain components of a sprint nucleic acid. More than one sprint nucleic acid may be used. The sprint nucleic acid may contain a donor binding site. The donor binding site may be complementary to the donor nucleic acid. The donor binding site may be partially complementary to the donor nucleic acid.
[0386] The sprint strand may be DNA or may contain DNA. The sprint strand may be RNA or may contain RNA. The sprint nucleic acid may be included as part of the exogenous first integrated nucleic acid. The sprint nucleic acid may be included as a strand of the double-stranded exogenous first integrated nucleic acid. The sprint nucleic acid may be included as part of the guide nucleic acid.
[0387] The nucleic acid system comprises (a) a guide nucleic acid comprising (i) a spacer complementary to the region of the genomic locus of the genome strand, (ii) a scaffold for forming a complex with an RNA guide endonuclease, (iii) an optional donor binding site at least partially complementary to an exogenous first integrated nucleic acid, and (iv) a flap binding site at or adjacent to the genomic locus that is at least partially identical or complementary to the genomic flap; and (b) an exogenous first integrated nucleic acid comprising a 5' end ligated to the 3' end of the genome strand produced by the RNA guide endonuclease. The components of (i), (ii), (iii), or (iv) may be contained in a single guide nucleic acid, or may be divided among multiple guide nucleic acids, or may be contained collectively. 【038...
Claims
1. Endonuclease; DNA ligase; An embedded nucleic acid, wherein the embedded nucleic acid is configured to be ligated by the ligase to the chain breaks generated by the endonuclease in the target nucleic acid; and A sprint nucleic acid configured to have the 5' end of the embedded nucleic acid positioned at the 3' end of the strand break of the target nucleic acid generated by the endonuclease. A system that includes this.
2. The system according to claim 1, wherein the system includes a guide RNA and the endonuclease includes an RNA guide endonuclease.
3. The system according to claim 1, wherein the endonuclease is coupled to the DNA ligase.
4. The system according to claim 3, wherein the coupling is non-covalent.
5. The system according to claim 4, wherein the endonuclease bound to the DNA ligase comprises a first polypeptide comprising at least a portion of the endonuclease and a second polypeptide comprising at least a portion of the DNA ligase, and the first and second polypeptides are linked non-covalently.
6. The system according to claim 5, wherein the first polypeptide comprises a first heterodimerizing domain that binds to a second heterodimerizing domain, and the second polypeptide comprises the second heterodimerizing domain.
7. The system according to claim 6, wherein the heterodimer domain comprises a leucine zipper, a PDZ domain, streptavidin, a streptavidin-binding protein, a foldon domain, a hydrophobic moiety, or a functional binding fragment thereof.
8. The system according to claim 3, wherein the coupling is a covalent bond.
9. The system according to claim 4, comprising a fusion protein containing the endonuclease and the DNA ligase.
10. The system according to claim 9, wherein the endonuclease is at the amino (N) terminus relative to the DNA ligase in the fusion protein.
11. The system according to claim 9, wherein the endonuclease is carboxyl (C)-terminus relative to the DNA ligase in the fusion protein.
12. The system according to claim 9, wherein the fusion protein includes a linker comprising 1 to 100 amino acids that link the endonuclease and the DNA ligase.
13. The system according to claim 7, wherein the first polypeptide comprises a first intein bound to a second intein, and the second polypeptide comprises the second intein.
14. The system according to claim 3, wherein the ligase includes a hairpin binding motif, and the endonuclease and the DNA ligase are coupled with a nucleic acid that includes a scaffold that binds to the endonuclease and a hairpin that binds to the hairpin binding motif.
15. The system according to claim 14, wherein the hairpin binding motif comprises an MS2 coat protein (MCP) peptide, and the hairpin comprises an MS2 hairpin.
16. The system according to claim 3, wherein the endonuclease and the DNA ligase are coupled with a heterobifunctional molecule containing an endonuclease-binding domain and a DNA ligase-binding domain.
17. The system according to claim 16, wherein the heterobifunctional molecule includes a small molecule.
18. The system according to claim 1, wherein the sprint is part of the guide nucleic acid, or the chain of the incorporated nucleic acid includes the sprint.
19. The system according to claim 1, wherein the sprint nucleic acid comprises a donor nucleic acid binding site (DBS) at least partially complementary to the embedded nucleic acid, a flap binding site (FBS) at least partially complementary to the target nucleic acid, and a guide binding site (GBS) at least partially complementary to the guide nucleic acid.
20. A second embedded nucleic acid, configured to be ligated by the ligase to a chain break generated by the endonuclease in a second target nucleic acid chain; and A second sprint nucleic acid configured to position the 5' end of the second embedded nucleic acid at the 3' end of the strand break of the second strand of the target nucleic acid generated by the endonuclease. The system according to claim 1, including the following:
21. The system according to claim 20, wherein the system includes a guide RNA and the endonuclease includes an RNA guide endonuclease.
22. The system according to claim 20, wherein the second embedded nucleic acid is at least partially complementary to the first embedded nucleic acid.
23. The system according to claim 20, wherein the DBS of the first sprint and the DBS of the second sprint overlap by at least 10 bp, or at least 15 bp, or at least 20 bp, or at least 30 bp, or at least 40 bp.
24. A system comprising cells containing a heterologous RNA guide endonuclease, a DNA ligase, and an integrated recombinant sequence, wherein the integrated nucleic acid is configured to be ligated by the DNA ligase to strand breaks generated by the endonuclease in a target nucleic acid.
25. The system according to claim 24, wherein the DNA ligase is endogenous to the cell.
26. The system according to claim 24, wherein the DNA ligase is a heterogeneous DNA ligase.
27. The system according to any one of claims 1 to 26, wherein the endonuclease comprises an RNA guide endonuclease.
28. The system according to claim 27, wherein the endonuclease comprises a class II CRISPR / Cas endonuclease.
29. The system according to claim 28, wherein the endonuclease comprises Cas9 endonuclease.
30. The system according to claim 27, wherein the endonuclease comprises nicasse.
31. The system according to claim 28, wherein the endonuclease comprises Cas9 niccase.
32. The system according to any one of claims 1 to 31, wherein the endonuclease or DNA ligase comprises a nuclear localization signal, a chromatin modification domain, a cell-permeable peptide, a tag polypeptide, a nucleic acid binding domain, or streptavidin.
33. The system according to any one of claims 1 to 32, further comprising a guide nucleic acid.
34. The system according to any one of claims 1 to 33, wherein the embedded nucleic acid includes at least one binding site (att) for incorporating a recombinant sequence.
35. The system according to claim 34, wherein the incorporated nucleic acid includes a binding site on the bacterial portion (attB) incorporating the recombinant sequence.
36. The system according to claim 34, wherein the incorporated nucleic acid includes two attachment sites on the bacterial portion (attB) incorporating the recombinant sequence.
37. The system according to claim 34, wherein the incorporated nucleic acid includes a binding site on the phage portion (attP) into which the recombinant sequence is incorporated.
38. The system according to claim 34, wherein the incorporated nucleic acid includes two binding sites on the phage portion (attP) that incorporates the recombinant sequence.
39. The system according to claim 34, wherein the incorporated nucleic acid includes one attachment site on the bacterial portion (attB) incorporating the recombinant sequence and one attachment site on the phage portion (attP) incorporating the recombinant sequence.
40. Integrase and At least one congenerally recombinant sequence; and exogenous nucleic acid A second embedded nucleic acid including It further includes, The at least one homologous recombinant sequence of the second embedded nucleic acid is coupled with the at least one embedded recombinant sequence of the embedded nucleic acid, The integrase incorporates all or part of the second integrated nucleic acid into the target nucleic acid in the integrated recombinant sequence. The system according to any one of claims 1 to 39.
41. The system according to claim 40, wherein the second integrated nucleic acid comprises at least one binding site (att) homologous recombinant sequence.
42. The system according to claim 41, wherein the second incorporated nucleic acid includes a binding site on the bacterial portion (attB) homologous recombinant sequence.
43. The system according to claim 41, wherein the second incorporated nucleic acid includes two binding sites on the bacterial portion (attB) homologous recombinant sequence.
44. The system according to claim 41, wherein the second integrated nucleic acid includes a binding site on the phage portion (attP) homologous recombinant sequence.
45. The system according to claim 41, wherein the second integrated nucleic acid includes two binding sites on the phage portion (attP) homologous recombinant sequence.
46. The system according to claim 41, wherein the second incorporated nucleic acid includes one attachment site on the bacterial portion (attB) incorporating the recombinant sequence and one attachment site on the phage portion (attP) incorporating the recombinant sequence.
47. The system according to claim 40, wherein the integrase is coupled to the endonuclease or the ligase.
48. The system according to claim 40, wherein the integrase comprises serine integrase.
49. The system according to claim 48, wherein the serine integrase comprises PhiC31 bacteriophage integrase, Bxb1 mycobacteriophage integrase, Pseudomonas aeruginosa integrase (Pa01), Nocardia otitidiscaviarum integrase (No67), or Streptomyces ipomoeae integrase (Si74).
50. The system according to claim 40, wherein the integrase is coupled to a recombination direction factor (RDF).
51. The system according to any one of paragraphs 1 to 50, wherein the embedded nucleic acid or the second embedded nucleic acid includes a modified nucleotide.
52. The system according to claim 51, wherein the modified nucleotide includes a methylated nucleotide.
53. The system according to claim 51, wherein the modified nucleotide includes methylated cytosine (e.g., 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
54. The system according to claim 53, wherein the modified nucleotide comprises methylated cytosine 5-mC.
55. The system according to claim 51, wherein the modified nucleotide includes a 5' reverse dideoxy-T, a 3' phosphorylation, a 3'C3 spacer, a 3' reverse dT, or a combination thereof.
56. The system according to any one of claims 1 to 55, wherein the sprint includes modification.
57. The system according to claim 56, wherein the modification includes streptavidin operably coupled to the splint.
58. The system according to claim 57, wherein the sprint containing streptavidin is operably coupled to the DNA ligase, and the DNA ligase is operably coupled to biotin.
59. The system according to claim 56, wherein the modification includes biotin operably coupled to the splint.
60. The system according to claim 59, wherein the sprint containing biotin is operably coupled to the DNA ligase, and the DNA ligase is operably coupled to streptavidin.
61. The system according to any one of claims 1 to 60, wherein the endonuclease includes a fusion partner.
62. The system according to claim 61, wherein the fusion partner comprises streptavidin; Rad51 DNA repair protein (rad51DBD) or a fragment thereof; high mobility group nucleosome-binding domain 1 (HN1) or a fragment thereof; histone H1 central globular domain (H1G) or a fragment thereof; Brex27 or a fragment thereof; or a combination thereof.
63. A method for incorporating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising introducing a nucleic acid system into the host cell, wherein the method is a) i. Spacers complementary to the genomic locus regions of the genome strand, ii. Scaffold for complexing with endonuclease, iii. Any donor binding site that is at least partially complementary to the embedded nucleic acid, iv. Flap binding sites located at or adjacent to the genomic locus that are at least partially identical or complementary to the genomic flap. Guide nucleic acids including, b) A first integrated nucleic acid comprising at least one nucleic acid sequence encoding at least one integrated recombinant sequence, further comprising a 5' end ligated to the 3' end of the genome strand generated by the endonuclease, Introducing the host cell into a nucleic acid system containing the host cell A method for incorporating exogenous nucleic acids, including [specific exogenous nucleic acids], into target nucleic acids in host cells.
64. A method for incorporating an exogenous nucleic acid into a target nucleic acid in a host cell, comprising introducing a nucleic acid system into the host cell, wherein the method is a) Spacers that are complementary to the regions of genomic loci in the genome strand. A scaffold for complexing with endonucleases, and Optional sprint binding sites at least partially complementary to the sprint nucleic acid Guide nucleic acids including, b) A first integrated nucleic acid comprising at least one nucleic acid sequence encoding at least one integrated recombinant sequence, further comprising a 5' end ligated to the 3' end of the genome strand generated by the endonuclease, c) A sprint nucleic acid comprising a flap binding site at or adjacent to the genomic locus that is at least partially identical or complementary to the genomic flap, and an optional guide binding site (GBS) that is at least partially complementary to the guide nucleic acid, Introducing the host cell into a nucleic acid system containing the host cell A method for incorporating exogenous nucleic acids, including [specific exogenous nucleic acids], into target nucleic acids in host cells.
65. The method according to any one of claims 63 or 64, wherein the sprint nucleic acid further comprises a donor binding site (DBS) that is at least partially identical or complementary to a portion of the first integrated nucleic acid.
66. The method according to claim 64, wherein the GBS comprises a modified nucleotide.
67. The method according to claim 66, wherein the GBS includes a region in which alternating nucleotides or every three nucleotides contain LNAs.
68. The method according to claim 65, wherein the DBS includes a modified nucleotide.
69. The method according to claim 68, wherein the DBS includes a region in which alternating nucleotides or every three nucleotides contain LNA.
70. The method according to any one of claims 63 or 64, wherein the sprint nucleic acid includes a binding site having a modified nucleotide in the composition of the sprint nucleic acid shown in Table 12.
71. The method according to any one of claims 63 or 64, wherein the guide nucleic acid includes a sequence of linked nucleic acids between the scaffold and the donor binding site.
72. The method according to any one of claims 63 or 64, wherein the guide nucleic acid includes an MS2 binding loop within the scaffold.
73. The method according to claim 64, wherein the guide nucleic acid includes an MS2 binding loop between the scaffold and the donor binding site.
74. The method according to any one of claims 63 or 64, wherein the guide nucleic acid, the first embedded nucleic acid, or the sprint nucleic acid includes a modified internucleoside bond.
75. The method according to claim 74, wherein the modified nucleoside bond includes a phosphorothioate bond.
76. The method according to claim 75, wherein the modified nucleoside bond includes a phosphonoacetic acid bond.
77. The method according to claim 63 or 64, wherein the guide nucleic acid, the first integrated nucleic acid, or the sprint nucleic acid comprises a modified nucleoside.
78. The method according to claim 77, wherein the modified nucleoside includes locked nucleic acid (LNA), 2'-fluoro, 2'O-alkyl, methylated cytosine, reverse thymidine, or a combination thereof.
79. The method according to claim 63 or 64, wherein the endonuclease comprises an RNA guide endonuclease.
80. The method according to claim 79, wherein the endonuclease comprises a class II CRISPR / Cas endonuclease.
81. The method according to claim 80, wherein the endonuclease comprises Cas9 endonuclease.
82. The method according to claim 81, wherein the endonuclease comprises nicasse.
83. The method according to claim 82, wherein the endonuclease comprises Cas9 nicasse.
84. The method further comprises a second embedded nucleic acid, wherein the second embedded nucleic acid is At least one nucleic acid sequence encoding a homologous recombinant sequence Includes, The at least one homologous recombinant sequence of the second embedded nucleic acid is coupled with the at least one embedded recombinant sequence of the first embedded nucleic acid, The integrase incorporates all or part of the second integrated nucleic acid into the target nucleic acid in the integrated recombinant sequence. The method according to any one of claims 63 to 83.
85. The method according to claim 84, wherein the second integrated nucleic acid comprises at least one binding site (att) homologous recombinant sequence.
86. The method according to claim 85, wherein the second incorporated nucleic acid includes a binding site on the bacterial portion (attB) homologous recombinant sequence.
87. The method according to claim 85, wherein the second incorporated nucleic acid includes two binding sites on the bacterial portion (attB) homologous recombinant sequence.
88. The method according to claim 85, wherein the second integrated nucleic acid includes a binding site on the phage moiety (attP) homologous recombinant sequence.
89. The method according to claim 85, wherein the second integrated nucleic acid includes two binding sites on the phage moiety (attP) homologous recombinant sequence.
90. The method according to claim 85, wherein the second incorporated nucleic acid includes one attachment site on the bacterial portion (attB) incorporating the recombinant sequence and one attachment site on the phage portion (attP) incorporating the recombinant sequence.
91. The method according to claim 84, wherein the second integrated nucleic acid comprises at least one locus of X (crossover) in a P1 (LoxP) congeneral recombinant sequence.
92. The method according to claim 84, wherein the first embedded nucleic acid comprises at least one flippase recognition target (FRT) that incorporates a recombinant sequence.
93. The method according to any one of claims 84 to 92, wherein the second incorporated nucleic acid includes a regulatory sequence.
94. The method according to claim 93, wherein the regulatory array is a promoter.
95. The method according to any one of claims 84 to 94, wherein the integrase is serine integrase.
96. The method according to claim 95, wherein the serine integrase is PhiC31 bacteriophage integrase, Bxb1 mycobacteriophage integrase, Pseudomonas aeruginosa integrase (Pa01), Nocardia otitidiscaviarum integrase (No67), or Streptomyces ipomoeae integrase (Si74).
97. The method according to any one of claims 95 or 96, wherein the integrase is coupled to a recombination direction factor (RDF).
98. The method according to any one of claims 84 to 94, wherein the integrase is tyrosine integrase.
99. The method according to claim 98, wherein the tyrosine integrase is Cre recombinase.
100. The method according to claim 98, wherein the tyrosine integrase is flippase (Flp).
101. The method according to any one of claims 63 to 100, wherein the first embedded nucleic acid or the second embedded nucleic acid includes a modified nucleotide.
102. The method according to claim 101, wherein the modified nucleotide includes a methylated nucleotide.
103. The method according to claim 102, wherein the methylated nucleotide includes methylated cytosine (e.g., 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
104. The method according to claim 103, wherein the methylated nucleotide comprises methylated cytosine 5-mC.
105. The method according to claim 101, wherein the modified nucleotide includes a 5' reverse dideoxy-T, a 3' phosphorylation, a 3' C3 spacer, a 3' reverse dT, or a combination thereof.
106. The method according to any one of claims 63 to 105, wherein the sprint nucleic acid includes modification.
107. The method according to claim 106, wherein the sprint nucleic acid includes the sequence shown in Table 12.
108. The method according to claim 106, wherein the modification comprises streptavidin operably coupled to the sprint nucleic acid.
109. The method according to claim 108, wherein the sprint nucleic acid containing streptavidin is operably coupled with the DNA ligase, and the DNA ligase is operably coupled with biotin.
110. The method according to claim 106, wherein the modification comprises biotin operably coupled to the sprint nucleic acid.
111. The method according to claim 110, wherein the sprint nucleic acid containing biotin is operably coupled with the DNA ligase, and the DNA ligase is operably coupled with streptavidin.
112. The method according to any one of claims 63 to 111, wherein the endonuclease includes a fusion partner.
113. The method according to claim 112, wherein the fusion partner comprises streptavidin; Rad51 DNA repair protein (rad51DBD) or a fragment thereof; high mobility group nucleosome-binding domain 1 (HN1) or a fragment thereof; histone H1 central globular domain (H1G) or a fragment thereof; Brex27 or a fragment thereof; or a combination thereof.
114. An editing method comprising ligating an embedded recombinant sequence to a nick in a target nucleic acid, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease.
115. (i) Integrases; and (ii) Ligase or endonuclease A fusion protein containing [the specified ingredient].
116. The fusion protein according to claim 115, wherein the ligase or endonuclease comprises the ligase.
117. The fusion protein according to claim 115, wherein the ligase or endonuclease comprises the endonuclease.
118. The fusion protein according to claim 115, wherein the ligase or endonuclease comprises the ligase and the endonuclease.
119. A method comprising ligating an integration sequence to a nick of a target nucleic acid in a cell, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease, and the integration sequence comprises a modified nucleotide.
120. The method according to claim 119, wherein the incorporated nucleic acid includes a modified nucleotide.
121. The method according to claim 120, wherein the modified nucleotide includes a methylated nucleotide.
122. The method according to claim 120, wherein the modified nucleotide includes methylated cytosine (e.g., 5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), 5-carboxylcytosine (5-caC), N6-methyladenine (6-mA), or a combination thereof.
123. The method according to claim 121, wherein the methylated nucleotide comprises methylated cytosine 5-mC.
124. The method according to claim 119, wherein the modified nucleotide includes a 5' reverse dideoxy-T, a 3' phosphorylation, a 3' C3 spacer, a 3' reverse dT, or a combination thereof.
125. An editing method comprising ligating an integration sequence to a nick in a target nucleic acid, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease, and the integration of the integration sequence introduces a methylated nucleoside into the genome.
126. An editing method comprising ligating an integration sequence to a nick in a target nucleic acid, wherein the nick is generated by contacting the target nucleic acid with an RNA guide endonuclease, and the integration of the integration sequence removes a methylated nucleoside from the genome.
127. The process involves contacting a target nucleic acid within a cell with an endonuclease at a predetermined gene locus of the target nucleic acid, thereby introducing a nick to the predetermined gene locus of the target nucleic acid. Introducing pre-synthesized methylated embedded nucleic acids, Ligating the end of the incorporated nucleic acid to the end of the nick at the predetermined gene locus of the target nucleic acid, Editing methods, including those mentioned above.
128. Ligauze; Endonucleases that introduce nicks to specific gene loci of target nucleic acids; and A pre-synthesized methylated integrated nucleic acid, which includes a terminal that is ligated to the end of the nick by the ligase at the predetermined gene locus of the target nucleic acid. An editing system that includes this.
129. A method comprising contacting an endonuclease containing a first heterodimerized moiety with a sprint nucleic acid containing a second heterodimerized moiety.
130. The method according to claim 129, further comprising introducing a nick or strand break at a predetermined gene locus of a target nucleic acid by the endonuclease, and ligating an integrated nucleic acid to the nick or strand break, wherein the sprint nucleic acid binds to the integrated nucleic acid and the target nucleic acid.
131. The method according to claim 129 or 130, wherein the first heterodimerized portion contains biotin and the second heterodimerized portion contains streptavidin or avidin, or the second heterodimerized portion contains biotin and the first heterodimerized portion contains streptavidin or avidin.
132. Endonucleases bound to the first heterodimerized portion; and Sprint nucleic acid containing the second heterodimerized moiety A system that includes this.
133. The system according to claim 132, further comprising the ligase.
134. The method according to claim 133, wherein the endonuclease introduces a nick or strand break at a predetermined locus of the target nucleic acid, the ligase ligates the integrated nucleic acid to the nick or strand break, and the sprint nucleic acid binds to the integrated nucleic acid and the target nucleic acid.
135. The system according to any one of claims 132 to 134, wherein the first heterodimerized portion contains biotin and the second heterodimerized portion contains streptavidin or avidin, or the second heterodimerized portion contains biotin and the first heterodimerized portion contains streptavidin or avidin.