DNA ligase compositions, methods, and uses thereof

CN122514599APending Publication Date: 2026-08-04REVISON BIOTECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
REVISON BIOTECH
Filing Date
2024-09-20
Publication Date
2026-08-04

Smart Images

  • Figure CN122514599A_ABST
    Figure CN122514599A_ABST
Patent Text Reader

Abstract

DNA ligase compositions, guide polynucleotides, donor nucleic acids, systems, methods, and uses thereof are provided. Methods of modifying target nucleic acids and genetically modifying cells are described. These methods can be used for biotechnological applications, therapeutic treatments, and for generating cell therapies for treating various diseases and conditions. Also included are scaffolds and kits comprising the compositions and systems described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing This application claims the benefit of U.S. Provisional Application No. 63 / 584,712, filed September 22, 2023, which is incorporated herein by reference in its entirety.

[0002] sequence list This application includes a sequence list that has been electronically submitted in ASCII format and is hereby incorporated in its entirety by reference. The ASCII copy created on September 19, 2024, is named 219001_702601_SL.txt and has a size of 1,265,664 bytes.

[0003] background While many CRISPR systems have become useful tools for gene editing, they are limited by off-target effects, insert length, and inefficient nucleic acid editing with low fidelity. Therefore, there is a significant unmet need for gene modification systems that can address these issues.

[0004] Overview This document provides compositions comprising: (a) a donor nucleic acid; (b) a polynucleotide encoding a protein construct comprising a nicking enzyme or a variant thereof and a DNA ligase or a functional fragment thereof; and (c) a guiding polynucleotide comprising: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region comprising a secondary structure binding to an engineered protein construct comprising the nicking enzyme region; (iii) a ligation splint 2 region comprising a ligation splint region complementary to both the donor and target nucleic acids, and wherein the ligation splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region comprising: a deoxyribonucleotide and a ribonucleotide, wherein the ligation splint 1 region is complementary to the target nucleic acid.

[0005] This article provides engineered fusion proteins comprising: (a) a DNA ligase or a functional fragment thereof; and (b) an engineered nickase comprising three amino acid substitutions at positions 221, 394, and 840 of a nuclease containing the sequence of SEQ ID NO: 69.

[0006] This article provides a system for modifying target nucleic acids, the system comprising: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region contains a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region contains at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.

[0007] This article provides a system for modifying target nucleic acids, the system comprising: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region comprises: one or more ribonucleotides or one or more deoxyribonucleotides, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.

[0008] This article provides a system for modifying target nucleic acids, the system comprising: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region comprises: one or more deoxyribonucleotides, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.

[0009] This article provides a system for modifying target nucleic acids, wherein the system comprises: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region comprises: a deoxyribonucleotide and a ribonucleotide, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.

[0010] This document provides compositions comprising: the system provided herein, the engineered protein construct provided herein, the DNA ligase provided herein or a functional fragment thereof, or the guide polynucleotide provided herein.

[0011] This document provides compositions comprising: the system provided herein; and a delivery medium.

[0012] This article provides polynucleotides, wherein the polynucleotides encode the systems, engineered proteins, or guide polynucleotides provided herein.

[0013] This article provides a polynucleotide genome, wherein the polynucleotide genome encodes the systems, engineered proteins, or guide polynucleotides provided herein.

[0014] This article provides nanoparticles, wherein the nanoparticles comprise: the polynucleotides provided herein, the polynucleotide groups provided herein, the systems provided herein, the compositions provided herein, the cells provided herein, the carriers provided herein, or any part thereof.

[0015] This article provides vectors containing the polynucleotides or polynucleotide sequences provided herein.

[0016] This article provides cells, wherein the cells comprise: the systems provided herein, the engineered proteins provided herein, the compositions provided herein, the vectors provided herein, or the guide polynucleotides provided herein.

[0017] This document provides compositions comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises, in a 5' to 3' order: (a) a targeting region complementary to the target nucleic acid; (b) a protein-binding region comprising a protein-binding secondary structure; (c) a linker 2 region, wherein the linker 2 region is complementary to both the donor nucleic acid and the target nucleic acid and has at least one alteration relative to the target nucleic acid or at least one nucleic acid strand thereof; and (d) a linker 1 region, wherein the linker 1 region comprises one or more deoxyribonucleotides or one or more ribonucleotides.

[0018] This document provides compositions comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises, in a 5' to 3' order: (a) a targeting region complementary to the target nucleic acid; (b) a protein-binding region comprising a protein-binding secondary structure; (c) a linker 2 region, wherein the linker 2 region is complementary to both the donor and target nucleic acids and has at least one alteration relative to the target nucleic acid or at least one of its nucleic acid strands; and (d) a linker 1 region, wherein the linker 1 region comprises deoxyribonucleotides and ribonucleotides.

[0019] This article provides a method for linking donor nucleic acids to target nucleic acids, wherein the method comprises: contacting a cell or cell-free system with: (a) a donor nucleic acid, (b) a guiding polynucleotide or a polynucleotide encoding a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a target region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure for binding to a nuclease or nicking enzyme; and (iii) a linker 2 region, wherein the linker 2 region is complementary to the donor nucleic acid and complementary to the target nucleic acid and has at least one modification relative to the target nucleic acid; (iv) a linker 1 region, wherein the linker 1 region comprises one or more deoxyribonucleotides or one or more ribonucleotides, and (c) an engineered protein or a polynucleotide encoding an engineered protein, wherein the engineered protein comprises: (i) a nicking enzyme region; and (ii) The DNA ligase region includes: a region that guides the polynucleotide to form a complex with the engineered protein via a protein-binding region; a region that guides the target sequence to form a complex with the complementary strand; a region that cleaves the engineered protein in the target nucleic acid to produce a leading strand and a complementary strand; a region that connects the splint 1 to the leading strand to form a complex; and a region that attaches the donor nucleic acid to the leading strand, thereby connecting the donor nucleic acid to the target nucleic acid.

[0020] This article provides methods comprising administering the system, composition, or cell provided herein to a subject, organ, tissue, or cell, wherein a gene modifying the subject, organ, tissue, or cell is administered.

[0021] This article provides methods comprising: contacting cells or cell populations with a system or composition provided herein, thereby modifying genes in the cells.

[0022] This article provides cell populations prepared using the methods described herein.

[0023] This document provides a kit comprising: a first container and a second container, the first container comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure for binding to a nuclease or nicking enzyme; (iii) a ligation splint 2 region, wherein the ligation splint 2 region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint 2 region comprises at least one mismatched nucleobase relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region comprises a deoxyribonucleotide and a ribonucleotide, and the second container comprises: an engineered polypeptide comprising a nicking enzyme operatively linked to a DNA ligase.

[0024] This document provides a support, wherein the support comprises: the system or composition provided herein; and a solid surface, wherein the system or composition is fixed to the solid surface. Brief description of the attached diagram The novel features of the invention are specifically set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description and accompanying drawings, which illustrate illustrative embodiments utilizing the principles of the invention, in which: Figures 1A-1B A schematic diagram of the implementation scheme of the system provided in this paper is shown. Figure 1A A schematic diagram of a DNA editing process performed by an engineered protein comprising a Cas9-nicking enzyme-DNA ligase and a guide polynucleotide designed to form a complex with a double-stranded target nucleic acid and edit the target nucleic acid by linking the donor nucleic acid to the target nucleic acid sequence. Figure 1B This diagram illustrates a linker-editing guide polynucleotide (legRNA) used for DNA editing via DNA ligase. The legRNA comprises a linker clip 2 region (LS2) and a linker clip 1 region (LS1). Legend: D is a deoxyribonucleotide (DNA), and R is a ribonucleotide (RNA).

[0026] Figure 2 This is an image of a denatured urea-polyacrylamide gel showing targeted DNA editing using trans-splastic oligonucleotides via DNA ligase. Lanes 1 and 12 show oligonucleotide ladder bands of uncut (oligonucleotide 1, 54 nucleotides (nt)) and pre-cut (oligonucleotide 32, 34 nt) DNA. The bands in lane 2 represent the Cas9 nickase (H840A) control using a single guide RNA transcribed in vitro. The bands in lane 3 represent the positive control for 15 nt DNA insertion (49 nt) (oligonucleotide 32, oligonucleotide 11, oligonucleotide 22) and the oligonucleotide ladder bands for incorporation of uncut (oligonucleotide 1, 54 nt) and pre-cut (oligonucleotide 32, 34 nt) DNA. The bands in lane 4 represent positive controls with 90 nt of inserted DNA written to (49 nt) (oligonucleotides 32, 11, and 22) and oligonucleotide ladder bands for incorporation of uncut (oligonucleotide 1, 54 nt) and pre-cut (oligonucleotide 32, 34 nt) DNA. The bands in lane 5 represent negative control DNA ligase editing without Cas9 nickase (H840A). The bands in lanes 6-11 represent products of DNA ligation editing using a trans DNA ligation clip with oligonucleotide 22 and ligation donor oligonucleotides with oligonucleotide 11 (lane 6), oligonucleotide 12 (lane 7), oligonucleotide 13 (lane 8), oligonucleotide 14 (lane 9), oligonucleotide 15 (lane 10), or oligonucleotide 16 (lane 11).

[0027] Figure 3 These are images of denatured urea-polyacrylamide gels, showing targeted DNA editing using synthetic single-guide polynucleotides via DNA ligase. The bands in lanes 1 and 12 represent oligonucleotide ladder bands showing uncut (oligonucleotide 1, 54 nt) and pre-cut (oligonucleotide 2, 34 nt) DNA. The band in lane 2 represents the control using the Cas9 cleavage enzyme (H840A) with guide 3. The bands in lanes 3–11 represent the products of DNA ligation editing using guide 1 (lane 3), guide 2 (lane 4), guide 3 (lane 5), guide 4 (lane 6), guide 5 (lane 7), guide 6 (lane 8), guide 7 (lane 9), guide 8 (lane 10), or guide 9 (lane 11).

[0028] Figure 4This is a bar graph showing the genome editing efficiency of DNA ligases in human cells with three engineered protein constructs. The Y-axis represents the desired editing efficiency (%). The X-axis represents the engineered protein constructs (constructs 10-12; SEQ ID NO: 98-100).

[0029] Figure 5 This is a bar graph showing the efficiency of DNA ligase genome editing at FANCF site 1 (gRVB_3) in human cells using various combinations of Chlorella DNA ligase editor constructs (constructor 10) and clip and DNA donor modifications. The Y-axis represents the expected editing efficiency (%). The X-axis represents the combination of splint and DNA donor modification, including oRVB_61 / oRVB_95 (SEQ ID NO: 184 / SEQ ID NO: 218), oRVB_61 / oRVB_96 (SEQ ID NO: 184 / SEQ ID NO: 219), oRVB_61 / oRVB_97 (SEQ ID NO: 184 / SEQ ID NO: 220), oRVB_62 / oRVB_95 (SEQ ID NO: 185 / SEQ ID NO: 218), oRVB_62 / oRVB_96 (SEQ ID NO: 185 / SEQ ID NO: 219), and oRVB_62 / oRVB_97 (SEQ ID NO: 185 / SEQ ID NO: 220).

[0030] Figure 6 This is a bar graph showing the DNA ligase genome editing efficiency at HEK site 3 in human cells using the Chlorella DNA ligase editor construct (construct 10) and various splice constructs with and without RNA. The Y-axis represents the desired editing efficiency (%). The X-axis represents various splice constructs with and without RNA, including oRVB_81 / oRVB_82 (SEQ ID NO: 204 / SEQ ID NO: 205), oRVB_69 / oRVB_70 (SEQ ID NO: 192 / SEQ ID NO: 193), oRVB_77 / oRVB_78 (SEQ ID NO: 200 / SEQ ID NO: 201), oRVB_65 / oRVB_66 (SEQ ID NO: 188 / SEQ ID NO: 189), and oRVB_73 / oRVB_74 (SEQ ID NO: 196 / SEQ ID NO: 197).

[0031] Figure 7This is a bar graph showing the DNA ligase genome editing efficiency at AAV site 1 in human cells using the Chlorella DNA ligase editor construct (construct 10) and various clip constructs with and without RNA. The Y-axis represents the desired editing efficiency (%). The X-axis represents clip constructs with and without RNA, including oRVB_79 / oRVB_80 (SEQ ID NO: 202 / SEQ ID NO: 203), oRVB_75 / oRVB_76 (SEQ ID NO: 198 / SEQ ID NO: 199), oRVB_63 / oRVB_64 (SEQ ID NO: 186 / SEQ ID NO: 187), oRVB_67 / oRVB_68 (SEQ ID NO: 190 / SEQ ID NO: 191), and oRVB_71 / oRVB_72 (SEQ ID NO: 194 / SEQ ID NO: 195).

[0032] Figure 8 This is a bar graph showing the DNA ligase genome editing efficiency at AAV site 1 in human cells using the Chlorella DNA ligase editor construct (construct 10) and guide RNA (legRNA) with splint regions incorporating varying degrees of RNA. The Y-axis represents the desired editing efficiency (%). The X-axis represents the guide RNA constructs, including gRVB_9 / gRVB_13 (SEQ ID NO: 110 / SEQ ID NO: 114), gRVB_16 / gRVB_20 (SEQ ID NO: 117 / SEQ ID NO: 121), and gRVB_17 / gRVB_21 (SEQ ID NO: 118 / SEQ ID NO: 122).

[0033] Several aspects will now be described in more detail below. However, such aspects can be implemented in many different forms and should not be construed as limited to the implementation methods described herein.

[0034] Detailed description of the invention This document provides compositions, kits, methods, and uses for editing donor nucleic acid sequences and ligating them to target genes using DNA ligases. In brief, this document further describes: (1) gene editing systems; (2) delivery media and vectors; (3) cellular and cell-free systems; (4) pharmaceutical compositions, administration, and delivery; (5) scaffolds and systems; (6) kits; (7) gene editing activity; and (8) applications.

[0035] This document provides compositions and systems for gene editing, which can be used in a variety of applications such as gene editing, diagnostics, and biopharmaceutical manufacturing. The compositions and systems provided herein include engineered proteins containing nicking enzyme activity and DNA ligase activity that allow the engineered protein to bind to dsDNA and create nicks on a non-target DNA strand to modify target nucleic acids.

[0036] The compositions and systems provided herein also include a guide polynucleotide and a donor nucleic acid for incorporation into a target nucleic acid. The guide polynucleotide binds to the engineered protein and the target nucleic acid to promote loop formation between the cleaved target nucleic acid and the guide polynucleotide, thereby facilitating high-fidelity gene editing. The guide polynucleotide mediates hybridization of the donor nucleic acid, bringing the 3' end of the cleavage site of the target nucleic acid and the 5' end of the donor nucleic acid close to each other. The target nucleic acid and the guide polynucleotide form a complex to provide a substrate for DNA ligase to directly attach the donor nucleic acid sequence of the guide nucleic acid and the 3' end of the cleavage site. The compositions and systems provided herein allow for controlled editing results with limited off-target effects. Furthermore, the compositions and systems provided herein allow for the insertion of large nucleic acid sequences that can be precisely integrated into the target nucleic acid.

[0037] definition All definitions used herein should be understood to take precedence over dictionary definitions, definitions in incorporated documents, and / or the general meaning of the terms used in the definitions.

[0038] All references, patents, and patent applications disclosed herein are incorporated by reference to their respective subjects, and in some cases, these references, patents, and patent applications may cover the entire document. All references disclosed herein, including patent references and non-patent references, are incorporated herein by reference in their entirety as if each reference were incorporated individually. However, when patents, patent applications, or publications containing explicit definitions are incorporated by reference, these explicit definitions should be understood to apply to the incorporated patent, patent application, or publication in which they exist, and not necessarily to the text of this application, particularly the claims of this application; in such cases, the definitions provided herein are intended to supersede them.

[0039] Unless explicitly indicated otherwise, the indefinite articles “a” and “an” used herein in the specification and claims shall be understood to mean “at least one”.

[0040] The phrase “and / or” as used herein in the specification and claims should be understood to mean “one or two” of the elements so connected (i.e., elements that are jointly present in some cases and separately present in others). The use of “and / or” to list multiple elements should be interpreted in the same way, i.e., “one or more” of the elements so connected. In addition to the elements specifically indicated by the “and / or” clause, other elements may optionally be present, whether related to or unrelated to those specifically indicated. Thus, as a non-limiting example, references to “A and / or B” used in conjunction with open-ended language (such as “comprising / including”) may, in one embodiment, refer to only A (optionally including elements other than B); in another embodiment, refer to only B (optionally including elements other than A); in yet another embodiment, refer to both A and B (optionally including other elements); and so on.

[0041] As used herein in the specification and claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” should be interpreted as inclusive, i.e., including multiple elements or at least one of the elements in the list, but also including more than one, and optionally including additional unlisted items. Terms that explicitly indicate the opposite, such as “only one” or “exactly one”, or “consisting of” as used in the claims, will refer to the inclusion of multiple elements or exactly one element in the list of elements. In general, when preceded by exclusive terms such as “any,” “one of,” “only one of,” or “exact one of,” the term “or” as used herein should only be interpreted as indicating an exclusive alternative (i.e., “one or the other, but not both”). “Substantially consisting of” when used in the claims should have the ordinary meaning as it is used in the field of patent law.

[0042] As used herein, “optional” or “optionally” means that the situation described below may or may not occur, such that the description includes instances where the situation occurs and instances where the situation does not occur.

[0043] As used herein, the terms “about” or “approximately” mean a range of up to ±20% of a given value. Optionally, particularly, with respect to biological systems or methods, the term may mean within orders of magnitude of the value, preferably within 2 times. Where a particular value is described in the application and claims, the term “about” is implied unless otherwise stated and, in the context herein, means within an acceptable margin of error for the particular value.

[0044] The term "effective amount" or "therapeutic effective amount" refers to an amount sufficient to achieve or at least partially achieve the desired effect.

[0045] As used herein, the term “leading strand” and its grammatical equivalents refer to a single-stranded nucleic acid sequence that binds to either the linker splint 1 region (also referred to herein as the splint) or the linker splint 2 region of the guiding polynucleotide provided herein.

[0046] As used herein, the term “complementary strand” and its grammatical equivalents refer to a single-stranded nucleic acid sequence that contains a protospacer adjacent motif (PAM), is complementary to the leader strand, and / or binds to the target region of the guiding polynucleotide provided herein.

[0047] The term "effective amount" or "therapeutic effective amount" refers to an amount sufficient to achieve or at least partially achieve the desired effect.

[0048] (1) Gene editing system This document provides a system comprising: a donor nucleic acid, a guide polynucleotide, and an engineered protein. In some embodiments, the system is used in editing nucleic acids. In some embodiments, the system cleaves the target nucleic acid. In some embodiments, the system generates single-strand breaks in the target nucleic acid. In some embodiments, the system incorporates the donor nucleic acid into the target nucleic acid, thereby replacing an abnormal nucleic acid sequence relative to a reference sequence.

[0049] The system provided herein comprises the following components: (1) a donor nucleic acid; (2) a guide polynucleotide; (3) a nuclease, a nicking enzyme, or a variant thereof; and (4) a DNA ligase or a functional fragment thereof. An exemplary system is shown in Figure 1AThe system comprises engineered proteins including a Cas9 nickase, a DNA ligase, a donor nucleic acid, and a guide polynucleotide containing a DNA ligation splint complex to form a ribonucleoprotein (RNP). The DNA ligation splint may contain a ligation splint 2 region and / or a ligation splint 1 region as described herein. The RNP can target double-stranded DNA (dsDNA) in the presence of a protospacer adjacent motif (PAM) and high sequence homology between the guide polynucleotide and dsDNA. Following conjugation between the Cas9-ligase:guide polynucleotide complex and the dsDNA, the Cas9 nickase generates a single-strand break in the non-target strand (NTS), thereby allowing the release of single-stranded DNA (step 1). The guide polynucleotide DNA ligation splint can then hybridize with the single-stranded DNA along with the ligation donor nucleic acid, presenting a substrate for the DNA ligase to directly attach the 5' end of the donor nucleic acid to the 3' end of the NTS (step 2). Following the DNA ligation editing step, Cas9-ligase directs the dissociation of the polynucleotide complex with the target dsDNA, resulting in different conformations in which the DNA ligation-edited strand is either unhybridized or hybridized with the dsDNA (step 3). Subsequently, endogenous cellular enzymes for DNA repair incorporate either the original DNA strand or the newly attached DNA ligation-edited strand into the dsDNA, thereby producing unedited dsDNA (step 4) or edited dsDNA (step 5), respectively.

[0050] In some arrangements, the guiding polynucleotide used for DNA editing via DNA ligase comprises two components – (1) a splice 1 region (LS1) and (2) a splice 2 region (LS2). LS1 mediates hybridization with the non-target strand (NTS), while LS2 mediates hybridization with the ligation donor oligonucleotide. Figure 1B After LS1 and LS2 hybridize with the NTS and donor oligonucleotide, respectively, this presents a substrate for DNA ligase to directly attach the donor nucleic acid to the non-target strand (also referred to herein as the leading strand) via a DNA ligation event. LS1 contains RNA (typically at the 3' end) and DNA (typically at the 5' end) in varying proportions to enhance DNA ligation efficiency. LS2 typically contains the majority of DNA bases to facilitate efficient DNA ligation by DNA ligase. In some embodiments, the LS2 region contains only RNA. In some embodiments, the LS1 region contains only DNA.

[0051] The system presented in this paper is designed for high-fidelity and high-persistence insertion of nucleic acids. This system can be used to allow for the precise, targeted insertion of polynucleotides into target nucleic acids. In some embodiments, the system provided herein allows for the targeted insertion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, at least 80,000, at least 90,000, at least 100,000 or more polynucleotides. In some embodiments, the system provided herein allows for the targeted insertion of single-stranded polynucleotides of 10 kb or greater.

[0052] donor nucleic acid This document provides compositions and systems comprising donor nucleic acids. The donor nucleic acid provides a high-fidelity DNA insertion to introduce modifications into the target nucleic acid and can further correct for aberrations in the sequence relative to the wild-type sequence. In some embodiments, the wild-type sequence is a reference sequence, a nucleic acid sequence from a healthy subject, or a nucleic acid sequence from a healthy cell. In some embodiments, the donor nucleic acid comprises the reverse complementary sequence of the target nucleic acid. In some embodiments, the donor nucleic acid comprises one or more alterations to nucleosides, nucleosides, or nucleotides. In some embodiments, the donor nucleic acid comprises the reverse complementary sequence of a sequence encoding an intron or a variant thereof. In some embodiments, the donor nucleic acid comprises the reverse complementary sequence of a sequence encoding an exon or a variant thereof. In some embodiments, the donor nucleic acid comprises the reverse complementary sequences of sequences encoding both exons and introns. In some embodiments, the donor nucleic acid comprises the complementary sequence of a sequence encoding an intron or a variant thereof. In some embodiments, the donor nucleic acid comprises the complementary sequence of a sequence encoding an exon or a variant thereof. In some embodiments, the donor nucleic acid comprises the complementary sequence of sequences encoding both exons and introns. In some embodiments, the donor nucleic acid comprises the reverse complementary sequence of a sequence encoding a non-coding element (e.g., a promoter, enhancer) or a variant thereof. In some embodiments, the donor nucleic acid comprises the complementary sequence of a sequence encoding a non-coding element (e.g., a promoter, enhancer) or a variant thereof. In some embodiments, the donor nucleic acid comprises an alteration. In some embodiments, the donor nucleic acid comprises a sequence containing at least one nucleobase that is complementary to or mismatched with a sequence encoding a splice acceptor site. In some embodiments, the donor nucleic acid sequence comprises an A / C mismatch, A / T mismatch, A / G mismatch, T / C mismatch, T / G mismatch, T / A mismatch, C / G mismatch, C / A mismatch, C / T mismatch, G / C mismatch, G / T mismatch, G / A mismatch, or any combination thereof relative to the target nucleic acid sequence or complementary strand (also referred to as the lagging strand) provided herein. In some embodiments, the donor nucleic acid comprises an epigenetic alteration (e.g., methylation).

[0053] Guided polynucleotides This document provides compositions and systems comprising a guide polynucleotide or a polynucleotide encoding a guide polynucleotide. The guide polynucleotide provided herein binds to a target nucleic acid and a nuclease or nicking enzyme provided herein to form a complex. In some embodiments, the complex promotes the cleavage of the target nucleic acid.

[0054] In some embodiments, the guiding polynucleotide includes RNA nucleosides and DNA nucleosides. In some embodiments, the guiding polynucleotide includes RNA nucleotides and DNA nucleotides. In some embodiments, the guiding polynucleotide includes RNA nucleobases and DNA nucleobases. Examples of nucleobases include, but are not limited to, adenine (A), guanine (G), cytosine (C), thymine (T), and uracil (U). The nucleobases of the nucleotide may be independently selected from purines, pyrimidines, purine analogs, or pyrimidine analogs. In embodiments, nucleobases may include, for example, naturally occurring and synthetic derivatives of the bases. The guiding polynucleotides provided herein may contain non-naturally occurring sequences or engineered sequences. In some embodiments, the polynucleotide encoding the guiding polynucleotides provided herein contains a promoter region or localization sequence. For example, the U6 promoter can be used to drive the expression of the guiding polynucleotide in cells.

[0055] In some embodiments, the guiding polynucleotides or donor nucleic acids provided herein comprise modified nucleosides, modified nucleotides, or modified nucleosides. The modified nucleosides and modified nucleotides described herein that can be incorporated into the target nucleic acid may include modified nucleosides. In some embodiments, the nucleosides or nucleotides provided herein are chemically modified. Nucleosides can be modified or replaced to provide modified nucleosides and modified nucleotides that can be incorporated into the target nucleic acid.

[0056] In some embodiments, the modified nucleobase is a modified cytosine. Exemplary nucleobases and nucleosides having modified cytosine include, but are not limited to, 5-aza-cytidine, 6-aza-cytidine, pseudo-isocytidine, 3-methyl-cytidine (m3C), N4-acetyl-cytidine (ac4C), 5-formyl-cytidine (f5C), N4-methyl-cytidine (m4C), 5-methyl-cytidine (m5C), 5-halo-cytidine (e.g., 5-iodo-cytidine), 5-hydroxymethyl-cytidine (hm5C), 1-methyl-cytidine, etc. 2-Thio-cytidine, 2-thio-cytidine (s2C), 2-thio-5-methyl-cytidine, 4-thio-cytidine, 4-thio-1-methyl-cytidine, 4-thio-1-methyl-1-deazo-cytidine, 1-methyl-1-deazo-cytidine, zebularine, 5-aza-zebularine, 5-methyl-zebularine 5-aza-2-thio-zebularine, 2-thio-zebularine, 2-methoxy-cytidine, 2-methoxy-5-methyl-cytidine, 4-methoxy-pseudo-cytidine, 4-methoxy-1-methyl-pseudo-cytidine, lysicidine (k2C), o-thio-cytidine, 2'-O-methyl-cytidine (Cm), 5,2'-O-dimethyl-cytidine (m5Cm), N4-acetyl-2'-O-methyl 5-Formyl-2'-O-methyl-cytidine (ac4Cm), N4,2'-O-dimethyl-cytidine (m4Cm), 5-Formyl-2'-O-methyl-cytidine (f5Cm), N4,N4,2'-O-trimethyl-cytidine (m42Cm), 1-Thio-cytidine, 2'-F-arabinose-cytidine, 2'-F-cytidine and 2'-OH-arabinose-cytidine, 2'-OMe-exNA-cytidine and 2'-F-exNA-cytidine.

[0057] In some embodiments, the modified nucleobase is a modified adenine. Exemplary nucleobases and nucleosides having modified adenine include, but are not limited to, 2-amino-purine, 2,6-diaminopurine, 2-amino-6-halo-purine (e.g., 2-amino-6-chloro-purine), 6-halo-purine (e.g., 6-chloro-purine), 2-amino-6-methyl-purine, 8-azido-adenosine, 7-deadenine, 7-deadenine-8-aza-adenosine, 7-deadenine-2-amino-purine, 7-deadenine-8-aza-2-diamino-purine, 7-deadenine-2,6-diaminopurine, 7-deadenine-8-aza-2,6-diaminopurine, 1-methyl -Adenosine (m1A), 2-methyl-adenine (m2A), N6-methyl-adenosine (m6A), 2-methylthio-N6-methyl-adenosine (ms2m6A), N6-isopentenyl-adenosine (i6A), 2-methylthio-N6-isopentenyl-adenosine (ms2i6A), N6-(cis-hydroxyisopentenyl)adenosine (io6A), 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine (ms2io6A), N6-glycylcarbamoyl-adenosine (g6A), N6-threonylcarbamoyl-adenosine (t6A), N6-methyl-N6- Threonylcarbamoyl-adenosine (m6t6A), 2-methylthio-N6-threonylcarbamoyl-adenosine (ms2g6A), N6,N6-dimethyl-adenosine (m62A), N6-hydroxyn-valinecarbamoyl-adenosine (hn6A), 2-methylthio-N6-hydroxyn-valinecarbamoyl-adenosine (ms2hn6A), N6-acetyl-adenosine (ac6A), 7-methyl-adenosine, 2-methylthio-adenosine, 2-methoxy-adenosine, o-thio-adenosine, 2'-O-methyl-adenosine (Am), N6,2'-O-dimethyl -Adenosine (m6Am), N6-methyl-2'-deoxyadenosine, N6,N6,2'-O-trimethyl-adenosine (m62Am), 1,2'-O-dimethyl-adenosine (miAm), 2'-O-ribosyladenosine (phosphate) (Ar(p)), 2-amino-N6-methyl-purine, 1-thio-adenosine, 8-azido-adenosine, 2'-F-arabinose-adenosine, 2'-F-adenosine, 2'-OH-arabinose-adenosine, and N6-(19-amino-pentaenoyl)-adenosine, 2'-OMe-exNA-adenosine, and 2'-F-exNA-adenosine.

[0058] In some embodiments, the modified nucleobase is a modified guanine. Exemplary nucleobases and nucleosides having modified guanine include, but are not limited to, inosine (I), 1-methyl-inosine (mil), wyoside (imG), methyl wyoside (mimG), 4-demethyl wyoside (imG-14), isowyoside (imG2), weitingoside (yW), peroxyweitingoside (o2yW), hydroxyweitingoside (OHyW), unmodified hydroxyweitingoside (OHyW*), 7-deazo-guanosine, guanosine (Q), epoxy-guanosine (oQ), galactosyl-guanosine (galQ), mannosyl-guanosine (manQ), and 7-cyano-7-deazo-guanosine (pre Qo), 7-aminomethyl-7-deazo-guanosine (preQi), archaenoside (G+), 7-deazo-8-aza-guanosine, 6-thio-guanosine, 6-thio-7-deazo-guanosine, 6-thio-7-deazo-8-aza-guanosine, 7-methyl-guanosine (m7G), 6-thio-7-methyl-guanosine, 7-methyl-inosine, 6-methoxy-guanosine, 1-methyl-guanosine (m'G), N2-methyl-guanosine (m2G), N2,N2-dimethyl-guanosine (m22G), N2,7-dimethyl-guanosine (m2'7G), N2,N2,7-dimethyl-guanosine (m... 2,2,7G), 8-oxo-guanosine, 7-methyl-8-oxo-guanosine, 1-methyl-6-thio-guanosine, N2-methyl-6-thio-guanosine, N2,N2-dimethyl-6-thio-guanosine, o-thio-guanosine, 2'-O-methyl-guanosine (Gm), N2-methyl-2'-O-methyl-guanosine (m) 2"Gm), N2,N2-dimethyl-2'-O-methyl-guanosine (m22Gm), 1-methyl-2'-O-methyl-guanosine (m'Gm), N2,7-dimethyl-2'-O-methyl-guanosine (m",7Gm), 2'-O-methyl-inosine (Im), 1,2'-O-dimethyl-inosine (m'lm), 06-phenyl-2'-deoxyinosine, 2'-O-ribosylguanosine (phosphate) (Gr(p)), 1-thio-guanosine, O6-methyl-guanosine, O6-methyl-2'-deoxyguanosine, Z-F-arabinose-guanosine, 2'-F-guanosine, 2'-OMe-exNA-guanosine and 2'-F-exNA-guanosine.

[0059] In some embodiments, the modified nucleobase is a modified uracil. Exemplary nucleobases and nucleosides having modified uracil include, but are not limited to, pseudouridine (ψ), pyridin-4-ketoribonucleoside, 5-aza-uridine, 6-aza-uridine, 2-thio-5-aza-uridine, 2-thio-uridine (s2U), 4-thio-uridine (s4U), 4-thio-pseudouridine, 2-thio-pseudouridine, 5-hydroxy-uridine (ho5U), 5-aminoallyl-uridine, and 5-halo-uridine (e.g., 5-iodo-uridine or 5-bromo-uridine). 3-Methyluridine (m3U), 5-methoxyuridine (mo5U), uridine 5-hydroxyacetic acid (cmo5U), methyl uridine 5-hydroxyacetic acid (mcmo5U), 5-carboxymethyluridine (cm5U), 1-carboxymethyl-pseudouridine, 5-carboxyhydroxymethyluridine (chm5U), 5-carboxyhydroxymethyluridine methyl ester (mchm5U), 5-methoxycarbonylmethyluridine (mcm5U), 5-methoxycarbonylmethyl-2-thiouridine (mcm5) 5-Methylaminomethyluridine (nm5s2U), 5-methylaminomethyluridine (mnm5U), 5-methylaminomethyl-2-thiouridine (mnm5s2U), 5-methylaminomethyl-2-selenouridine (mnm5se2U), 5-carbamoylmethyluridine (ncm5U), 5-carboxymethylaminomethyluridine (cmnm5U), 5-carboxymethylaminomethyl-2-thiouridine (cmnm5s2U) 5-Propyno-uridine, 1-Propyno-pseudouridine, 5-Taurate methyl-pseudouridine (xcm5U), 1-Taurate methyl-pseudouridine, 5-Taurate methyl-2-thio-uridine (Tm5s2U), 1-Taurate methyl-4-thio-pseudouridine, 5-Methyl-uridine (m5U, i.e., with nucleobase deoxythymidine), 1-Methyl-pseudouridine (ηι'ψ), 5-Methyl-2-thio-uridine (m5s2U), 1-Methyl-4-thio-pseudouridine (m xi / ), 4-Thio-1-methyl-pseudouridine, 3-Methyl-pseudouridine (m ψ), 2-Thio-1-methyl-pseudouridine, 1-Methyl-1-deazo-pseudouridine, 2-Thio-1-methyl-1-deazo-pseudouridine, dihydrouridine (D), dihydropseudouridine, 5,6-Dihydrouridine, 5-Methyl-Dihydrouridine (m5D), 2-Thio-Dihydrouridine, 2-Thio-Dihydropseuuridine, 2-Methoxy-uridine, 2-Methoxy-4-Thio-uridine, 4-Methoxy-pseuuridine, 4-Methoxy-2-Thio-pseuuridine, N1-Methyl-pseuuridine, 3-(3-amino-3-carboxypropyl)uridine (acp3U), 1-Methyl-3-(3-amino-3-carboxypropyl)pseuuridine (acp3) ψ), 5-(isopentenylaminomethyl)uridine (inm5U), 5-(isopentenylaminomethyl)-2-thiouridine (inm5s2U), o-thiouridine, 2'-O-methyluridine (Um), 5,2'-O-dimethyluridine (m5Um), 2'-O-methyl-pseudouridine (ψηι), 2-thio-2'-O-methyluridine (s2Um), 5-methoxycarbonylmethyl-2'-O-methyluridine (mem5Um), 5-carbamoylmethyl-2'-O-methyluridine (ncm5Um), 5-carboxymethylaminomethyl-2 '-O-methyluridine (cmnm5Um), 3,2'-O-dimethyluridine (m3Um), 5-(isopentenylaminomethyl)-2'-O-methyluridine (inm5Um), 1-thiouridine, deoxythymidine, 2'-F-arabinose-uridine, 2'-F-uridine, 2'-OH-arabinose-uridine, 5-(2-methoxycarbonylvinyl)uridine, 5-[3-(1-E-propenylamino)]uridine, pyrazolo[3,4-d]pyrimidine, xanthine, hypoxanthine, 2'-OMe-exNA-uridine, and 2'-F-exNA-uridine.

[0060] In some embodiments, the modified nucleobase is a modified thymine. In some embodiments, the modified nucleoside is a modified thymidine. Non-limiting examples of modified thymine and thymidine include: 6-(azo)thymine, 3'-azido-3'-deoxythymidine, 2',3'-didehydro-2',3'-dideoxythymidine; 1-(2,3-dideoxy-β-D-glycerolpent-2-enfuranosyl)thymidine, 3-(2-chloroethyl)thymidine, 3'-fluoro-3'-deoxythymidine, β-L-2'-deoxythymidine, thieno[3,4-d]- Pyrimidine T-mimetic deoxynucleoside, 1-(2-deoxy-β-D-threo-pentafuranosyl)thymidine, 5-ethynyl-2'-deoxyuridine, bromodeoxyuridine, tritylthymidine, 5-chlorodeoxyuridine (CldU), 5-iododeoxyuridine (IdU), 2-thiothymidine triphosphate, 5-(α-tert-butyl-o-bromobenzyloxy)methyl-2'-deoxyuridine, 5-(α-methylbenzyloxy)methyluracil, and 5-ethyldeoxyuridine.

[0061] Nucleic acids can be modified using a variety of chemicals and modifications. In some embodiments, the conventional internucleotide linkage between nucleotides can be altered by monothiolation or dithiolation of the phosphodiester bond to produce thiophosphate or dithiophosphate, respectively. Other modifications to the internucleotide linkage can include amidation or peptide linkers. Ribose can be modified by substituting the 2'-O moiety with a lower alkyl group (C1-4, such as 2'-O-Me), alkenyl (C2-4), alkynyl (C2-4), methoxyethyl (2'-MOE), or other substituents. In some cases, the substituent for the 2'OH group can include a methyl, methoxyethyl, or 3,3'-dimethylallyl group. In some cases, locked nucleic acid sequences (LNAs) containing 2'-4' intramolecular bridges (such as a methylene bridge between the 2' oxygen and the 4' carbon) can be applied. Purine and / or pyrimidine nucleobases can be modified to alter their properties, for example, through amination or deamination of heterocycles. Many of these modified nucleotides and their corresponding ribonucleotides are available from commercial suppliers. If desired, the guiding polynucleotide and / or donor nucleic acid may contain aminophosphate, thiophosphate, and / or methylphosphonate bonds. Several suitable methods can be used to generate nucleic acid molecules and nucleic acids containing modified nucleotides. For example, guiding polynucleotides and / or donor nucleic acids containing modified nucleotides can be prepared by transcribing DNA encoding the guiding polynucleotide and / or donor nucleic acid using a suitable DNA-dependent RNA polymerase, such as T7 phage RNA polymerase, SP6 phage RNA polymerase, T3 phage RNA polymerase, or mutants of these polymerases. In some embodiments, guiding polynucleotides and / or donor nucleic acids containing modified nucleotides can be prepared by in vitro transcription. The transcription reaction may contain nucleotides and modified nucleotides, as well as other components supporting the activity of the selected polymerase, such as suitable buffers and suitable salts. Nucleotide analogs can be engineered to incorporate into directing polynucleotides, for example, to alter the stability of such RNA / DNA molecules or increase their resistance to RNases.

[0062] In some embodiments, the guiding polynucleotide comprises a secondary or tertiary structure. In some embodiments, the secondary structure includes protrusions, stems, stem-loops, loops, tetraloops, hairpins, wobbling base pairs, pseudoknots, nexus, or combinations thereof. In some embodiments, the protrusion, stem-loop, or hairpin comprises an unpaired nucleotide region within the nucleic acid double strand.

[0063] In some embodiments, the guiding polynucleotide provided herein comprises a targeting region complementary to the target nucleic acid. In some embodiments, the targeting region comprises RNA. In some embodiments, the targeting region has at least 80% complementarity with the target nucleic acid. In some embodiments, the targeting region has at least 85% complementarity with the target nucleic acid. In some embodiments, the targeting region has at least 90% complementarity with the target nucleic acid. In some embodiments, the targeting region has at least 95% complementarity with the target nucleic acid. In some embodiments, the targeting region has at least 99% complementarity with the target nucleic acid. In some embodiments, the targeting region has 100% complementarity with the target nucleic acid.

[0064] In some embodiments, the target region comprises at least 10 ribonucleotides, at least 15 ribonucleotides, at least 20 ribonucleotides, at least 25 ribonucleotides, at least 30 ribonucleotides, at least 35 ribonucleotides, at least 40 ribonucleotides, at least 45 ribonucleotides, at least 50 ribonucleotides, or more. In some embodiments, the target region comprises at least about 10 to at most 15 ribonucleotides to enhance specificity. In some embodiments, the target region comprises at least about 20 to 30 ribonucleotides to enhance structural stability and specificity.

[0065] In some embodiments, the target region hybridizes with the target nucleic acid. In some embodiments, a nuclease or nicking enzyme cleaves the target nucleic acid to produce a leading strand and a complementary strand (also referred to herein as a lagging strand). In some embodiments, after the engineered protein construct or nicking enzyme provided herein cleaves the target nucleic acid, the target region binds to at least a portion of the cleaved target nucleic acid. In some embodiments, the target region binds to the complementary strand. In some embodiments, the complementary strand comprises a protospacer adjacent motif (PAM). The PAM is a short nucleic acid sequence (typically 2-6 base pairs in length) located before the region targeted for cleavage by the nuclease or nicking enzyme provided herein. In some embodiments, the target region binds to a target gene sequence within at least 20 nucleotides, at least 15 nucleotides, at least 10 nucleotides, or at least 5 nucleotides of the protospacer adjacent motif (PAM). The guiding polynucleotide may bind 5' upstream or 3' downstream of the protospacer adjacent motif (PAM) sequence via the target region.

[0066] In some embodiments, the guiding polynucleotide includes a protein-binding region. In some embodiments, the protein-binding region includes a protein-binding secondary structure. In some embodiments, the protein-binding region binds to a nuclease, nickase, endonuclease, or exonuclease. In some embodiments, the protein-binding region binds to a Cas protein. When the guiding polynucleotide is complexed with the engineered protein provided herein, the secondary structure can minimize the possibility that the guiding polynucleotide interferes with the activity of the nuclease or nickase. In some embodiments, the secondary structure includes: a protrusion, a stem, a loop, a hairpin, a wobbling base pair, a pseudoknot, or a combination thereof.

[0067] In some embodiments, the guiding polynucleotide includes a linker splint 2 (LS2) region. In some embodiments, the linker splint 2 region is complementary to the donor nucleic acid. In some embodiments, the linker splint 2 region is complementary to the target nucleic acid. In some embodiments, the LS2 region is located at the 5' of the LS1 region. In some embodiments, the LS2 region is located within a protein-binding region. In some embodiments, the LS2 region contains a deoxyribonucleotide. In some embodiments, the LS2 region is located within the protein-binding region of the guiding polynucleotide. In some embodiments, the LS2 region includes secondary or tertiary structures that enhance the stability of the guiding polynucleotide. In some embodiments, the linker splint 2 region includes at least one mismatched nucleotide relative to the target nucleic acid. In some embodiments, the LS2 region includes a length of at least about 3 nucleotides to at most 100,000 nucleotides. In some embodiments, the linker splint 2 region includes at least about 5 nucleotides to at most 10,000 nucleotides. In some embodiments, the linker splint 2 region includes at least about 7 nucleotides to at most 10,000 nucleotides.

[0068] In some implementations, the LS2 region contains the reverse complementary sequence of the target nucleic acid or its strand. The LS2 region and / or donor nucleic acid can be engineered to correct for aberrations in the target sequence relative to the wild-type sequence. For example, the wild-type sequence may include a reference sequence, a nucleic acid sequence from a healthy subject, or a nucleic acid sequence from a healthy cell. Methods for obtaining the reference sequence of the polynucleotide may include sequencing or sequence alignment tools and databases such as NCBI BLAST or UniProt.

[0069] In some embodiments, the linker 2 region contains one or more alterations to a nucleobase, nucleoside, or nucleotide. In some embodiments, the linker 2 region contains the inverse complementary sequence of a sequence encoding an intron or a variant thereof. In some embodiments, the linker 2 region contains the inverse complementary sequence of a sequence encoding an exon or a variant thereof. In some embodiments, the linker 2 region contains the inverse complementary sequences of sequences encoding both exons and introns. In some embodiments, the linker 2 region contains the complementary sequence of a sequence encoding an intron or a variant thereof. In some embodiments, the linker 2 region contains the complementary sequence of a sequence encoding an exon or a variant thereof. In some embodiments, the linker 2 region contains the complementary sequence of sequences encoding both exons and introns. In some embodiments, the linker 2 region contains the inverse complementary sequence of a sequence encoding a non-coding element (e.g., a promoter, enhancer) or a variant thereof. In some embodiments, the linker 2 region contains the complementary sequence of a sequence encoding a non-coding element (e.g., a promoter, enhancer) or a variant thereof.

[0070] In some embodiments, the guiding polynucleotides provided herein contain one or more alterations. In some embodiments, one or more alterations include changes to at least one nucleotide, nucleoside, or nucleotide in the LS2 or LS1 region. In some embodiments, the linker splice 2 region contains a sequence containing at least one nucleotide that is complementary to or mismatched with the sequence encoding the splice acceptor site. In some embodiments, the linker splice 2 region contains A / C mismatches, A / T mismatches, A / G mismatches, T / C mismatches, T / G mismatches, T / A mismatches, C / G mismatches, C / A mismatches, C / T mismatches, G / C mismatches, G / T mismatches, G / A mismatches, or any combination thereof, relative to the target nucleic acid sequence or the complementary strand (also known as the lagging strand) provided herein.

[0071] In some embodiments, the guiding polynucleotide provided herein includes a splint 1 region. In some embodiments, the splint 1 region includes one ribonucleotide, more ribonucleotides, or an RNA region. In some embodiments, the splint 1 region includes one deoxyribonucleotide, more deoxyribonucleotides, or a DNA region. In some embodiments, the splint 1 region includes both ribonucleotides and deoxyribonucleotides. In some embodiments, the splint 1 region includes at least one ribonucleotide and at least one deoxyribonucleotide.

[0072] In some embodiments, the LS1 region of the guide polynucleotide forms a DNA-RNA (DR) loop upon association with the target nucleic acid. The DR loop provides a structural framework and stability for the formation of a complex between the target nucleic acid, the guide polynucleotide, and the engineered protein provided herein. In some embodiments, the LS1 forms a DR loop upon association with the target nucleic acid and the engineered protein provided herein.

[0073] In some embodiments, the LS1 region forms a DNA (D) loop upon association with the target nucleic acid. The D loop forms after the target DNA is cleaved by a nicking enzyme. A D loop is a DNA structure in which the two strands of a double-stranded DNA molecule are separated by a third DNA strand. In some embodiments, the third strand of the DNA in the D loop is either the LS2 or LS1 region, which directs the polynucleotide.

[0074] In some embodiments, the LS1 region forms an R-loop upon association with the target nucleic acid. In some embodiments, the LS1 region forms an RNA (R)-loop upon association with the target nucleic acid and the engineered protein provided herein. The R-loop can form between the complementary strand of the target DNA, the RNA portion of the LS1 region that guides the polynucleotide, and the leading strand of the target DNA after the target nucleic acid has been cleaved by a nickase.

[0075] In some embodiments, after the engineered protein construct provided herein cleaves the target nucleic acid, the LS1 region hybridizes with the leading strand of the DNA. In some embodiments, the ribonucleoside, ribonucleotide, or RNA region within the LS1 region hybridizes with the leading strand. In some embodiments, the LS1 region is located at the 3' of the LS2 region. In some embodiments, the LS1 region is located within the protein-binding region.

[0076] In some implementations, the connecting plate 1 region contains the following ratio of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1 up to 20:1. In some implementations, the connecting plate region 1 contains the following ratios of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 2:1, 2:3, 2:5, 2:7, 2:9, 2:11, 2:13, 2:15, 2:17, 2:19, 3:1, 3:2, 3:4, 3:5, 3:7, 3:8, 3:10, 3:11, 3:13, 3:14, 3:15, 3:16. 3:17, 3:19, 4:1, 4:3, 4:5, 4:7, 4:9, 4:11, 4:13, 4:15, 4:17, 4:19, 5:1, 5:2, 5:3, 5:4, 5:6, 5:7, 5:8, 5:9, 5:11, 5:12, 5:13, 5:14, 5:16, 6:1, 6:5, 6:7, 6:9, 6:11, 6:13, 6:15, 7:1, 7:2, 7:3, 7:4, 7:5, 7:6, 7:8, 7:9, 7:10, 7:11, 7:12, 7:13, 7:15, 8:1, 8:3, 8:5, 8:7, 8:9, 8:11, 8:13, 8:15, 9 :1, 9:2, 9:4, 9:5, 9:7, 9:8, 9:10, 9:11, 9:13, 9:15, 9:17, 9:19, 9:20, 10:1, 10:3, 10:7, 10:9, 10:11, 10:13, 10:15, 10:17, 10:19, 11:1, 11:2, 11:3, 11:4, 11:5, 11:6, 11:7, 11:8, 11:9, 11:10, 11:12, 11:13, 11:15, 12:1, 12:5, 12:7, 12:9, 12:11, 12:13, 13:1, 13:2, 13:3, 13:4, 13:5, 1 3:6, 13:7, 13:8, 13:9, 13:10, 13:11, 13:12, 13:14, 14:1, 14:3, 14:5, 14:9, 14:11, 14:13, 15:1, 15:2, 15:4, 15:6, 15:8, 15:11, 15:13, 16:1, 16:3, 16:5, 16:7, 16:9, 16:11, 16:13, 16:15, 17:1, 17:2, 17:3, 17:4, 17:5, 17:6, 17:7, 17:8, 17:9, 17:10, 17:11, 17:12, 17:13, 17:14, 17:1517:16, 18:1, 18:5, 18:7, 18:11, 18:13, 18:17, 19:1, 19:2, 19:3, 19:4, 19:5, 19:6, 19:7, 19:8, 19:9, 19:10, 19:11, 19:12, 19:13, 19:14, 19:15, 19:16, 19:17, 19:18, or 20:1. In some embodiments, region 1 of the connecting clamp contains at least about 3 nucleobases and at most about 20 nucleobases.

[0077] The interactions between donor nucleic acids, the guide polynucleotides provided herein, and the engineered proteins provided herein are shown in Figure 1A In the middle. An exemplary guiding polynucleotide structure is shown in Figure 1B Therefore, in some embodiments, the guiding polynucleotide provided herein comprises, in the 5' to 3' orientation: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the nicking enzyme region; (iii) a linker 2 region, which is complementary to the donor nucleic acid and the target nucleic acid and has at least one mismatched nucleotide base relative to the target nucleic acid; and (iv) a linker 1 region, wherein the linker 1 region comprises: a DNA nucleoside and an RNA nucleoside, wherein the linker 1 region is complementary to the target nucleic acid. In some embodiments, upon introduction into a cell or cell-free system, the compositions or systems provided herein incorporate the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid. In some embodiments, the newly inserted nucleic acid comprises a single-stranded DNA sequence substantially complementary to the complementary strand (which contains a PAM sequence) and also comprises at least one mismatched nucleotide base relative to the complementary strand. In some embodiments, the linker 1 region forms an R-loop upon association with the target nucleic acid and DNA ligase.

[0078] In some embodiments, the guiding polynucleotide comprises a sequence or a portion thereof from Table 3 or Table 7. In some embodiments, the guiding polynucleotide comprises a sequence that is at least 75%, 80%, 85%, 90%, 95%, 99%, or 100% identical to any one of SEQ ID NO: 45 to 57 or SEQ ID NO: 102 to 123.

[0079] The guide polynucleotides provided in this article can be generated, for example, through phosphoramide chemical synthesis, in vitro transcription (IVT) technology, M13 phage methods, DNA isolation, RNA isolation, tag fragmentation, and combinations thereof. Furthermore, guide polynucleotides can be generated and expressed by transducing or transfecting cells with plasmid DNA or vectors containing polynucleotides encoding expression cassettes.

[0080] In some embodiments, the guiding polynucleotide includes a 5' cap. A free 5' hydroxyl group can be capped by acetylation. In some embodiments, the guiding polynucleotide includes a 3' polyadenylated tail. In some embodiments, the guiding polynucleotide includes a polynucleotide linker or spacer region.

[0081] Engineered proteins that interact with the guiding polynucleotides provided herein can be used in the systems and compositions provided herein to insert nucleic acids. Engineered proteins that can be used to edit nucleic acids and target the cleavage of specific target sequences are described in further detail below.

[0082] engineered proteins This document provides engineered proteins that specifically bind to target nucleic acids by hybridizing with a guide polynucleotide to a target sequence or a polynucleotide encoding an engineered protein provided herein. The engineered proteins are also linked to nucleic acid donor oligonucleotides for incorporation into the target nucleic acid. The engineered proteins provided herein can be fused with one or more additional protein constructs for a specific application. In some embodiments, the specific application is DNA ligation or nucleic acid cleavage. Polymerization of the engineered protein constructs provided herein can be achieved by direct fusion with another engineered protein construct. In some embodiments, each protein construct is operatively linked via an adapter peptide.

[0083] This article provides systems, compositions, and engineered proteins containing nuclease or nicking enzyme regions. Nucleases are enzymes that cleave nucleic acids. For example, nucleases can create single-strand or double-strand breaks in target nucleic acids. Nucleases and nicking enzymes target specific nucleic acid sequences by binding to guide polynucleotides that hybridize with the target nucleic acid. Various types of nucleases can be used in the systems and compositions provided herein.

[0084] This document provides systems, compositions, and engineered proteins comprising a nicking enzyme or a nicking enzyme region. In some embodiments, the engineered protein comprises a nicking enzyme. In some embodiments, the engineered protein provided herein comprises an engineered protein construct comprising a nicking enzyme region. A nicking enzyme is an enzyme that cleaves one strand of double-stranded DNA.

[0085] In some embodiments, the nicking enzymes provided herein cleave one strand of a DNA duplex to produce a nicked DNA molecule. Exemplary amino acid sequences for nucleases and nicking enzymes include, but are not limited to, for example, NCBI gene ID:1238121: NP_858382.1 [Conjugation transfer nicking enzyme / helicase TraI (plasmid) [Shigella freundii ( Shigella flexneri ) 2a str. 301]] UniProt: Q8YS92 [DNA nicking enzyme - Nostoc sp.( Nostoc sp .)(strain PCC 7120 / SAG25.82 / UTEX 2576)]: MVSTLDDTKRNAIAEKLADAKLLQELIIENQERFLRESTDNEISNRIRDFLEDDRKNLGIIETVIVQYGIQKEPRQTVREMVDQVRQLMQGSQLNFFEKVAQHELLKHKQVMSGLLVHKAAQKVGADVLAAIGPLNTVNFENRAHQEQLKGILEILGVRELTGQDADQGIWGRVQDAIAAFSGAVGSAVTQGSDKQDMNIQDVIRMDHNKVNILFTELQQSNDPQKIQEYFGQIYKDLTAHAEAEEEVLYPRVRSFYGEGDTQELYDEQSEMKRLLEQIKAISPSAPEFKDRVRQLADIVMDHVRQEESTLFAAIRNNLSSEQTEQWATEFKAAKSKIQQRLGGQATGAGV (SEQ ID NO: 68); Streptococcus pyogenes( Streptococcus pyogenes ) Cas9: Streptococcus pyogenes Cas9 H840A nickase ( Bold / Underlined indicating amino acid substitution): MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 70); Streptococcus pyogenes Cas9 R221K, N394K, H840A nickase ( Bold / Underlined indicating amino acid substitutions): MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSR KLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKL K REDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVD AIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (SEQ ID NO: 71); Staphylococcus aureus ( Staphylococcus aureus ) Cas9 N580A nickase ( Bold / Underlined indicating an amino acid substitution): MKRNYILGLDIGITSVGYGIIDYETRDVIDAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLFDYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRRGVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKKDGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFK QKKKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQISNLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRKLVPKKVDLSQQK EIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIELAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLIEKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFDNSFNNKVLVKQEE ASKKGNRTPFQYLSSSDSKISYETFKKHILNLAKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRGLMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGYK HHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESMPEIETEQEYKEIFITPHQIKDFKDYKYSHRVDKKPNRELINDTLYSTRKDDKGNTLIVNNLNGLYDKDNDKLKK LINKSPEKLLMYHHDPQTYQKLKLIMEQYGDEKNPLYKYYEETGNYLTKYSKKDNGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDVYLDNGVYKFVTVKNLDVIKKE NYYEVNSKCYEEAKKLKKISNQAEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKKG(SEQ ID NO: 72); and The new culprit is Francisella foenum-graecum ( Francisella novicida Cas9 H969A nickase (bold / underlined to indicate amino acid substitutions): MNFKILPIAIDLGVKNTGVFSAFYQKGTSLERLDNKNGKVYELSKDSYTLLMNNRTARRHQRRGIDRKQLVKRLFKLIWTEQLNLEWDKDTQQAISFLFNRRGFSFITDGYSPEYLNIVPEQVKAILMDIFDDYNGEDDLDSYLKLATEQESKISEIYNKLMQKILEFKLMKLCTDIKDDKVSTKTLKEITSYEFELLADYLANYSESLKTQKFSYTDKQGNLKELSYYHHDKYNIQEFLKRHATINDRILDTLLTDDLDIWNFNFEKFDFDKNEEKLQNQEDKDHIQAHLHHFVFAVNKIKSEMASGGRHRSQYFQEITNVLDENNHQEGYLKNFCENLHNKKYSNLSVKNLVNLIGNLSNLELKPLRKYFNDKIHAKADHWDEQKFTETYCHWILGEWRVGVKDQDKKDGAKYSYKDLCNELKQKVTKAGLVDFLLELDPCRTIPPYLDNNNRKPPKCQSLILNPKFLDNQYPNWQQYLQELKKLQSIQNYLDSFETDLKVLKSSKDQPYFVEYKSSNQQIASGQRDYKDLDARILQFIFDRVKASDELLLNEIYFQAKKLKQKASSELEKLESSKKLDEVIANSQLSQILKSQHTNGIFEQGTFLHLVCKYYKQRQRARDSRLYIMPEYRYDKKLHKYNNTGRFDDDNQLLTYCNHKPRQKRYQLLNDLAGVLQVSPNFLKDKIGSDDDLFISKWLVEHIRGFKKACEDSLKIQKDNRGLLNHKINIARNTKGKCEKEIFNLICKIEGSEDKKGNYKHGLAYELGVLLFGEPNEASKPEFDRKIKKFNSIYSFAQIQQIAFAERKGNANTCAVCSADNAHRMQQIKITEPVEDNKDKIILSAKAQRLPAIPTRIVDGAVKKMATILAKNIVDDNWQNIKQVLSAKHQLHIPIITESNAFEFEPALADVKGKSLKDRRKKALERISPENIFKDKNNRIKEFAKGISAYSGANLTDGDFDGAKEELD AIIPRSHKKYGTLNDEANLICVTRGDNKNKGNRIFCLRDLADNYKLKQFETTDDLEIEKKIADTIWDANKKDFKFGNYRSFINLTPQEQKAFRALFLADENPIKQAVIRAINNRNRTFVNGTQRYFAEVLANNIYLRAKKENLNTDKISFDYFGIPTIGNGRGIA EIRQLYEKVDSDIQAYAKGDKPQASYSHLIDAMLAFCIAADEHRNDGSIGLEIDKNYSLYPLDKNTGEVFTKDIFSQIKITDNEFSDKKLVRKKAIEGFNTHRQMTRDGIYAENYLPILIHKELNEVRKGYTWKNSEEIKIFKGKKYDIQQLNNLVYCLKFVDKP ISIDIQISTLEELRNILTTNNIAATAEYYYINLKTQKLHEYYIENYNTALGYKKYSKEMEFLRSLAYRSERVKIKSIDDVKQVLDKDSNFIIGKITLPFKKEWQRLYREWQNTTIKDDYEFLKSFFNVKSITKLHKKVRKDFSLPISTNEGKFLVKRKTWDNNFIYQILNDSDSRADGTKPFIPAFDISKNEIVEAIIDSFTSKNIFWLPKNIELQKVDNKNIFAIDTSKWFEVETPSDLRDIGIATIQYKIDNNSRPKVRVKLDYVIDDDSKINYFMNHSLLKSRYPDKVLEILKQSTIIEFESSGFNKTIKEMLGMKLAGIYNETSNN (SEQ ID NO: 73). In some embodiments, the systematic or engineered protein provided herein comprises a sequence that is at least 85% identical to any one of SEQ ID NOs: 67-73. In some embodiments, the system or engineered protein provided herein comprises a sequence that is at least 90% identical to any one of SEQ ID NO: 67-73. In some embodiments, the system or engineered protein provided herein comprises a sequence that is at least 95% identical to any one of SEQ ID NO: 67-73. In some embodiments, the system or engineered protein provided herein comprises a sequence that is at least 99% identical to any one of SEQ ID NO: 67-73. In some embodiments, the system or engineered protein provided herein comprises any one of SEQ ID NO: 67-73.

[0086] In some embodiments, the engineered protein comprises a Cas protein or a fragment thereof. In some embodiments, the engineered protein comprises a Cas protein or a fragment thereof containing at least one amino acid substitution relative to a reference amino acid sequence (e.g., wild-type sequence) of the Cas protein or a fragment thereof. In some embodiments, Cas is catalytically inactivated Cas or partially inactivated Cas. For example, partially inactivated Cas may comprise a nicking enzyme. In some embodiments, catalytically inactivated Cas or partially inactivated Cas is selected from the group consisting of catalytically inactivated derivatives: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas11, Cas12a (Cpf1), Cas... 12b, Cas13, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr 1. Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, C sf1, Csf2, CsO, Csf4, c2c1, c2c3, Cas9HiFi, xCas9, CasX, CasY, CasRX, SpCas9-VQR, SpCas9-VRQR, SpCas9-VRER, SaCas9-KKH, SpCas9-NG, SpCas9-NRRH, SpCas9-NRTH, SpCas9-NRCH, iSpyMac, St1Cas9 LMD9-LMG18311, St1Cas9 LMD9-CNRZ1066, St1Cas9-KQKL, and their variants, fragments, mutants, or derivatives. In some embodiments, the Cas protein or mutant Cas protein provided herein is a type V Cas protein. In some embodiments, type V Cas proteins include: Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas14, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, and their variants, fragments, mutants, or derivatives. In some embodiments, the Cas protein is a chimeric Cas protein comprising domains from the different Cas proteins provided herein.

[0087] In some embodiments, the Cas protein is derived from bacteria. In some embodiments, the bacteria belong to the following genera: Streptococcus (…). Streptococcus Francisella spp. Francisella Lactococcus spp. Lactococcus Lactobacillus () Lactobacillus), Pseudomonas spp. Pseudomonas ), Bacillus spp. Geobacillus Clostridium ( Clostridium Streptomyces ( Streptomyces ), genus Actinomycetes ( Actinoplanes Synechocybe ( Synechococcus Corynebacterium spp. Corynebacterium ), halophilic bacteria ( Haloferax ), genus *Saltboxia* Haloarcula ), Methanococcus spp. Methanococcus ), Neisseria spp. Neisseria ), Campylobacter spp. Campylobacter ) or Staphylococcus spp. Staphylococcus ).

[0088] In some embodiments, the Cas protein is derived from bacteria, wherein the bacteria is *Streptococcus pyogenes* (e.g., SpCas9). In some embodiments, the Cas protein is *Streptococcus pyogenes* Cas9 (SpCas9), wherein *Streptococcus pyogenes* Cas9 contains at least 90% of the sequence identical to SEQ ID NO: 69. In some embodiments, the Cas protein is *Streptococcus pyogenes* Cas9, wherein *Streptococcus pyogenes* Cas9 contains at least 95% of the sequence identical to SEQ ID NO: 69. In some embodiments, the Cas protein is *Streptococcus pyogenes* Cas9, wherein *Streptococcus pyogenes* Cas9 contains the same sequence as SEQ ID NO: 69. In some embodiments, *Streptococcus pyogenes* Cas9 contains a mutation. In some embodiments, *Streptococcus pyogenes* Cas9 contains a mutation, wherein the mutation comprises one or more amino acid substitutions.

[0089] In some embodiments, the nuclease or nickase region of the protein construct binds to the guide polynucleotide provided herein. In some embodiments, the nuclease or nickase region of the protein construct binds to the protein-binding region of the guide polynucleotide provided herein. In some embodiments, the nuclease or nickase region of the protein construct binds to the target nucleic acid to form a complex. In some embodiments, the complex also includes a portion of the guide polynucleotide sequence. For example, the protein-binding region and the target region of the guide polynucleotide form a complex between the nuclease or nickase region of the engineered protein provided herein and the target nucleic acid via a target sequence.

[0090] The binding of nucleases, nicking enzymes, or fragments thereof is mediated by complete or partial complementarity of the target region of the guide polynucleotide provided herein. In some embodiments, an engineered protein construct (e.g., a nicking enzyme or nuclease) binds to the guide polynucleotide provided herein, wherein the guide polynucleotide binds to or is adjacent to a PAM sequence. Enzymes and proteins from different bacterial species may recognize different sequence motifs or PAMs. In some embodiments, the engineered protein construct binds to the guide polynucleotide, which binds to the target nucleic acid within 5, 10, 15, or 20 nucleotides of the PAM sequence.

[0091] After binding to the target nucleic acid via a guide sequence, the nucleases and nicking enzymes provided herein specifically cleave the target nucleic acid at the distal or proximal end of the PAM. In some embodiments, the nuclease or nicking enzyme region of the protein construct generates single-strand breaks in the target nucleic acid. In some embodiments, the nicking enzyme region generates double-strand breaks in the target nucleic acid. In some embodiments, the nucleases and nicking enzymes provided herein generate blunt ends when cleaving the target nucleic acid. In some embodiments, the nucleases provided herein generate staggered ends when cleaving the target nucleic acid, which results in higher integration rates of the synthesized DNA, thereby improving gene editing efficiency. In some embodiments, the nucleases and nicking enzymes provided herein cleave the non-target strand (NTS). In some embodiments, the nucleases and nicking enzymes provided herein cleave the complementary strand.

[0092] This document provides systems, compositions, and engineered proteins comprising a DNA ligase or a functional fragment thereof. In some embodiments, the engineered protein comprises a DNA ligase or a DNA ligase region. In some embodiments, the DNA ligase region is operatively ligated to a nuclease or nicking enzyme region or nicking enzyme region provided herein. In some embodiments, the DNA ligase is an ATP-dependent ligase. In some embodiments, the DNA ligase is a bacterial, eukaryotic, insect, or plant DNA ligase. In some embodiments, the DNA ligase is selected from the group consisting of *Escherichia coli* (…). E. coli DNA ligase, Taq DNA ligase, T4 DNA ligase, human DNA ligase I, human DNA ligase III, human DNA ligase IV. In some embodiments, the DNA ligase is T4 DNA ligase. Exemplary DNA ligase sequences are listed in Table 1.

[0093] Table 1. DNA ligase sequences In some embodiments, the DNA ligase, DNA ligase region, or DNA ligase fragment binds to the ligation clip 1 region of the guiding polynucleotide. In some embodiments, the DNA ligase does not bind to the target nucleic acid sequence. In some embodiments, the DNA ligase attaches the 5' end of the donor nucleic acid to the 3' end of the leading strand. The donor nucleic acid can replace at least a portion of the target sequence in the double-stranded target nucleic acid, thereby editing the double-stranded target nucleic acid.

[0094] In some embodiments, the engineered proteins provided herein further comprise one or more additional protein regions or protein constructs, wherein one or more additional protein regions or protein constructs comprise: nucleases, acetyltransferases, ATPases, Argonaute proteins, base editors, Cas peptides, catalytically inactivated Cas peptides, deacetylases, deaminases, uncapping proteins, endonucleases, exonucleases, helicases, ligases, megnucleases, methyltransferases, cleavage enzymes, polymerases, proteases, recombinases, restriction enzymes, ribonucleoproteins (RNPs), self-cleaving protein sequences, splicing factors, transcription activators, transcription activator-like effector nucleases (TALENs), transcription repressors, transposases, zinc fingers, or any combination thereof. In some embodiments, the compositions and systems provided herein comprise additional engineered proteins. In some implementations, additional engineered proteins include: nucleases, acetyltransferases, acetyltransferases, ATPases, Argonaute proteins, base editors, Cas peptides, catalytically inactivated Cas peptides, deacetylases, deaminases, uncapping proteins, endonucleases, exonucleases, helicases, ligases, megnucleases, methyltransferases, cleavage enzymes, polymerases, proteases, recombinases, restriction enzymes, ribonucleoproteins (RNPs), self-cleaving protein sequences, splicing factors, transcription activators, transcription activator-like effector nucleases (TALENs), transcription repressors, transposases, zinc fingers, or any combination thereof.

[0095] In some embodiments, the engineered proteins provided herein comprise a protein construct that regulates transcription. In some embodiments, the engineered proteins provided herein comprise a transcriptional repressor protein construct. In some embodiments, the engineered proteins provided herein comprise a zinc finger protein construct. In some embodiments, the zinc finger protein construct comprises a Krüppel-associated box (KRAB) protein or a functional fragment thereof. In some embodiments, the engineered protein further comprises a KRAB domain that binds to a transcriptional co-repressor protein. In some embodiments, the engineered proteins provided herein comprise a SUMO protein construct.

[0096] In some embodiments, the engineered proteins provided herein comprise transcription activator protein constructs. Transcription activator protein constructs recruit transcription factors from host cells to target nucleic acids to regulate target nucleic acid expression. Exemplary transcription activators include VP64, VP16, VP160, VP48, VP96, p65, Rta, VPR, hsf1, and p300. In some embodiments, the engineered proteins provided herein comprise one or more VP16 protein constructs. In some embodiments, the engineered proteins provided herein comprise one or more VP64 protein constructs. In some embodiments, the engineered proteins provided herein comprise one or more VPR protein constructs. In some embodiments, the engineered proteins provided herein comprise one or more SunTag protein constructs. In some embodiments, the engineered proteins provided herein comprise VP64, p65, and HSF1 (SunTag-p65-heat shock factor 1 or SPH). In some embodiments, the engineered proteins provided herein comprise CREB-binding protein (CBP). CBP can be used to recruit transcriptional mechanisms and function as a histone acetyltransferase (HAT) that alters chromatin structure.

[0097] In some embodiments, the engineered proteins provided herein also include aptamers. In some embodiments, the aptamers bind to one or more MS2 proteins.

[0098] In some embodiments, the engineered proteins provided herein comprise self-cleaving protein sequences. Non-limiting examples of self-cleaving protein sequences include E2A, P2A, and T2A. In some embodiments, the engineered proteins provided herein also comprise antibiotic resistance protein constructs or antibiotic resistance selection markers. Non-limiting examples of antibiotic resistance proteins and selection markers include: aminoglycoside acetyltransferases, rifampicin ADP-ribosyltransferases, dihydrofolate reductases, multidrug and toxic compound extrusion transporters, antibiotic resistance ATP-binding cassette family F (ARE ABC-F) proteins, β-lactamases, blastcin-S deaminases, penicillin-binding proteins (PBPs), and puromycin-N-acetyltransferases. In some embodiments, the ARE ABC-F protein is MsrE, Erm, Vga, Lsa, Sal, or OptrA.

[0099] In some embodiments, the engineered protein provided herein comprises a base editor. In some embodiments, the base editor is selected from the group consisting of: adenine base editor, adenosine base editor, cytidine deaminase, cytosine-to-guanine base editor, and deaminase dimer. In some embodiments, the cytidine deaminase is activation-induced deaminase (AID), APOBEC deaminase, APOBEC3G, APOBEC1, cytidine deaminase 1 (CDA1), or a functional fragment or derivative thereof. In some embodiments, the adenosine base editor is ecTadA, saTadA, or a functional fragment or derivative thereof. In some embodiments, the engineered protein provided herein also comprises uracil-DNA glycosylase, uracil-DNA glycosylase inhibitor, or a functional fragment or derivative thereof.

[0100] In some embodiments, the systems, compositions, or engineered proteins provided herein further comprise nuclear localization sequences (NLS). The NLS targets the protein to the cell nucleus, localizing the engineered proteins provided herein to a target nucleic acid immediately adjacent to the cell nucleus. In some embodiments, the systems, compositions, or engineered proteins comprise more than one nuclear localization sequence (NLS). In some embodiments, the NLS is derived from simian vacuolating virus 40 (SV40). In some embodiments, the NLS comprises a single SV40 NLS. In some embodiments, the NLS comprises a two-part SV40 NLS. In some embodiments, the NLS comprises PKKKRKV (SEQ ID NO: 74) or KRTADGSEFEPKKKRKV (SEQ ID NO: 75). In some embodiments, the NLS comprises a nucleoplasmic protein sequence. In some embodiments, the nucleoplasmic protein sequence comprises: KRPAATKKAGQAKKKK (SEQ ID NO: 76).

[0101] In some embodiments, the systems, compositions, or engineered proteins provided herein further comprise adapters. An adapter is a molecular entity that can directly or indirectly connect two parts of a composition. For example, an adapter can link a first protein construct to a second protein construct, and vice versa, and so on. The adapter can be configured to meet specific needs. For example, the adapter can be configured to improve stability or achieve an optimal length between two amino acid sequences. In some embodiments, the adapter can be configured to allow nucleases or nicking enzymes to polymerize with the DNA ligases provided herein (e.g., from monomers to dimers, trimers, tetramers, pentamers, or higher polymeric complexes) while retaining biological activity. In some embodiments, the biological activity includes the cleavage of a target nucleic acid or a group of target nucleic acids. In some embodiments, the adapter may be selected from (GS). n (SEQ ID NO: 229) or (GGS)n (SEQ ID NO: 230), where n is 1, 2, 3, 4, or 5. Other exemplary connectors are provided in Table 2.

[0102] Table 2. Connector Sequence In some embodiments, the adapter is configured to facilitate the expression and purification of the nucleases or nicking enzymes provided herein. In some embodiments, the engineered proteins provided herein contain cleavable or non-cleavable adapters between different protein constructs or domains of the engineered protein. For example, the adapter may be a peptide adapter, such as an adapter of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids. Typically, the number of additional components in an engineered protein will depend on many factors, including, for example, the size or molecular weight of the protein used for delivery to the cell, or the desired translocation of the protein in the cell type of interest.

[0103] This document provides engineered protein constructs and polynucleotides encoding engineered proteins, wherein the engineered protein constructs comprise engineered nickases and DNA ligases or functional fragments thereof. In some embodiments, the engineered protein constructs further comprise a linker between a nickase region or nuclease region and a DNA ligase or functional fragment thereof. In some embodiments, the linker comprises an amino acid sequence that is at least 95% identical to any one of SEQ ID NO: 77-88. In some embodiments, the linker comprises an amino acid sequence that is at least 99% identical to any one of SEQ ID NO: 77-88. In some embodiments, the linker comprises any one of SEQ ID NO: 77-88.

[0104] Combination composition This document provides compositions comprising: (a) a donor nucleic acid; (b) a polynucleotide encoding a nickase or a variant thereof; (c) a polynucleotide encoding a protein construct comprising a nickase or a variant thereof and a DNA ligase or a functional fragment thereof; and (d) a guide polynucleotide described herein. In some embodiments, the nickase is an engineered Cas protein or a functional variant thereof. In some embodiments, the engineered Cas protein or a functional variant thereof comprises at least one amino acid substitution corresponding to position 840 of SEQ ID NO: 69. In some embodiments, the engineered Cas protein or a functional variant thereof further comprises at least one amino acid substitution corresponding to positions 221, 394, or a combination thereof of SEQ ID NO: 69. In some embodiments, the engineered Cas protein or a functional variant thereof comprises an amino acid substitution of R221K, N394K, H840A, or any combination thereof. In some embodiments, the engineered Cas protein or a functional variant thereof comprises an amino acid substitution of H840A. In some embodiments, the engineered Cas protein or a functional variant thereof comprises an amino acid substitution of R221K. In some embodiments, the engineered Cas protein or a functional variant thereof comprises an amino acid substitution of N394K. In some embodiments, the engineered Cas protein or a functional variant thereof comprises 97%, 98%, 99%, or 100% identical to the sequence of SEQ ID NO: 70 or SEQ ID NO: 71. In some embodiments, the engineered Cas protein or a functional variant thereof comprises the sequence of SEQ ID NO: 70 or SEQ ID NO: 71. In some embodiments, the DNA ligase or functional fragment comprises *E. coli* DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, *Chlorella virus* DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human ligase V, or variants or combinations thereof. In some embodiments, the DNA ligase or a functional fragment thereof comprises *Chlorella virus* DNA ligase, T4 DNA ligase, or human DNA ligase IV. In some embodiments, the DNA ligase or a functional fragment thereof comprises at least 97%, at least 98%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 58 to 66. In some embodiments, the DNA ligase or a functional fragment thereof comprises the sequence of SEQ ID NO: 58 to 66. In some embodiments, the DNA ligase or a functional fragment thereof comprises a Chlorella virus DNA ligase comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 58.In some embodiments, the DNA ligase or a functional fragment thereof includes a Chlorella virus DNA ligase containing the sequence of SEQ ID NO: 58.

[0105] This document provides engineered fusion proteins comprising (a) a DNA ligase or a functional fragment thereof; and (b) an engineered nickase comprising three amino acid substitutions at positions 221, 394, and 840 of a nuclease corresponding to the sequence comprising SEQ ID NO: 69. In some embodiments, the substitutions result in enhanced nickase activity compared to other engineered nickases that are equivalent in other respects. In some embodiments, the engineered nickase comprises amino acid substitutions of R221K, N394K, and H840A. In some embodiments, the engineered nickase or a functional variant thereof comprises a sequence that is 97%, 98%, 99%, or 100% identical to SEQ ID NO: 71. In some embodiments, the engineered nickase or a functional variant thereof comprises the sequence of SEQ ID NO: 71. In some embodiments, the DNA ligase or functional fragment includes *E. coli* DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, *Chlorella virus* DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human ligase V, variants or combinations thereof. In some embodiments, the DNA ligase or functional fragment thereof includes *Chlorella virus* DNA ligase, T4 DNA ligase, or human DNA ligase IV. In some embodiments, the DNA ligase or functional fragment thereof comprises 97%, 98%, 99%, or 100% identical to the sequence in SEQ ID NO: 58 to 66. In some embodiments, the DNA ligase or functional fragment thereof comprises the sequence in SEQ ID NO: 58 to 66. In some embodiments, the DNA ligase or functional fragment thereof includes a *Chlorella virus* DNA ligase comprising at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the sequence in SEQ ID NO: 58. In some embodiments, the DNA ligase or a functional fragment thereof comprises a Chlorella virus DNA ligase containing the sequence of SEQ ID NO: 58. In some embodiments, the DNA ligase or a functional fragment thereof comprises a Chlorella virus DNA ligase containing at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical sequence to SEQ ID NO: 58. In some embodiments, the DNA ligase or a functional fragment thereof comprises a Chlorella virus DNA ligase containing the sequence of SEQ ID NO: 58. In some embodiments, the engineered fusion protein further comprises an adapter, a nuclear localization sequence (NLS), or another protein construct.In some implementations, additional protein constructs include cell-targeting moieties, receptor-targeting moieties, regulatory elements, nucleases, acetyltransferases, ATPases, Argonaute proteins, base editors, Cas peptides, catalytically inactivated Cas peptides, deacetylases, deaminases, uncapping proteins, endonucleases, exonucleases, helicases, ligases, megnucleases, methyltransferases, nickases, polymerases, proteases, recombinases, restriction enzymes, ribonucleoproteins (RNPs), self-cleaving protein sequences, splicing factors, transcription activators, transcription activator-like effector nucleases (TALENs), transcription repressors, transposases, zinc fingers, or any combination thereof.

[0106] This document provides engineered fusion proteins comprising at least 90% identical amino acid sequences to any one of SEQ ID NO: 89-101. This document provides engineered fusion proteins comprising at least 95% identical amino acid sequences to any one of SEQ ID NO: 89-101. This document provides engineered fusion proteins comprising at least 96% identical amino acid sequences to any one of SEQ ID NO: 89-101. This document provides engineered fusion proteins comprising at least 97% identical amino acid sequences to any one of SEQ ID NO: 89-101. This document provides engineered fusion proteins comprising at least 98% identical amino acid sequences to any one of SEQ ID NO: 89-101. This document provides engineered fusion proteins comprising at least 99% identical amino acid sequences to any one of SEQ ID NO: 89-101. This document provides engineered fusion proteins comprising at least 95% identical amino acid sequences to any one of SEQ ID NO: 89-101, wherein the engineered fusion proteins contain at least one, at least two, at least three, at least four, at least five, at least six, or at least seven amino acid substitutions in the DNA ligase region. This document also provides engineered fusion proteins comprising the amino acid sequences of SEQ ID NO: 89-101.

[0107] This document provides nucleic acids and polynucleotides encoding any of the engineered fusion proteins provided herein. In some embodiments, the polynucleotide encoding the engineered fusion protein includes DNA. In some embodiments, the polynucleotide encoding the engineered fusion protein includes RNA. This document provides vectors comprising any nucleic acid or polynucleotide provided herein, or more than one nucleic acid or polynucleotide as provided herein. In some embodiments, the nucleic acids, vectors, or viral vectors provided herein also comprise the guide polynucleotides provided herein. In some embodiments, the nucleic acids, vectors, or viral vectors provided herein also comprise DNA encoding the guide polynucleotides provided herein.

[0108] target nucleic acid The guiding polynucleotide provided herein may contain a degree of complementarity with the target polynucleotide sequence of interest or its strand. The targeting region of the guiding polynucleotide provided herein may contain at least partial sequence complementarity with the target polynucleotide. The targeting sequence may have a degree of sequence complementarity with the target nucleic acid sufficient to enable hybridization between the guiding polynucleotide and the target polynucleotide. In some cases, the targeting sequence contains 95%, 96%, 97%, 98%, 99%, or 100% sequence complementarity with the target polynucleotide. Sequence complementarity can be determined using alignment methods known in the art, such as sequence alignment using publicly available software such as BLAST, Align, or ClustalW2. In some embodiments, the target nucleic acid provided herein comprises a gene or polynucleotide containing DNA. In some embodiments, the gene or polynucleotide comprises a mammalian gene or polynucleotide. In some embodiments, the gene or polynucleotide comprises a human gene or polynucleotide. In some embodiments, the target nucleic acid provided herein comprises DNA. In some embodiments, the target nucleic acid provided herein comprises single-stranded DNA. In some embodiments, the target nucleic acid provided herein comprises double-stranded DNA. In some embodiments, the target nucleic acid provided herein comprises RNA. In some embodiments, the target nucleic acid provided herein includes single-stranded RNA. In some embodiments, the target nucleic acid provided herein includes double-stranded RNA.

[0109] In some embodiments, the target nucleic acid or complementary strand provided herein contains a protospacer adjacent motif (“PAM”), wherein the PAM is a short, T-rich sequence. In some embodiments, cleavage of the target nucleic acid by the engineered protein provided herein occurs downstream of or 3' of the PAM sequence. In some embodiments, cleavage of the target nucleic acid occurs upstream of or 5' of the PAM sequence. In some embodiments, the PAM is an NGG PAM sequence. In some embodiments, the PAM includes NGAN, NGNG, NGAG, NGCG, where N is A, G, C, or T. In some embodiments, the PAM is a T-rich PAM. In some embodiments, the PAM has the nucleotide sequence (T)XN, where X is the number of thymine (e.g., 1-10), and N is A, G, C, or T. In some embodiments, X equals 2, and therefore, the PAM is TTN. In some embodiments, X is 3, and therefore, the PAM is TTTN. In some implementations, the nuclease or nickase provided herein recognizes the sequence motif TTTN and directs the cleavage of target nucleic acid sequences 1-24 bp downstream of that sequence (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24).

[0110] Without limitation, the target nucleic acid sequence can be derived from any cell or organism. Determining the appropriate sequence for the binding of the guide polynucleotide to the target nucleic acid provided herein will depend on the desired target sequence and the structure of the nuclease or nicking enzyme provided herein. Methods for designing targeting regions of guide polynucleotides for gene editing are known in the art and include, for example, the use of software and databases such as Breaking-Cas, Cas-OFFinder, CRISPR-DT, CHOPCHOP, CCTOP, CRISPick, or CRISPOR.

[0111] In some implementations, the target nucleic acid comprises double-stranded DNA (dsDNA). In the case of dsDNA, the polynucleotide and cleavage enzyme regions of the engineered protein provided herein bind to the target nucleic acid and generate single-strand breaks in the dsDNA. Cleavage can occur on either the bottom or top strand of the double-stranded DNA. Cleavage of dsDNA produces DNA replication forks as well as leading and complementary strands (also known as lagging strands). The guiding polynucleotide binds to both strands of the target DNA via different regions of the guiding polynucleotide. For example, after cleavage of dsDNA by the cleavage enzyme or cleavage enzyme region of the engineered polynucleotide, the ligation splint 1 region binds to the leading strand of the DNA. The target region of the guiding polynucleotide contains RNA and binds to the complementary strand containing a PAM sequence. The ligation splint 2 region mediates hybridization with the ligation donor nucleic acid. After the ligation splint 1 region binds to the leading strand and the donor nucleic acid hybridizes with the ligation splint 2 region, the DNA ligase of the engineered protein catalyzes the formation of a phosphodiester bond between the 3' hydroxyl terminus of the leading strand and the 5' phosphate terminus of the donor nucleic acid. DNA ligase does not bind to the complementary strand containing the PAM sequence on the target nucleic acid. After the DNA ligase completes the ligation, it directs the polynucleotide to dissociate from the nicked complementary strand of the target nucleic acid, and the donor nucleic acid is incorporated into the target nucleic acid. In some cases, DNA bumps form, where DNA editing is included in the ligated DNA sequence because the new sequence differs from the complementary target DNA strand.

[0112] Additional DNA mismatch repair proteins recognize bumps and can facilitate the correction and modification of abnormal sequences (relative to wild-type reference sequences) in the target nucleic acid. In some embodiments, the target nucleic acid can be repaired by DNA mismatch repair (MMR) proteins, homology-directed repair (HDR) proteins, or non-homologous end joining (NHEJ) repair proteins. In some embodiments, the MMR, HDR, and / or NHEJ proteins are endogenous to the cell. In some embodiments, the MMR, HDR, and / or NHEJ proteins are introduced into cells or cell-free systems. In some cases, donor nucleic acids or LS2 regions can be integrated into the genome when the cleavage enzyme cleaves the bottom strand of the DNA (e.g., the complementary strand) away from editing.

[0113] (2) Delivery medium and carrier This document provides compositions comprising the guide polynucleotides and delivery mediators provided herein. This document provides compositions comprising the engineered proteins and delivery mediators provided herein. This document provides systems and one or more delivery mediators provided herein. The compositions and cells provided herein can be delivered to target cells, tissues, organs, or subjects by any suitable means.

[0114] The engineered proteins, guide polynucleotides, and any polynucleotides encoding engineered proteins or guide polynucleotides provided herein may be mixed with delivery media that allow the system to be delivered to cells, tissues, or subjects. Polynucleotides and polynucleotide groups encoding the engineered proteins and / or guide polynucleotides provided herein include DNA, RNA, or both DNA and RNA. In some embodiments, RNA encodes the engineered proteins or guide polynucleotides provided herein. For example, RNA delivery of a protein can improve the expression of the engineered protein in human cells. In some embodiments, the polynucleotides provided herein include self-replicating RNA or viral RNA.

[0115] In some implementations, the delivery medium is a liposome. Liposomes are formed from phospholipids dispersed in an aqueous medium and spontaneously form multilayered concentric bilayer vesicles (also known as multilayered vesicles (MLVs)). MLVs typically have a diameter of 25 nm to 4 μm. Acoustic treatment of MLVs results in the formation of small monolayered vesicles (SUVs) containing an aqueous solution in the core, with diameters ranging from 200 to 500 angstroms. Liposomes interact with cells via various mechanisms: endocytosis by phagocytes of the reticuloendothelial system, such as macrophages and neutrophils; adsorption to the cell surface via nonspecific weak hydrophobic or electrostatic forces, or via specific interactions with cell surface components; fusion with the plasma cell membrane by inserting the lipid bilayer of the liposome into the plasma membrane, while simultaneously releasing the liposome contents into the cytoplasm; and transfer of liposome lipids to the cell membrane or subcellular membrane, or vice versa, without any association of the liposome contents. Modifying the liposome formulation can alter the mechanism of action, although more than one mechanism may act simultaneously. Nanocapsules can generally capture compounds in a stable and reproducible manner. To avoid side effects caused by intracellular polymer overload, such ultrafine particles (approximately 0.1 μm in size) should be designed using polymers that can degrade in vivo. Biodegradable polycyanoacrylate nanoparticles can also be used as delivery media.

[0116] In some implementations, the delivery medium is phospholipids. When dispersed in water, phospholipids can form a variety of structures besides liposomes, depending on the molar ratio of lipid to water. At low ratios, liposomes form. The physical characteristics of liposomes depend on pH, ionic strength, and the presence of divalent cations. Liposomes can exhibit low permeability to ions and polar substances, but undergo a phase transition at elevated temperatures, which significantly alters their permeability. The phase transition involves a change from a closely packed, ordered structure (called the gel state) to a loosely packed, less ordered structure (called the fluid state). This occurs at the characteristic phase transition temperature and leads to increased permeability to ions, sugars, and drugs.

[0117] In some embodiments, the delivery medium is nanoparticles. The tissue-specific nanoparticle delivery vehicles provided herein can also be used as pharmaceutically acceptable delivery vehicles. In some embodiments, the nanoparticles are gold nanoparticles, platinum nanoparticles, iron oxide nanoparticles, lipid nanoparticles, selenium nanoparticles, tumor-targeting ethylene glycol chitosan nanoparticles (CNP), cathepsin B-sensitive nanoparticles, hyaluronic acid nanoparticles, paramagnetic nanoparticles, or polymer nanoparticles. In some embodiments, the delivery medium is lipid nanoparticles.

[0118] This document provides compositions comprising lipid nanoparticles, the lipid nanoparticles comprising: the system provided herein, a polynucleotide sequence encoding the system provided herein, a carrier provided herein, an engineered protein provided herein, a guide polynucleotide or a portion thereof, or any composition provided herein. In some embodiments, the lipid nanoparticles are solid lipid nanoparticles (SLNs) or nanostructured lipid carriers (NLCs).

[0119] The lipid nanoparticles provided herein may comprise cationic lipids, ionizable lipids, lipid mixtures, zwitterionic lipids, and / or phospholipids. In some embodiments, the lipid nanoparticles comprise cationic lipids selected from the group consisting of: 1,2-di-O-octadecenyl-3-trimethylammonium-propane (DOTMA), 1,2-dioleoyl- sn-Glycerol-3-phosphate ethanolamine (DOPE), 1,2-dioleoyl-3-trimethylammonium-propane (DOTAP), dimethyl dioctadecyl ammonium bromide (DDAB), and ethylphosphatidylcholine (ePC). In some embodiments, the lipid nanoparticles comprise ionizable lipids selected from the group consisting of: 2S)-2,5-bis(3-aminopropylamino)-N-[2-(octadecylamino)acetyl]pentanamide (DOGS; Transfectam), N1-[2-((1S)-1-[(3-aminopropyl)amino]-4-[di(3-aminopropyl)amino]butylcarbamoyl)ethyl]-3,4-di[oleoyloxy]-benzamide (MVL5), DC-cholesterol and N4-cholesterolyl-spermine (GL67), 9Z,12Z-octadecadienoic acid, 3-[4,4-bis(octyloxy)-1-oxobutoxy]-2-[[[[3-(diethylamino)propoxy]carbonyl]oxy]methyl [Lipid 5] propyl ester (LP01), heptadecan-9-yl 8-[2-hydroxyethyl-(6-oxo-6-undecyloxyhexyl)amino]octanoate (SM-102), [(4-hydroxybutyl)azanidinediyl]bis(hexane-6,1-diyl)bis(2-hexyldecanoate) (ALC-0315), bis(2-butyloctyl)10-(N-(3-(dimethylamino)propyl)nonanoylamino)nonadecanate (Lipid A9), 5-(dimethylamino)valerate, (6Z)-1,2-di-(4Z)-4-decen-1-yl-6-dodecen-1-yl ester (Lipid CL1), 7-[(2-hydroxyethyl)[8-(nonoxy)-8-oxooctyl]amino]heptyl 2-octyldecanoate (Lipid 5). In some embodiments, the lipid nanoparticles contain phosphatidylcholine. In some embodiments, the lipid nanoparticles comprise cholesterol or cholesterol analogs. In some embodiments, the lipid nanoparticles comprise cholesterol analogs selected from the group consisting of: β-sitosterol, vitamin D3, vitamin D2, calcipotriol, stigmasterol, betulin, lupeol, ursolic acid, oleanolic acid, stigmasterol, campesterol, fucosterol, brassicerol, and ergosterol. In some embodiments, the lipid nanoparticles comprise polyethylene glycol (PEG). In some embodiments, the PEG is 1,2-dimyristoyl-racemic-glycero-3-methoxy polyethylene glycol-2000 (PEG). 2000 -DMG) or 1,2-distearyl-racemic-glycero-3-methoxy polyethylene glycol-2000 (PEG) 2000-DSG). In some embodiments, the lipid nanoparticles comprise N-acetylgalactosamine (GalNAc). In some embodiments, GalNAc is 1,2-distearyl-sn-glycero-3-phosphoethanolamine-N-[tris-GalNAc-GABA-(polyethylene glycol)-2000](tris-GalNAc-PEG2000-DSPE).

[0120] In some embodiments, the composition containing lipid nanoparticles is in emulsion form. In some embodiments, the composition containing lipid nanoparticles is in liquid form. In some embodiments, the composition containing lipid nanoparticles is in gel form. In some embodiments, the composition containing lipid nanoparticles is in solid form. In some embodiments, the system provided herein, the polynucleotide set encoding the system provided herein, the carrier provided herein, the engineered protein or guide polynucleotide provided herein, or a portion thereof, is encapsulated with lipid nanoparticles. In some embodiments, the system provided herein, the polynucleotide set encoding the system provided herein, the carrier provided herein, the engineered protein or guide polynucleotide provided herein, or a portion thereof, is complexed with lipid nanoparticles.

[0121] The compositions provided herein can be delivered to cellular systems using vectors, such as those containing a polynucleotide sequence encoding the system, guide polynucleotide, engineered protein, or composition provided herein. In some embodiments, the system as described herein can be delivered without a viral vector. Any vector system can be used, including but not limited to plasmid vectors, viral vectors, and oncolytic virus vectors. Furthermore, any of these vectors may contain one or more transcription factors, transgenes, or molecular tags.

[0122] In some implementations, the vectors provided herein are viral vectors. Exemplary viral vectors include, but are not limited to, lentiviral vectors, retroviral vectors, adeno-associated virus (AAV) vectors, adenovirus vectors, herpes simplex virus vectors, alphavirus vectors, flavivirus vectors, rhabdovirus vectors, measles virus vectors, Newcastle disease virus vectors, poxvirus vectors, microRNA viral vectors, and oncolytic virus vectors.

[0123] In some embodiments, the viral vector comprises AAV. The AAV may have one or more wild-type AAV genes that are completely or partially deleted. For example, the rep and / or cap genes of the AAV may be completely or partially deleted, but functional flanking ITR sequences may still be retained. Functional ITR sequences are essential for the rescue, replication, and packaging of AAV virions. The ITR only needs to provide for functional rescue, replication, and packaging; the sequence does not need to be a wild-type polynucleotide sequence and, for example, can be altered by nucleotide insertion, deletion, or substitution. Recombinant AAV vectors (rAAV) contain infectious, replication-defective viruses consisting of an AAV protein shell encapsulating heterologous nucleotide sequences of interest flanked by AAV ITRs. rAAV vectors are generated in suitable host cells containing the AAV vector, AAV helper functions, and accessory functions. In this way, the host cell is able to encode the AAV polypeptide required to package the AAV vector (containing the recombinant nucleotide sequences of interest) into infectious recombinant virion particles for subsequent gene delivery. In some embodiments, the AAV or rAAV provided herein includes serotypes of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAVrh10, or any combination thereof. In some embodiments, the delivery medium includes a hybrid AAV-lipid nanoparticle delivery system.

[0124] In some embodiments, the viral vector is a lentiviral vector. In some embodiments, the lentiviral vector is selected from the group consisting of: human immunodeficiency virus type 1 (HIV-1); human immunodeficiency virus type 2 (HIV-2); Vesnamedi virus (VMV); caprine arthritis-encephalitis virus (CAEV); equine infectious anemia virus (EIAV); feline immunodeficiency virus (FIV); bovine immunodeficiency virus (BIV); and simian immunodeficiency virus (SIV), fragments, derivatives, or variants thereof.

[0125] Conventional virus- and non-viral gene transfer methods can be used to introduce polynucleotides encoding the compositions, systems, guide polynucleotides, or engineered proteins provided herein into cells and target tissues. Exemplary non-viral vector delivery systems may include DNA plasmids, naked nucleic acids, and nucleic acids complexed with delivery media such as liposomes or poloxamers. Viral vector delivery systems may also include DNA viruses and RNA viruses that have an appended or integrated genome after delivery to cells.

[0126] Non-viral methods for nucleic acid delivery include electroporation, lipid transfection, nuclear transfection, gold nanoparticle delivery, microinjection, gene gun, viral microsomes, liposomes, immunoliposomes, polycationic or lipid:nucleic acid conjugates, naked DNA, mRNA, artificial viruses, and agent-enhanced DNA uptake. Acoustic perforation using systems such as the Sonitron 2000 system (Rich-Mar) can also be used for nucleic acid delivery. Other exemplary nucleic acid delivery systems include those offered by Lonza Nucleofactor Technologies (Cologne, Germany), Life Technologies (Frederick, Md.), MAXCYTE, Inc. (Rockville, Md.), BTX Molecular Delivery Systems (Holliston, Mass.), and Copernicus Therapeutics Inc. Lipid transfection reagents are commercially available (e.g., TRANSFECTAM® and LIPOFECTIN®).

[0127] The compositions and systems provided herein can be delivered to cells (ex vivo administration) or target tissues (in vivo administration). Other delivery methods include packaging the polynucleotide to be delivered into an EnGeneIC delivery vehicle (EDV). These EDVs are specifically delivered to the target tissue using a bispecific antibody, wherein one arm of the antibody is specific to the target tissue and the other arm is specific to the EDV. The antibody carries the EDV to the surface of the target cells and then carries the EDV into the cells via endocytosis.

[0128] Vectors containing viral and nonviral vectors encoding nucleic acids encoded by the nucleic acid editing system provided herein can also be directly applied to organisms to transduce cells in vivo. Optionally, naked DNA or mRNA can be applied. Application is carried out via any route commonly used to introduce molecules into final contact with blood or tissue cells, including but not limited to injection, infusion, surface application, and electroporation. More than one route can be used to apply a particular composition.

[0129] In some embodiments, the compositions, engineered proteins, guide polynucleotides, or systems provided herein can shuttle to the cell nucleus. For example, the vector may contain a nuclear localization sequence (NLS). The vectors or any compositions provided herein may also shuttle via proteins or protein complexes. In some embodiments, the compositions or systems provided herein can be introduced into cells or target tissues via small circular vectors.

[0130] In some embodiments, the vector or polynucleotide provided herein may be pre-complexed with the engineered protein provided herein prior to electroporation into cells. The engineered protein that can be used for shuttle may be a nicking enzyme or a catalytically inactivated Cas protein. The nuclease that can be used for shuttle may be a protein with nuclease-competent activity. In some embodiments, the engineered protein provided herein may be pre-mixed with the guide polynucleotide provided herein and any other elements, such as transgenes or other engineered proteins.

[0131] Cells can be transfected using mutant or chimeric adeno-associated virus vectors encoding the systems or compositions provided herein. For example, the AAV vector concentration can be from about 0.5 nanograms to as high as 50 micrograms.

[0132] The systems or compositions provided herein can also be introduced into cells via electroporation. The amount of polynucleotides that can be introduced into cells via electroporation can be varied to optimize transfection efficiency and / or cell viability. In some embodiments, less than about 100 picograms of nucleic acid may be added to each cell sample (this may include one or more cells being electroporated). In some embodiments, at least about 100 picograms, at least about 200 picograms, at least about 300 picograms, at least about 400 picograms, at least about 500 picograms, at least about 600 picograms, at least about 700 picograms, at least about 800 picograms, at least about 900 picograms, at least about 1 microgram, at least about 1.5 micrograms, at least about 2 micrograms, at least about 2.5 micrograms, at least about 3 micrograms, at least about 3.5 micrograms, at least about 4 micrograms, at least about 4.5 micrograms, at least about 5 micrograms, at least about 5.5 micrograms, at least about 6 micrograms, at least about 6 micrograms, at least about 6 micrograms, at least about 6 micrograms, at least about 1 microgram, at least about 1.5 micrograms, at least about 2 micrograms, at least about 2.5 micrograms, at least about 3 micrograms, at least about 3.5 micrograms, at least about 4 micrograms, at least about 4.5 micrograms, at least about 5 micrograms, at least about 5.5 micrograms, at least about 6 micrograms, at least about 6 micrograms, at least about 6 micrograms, at least about 1 microgram, at least about 1.5 micrograms, at least about 2 micrograms, at least about 2.5 micrograms, at least about 3 micrograms, at least about Micrograms, at least about 6.5 micrograms, at least about 7 micrograms, at least about 7.5 micrograms, at least about 8 micrograms, at least about 8.5 micrograms, at least about 9 micrograms, at least about 9.5 micrograms, at least about 10 micrograms, at least about 11 micrograms, at least about 12 micrograms, at least about 13 micrograms, at least about 14 micrograms, at least about 15 micrograms, at least about 20 micrograms, at least about 25 micrograms, at least about 30 micrograms, at least about 35 micrograms, at least about 40 micrograms, at least about 45 micrograms, or at least about 50 micrograms of nucleic acid may be added to each cell sample. For example, 1 microgram of the nucleic acid, polynucleotide, vector, or composition provided herein may be added to each cell sample for electroporation. In some embodiments, the amount of nucleic acid required for optimal transfection efficiency and / or cell viability may be cell type specific. In some embodiments, the amount of nucleic acid used for each sample may directly correspond to transfection efficiency and / or cell viability. Using any nucleic acid delivery platform described herein, such as nuclear transfection or electroporation, the transfection efficiency of cells can be, or can be approximately, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or more than 99.9%.

[0133] Viral particles, such as AAV, can be used to deliver viral vectors containing genes or transgenes of interest into cells, either ex vivo or in vivo. In some embodiments, mutant adeno-associated virus vectors or chimeric adeno-associated virus vectors, as disclosed herein, can be measured in pfu (plaque-forming units). In some embodiments, the pfu of the recombinant virus or mutant adeno-associated virus vector or chimeric adeno-associated virus vector of the compositions and methods of this disclosure can be about 10. 8 Approximately 5×10 10pfu. In some implementations, the recombinant virus of this disclosure is at least about 1 × 10⁻⁶. 8 2×10 8 3×10 8 4×10 8 5×10 8 6×10 8 7×10 8 8×10 8 9×10 8 1×10 9 2×10 9 3×10 9 4×10 9 5×10 9 6×10 9 7×10 9 8×10 9 9×10 9 1×10 10 2×10 10 3×10 10 4×10 10 and 5×10 10 PFU. In some implementations, the recombinant virus of this disclosure is at most about 1 × 10⁻⁶. 8 2×10 8 3×10 8 4×10 8 5×10 8 6×10 8 7×10 8 8×10 8 9×10 8 1×10 9 2×10 9 3×10 9 4×10 9 5×10 9 6×10 9 7×10 9 8×10 9 9×10 9 1×10 10 2×10 10 3×10 10 4×10 10 and 5×10 10 PFU. In some aspects, the mutated adeno-associated virus vector or chimeric adeno-associated virus vector of this disclosure can be used as a vector for genome measurement. In some embodiments, the recombinant virus of this disclosure is 1 × 10⁻⁶. 10 Up to 3×10 12 One vector genome, or 1×109 Up to 3×10 13 One vector genome, or 1×10 8 Up to 3×10 14 One vector genome, or at least about 1 × 10⁻⁶ 1 1×10 2 1×10 3 1×10 4 1×10 5 1×10 6 1×10 7 1×10 8 1×10 9 1×10 10 1×10 11 1×10 12 1×10 13 1×10 14 1×10 15 1×10 16 1×10 17 and 1×10 18 One vector genome, or 1×10 8 Up to 3×10 14 One vector genome, or at most about 1×10 1 1×10 2 1×10 3 1×10 4 1×10 5 1×10 6 1×10 7 1×10 8 1×10 9 1×10 10 1×10 11 1×10 12 1×10 13 1×10 14 1×10 15 1×10 16 1×10 17 and 1×10 18 One vector genome.

[0134] In some embodiments, the multiple of infection (MOI) can be used to measure the mutated adeno-associated virus vector or chimeric adeno-associated virus vector of this disclosure. In some embodiments, MOI can refer to the ratio or fold increase in the number of cells to which the vector or viral genome and nucleic acid can be delivered. In some embodiments, MOI can be 1 × 10⁻⁶. 6GC / mL. In some implementations, the MOI can be 1×10⁻⁶. 5 GC / mL to 1×10 7 GC / mL. In some implementations, the MOI can be 1×10⁻⁶. 4 GC / mL to 1×10 8 GC / mL. In some embodiments, the recombinant virus of this disclosure is at least about 1 × 10⁻⁶. 1 GC / mL, 1×10 2 GC / mL, 1×10 3 GC / mL, 1×10 4 GC / mL, 1×10 5 GC / mL, 1×10 6 GC / mL, 1×10 7 GC / mL, 1×10 8 GC / mL, 1×10 9 GC / mL, 1×10 10 GC / mL, 1×10 11 GC / mL, 1×10 12 GC / mL, 1×10 13 GC / mL, 1×10 14 GC / mL, 1×10 15 GC / mL, 1×10 16 GC / mL, 1×10 17 GC / mL and 1×10 18 GC / mL MOI. In some embodiments, the mutated adeno-associated virus or chimeric adeno-associated virus of this disclosure is approximately 1 × 10⁻⁶. 8 GC / mL to approximately 3 × 10 14 GC / mL, or at most about 1×10 1 GC / mL, 1×10 2 GC / mL, 1×10 3 GC / mL, 1×10 4 GC / mL, 1×10 5 GC / mL, 1×10 6 GC / mL, 1×10 7 GC / mL, 1×10 8 GC / mL, 1×10 9 GC / mL, 1×10 10 GC / mL, 1×10 11 GC / mL, 1×10 12 GC / mL, 1×10 13 GC / mL, 1×1014 GC / mL, 1×10 15 GC / mL, 1×10 16 GC / mL, 1×10 17 GC / mL and 1×10 18 GC / mL MOI.

[0135] In some respects, non-viral vectors or nucleic acids can be delivered without using mutated adeno-associated virus vectors or chimeric adeno-associated virus vectors, and can be measured according to the amount of nucleic acid. Generally, any suitable amount of nucleic acid can be used with the compositions and methods of this disclosure. In some implementation schemes, the nucleic acid can be at least about 1 pg, 10 pg, 100 pg, 1 pg, 10 pg, 100 pg, 200 pg, 300 pg, 400 pg, 500 pg, 600 pg, 700 pg, 800 pg, 900 pg, 1 μg, 10 μg, 100 μg, 200 μg, 300 μg, 400 μg, 500 μg, 600 μg, 700 μg, 800 μg, 900 μg, 1 ng, 10 ng, 100 ng, 200 ng, 300 ng, 400 ng, 500 ng, 600 ng, 700 ng, 800 ng, 900 ng, 1 mg, 10 mg, 100 mg, 200 mg, 300 mg, 400 mg, 500 mg, etc. mg, 600 mg, 700 mg, 800 mg, 900 mg, 1g, 2g, 3g, 4g or 5g. In some implementation schemes, nucleic acids can be up to about 1 pg, 10 pg, 100 pg, 1 pg, 10 pg, 100 pg, 200 pg, 300 pg, 400 pg, 500 pg, 600 pg, 700 pg, 800 pg, 900 pg, 1 μg, 10 μg, 100 μg, 200 μg, 300 μg, 400 μg, 500 μg, 600 μg, 700 μg, 800 μg, 900 μg, 1 ng, 10 ng, 100 ng, 200 ng, 300 ng, 400 ng, 500 ng, 600 ng, 700 ng, 800 ng, 900 ng, 1 mg, 10 mg, 100 mg, 200 mg, 300 mg, 400 mg, 500 mg, etc. mg, 600 mg, 700 mg, 800 mg, 900 mg, 1 g, 2 g, 3 g, 4 g or 5 g.

[0136] The proteins, vectors, plasmids, compositions, systems, engineered proteins, and guide polynucleotides provided herein can be delivered by any suitable method, including transfection, electroporation, liposome delivery, membrane fusion technology, high-speed DNA-coated particles, viral infection, and protoplast fusion. Methods for constructing any embodiment of the compositions provided herein include genetic engineering, recombinant engineering, and synthetic techniques.

[0137] Engineered proteins, guide polynucleotides, or polynucleotides encoding compositions provided herein can be delivered to cells via electroporation. Electroporation using systems such as the NEON® transfection system (ThermoFisher Scientific) or LonzaNucleofactor technologies can also be used to deliver nucleic acids and proteins into cells. For example, engineered proteins provided herein can be purified and complexed with suitable guide polynucleotides for delivery into cells. Electroporation parameters can be tuned to optimize delivery efficiency and / or cell viability. Electroporation devices can have pulse settings in various waveforms, such as exponential decay, time constant, and square wave. Each cell type has a unique optimal field strength (E) that depends on the applied pulse parameters (such as voltage, capacitance, and resistance). The application of the optimal field strength induces permeabilization by generating a transmembrane voltage, which allows nucleic acids to cross the cell membrane. In some embodiments, the electroporation pulse voltage, pulse width, number of pulses, cell density, and tip type can be tuned to optimize transfection efficiency and / or cell viability.

[0138] (3) Cellular and cell-free systems This document provides cells comprising the systems, guide polynucleotides, compositions, or engineered proteins provided herein. In some embodiments, polynucleotides encoding engineered proteins provided herein and polynucleotides encoding guide polynucleotides provided herein are administered to cells or cell populations.

[0139] The compositions, polynucleotides, engineered proteins, guide polynucleotides, and systems provided herein can be delivered to any suitable cell. In some embodiments, the compositions, polynucleotides, engineered proteins, guide polynucleotides, and systems provided herein regulate the genes, proteins, and / or functional phenotypes of a cell.

[0140] Suitable cells may include, but are not limited to, eukaryotic and prokaryotic cells and / or cell lines. Suitable cells may be, for example, primary human cells. Primary cells can be taken directly from living tissue (i.e., biopsy material) and used to establish cells for in vitro growth. ,Compared to continuous tumorigenesis cell lines or artificially immortalized cell lines, these primary cells undergo very little population doubling and are therefore more representative of the major functional components and characteristics of the tissues from which they originate. Primary cells can be obtained from a variety of sources, such as organs, vascular systems, erythrocyte sedimentation rate (ESR) amber layer, whole blood, apheresis components, plasma, bone marrow, tumors, cell banks, cryopreservation banks, or blood samples. Primary cells can be stem cells.

[0141] Suitable cells that can interact with the compositions, polynucleotides, engineered proteins, or guiding polynucleotides or systems provided herein include, but are not limited to: epithelial cells, fibroblasts, nerve cells, keratinocytes, hematopoietic cells, melanocytes, chondrocytes, leukocytes, lymphocytes (B, NK, and T), macrophages, monocytes, mononuclear cells, cardiomyocytes, other muscle cells, granulosa cells, cumulus cells, epidermal cells, endothelial cells, pancreatic islet cells, blood cells, blood progenitor cells, bone cells, bone progenitor cells, neuronal stem cells, primitive stem cells, hepatocytes, keratinocytes, umbilical vein endothelial cells, aortic endothelial cells, microvascular endothelial cells, fibroblasts, hepatic stellate cells, aortic smooth muscle cells, cardiomyocytes, neurons, Kupffer cells, smooth muscle cells, and Schwann cells. Cells and epithelial cells, erythrocytes, platelets, neutrophils, lymphocytes, monocytes, eosinophils, basophils, adipocytes, chondrocytes, pancreatic islet cells, thyroid cells, parathyroid cells, parotid gland cells, tumor cells, glial cells, astrocytes, red blood cells, white blood cells, macrophages, epithelial cells, somatic cells, pituitary cells, adrenal cells, hair cells, bladder cells, kidney cells, retinal cells, rod cells, cone cells, heart cells, pacemaker cells, spleen cells, antigen-presenting cells, memory cells, T cells, B cells, plasma cells, muscle cells, ovarian cells, uterine cells, prostate cells, vaginal epithelial cells, sperm cells, testicular cells, germ cells, oocytes, Leydig cells, peritubular cells, Sertoli cells Cells, including luteal cells, cervical cells, endometrial cells, mammary cells, follicular cells, mucous cells, ciliated cells, non-keratinized epithelial cells, keratinized epithelial cells, lung cells, goblet cells, columnar epithelial cells, dopaminergic cells, squamous epithelial cells, osteocytes, osteoblasts, osteoclasts, dopaminergic cells, embryonic stem cells, fibroblasts, and fetal fibroblasts. Furthermore, one or more types of cells may be, for example, pancreatic islet cells and / or cell clusters, including but not limited to pancreatic α cells, pancreatic β cells, pancreatic δ cells, pancreatic F cells (also known as pancreatic polypeptide cells or PP cells), or pancreatic ε cells.

[0142] Suitable cells also include stem cells, such as embryonic stem cells, induced pluripotent stem cells, hematopoietic stem cells, neuronal stem cells, and mesenchymal stem cells. Suitable cells can include any number of primary cells, such as human cells, non-human cells, and / or mouse cells. Suitable cells can be progenitor cells. Suitable cells can be derived from the subject to be treated. For example, the subject to be treated can be a subject with a disease, a subject requiring treatment, or a subject with compromised immunity. Suitable cells can be derived from human donors.

[0143] In some embodiments, cells are genetically modified using the methods, systems, and compositions provided herein. The cells provided herein can be administered to subjects with appropriate needs, such as those suffering from diseases or conditions requiring treatment.

[0144] Methods for obtaining suitable cells (such as primary human cells) may include cell selection. In some embodiments, the cells may contain markers that allow for selection of the cells. For example, such markers may include GFP, resistance genes (e.g., genes conferring antibiotic resistance), cell surface markers, or endogenous tags. Cell selection can be performed using any endogenous marker. Suitable cells can be selected using any technique. Such techniques may include flow cytometry and / or magnetic column chromatography. The selected cells may also be expanded to a large scale.

[0145] Delivery media for in vivo and in vitro use may include pharmaceutically acceptable carriers. Pharmaceutically acceptable carriers are determined in part by the specific composition being administered (e.g., polynucleotide, protein, carrier, or cell) and by the specific method used to administer the composition.

[0146] This document provides cell-free systems comprising the systems, compositions, and guide polynucleotides or engineered proteins provided herein. Cell-free systems contain components sufficient to carry out synthetic reactions and, in some cases, retain the biological activity of DNA ligases and nickases when stored at room temperature for a period of time. In some embodiments, the cell-free system comprises a set of reagents capable of providing or supporting a biosynthetic reaction. Non-limiting examples of biosynthetic reactions include combinations of in vitro DNA replication, transcription, translation, or reactions in the absence of cells. Cell-free systems can be prepared using enzymes, coenzymes, and other subcellular components isolated or purified from eukaryotic or prokaryotic cells (including recombinant cells), or as extracts or fractions of such cells. Cell-free systems can be derived from a variety of sources, including but not limited to eukaryotic and prokaryotic cells, such as bacteria, including but not limited to: *Escherichia coli*, thermophilic bacteria, etc., wheat germ, rabbit reticulocytes, mouse L cells, Ehrlich ascites carcinoma cells, HeLa cells, CHO cells, and budding yeast, etc. In some embodiments, the cell-free system is lyophilized. In some embodiments, the cell-free system also comprises a scaffold.

[0147] (4) Pharmaceutical composition, administration and delivery This document provides a pharmaceutical composition comprising: the composition provided herein, a guide polynucleotide, an engineered protein or polynucleotide; and a pharmaceutically acceptable diluent, carrier or excipient. This document also provides a pharmaceutical composition comprising the carrier provided herein; and a pharmaceutically acceptable diluent, carrier or excipient.

[0148] In some embodiments, the compositions provided herein (e.g., a carrier comprising the system or composition provided herein) are combined with pharmaceutically acceptable salts, excipients, and / or carriers to form a pharmaceutical composition. The drug salts, excipients, and carriers can be selected based on the route of administration, target tissue location, and time course of drug delivery. Pharmaceutically acceptable carriers or excipients may include solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, and absorption delay agents compatible with drug administration.

[0149] In some embodiments, the pharmaceutical composition is in the form of a solid, semi-solid, liquid, or gas (aerosol). Injectable products, such as sterile injectable aqueous or oily suspensions, can be formulated using suitable dispersants or wetting agents and suspending agents according to known techniques. Sterile injectable products can also be sterile injectable solutions, suspensions, or emulsions in non-toxic, parenteral diluents or solvents. Acceptable media and solvents that can be used are water, Ringer's solution (USP), and isotonic sodium chloride solution, etc. Furthermore, sterile, non-volatile oils are routinely used as solvents or suspension media. For this purpose, any mild, non-volatile oil, including synthetic monoglycerides or diglycerides, can be used. Additionally, fatty acids such as oleic acid are used in the preparation of injectable formulations. Injectable formulations can be sterilized, for example, by filtration through a bacterial retention filter, or by incorporating a sterilizing agent in the form of a sterile solid composition, which can be dissolved or dispersed in sterile water or other sterile injectable media prior to use.

[0150] For ease of administration and uniform dosage, the compositions and systems provided herein can be formulated in dose-unit form. Dose-unit form is a physically discrete unit of the composition provided herein suitable for the subject to be treated. For any composition provided herein, the therapeutically effective dose can be initially estimated in cell culture assays or animal models such as mice, rabbits, dogs, pigs, or non-human primates. Animal models are also used to obtain desirable concentration ranges and routes of administration. Such information can then be used to determine the useful dose and route of administration for human use. The therapeutic efficacy and toxicity of the compositions provided herein can be determined using standard pharmaceutical procedures in cell culture or laboratory animals, such as ED. 50 (This dose is effective in treating 50% of the population) and LD50 50 (This dose is fatal to 50% of the population). The dose ratio of toxicity to therapeutic effect is the therapeutic index, and the therapeutic index can be expressed as the ratio LD50. 50 / ED 50 In some implementations, pharmaceutical compositions exhibiting a high therapeutic index may be useful. Data obtained from cell culture assays and animal studies can be used to formulate dosage ranges for human use.

[0151] This document provides pharmaceutical compositions for administration to a subject with a corresponding need. In some embodiments, the pharmaceutical composition is a treatment for the disease or condition described herein. In some embodiments, the pharmaceutical compositions described herein are in a form that allows administration of the compositions described herein to a subject. In some embodiments, the pharmaceutical compositions are formulated for intratumoral delivery. In some embodiments, the administration of the pharmaceutical compositions described herein is local or systemic. In some embodiments, the pharmaceutical compositions described herein are formulated for administration / for administration via intratumoral, subcutaneous, intradermal, intramuscular, inhalation, intravenous, intraperitoneal, or intracranial routes. In some embodiments, administration is performed every 1 hour, 2 hours, 4 hours, 6 hours, 8 hours, 12 hours, 24 hours, 36 hours, or 48 hours. In some implementations, application is performed for at least about 5 hours, at least about 10 hours, at least about 12 hours, at least about 15 hours, at least about 20 hours, at least about 24 hours (1 day), at least about 48 hours (2 days), at least about 72 hours (3 days), at least about 96 hours (4 days), at least about 120 hours (5 days), at least about 144 hours (6 days), at least about 168 hours (7 days), at least about 336 hours (14 days), at least about 504 hours (21 days), at least about 672 hours (28 days), or up to 744 hours (31 days). In some implementations, application is performed every 744 hours (once a month) or once a year (365 days).

[0152] The vector can be delivered in vivo by administration to an individual subject, typically via systemic administration, such as intravenous, intraperitoneal, intramuscular, subcutaneous, or intracranial infusion. As described below, the vector can be delivered via surface administration. Alternatively, the vector can be delivered ex vivo to cells, such as cells explanted from an individual subject (e.g., lymphocytes, T cells, bone marrow aspirate, or tissue biopsy), followed by re-implantation of the cells into the subject, typically after selection of vector-incorporated cells. Cells can be expanded in cell cultures or in bioreactors before or after selection.

[0153] This document provides cells expressing the systems or compositions provided herein. In some embodiments, the cells are contacted in vitro or in vitro with nucleic acids encoding the compositions or systems provided herein. In some embodiments, the cells are contacted in vitro or in vitro with a vector encoding an engineered protein provided herein, the system provided herein, or the composition provided herein.

[0154] (5) Support and System This document provides scaffolds, wherein the scaffolds contain any compositions, systems, or guide polynucleotides provided herein. The scaffolds provided herein can be used to detect newly edited nucleic acids or to identify target nucleic acids provided herein. In some embodiments, the compositions provided herein are immobilized to the scaffold. In some embodiments, the scaffold includes a surface. In some embodiments, the surface includes a solid, semi-solid, or gel surface. In some embodiments, the scaffold includes a reaction chip, paper, quartz microfibers, a mixture of cellulose esters, porous alumina, a patterned surface, a tube, a pore, or a matrix. In some embodiments, the scaffold includes a patterned surface adapted to immobilize molecules in an ordered pattern. In some embodiments, a patterned surface refers to an arrangement of different regions in or on an exposed layer of the scaffold. In some embodiments, the scaffold includes an array of pores or recesses in the surface. The composition and geometry of the scaffold can vary depending on its intended use. In some embodiments, the scaffold is a planar structure, such as a slide, chip, microchip, and / or array. Thus, the surface of the scaffold can be in the form of a planar layer. In some embodiments, the scaffold includes one or more surfaces of a flow cell. A flow cell is a chamber comprising a solid surface through which one or more fluid reagents can flow. In some embodiments, the support or its surface is non-planar, such as the inner or outer surface of a tube or container. In some embodiments, the support comprises microspheres or beads. Microspheres, beads, or particles can be made of a variety of materials, including but not limited to plastics, ceramics, glass, and polystyrene. In some embodiments, the microspheres are magnetic microspheres or beads. Optionally or additionally, the beads can be porous. Bead sizes range from nanometers (nm) (e.g., about 100 nm) to millimeters (e.g., about 1 mm).

[0155] In some embodiments, the scaffold includes an engineered proteome, the polynucleotide set provided herein, the system set provided herein, or any combination thereof. Systems further comprising a scaffold are provided herein, wherein the scaffold includes a surface. In some embodiments, the scaffold also includes the cell-free system provided herein. In some embodiments, the systems and scaffolds provided herein also include reagents for nucleic acid amplification. In some embodiments, the systems and scaffolds provided herein also include reagents for DNA replication. In some embodiments, the systems provided herein also include: (a) the scaffold provided herein; (b) a reporter molecule; and (c) a detector. In some embodiments, the reporter molecule generates a detectable signal that is detected by the detector when the target nucleic acid forms a complex with the scaffold, the guide polynucleotide provided herein, or the engineered protein provided herein. In some embodiments, the reporter molecule is selected from the group consisting of fluorophores, dyes, peptides, antibodies, nucleic acids, and any combination thereof. In some embodiments, the detectable signal is a calorimetric signal, a potentiometric signal, an amperometric signal, an optical signal, or a piezoelectric signal. The reporter molecule can be used to identify cells containing novel nucleic acids edited by the engineered protein provided herein. The reporter molecule can also be used, for example, for cell sorting, nucleic acid isolation, nucleic acid sequencing, or immunochemical techniques.

[0156] (6) Reagent kit This document provides kits comprising: systems, compositions, polynucleotides, vectors, guide polynucleotides or engineered mRNAs or proteins provided herein; and packaging and materials for use therewith. In some embodiments, the kit also comprises a scaffold provided herein. In some embodiments, the kit also comprises a cell-free system. In some embodiments, the kit also comprises a cell population. In some embodiments, the cells are stored in a cryopreservation medium. In some embodiments, the cryopreservation medium comprises dimethyl sulfoxide (DMSO). In some embodiments, the cryopreservation medium comprises a buffer, an isotonic agent, or an apoptosis inhibitor. Non-limiting examples of buffer elements include: citrate, phosphate, succinate, tartrate, fumarate, gluconate, oxalate, lactate, acetate, histidine, and tris. Non-limiting examples of isotonic agents include, for example, citrate, phosphate, succinate, tartrate, fumarate, gluconate, oxalate, lactate, acetate, histidine, and tris. Other isotonic agents include sodium chloride, potassium chloride, boric acid, sodium borate, mannitol, glycerol, propylene glycol, polyethylene glycol, maltose, sucrose, erythritol, arabinitol, xylitol, sorbitol, trehalose, and glucose. Apoptosis inhibitors may include, for example, Rho-associated kinase (ROCK) inhibitors, catalase, and zVAD-fmk. In some embodiments, the kit contains reagents. In some embodiments, the reagents include sugars and sugar derivatives. In some embodiments, the sugar derivatives include sodium carboxymethyl cellulose or cellulose acetate. In some embodiments, the reagents include detergents, glycols, polyols, esters, buffers, alginate, and / or organic solvents.

[0157] In some embodiments, the formulation of the compositions described herein is prepared in a single container for administration to cells, cell-free systems, or subjects. In some embodiments, the formulation of the compositions provided herein is prepared in two containers for administration, separating the guide polynucleotide or the polynucleotide encoding the guide polynucleotide and / or the polynucleotide encoding the engineered protein provided herein. As used herein, “container” includes a vessel, vial, ampoule, tube, cup, box, bottle, flask, wide-mouth bottle, dish, pore of a single-hole or multi-hole device, reservoir, can, etc., or other device in which the compositions disclosed herein can be placed, stored, and / or transported, and in which the contents can be removed. Examples of such containers include sealed or resealable tubes and ampoules of glass and / or plastic, including those with rubber diaphragms or other sealing devices compatible with the use of needles and syringes to remove the contents. In some embodiments, the container is RNase-free.

[0158] This document provides a kit comprising: a first container and a second container, the first container comprising: a donor nucleic acid, a guide polynucleotide, or a polynucleotide encoding a guide polynucleotide, wherein the guide polynucleotide comprises: (i) a target region complementary to the target nucleic acid; (ii) a protein-binding region comprising a secondary structure binding to a nuclease or nicking enzyme; and (iii) a ligation splice 2 region complementary to the donor nucleic acid, the target nucleic acid, and at least one mismatched nucleobase relative to the target nucleic acid; and (iv) a ligation splice 1 region comprising: DNA nucleotides and RNA nucleotides, the second container comprising: an engineered protein or a polynucleotide encoding an engineered protein, wherein the engineered protein comprises: a nicking enzyme operatively linked to a DNA ligase. In some embodiments, the kit provided herein also comprises reagents for nucleic acid amplification, transcription, translation, or nucleic acid isolation. In some embodiments, the kit provided herein also comprises a reporter molecule provided herein.

[0159] (7) Methods for determining gene editing activity This document provides methods for determining the gene-editing activity, efficiency, and selectivity of engineered proteins, compositions, or systems provided herein for target nucleic acid sequences. In some embodiments, the target nucleic acid is cleaved by a nuclease or nicking enzyme provided herein upon binding of the directed polynucleotide to the target nucleic acid sequence. In some embodiments, the target nucleic acid is DNA. In some embodiments, the nicking enzyme region of the engineered protein provided herein cleaves DNA, resulting in single-strand breaks. In some embodiments, the nicking enzyme region of the engineered protein provided herein cleaves the target DNA via interleaved DNA single-strand or double-strand breaks. The ability of the nuclease or nicking enzyme provided herein to recognize PAM sequences can be determined by in vitro selection assays.

[0160] The activity of the system provided herein can be determined using cells expressing reporter proteins or containing reporter genes. For example, a reporter gene can be engineered to contain obstructions, such as stop codons, frameshift mutations, spacer regions, adapters, or transcription terminators; the system can then be used to remove the obstructions, and the resulting functional reporter protein can be detected. Similarly, a reporter gene can be introduced into a target nucleic acid to detect the target. In some embodiments, the reporter gene can be engineered such that specific sequence modifications are required to restore the function of the reporter protein. In other embodiments, the reporter gene can be engineered such that any insertion or deletion resulting in a frameshift of one or two bases in the target nucleic acid is sufficient to restore the function of the reporter protein. Examples of reporter proteins encoded by reporter genes include colorimetric enzymes, metabolic enzymes, fluorescent proteins, antibiotic resistance-associated enzymes and transporters, and luminescent enzymes. Examples of such reporter proteins include β-galactosidase, chloramphenicol acetyltransferase, green fluorescent protein, red fluorescent protein, and firefly luciferase and kidney luciferase. Different detection methods can be used for different reporter proteins. For example, reporter proteins can affect cell viability, cell growth, fluorescence, luminescence, or the expression of detectable products. In some embodiments, reporter proteins can be detected using colorimetric assays. In some embodiments, the reporter protein can be a fluorescent protein, and DNA editing can be determined by measuring the level of fluorescence in treated cells or the number of treated cells with at least a threshold level of fluorescence. In some embodiments, the transcript levels of the reporter gene can be assessed. In other embodiments, the reporter gene can be assessed by sequencing.

[0161] The integration of donor nucleic acids ligated by DNA ligase into target nucleic acids can be measured using any technique. For example, integration can be measured by denaturing urea-polyacrylamide gel electrophoresis, PAGE gel electrophoresis, flow cytometry, Surveyor nuclease assay, trace-through-deposition insertion / deletion (TIDE), ligation PCR, droplet digital PCR, or any combination thereof. In other embodiments, transgenic integration can be measured by PCR or droplet digital PCR. TIDE analysis can also be performed on engineered cells. In vitro cell transfection can also be used for diagnostics, research, or gene therapy. In some embodiments, transfected cells are re-infused into a host organism. In some embodiments, cells are isolated from a subject organism, transfected with nucleic acids, and re-infused into the subject.

[0162] The amount of genetically modified cells necessary for therapeutic efficacy in a subject can vary depending on cell viability and the efficiency of cell genetic modification. For example, the efficiency of transgene integration into one or more cells can determine the number of cells administered to the subject. In some embodiments, the product of genetically modified cell viability and transgene integration efficiency (e.g., proliferation) can correspond to a therapeutic aliquot of cells available for administration to the subject. In some embodiments, an increase in genetically modified cell viability can correspond to a reduction in the amount of cells necessary for therapeutic efficacy in the subject. In some embodiments, an increase in the efficiency of transgene integration into one or more cells can correspond to a reduction in the amount of cells necessary for therapeutic efficacy in the subject. In some embodiments, determining the amount of cells necessary for therapeutic efficacy can include determining a function corresponding to changes in cell viability over time. In some embodiments, determining the amount of cells necessary for therapeutic efficacy can include determining a function corresponding to changes in the efficiency of transgene integration into one or more cells relative to time-related variables. Variables can include, for example, cell culture time, electroporation time, and cell stimulation time.

[0163] Various methods can be used to quantify non-homologous end joining (NHEJ) and homologous directed repair (HDR). For example, the percentage of NHEJ, HDR, or a combination of both can be determined by co-delivering gene-editing molecules (such as the guide polynucleotides and engineered proteins provided herein) with donor nucleic acids encoding promoter-free tags or markers into cells. In some embodiments, the target marker is green fluorescent protein (GFP) or a polynucleotide encoding GFP. After a period of time (e.g., approximately 72 to 96 hours), flow cytometry can be performed to quantify the total cell count (N). 总 The number of tag-positive cells and tag / GFP-negative cells can be measured. In tag-negative cells, next-generation sequencing can be performed to identify cells without mutations and those with mutations. HDR efficiency and NHEJ efficiency can be calculated from the assays.

[0164] Additional assays for determining the gene editing efficiency of the systems or compositions provided herein may include, but are not limited to: RT-PCR, nucleic acid sequencing, T7 endonuclease 1 (T7E1) mismatch detection assay, trace-by-deletion (TIDE) analysis, and detection of insertions / deletions by amplicon analysis (IDAA) assay. Insertion / deletion patterns induced at target sites of programmable nucleases or nickases can also be determined by PCR amplification of the corresponding regions and subsequent next-generation sequencing. To obtain a quantitative single-cell view of gene editing efficiency, cell surface markers can be targeted, and signal loss due to insertion / deletion formation can be quantified by flow cytometry. Single-cell sequencing can be used to assess genome editing efficiency.

[0165] (8) Application This document provides methods for modifying target nucleic acids or producing alterations in target nucleic acids or genes within cells. In some embodiments, the method includes administering the system, composition, or cell provided herein to a cell, tissue, or subject, wherein the administration produces an alteration in the target nucleic acid. In some embodiments, the number of alterations in the target nucleic acid is at least one alteration. In some embodiments, the percentage of the alteration in the target nucleic acid molecule is at least 0.5%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%. In some embodiments, the method of modifying the target nucleic acid further includes mutation of the target nucleic acid sequence. In some embodiments, the alteration or modification of the target nucleic acid includes: insertion, deletion, substitution, copy number change, point mutation, frameshift mutation, missense mutation, nonsense mutation, stop codon mutation, epigenetic marker, or any combination thereof.

[0166] In some embodiments, alterations are made in the target nucleic acid to form edited nucleic acids. In some embodiments, the edited nucleic acids restore the expression of wild-type proteins encoded by genes, relative to comparable cells or cell populations that have not been exposed to the system or composition provided herein.

[0167] This document provides a method for modifying cells in vitro. In some embodiments, the method includes contacting cells with a composition provided herein under conditions that allow nucleases or nicking enzymes to cleave target nucleic acid molecules, thereby modifying the cells. This document also provides a method for modifying cells in vitro. In some embodiments, the method includes contacting cells with a composition provided herein under conditions that allow DNA to be linked to a novel nucleic acid sequence for incorporation into a target nucleic acid, thereby modifying the target nucleic acid and the cells.

[0168] This article also provides methods for in vitro cell modification, including: Cells are contacted with a ribonucleoprotein (RNP) complex, wherein the RNP complex comprises: (i) an engineered protein provided herein; and (ii) a polynucleotide that binds to a target nucleic acid, wherein, upon contacting the cell with the RNP complex, the engineered protein cleaves the target nucleic acid molecule and attaches a new nucleic acid, thereby modifying the cell. In some embodiments, the cells are immune cells or stem cells. In some embodiments, immune cells are leukocytes, lymphocytes, natural killer cells, dendritic cells, macrophages, myeloid cells, T cells, B cells, stem cells, induced pluripotent stem cells, cancer cells, or endothelial cells. In some embodiments, stem cells are embryonic stem cells, induced pluripotent stem cells (iPSCs), or adult stem cells.

[0169] The methods for modifying target nucleic acids provided herein can be used in agricultural applications, such as producing improved plant species. These methods can also be used in biomedical applications, drug screening platforms, and therapeutic applications. For example, the compositions and systems provided herein can be used to generate cell therapies for treating diseases or disorders.

[0170] This article provides a method for modifying nucleic acids, comprising: contacting a cell or cell-free system with: (a) a donor nucleic acid; (b) a directing polynucleotide or a polynucleotide encoding a directing polynucleotide, wherein the directing polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region comprising a secondary structure binding to a nicking enzyme; and (iii) a linker 2 region complementary to the donor nucleic acid, the target nucleic acid, and having at least one mismatched nucleobase relative to the target nucleic acid; and (iv) a linker 1 region comprising: a deoxyribonucleotide and a ribonucleotide; and (c) an engineered protein or a polynucleotide encoding an engineered protein, wherein the engineered protein comprises: (i) a nicking enzyme region; and (ii) A DNA ligase region; wherein: a guide polynucleotide to form a complex with an engineered protein via a protein-binding region; a target sequence to form a complex with the complementary strand of the target nucleic acid; a nicking enzyme region of the engineered protein to generate a leading strand in the target nucleic acid; a ligation splint 1 region to form a complex with the leading strand; and a ligation splint 2 region to form a complex with the donor nucleic acid; wherein the DNA ligase region catalyzes the formation of a phosphodiester bond between the 3' hydroxyl terminus of the leading strand and the 5' phosphate terminus of the donor nucleic acid, thereby inserting a new nucleic acid. In some embodiments, the target region dissociates from the complementary strand. In some embodiments, the new nucleic acid is incorporated into the target nucleic acid by hybridization with the complementary strand. In some embodiments, the incorporation of the donor nucleic acid recruits DNA repair proteins that edit the nucleotides of the target nucleic acid to the target nucleic acid.

[0171] This document provides a method for detecting nucleic acids in a test sample, the method comprising: (a) immobilizing a guide polynucleotide onto a scaffold provided herein; and (b) contacting the guide polynucleotide with: (i) a test sample; and (ii) an engineered protein provided herein. In some embodiments, the test sample contains a target nucleic acid capable of binding to the guide polynucleotide. In some embodiments, a complex is formed between the guide polynucleotide, the engineered protein provided herein, and the target nucleic acid. In some embodiments, after the complex is formed, the engineered protein (e.g., a nicking enzyme region) cleaves the target nucleic acid. In some embodiments, the method further includes detecting a signal indicating cleavage of the target nucleic acid molecule or detecting a new nucleic acid strand containing a reporter molecule, thereby detecting the target nucleic acid in the sample.

[0172] In some embodiments, the method further includes amplifying the target nucleic acid prior to the detection step. In some embodiments, amplification includes polymerase chain reaction (PCR), sequence-based amplification (NASBA), recombinase polymerase amplification (RPA), loop-mediated isothermal amplification (LAMP), strand displacement amplification (SDA), helicase-dependent amplification (HDA), nickase amplification reaction (NEAR), multiple displacement amplification (MDA), rolling circle amplification (RCA), modified multiple displacement amplification (IMDA), SMART (Simple Method for Amplifying RNA Targets), single primer isothermal amplification (SPIA), ligase reaction (LCR), transcription-mediated amplification (TMA), branching amplification method (RAM), or any combination thereof. In some embodiments, the method further includes endonuclease mismatch detection, immunoassay, gel electrophoresis, plasmid interference detection, nucleic acid sequencing, or any combination thereof. In some embodiments, detection includes calorimetry, potentiometry, amperometric detection, optical detection, piezoelectric detection, or any combination thereof.

[0173] This document provides methods for treating a disease or condition in a subject with an appropriate need. In some embodiments, the subject has, is suspected of having, or has been diagnosed with a disease or condition. In some embodiments, the method includes administering the system, composition, carrier, pharmaceutical composition, or polynucleotide provided herein to the subject. In some embodiments, the administration is local or systemic. In some embodiments, the administration is intranasal, subcutaneous, intravenous, inhaled, intramuscular, intratumoral, peritumoral, intrathecal, vaginal, or intradermal. In some embodiments, the method further includes administering a therapeutic agent to the subject.

[0174] Use of the compositions and systems provided herein in the preparation of medicaments is provided. Use of the compositions described herein in the preparation of medicaments for the therapeutic and / or prophylactic treatment of the diseases or conditions described herein is also provided. In some embodiments, the disease or condition is caused by a mutated disease-related gene. In some embodiments, the disease-related gene is any gene associated with an increased risk of having or developing a disease. In some embodiments, the disease-related gene is any gene or polynucleotide that produces a transcriptional or translational product at an abnormal level or in an abnormal form in cells derived from disease-affected tissues or cells, compared to non-disease control tissues or cells. In some embodiments, the disease-related gene is a gene that becomes expressed at an abnormally high level. In some embodiments, the disease-related gene is a gene that becomes expressed at an abnormally low level, wherein the altered expression is associated with the occurrence and / or progression of the disease. In some embodiments, the disease-related gene is a gene having one or more mutations or genetic variations in linkage disequilibrium with one or more genes responsible for the etiology of the disease or with the etiology of the disease. The transcriptional or translational product may be known or unknown and may be at normal or abnormal levels.

[0175] In some embodiments, the disease or condition is caused by mutations associated with DNA repeat instability and neurological disorders. Specific aspects of tandem repeat sequences have been identified as responsible for more than twenty human diseases. Systems can be used to correct these defects in genomic instability. In some embodiments, the disease or condition is a neurological disorder, cardiovascular disease, cancer, respiratory disease, diabetes, obesity, eye disease, hearing loss, blindness, or a rare genetic disease or disorder. In some embodiments, the disease or condition is age-related macular degeneration, schizophrenia disorder, trinucleotide repeat disorder, or Fragile X syndrome. In some embodiments, the disease or condition is a secretase-related disorder. In some embodiments, the disease or condition is a prion-related condition. In some embodiments, the disease or condition is ALS. In some embodiments, the disease or condition is drug addiction. In some embodiments, the disease or condition is autism. In some embodiments, the disease or condition is Alzheimer's disease. In some embodiments, the disease or condition is inflammation. In some embodiments, the disease or condition is Parkinson's disease. Other examples of diseases and conditions that can be treated with the systems and compositions provided herein include, but are not limited to: Aieardi-Goutieres syndrome; Alexander disease; Allan-Herndon-Dudley syndrome; POLG-related disorders; α-mannosin storage disorders (types II and III); Alstrom syndrome; Angelman syndrome; ataxia-telangiectasia; neuronal ceroid lipofuscin deposition; β-thalassemia; bilateral optic atrophy and (infant) optic atrophy type 1; retinoblastoma (bilateral); Canavan disease; Cerebrooculofacial Syndrome 1 (COFS1); Cerebrotcndinous Xanthomatosis; Cornelia de... Lange syndrome; MAPT-related disorders; hereditary prion diseases; Dravet syndrome; early-onset familial Alzheimer's disease; Friedreich's ataxia (FRDA); Fryns syndrome; fucoside storage disease; Fukuyama congenital muscular dystrophy; galactosylsialic acid storage disease; Gaucher disease; organic acidemia; hemophagocytic lymphohistiocytosis; progeria syndrome; mucolipidemia II; infantile free sialic acid storage disease; PLA2G6-related neurodegeneration;Jervell and Lange-Nielsen syndrome; junctional epidermolysis bullosa; Huntington's disease; Krabbe disease (infant); mitochondrial DNA-related Leigh syndrome and NARP; Lesch-Nyhan syndrome; LIS1-related anegyriosis; Lowe syndrome; maple syrup diabetes; MECP2 duplication syndrome; ATP7A-related copper transport disorder; LAMA2-related muscular dystrophy; arylsulfatase A deficiency; mucopolysaccharidosis type I, II, or III; peroxisome biogenesis disorder, Zellweger syndrome spectrum; neurodegeneration with brain iron accumulation disorder; acid sphingomyelinase deficiency; Niemann-Pick disease type C; glycine encephalopathy; ARX-related disorders; urea cycle disorders; COL Osteogenesis imperfecta associated with 1A1 / 2; Mitochondrial DNA deletion syndrome; PLP1-related disorders; Perry syndrome; Phelan-McDermid syndrome; Glycogen storage disease type 11 (Pompe disease) (infant); MAPT-related disorders; MECP2-related disorders; Rhizomelic Chondrodysplasia Punctata type 1; Roberts syndrome; Sandhof disease; Schindler's disease type 1; Adenosine deaminase deficiency; Smith-Lemli-Opitz syndrome; Spinal muscular atrophy; Spinocerebellar ataxia with infantile onset; Hexosamine A deficiency; Lethal osteodystrophy type 1; Type VI collagen-related disorders; Usher syndrome type 1; Congenital muscular dystrophy; Wolf-Hirschhorn syndrome; Lysosomal acid lipase deficiency; and xeroderma pigmentosum. In some implementations, the subjects are mammals. In some implementations, the subjects are humans.

[0176] Exemplary Implementation This document provides compositions comprising: (a) a donor nucleic acid; (b) a polynucleotide encoding a protein construct comprising a nicking enzyme or a variant thereof and a DNA ligase or a functional fragment thereof; and (c) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region comprising a secondary structure binding to an engineered protein construct comprising the nicking enzyme region; (iii) a ligase splint 2 region, wherein the ligase splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligase splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligase splint 1 region, wherein the ligase splint 1 region comprises: a deoxyribonucleotide and a ribonucleotide, wherein the ligase splint 1 region is complementary to the target nucleic acid. This document provides compositions wherein the ligase splint 1 region forms a DNA-RNA (DR) loop upon association with the target nucleic acid and the DNA ligase. This document provides compositions wherein the ligase splint 1 region comprises an RNA to DNA ratio of 1:1 up to 20:1. This document provides compositions wherein region 1 of the connecting clip contains the following ratios of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 2:1, 2:3, 2:5, 2:7, 2:9, 2:1 1, 2:13, 2:15, 2:17, 2:19, 3:1, 3:2, 3:4, 3:5, 3:7, 3:8, 3:10, 3:11, 3:13, 3:14, 3:15, 3:16, 3:17, 3:19, 4:1, 4:3, 4:5, 4:7, 4:9, 4:11, 4:13, 4:15, 4:17, 4:19, 5:1, 5:2, 5:3, 5:4, 5:6, 5:7, 5:8, 5: 9, 5:11, 5:12, 5:13, 5:14, 5:16, 6:1, 6:5, 6:7, 6:9, 6:11, 6:13, 6:15, 7:1, 7:2, 7:3, 7:4, 7:5, 7:6, 7:8, 7:9, 7:10, 7:11, 7:12, 7:13, 7:15, 8:1, 8:3, 8:5, 8:7, 8:9, 8:11, 8:13, 8:15, 9:1, 9:2, 9:4 9:5, 9:7, 9:8, 9:10, 9:11, 9:13, 9:15, 9:17, 9:19, 9:20, 10:1, 10:3, 10:7, 10:9, 10:11, 10:13, 10:15, 10:17, 10:19, 11:1, 11:2, 11:3, 11:4, 11:5, 11:6, 11:7, 11:8, 11:9, 11:10, 11:12, 11:1311:15、12:1、12:5、12:7、12:9、12:11、12:13、13:1、13:2、13:3、13:4、13:5、13:6、13:7、13:8、13:9、13:10、13:11、13:12、13:14、14:1、14:3、14:5、14:9、14:11、14:13、15:1、15:2、15:4、15:6、15:8、15:11、15:13、16:1、16:3、16:5、16:7、16:9、16:11、16:13、16:15、17:1、17:2、17:3、17:4、17:5、17:6、17:7、17:8、17:9、17:10、17:11、17:12、17:13、17:14、17:15、17:16、18:1、18:5、18:7、18:11、18:13、18:17、19:1、19:2、19:3、19:4、19:5、19:6、19:7、19:8、19:9、19:10、19:11、19:12、19:13、19:14、19:15、19:16、19:17、19:18 or 20:1. This document provides compositions wherein the connecting splint 2 region contains at least about 5 nucleotides and up to 10,000 nucleotides. This document provides compositions wherein the connecting splint 2 region contains at least about 7 nucleotides and up to 1,000 nucleotides. This document provides compositions wherein the connecting splint 1 region contains at least about 5 nucleotides and up to 20 nucleotides. This document provides compositions wherein the connecting splint 2 region contains an inverse complementary sequence of a non-coding polynucleotide sequence or a variant thereof. This document provides compositions wherein the connecting splint 2 region contains an inverse complementary sequence of a coding region encoding a polynucleotide sequence or a variant thereof. This document provides compositions wherein the connecting splint 2 region contains an inverse complementary sequence of a sequence encoding an exon or intron. This document provides compositions wherein the connecting splint 2 region contains a complementary sequence of a sequence encoding a non-coding polynucleotide sequence or a variant thereof. This document provides compositions wherein the connecting splint 2 region contains a complementary sequence of a coding region encoding a polynucleotide sequence or a variant thereof. This document provides compositions wherein the connecting splint 2 region contains a complementary sequence of a sequence encoding an exon or intron. This document provides compositions wherein the connecting splint region 2 comprises a sequence containing at least one nucleobase complementary to or mismatched with a sequence encoding a splice acceptor site. This document provides compositions wherein the nickase is an engineered Cas protein or a functional variant thereof. This document provides compositions wherein the engineered Cas protein or a functional variant thereof comprises at least one amino acid substitution corresponding to position 840 of SEQ ID NO: 69. This document provides compositions wherein the engineered Cas protein or a functional variant thereof further comprises at least one amino acid substitution corresponding to positions 221, 394, or a combination thereof of SEQ ID NO: 69. This document provides compositions wherein the engineered Cas protein or a functional variant thereof comprises amino acid substitutions of R221K, N394K, H840A, or any combination thereof. This document provides compositions wherein the engineered Cas protein or a functional variant thereof comprises a sequence that is 97%, 98%, 99%, or 100% identical to SEQ ID NO: 70 or SEQ ID NO: 71. This document provides compositions wherein the engineered Cas protein or a functional variant thereof comprises the sequence of SEQ ID NO: 70 or SEQ ID NO: 71. This document provides compositions in which the DNA ligase or functional fragment comprises *E. coli* DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, *Chlorella virus* DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human ligase V, or variants or combinations thereof.T4 DNA ligase or human DNA ligase IV. This document provides compositions wherein the DNA ligase or a functional fragment thereof comprises 90%, 91%, 92%, 93%, 94%, 95%, 97%, 98%, 99%, or 100% of the sequence SEQ ID NO: 58-66. This document provides compositions wherein the DNA ligase or a functional fragment thereof comprises the sequence of SEQ ID NO: 58-66. This document provides compositions wherein the protein construct further comprises an adapter, a nuclear localization sequence (NLS), or a combination thereof. This document provides compositions wherein the DNA ligase or a functional fragment thereof is ligated with a nicking enzyme or a variant thereof. This document provides compositions wherein the protein construct comprises 90%, 91%, 92%, 93%, 94%, 95%, 97%, 98%, 99%, or 100% of the sequence SEQ ID NO: 89-101.

[0177] This document provides engineered fusion proteins comprising: (a) a DNA ligase or a functional fragment thereof; and (b) an engineered nickase comprising three amino acid substitutions at positions 221, 394, and 840 of a nuclease containing the sequence of SEQ ID NO: 69. This document provides engineered fusion proteins wherein the substitutions result in enhanced nickase activity compared to other engineered nickases that are equivalent in other respects. This document provides engineered fusion proteins wherein the engineered nickase comprises amino acid substitutions of R221K, N394K, and H840A. This document provides engineered fusion proteins wherein the engineered nickase or a functional variant thereof comprises a sequence that is 97%, 98%, 99%, or 100% identical to SEQ ID NO: 71. This document provides engineered fusion proteins wherein the engineered nickase or a functional variant thereof comprises the sequence of SEQ ID NO: 71. This document provides engineered fusion proteins in which the DNA ligase or functional fragment includes *E. coli* DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, *Chlorella virus* DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human ligase V, or variants or combinations thereof. This document provides engineered fusion proteins in which the DNA ligase or functional fragment thereof includes *Chlorella virus* DNA ligase, T4 DNA ligase, or human DNA ligase IV. This document provides engineered fusion proteins in which the DNA ligase or functional fragment thereof contains a sequence that is 97%, 98%, 99%, or 100% identical to SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 61, or SEQ ID NO: 65. This document provides engineered fusion proteins in which the DNA ligase or functional fragment thereof contains a sequence of SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 61, or SEQ ID NO: 65. This article provides engineered fusion proteins, wherein the engineered fusion proteins further comprise adapters, nuclear localization sequences (NLS), or other protein constructs. This article provides engineered fusion proteins, wherein the engineered fusion proteins comprise 97%, 98%, 99%, or 100% identical amino acid sequences to any one of SEQ ID NO: 89-101. This article provides engineered fusion proteins, wherein the engineered fusion proteins comprise the amino acid sequences of any one of SEQ ID NO: 89-101.This article provides engineered fusion proteins, in which additional protein constructs include cell-targeting moieties, receptor-targeting moieties, regulatory elements, nucleases, acetyltransferases, ATPases, Argonaute proteins, base editors, Cas peptides, catalytically inactivated Cas peptides, deacetylases, deaminases, uncapping proteins, endonucleases, exonucleases, helicases, ligases, megnucleases, methyltransferases, nickases, polymerases, proteases, recombinases, restriction enzymes, ribonucleoproteins (RNPs), self-cleaving protein sequences, splicing factors, transcription activators, transcription activator-like effector nucleases (TALENs), transcription repressors, transposases, zinc fingers, or any combination thereof.

[0178] This article provides nucleic acids, wherein the nucleic acids encode engineered fusion proteins as provided herein. This article provides nucleic acids, wherein the nucleic acids encoding engineered fusion proteins include RNA. This article provides nucleic acids, wherein the nucleic acids encoding engineered fusion proteins include DNA.

[0179] This article provides vectors containing the nucleic acids provided herein. This article provides vectors, including viral vectors. This article provides vectors, including lentiviral vectors, retroviral vectors, adeno-associated virus (AAV) vectors, adenovirus vectors, herpes simplex virus vectors, alphavirus vectors, flavivirus vectors, rhabdovirus vectors, measles virus vectors, Newcastle disease virus vectors, poxvirus vectors, microRNA viral vectors, or oncolytic virus vectors.

[0180] This article provides a system for modifying target nucleic acids, the system comprising: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region contains a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region contains at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid. This article provides a system for modifying target nucleic acids, the system comprising: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region comprises: one or more ribonucleotides or one or more deoxyribonucleotides, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid. This article provides a system for modifying target nucleic acids, the system comprising: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region comprises: one or more deoxyribonucleotides, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.This article provides a system for modifying target nucleic acids, wherein the system comprises: (a) a donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region; (c) a DNA ligase or a functional fragment thereof; and (d) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) a ligation splint 2 region, wherein the ligation splint region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the ligation splint region comprises at least one modification relative to the target nucleic acid; and (iv) a ligation splint 1 region, wherein the ligation splint 1 region comprises: a deoxyribonucleotide and a ribonucleotide, wherein the ligation splint 1 region is complementary to the target nucleic acid, wherein, upon introduction into a cellular or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.

[0181] This article also provides a system in which the ligation splice region 1 forms a DNA-RNA (DR) loop upon association with the target nucleic acid and DNA ligase. This article also provides a system in which the ligation splice region 1 contains a ribonucleic acid (RNA) to deoxyribonucleic acid (DNA) ratio of 1:1 up to 20:1. This article also provides a system in which the ligation splice region 1 contains a ribonucleic acid (RNA) to deoxyribonucleic acid (DNA) ratio of 1:0 up to 20:0. This article also provides a system in which the connecting plate region 1 contains the following ratios of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 2:1, 2:3, 2:5, 2:7, 2:9, 2:11, 2:13, 2:15, 2:17, 2:19, 3:1, 3:2, 3:4, 3:5, 3:7, 3:8, 3 :10, 3:11, 3:13, 3:14, 3:15, 3:16, 3:17, 3:19, 4:1, 4:3, 4:5, 4:7, 4:9, 4:11, 4:13, 4:15, 4:17, 4:19, 5:1, 5:2, 5:3, 5:4, 5:6, 5:7, 5:8, 5:9, 5:11, 5:12, 5:13, 5:14, 5:16, 6:1, 6:5, 6:7, 6:9, 6:11, 6:13, 6:15, 7:1, 7:2, 7:3, 7:4, 7:5, 7:6, 7:8, 7:9, 7:10, 7:11, 7:12, 7:13, 7:15, 8:1, 8:3, 8:5, 8:7, 8:9, 8:11, 8:13, 8:15, 9:1, 9:2, 9:4, 9:5, 9:7, 9:8, 9:10, 9:11, 9:13, 9:15, 9:17, 9:19, 9:20, 10:1, 10:3, 10:7, 10:9, 10:11, 10:13, 10:15, 10:17, 10:19, 11:1, 11:2, 11:3, 11:4, 11:5, 11:6, 11:7, 11:8, 11:9 11:10, 11:12, 11:13, 11:15, 12:1, 12:5, 12:7, 12:9, 12:11, 12:13, 13:1, 13:2, 13:3, 13:4, 13:5, 13:6, 13:7, 13:8, 13:9, 13:10, 13:11, 13:12, 13:14, 14:1, 14:3, 14:5, 14:9, 14:11, 14:13, 15:1, 15:2, 15:4, 15:6, 15:8, 15:11, 15:13, 16:1, 16:3, 16:516:7, 16:9, 16:11, 16:13, 16:15, 17:1, 17:2, 17:3, 17:4, 17:5, 17:6, 17:7, 17:8, 17:9, 17:10, 17:11, 17:12, 17:13, 17:14, 17:15, 17:16, 18:1, 18:5, 1 The ratios are 8:7, 18:11, 18:13, 18:17, 19:1, 19:2, 19:3, 19:4, 19:5, 19:6, 19:7, 19:8, 19:9, 19:10, 19:11, 19:12, 19:13, 19:14, 19:15, 19:16, 19:17, 19:18, or 20:1. This article also provides a system in which, after cleavage of the target nucleic acid by the engineered protein construct, the splint 1 region hybridizes with the leading strand. This article also provides a system in which the ribonucleotides and / or deoxyribonucleotides of the splint 1 region hybridize with the leading strand. This article also provides a system in which the splint 2 region contains at least about 5 nucleotides and up to 10,000 nucleotides. This article also provides a system in which the splint 2 region contains at least about 7 nucleotides and up to 1,000 nucleotides. This article also provides a system in which the connecting splint 1 region contains at least about 5 nucleotides and up to 20 nucleotides. This article also provides a system in which the connecting splint 2 region contains the inverse complementary sequence of a non-coding polynucleotide sequence or a variant thereof. This article also provides a system in which the connecting splint 2 region contains the inverse complementary sequence of a coding region of a polynucleotide sequence or a variant thereof. This article also provides a system in which the connecting splint 2 region contains the inverse complementary sequence of a sequence encoding an exon or intron. This article also provides a system in which the connecting splint 2 region contains the complementary sequence of a sequence encoding a non-coding polynucleotide sequence or a variant thereof. This article also provides a system in which the connecting splint 2 region contains the complementary sequence of a coding region of a polynucleotide sequence or a variant thereof. This article also provides a system in which the connecting splint 2 region contains the complementary sequence of a sequence encoding an exon or intron. This article also provides a system in which the connecting splint 2 region contains a sequence containing at least one nucleobase that is complementary or mismatched with a sequence encoding a splice acceptor site. This article also provides a system in which the targeting region hybridizes with the complementary strand of the target nucleic acid, and the cleavage enzyme region cleaves the target nucleic acid. This article also provides a system in which the target region hybridizes with the complementary strand of the target nucleic acid within at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides of the PAM sequence. This article also provides a system in which the engineered protein construct contains a Cas protein or a mutant Cas protein. This article also provides a system in which the Cas protein or the mutant Cas protein is a type V Cas protein. This article also provides a system in which type V Cas proteins include: Cas12a, Cas12b, Cas12c, Cas12d, Cas12e,Cas14, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k. This article also provides systems in which Cas proteins or mutant Cas proteins include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas11, Cas12, Cas13, Cas14, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, C sb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, c2c1, c2c3, Cas9HiFi, xCas9, CasX, CasY, CasRX, SpCas9-VQR, SpCas9-VRQR, SpCas9-VRER, SaCas9-KKH, SpCas9-NG, SpCas9-NRRH, SpCas9-NRTH, SpCas9-NRCH, iSpyMac, St1Cas9 LMD9-LMG18311, St1Cas9LMD9-CNRZ1066, St1Cas9-KQKL, their variants, or any combination thereof. Systems in which the Cas protein or mutant Cas protein is a type II Cas protein are also provided herein. This article also provides systems in which type II Cas proteins include Cas9, Cas1, Cas2, or Csn2. This article also provides systems in which engineered protein constructs cleave the target nucleic acid 5' upstream of the PAM sequence. This article also provides systems in which engineered protein constructs cleave the target nucleic acid to produce two single strands of DNA or two single strands of an RNA duplex. This article also provides systems in which engineered protein constructs generate single-strand breaks in the target nucleic acid. This article also provides systems in which DNA ligases or functional fragments thereof include *E. coli* DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, *Chlorella virus* DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human DNA ligase V, or variants or combinations thereof. This article also provides systems in which DNA ligases or functional fragments thereof include *Archaeoptera scintillans* (… Archaeoglobus fulgidus DNA ligase, Archaeocarpus scintillans DNA ligase, and Coccidia flavovirens ( Pyrococcus furiosus DNA ligase, thermoautotrophic methanotherapeutic bacteria ( Methanothermobacter thermautotrophicus DNA ligase, deep-sea fireball bacteria ( Pyrococcus abyssi DNA ligase, sulfur-bearing leaf fungus ( Sulfolobus solfataricus This document describes a system in which the ligase hybridizes with the donor nucleic acid in region 2 of the ligase. It also describes a system in which the DNA ligase does not bind to the target nucleic acid. Furthermore, it describes a system in which the DNA ligase attaches the donor nucleic acid to the leading strand. It also describes a system in which the novel DNA strand has at least 50% to 99.99% complementarity with the target nucleic acid sequence. It also describes a system in which the novel DNA strand contains nucleobases mismatched relative to the complementary strand of the target nucleic acid. It also describes a system in which the secondary structure of the protein-binding region guiding the polynucleotide includes: protrusions, stems, loops, hairpins, wobbling base pairs, pseudojunctions, or combinations thereof. It also describes a system in which the target nucleic acid comprises DNA. It also describes a system in which the target nucleic acid comprises RNA. It also describes a system in which the system further includes additional protein constructs, adapters, or nuclear localization sequences (NLS). Finally, it describes a system in which the protein construct and the DNA ligase are linked together via a polypeptide adapter.

[0182] This document provides compositions comprising: the system provided herein, the engineered protein construct provided herein, the DNA ligase provided herein or a functional fragment thereof, or the guide polynucleotide provided herein.

[0183] This document provides compositions comprising: the system provided herein; and a delivery medium.

[0184] This article provides polynucleotides, wherein the polynucleotides encode the systems, engineered proteins, or guide polynucleotides provided herein.

[0185] This article provides a polynucleotide genome, wherein the polynucleotide genome encodes the systems, engineered proteins, or guide polynucleotides provided herein.

[0186] This article provides nanoparticles comprising: polynucleotides, polynucleotide sequences, systems, compositions, cells, carriers, or any portion thereof, as provided herein. This article also provides nanoparticles that are lipid nanoparticles.

[0187] This article provides vectors containing the polynucleotides or polynucleotide sequences provided herein.

[0188] This article provides cells, wherein the cells comprise: the systems provided herein, the engineered proteins provided herein, the compositions provided herein, the vectors provided herein, or the guide polynucleotides provided herein.

[0189] This document provides compositions comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises, in 5' to 3' order: (a) a targeting region complementary to the target nucleic acid; (b) a protein-binding region, wherein the protein-binding region comprises a protein-binding secondary structure; (c) a linker 2 region, wherein the linker 2 region is complementary to both the donor nucleic acid and the target nucleic acid and has at least one alteration relative to the target nucleic acid or at least one nucleic acid strand thereof; and (d) a linker 1 region. This document provides compositions comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises, in a 5' to 3' order: (a) a targeting region complementary to the target nucleic acid; (b) a protein-binding region comprising a protein-binding secondary structure; (c) a linker 2 region, wherein the linker 2 region is complementary to both the donor and target nucleic acids and has at least one alteration relative to the target nucleic acid or at least one of its nucleic acid strands; and (d) a linker 1 region, wherein the linker 1 region comprises one or more ribonucleotides or one or more deoxyribonucleotides. This document provides compositions comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises, in a 5' to 3' order: (a) a targeting region complementary to the target nucleic acid; (b) a protein-binding region comprising a protein-binding secondary structure; (c) a linker 2 region, wherein the linker 2 region is complementary to both the donor nucleic acid and the target nucleic acid and has at least one alteration relative to the target nucleic acid or at least one of its nucleic acid strands; and (d) a linker 1 region, wherein the linker 1 region comprises one or more deoxyribonucleotides. This document provides compositions comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises, in a 5' to 3' sequence: (a) a targeting region complementary to the target nucleic acid; (b) a protein-binding region comprising a protein-binding secondary structure; (c) a linker 2 region, wherein the linker 2 region is complementary to both the donor nucleic acid and the target nucleic acid and has at least one alteration relative to the target nucleic acid or at least one of its nucleic acid strands; and (d) a linker 1 region, wherein the linker 1 region comprises a deoxyribonucleotide and a ribonucleotide. This document also provides compositions wherein the target nucleic acid comprises a double-stranded DNA (dsDNA) or single-stranded RNA (ssRNA) duplex cleaved by a nicking enzyme. This document also provides compositions wherein the protein-binding region binds to a nicking enzyme or an endonuclease. This document also provides compositions wherein the secondary structure includes protrusions, stems, loops, hairpins, wobbling base pairs,False knots or combinations thereof. This document also provides compositions wherein the connecting splint 2 region comprises at least about 5 nucleotides and up to 10,000 nucleotides. This document also provides compositions wherein the connecting splint 2 region comprises at least about 7 nucleotides and up to 1,000 nucleotides. This document also provides compositions wherein the connecting splint 2 region comprises at least one change of a mismatched nucleotide. This document also provides compositions wherein mismatched nucleotides include A / C mismatches, A / T mismatches, A / G mismatches, and T / C mismatches, T / G mismatches, T / A mismatches, C / G mismatches, C / A mismatches, C / T mismatches, G / C mismatches, G / T mismatches, G / A mismatches, or combinations thereof relative to the target nucleic acid. This document also provides compositions wherein the donor nucleic acid comprises the inverse complementary sequence of a non-coding region of a multinucleotide or a variant thereof. This document also provides compositions wherein the donor nucleic acid comprises the inverse complementary sequence of a coding region encoding a multinucleotide sequence or a variant thereof. This document also provides compositions in which the donor nucleic acid comprises an inverse complementary sequence encoding an exon or intron. This document also provides compositions in which the donor nucleic acid comprises a sequence containing at least one nucleobase complementary to or mismatched with a sequence encoding a splice acceptor site. This document also provides compositions in which the donor nucleic acid comprises a complementary sequence encoding an intron or a variant thereof. This document also provides compositions in which the donor nucleic acid comprises a complementary sequence encoding an exon or a variant thereof. This document also provides compositions in which the donor nucleic acid comprises complementary sequences encoding both exons and introns. This document also provides compositions in which the ligase clip 1 region forms a DNA-RNA (DR) loop upon association with the target nucleic acid and DNA ligase. This document also provides compositions in which the ligase clip 1 region comprises an RNA to DNA ratio of 1:1 up to 20:1. This document also provides compositions in which the ligase clip 1 region comprises an RNA to DNA ratio of 1:0 up to 20:0. This document also provides compositions wherein region 1 of the connecting clip comprises the following ratios of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 2:1, 2:3, 2:5, 2:7, 2:9, 2:11, 2:13, 2:15, 2:17. 2:19, 3:1, 3:2, 3:4, 3:5, 3:7, 3:8, 3:10, 3:11, 3:13, 3:14, 3:15, 3:16, 3:17, 3:19, 4:1, 4:3, 4:5, 4:7, 4:9, 4:11, 4:13, 4:15, 4:17, 4:19, 5:1, 5:2, 5:3, 5:4, 5:6, 5:7, 5:8, 5:9, 5:11, 5:12, 5:13, 5:14, 5:16, 6:1, 6:56:7、6:9、6:11、6:13、6:15、7:1、7:2、7:3、7:4、7:5、7:6、7:8、7:9、7:10、7:11、7:12、7:13、7:15、8:1、8:3、8:5、8:7、8:9、8:11、8:13、8:15、9:1、9:2、9:4、9:5、9:7、9:8、9:10、9:11、9:13、9:15、9:17、9:19、9:20、10:1、10:3、10:7、10:9、10:11、10:13、10:15、10:17、10:19、11:1、11:2、11:3、11:4、11:5、11:6、11:7、11:8、11:9、11:10、11:12、11:13、11:15、12:1、12:5、12:7、12:9、12:11、12:13、13:1、13:2、13:3、13:4、13:5、13:6、13:7、13:8、13:9、13:10、13:11、13:12、13:14、14:1、14:3、14:5、14:9、14:11、14:13、15:1、15:2、15:4、15:6、15:8、15:11、15:13、16:1、16:3、16:5、16:7、16:9、16:11、16:13、16:15、17:1、17:2、17:3、17:4、17:5、17:6、17:7、17:8、17:9、17:10、17:11、17:12、17:13、17:14、17:15、17:16、18:1、18:5、18:7、18:11、18:13、18:17、19:1、19:2、19:3、19:4、19:5、19:6、19:7、19:8、19:9、19:10、19:11、19:12、19:13、19:14、19:15、19:16、19:17、19:18 or 20:1. This document also provides compositions in which the splice region 1 hybridizes with the target nucleic acid after cleavage by a nicking enzyme, wherein the nicking enzyme produces a leading strand and a complementary strand. This document also provides compositions in which ribonucleotides hybridize with the leading strand. This document also provides compositions in which the splice region 1 comprises at least about 5 nucleotides and up to 10,000 nucleotides. This document also provides compositions in which the splice region 1 comprises at least about 7 nucleotides and up to 1,000 nucleotides. This document also provides compositions in which the splice region 1 comprises at least about 5 nucleotides and up to 20 nucleotides. This document also provides compositions in which the splice region 1 comprises an inverse complementary sequence encoding an intron or a variant thereof. This document also provides compositions in which the splice region 1 comprises an inverse complementary sequence encoding an exon or a variant thereof. This document also provides compositions in which the splice region 1 comprises an inverse complementary sequence encoding both exons and introns. This document also provides compositions in which the splice region 1 comprises a sequence encoding a splice acceptor site or a nucleotide complementary to a splice acceptor site. This document also provides compositions wherein the connecting splint region 1 contains a complementary sequence encoding an intron or a variant thereof. This document also provides compositions wherein the connecting splint region 1 contains a complementary sequence encoding an exon or a variant thereof. This document also provides compositions wherein the connecting splint region 1 contains complementary sequences encoding both exons and introns. This document also provides compositions wherein the connecting splint region 1 contains nucleobases mismatched relative to splice acceptor sites in a target nucleic acid sequence. This document also provides compositions wherein a ribonucleotide is located at the 3' end of the connecting splint region 1. This document also provides compositions wherein the connecting splint region 1 contains 1 ribonucleotide, 2 ribonucleotides, 3 ribonucleotides, 4 ribonucleotides, 5 ribonucleotides, or up to 10 ribonucleotides. This document also provides compositions wherein the length of the connecting splint region 1 is at least about 5 nucleotides to at most 10 nucleotides. This document also provides compositions wherein the composition further comprises an engineered protein or a polypeptide encoding an engineered protein. This document also provides compositions wherein the engineered protein comprises: (a) a nicking enzyme region; and (b) a DNA ligase region. This document also provides compositions in which the DNA ligase region binds to either the ligation clip 1 region or the ligation clip 2 region that guides the polynucleotide. This document also provides compositions in which the DNA ligase region ligates the donor nucleic acid to the 3' end of the leading strand. This document also provides compositions in which the nicking enzyme region contains a mutant Cas protein that produces a single-strand break in the target nucleic acid. This document also provides compositions in which engineered Cas proteins include: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas11, Cas12, Cas13, Cas14, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2,Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, c2c1, c2c3, Cas9HiFi, xCas9, CasX, CasY, CasRX, variants, or any combination thereof. Compositions are also provided herein, wherein the composition further comprises a delivery medium. This document also provides compositions in which the delivery medium includes carriers, lipids, nanoparticles, plasmids, viruses, liposomes, extracellular vesicles, emulsions, peptides, sugars, polymers, chitosan, polyethyleneimine (PEI), poly(lactide-co-glycolic acid) (PLGA), poly-L-lysine (PLL), or combinations thereof.

[0190] This article provides a method for linking donor nucleic acids to target nucleic acids, wherein the method comprises: contacting a cell or cell-free system with: (a) a donor nucleic acid, (b) a guiding polynucleotide or a polynucleotide encoding a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a target region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure for binding to a nuclease or nicking enzyme; and (iii) a linker 2 region, wherein the linker 2 region is complementary to the donor nucleic acid and complementary to the target nucleic acid and has at least one modification relative to the target nucleic acid; (iv) a linker 1 region, wherein the linker 1 region comprises one or more deoxyribonucleotides or one or more ribonucleotides, and (c) an engineered protein or a polynucleotide encoding an engineered protein, wherein the engineered protein comprises: (i) a nicking enzyme region; and (ii) The DNA ligase region includes: a region that guides the formation of a complex between a polynucleotide and an engineered protein via a protein-binding region; a region that guides the formation of a complex between the target sequence and the complementary strand; a nicking enzyme region of the engineered protein that induces a break in the target nucleic acid to produce a leading strand and a complementary strand; a splint 1 region that forms a complex with the leading strand; and a region that attaches the donor nucleic acid to the leading strand, thereby ligating the donor nucleic acid to the target nucleic acid. This paper also provides a method in which the target strand and the complementary strand dissociate. Furthermore, this paper provides a method in which a novel nucleic acid is incorporated into the target nucleic acid by hybridization with the complementary strand.

[0191] This document provides methods comprising administering to a subject, organ, tissue, or cell a system, composition, or cell provided herein, wherein the administration modifies a gene in the subject, organ, tissue, or cell. This document also provides methods wherein the alteration includes: insertion, deletion, substitution, copy number change, point mutation, frameshift mutation, missense mutation, nonsense mutation, stop codon mutation, epigenetic marker, or any combination thereof. This document also provides methods wherein the administration is local or systemic. This document also provides methods wherein the administration is intranasal, subcutaneous, intravenous, inhaled, intramuscular, intratumoral, peritumoral, intrathecal, vaginal, or intradermal. This document also provides methods wherein the subject is a mammal. This document also provides methods wherein the subject has, is suspected of having, or has been diagnosed with a disease or condition. This document also provides methods wherein the disease or condition includes a genetic disease or condition. This document also provides methods wherein the method further includes administering a therapeutic agent to the subject.

[0192] This document provides methods comprising: contacting cells or cell populations with the system or composition provided herein to modify genes in the cells. This document also provides methods wherein the contact is performed in vitro, in vivo, or ex vivo. This document further provides methods wherein gene alterations include: insertions, deletions, substitutions, copy number changes, point mutations, frameshift mutations, missense mutations, nonsense mutations, stop codon mutations, epigenetic markers, or any combination thereof. This document also provides methods wherein, relative to equivalent cells or cell populations that have not been contacted with the system or composition, the gene alteration restores the expression of a wild-type protein encoded by the gene. This document further provides methods wherein the cells include eukaryotic or prokaryotic cells. This document further provides methods wherein the cells include mammalian cells. This document further provides methods wherein the cell population includes a human leukocyte population, a stem cell population, or a bacterial population.

[0193] This article provides cell populations prepared using the methods described herein.

[0194] This document provides a kit comprising: the system provided herein, the composition provided herein, the vector provided herein, the polynucleotide or the cell provided herein, and packaging and materials for use herein.

[0195] This document provides a kit comprising: a first container and a second container, the first container comprising: a donor nucleic acid, and a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) a targeting region complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure for binding to a nuclease or nicking enzyme; (iii) a ligation splice 2 region, wherein the ligation splice 2 region is complementary to both the donor and target nucleic acids, and wherein the ligation splice 2 region comprises at least one mismatched nucleobase relative to the target nucleic acid; and (iv) a ligation splice 1 region, wherein the ligation splice 1 region comprises a deoxyribonucleotide and a ribonucleotide, the second container comprising: an engineered polypeptide comprising a nicking enzyme operatively linked to a DNA ligase. This document also provides a kit comprising reagents for nucleic acid amplification, transcription, translation, or nucleic acid isolation. This document also provides a kit comprising a reporter molecule.

[0196] This document provides a support, wherein the support comprises: the system or composition provided herein; and a solid surface, wherein the system or composition is fixed to the solid surface.

[0197] Example Example 1. Editing double-stranded targets using a DNA ligase system A DNA editing system comprising a Cas9 nickase (H840A), a DNA ligase, a guide polynucleotide, and a donor nucleic acid was prepared and evaluated using different double-stranded DNA targets (double-stranded DNA substrates). Specifically, double-stranded DNA substrates, guide polynucleotides, and RNPs consisting of the Cas9 H840A nickase, DNA ligase, and guide polynucleotides were generated and evaluated, as described below.

[0198] Double-stranded DNA substrate production 5'-FAM-labeled double-stranded DNA (dsDNA) was obtained by annealing two oligonucleotides, oligonucleotide 1 and oligonucleotide 2, in a 1:1 ratio. The oligonucleotides were resuspended in nuclease-free water to a concentration of 100 μM. Subsequently, 10 μL of 5'-FAM-labeled oligonucleotide 1 and 10 μL of oligonucleotide 2 were mixed with 10 μL of NEBuffer r3.1 (10×) and 70 μL of nuclease-free water in a PCR tube and incubated at 95°C for 5 minutes, followed by slow cooling to room temperature. Annealed FAM-labeled dsDNA was prepared at a concentration of 10 μM or 10 picomoles / μL. The oligonucleotide sequences are listed in Table 3.

[0199] In vitro guided polynucleotide productionIn summary, PCR (Q5® Hot Start High Fidelity 2× Master Mixture, NEB) was performed using oligonucleotides (oligonucleotide 3, oligonucleotide 4, oligonucleotide 5, and oligonucleotide 6) to generate a double-stranded DNA template containing the T7 promoter for in vitro transcription. DNA from the PCR reaction was purified using AMPure XP beads (Beckman Coulter) and its concentration was measured using Nanodrop. 75 ng of DNA template was used for the in vitro transcription reaction using the HiScribe® T7 Rapid High Yield RNA Synthesis Kit (NEB). The guiding nucleic acid products were purified using the Monarch® RNA Cleanup Kit (NEB) and their concentration was measured using Nanodrop (Thermo Fisher Scientific). The oligonucleotide sequences are listed in Table 3. In Table 3, A, G, C, and T are deoxyribonucleotides (DNA), and rA, rG, rC, and rU are ribonucleotides (RNA).

[0200] Table 3. Oligonucleotides Synthesis-guided polynucleotide production Design and chemically synthesize all the synthetic guides listed in Table 4. Purify the guides by HPLC, and verify the quality of each individual guide polynucleotide by mass spectrometry (LC-MS). In Table 4, A, C, G, and U are ribonucleotides (RNA), and dA, dC, dG, and dT are deoxyribonucleotides (DNA).

[0201] Table 4. Synthesis-Guided Polynucleotides In vitro DNA nicking, DNA ligation via DNA ligase, and analysisFirst, CRISPR ribonucleoproteins (RNPs) were prepared by combining 61 picomol of Streptococcus pyogenes Cas9 H840A nickase with 80 picomol of guide RNA and incubating at room temperature for 10 minutes to allow RNP complexation. Cas9 nickase and DNA ligase activities were performed in a total 7 μL volume containing 1 μL of FAM dsDNA (10 picomol), 100 picomol of oligonucleotides for trans DNA editing (for cis DNA writing, incorporating the DNA template into the guide polynucleotide sequence), 0.7 μL of 10× T4 DNA ligase reaction buffer (NEB), 0.5 μL of DNA ligase, 2 μL of Cas9 nickase-guided RNP, and the remaining volume of nuclease-free water, making a final volume of 7 μL. Various lengths of ligation donor oligomers were provided at 100 picomol / reaction for trans DNA editing. The reaction was incubated at 37°C for 1 hour, and the sample was treated with 0.5 μL of proteinase K solution (20 mg / mL, Qiagen) and incubated at 56°C for 30 min. The sample was then heat-inactivated at 95°C for 10 min. The reaction product was combined with gel loading buffer II (2X, ThermoFisherScientific) and denatured at 95°C for 5 min, followed by separation on a denaturing polyacrylamide gel (15% TBE-urea, 60°C, 150V) for 1 hour. The DNA product was visualized using a Life Technologies gel imaging system via FAM fluorescence signals.

[0202] Editing products from targeted DNA editing via DNA ligase were analyzed using denaturing urea-polyacrylamide gel electrophoresis, such as... Figure 2 As shown in the diagram. Lanes 6-11 demonstrate the successful insertion of nucleotides of different lengths, including 15 nt (lane 6, bottom band), 25 nt (lane 7, top band), 30 nt (lane 8, top band), 50 nt (lane 9, top band), 70 nt (lane 10, top band) and 90 nt (lane 11, top band).

[0203] To determine the effect of the LS1 composition on the editing efficiency of a DNA editing system, denaturing urea-polyacrylamide gel electrophoresis was used to analyze the DNA ligation editing activity via DNA ligase, wherein the DNA ligation splint (LS) was incorporated into different synthetic guide polynucleotides (guides 1 to 9 (SEQ ID NO: 45-53)), such as... Figure 3As shown in the diagram. Specifically, guides 1 through 9 all have a connecting clip 2 (LS2) with the following structure: 5' DDDDDDDDDDDDDDD 3' (D represents DNA), while LS1 consists of all DNA bases (guides 1, 2, and 3 (SEQ ID NO: 45-47)) or both RNA and DNA (guides 4, 5, 6, 7, 8, and 9 (SEQ ID NO: 48-53)). The LS1 in guides 1, 2, and 3 has the following structure: 5′ DDDDD 3′, 5′DDDDDD 3′, and 5′ DDDDDDD 3′, respectively (D represents DNA, and R represents RNA). The LS1 in guides 4, 5, 6, 7, 8, and 9 has the following architecture: 5' DDDDDRR 3', 5' DDDDRRR 3', 5' DDDRRRR 3', 5' DDRRRRR 3', 5' DRRRRRR 3', and 5' RRRRRRR 3', respectively. Lanes 3-11 have a linker oligonucleotide (oligonucleotide 40) for adding 15 nt to the nicked DNA to produce a 49 nt edited product.

[0204] Figure 3 As shown, most of the nicked products are converted into ligation-edited products. However, increasing the length of LS1 further maximizes efficiency. When the amount of DNA in LS1 is less than 3, the activity of T4 DNA ligase for editing decreases significantly, indicating that at least 3 DNA nucleotides in LS1 are required for effective DNA editing by T4 DNA ligase.

[0205] Example 2. Plasmid Cloning Backbone plasmids containing the T7 RNA polymerase promoter, 5' UTR sequence, and 3' UTR sequence were constructed to clone various editor constructs. Gene fragments containing various Cas enzymes and DNA ligases were assembled into the backbone plasmids using ligation-based cloning or Gibson assembly for use as templates for in vitro transcription. All plasmids were prepared using the Qiagen Midi / Maxi Plus kit (Midi or Maxi). The amino acid sequences of the resulting engineered protein constructs are provided in Table 5 below.

[0206] Table 5. Engineered Protein Constructs (DNA Ligase Genome Editor) Example 3. HEK293T culture and electroporation.

[0207] Frozen vials of the HEK293T (ATCC CRL-3216) cell line were thawed at 37°C and cultured in Dulbecco modified Eagle medium supplemented with 10% (v / v) fetal bovine serum (FBS) at 37°C and 5% CO2. Cells were grown to 90% confluence before passage. Cells were passaged three times before any experiment. Before electroporation, 48-well plates were coated with poly-D-lysine for 1 hour and then washed three times with PBS. After PBS washing, 400 μL of antibiotic-free complete medium was added to each well and incubated at 37°C and 5% CO2 before use. All electroporation was performed using the Neon NxT electroporation system. For electroporation, 1 μg of purified mRNA, 100 pmol of each legRNA and splint, 100 pmol of donor DNA and 100,000 cells were mixed in buffer R and electroporated with 10 μl of Neon tip. An electroporation setup of 1150V, 20ms, and 2 pulses was used for all electroporation wells. After electroporation, cells were added to individual wells of a 48-well plate containing cell culture medium and incubated at 37°C and 5% CO2. Genomic DNA was extracted 48 hours after electroporation for NGS analysis.

[0208] Example 4. In vitro transcribed (IVT) mRNA Plasmids containing the genes of interest were completely digested and linearized using BsmBI (New England Biolabs) before being used for in vitro transcription (IVT). IVT reactions were performed using the NEB HiScribe T7 High-Yield RNA Synthesis Kit (New England Biolabs). In summary, IVT reactions were performed at 37°C with the addition of CleanCap reagent AG (Trilink Biotechnologies), in which N1-methyl-pseudo-UTP (Trilink Biotechnologies) was used to replace 100% of the UTP. IVT reactions were terminated after 2 hours. After IVT, each reaction was incubated with DNase I (New England Biolabs) for 15 minutes. RNA was then purified using the Monarch RNA Cleanup Kit (New England Biolabs).

[0209] Example 5. Next-Generation Sequencing (NGS) Library Preparation and Analysis.

[0210] PCR primers containing Illumina-compatible adaptor sequences (see Table 7) were used to amplify specific genomic regions of interest. Following standard PCR protocols using a Q5 hot-start high-fidelity 2× master mix (New England Biolabs), the resulting PCR products were cleaned with 0.7× Ampure XP beads (Beckman Coulter). The purified PCR products were sent to the AmpExpress service at Quintara Biosciences. Amplicon sequencing data were analyzed using CRISPResso. Editing efficiency was calculated from an allele frequency table file based on the number of sequencing reads containing the desired edit of interest relative to the total number of sequencing reads in a given sample.

[0211] Example 6. Gene editing in HEK293 cells using engineered proteins of DNA ligase and nickase. To evaluate a gene editing system comprising a Cas9 nickase, a DNA ligase, a guide polynucleotide, and a donor nucleic acid in human cells, mRNA encoding the engineered protein constructs in Table 5 was transcribed in vitro according to the method described in Example 4, and the synthetic guide RNA in Table 6 was prepared by chemical synthesis. The mRNA, synthetic guide RNA, and oligonucleotide were delivered to HEK293T cells as described in Example 3.

[0212] The engineered protein constructs (constructs 1-13; SEQ ID NO: 89-101) were screened for precise genome editing activity at the FANCF site 1, the genome target site, in HEK293T cells. The results showed that the DNA ligase genome editor facilitated precise genome editing in human cells. Among the thirteen engineered protein constructs, the construct containing T4 DNA ligase, Chlorella DNA ligase, and human DNA ligase IV editors enabled precise installation of a 3 bp substitution edit (+2 C>T; +4–5 TG>AC) at the FANCF site 1 in HEK293T cells. Figure 4 ).

[0213] To verify that the effects of splint and DNA donor modifications on enhancing DNA ligase editing activity in human cells, various splint and DNA donor modifications (Table 7, SEQ ID NO: 184-228) were tested. Ligase editing at the genomic target FANCF site 1 (gRVB_3) using a Chlorella DNA ligase editor construct (constructor 10, SEQ ID NO: 98) was performed to install a 3 bp substitution edit (+2 C>T; +4–5 TG>AC). The editing efficiencies of gene editing systems using various combinations of splint and DNA donor are shown in the figure. Figure 5 The results showed that modifications such as phosphate thioester bonds (PS) and locked nucleic acids (LNAs) (as shown in the combination of oRVB_61 and oRVB_97) also increased ligase editing activity by up to 21.5 times compared to combinations without such modifications (oRVB_62 / oRVB_95).

[0214] To verify the effectiveness of incorporating RNA into a guide RNA clip during DNA ligase editing, a gene editing system containing a Chlorella DNA ligase editor construct (construct 10) at HEK site 3 and AAV site 1 in HEK293T cells and various ligase-edited guide RNAs was tested. Delivery of in vitro transcribed mRNA encoding the ligase editor, synthesized guide RNA, DNA clip, and donor mRNA into HEK293T cells was performed via electroporation as described in Example 4.

[0215] mRNA was delivered to HEK293T cells as described in Example 3. Figure 6 As shown, compared to guide RNA clips without RNA incorporation (oRVB_81 / oRVB_82), the RNA-containing ligase editing guide RNA (gRVB_5 / gRVB_6) within clips (oRVB_73 / oRVB_74) enhanced the precision of genome editing at the HEK site for insertion at the attB site by up to 36.9-fold. Similarly, as Figure 7As shown, compared to guide RNA clips without RNA incorporation (oRVB_79 / oRVB_80), ligase editing guide RNAs containing RNA within clips (oRVB_71 / oRVB_72) (gRVB_1 / gRVB_2) enhance the precise genome editing at AAV site 1 for insertion attB site by up to 10.5-fold.

[0216] To verify the effectiveness of the amount of RNA incorporated into the ligase editor-guided RNA (legRNA) splint region, gene editing systems containing a Chlorella DNA ligase editor construct (constructor 10) and ligase editor-guided RNA (legRNA) with varying degrees of RNA incorporation into the splint region were tested to determine the ligase editing activity at the AAV site 1, the target site in the human cell genome. The desired editing efficiency refers to the efficiency of inserting the attB site at AAV site 1. Delivery of in vitro transcribed mRNA encoding the ligase editor, synthesized guide RNA, and DNA donor into HEK293T cells was performed via electroporation as described in Example 4. Figure 8 As shown, incorporating RNA to varying degrees into the splint region of a ligase editor guide RNA (legRNA) with a single guide RNA architecture promotes ligase editing activity at the AAV site 1, a genomic target, in human cells.

[0217] Table 6. Synthesis-Guiding Polynucleotides Table 7. Oligonucleotides In Tables 6 and 7, A, C, G, and U are ribonucleotides (RNA), and dA, dC, dG, and dT are deoxyribonucleotides (DNA), +A, +T, +C, and +G are locked nucleic acids (LNA), rA, rU, rC, and rG are ribonucleotides (RNA), * indicates a phosphate thioester (PS) bond, mA, mU, mC, and mG correspond to 2'-O-methyl RNA, / iMe-Dc / corresponds to 5'-methyl dC, / 3Me-dC / corresponds to 3'-5-methyl dC, and / 5Phos / corresponds to 5' phosphorylation.

[0218] While preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. The following claims are intended to define the scope of the invention and thereby cover the methods and structures within the scope of these claims and their equivalents.

Claims

1. A composition comprising: (a) Donor nucleic acid; (b) A polynucleotide encoding a protein construct comprising a nicking enzyme or a variant thereof and a DNA ligase or a functional fragment thereof; and (c) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) Target regions that are complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to an engineered protein construct containing a nicking enzyme region; (iii) A connecting clamp region 2, wherein the connecting clamp region is complementary to both the donor nucleic acid and the target nucleic acid, and wherein the connecting clamp region includes at least one modification relative to the target nucleic acid; and (iv) Connecting splint region 1, wherein the connecting splint region 1 comprises: deoxyribonucleotide and ribonucleotide, wherein the connecting splint region 1 is complementary to the target nucleic acid.

2. The composition according to claim 1, wherein the connecting clamp region 1 forms a DNA-RNA (DR) loop after associating with the target nucleic acid and DNA ligase.

3. The composition of claim 1, wherein the ligase clip region 1 comprises an RNA to DNA ratio of 1:1 up to 20:

1.

4. The composition according to claim 3, wherein the connecting clamp region 1 comprises the following ratios of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 2:1, 2:3, 2:5, 2:7, 2:9, 2:11, 2:13, 2:15, 2:17, 2:19, 3:1, 3:2, 3:4, 3:5, 3:7, 3:8, 3:10, 3:11, 3:13, 3:14, 3:15, 3:16, 3:1 7, 3:19, 4:1, 4:3, 4:5, 4:7, 4:9, 4:11, 4:13, 4:15, 4:17, 4:19, 5:1, 5:2, 5:3, 5:4, 5:6, 5:7, 5:8, 5:9, 5:11, 5:12, 5:13, 5:14, 5:16, 6:1, 6:5, 6:7, 6:9 6:11, 6:13, 6:15, 7:1, 7:2, 7:3, 7:4, 7:5, 7:6, 7:8, 7:9, 7:10, 7:11, 7:12, 7:13, 7:15, 8:1, 8:3, 8:5, 8:7, 8:9, 8:11, 8:13, 8:15, 9:1, 9:2, 9:4, 9:5, 9 :7, 9:8, 9:10, 9:11, 9:13, 9:15, 9:17, 9:19, 9:20, 10:1, 10:3, 10:7, 10:9, 10:11, 10:13, 10:15, 10:17, 10:19, 11:1, 11:2, 11:3, 11:4, 11:5, 11:6, 11 :7, 11:8, 11:9, 11:10, 11:12, 11:13, 11:15, 12:1, 12:5, 12:7, 12:9, 12:11, 12:13, 13:1, 13:2, 13:3, 13:4, 13:5, 13:6, 13:7, 13:8, 13:9, 13:10, 13:11 13:12, 13:14, 14:1, 14:3, 14:5, 14:9, 14:11, 14:13, 15:1, 15:2, 15:4, 15:6, 15:8, 15:11, 15:13, 16:1, 16:3, 16:5, 16:7, 16:9, 16:11, 16:13, 16:15 17:1, 17:2, 17:3, 17:4, 17:5, 17:6, 17:7, 17:8, 17:9, 17:10, 17:11, 17:12, 17:13, 17:14, 17:15, 17:16, 18:1, 18:5, 18:7, 18:11, 18:13, 18:17, 19:119:2, 19:3, 19:4, 19:5, 19:6, 19:7, 19:8, 19:9, 19:10, 19:11, 19:12, 19:13, 19:14, 19:15, 19:16, 19:17, 19:18 or 20:

1.

5. The composition of claim 1, wherein the connecting clamp 2 region comprises at least about 5 nucleotides to at most 10,000 nucleotides.

6. The composition of claim 1, wherein the connecting plate 2 region comprises at least about 7 nucleotides to at most 1,000 nucleotides.

7. The composition of claim 1, wherein the connecting clamp region 1 comprises at least about 5 nucleotides to at most 20 nucleotides.

8. The composition of claim 1, wherein the connecting clip 2 region comprises an inverse complementary sequence of a non-coding polynucleotide sequence or a variant thereof.

9. The composition of claim 1, wherein the connecting clamp region 2 comprises the inverse complementary sequence of a coding region encoding a polynucleotide sequence or a variant thereof.

10. The composition of claim 1, wherein the connecting clamp 2 region comprises an inverse complementary sequence of a sequence encoding an exon or intron.

11. The composition of claim 1, wherein the connecting clamp region 2 comprises a complementary sequence to a sequence encoding a non-coding polynucleotide sequence or a variant thereof.

12. The composition of claim 1, wherein the connecting clamp region 2 comprises a complementary sequence to a coding region encoding a polynucleotide sequence or a variant thereof.

13. The composition of claim 1, wherein the connecting clamp 2 region comprises a complementary sequence to a sequence encoding an exon or intron.

14. The composition of claim 1, wherein the connecting clamp 2 region comprises a sequence containing at least one nucleobase that is complementary to or mismatched with the sequence encoding the splice acceptor site.

15. The composition of claim 1, wherein the nicking enzyme is an engineered Cas protein or a functional variant thereof.

16. The composition of claim 15, wherein the engineered Cas protein or a functional variant thereof comprises at least one amino acid substitution corresponding to position 840 of SEQ ID NO:

69.

17. The composition of claim 16, wherein the engineered Cas protein or a functional variant thereof further comprises at least one amino acid substitution corresponding to position 221, position 394 or a combination thereof of SEQ ID NO:

69.

18. The composition of claim 16, wherein the engineered Cas protein or a functional variant thereof comprises amino acid substitutions of R221K, N394K, H840A or any combination thereof.

19. The composition of claim 15, wherein the engineered Cas protein or a functional variant thereof comprises 97%, 98%, 99%, or 100% identical sequence to SEQ ID NO: 70 or SEQ ID NO:

71.

20. The composition of claim 15, wherein the engineered Cas protein or a functional variant thereof comprises the sequence of SEQ ID NO: 70 or SEQ ID NO:

71.

21. The composition according to claim 1, wherein the DNA ligase or functional fragment comprises *Escherichia coli* (…). E. coli DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Chlorella virus DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human ligase V, and their variants or combinations.

22. The composition of claim 19, wherein the DNA ligase or a functional fragment thereof comprises Chlorella virus DNA ligase, T4 DNA ligase or human DNA ligase IV.

23. The composition of claim 19, wherein the DNA ligase or a functional fragment thereof comprises the same sequence as SEQ ID NO: 58-66 90%, 91%, 92%, 93%, 94%, 95%, 97%, 98%, 99%, or 100%.

24. The composition of claim 19, wherein the DNA ligase or a functional fragment thereof comprises the sequence of SEQ ID NO:58-66.

25. The composition of claim 1, wherein the protein construct further comprises a linker, a nuclear localization sequence (NLS), or a combination thereof.

26. The composition of claim 25, wherein the DNA ligase or a functional fragment thereof is ligated to the nicking enzyme or a variant thereof.

27. The composition of claim 25, wherein the protein construct comprises the same sequence as SEQ ID NO: 89-10190%, 91%, 92%, 93%, 94%, 95%, 97%, 98%, 99%, or 100%.

28. An engineered fusion protein comprising: (a) DNA ligase or a functional fragment thereof; and (b) An engineered nickase comprising three amino acid substitutions at positions 221, 394, and 840 of a nuclease containing the sequence of SEQ ID NO:

69.

29. The engineered fusion protein of claim 28, wherein the substitution results in enhanced nicking enzyme activity compared to other equivalent engineered nicking enzymes.

30. The engineered fusion protein of claim 28, wherein the engineered nickase comprises amino acid substitutions of R221K, N394K, and H840A.

31. The engineered fusion protein of claim 28, wherein the engineered nickase or a functional variant thereof comprises 97%, 98%, 99% or 100% of the same sequence as SEQ ID NO:

71.

32. The engineered fusion protein of claim 28, wherein the engineered nickase or a functional variant thereof comprises the sequence of SEQ ID NO:

71.

33. The engineered fusion protein of claim 32, wherein the DNA ligase or functional fragment comprises Escherichia coli DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Chlorella virus DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human ligase V, or variants or combinations thereof.

34. The engineered fusion protein of claim 32, wherein the DNA ligase or a functional fragment thereof comprises Chlorella virus DNA ligase, T4 DNA ligase or human DNA ligase IV.

35. The engineered fusion protein of claim 34, wherein the DNA ligase or a functional fragment thereof comprises 97%, 98%, 99%, or 100% identical sequences to SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 61, or SEQ ID NO:

65.

36. The engineered fusion protein of claim 34, wherein the DNA ligase or a functional fragment thereof comprises the sequence of SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 61 or SEQ ID NO:

65.

37. The engineered fusion protein according to any one of claims 28 to 36 further comprises an adapter, a nuclear localization sequence (NLS), or another protein construct.

38. The engineered fusion protein according to claim 37, comprising 97%, 98%, 99% or 100% of the same amino acid sequence as any one of SEQ ID NO: 89-101.

39. The engineered fusion protein according to claim 37, comprising the amino acid sequence of any one of SEQ ID NO: 89-101.

40. The engineered fusion protein of claim 37, wherein the additional protein construct comprises a cell-targeting portion, a receptor-targeting portion, a regulatory element, a nuclease, an acetyltransferase, an acetyltransferase, an ATPase, an Argonaute protein, a base editor, a Cas polypeptide, a catalytically inactivated Cas polypeptide, a deacetylase, a deaminase, a decapping protein, an endonuclease, an exonuclease, a helicase, a ligase, a meganuclease, a methyltransferase, a nickase, a polymerase, a protease, a recombinase, a restriction enzyme, a ribonucleoprotein (RNP), a self-cleaving protein sequence, a splicing factor, a transcription activator, a transcription activator-like effector nuclease (TALEN), a transcription repressor, a transposase, a zinc finger, or any combination thereof.

41. A nucleic acid that encodes an engineered fusion protein according to any one of claims 28 to 40.

42. The nucleic acid of claim 41, wherein the nucleic acid encoding the engineered fusion protein comprises RNA.

43. The nucleic acid of claim 41, wherein the nucleic acid encoding the engineered fusion protein comprises DNA.

44. A vector comprising a nucleic acid according to any one of claims 41 to 43.

45. The vector according to claim 44, wherein the vector comprises a viral vector.

46. ​​The vector according to claim 45, wherein the viral vector includes a lentiviral vector, a retroviral vector, adeno-associated virus (AAV) vector, adenovirus vector, herpes simplex virus vector, alphavirus vector, flavivirus vector, rhabdovirus vector, measles virus vector, Newcastle disease virus vector, poxvirus vector, microRNA virus vector, or oncolytic virus vector.

47. A system for modifying a target nucleic acid, the system comprising: (a) Donor nucleic acid; (b) A polynucleotide encoding an engineered protein construct containing a nicking enzyme region or a variant thereof and a DNA ligase or a functional fragment thereof; and (c) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) Target regions that are complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) Connecting clamp region 2, wherein the connecting clamp region is complementary to the donor nucleic acid and to the target nucleic acid, and wherein the connecting clamp region includes at least one change relative to the target nucleic acid; and (iv) Connecting splint region 1, wherein the connecting splint region 1 comprises: deoxyribonucleotides and ribonucleotides, wherein the connecting splint region 1 is complementary to the target nucleic acid. In this system, after being introduced into a cell, cell nucleus, or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.

48. The system of claim 47, wherein the connecting clamp region 1 forms a DNA-RNA (DR) loop after associating with the target nucleic acid and the DNA ligase.

49. The system of claim 47, wherein the connecting clamp 1 region comprises a ribonucleic acid (RNA) to deoxyribonucleic acid (DNA) ratio of 1:1 up to 20:

1.

50. The system of claim 47, wherein the connecting clamp 1 region comprises the following ratios of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 2:1, 2:3, 2:5, 2:7, 2:9, 2:11, 2:13, 2:15, 2:17, 2:19, 3:1, 3:2, 3:4, 3:5, 3:7, 3:8, 3:10, 3:11, 3:13, 3:14, 3:15, 3:16, 3: 17, 3:19, 4:1, 4:3, 4:5, 4:7, 4:9, 4:11, 4:13, 4:15, 4:17, 4:19, 5:1, 5:2, 5:3, 5:4, 5:6, 5:7, 5:8, 5:9, 5:11, 5:12, 5:13, 5:14, 5:16, 6:1, 6:5, 6:7, 6:9 6:11, 6:13, 6:15, 7:1, 7:2, 7:3, 7:4, 7:5, 7:6, 7:8, 7:9, 7:10, 7:11, 7:12, 7:13, 7:15, 8:1, 8:3, 8:5, 8:7, 8:9, 8:11, 8:13, 8:15, 9:1, 9:2, 9:4, 9:5, 9 :7, 9:8, 9:10, 9:11, 9:13, 9:15, 9:17, 9:19, 9:20, 10:1, 10:3, 10:7, 10:9, 10:11, 10:13, 10:15, 10:17, 10:19, 11:1, 11:2, 11:3, 11:4, 11:5, 11:6, 11 :7, 11:8, 11:9, 11:10, 11:12, 11:13, 11:15, 12:1, 12:5, 12:7, 12:9, 12:11, 12:13, 13:1, 13:2, 13:3, 13:4, 13:5, 13:6, 13:7, 13:8, 13:9, 13:10, 13:11 13:12, 13:14, 14:1, 14:3, 14:5, 14:9, 14:11, 14:13, 15:1, 15:2, 15:4, 15:6, 15:8, 15:11, 15:13, 16:1, 16:3, 16:5, 16:7, 16:9, 16:11, 16:13, 16:15 17:1, 17:2, 17:3, 17:4, 17:5, 17:6, 17:7, 17:8, 17:9, 17:10, 17:11, 17:12, 17:13, 17:14, 17:15, 17:16, 18:1, 18:5, 18:7, 18:11, 18:13, 18:17, 19:119:2, 19:3, 19:4, 19:5, 19:6, 19:7, 19:8, 19:9, 19:10, 19:11, 19:12, 19:13, 19:14, 19:15, 19:16, 19:17, 19:18 or 20:

1.

51. The system of claim 47, wherein after the engineered protein construct cleaves the target nucleic acid, the connecting splint region 1 hybridizes with the leading strand.

52. The system of claim 47, wherein the ribonucleotides and / or deoxyribonucleotides of the connecting splice 1 region hybridize with the leading strand.

53. The system of claim 47, wherein the connecting clamp 2 region comprises at least about 5 nucleotides to at most 10,000 nucleotides.

54. The system of claim 47, wherein the connecting clamp 2 region comprises at least about 7 nucleotides to at most 1,000 nucleotides.

55. The system of claim 47, wherein the connecting clamp region 1 comprises at least about 5 nucleotides to at most 20 nucleotides.

56. The system of claim 47, wherein the connecting clamp region 2 comprises an inverse complementary sequence of a non-coding polynucleotide sequence or a variant thereof.

57. The system of claim 47, wherein the connecting clamp region 2 comprises the inverse complementary sequence of a coding region encoding a polynucleotide sequence or a variant thereof.

58. The system of claim 47, wherein the connecting clamp 2 region comprises an inverse complementary sequence of a sequence encoding an exon or intron.

59. The system of claim 47, wherein the connecting clamp region 2 comprises a complementary sequence to a sequence encoding a non-coding polynucleotide sequence or a variant thereof.

60. The system of claim 47, wherein the connecting clamp region 2 comprises a complementary sequence to a coding region encoding a polynucleotide sequence or a variant thereof.

61. The system of claim 47, wherein the connecting clamp 2 region comprises a complementary sequence of a sequence encoding an exon or intron.

62. The system of claim 47, wherein the connecting clamp 2 region comprises a sequence containing at least one nucleobase that is complementary to or mismatched with the sequence encoding the splice acceptor site.

63. The system of claim 47, wherein the targeting region hybridizes with the complementary strand of the target nucleic acid, and the cleavage enzyme region cleaves the target nucleic acid.

64. The system of claim 47, wherein the targeting region hybridizes with the complementary strand of the target nucleic acid within at least about 5 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides of the PAM sequence.

65. The system of claim 47, wherein the engineered protein construct comprises an engineered Cas protein.

66. The system of claim 65, wherein the engineered Cas protein is a type V Cas protein.

67. The system of claim 66, wherein the V-type Cas protein comprises: Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas14, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k.

68. The system of claim 66, wherein the engineered Cas protein comprises: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas11, Cas12, Cas13, Cas14, Csy1, Csy2, Csy3, Cse1 , Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx1 4. Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, c2c1, c2c3, Cas9HiFi, xCas9, CasX, CasY, CasRX, SpCas9-VQR, SpCas9-VRQR, SpCas9-VRER, SaCas9-KKH, SpCas9-NG, SpCas9-NRRH, SpCas9-NRTH, SpCas9-NRCH, iSpyMac, St1Cas9 LMD9-LMG18311, St1Cas9 LMD9-CNRZ1066, St1Cas9-KQKL, their variants or any combination thereof.

69. The system of claim 66, wherein the engineered Cas protein is a type II Cas protein.

70. The system of claim 69, wherein the type II Cas protein comprises: Cas9, Cas1, Cas2 or Csn2 or their variants.

71. The system according to any one of claims 47 to 70, wherein the engineered protein construct cleaves the target nucleic acid upstream of the 5' of the PAM sequence.

72. The system according to any one of claims 47 to 71, wherein the engineered protein construct cleaves the target nucleic acid to produce two single strands of DNA or two single strands of an RNA duplex.

73. The system according to any one of claims 47 to 72, wherein the engineered protein construct generates single-strand breaks in the target nucleic acid.

74. The system according to any one of claims 47 to 73, wherein the DNA ligase or a functional fragment thereof comprises Escherichia coli DNA ligase, Taq DNA ligase, T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, Chlorella virus DNA ligase, human DNA ligase I, human DNA ligase II, human DNA ligase III, human DNA ligase IV, human ligase V, variants or combinations thereof.

75. The system according to any one of claims 47 to 74, wherein the DNA ligase or a functional fragment thereof comprises Archaeocystis scintillans (… Archaeoglobus fulgidus DNA ligase, Archaeocarpus scintillans DNA ligase, and Coccidia flavovirens ( Pyrococcus furiosus DNA ligase, thermoautotrophic methanotherapeutic bacteria ( Methanothermobacter thermautotrophicus DNA ligase, deep-sea fireball bacteria ( Pyrococcus abyssi DNA ligase, sulfur-bearing leaf fungus ( Sulfolobus solfataricus DNA ligase, Haemophilus influenzae ( Haemophilus influenzae DNA ligase or herpes simplex virus (HSV) DNA ligase.

76. The system according to any one of claims 47 to 75, wherein the connecting clamp region 2 hybridizes with the donor nucleic acid.

77. The system according to any one of claims 47 to 76, wherein the DNA ligase does not bind to the target nucleic acid.

78. The system according to any one of claims 47 to 77, wherein the DNA ligase attaches the donor nucleic acid to the leading strand.

79. The system of claim 78, wherein the novel DNA strand has at least 50% to 99.99% complementarity with the target nucleic acid sequence.

80. The system of claim 78, wherein the new DNA strand comprises nucleobases that are complementary to the target nucleic acid.

81. The system according to any one of claims 47 to 80, wherein the secondary structure of the protein-binding region guiding the polynucleotide comprises: A protrusion, stem, loop, hairpin, oscillating base pair, pseudoknot, or a combination thereof.

82. The system according to any one of claims 47 to 81, wherein the target nucleic acid comprises DNA.

83. The system according to any one of claims 47 to 82, wherein the target nucleic acid comprises RNA.

84. The system according to any one of claims 47 to 83 further comprises additional protein constructs, adapters, or nuclear localization sequences (NLS).

85. The system according to any one of claims 47 to 84, wherein the nicking enzyme region or a variant thereof and the DNA ligase are linked together by a polypeptide adapter.

86. The system of claim 85, wherein the protein construct comprises the sequence of SEQ ID NO: 89-101.

87. A system for modifying a target nucleic acid, the system comprising: (a) Donor nucleic acid; (b) an engineered protein construct containing a nicking enzyme region or a variant thereof and a DNA ligase or a functional fragment thereof; and (c) a guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) Target regions that are complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure that binds to the engineered protein construct containing the nicking enzyme region; (iii) Connecting clamp region 2, wherein the connecting clamp region is complementary to the donor nucleic acid and to the target nucleic acid, and wherein the connecting clamp region includes at least one change relative to the target nucleic acid; and (iv) Connecting splint region 1, wherein the connecting splint region 1 comprises: deoxyribonucleotides and ribonucleotides, wherein the connecting splint region 1 is complementary to the target nucleic acid. In this system, after being introduced into a cell, cell nucleus, or cell-free system, the system incorporates the donor nucleic acid into the target nucleic acid, thereby modifying the target nucleic acid.

88. A composition comprising the system according to any one of claims 47 to 87; and a delivery medium.

89. A polynucleotide or polynucleotide group, said polynucleotide or polynucleotide group encoding a system according to any one of claims 47 to 87.

90. A vector comprising the polynucleotide of claim 89 or the polynucleotide group of claim 89.

91. A nanoparticle comprising the polynucleotide of claim 89 or the polynucleotide group of claim 89.

92. A cell comprising the system according to any one of claims 47 to 87.

93. A composition comprising: Donor nucleic acid, and The directing polynucleotide or the polynucleotide encoding the directing polynucleotide, wherein the directing polynucleotide comprises, in a 5' to 3' sequence: (a) Target regions that are complementary to the target nucleic acid; (b) A protein-binding region, wherein the protein-binding region includes a secondary structure that binds to a protein; (c) A connecting clamp region 2, wherein the connecting clamp region 2 is complementary to the donor nucleic acid and to the target nucleic acid, and has at least one alteration relative to the target nucleic acid or at least one nucleic acid chain thereof; and (d) Connecting clamp 1 region, wherein the connecting clamp 1 region comprises: deoxyribonucleotides and ribonucleotides.

94. The composition of claim 93, wherein the target nucleic acid comprises a double-stranded DNA (dsDNA) or a single-stranded RNA (ssRNA) duplex cleaved by a nicking enzyme.

95. The composition of claim 93, wherein the protein binding region is bound to a nicking enzyme or an endonuclease.

96. The composition of claim 93, wherein the secondary structure comprises a protrusion, a stem, a ring, a hairpin, a wobbling base pair, a pseudoknot, or a combination thereof.

97. The composition of claim 93, wherein the connecting clamp 2 region comprises at least about 5 nucleotides to at most 10,000 nucleotides.

98. The composition of claim 93, wherein the connecting clamp 2 region comprises at least about 7 nucleotides to at most 1,000 nucleotides.

99. The composition of claim 93, wherein the connecting plate 2 region comprises at least one change of mismatched nucleobases.

100. The composition of claim 99, wherein the mismatched nucleobases comprise A / C mismatch, A / T mismatch, A / G mismatch, and T / C mismatch, T / G mismatch, T / A mismatch, C / G mismatch, C / A mismatch, C / T mismatch, G / C mismatch, G / T mismatch, G / A mismatch, or combinations thereof, relative to the target nucleic acid.

101. The composition of claim 93, wherein the donor nucleic acid comprises the inverse complementary sequence of a non-coding region of a polynucleotide or a variant thereof.

102. The composition of claim 93, wherein the donor nucleic acid comprises the inverse complementary sequence of a coding region encoding a polynucleotide or a variant thereof.

103. The composition of claim 93, wherein the donor nucleic acid comprises an inverse complementary sequence of a sequence encoding an exon or intron.

104. The composition of claim 93, wherein the donor nucleic acid comprises a sequence containing at least one nucleobase that is complementary to or mismatched with the sequence encoding the splice acceptor site.

105. The composition of claim 93, wherein the donor nucleic acid comprises a complementary sequence encoding an intron or a variant thereof.

106. The composition of claim 93, wherein the donor nucleic acid comprises a complementary sequence encoding an exon or a variant thereof.

107. The composition of claim 93, wherein the donor nucleic acid comprises a complementary sequence encoding exons and introns.

108. The composition of claim 93, wherein the connecting clamp region 1 forms a DNA-RNA (DR) loop after associating with the target nucleic acid and DNA ligase.

109. The composition of claim 93, wherein the ligase clip region 1 comprises an RNA to DNA ratio of 1:1 up to 20:

1.

110. The composition according to claim 93, wherein the connecting clamp region 1 comprises the following ratios of ribonucleic acid (RNA) to deoxyribonucleic acid (DNA): 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 2:1, 2:3, 2:5, 2:7, 2:9, 2:11, 2:13, 2:15, 2:17, 2:19, 3:1, 3:2, 3:4, 3:5, 3:7, 3:8, 3:10, 3:11, 3:13, 3:14, 3:15, 3:16, 3:17, 3:18, 3:19, 3:10, 3:11, 3:13, 3:14, 3:15, 3:16 ... :17, 3:19, 4:1, 4:3, 4:5, 4:7, 4:9, 4:11, 4:13, 4:15, 4:17, 4:19, 5:1, 5:2, 5:3, 5:4, 5:6, 5:7, 5:8, 5:9, 5:11, 5:12, 5:13, 5:14, 5:16, 6:1, 6:5, 6:7, 6: 9, 6:11, 6:13, 6:15, 7:1, 7:2, 7:3, 7:4, 7:5, 7:6, 7:8, 7:9, 7:10, 7:11, 7:12, 7:13, 7:15, 8:1, 8:3, 8:5, 8:7, 8:9, 8:11, 8:13, 8:15, 9:1, 9:2, 9:4, 9:5 9:7, 9:8, 9:10, 9:11, 9:13, 9:15, 9:17, 9:19, 9:20, 10:1, 10:3, 10:7, 10:9, 10:11, 10:13, 10:15, 10:17, 10:19, 11:1, 11:2, 11:3, 11:4, 11:5, 11:6, 11 :7, 11:8, 11:9, 11:10, 11:12, 11:13, 11:15, 12:1, 12:5, 12:7, 12:9, 12:11, 12:13, 13:1, 13:2, 13:3, 13:4, 13:5, 13:6, 13:7, 13:8, 13:9, 13:10, 13:11 13:12, 13:14, 14:1, 14:3, 14:5, 14:9, 14:11, 14:13, 15:1, 15:2, 15:4, 15:6, 15:8, 15:11, 15:13, 16:1, 16:3, 16:5, 16:7, 16:9, 16:11, 16:13, 16:15 17:1, 17:2, 17:3, 17:4, 17:5, 17:6, 17:7, 17:8, 17:9, 17:10, 17:11, 17:12, 17:13, 17:14, 17:15, 17:16, 18:1, 18:5, 18:7, 18:11, 18:13, 18:17, 19:119:2, 19:3, 19:4, 19:5, 19:6, 19:7, 19:8, 19:9, 19:10, 19:11, 19:12, 19:13, 19:14, 19:15, 19:16, 19:17, 19:18 or 20:

1.

111. The composition of claim 93, wherein the connecting splint 1 region hybridizes with the target nucleic acid after the cleavage enzyme cleaves the target nucleic acid, wherein the cleavage enzyme produces a leading strand and a complementary strand.

112. The composition of claim 111, wherein the ribonucleotide hybridizes with the leading strand.

113. The composition of claim 93, wherein the connecting plate region 1 comprises at least about 5 nucleotides to at most 10,000 nucleotides.

114. The composition of claim 93, wherein the connecting plate region 1 comprises at least about 7 nucleotides to at most 1,000 nucleotides.

115. The composition of claim 93, wherein the connecting clamp region 1 comprises at least about 5 nucleotides to at most 20 nucleotides.

116. The composition of claim 93, wherein the connecting clamp region 1 comprises an inverse complementary sequence encoding an intron or a variant thereof.

117. The composition of claim 93, wherein the connecting clamp region 1 comprises an inverse complementary sequence encoding a sequence of an exon or a variant thereof.

118. The composition of claim 93, wherein the connecting clamp region 1 comprises an inverse complementary sequence of sequences encoding exons and introns.

119. The composition of claim 93, wherein the connecting clamp region 1 comprises a sequence encoding a splice acceptor site or a nucleotide complementary to the splice acceptor site.

120. The composition of claim 93, wherein the connecting clamp region 1 comprises a complementary sequence encoding a sequence of an intron or a variant thereof.

121. The composition of claim 93, wherein the connecting clamp region 1 comprises a complementary sequence encoding a sequence of an exon or a variant thereof.

122. The composition of claim 93, wherein the connecting clamp region 1 comprises a complementary sequence encoding the sequences of exons and introns.

123. The composition of claim 93, wherein the connecting clip region 1 contains nucleobases mismatched with splice acceptor sites in the sequence of the target nucleic acid.

124. The composition of claim 93, wherein the ribonucleotide is located at the 3' end of the connecting clamp region 1.

125. The composition of claim 93, wherein the connecting plate region 1 comprises 1 ribonucleotide, 2 ribonucleotides, 3 ribonucleotides, 4 ribonucleotides, 5 ribonucleotides, or up to 10 ribonucleotides.

126. The composition of claim 93, wherein the length of the connecting clamp region 1 is at least about 5 nucleotides to at most 10 nucleotides.

127. The composition according to any one of claims 93 to 126 further comprises an engineered protein or a polypeptide encoding said engineered protein.

128. The composition of claim 127, wherein the engineered protein comprises: (a) the nicking enzyme region; and (b) DNA ligase region.

129. The composition of claim 128, wherein the DNA ligase region binds to the linker region 1 or the linker region 2 of the guiding polynucleotide.

130. The composition of claim 128, wherein the DNA ligase region ligates the donor nucleic acid to the 3' end of the leader strand.

131. The composition of claim 128, wherein the nicking enzyme region comprises an engineered Cas protein that generates single-strand breaks in the target nucleic acid.

132. The composition of claim 132, wherein the engineered Cas protein comprises: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas11, Cas12, Cas13, Cas14, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, C mr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, c2c1, c2c3, Cas9HiFi, xCas9, CasX, CasY, CasRX, variants or any combination thereof.

133. The composition according to any one of claims 128 to 132, further comprising a delivery medium.

134. The composition of claim 133, wherein the delivery medium comprises a carrier, lipid, nanoparticle, plasmid, virus, liposome, extracellular vesicle, emulsion, peptide, sugar, polymer, chitosan, polyethyleneimine (PEI), poly(lactide-co-glycolic acid) (PLGA), poly-L-lysine (PLL), or a combination thereof.

135. A cell comprising the composition according to any one of claims 128 to 134.

136. A method for linking donor nucleic acid to target nucleic acid, the method comprising: Expose cellular or cell-free systems to the following: (a) Donor nucleic acid; (b) a guiding polynucleotide or a polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) Target regions that are complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region comprises a secondary structure for binding to a nuclease; and (iii) Connecting clamp region 2, wherein the connecting clamp region 2 is complementary to the donor nucleic acid and is complementary to the target nucleic acid and has at least one change relative to the target nucleic acid; (iv) Connecting splice 1 region, wherein the connecting splice 1 region comprises: deoxyribonucleotides and ribonucleotides; (c) an engineered protein or a polynucleotide encoding said engineered protein, wherein said engineered protein comprises: (i) nuclease region; and (ii) DNA ligase region; in: The guiding polynucleotide forms a complex with the engineered protein via the protein-binding region. The nuclease region of the engineered protein breaks down in the target nucleic acid to produce a leading strand and a complementary strand. The targeting sequence forms a complex with the complementary strand. The connecting clamp 1 region forms a complex with the leading chain, and The DNA ligase region therein attaches the donor nucleic acid to the leader strand, thereby linking the donor nucleic acid to the target nucleic acid.

137. The method of claim 136, wherein the target chain is dissociated from the complementary chain.

138. The method of claim 136, wherein the novel nucleic acid is incorporated into the target nucleic acid by hybridization with the complementary strand.

139. The method of claim 136, wherein the incorporation of the new nucleic acid recruits a DNA repair protein to the target nucleic acid, the DNA repair protein editing the nucleobases of the target nucleic acid.

140. A method comprising: The administration of a system according to any one of claims 47 to 87, a composition according to any one of claims 1 to 24 or 93 to 134, or a cell according to claim 92 or 135 to a subject, organ, tissue, or cell, wherein the administration modifies genes in the subject, organ, tissue, or cell.

141. The method of claim 140, wherein the alteration of the gene comprises: Insertion, deletion, substitution, copy number change, point mutation, frameshift mutation, missense mutation, nonsense mutation, stop codon mutation, epigenetic marker, or any combination thereof.

142. The method of claim 140, wherein the application is local or systemic.

143. The method of claim 140, wherein the application is intranasal, subcutaneous, intravenous, inhaled, intramuscular, intratumoral, peritumoral, intrathecal, vaginal, or intradermal.

144. The method of claim 140, wherein the subject is a mammal.

145. The method of claim 140, wherein the subject has, is suspected of having, or has been diagnosed with a disease or condition.

146. The method of claim 145, wherein the disease or condition includes a genetic disease or condition.

147. The method according to any one of claims 140 to 146, the method further comprising administering a therapeutic agent to the subject.

148. A method comprising: Contacting cells or cell populations with the system according to any one of claims 47 to 87 or the composition according to any one of claims 1 to 24 or 93 to 134, thereby producing alterations in the target nucleic acid.

149. The method of claim 148, wherein the contact is performed in vitro, in vivo, or ex vivo.

150. The method of claim 148, wherein the change comprises: Insertion, deletion, substitution, copy number change, point mutation, frameshift mutation, missense mutation, nonsense mutation, stop codon mutation, epigenetic marker, or any combination thereof.

151. The method of claim 148, wherein the alteration restores the expression of the wild-type protein encoded by the gene relative to equivalent cells or cell populations that have not been in contact with the system or the composition.

152. The method of claim 148, wherein the cell comprises a eukaryotic cell or a prokaryotic cell.

153. The method of claim 148, wherein the cell comprises a mammalian cell.

154. The method of claim 148, wherein the cell population comprises a human leukocyte population, a stem cell population, or a bacterial population.

155. A cell population prepared by the method according to any one of claims 148 to 154.

156. A kit comprising: a system according to any one of claims 47 to 87 or a composition according to any one of claims 1 to 24 or 93 to 134, packaging and materials for use therein.

157. A reagent kit comprising: The first container contains: Donor nucleic acid, and The guiding polynucleotide or the polynucleotide encoding the guiding polynucleotide, wherein the guiding polynucleotide comprises: (i) A target region that is complementary to the target nucleic acid; (ii) a protein-binding region, wherein the protein-binding region contains a secondary structure for binding to a nuclease; (iii) A connecting clamp region 2, wherein the connecting clamp region 2 is complementary to the donor nucleic acid and to the target nucleic acid, and wherein the connecting clamp region 2 contains at least one mismatched nucleobase relative to the target nucleic acid; and (iv) Connecting splice region 1, wherein the connecting splice region 1 comprises: deoxyribonucleotides and ribonucleotides, and The second container contains: An engineered polypeptide comprising a nicking enzyme operatively linked to a DNA ligase.

158. The kit according to claim 157 further comprises reagents for nucleic acid amplification, transcription, translation, or nucleic acid isolation.

159. The kit according to any one of claims 157 or 158 further comprises a reporter molecule.

160. A support, the support comprising: The system according to any one of claims 47 to 87 or the composition according to any one of claims 1 to 24 or 93 to 134; and a solid surface, wherein the system or the composition is fixed on the solid surface.

161. The scaffold of claim 160, wherein the scaffold comprises a reaction chip, paper, quartz microfibers, a mixture of cellulose esters, porous alumina, a patterned surface, a tube, a pore, or a matrix.

162. The stent of claim 160 or claim 161, wherein the stent further comprises a reporter molecule that generates a detectable signal detectable by a detector.