Class 2, V-type CRISPR system

JP2026509249A5Pending Publication Date: 2026-03-27METAGENOMI THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-08
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems are insufficient in terms of specific targeting and nucleic acid editing efficiency, making it difficult to efficiently modify specific gene sequences.

Method used

An engineered nuclease system was developed, comprising a nuclease with high sequence homology and a guide polynucleotide, which can form a complex with the target nucleic acid sequence and bind specifically to DNA methyltransferase to achieve precise nucleic acid modification.

Benefits of technology

It enables efficient modification of specific gene sequences, improves the accuracy and efficiency of nucleic acid editing, and is applicable to gene modification in various cell types.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Methods, compositions, and systems derived from uncultured microorganisms useful for gene editing are described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application claims the benefits and priority of U.S. Provisional Patent Application No. 63 / 489,162 filed on 8 March 2023, U.S. Provisional Patent Application No. 63 / 504,422 filed on 25 May 2023, and U.S. Provisional Patent Application No. 63 / 587,655 filed on 3 October 2023, each of which is incorporated herein by reference in its entirety. [Overview of the project]

[0002] In certain embodiments, an engineered nuclease system is described herein, comprising: a) an endonuclease containing a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1-325, 420-431, 476-624, 629, 1065-1090, 1114-1118, and 1746-1752; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease contains a sequence having at least 90% sequence identity with any one of SEQ ID NOs: 1-325, 420-431, 476-624, 629, 1065-1090, 1114-1118, and 1746-1752. In some embodiments, the endonuclease contains a sequence having 100% sequence identity with any one of SEQ ID NOs: 1-325, 420-431, 476-624, 629, 1065-1090, 1114-1118, and 1746-1752. In some embodiments, the manipulated guide polynucleotide contains crRNA and tracrRNA. In some embodiments, the manipulated guide polynucleotide contains SEQ ID NOs: 333-335, 355-357, 410-411, 346-347, 368-369, 412-413, 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 44 The sequence contains at least 90% sequence identity with any one of the following: 2, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1091-1113, 1119-1120, 1697-1731, and 1876.In some embodiments, the manipulated guide polynucleotides are SEQ ID NOs: 333-335, 355-357, 410-411, 346-347, 368-369, 412-413, 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, The sequence includes sequences having 100% sequence identity with any one of 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1091-1113, 1119-1120, 1697-1731, and 1876. In some embodiments, the manipulated guide polynucleotide is a single guide nucleic acid. In some embodiments, the manipulated guide polynucleotide is a dual guide nucleic acid. In some embodiments, the manipulated guide polynucleotide is RNA. In some embodiments, the endonuclease is non-covalently bound to the manipulated guide polynucleotide. In some embodiments, the endonuclease is covalently bound to the manipulated guide polynucleotide. In some embodiments, the endonuclease is fused to the manipulated guide polynucleotide. In some embodiments, the endonuclease system further comprises a DNA methyltransferase. In some embodiments, the DNA methyltransferase is non-covalently bonded to the endonuclease. In some embodiments, the DNA methyltransferase is fused to the endonuclease in a single polypeptide. In some embodiments, the DNA methyltransferase comprises Dmnt3A or Dnmt3L.

[0003] In a particular embodiment, an engineered nuclease system is described herein, comprising: a) an endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 6 to 14; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 333 to 335 and 355 to 357.

[0004] In a particular embodiment, an engineered nuclease system is described herein, comprising: a) an endonuclease comprising a sequence having at least 80% sequence identity with sequence number 15; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 410 to 411.

[0005] In a particular embodiment, an engineered nuclease system is described herein, comprising: a) an endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 16-29; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 346-347, 368-369, and 412-413.

[0006] In a particular embodiment, a) an endonuclease comprising a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises SEQ ID NOs: 326-332, 336-345, 348-354, 3 This specification describes an engineered nuclease system comprising an engineered guide polynucleotide having at least 80% sequence identity with any one of 58-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876.

[0007] In certain embodiments, an engineered nuclease system is described herein, comprising: a) an endonuclease containing a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1065-1090, 1114-1118, and 1746-1752; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1091-1113, 1119-1120, and 1876. In some embodiments, the engineered guide polynucleotide is a single guide nucleic acid. In some embodiments, the engineered guide polynucleotide is a dual guide nucleic acid. In some embodiments, the engineered guide polynucleotide is RNA. In some embodiments, the endonuclease is non-covalently bonded to the engineered guide polynucleotide. In some embodiments, the endonuclease is covalently bonded to the engineered guide polynucleotide. In some embodiments, the endonuclease is fused to an engineered guide polynucleotide. In some embodiments, the endonuclease system further comprises a DNA methyltransferase. In some embodiments, the DNA methyltransferase is non-covalently bonded to the endonuclease. In some embodiments, the DNA methyltransferase is fused to the endonuclease in a single polypeptide. In some embodiments, the DNA methyltransferase comprises Dmnt3A or Dnmt3L.

[0008] In certain embodiments, a method for modifying a target nucleic acid sequence is described herein, comprising contacting the target nucleic acid sequence with an engineered nuclease system described herein. In some embodiments, modifying a target nucleic acid sequence involves binding, nicking, or cleaving the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence includes genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the modification is in vitro. In some embodiments, the modification is in vivo. In some embodiments, the modification is ex vivo.

[0009] In certain embodiments, a method for modifying a target nucleic acid sequence in mammalian cells is described herein, comprising contacting the mammalian cells with an engineered nuclease system described herein. In some embodiments, the method further includes selecting cells containing the modification.

[0010] In certain embodiments, a method for modifying TRAC is described herein, comprising contacting TRAC with an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 767-798. In some embodiments, the target nucleic acid sequence comprises a sequence having any one of SEQ ID NOs: 799-830.

[0011] In a particular embodiment, a method for modifying APOA1 is described herein, comprising contacting APOA1 with an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the target nucleic acid sequence includes a sequence having one of the sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676.

[0012] In a particular embodiment, a method for modifying AAVS1, comprising contacting APOA1 using an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the target nucleic acid sequence includes a sequence having one of the sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806.

[0013] In a particular embodiment, a method for modifying albumin, comprising contacting albumin with an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the target nucleic acid sequence includes a sequence having any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994.

[0014] In certain embodiments, nucleic acids encoding the manipulated nuclease system described herein are described herein.

[0015] In certain embodiments, cells comprising the engineered nuclease system described herein are described herein. In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are immortalized cells. In some embodiments, the cells are insect cells. In some embodiments, the cells are yeast cells. In some embodiments, the cells are plant cells. In some embodiments, the cells are fungal cells. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the cells are engineered cells. In some embodiments, the cells are stable cells. In some embodiments, the cells are primary cells. In some embodiments, the primary cells are T cells. In some embodiments, the primary cells are hematopoietic stem cells (HSCs). [Brief explanation of the drawing]

[0016] Novel features of this disclosure are specifically described in the appended claims. A better understanding of the features and advantages of this disclosure will be obtained by referring to the following detailed description illustrating exemplary embodiments in which the principles of this disclosure are utilized, and to the appended drawings (also referred to herein as "Figure" and "FIG.").

[0017] [Figure 1] Prior to this disclosure, a typical organization of the different classes and types of CRISPR / Cas loci described above is shown. [Figure 2A]An overview of the MG119 family is shown. Figure 2A shows multiple alignments of representative MG119 effectors, illustrating the domain composition and the conservation of RuvC catalytic residues important for double-strand DNA cleavage activity. Figure 2A discloses MG119-3 (SEQ ID NO: 32), MG119-4 (SEQ ID NO: 33), MG119- (SEQ ID NO: 31), MG119-1 (SEQ ID NO: 30), and MG119-5 (SEQ ID NO: 34). [Figure 2B] An overview of the MG119 family is shown. Figure 2B depicts a CRISPR-containing contig with a genomic context surrounding a CRISPR array and Cas effector (example of MG119-1). [Figure 2C] This shows an overview of the MG119 family. Figure 2C shows the folding of a direct iteration of MG119-1 (sequence number 1934). [Figure 2D] An overview of the MG119 family is shown. Figure 2D shows the single guide RNA (SEQ ID NO: 1935) designed for MG119-1. [Figure 3A] An overview of the MG90 family is presented. Figure 3A shows multiple alignments of representative MG90 effectors, illustrating domain composition and the conservation of RuvC catalytic residues important for double-strand DNA cleavage activity. Figure 3A discloses SEQ ID NOs. 26 (MG90-21), 20 (MG90-8), 17 (MG90-5), 23 (MG90-18), and 16 (MG90-3). [Figure 3B] An overview of the MG90 family is shown. Figure 3B depicts a CRISPR-containing contig with a genomic context surrounding a CRISPR array and Cas effector (example of MG90-5). [Figure 3C] An overview of the MG90 family is shown. Figure 3C shows the folding of a direct repeat of MG90-5. [Figure 4A]Shows an overview of the MG126 family. Figure 4A shows a multiple alignment of representatives of the MG126 effector, showing domain composition and conservation of the RuvC catalytic residues important for the function of double-strand DNA cleavage activity. Figure 4A discloses MG126-3 (SEQ ID NO: 319), MG126-4 (SEQ ID NO: 321), and MG126-7 (SEQ ID NO: 324). [Figure 4B] Shows an overview of the MG126 family. Figure 4B shows a depiction of a CRISPR-containing contig having a genomic context surrounding the CRISPR array and the Cas effector (example of MG126-4). [Figure 4C] Shows an overview of the MG126 family. Figure 4C shows the folding of the direct repeats (SEQ ID NO: 1936) of MG126-4. [Figure 5A] Shows an overview of the MG118 family. Figure 5A shows a multiple alignment of representatives of the MG118 effector, showing domain composition and conservation of the RuvC catalytic residues important for the function of double-strand DNA cleavage activity. Figure 5A discloses SEQ ID NO: 15 (MG118-1). [Figure 5B] Shows an overview of the MG118 family. Figure 5B shows a depiction of a CRISPR-containing contig having a genomic context surrounding the CRISPR array and the Cas effector (example of MG118-1). [Figure 5C] Shows an overview of the MG118 family. Figure 5C shows the folding of the direct repeats (SEQ ID NO: 1937) of MG118-1. [Figure 6A] Shows an overview of the MG122 family. Figure 6A shows a multiple alignment of representatives of the MG122 effector, showing domain composition and conservation of the RuvC catalytic residues important for the function of double-strand DNA cleavage activity. Figure 6A discloses SEQ ID NO: 3 (MG122-3), SEQ ID NO: 4 (MG122-4), and SEQ ID NO: 5 (MG122-5). [Figure 6B] Shows an overview of the MG122 family. Figure 6B shows a depiction of a CRISPR-containing contig having a genomic context surrounding the CRISPR array and the Cas effector (example of MG122-4). [Figure 6C]An overview of the MG122 family is shown. Figure 6C shows the folding of a direct iteration of MG122-4 (sequence number 1938). [Figure 7A] An overview of the MG120 family is shown. Figure 7A shows multiple alignments of representative MG120 effectors, illustrating the domain composition and the conservation of RuvC catalytic residues important for double-strand DNA cleavage activity. Figure 7A discloses SEQ ID NOs. 6 (MG120-1), 7 (MG120-2), 8 (MG120-3), 9 (MG120-4), 10 (MG120-5), 11 (MG120-6), 12 (MG120-7), and 14 (MG120-9). [Figure 7B] An overview of the MG120 family is shown. Figure 7B depicts a CRISPR-containing contig with a genomic context surrounding a CRISPR array and Cas effector (example of MG120-1). [Figure 7C] An overview of the MG120 family is shown. Figure 7C shows the folding of a direct iteration of MG120-1 (sequence number 1939). [Figure 8A] An overview of the MG91 family is shown. Figure 8A depicts a CRISPR-containing contig with a genomic context surrounding a CRISPR array and Cas effector (example: MG91B-24). [Figure 8B] This provides an overview of the MG91 family. Figure 8B shows the folding of a direct iteration (sequence number 1940) of MG91B-24. [Figure 8C] An overview of the MG91 family is shown. Figure 8C depicts a CRISPR-containing contig with a genomic context surrounding a CRISPR array and Cas effector (example: MG91C-10). [Figure 8D] An overview of the MG91 family is shown. Figure 8D shows the folding of a direct iteration of MG91C-10 (sequence number 1941). [Figure 9]This study demonstrates the in vitro activity of MG119-2 using a TXTL assay. MG119-2 was tested for dsDNA cleavage using two intergene sequences from an MG119-2 contig, a minimal array (MA) sequence containing forward or reverse-oriented repeats, and a PAM library target plasmid. In lane 1, positive intergene enrichment was observed as amplified cleavage products containing intergene (IG) sequence 1 and a minimal array with forward-oriented repeats. Lanes 3 and 7 are negative controls with IG omitted, and lane 4 is a third negative control with both array and IG omitted. [Figure 10A] The sequence logo for MG119-2 PAM(5'-nTnn-3'), determined via next-generation sequencing (NGS) of the cleavage product obtained from the in vitro cleavage assay, is shown. [Figure 10B] The histogram of the transection site (23 bd away from PAM) is shown. [Figure 11A] Examples of active MG119 nucleases and their sgRNA designs are shown. Figure 11A shows the predicted folding of a single guide RNA sequence without spacers. The blue circle represents the first 5' nucleotide of the tracrRNA, and the red circle represents the 3' nucleotide of the repeat. The tracrRNA and repeat sequence are looped in a GAAA tetraloop. The repeat-anti-repeat fold is at the 3' end of each structure. Three different RNA structures of active guides within the same family are shown. From left to right: The MG119-28 guide has four hairpins, three smaller hairpins at the 5' end, and a very long hairpin with two bulges next to the repeat-anti-repeat fold. The MG119-83 sgRNA has three small hairpins, and the repeat-anti-repeat has two bulges. MG119-118 has four hairpins, the second hairpin branching into three hairpins from its 5' end, while the third hairpin and the repeat / anti-repeat have one bulge. This guide also has several pairing nucleotides between the 5' end of the tracr and the 3' end of the repeat. [Figure 11B]Examples of active MG119 nucleases and their sgRNA designs are shown. Figure 11B shows in vitro cleavage assay amplification products on a 2% agarose gel. The low molecular weight DNA ladder is located in lanes 1, 7, and 11. Contents of other lanes from left to right: (2) MG119-28 nuclease only, MG119-28 nuclease plus (3) sgRNA1 with U67 spacer, (4) sgRNA1 with U40 spacer, (5) sgRNA2 with U67 spacer, and (6) sgRNA2 with U40 spacer, (8) MG119-83 nuclease only, MG119-83 nuclease plus (9) sgRNA1 with U67 spacer, and (10) sgRNA1 with U40 spacer, (12) MG119-118 nuclease only, MG119-118 nuclease plus (13) sgRNA1 with U67 spacer, and (14) sgRNA1 with U40 spacer. The resulting amplicon product is 188 bp for the U67 spacer carrying guide, or 205 bp for the U40 spacer carrying guide. [Figure 12-1] The sequence logo for the protospacer adjacent motif (PAM) of the active MG119 nuclease is shown. [Figure 12-2] The sequence logo for the protospacer adjacent motif (PAM) of the active MG119 nuclease is shown. [Figure 13A] Figure 13A shows an exemplary SDS-PAGE gel and size exclusion chromatography (SEC) A280 trace of the protein purification steps. Figure 13A shows (1) dissolution after sonication, (2) centrifugation after clarification, (3) Ni-NTA gravity column flow-through, (4) eluate from Ni-NTA resin, and (5) MG119-28△ purification using the sample recovered from the concentrated sample. [Figure 13B] Figure 13B shows an exemplary SDS-PAGE gel and size exclusion chromatography (SEC) A280 trace from the protein purification step. The S200i 10 / 300 GL column SEC A280 trace was shown. Peak fractions were pooled and concentrated. [Figure 13C]Figures 13C and 13D show exemplary SDS-PAGE gels and size exclusion chromatography (SEC) A280 traces of the protein purification steps. (1) Dissolution after sonication, (2) Centrifugation after clarification, (3) Ni-NTA gravity column flow-through, (4) Eluten from Ni-NTA resin, (5) Concentrated protein, (6) Concentrated protein cleaved overnight with TEV protease, (7) and aggregates centrifuged (21,000 × g, 4°C, 10 min) for pelletization, (8) Amylose column flow-through, (9) Flow-through centrifuged for pelletization of aggregates (21,000 × g, 4°C, 10 min), and (10) Purification of MBP-tagged / cleaved MG119-28Δ using samples recovered from the concentrated flow-through. [Figure 13D] Figures 13C and 13D show exemplary SDS-PAGE gels and size exclusion chromatography (SEC) A280 traces of the protein purification steps. (1) Dissolution after sonication, (2) Centrifugation after clarification, (3) Ni-NTA gravity column flow-through, (4) Eluten from Ni-NTA resin, (5) Concentrated protein, (6) Concentrated protein cleaved overnight with TEV protease, (7) and aggregates centrifuged (21,000 × g, 4°C, 10 min) for pelletization, (8) Amylose column flow-through, (9) Flow-through centrifuged for pelletization of aggregates (21,000 × g, 4°C, 10 min), and (10) Purification of MBP-tagged / cleaved MG119-28Δ using samples recovered from the concentrated flow-through. [Figure 13E] Figure 13E shows an exemplary SDS-PAGE gel and size exclusion chromatography (SEC) A280 trace from the protein purification step. [Figure 13F]Figure 13F shows exemplary SDS-PAGE gel and size exclusion chromatography (SEC) A280 traces from the protein purification step. Of the five MG119 candidates expressed in both pMGB and pMGB△ expression vectors, all showed higher yields in the pMGB△ vector. [Figure 14A] An example of in vitro cleavage efficiency using purified protein is shown. Figure 14A shows an agarose gel illustrating titration of RNP:substrate ratios and increased substrate cleavage at higher ratios. [Figure 14B] An example of in vitro cleavage efficiency using purified protein is shown. Figure 14B shows the percentage of cleaved substrate determined for each lane using density measurements. The cleavage fractions were plotted in Prism 8, and the protein activity fraction was calculated using the slope of the linear range of cleavage. This assay used MG119-28 expressed with a pMGB△ backbone. [Figure 15A] Examples of in vitro cleavage and editing efficiency of mouse Hepa1-6 cell DNA are shown. Figure 15A shows the cleavage percentage of MG119-28 using four chemically modified guides targeting the mouse albumin gene at intron 1 (Table 6). Two concentrations of nuclease were tested: 15.6 nM (black bars) and 7.8 nM (white bars). Cleavage was normalized to an untargeted control. MG119-28 can cleave Hepa1-6 gDNA up to an average of 60% at sgRNA4 in 15.6 nM RNP and up to 33% in 7.8 nM RNP. [Figure 15B] This shows examples of in vitro cleavage and editing efficiency of mouse Hepa1-6 cell DNA. Figure 15B shows the percentage of indels generated by MG119-28 in Hepa1-6 cells normalized to the apo reaction. Each condition was performed three times (triplicate). An average of 25.12% of the sequenced reads were edited with sgRNA3. sgRNA3 is consistently active in vitro and in cells, as shown here. The next best guide in cells is sgRNA4, with an average editing rate of 4.11%. The observed edits are mainly 4–24 bp deletions. [Figure 16A] This shows the genomic context of representative Cas nuclease genes in the MG191 family. Figure 16A shows the CRISPR array (indicated by repeats and spacers) and the nuclease gene (dark gray arrow). [Figure 16B] This shows the genomic context of representative Cas nuclease genes in the MG191 family. Figure 16B shows an example of a crRNA fold. [Figure 17A] Examples of proteins purified without the fusion protein (SUMO-MG119-1) (MG119-1△) and proteins purified with the fusion protein (SUMO-MG119-1) are shown. Figure 17A: MG119-1 (approximately 56 kDa) can be expressed and purified without the fusion protein, but the resulting sample has a relatively low yield and is impure. [Figure 17B] Examples of proteins purified without the fusion protein (SUMO-MG119-1) (MG119-1△) and proteins purified with the fusion protein (SUMO-MG119-1) are shown. Figure 17B: Inclusion of the SUMO domain (SUMO-MG119-1, approximately 63 kDa) as the N-terminal fusion protein increases both the expression yield and purity of the sample. [Figure 18A] (Figure 18A) shows the sequence logos of the protospacer adjacent motif (PAM) of nuclease MG119-137 from the target strand and (Figure 18B) from the non-target strand. [Figure 18B] (Figure 18A) shows the sequence logos of the protospacer adjacent motif (PAM) of nuclease MG119-137 from the target strand and (Figure 18B) from the non-target strand. [Figure 19] An in vitro cleavage assay for determining spacer length preference is shown. Agarose gel containing cleavage products including linearized plasmids from dsDNA cleavage and nicked plasmids from ssDNA cleavage. [Figure 20A]An example of single-guide engineering (MG119-2 sgRNA2) is shown. Figure 20A shows round 1 guide engineering. Truncations for sequence-optimized WT sgRNA were designed, complexed with a constant ratio of sgRNA:effector, and tested with a constant ratio of RNP:substrate DNA. Cleavage was measured via densitometry and normalized by the degree of cleavage from non-truncated sequences. In this assay, sgRNA2_4 and sgRNA2_6 successfully truncated (≥80% cleavage from non-truncated sequences). [Figure 20B] An example of single-guide engineering (MG119-2 sgRNA2) is shown. Figure 20B shows round 2 guide engineering. Further truncation of the MG119-2 sgRNA2 guide was performed by combining sgRNA2_4 and sgRNA2_6 deletions with additional truncation. The shortest guide that satisfies a relative cleavage threshold of ≥80% is sgRNA2_4.6.11(96nt). [Figure 21A] Figure 21A: This shows the percentage of indels in K562 cells. The percentage represents the proportion of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions. Each bar represents the indel % of the indicated MG119-2 guide targeting exon 3 of the human TRAC gene. Black bars represent data from two independent replications of mRNA+sgRNA nucleofection. Gray bars represent data from a single RNP nucleofection. [Figure 21B] Figure 21B: This shows the percentage of indels in K562 cells. The percentage represents the proportion of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions. Each bar represents the indel % of the indicated MG119-125 guide targeting exon 3 of the human APOA1 gene. Displayed in descending order of indel %. [Figure 21C]Figure 21C: Percentage of indels in K562 cells. The percentage represents the proportion of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions. Each bar represents the percentage of indels for the indicated MG119-129 guide targeting exon 3 of the human APOA1 gene. [Figure 22A] Figure 22A: MG119-2 activity at hAPOA1 exon 3, hALB intron 1, and AAVS1 is shown as a percentage of indels in K562 cells. The percentage of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions is shown. Black bars represent data from mRNA nucleofection (n=1). Figure 22A: Each bar represents the percentage of indels of an MG119-2 guide targeting exon 3 of the human APOA gene. [Figure 22B] Figure 22B: MG119-2 activity at hAPOA1 exon 3, hALB intron 1, and AAVS1 is shown as a percentage of indels in K562 cells. The percentage of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions is shown. Black bars represent data from mRNA nucleofection (n=1). Figure 22B: Each bar represents the percentage of indels of an MG119-2 guide targeting intron 1 of the human ALB gene. [Figure 22C] Figure 22C: MG119-2 activity in hAPOA1 exon 3, hALB intron 1, and AAVS1 is shown as a percentage of indels in K562 cells. The percentage of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions is shown. Black bars represent data from mRNA nucleofection (n=1). Figure 22C: Each bar represents the percentage of indels of an MG119-2 guide targeting the AAVS1 gene. [Figure 23A]Figure 23A: MG119-28 activity at hAPOA1 exon 3, hALB intron 1, and AAVS1 is shown as a percentage of indels in K562 cells. The percentage of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions is shown. Black bars represent data from mRNA nucleofection (n=1). Figure 23A: Each bar represents the indel % of the indicated MG119-28 guide targeting exon 3 of the human APOA gene. [Figure 23B] Figure 23B: MG119-28 activity at hAPOA1 exon 3, hALB intron 1, and AAVS1 is shown as a percentage of indels in K562 cells. The percentage of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions is shown. Black bars represent data from mRNA nucleofection (n=1). Figure 23B: Each bar represents the indel % of the indicated MG119-28 guide targeting intron 1 of the human ALB gene. [Figure 23C] Figure 23C: MG119-28 activity in hAPOA1 exon 3, hALB intron 1, and AAVS1 is shown as a percentage of indels in K562 cells. The percentage of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions is shown. Black bars represent data from mRNA nucleofection (n=1). Figure 23C: Each bar represents the indel % of the indicated MG119-28 guide targeting the AAVS1 gene. [Figure 24] The graph shows MG119-32 activity in hALB intron 1, hAPOA1 exon 3, and AAVS1, expressed as a percentage of indels in K562 cells. The percentage represents the proportion of amplicons (indel %) from NGS amplicon sequencing containing insertions or deletions. Each bar represents the indel % of the indicated MG119-32 guide targeting intron 1 of the human ALB gene, exon 3 of the human APOA gene, and the AAVS1 gene. Black bars represent data from RNP nucleofection (n=2). [Figure 25A] This shows the engineering of MG119-2 splitting guides for 119-2_sg2_WT. Figure 25A: Five splitting guides for MG119-2 guides for 119-2_sg2_WT were designed and synthesized using two fragments, namely fragment "A" (5' half) and fragment "B" (3' half). In some cases (v4 and v5), small portions of the guide were excluded in the final annealed construct. The final designs were mapped to the reference sequence (sequence number 655) using the parental sgRNA. [Figure 25B] This shows the MG119-2 splitting guide engineering for 119-2_sg2_WT. Figure 25B shows an agarose gel with the RNP:substrate ratio, and the corresponding graph shows the percentage of substrate cleaved by each splitting guide version, determined for each lane using densitometry. [Figure 26A] This shows the MG119-2 splitting guide engineering for 119-2_sg2_4.6.11. Figure 26A: Five splitting guides were designed for MG119-2_sg2_4.6.11 and synthesized using two fragments, namely fragment "A" (5' half) and fragment "B" (3' half). In some cases (v4 and v5), small portions of the guide were excluded in the final annealed construct. The final designs were mapped to the reference sequence (sequence number 668) using the parental sgRNA. [Figure 26B] This shows the MG119-2 splitting guide engineering for 119-2_sg2_4.6.11. Figure 26B shows an agarose gel with RNP:substrate ratios, and the corresponding graph shows the percentage of substrate cleaved by each splitting guide version, determined for each lane using densitometry. Designs v3 and v4 have activity equivalent to the sgRNA parent. [Figure 27A]This shows the engineering of MG119-28 splitting guides for 119-28_sg1_WT. Figure 27A: Four splitting guides for MG119-28 guides for 119-28_sg1_WT were designed and synthesized using two fragments, namely fragment "A" (5' half) and fragment "B" (3' half). In some cases (v2 and v3), small portions of the guide were excluded in the final annealed construct. The final designs were mapped to the reference sequence (sequence number 686) using the parental sgRNA. [Figure 27B] This shows the MG119-28 splitting guide engineering for 119-28_sg1_WT. Figure 27B shows an agarose gel with RNP:substrate ratios, and the corresponding graph shows the percentage of substrate cleaved by each splitting guide version, determined for each lane using densitometry. Activity testing of the splitting 119-28_sg1_WT guides shows that splitting guides v1, v3, and v4 are highly active compared to the sgRNA parent. [Figure 28A] This shows the engineering of MG119-28 splitting guides for 119-28_sg1_8.5. Figure 28A: Four splitting guides were designed for MG119-28_sg1_8.5 and synthesized using two fragments, namely fragment "A" (5' half) and fragment "B" (3' half). In some cases (v2 and v3), small portions of the guide were excluded in the final annealed construct. The final designs were mapped to the reference sequence (SEQ ID NO: 704) using the parental sgRNA. [Figure 28B] This shows the MG119-28 splitting guide engineering for 119-28_sg1_8.5. Figure 28B shows an agarose gel showing the RNP:substrate ratio, and the corresponding graph shows the percentage of substrate cleaved by each splitting guide version, determined for each lane using densitometry. Activity testing of the splitting 119-28_sg1_8.5 guides shows that splitting guides v1, v3, and v4 are highly active compared to the sgRNA parent. [Figure 29A]This shows the optimization of the guide length for MG119-32 sgRNA1 round 1. Figure 29A: MG119-32 sg1 guide truncation design aligned to untrimmed WT sg1 sgRNA. [Figure 29B] Figure 29B shows the optimization of the guide length for MG119-32 sgRNA1 round 1. The relative activity of sg1 guide truncation is shown. The data suggest that designs 5, 6, 7, and 8 have better or equivalent activity than untrimmed WT sg1 sgRNA. [Figure 29C] This shows the optimization of the guide length for MG119-32 sgRNA2 round 1. Figure 29C: MG119-32 sg2 guide truncation design aligned to untrimmed WT sg2 sgRNA. [Figure 29D] This shows the optimization of the guide length for MG119-32 sgRNA2 round 1. Figure 29D shows the relative activity of sg2 guided truncations. The data suggest that designs 6, 7, and 8 have better or equivalent activity than untrimmed WT sg2 sgRNA. [Figure 30A] This shows the optimization of the guide length for MG119-32 sgRNA1 Round 2. Figure 30A: Successful guide from Round 1 and new Round 2 MG119-32 sg1 guide truncation design aligned to untrimmed WT sg1 sgRNA. [Figure 30B] This shows the optimization of the guide length for MG119-32 sgRNA1 round 2. Figure 30B shows the relative activity of sg1 guide truncations, including the successful guide from round 1 (light gray) and the new guide from round 2 (dark gray). The data suggest that 119-32_sg1_5.6 and 119-32_sg1_5.7 have better or equivalent activity than untrimmed WT sg1 sgRNA. [Figure 31A]The monitoring of nicking activity throughout the guide engineering process is shown. Figure 31A: Shows the cleavage reactions for MG119-2 and guide constructs (SEQ ID NOs. 434, 655, 659, 663, and 668). [Figure 31B] Monitoring of nicking activity throughout the guide engineering process is shown. Figure 31B: Shows the cleavage reactions for MG119-28 and guide constructs (SEQ ID NOs. 686, 691, 703, and 704). For both the effector and guide panels, the intensity of the band corresponding to the nicked product does not change significantly compared to WT sgRNA. [Figure 32] This shows a phylogenetic tree of representative Cas12 nucleases (600-1100aa) that highlight the novel MG family. [Figure 33] The image shows the amplification products of the in vitro cleavage assay by automated gel electrophoresis. An example of the active MG191 protein successfully cleaved from the PAM library produced a band of approximately 205 bp. [Figure 34] The sequence logo of the protospacer adjacent motif (PAM) obtained from NGS sequencing of the amplified cleavage site is shown. [Figure 35A] Examples of purification and activity analysis of MG191-15 are shown. Figure 35A: Protein expression induction and purification were monitored by gel electrophoresis. The predicted protein MW was approximately 136 kDa. [Figure 35B] Examples of purification and activity analysis of MG191-15 are shown. Figure 35B: Concentrated protein was passed through an S200i 10 300 SEC column, the peak fraction was collected and concentrated (shaded box). [Figure 35C] Examples of purification and activity analysis of MG191-15 are shown. Figure 35C: The normalized protein content of each protein preparation was run on an SDS-PAGE protein gel for purification analysis via densitometry. [Figure 35D]Examples of purification and activity analysis of MG191-15 are shown. Figure 35D: Protein activity was evaluated by in vitro cleavage using linear DNA as the substrate. Lanes: (1) Substrate only, (2) Substrate + Apo (unguided) protein, (3) Substrate + 5 × molar excess RNP, (4) Substrate + 25 × molar excess RNP. [Figure 35E] Examples of purification and activity analysis of MG191-15 are shown. Figure 35E: Automated electrophoresis gel showing nuclease activity for MG191-18, MG191-25, and MG191-29. Activity was detected by the presence of a cleaved product band amplified by approximately 250 bp (black arrow) in the presence of precRNA (+MA). [Figure 36] This study demonstrates the intracellular activity of MG119-28 using truncated sgRNA. The MG119-28 truncated guide was screened in K562 cells for activity at nine sites targeting hAAVS1. mRNA and the guide were transfected into cells via nucleofection. After 3 days, the edited sites were amplified for NGS and sequenced for indel analysis. All conditions were repeated twice. [Figure 37] This paper describes dose titrations of RNPs to test the best truncate scaffold efficacy. MG119-28 WT (black circles) and sg1_8 truncate guide (triangles) RNPs were screened in K562 cells at three doses across nine sites targeting hAAVS1. RNPs were transfected into cells via nucleofection. After 3 days, the edited sites were amplified for NGS and sequenced for indel analysis. [Figure 38] The study demonstrated cellular activity associated with incremental truncation of MG119-28 sgRNA1. Eight truncation designs were generated by truncating the MG119-28 wild-type guide at two stem-loop structures of repeat-anti-repeat regions. The guide, carrying an AAVS1-targeting spacer, was transfected to K562 by nucleofection. Genomic DNA was analyzed for indels using NGS. [Figure 39]This study demonstrates the in vitro cleavage activity of modern and ancestral MG191 nucleases. The minimal array used in this experiment consists of a repeat sequence, a spacer targeting an 8N PAM plasmid library, and another repeat sequence. "A" indicates Apo conditions, where the minimal array is omitted. Each reactant is analyzed using Tapestation to detect cleavage products of 200–300 bp. Ancestral candidates were tested with minimal arrays derived from closely related homologs. Modern candidates were tested with minimal arrays containing repeat sequences in two orientations (forward and reverse). Cleavage was observed in MG191-51 and 191-51 ancestral candidates, as well as in modern candidates MG191-1, MG191-2, MG191-5, and MG191-28. [Figure 40A] The sequence logos from NGS sequencing of in vitro cleavage active products are shown. Figure 40A: The ancestral MG191-53 was active against crRNA from the modern MG191 nuclease, with activity indicated at the top of each sequence logo. Ancestral candidates were tested in minimal arrays, and the active nucleases showed that they could process crevices from the modern nuclease. MG191-53 prefers the PAM of TtR, with lowercase letters indicating a low signal. [Figure 40B] The sequence logo from NGS sequencing of the in vitro cleavage active product is shown. Figure 40B: The modern MG191-5 nuclease is active against its corresponding crRNA. MG191-5 prefers the PAM of TRR. [Figure 41] Gel electrophoresis analysis shows that MG191 crRNA is compatible with multiple MG191 candidates in vitro. Nucleases with similar known PAMs can be activated with the same crRNA. Double-strand breaks are tested by incubating each nuclease with multiple minimal arrays consisting of repeat sequences, spacer sequences targeting PAM plasmid libraries, and repeat sequences. Break products of 200–300 bp in length are analyzed and imaged on a 1.5% agarose gel. Apo conditions are tested without minimal arrays. PC = positive control nuclease and crRNA. [Figure 42A]Figure 42A shows the SEC purification of Sumo-MG191-12 and the activity analysis of SUMO-MG191-12 and SUMO-MG191-25. SUMO-fusion enriched proteins after IMAC were run on an S200i 10 300 SEC column, the peak fraction was collected and concentrated (shaded box, left), and then run on an SDS-PAGE protein gel (right). [Figure 42B] Figure 42B shows the SEC purification of Sumo-MG191-12 and the activity analysis of SUMO-MG191-12 and SUMO-MG191-25. Figure 42B: Protein activity is evaluated by in vitro cleavage reaction using 521 bp linear DNA as the substrate. Lanes: (1) Substrate only, (2)-(5) Substrate + 1×, 5×, 25×, 40× molar excess RNP relative to the substrate. Figure 42C: Protein activity is evaluated by in vitro cleavage reaction using 2,281 bp plasmid DNA as the substrate. Lanes: (1) Substrate only, (2) No guide: 40× molar excess protein relative to the substrate without SgRNA, (3)-(5) Substrate + 1×, 5×, 25×, 40× molar excess RNP relative to the substrate. [Figure 42C] Figure 42B shows the SEC purification of Sumo-MG191-12 and the activity analysis of SUMO-MG191-12 and SUMO-MG191-25. Figure 42B: Protein activity is evaluated by in vitro cleavage reaction using 521 bp linear DNA as the substrate. Lanes: (1) Substrate only, (2)-(5) Substrate + 1×, 5×, 25×, 40× molar excess RNP relative to the substrate. Figure 42C: Protein activity is evaluated by in vitro cleavage reaction using 2,281 bp plasmid DNA as the substrate. Lanes: (1) Substrate only, (2) No guide: 40× molar excess protein relative to the substrate without SgRNA, (3)-(5) Substrate + 1×, 5×, 25×, 40× molar excess RNP relative to the substrate.

[0018] A brief explanation of sequence listings The sequence listings submitted herein provide exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems described herein. The following is an exemplary description of some of these sequences.

[0019] MG122 Sequence IDs 1-5 show the complete amino acid sequence of MG122 nuclease.

[0020] MG120 Sequence IDs 6-14 show the complete amino acid sequence of MG120 nuclease.

[0021] Sequence IDs 333-335 and 355-357 show the nucleotide sequences of MG120 tracrRNA derived from the same locus as the MG120 Cas effector.

[0022] Sequence IDs 374-375 and 389-390 show the nucleotide sequences of the MG120 minimum array.

[0023] MG118 Sequence ID 15 shows the complete amino acid sequence of MG118 nuclease.

[0024] Sequence ID 376 shows the nucleotide sequence of the MG118 minimal array.

[0025] Sequence ID 391 shows the nucleotide sequence of the MG118 minimal array.

[0026] Sequence IDs 400-401 show the nucleotide sequences of MG118-targeted CRISPR repeats.

[0027] Sequence IDs 410-411 show the nucleotide sequences of MG118 crRNA.

[0028] MG90 Sequence IDs 16-29 show the complete amino acid sequence of MG90 nuclease.

[0029] Sequence IDs 346-347 and 368-369 show the nucleotide sequences of MG90 tracrRNA derived from the same gene locus as the MG90 Cas effector.

[0030] Sequence IDs 383-384 and 398-399 show the nucleotide sequences of the MG90 minimum array.

[0031] Sequence IDs 402-403 show the nucleotide sequences of MG90-targeted CRISPR repeats.

[0032] Sequence IDs 412-413 show the nucleotide sequences of MG90 sgRNA.

[0033] MG119 Sequence IDs 30-150, 420-431, 476-624, and 629 show the complete amino acid sequence of the MG119 nuclease.

[0034] Sequence IDs 326-332, 336-345, 348-354, and 358-367 show the nucleotide sequences of MG119 tracrRNA derived from the same locus as the MG119 Cas effector.

[0035] Sequence numbers 370-373, 377-382, 385-388, and 392-397 show the nucleotide sequences of the MG119 minimal array.

[0036] Sequence IDs 404-409 show the nucleotide sequences of MG119-targeted CRISPR repeats.

[0037] Sequence numbers 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731 show the nucleotide sequences of MG119 sgRNA.

[0038] Sequence numbers 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, and 646 show the nucleotide sequences of MG119 PAM.

[0039] MG91B Sequence IDs 151-291 show the complete amino acid sequences of MG91B nuclease.

[0040] MG91C Sequence IDs 292-318 show the complete amino acid sequence of MG91C nuclease.

[0041] MG91A Sequence ID 319 shows the complete amino acid sequence of MG91A nuclease.

[0042] MG126 Sequence IDs 320-325 show the complete amino acid sequence of MG126 nuclease.

[0043] MG191 Sequence IDs 1065-1090, 1114-1118, and 1142-1172 show the complete amino acid sequences of the MG191 nuclease.

[0044] Sequence IDs 1091-1113 and 1119-1120 show the nucleotide sequences of MG191 crRNA.

[0045] Sequence IDs 1732-1738 show the nucleotide sequences of the MG191 minimal array.

[0046] Sequence IDs 1739-1745 and 1837-1841 show the nucleotide sequences of MG191 PAM.

[0047] Sequence ID 1873 shows the nucleotide sequence of the active MG191 repeat sequence.

[0048] Sequence ID 1874 shows the nucleotide sequence of the MG191 minimal array.

[0049] Sequence ID 1875 shows the nucleotide sequence of the MG191 PAM sequence.

[0050] Sequence ID 1876 shows the nucleotide sequence of MG191.

[0051] MG185 Sequence ID 1746 shows the complete amino acid sequence of MG185 nuclease.

[0052] MG186 Sequence IDs 1747-1748 show the complete amino acid sequence of MG186 nuclease.

[0053] MG187 Sequence IDs 1749-1750 show the complete amino acid sequence of MG187 nuclease.

[0054] MG188 Sequence IDs 1751-1752 show the complete amino acid sequence of MG188 nuclease.

[0055] Nuclear localization signals Sequence IDs 630-645 and 1904-1933 show the amino acid sequences of nuclear localization signals.

[0056] TRAC targeting Sequence IDs 767–798 show the nucleotide sequences of sgRNAs engineered to function with the MG119-2 nuclease in order to target TRAC.

[0057] Sequence IDs 799-830 show the DNA sequences of the TRAC target sites.

[0058] APOA1 targeting using MG119-125 Sequence IDs 831-904 show the nucleotide sequences of sgRNAs engineered to function with the MG119-125 nuclease to target APOA1.

[0059] Sequence IDs 905-978 show the DNA sequences of the APOA1 target site.

[0060] APOA1 targeting using MG119-129 Sequence IDs 979-1021 show the nucleotide sequences of sgRNAs engineered to function with the MG119-129 nuclease to target APOA1.

[0061] Sequence IDs 1022-1064 show the DNA sequences of the APOA1 target site.

[0062] Albumin (ALB) targeting using MG119-2 Sequence IDs 1121–1169 show the nucleotide sequences of sgRNAs engineered to function with the MG119-2 nuclease in order to target the albumin gene.

[0063] Sequence IDs 1170-1218 show the DNA sequences of the albumin gene target sites.

[0064] APOA1 targeting using MG119-2 Sequence IDs 1219–1237 show the nucleotide sequences of sgRNAs engineered to function with the MG119-2 nuclease in order to target APOA1.

[0065] Sequence IDs 1238-1256 show the DNA sequences of the APOA1 target site.

[0066] Targeting AAVS1 using MG119-2 Sequence IDs 1257–1324 show the nucleotide sequences of sgRNAs engineered to function with the MG119-2 nuclease in order to target AAVS1.

[0067] Sequence IDs 1325-1392 show the DNA sequences of the AAVS1 target site.

[0068] Albumin (ALB) targeting using MG119-28 Sequence IDs 1393–1441, 1887, 1889, 1891, and 1892–1893 show the nucleotide sequences of sgRNAs engineered to function with the MG119-28 nuclease to target the albumin gene.

[0069] Sequence IDs 1442-1490, 1888, 1890, and 1994 show the DNA sequences of the albumin gene target sites.

[0070] APOA1 targeting using MG119-28 Sequence IDs 1491–1506 show the nucleotide sequences of sgRNAs engineered to function with the MG119-28 nuclease in order to target APOA1.

[0071] Sequence IDs 1507-1522 show the DNA sequences of the APOA1 target site.

[0072] Targeting AAVS1 using MG119-28 Sequence IDs 1523-1562 and 1753-1779 show the nucleotide sequences of sgRNAs engineered to function with the MG119-28 nuclease in order to target AAVS1.

[0073] Sequence IDs 1563-1602 and 1780-1806 show the DNA sequences of the AAVS1 target site.

[0074] Albumin (ALB) targeting using MG119-32 Sequence IDs 1603–1632 show the nucleotide sequences of sgRNAs engineered to function with the MG119-32 nuclease in order to target albumin.

[0075] Sequence IDs 1633-1662 show the DNA sequences of albumin target sites.

[0076] APOA1 targeting using MG119-32 Sequence IDs 1663–1669 show the nucleotide sequences of sgRNAs engineered to function with the MG119-32 nuclease in order to target APOA1.

[0077] Sequence IDs 1670-1676 show the DNA sequences of the APOA1 target site.

[0078] Targeting AAVS1 using MG119-32 Sequence IDs 1677–1686 show the nucleotide sequences of sgRNAs engineered to function with the MG119-32 nuclease in order to target AAVS1.

[0079] Sequence IDs 1687-1696 show the DNA sequences of the AAVS1 target site.

[0080] U67 Spacer DNA Sequence IDs 1877 and 1891 show the nucleotide sequences of the U67 spacer DNA.

[0081] U40 Spacer DNA Sequence ID 1878 shows the nucleotide sequence of the U40 spacer DNA.

[0082] Oligonucleotide primers for nucleic acid synthesis Sequence IDs 1895, 1896, 1897, 1898, and 1899 show the nucleotide sequences of 611F_HE, 869R_HE, 680_HE Taqman probes, 611 NGS, and 927 R NGS DNA primers for nucleic acid synthesis.

[0083] SUMO protein sequence Sequence ID 1900 shows the amino acid sequence of the SUMO protein.

[0084] Precision protease sequence Sequence ID 1902 shows the amino acid sequence of the precision protease region. [Modes for carrying out the invention]

[0085] While various embodiments of the Disclosure are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Numerous variations, modifications, and substitutions can be conceived by those skilled in the art without departing from the Disclosure. It should be understood that various alternatives to the embodiments of the Disclosure described herein may be used.

[0086] Whenever the terms "no more than," "less than," or "less than or equal to" are placed before the first number in a set of two or more numbers, the terms "no more than," "less than," or "less than" apply to each number in that set. For example, 3, 2, or 1 or less is equivalent to 3 or less, 2 or less, or 1 or less.

[0087] The practices of some of the methods disclosed herein employ techniques from immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, unless otherwise indicated.

[0088] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context otherwise explicitly indicates. Furthermore, the terms "including," "includes," "having," "has," and "with," or their variations thereof, are intended to be as inclusive as the term "comprising," to the extent that they are used in either the detailed description or the claims.

[0089] The terms “about” or “approximately” mean within an acceptable range of error for a particular value as determined by those skilled in the art, which depends in part on how the value is measured or determined, i.e., the limits of the measuring system. For example, “about” may mean within one or more standard deviations according to the practice of the art. Alternatively, “about” may mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0090] As used in this disclosure, the term “nucleotide” refers to a base-sugar-phosphate combination. Nucleotides intended to be used include naturally occurring and synthetic nucleotides. A nucleotide is the monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes ribonucleoside triphosphates such as adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, e.g., dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used in this disclosure, the term nucleotide encompasses dideoxyribonucleoside triphosphate (ddNTP) and its derivatives. Exemplary examples of ddNTPs include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or may be labeled in a detectable manner, such as by using optically detectable moieties (e.g., fluorophores) or moieties containing quantum dots. Examples of detectable labels include radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzymatic labels. Examples of fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer (Foster City, Calif), and FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, fluorescein-15-dATP, fluorescein-12-dUTP, tetramethylrhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, available from Boehringer (Mannheim, Indianapolis, India), and Molecular Chromosome-labeled nucleotides available from Probes (Eugene, Oreg) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0091] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to refer to polymeric forms of nucleotides of any length, which are either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-stranded, double-stranded, or multi-stranded forms. Polynucleotides as intended include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, multiple loci (one locus) defined by binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. When referring to T in polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA. Polynucleotides can be exogenous or endogenous to cells and / or exist in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., modified backbone, sugar, or nucleic acid base). Where present, modifications to the nucleotide structure are conferred before or after polymer assembly. Non-exclusive examples of modifications include 5-bromouracil, peptide nucleic acids, heteronucleotides, morpholino, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein bound to sugar), thiol-containing nucleotides, biotin-bound nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosin, and waiosin. Nucleotide sequences can be interrupted by non-nucleotide components.

[0092] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to refer to polymers of at least two amino acid residues joined by peptide bonds. These terms do not imply a specific length of the polymer and are not intended to imply or distinguish whether peptides are produced using recombinant techniques, chemical or enzymatic synthesis, or naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The terms include amino acid chains of any length, including full-length proteins and proteins (e.g., domains) with or without secondary or tertiary structures. The terms also encompass amino acid polymers modified by any other operations, such as disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with labeling components. As used in this disclosure, the terms “amino acid” and “multiple amino acids” refer to natural and non-natural amino acids, including but not limited to modified amino acids. Modified amino acids include amino acids that have been chemically modified to include a group or chemical moiety that does not naturally occur on the amino acid. The term "amino acid" includes both D-amino acids and L-amino acids.

[0093] As used herein, “operably linked,” “operable linkage,” “operatively linked,” or their grammatical equivalents refer to the arrangement of gene elements, such as promoters, enhancers, polyadenylation sequences, etc., where the operation (e.g., movement or activation) of a first gene element has some effect on a second gene element. The effect on the second gene element may, but does not have to be, the same type as the operation of the first gene element. For example, if the movement of the first element causes the activation of the second element, the two gene elements are operably linked. For example, if a regulatory element, which may include a promoter sequence and / or an enhancer sequence, helps initiate transcription of a coding sequence, the regulatory element is operably linked to the coding region. Intervening residues may exist between the regulatory element and the coding region as long as this functional relationship is maintained.

[0094] The terms "transfection" or "transfected" refer to the introduction of nucleic acids into cells by non-viral or virus-based methods. Nucleic acid molecules can be complete proteins or gene sequences that encode functional portions thereof.

[0095] As used in this disclosure, “unnatural” means a nucleic acid or polypeptide sequence that does not exist in nature. Unnatural means a nucleic acid or polypeptide sequence that does not exist in nature, including modifications such as mutations, insertions, or deletions. The term unnatural encompasses fusion nucleic acids or polypeptides in which the unnatural sequence encodes or exhibits the activity (e.g., enzyme activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.) of the nucleic acid or polypeptide sequence to which the unnatural sequence is fused. Unnatural nucleic acids or polypeptide sequences include those that are genetically engineered to link a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to produce a chimeric nucleic acid or polypeptide sequence encoding a chimeric nucleic acid or polypeptide.

[0096] As used in this disclosure, the term “promoter” refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and may be located adjacent to or overlapping with a nucleotide or region of nucleotides from which RNA transcription is initiated. A promoter may include a specific DNA sequence that binds to a protein factor, often referred to as a transcription factor, which facilitates the binding of RNA polymerase to DNA, resulting in gene transcription. Basic eukaryotic promoters typically, but not necessarily, include a TATA-box and / or CAAT-box.

[0097] As used herein, the term “expression” refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (into mRNA or other RNA transcripts, etc.), and / or the process by which the transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides may be collectively referred to as “gene products.” When the polynucleotide is derived from genomic DNA, the term “expression” includes the splicing of mRNA in eukaryotic cells.

[0098] As used in this disclosure, “vector” means a macromolecule or aggregate of macromolecules that contains or is related to a polynucleotide and mediates the delivery of a polynucleotide to a cell. Examples of vectors include nuclear-based vectors (e.g., plasmids and viral vectors) and liposomes. Exemplary nucleic acid-based vectors generally include genetic elements, such as regulatory elements, that are operably linked to a gene to promote the expression of the gene in a target.

[0099] As used in this disclosure, “expression cassette” and “nucleic acid cassette” are used interchangeably to refer to components of a vector that include a combination of nucleic acid sequences or elements (e.g., therapeutic genes, promoters, and terminators) that are expressed or operably linked for expression. The terms encompass an expression cassette that includes a combination of regulatory elements operably linked for expression and one or more genes.

[0100] A “functional fragment” of a DNA or protein sequence refers to a fragment that possesses biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes its ability to influence expression in a manner attributable to the full-length sequence.

[0101] The terms “engineered,” “synthetic,” and “artificial” are used interchangeably herein to refer to objects modified by human intervention. For example, these terms refer to polynucleotides or polypeptides that do not exist in nature. Engineered peptides have, but do not require, low sequence identity with respect to naturally occurring human proteins (e.g., less than 50%, less than 25%, less than 10%, less than 5%, less than 1%). For example, the VPR domain and VP64 domain are synthetic transactivation domains. Non-limiting examples include: nucleic acids modified by altering their sequence to a sequence that does not occur in nature; nucleic acids modified by ligating them to nucleic acids that are not related in nature so that the ligated product possesses a function not present in the original nucleic acid; engineered nucleic acids synthesized in vitro using sequences that do not exist in nature; proteins modified by altering their amino acid sequence to a sequence that does not exist in nature; engineered proteins that acquire new functions or properties. An “engineered” system includes at least one engineered component.

[0102] As used in this disclosure, “guide nucleic acid” or “guide polynucleotide” means a nucleic acid that hybridizes to a target nucleic acid and thereby directs the associated nuclease to the target nucleic acid. Guide nucleic acids are, but are not limited to, RNA (guide RNA or gRNA), DNA, or a mixture of RNA and DNA. Guide nucleic acids may include crRNA or tracrRNA, or a combination of both. The term guide nucleic acid encompasses engineered guide nucleic acids and programmable guide nucleic acids that specifically bind to the target nucleic acid. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. A double-stranded target polynucleotide chain that is complementary to the guide nucleic acid and hybridizes with it is called the complementary chain. A double-stranded target polynucleotide chain that is complementary to the complementary chain and therefore not complementary to the guide nucleic acid is called the non-complementary chain. A guide nucleic acid having a polynucleotide chain is called a “single guide nucleic acid”. A guide nucleic acid having two polynucleotide chains is called a “double guide nucleic acid”. Unless otherwise specified, the term “guide nucleic acid” is inclusive and refers to both single guide nucleic acids and double guide nucleic acids. A guide nucleic acid may include a segment referred to as a “nucleic acid targeting segment,” “nucleic acid targeting sequence,” or “spacer.” A nucleic acid targeting segment may include a subsegment referred to as a “protein-binding segment,” “protein-binding sequence,” or “Cas protein-binding segment.”

[0103] As used herein, the term “Cas12a” refers to a family of class 2, VA-type Cas endonucleases that (a) use a relatively small guide RNA (approximately 42–44 nucleotides) processed by the nuclease itself after transcription from a CRISPR array, and (b) cleave DNA in a manner that leaves alternating cleavage sites. Further characteristics of this enzyme family can be found, for example, in Zetsche B, Heidenreich M, Mohanraju P, et al. Nat Biotechnol 2017;35:31-34, and Zetsche B, Gootenberg JS, Abudayyeh OO, et al. Cell 2015;163:759-771.

[0104] The terms "tracrRNA" or "tracr sequence" refer to transactivated CRISPR RNA. TracrRNA interacts with CRISPR(cr)RNA to form a guide nucleic acid (e.g., guide RNA or gRNA) that can hybridize to a target nucleic acid, thereby directing the associated nuclease to the target nucleic acid.

[0105] The terms “sequence identity” or “identity rate” in relation to two or more nucleic acid or polypeptide sequences refer to two or more sequences (e.g., in pairwise alignment) or (e.g., in multiple sequence alignment) that are identical or have a specific proportion of identical amino acid residues or nucleotides when compared and aligned for maximum match across a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP, which uses the BLOSUM62 scoring matrix with parameters of word length (W) of 3, expected value (E) of 10, and gap cost set by presence of 11 and extension of 1, and also uses conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP, which uses PAM30 scoring with parameters of word length (W) of 2, expected value (E) of 1,000,000, and gap cost set by 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW, which uses Smith-Waterman homology search algorithm parameters of 2 for match, -1 for mismatch, and -1 for gap; MUSCLE, which uses the default parameters; MAFFT, which uses parameters of 2 for retrieval and 1,000 for maximum repeats; Novafold, which uses the default parameters; and HMMER hmmalign, which uses the default parameters.

[0106] In the context of two or more nucleic acid sequences or polypeptide sequences, the term “optimally aligned” refers to two (e.g., in pairwise alignment) or more (e.g., in multiple sequence alignment) sequences aligned by the greatest match of amino acid residues or nucleotides, as identified by the alignment that produces the highest or “optimized” identity percentage score.

[0107] Any variant of the enzymes described herein having one or more conserved amino acid substitutions is included in this disclosure. Such conserved substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids having similar hydrophobicity, polarity, and R-chain length with each other. In addition, or alternatively, by comparing aligned sequences of homologous proteins from different species, conserved substitutions can be identified by finding interspecies mutated amino acid residues (e.g., non-conserved residues) without altering the fundamental function of the encoded protein. Such conservatively substituted variants are composed of at least about 20%, at least about 25%, at least about 30%, at least about 35% of any one of the endonuclease protein sequences described herein (e.g., the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, MG126, or MG191 family endonucleases, or any other family nucleases described herein). The variants may include variants having at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identity. In some embodiments, such conservatively substituted variants are functional variants. Such functional variants may include sequences with substitutions such that the activity of one or more key active site residues or guide RNA binding residues of the endonuclease is not disrupted. In some embodiments, any functional variant of any of the proteins described herein lacks at least one substitution of a conserved or functional residue that is called out in Figure 2A, Figure 3A, Figure 4A, Figure 5A, or Figure 6A.In some embodiments, any functional variant of the proteins described herein lacks substitutions for all of the conserved or functional residues called out in Figures 2A, 3A, 4A, 5A, or 6A.

[0108] The disclosure also includes variants of any of the enzymes described herein (e.g., deactivation variants) having substitutions of one or more catalytic residues to reduce or eliminate the activity of the enzyme. In some embodiments, the deactivation variants as proteins described herein include disruptive substitutions of at least one, at least two, or all three catalytic residues that are called out in Figures 2A, 3A, 4A, 5A, or 6A.

[0109] Tables of conserved substitutions that provide functionally similar amino acids are available from various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993))). The following eight groups each contain amino acids that are conserved substitutions with each other: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), leucine (L), methionine (M), valine (V); 6) Phenylalanine (F), tyrosine (Y), tryptophan (W); 7) Serine (S), threonine (T); and 8) Cysteine ​​(C), Methionine (M)

[0110] overview The discovery of novel Cas enzymes with unique functionalities and structures offers the potential to further improve gene editing techniques, speed, specificity, functionality, and ease of use. Relatively few functionally characterized CRISPR / Cas enzymes exist in the literature compared to the predicted prevalence of clustered and regularly arranged short palindromic repeat (CRISPR) systems in microorganisms and the complete diversity of microbial species. This is partly because a vast number of microbial species may not be readily cultured under laboratory conditions. Metagenomic sequencing from natural environmental niches containing numerous microbial species could dramatically increase the number of novel characterized CRISPR / Cas systems and potentially accelerate the discovery of new oligonucleotide editing functions. A fruitful recent example of such an approach is demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.

[0111] The CRISPR / Cas system is an RNA-directed nuclease complex that functions as an adaptive immune system in microorganisms. In their natural context, the CRISPR / Cas system arises in a CRISPR (clustered and regularly arranged short palindromic repeat) operon or locus, which generally consists of two parts: (i) an array of short repeat sequences (30-40 bp) separated by short spacer sequences that encode an RNA-based targeting element, and (ii) an ORF encoding a Cas nuclease. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target nucleic acid and a crRNA guide, and (ii) the presence of a protospacer-adjacent motif (PAM) sequence within a specific vicinity of the target nucleic acid sequence that depends on a specific Cas nuclease (PAMs are typically sequences not commonly expressed in the host genome). Depending on the precise function and structure of the system, CRISPR-Cas systems are generally classified into two classes, five types, and sixteen subtypes based on shared functional characteristics and evolutionary similarities (see Figure 1).

[0112] Class 1 CRISPR-Cas systems have large, multi-subunit effector complexes and include type I, type III, and type IV Cas nucleases. Class 2 CRISPR-Cas systems generally have single polypeptide multi-domain nuclease effectors and include type II, type V, and type VI Cas nucleases.

[0113] The type II CRISPR-Cas system is considered the simplest in terms of components. In the type II CRISPR-Cas system, processing the CRISPR array into mature crRNA does not require the presence of a special endonuclease subunit, but rather a small transcoding crRNA (tracrRNA) having a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to produce a mature effector enzyme loaded with both tracrRNA and crRNA. The Cas II nuclease is identified as a DNA nuclease. The type II effector generally exhibits a structure containing a RuvC-like endonuclease domain that fits into an RNase H fold, which has an unrelated HNH nuclease domain inserted within the fold of the RuvC-like nuclease domain. The RuvC-like domain is involved in cleaving target (e.g., crRNA-complementary) DNA strands, while the HNH domain is involved in cleaving displaced DNA strands.

[0114] Type V CRISPR-Cas systems are characterized by a nuclease effector structure (e.g., Cas12) similar to that of type II effectors, containing a RuvC-like domain. Like type II, most (though not all) type V CRISPR systems use tracrRNA to process precursor crRNA into mature crRNA. However, unlike type II systems, which require RNAse III to cleave precursor crRNA into multiple crRNAs, type V systems can cleave precursor crRNA using the effector nuclease itself. Similar to type II CRISPR-Cas systems, type V CRISPR-Cas systems are also identified as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to possess robust single-strand nonspecific deoxyribonuclease activity, activated by the first crRNA-directed cleavage of a double-stranded target sequence.

[0115] CRISPR-Cas systems have emerged in recent years as a choice for gene editing technology due to their targeting potential and ease of use. The most commonly used systems are class 2 type II SpCas9 and class 2 type VA Cas12a (formerly Cpf1). VA systems, in particular, are becoming more widely used because their reported specificity in cells is higher than other nucleases, and they have fewer or no off-target effects. VA systems also have the advantage that their guide RNA is small (42-44 nucleotides compared to approximately 100 nt for SpCas9) and is processed by the nuclease itself after transcription from the CRISPR array, simplifying multiplexing applications involving multiple gene editing. Furthermore, VA systems have staggered cleavage sites, which can facilitate directed repair pathways such as microhomology-dependent targeted integration (MITI).

[0116] The most commonly used VA-type enzymes require a 5' protospacer adjacency motif (PAM) adjacent to the selected target site: 5'-TTTV-3' for Lachnospiraceae bacterium ND2006 LbCas12a and Acidaminococcus species AsCas12a; and 5'-TTV-3' for Francisella novicida FnCas12a. Recent investigations of orthologs have revealed proteins with less restrictive PAM sequences that are also active in mammalian cell cultures, such as YTV, YYN, or TTN. However, these enzymes do not fully encompass the biodiversity and targetability of V-type enzymes and may not represent all possible activity and PAM sequence requirements. In certain embodiments, improved V-type nucleases that are highly targetable, compact, and precise for use in systems and methods for gene editing are described herein.

[0117] MG Enzyme In certain embodiments, an engineered nuclease system comprising an endonuclease and an engineered guide polynucleotide is described herein. In some embodiments, the endonuclease is a class 2, type V endonuclease. In some embodiments, the endonuclease is a double-stranded nuclease. In some embodiments, the endonuclease is catalytically dead. In some embodiments, the endonuclease is modified. In some embodiments, the endonuclease is modified to yield an endonuclease having nickasase activity. In some embodiments, the modified endonuclease is a site-specific nickasase.

[0118] In some embodiments, the endonuclease is MG90 endonuclease (see Figures 3A-3C). In some embodiments, the endonuclease is MG90A endonuclease. In some embodiments, the endonuclease is MG90B endonuclease. In some embodiments, the endonuclease is MG90C endonuclease. In some embodiments, the endonuclease is MG91 endonuclease (see Figures 8A-8B). In some embodiments, the endonuclease is MG118 endonuclease (see Figures 5A-5C). In some embodiments, the endonuclease is MG119 endonuclease (see Figures 2A-2D). In some embodiments, the endonuclease is MG120 endonuclease (see Figures 7A-7C). In some embodiments, the endonuclease is MG122 endonuclease (see Figures 6A-6C). In some embodiments, the endonuclease is MG126 endonuclease (see Figures 4A-4C). In some embodiments, the endonuclease is MG191 endonuclease.

[0119] In some embodiments, the nuclease is less than approximately 1,000 amino acids in length. In some embodiments, the nuclease is less than approximately 900 amino acids in length. In some embodiments, the nuclease is less than approximately 850 amino acids in length. In some embodiments, the nuclease is less than approximately 800 amino acids in length. In some embodiments, the nuclease is less than approximately 750 amino acids in length. In some embodiments, the nuclease is less than approximately 700 amino acids in length. In some embodiments, the nuclease is less than approximately 650 amino acids in length. In some embodiments, the nuclease is less than approximately 600 amino acids in length. In some embodiments, the nuclease is less than approximately 550 amino acids in length. In some embodiments, the nuclease is less than approximately 500 amino acids in length.

[0120] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 1 to 5. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with any one of SEQ ID NOs: 1 to 5. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of SEQ ID NOs: 1 to 5. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with any one of SEQ ID NOs: 1 to 5. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with any one of sequence numbers 1 to 5. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with any one of sequence numbers 1 to 5. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with any one of sequence numbers 1 to 5. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with any one of sequence numbers 1 to 5. In some embodiments, the endonuclease includes a sequence having at least about 97% identity with any one of sequence numbers 1 to 5. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with any one of sequence numbers 1 to 5. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with any one of sequence numbers 1 to 5. In some embodiments, the endonuclease comprises a sequence having 100% identity with any one of sequence numbers 1 to 5.

[0121] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 6-14. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with any one of SEQ ID NOs: 6-14. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of SEQ ID NOs: 6-14. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with any one of SEQ ID NOs: 6-14. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with any one of sequence numbers 6-14. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with any one of sequence numbers 6-14. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with any one of sequence numbers 6-14. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with any one of sequence numbers 6-14. In some embodiments, the endonuclease includes a sequence having at least about 97% identity with any one of sequence numbers 6-14. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with any one of sequence numbers 6-14. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with any one of sequence numbers 6-14. In some embodiments, the endonuclease contains a sequence that is 100% identical to any one of sequence numbers 6-14.

[0122] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 85% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 90% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 95% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 96% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 97% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 98% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having at least about 99% identity with SEQ ID NO: 15. In some embodiments, the endonuclease contains a sequence having 100% identity with SEQ ID NO: 15.

[0123] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 16-29. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with any one of SEQ ID NOs: 16-29. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of SEQ ID NOs: 16-29. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with any one of SEQ ID NOs: 16-29. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with any one of sequence numbers 16-29. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with any one of sequence numbers 16-29. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with any one of sequence numbers 16-29. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with any one of sequence numbers 16-29. In some embodiments, the endonuclease includes a sequence having at least about 97% identity with any one of sequence numbers 16-29. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with any one of sequence numbers 16-29. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with any one of sequence numbers 16-29. In some embodiments, the endonuclease comprises a sequence having 100% identity with any one of sequence numbers 16 to 29.

[0124] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease includes a sequence having at least about 75% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease includes a sequence having at least about 80% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease contains a sequence having at least about 95% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease contains a sequence having at least about 96% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease contains a sequence having at least about 97% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629.In some embodiments, the endonuclease includes a sequence having at least about 98% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629. In some embodiments, the endonuclease includes a sequence having 100% identity with one of sequence numbers 30-150, 420-431, 476-624, and 629.

[0125] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 151-291. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with any one of SEQ ID NOs: 151-291. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of SEQ ID NOs: 151-291. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with any one of SEQ ID NOs: 151-291. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with any one of sequence numbers 151 to 291. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with any one of sequence numbers 151 to 291. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with any one of sequence numbers 151 to 291. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with any one of sequence numbers 151 to 291. In some embodiments, the endonuclease includes a sequence having at least about 97% identity with any one of sequence numbers 151 to 291. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with any one of sequence numbers 151 to 291. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with any one of sequence numbers 151 to 291. In some embodiments, the endonuclease comprises a sequence having 100% identity with any one of sequence numbers 151 to 291.

[0126] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 292-318. In some embodiments, the endonuclease contains a sequence having at least 70% identity with any one of SEQ ID NOs: 292-318. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of SEQ ID NOs: 292-318. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with any one of SEQ ID NOs: 292-318. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with any one of sequence numbers 292-318. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with any one of sequence numbers 292-318. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with any one of sequence numbers 292-318. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with any one of sequence numbers 292-318. In some embodiments, the endonuclease includes a sequence having at least about 97% identity with any one of sequence numbers 292-318. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with any one of sequence numbers 292-318. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with any one of sequence numbers 292-318. In some embodiments, the endonuclease comprises a sequence having 100% identity with any one of sequence numbers 292 to 318.

[0127] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 85% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 90% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 95% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 96% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 97% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 98% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having at least about 99% identity with SEQ ID NO: 319. In some embodiments, the endonuclease contains a sequence having 100% identity with SEQ ID NO: 319.

[0128] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 320-325. In some embodiments, the endonuclease contains a sequence having at least 70% identity with any one of SEQ ID NOs: 320-325. In some embodiments, the endonuclease contains a sequence having at least 75% identity with any one of SEQ ID NOs: 320-325. In some embodiments, the endonuclease contains a sequence having at least 80% identity with any one of SEQ ID NOs: 320-325. In some embodiments, the endonuclease includes a sequence having at least 85% identity with any one of sequence numbers 320-325. In some embodiments, the endonuclease includes a sequence having at least 90% identity with any one of sequence numbers 320-325. In some embodiments, the endonuclease includes a sequence having at least 95% identity with any one of sequence numbers 320-325. In some embodiments, the endonuclease includes a sequence having at least 96% identity with any one of sequence numbers 320-325. In some embodiments, the endonuclease includes a sequence having at least 97% identity with any one of sequence numbers 320-325. In some embodiments, the endonuclease includes a sequence having at least 98% identity with any one of sequence numbers 320-325. In some embodiments, the endonuclease includes a sequence having at least 99% identity with any one of sequence numbers 320-325. In some embodiments, the endonuclease comprises a sequence having 100% identity with any one of sequence numbers 320-325.

[0129] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 85% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 90% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 95% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 96% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 97% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having at least about 98% identity with one of sequence numbers 1065-1090 and 1114-1118.In some embodiments, the endonuclease contains a sequence having at least about 99% identity with one of sequence numbers 1065-1090 and 1114-1118. In some embodiments, the endonuclease contains a sequence having 100% identity with one of sequence numbers 1065-1090 and 1114-1118.

[0130] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 85% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 90% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 95% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 96% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 97% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 98% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having at least about 99% identity with SEQ ID NO: 1746. In some embodiments, the endonuclease contains a sequence having 100% identity with SEQ ID NO: 1746.

[0131] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 80% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 85% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 90% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 95% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 96% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 97% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 98% identity with any one of sequence numbers 1747-1748. In some embodiments, the endonuclease contains a sequence having at least about 99% identity with any one of sequence numbers 1747-1748.In some embodiments, the endonuclease contains a sequence that is 100% identical to any one of sequence numbers 1747-1748.

[0132] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs: 1749-1750. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with any one of SEQ ID NOs: 1749-1750. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of SEQ ID NOs: 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 80% identity with any one of sequence numbers 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with any one of sequence numbers 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with any one of sequence numbers 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with any one of sequence numbers 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with any one of sequence numbers 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 97% identity with any one of sequence numbers 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with any one of sequence numbers 1749-1750. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with any one of sequence numbers 1749-1750.In some embodiments, the endonuclease contains a sequence that is 100% identical to any one of sequence numbers 1749-1750.

[0133] In some embodiments, the endonuclease contains a sequence having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease contains a sequence having at least about 70% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease contains a sequence having at least about 75% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 80% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 85% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 90% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 95% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 96% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 97% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 98% identity with any one of sequence numbers 1751-1752. In some embodiments, the endonuclease includes a sequence having at least about 99% identity with any one of sequence numbers 1751-1752.In some embodiments, the endonuclease comprises a sequence having 100% identity with any one of sequence numbers 1751 to 1752.

[0134] In some embodiments, the nuclease contains a RuvC catalytic residue. In some embodiments, the nuclease does not require tracrRNA. In some embodiments, the nuclease contains a RuvC catalytic residue and does not require tracrRNA.

[0135] In some embodiments, the nuclease includes a PAM interaction domain. In some embodiments, the PAM includes one of the sequences SEQ ID NOs: 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, and 646.

[0136] In some embodiments, the manipulated nuclease system is discovered via metagenomic sequencing. In some embodiments, metagenomic sequencing is performed on samples taken from various environments. In some embodiments, the environment is a human microbiome, an animal microbiome, a high-temperature environment, a low-temperature environment, or sediment.

[0137] In some embodiments, the endonuclease includes a fusion of nuclear localization sequences (NLS). In some embodiments, the NLS is located at the N-terminus of the endonuclease. In some embodiments, the NLS is located at the C-terminus of the endonuclease. In some embodiments, the NLS is located at both the N-terminus and the C-terminus of the endonuclease.

[0138] In some embodiments, the NLS includes a sequence that has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the sequences 630–645 and 1904–1933. In some cases, the NLS contains sequences that are at least approximately 80% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 85% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 90% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 91% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 92% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 93% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 94% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 95% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 96% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 97% identical to sequence numbers 630-645 and 1904-1933.In some cases, the NLS contains sequences that are at least approximately 98% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are at least approximately 99% identical to sequence numbers 630-645 and 1904-1933. In some cases, the NLS contains sequences that are 100% identical to sequence numbers 630-645 and 1904-1933.

[0139] [Table 1-1]

[0140] [Table 1-2]

[0141] In some cases, the manipulated nuclease system further comprises a single-stranded or double-stranded DNA repair template.

[0142] In some cases, the single-stranded or double-stranded DNA repair template includes a first homologous arm from 5' to 3' containing a sequence of at least 20 nucleotides on the 5' side relative to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm containing a sequence of at least 20 nucleotides on the 3' side relative to the target sequence.

[0143] In some cases, the first homologous arm contains a sequence of at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides. In some cases, the second homologous arm contains a sequence of at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides.

[0144] In some cases, the first and second homologous arms are homologous to the prokaryotic genome sequence. In some cases, the first and second homologous arms are homologous to the bacterial genome sequence. In some cases, the first and second homologous arms are homologous to the fungal genome sequence. In some cases, the first and second homologous arms are homologous to the eukaryotic genome sequence.

[0145] In some embodiments, the manipulated nuclease system further includes a DNA repair template comprising a double-stranded DNA segment adjacent to one or two single-stranded DNA segments. In some embodiments, the single-stranded DNA segment is conjugated to the 5' end of the double-stranded DNA segment. In some embodiments, two single-stranded DNA segments are adjacent to the double-stranded DNA segment. In some embodiments, the single-stranded DNA segment is conjugated to the 3' end of the double-stranded DNA segment. In some embodiments, the stranded DNA segment has a length of 4 to 10 nucleotide bases. In some embodiments, the single-stranded DNA segment has a nucleotide sequence complementary to the sequence in the spacer sequence.

[0146] In some cases, a single-stranded DNA segment has a length of 1 to 15 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 4 to 10 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 4 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 5 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 6 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 7 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 8 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 9 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 10 nucleotide bases.

[0147] In some cases, the single-stranded DNA segment has a nucleotide sequence that is complementary to the sequence in the spacer sequence. In some cases, the double-stranded DNA sequence includes a barcode, open reading frame, enhancer, promoter, protein-coding sequence, miRNA-coding sequence, RNA-coding sequence, or transgene. In some embodiments, the double-stranded DNA sequence is adjacent to a nuclease cleavage site. In some embodiments, the nuclease cleavage site includes a spacer and a PAM sequence. In some embodiments, the PAM includes one of the sequences SEQ ID NOs: 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, 475, and 646.

[0148] In some cases, the manipulated nuclease system is Mg 2+ This further includes the sources of supply.

[0149] In some cases, the sequence can be determined by the BLASTP, CLUSTALW, MUSCLE, or MAFFT algorithms, or the CLUSTALW algorithm with Smith-Waterman homology search algorithm parameters. In some cases, the sequence can be determined by the BLASTP homology search algorithm, which uses a BLOSUM62 scoring matrix with 3 word length (W), 10 expected value (E) parameters, and 11 existence and 1 extension for gap cost, and also uses conditional composition score matrix adjustments.

[0150] Guide polynucleotides In some embodiments, the engineered nuclease systems disclosed herein include engineered guide polynucleotides, such as guide ribonucleic acid (gRNA), single gRNA, or dual guide RNA.

[0151] In some embodiments, the manipulated polynucleotide includes a sequence having at least about 70% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 75% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 80% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 85% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 90% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 95% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 96% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 97% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 98% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 99% identity with one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated polynucleotide contains 100% identity with any one of SEQ ID NOs. 333-335 and 355-357.

[0152] In some embodiments, the manipulated polynucleotide includes sequences having at least about 70% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotide includes sequences having at least about 75% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotide includes sequences having at least about 80% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotide includes sequences having at least about 85% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotide includes sequences having at least about 90% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotide includes sequences having at least about 95% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotide includes sequences having at least about 96% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotide includes sequences having at least about 97% identity with SEQ ID NOs. 410-411. In some embodiments, the manipulated polynucleotides contain sequences that are at least about 98% identical to sequence numbers 410-411. In some embodiments, the manipulated polynucleotides contain sequences that are at least about 99% identical to sequence numbers 410-411. In some embodiments, the manipulated polynucleotides contain 100% identical to sequence numbers 410-411.

[0153] In some embodiments, the manipulated polynucleotide includes a sequence having at least about 70% identity with one of sequence numbers 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 75% identity with one of sequence numbers 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 80% identity with one of sequence numbers 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 85% identity with one of sequence numbers 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 90% identity with one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 95% identity with one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 96% identity with one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 97% identity with one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide contains a sequence having at least about 98% identity with one of sequence numbers 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide contains a sequence having at least about 99% identity with one of sequence numbers 346-347, 368-369, and 412-413. In some embodiments, the manipulated polynucleotide contains 100% identity with one of sequence numbers 346-347, 368-369, and 412-413.

[0154] In some embodiments, the manipulated polynucleotides include sequences having at least about 70% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotides include sequences having at least about 75% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotides include sequences having at least about 80% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotides include sequences having at least about 85% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876.In some embodiments, the manipulated polynucleotides include sequences having at least about 90% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotides include sequences having at least about 95% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotides include sequences having at least about 96% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotides include sequences having at least about 97% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876.In some embodiments, the manipulated polynucleotides include sequences having at least about 98% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotides include sequences that have at least about 99% identity with any one of the following sequence numbers: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876. In some embodiments, the manipulated polynucleotide contains 100% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1697-1731, and 1876.

[0155] In some embodiments, the manipulated polynucleotide includes a sequence having at least about 70% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 75% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 80% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 85% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 90% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 95% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 96% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide includes a sequence having at least about 97% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide contains a sequence having at least about 98% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide contains a sequence having at least about 99% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated polynucleotide contains 100% identity with one of sequence numbers 1091-1113, 1119-1120, and 1876.

[0156] In some embodiments, the engineered guide polynucleotide targets a gene in a cell. In some embodiments, the engineered guide polynucleotide targets a gene in a mammalian cell. In some embodiments, the mammalian cell is a pig, cattle, goat, sheep, rodent, rat, mouse, non-human primate, or human cell. In some embodiments, the target gene or target locus is albumin, TRAC, AAVS1, or APOA1.

[0157] In some embodiments, the target gene is TRAC. In some embodiments, the guide polynucleotide targeting TRAC is encoded by one of sequence numbers 767-798, or by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with one of sequence numbers 767-798. In some embodiments, the guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of sequence numbers 767 to 798. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity with any one of sequence numbers 767 to 798. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity with one of sequence numbers 767-798. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity with one of sequence numbers 767-798. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity with one of sequence numbers 767-798. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity with one of sequence numbers 767-798. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity with one of sequence numbers 767-798.In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity with one of sequence numbers 767-798. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity with one of sequence numbers 767-798. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity with one of sequence numbers 767-798.

[0158] In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a target nucleic acid sequence (e.g., SEQ ID NOs. 767-798) within or in the introns of the TRAC gene. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to any one of SEQ ID NOs. 767-798, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs. 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity with any one of SEQ ID NOs. 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 85% identity with any one of SEQ ID NOs. 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity with any one of SEQ ID NOs. 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 95% identity with any one of SEQ ID NOs. 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity with any one of SEQ ID NOs. 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity with any one of SEQ ID NOs. 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 98% identity with any one of SEQ ID NOs. 767-798.In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 99% identity with any one of sequence numbers 767-798. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having 100% identity with any one of sequence numbers 767-798.

[0159] In some embodiments, the guide polynucleotide hybridizes or targets a sequence (e.g., SEQ ID NOs. 799-830) within or within the introns of the TRAC gene. In some embodiments, the guide polynucleotide hybridizes or targets a sequence with at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes or targets a sequence with at least approximately 80% identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes or targets a sequence with at least approximately 85% identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 90% identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 95% identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 96% identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 97% identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 98% identity with any one of SEQ ID NOs. 799-830. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having at least about 99% identity with any one of sequence numbers 799-830.In some embodiments, the guide polynucleotide hybridizes to or targets a sequence that has 100% identity with any one of sequence numbers 799-830.

[0160] In some embodiments, the target gene is APOA1. In some embodiments, the guide polynucleotide targeting APOA1 is encoded by one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669, or by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669.In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669.

[0161] In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a target nucleic acid sequence (e.g., SEQ ID NOs: 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669) within the APOA1 gene or introns of the APOA1 gene. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 85% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 95% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669.In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 98% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 99% identity with one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having 100% identity with any one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669.

[0162] In some embodiments, the guide polynucleotide hybridizes or targets sequences within the APOA1 gene or within the introns of the APOA1 gene (e.g., SEQ ID NOs. 905-978, 1022-1064, 1238-1256, and 1670-1676). In some embodiments, the guide polynucleotide hybridizes or targets sequences of any one of SEQ ID NOs. 905-978, 1022-1064, 1238-1256, and 1670-1676, or sequences having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs. 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 80% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 85% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 90% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 95% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 96% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676.In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 97% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 98% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 99% identity with one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having 100% identity with any one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676.

[0163] In some embodiments, the target gene is AAVS1. In some embodiments, the guide polynucleotide targeting AAVS1 is encoded by one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779, or by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide comprises a sequence containing at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779.In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity with one of sequence numbers 1257-1324, 1523-1562, and 1677-1686. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779.

[0164] In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to a target nucleic acid sequence (e.g., SEQ ID NOs. 1257-1324, 1523-1562, 1677-1686, and 1753-1779) within the AAVS1 gene or introns of the AAVS1 gene. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence complementary to any one of SEQ ID NOs. 1257-1324, 1523-1562, 1677-1686, and 1753-1779, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs. 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 85% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 95% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779.In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 98% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 99% identity with one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having 100% identity with any one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779.

[0165] In some embodiments, the guide polynucleotide hybridizes or targets sequences within the AAVS1 gene or within the introns of the AAVS1 gene (e.g., SEQ ID NOs. 1325-1392, 1563-1602, 1687-1696, and 1780-1806). In some embodiments, the guide polynucleotide hybridizes or targets sequences corresponding to any one of SEQ ID NOs. 1325-1392, 1563-1602, 1687-1696, and 1780-1806, or sequences having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs. 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 80% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 85% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 90% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 95% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 96% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806.In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 97% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 98% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 99% identity with one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence having 100% identity with any one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806.

[0166] In some embodiments, the target gene is albumin. In some embodiments, the guide polynucleotide targeting albumin is encoded by one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893, or by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide includes a sequence comprising at least 46 to 80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of SEQ ID NOs. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893.In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity with one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893.

[0167] In some embodiments, the guide polynucleotide hybridizes or targets a target nucleic acid sequence (e.g., SEQ ID NOs: 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893) within the albumin gene or introns of the albumin gene. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 85% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 95% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893.In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 98% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 99% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having 100% identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893.

[0168] In some embodiments, the guide polynucleotide hybridizes or targets sequences within the albumin gene or within the introns of the albumin gene (e.g., SEQ ID NOs: 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994). In some embodiments, the guide polynucleotide hybridizes or targets sequences of any one of SEQ ID NOs: 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994, or sequences having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 80% identity with any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 85% identity with any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 90% identity with any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 95% identity with any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 96% identity with any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994.In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 97% identity with any one of SEQ ID NOs: 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 98% identity with any one of SEQ ID NOs: 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 99% identity with any one of SEQ ID NOs: 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994. In some embodiments, the guide polynucleotide hybridizes to or targets a sequence that has 100% identity with any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994.

[0169] In some embodiments, the engineered guide polynucleotide is configured to form a complex with an endonuclease. In some cases, the engineered guide polynucleotide includes a spacer sequence. In some cases, the spacer sequence is configured to hybridize to a target nucleic acid sequence. In some cases, the endonuclease is configured to bind to a protospacer-adjacent motif (PAM) sequence.

[0170] In some embodiments, the manipulated guide polynucleotide comprises a DNA targeting segment containing a nucleotide sequence complementary to the target sequence of the target nucleic acid site, and a protein-binding segment containing two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) double helix. In some embodiments, the two complementary stretches of nucleotides are covalently bonded to each other by an intervening nucleotide. In some embodiments, the manipulated guide polynucleotide can form an endonuclease (e.g., a class 2, type V Cas endonuclease). In some embodiments, the DNA targeting segment is positioned at the 3' end of both of the two complementary stretches of nucleotides.

[0171] In some cases, the endonuclease is not a Cpf1 or Cms1 endonuclease.

[0172] In some cases, the endonuclease is configured to bind to the manipulated guide polynucleotide. In some cases, the Cas endonuclease is configured to bind to the manipulated guide polynucleotide. In some cases, class 2, type V Cas endonuclease is configured to bind to the manipulated guide polynucleotide. In some cases, class 2, type V, subtype Cas endonuclease is configured to bind to the manipulated guide polynucleotide.

[0173] In some embodiments, the guide polynucleotide is configured to form a complex with an endonuclease. In some embodiments, the guide polynucleotide binds to the endonuclease to form a complex. In some embodiments, the guide polynucleotide binds to the endonuclease (e.g., non-covalently through electrostatic interactions or hydrogen bonds) to form a complex. In some embodiments, the guide polynucleotide is fused to the endonuclease to form a complex.

[0174] In some cases, the guide polynucleotide contains a sequence complementary to the eukaryotic, fungal, plant, mammalian, or human genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the eukaryotic genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the fungal genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the plant genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the mammalian genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the human genome polynucleotide sequence.

[0175] In some cases, the guide polynucleotide includes a hairpin containing at least 8 base-pair ribonucleotides. In some cases, the guide polynucleotide includes a hairpin containing at least 9 base-pair ribonucleotides. In some cases, the guide polynucleotide includes a hairpin containing at least 10 base-pair ribonucleotides. In some cases, the guide polynucleotide includes a hairpin containing at least 11 base-pair ribonucleotides. In some cases, the guide polynucleotide includes a hairpin containing at least 12 base-pair ribonucleotides.

[0176] In some cases, the guide polynucleotide is 30 to 250 nucleotides long. In some cases, the guide polynucleotide is 42 to 44 nucleotides long. In some cases, the guide polynucleotide is 42 nucleotides long. In some cases, the guide polynucleotide is 43 nucleotides long. In some cases, the guide polynucleotide is 44 nucleotides long. In some cases, the guide polynucleotide is 85 to 245 nucleotides long. In some cases, the guide polynucleotide is more than 90 nucleotides long. In some cases, the guide polynucleotide is less than 245 nucleotides long.

[0177] In some embodiments, the guide polynucleotide comprises a synthetic or modified nucleotide. In some embodiments, the guide polynucleotide comprises one or more internucleoside linkers modified from natural phosphodiesters. In some embodiments, all or the entire sequence of internucleoside linkers of the guide polynucleotide is modified. For example, in some embodiments, the internucleoside linkage comprises sulfur (S), such as a phosphorothioate internucleoside linkage.

[0178] In some embodiments, the disclosure provides an engineered guide polynucleotide comprising a DNA targeting segment. In some cases, the DNA targeting segment comprises a nucleotide sequence complementary to a target sequence. In some cases, the target sequence is located in a target DNA molecule. In some cases, the engineered guide polynucleotide comprises a protein-binding segment. In some cases, the protein-binding segment comprises two complementary stretches of nucleotides. In some cases, the two complementary stretches of nucleotides hybridize to form a double-stranded RNA (dsRNA) double helix. In some cases, the two complementary stretches of nucleotides are covalently bonded to each other by an intervening nucleotide. In some cases, the engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease. In some cases, the endonuclease has at least approximately 20%, at least approximately 25%, at least approximately 30%, at least approximately 35%, at least approximately 40%, at least approximately 45%, at least approximately 50%, at least approximately 55%, at least approximately 60%, at least approximately 65%, at least approximately 70%, at least approximately 75%, at least approximately 80%, at least approximately 85%, at least approximately 90%, at least approximately 91%, at least approximately 92%, at least approximately 93%, at least approximately 94%, at least approximately 95%, at least approximately 96%, at least approximately 97%, at least approximately 98%, or at least approximately 99% identity with one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629. In some cases, the complex targets a target sequence of the target DNA molecule. In some cases, the DNA targeting segment is located at the 3' end of both of the two complementary stretches of the nucleotide.

[0179] In some cases, a double-stranded RNA (dsRNA) double helix contains at least 8 ribonucleotides. In some cases, a double-stranded RNA (dsRNA) double helix contains at least 9 ribonucleotides. In some cases, a double-stranded RNA (dsRNA) double helix contains at least 10 ribonucleotides. In some cases, a double-stranded RNA (dsRNA) double helix contains at least 11 ribonucleotides. In some cases, a double-stranded RNA (dsRNA) double helix contains at least 12 ribonucleotides.

[0180] In some embodiments, the guide polynucleotide comprises modifications to a ribose sugar or nucleic acid base. In some embodiments, the guide polynucleotide comprises one or more nucleosides containing a modified sugar moiety, where the modified sugar moiety is a modification of the sugar moiety compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, substitution with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acid (LNA)), or an unbound ribose ring typically lacking a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside in which the sugar moiety is replaced with a non-sugar moiety, e.g., a peptide nucleic acid (PNA) or a morpholino nucleic acid.

[0181] In some embodiments, the guide polynucleotide includes one or more modified sugars. In some embodiments, the sugar modification includes modifications made by altering substituents on the ribose ring to non-hydrogen groups or 2'-OH groups naturally found in DNA and RNA nucleosides. In some embodiments, substituents are introduced at the 2', 3', 4', or 5' positions, or combinations thereof. In some embodiments, the nucleoside having a modified sugar moiety includes 2'-modified nucleosides, e.g., 2'-substituted nucleosides. In some embodiments, the 2'-sugar-modified nucleoside is a nucleoside having a substituent other than -H or -OH at the 2' position (2'-substituted nucleoside), or includes a 2'-linked biradical and includes 2'-substituted nucleosides and LNA (2'-4' biradical-bridged) nucleosides. Examples of 2'-substituted nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group involves a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).

[0182] In some embodiments, the guide polynucleotide contains one or more modified sugars. In some embodiments, the guide polynucleotide contains only modified sugars. In certain embodiments, the guide polynucleotide contains more than 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar contains a 2'-O-methoxyethyl group. In some embodiments, the guide polynucleotide contains both internucleoside linker modification and nucleoside modification.

[0183] In some cases, the guide polynucleotide contains a sequence complementary to the eukaryotic, fungal, plant, mammalian, or human genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the eukaryotic genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the fungal genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the plant genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the mammalian genome polynucleotide sequence. In some cases, the guide polynucleotide contains a sequence complementary to the human genome polynucleotide sequence.

[0184] In some cases, the guide polynucleotide is 30 to 400 nucleotides long. In some cases, the guide polynucleotide is 85 to 245 nucleotides long. In some cases, the guide polynucleotide is more than 90 nucleotides long. In some cases, the guide polynucleotide is less than 245 nucleotides long. In some embodiments, the guide polynucleotide is 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides long. In the embodiment where the guide polynucleotides are added, the number of guide polynucleotides is approximately 30-40, 30-50, 30-60, 30-70, 30-80, 30-90, 30-100, 30-120, 30-140, 30-160, 30-180, 30-200, 30-220, 30-240, 50-60, 50-70, 50-80, 50-90, and 50-100. The lengths of the nucleotides are approximately 50-120, 50-140, 50-160, 50-180, 50-200, 50-220, 50-240, 100-120, 100-140, 100-160, 100-180, 100-200, 100-220, 100-240, 160-180, 160-200, 160-220, or 160-240.

[0185] In some embodiments, the sequence may be determined by the BLASTP, CLUSTALW, MUSCLE, or MAFFT algorithm, or by the CLUSTALW algorithm with Smith-Waterman homology search algorithm parameters. In some embodiments, the sequence is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix with 3 word length (W), 10 expected value (E) parameters, and 11 existence, 1 extension for gap cost, and with conditional composition score matrix adjustment.

[0186] MG series In certain embodiments, an engineered nuclease system comprising an endonuclease and an engineered guide polynucleotide is described herein. In some embodiments, the engineered guide polynucleotide comprises tracrRNA. In some embodiments, the engineered guide polynucleotide comprises a guide nucleic acid (e.g., gRNA). When T is referred to in polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA.

[0187] In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 70% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 70% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 75% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 75% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 80% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 80% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 85% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 85% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 90% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 90% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 95% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 95% identity with any one of SEQ ID NOs: 333-335 and 355-357.In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 96% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 96% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 97% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 97% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 98% identity with any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide containing a sequence having at least about 98% identity with any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence that is at least about 99% identical to any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide having a sequence that is at least about 99% identical to any one of SEQ ID NOs: 333-335 and 355-357. In some embodiments, the manipulated nuclease system comprises an endonuclease having 100% identity to any one of SEQ ID NOs: 6-14, and a manipulated polynucleotide having 100% identity to any one of SEQ ID NOs: 333-335 and 355-357.

[0188] In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 70% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 70% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 75% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 75% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 80% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 80% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 85% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 85% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 90% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 90% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 95% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 95% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 96% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 96% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 97% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 97% identity with SEQ ID NOs: 410-411.In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 98% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 98% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing a sequence having at least about 99% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing a sequence having at least about 99% identity with SEQ ID NOs: 410-411. In some embodiments, the manipulated nuclease system includes an endonuclease containing 100% identity with SEQ ID NO: 15 and a manipulated polynucleotide containing 100% identity with SEQ ID NOs: 410-411.

[0189] In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 70% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 70% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 75% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 75% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 80% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 80% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 85% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 85% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 90% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 90% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 95% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 95% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413.In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 96% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 96% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 97% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 97% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 98% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 98% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 99% identity with any one of SEQ ID NOs: 16-29, and a manipulated polynucleotide containing a sequence having at least about 99% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. In some embodiments, the manipulated nuclease system comprises an endonuclease having 100% identity with any one of SEQ ID NOs: 16-29 and a manipulated polynucleotide having 100% identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413.

[0190] In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 70% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 70% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 75% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 75% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 80% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 80% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731.In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 85% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 85% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 90% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 90% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 95% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 95% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731.In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 96% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 96% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence that is at least about 97% identical to any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence that is at least about 97% identical to any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence having at least about 98% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence having at least about 98% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731.In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence that is at least about 99% identical to any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having a sequence that is at least about 99% identical to any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731. In some embodiments, the manipulated nuclease system comprises an endonuclease having 100% identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629, and a manipulated polynucleotide having 100% identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731.

[0191] In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 70% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 70% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 75% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 75% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 80% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 80% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 85% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 85% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence that is at least about 90% identical to any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide having a sequence that is at least about 90% identical to any one of sequence numbers 1091-1113, 1119-1120, and 1876.In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 95% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 95% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 96% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 96% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 97% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 97% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease containing a sequence having at least about 98% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide containing a sequence having at least about 98% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876. In some embodiments, the manipulated nuclease system comprises an endonuclease having a sequence that is at least about 99% identical to any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide having a sequence that is at least about 99% identical to any one of sequence numbers 1091-1113, 1119-1120, and 1876.In some embodiments, the manipulated nuclease system comprises an endonuclease having 100% identity with any one of sequence numbers 1065-1090 and 1114-1118, and a manipulated polynucleotide having 100% identity with any one of sequence numbers 1091-1113, 1119-1120, and 1876.

[0192] In this specification, in some embodiments, engineered nuclease systems comprising an endonuclease and a DNA methyltransferase provided herein are further described. In some embodiments, the DNA methyltransferase is non-covalently bound to the endonuclease. In some embodiments, the DNA methyltransferase is fused to the endonuclease in a single polypeptide. In some embodiments, the DNA methyltransferase comprises Dmnt3A or Dnmt3L. In some embodiments, the engineered nuclease system further comprises a KRAB domain. In some embodiments, the KRAB domain is non-covalently bound to the endonuclease or DNA methyltransferase. In some embodiments, the KRAB domain is covalently bound to the endonuclease or DNA methyltransferase. In some embodiments, the KRAB domain is fused to the endonuclease or DNA methyltransferase in a single polypeptide.

[0193] cell In certain embodiments, cells comprising a class 2, type V nuclease system described herein are described herein.

[0194] In some embodiments, the cells include eukaryotic cells (e.g., plant cells, animal cells, protist cells, or fungal cells), mammalian cells (Chinese hamster ovary (CHO) cells, baby hamster kidney (BHK), human fetal kidney (HEK), mouse myeloma (NS0), or human retinal cells), immortalized cells (e.g., HeLa cells, COS cells, HEK-293T cells, MDCK cells, 3T3 cells, PC12 cells, Huh7 cells, HepG2 cells, K562 cells, N2a cells, or SY5Y cells), insect cells (e.g., Spodoptera frugiperda cells, Trichoplusia ni cells, Drosophila melanogaster cells, S2 cells, or Heliothis virescens cells), and yeast cells (e.g., Saccharomyces). These include cerevisiae cells, Cryptococcus cells, or Candida cells, plant cells (e.g., parenchymal cells, plagioclase cells, or plagioclase cells), fungal cells (e.g., Saccharomyces cerevisiae cells, Cryptococcus cells, or Candida cells), or prokaryotic cells (e.g., E. coli cells, Streptococcus bacterial cells, Streptomyces soil bacterial cells, or archaeal cells). In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are immortalized cells. In some embodiments, the cells are insect cells. In some embodiments, the cells are yeast cells. In some embodiments, the cells are plant cells. In some embodiments, the cells are fungal cells. In some embodiments, the cells are prokaryotic cells.

[0195] In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the primary cells are T cells. In some embodiments, the primary cells are hematopoietic stem cells (HSCs).

[0196] Delivery and vector In some embodiments, nucleic acid sequences encoding the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, or MG126 systems, including class 2, type V effector gRNAs, or gene editing systems disclosed herein, are disclosed herein.

[0197] In some embodiments, the nucleic acid encoding the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, or MG126 series is DNA, such as linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid encoding the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, or MG126 series is RNA, such as mRNA.

[0198] In some embodiments, nucleic acids encoding the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, or MG126 series are delivered by nucleic acid-based vectors. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., a circular DNA molecule that can autonomously replicate inside a cell), a cosmid (e.g., a pWE or sCos vector), an artificial chromosome, a human artificial chromosome (HAC), a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), a P1-derived artificial chromosome (PAC), a phagemid, a phage derivative, a bacmid, or a virus. In some embodiments, the nucleic acid-based vectors are pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4, pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, and pRI 101-AN. The selection is made from a list consisting of DNA, pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.

[0199] In some embodiments, the nucleic acid-based vector includes a promoter. In some embodiments, the promoter is selected from the group consisting of mini-promoters, inducible promoters, constitutive promoters, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments, the promoter is the U6 promoter. In some embodiments, the promoter is the CAG promoter.

[0200] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpesvirus, poxvirus, anerovirus, bocavirus, vacciniavirus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anerovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vacciniavirus. In some embodiments, the virus is a retrovirus.

[0201] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh 10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, The embodiments include AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or derivatives thereof. In some embodiments, the herpesvirus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0202] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof.In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.

[0203] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.

[0204] In some embodiments, nucleic acids encoding class 2, type V effectors, or genome editing systems are delivered by non-nucleic acid-based delivery systems (e.g., non-viral delivery systems). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is associated with lipids. In some embodiments, the nucleic acid associated with lipids is encapsulated within the aqueous interior of a liposome, dispersed within the lipid bilayer of a liposome, attached to a liposome via binding molecules associated with both the liposome and the nucleic acid, confined within a liposome, complexed with a liposome, dispersed in a lipid-containing solution, mixed with lipids, combined with lipids, contained as a suspension in lipids, contained in micelles, complexed with micelles, or otherwise associated with lipids. In some embodiments, the nucleic acid is contained in lipid nanoparticles (LNPs).

[0205] In some embodiments, the class 2, type V effector or genome editing system is introduced into the cell in any preferred manner, either stably or transiently. In some embodiments, the class 2, type V effector or genome editing system is transfected into the cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct encoding the class 2, type V effector or genome editing system. For example, the cell is transduced (e.g., using a virus encoding the class 2, type V effector or genome editing system) or transfected (e.g., using a plasmid encoding the class 2, type V effector or genome editing system) or translated into the class 2, type V effector or genome editing system with a nucleic acid encoding the class 2, type V effector or genome editing system. In some embodiments, the transduction is stably or transiently transduced. In some embodiments, cells expressing a class 2, type V effector or genome editing system, or containing a class 2, type V effector or genome editing system, are transfected or transfected with one or more gRNA molecules, for example, when the class 2, type V effector or genome editing system contains a CRISPR nuclease. In some embodiments, plasmids expressing a class 2, type V effector or genome editing system are introduced into cells by electroporation, transient (e.g., lipofection), stable genome integration (e.g., piggybac), and viral transduction (e.g., lentivirus or AAV), or other methods known to those skilled in the art. In some embodiments, the gene editing system is introduced into cells as one or more polypeptides. In some embodiments, delivery is achieved through the use of an RNP complex. For example, methods for delivering polypeptides and / or RNPs to cells by electroporation or cell compression are known in the art.

[0206] Exemplary methods for nucleic acid delivery include lipofection, nucleofection, electroporation, stable genome integration (e.g., piggybac), microinjection, biolistek, virosomes, liposomes, immunoliposomes, polycationic or lipid nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced incorporation of DNA. Lipofection is described, for example, in U.S. Patents 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam®, Lipofectin®, and SF Cell Line 4D-Nucleofector X Kit® (Lonza)). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those of WO91 / 17424 and WO91 / 16024. In some embodiments, delivery is to cells (e.g., in vitro or ex vivo administration) or target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in liposomes or nanoparticles that specifically target host cells.

[0207] Additional methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, US2003 / 0087817.

[0208] In some embodiments, this disclosure provides cells containing a vector or nucleic acid described herein. In some embodiments, the cells express a gene editing system or a part thereof. In some embodiments, the cells are human cells. In some embodiments, the cells are genome-edited ex vivo. In some embodiments, the cells are genome-edited in vivo.

[0209] How to use The gene editing system of the present disclosure can be used for various applications such as nucleic acid editing (e.g., gene editing) or binding to nucleic acid molecules (e.g., sequence-specific binding). Such systems can, for example, correct (e.g., remove or replace) genetically inherited mutations that can cause diseases in a subject, inactivate genes to confirm their function in cells, as diagnostic tools to detect gene elements that cause diseases (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivated enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate a virus by targeting its genome or to prevent it from infecting host cells, to add genes or modify metabolic pathways to engineer organisms to produce beneficial small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and / or to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors.

[0210] In certain embodiments, methods for modifying a target nucleic acid site are described herein, including providing a Class 2, Type V effector, or genome editing system disclosed herein. In some embodiments, modifying a target nucleic acid site includes binding, nicking, cleaving, marking, modifying, or translocating the target nucleic acid site. In some cases, the endonuclease induces a single-strand break or a double-strand break at or proximal to the target nucleic acid site. In some cases, the endonuclease induces alternating single-strand breaks within or 3' to the target nucleic acid site.

[0211] In some embodiments, the target nucleic acid is double-stranded. In some embodiments, the target nucleic acid is single-stranded. In some cases, the target nucleic acid site contains deoxyribonucleic acid (DNA). In some cases, the target nucleic acid site is double-stranded DNA. In some cases, the target nucleic acid site contains ribonucleic acid (RNA). In some cases, the target nucleic acid site contains genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some cases, the target nucleic acid site is in vitro. In some cases, the target nucleic acid site is intracellular. In some cases, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is genome-edited ex vivo. In some embodiments, the cell is genome-edited in vivo.

[0212] In some embodiments, the method is used to introduce a modification into the genome of a cell. In some embodiments, the modification is an insertion, a deletion, or a mutation. In some embodiments, the method is used to introduce site-specific insertions, deletions, and / or mutations (e.g., insertions and mutations) into the genome of a cell. In some embodiments, the method is used in combination with a nucleic acid template to facilitate site-specific insertion into the genome of a cell.

[0213] In some embodiments, the method is used to bind, cleave, mark, or modify double-stranded deoxyribonucleic acid polynucleotides. In some embodiments, the method comprises contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 Cas endonuclease. In some cases, the endonuclease is a class 2, V-type Cas endonuclease. In some cases, the endonuclease is a class 2, V-type, sub-type Cas endonuclease. In some cases, the endonuclease is complexed with an engineered guide RNA. In some cases, the engineered guide RNA is configured to bind to the endonuclease. In some cases, the engineered guide RNA is configured to bind to the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the engineered guide RNA is configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide. In some cases, double-stranded deoxyribonucleic acid polynucleotides contain protospacer-adjacent motifs (PAMs).

[0214] In certain embodiments, a method for modifying TRAC is described herein, comprising contacting TRAC with an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 767-798. In some embodiments, the engineered guide polynucleotide is encoded by any one of SEQ ID NOs: 767-798, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 767-798. In some embodiments, the manipulated guide polynucleotide is configured to hybridize to or target a sequence complementary to a sequence containing any one of SEQ ID NOs. 767-798, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs. 767-798. In some embodiments, the target nucleic acid sequence contains a sequence having any one of SEQ ID NOs. 799-830. In some embodiments, the target nucleic acid sequence contains a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs. 799-830.

[0215] In a particular embodiment, a method for modifying APOA1 is described herein, comprising contacting APOA1 with an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the manipulated guide polynucleotide is encoded by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 767-798, or any one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the manipulated guide polynucleotide is configured to hybridize with or target a sequence complementary to a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669. In some embodiments, the target nucleic acid sequence includes a sequence having at least one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676. In some embodiments, the target nucleic acid sequence includes a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 905-978, 1022-1064, 1238-1256, and 1670-1676.

[0216] In a particular embodiment, a method for modifying AAVS1, comprising contacting APOA1 using an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the manipulated guide polynucleotide is encoded by one of the sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779, or by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with one of the sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the manipulated guide polynucleotide is configured to hybridize with or target a sequence complementary to a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 1257-1324, 1523-1562, 1677-1686, and 1753-1779. In some embodiments, the target nucleic acid sequence includes a sequence having at least one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806. In some embodiments, the target nucleic acid sequence includes a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 1325-1392, 1563-1602, 1687-1696, and 1780-1806.

[0217] In a particular embodiment, a method for modifying albumin, comprising contacting albumin with an engineered nuclease system, wherein the engineered nuclease system comprises: a) an endonuclease having at least 80% sequence identity with any one of SEQ ID NOs: 30-150, 420-431, 476-624, and 629; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide has at least 80% sequence identity with any one of SEQ ID NOs: 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the manipulated guide polynucleotide is encoded by any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893, or by a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893. In some embodiments, the manipulated guide polynucleotide is configured to hybridize with a sequence complementary to a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1893, or to target a sequence. In some embodiments, the target nucleic acid sequence includes a sequence having at least one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994.In some embodiments, the target nucleic acid sequence includes a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 1170-1218, 1442-1490, 1633-1662, 1888, 1890, and 1994.

[0218] In certain embodiments, methods for producing the engineered nuclease system or its components described herein are further described herein. In some embodiments, the method includes culturing host cells using the engineered nuclease system or its components described herein.

[0219] In some embodiments, the host cell is a bacterial cell. In some embodiments, the bacterial cell is Bifidobacterium longum, Bifidobacterium lactis, Bifidobacterium animalis, Bifidobacterium breve, Bifidobacterium infantis, Bifidobacterium adolescentis, Lactobacillus acidophilus, Lactobacillus casei, Lactobacillus paracasei, Lactobacillus salivarius, Lactobacillus reuteri, Lactobacillus rhamnosus, Lactobacillus johnsonii, Lactobacillus plantarum, Lactobacillus fermentum, Lactococcus lactis, Streptococcus thermophilus, Lactococcus lactis, Lactococcus diacetylactis, Lactococcus cremoris, Lactobacillus bulgaricus, Lactobacillus helveticus, Lactobacillus The host cell is *Escherichia delbrueckii* or *Escherichia coli*. In some embodiments, the host cell is *E. coli* cells. In some embodiments, the *E. coli* cells are λDE3 lysogen or the BL21(DE3) strain. In some embodiments, the *E. coli* cells have the ompT lon genotype.

[0220] In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is a λDE3 lysogen, or the E. coli cell is the BL21(DE3) strain. In some embodiments, the E. coli cell has the ompT lon genotype.

[0221] In some embodiments, an open reading frame is operably linked to a promoter sequence. In some embodiments, the promoter is selected from the group consisting of mini-promoters, inducible promoters, constitutive promoters, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and derivatives thereof.

[0222] In some embodiments, the open reading frame includes the T7 promoter sequence, T7-lac promoter sequence, lac promoter sequence, tac promoter sequence, trc promoter sequence, ParaBAD promoter sequence, PrhaBAD promoter sequence, T5 promoter sequence, cspA promoter sequence, and araP BAD It is operably linked to a promoter, a strong leftward promoter (pL promoter) from a phage lambda, or any combination thereof.

[0223] In some embodiments, the open reading frame includes a sequence encoding an affinity tag in-frame to a sequence encoding an engineered nuclease system or its components as described herein. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose-binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is in-frame to a sequence encoding an engineered nuclease system or its components as described herein via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof.

[0224] In some embodiments, the open reading frame is codon-optimized for expression in host cells. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is incorporated into the genome of a host cell.

[0225] In some embodiments, the present disclosure provides a culture comprising host cells described herein in a suitable liquid medium.

[0226] In some embodiments, the disclosure provides a method for producing an engineered nuclease system or its components as described herein, comprising culturing one of the host cells described herein in a suitable growth medium. In some embodiments, the method further comprises inducing the expression of an engineered nuclease system or its components as described herein by adding an additional chemical agent or an increased amount of nutrients. In some embodiments, the additional chemical agent or increased amount of nutrients comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to IMAC or ion affinity chromatography. In some embodiments, the open reading frame comprises a sequence encoding an IMAC affinity tag in-frame linked to a sequence encoding an engineered nuclease system or its components as described herein. In some embodiments, the IMAC affinity tag is in-frame linked to a sequence encoding an engineered nuclease system or its components as described herein, via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site includes a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the method further includes cleaving the IMAC affinity tag by contacting the engineered nuclease system or its components as described herein with a protease corresponding to the protease cleavage site. In some embodiments, the method further includes removing the affinity tag from a composition containing an engineered nuclease system or its components as described herein by performing subtractive IMAC affinity chromatography.

[0227] kit In some embodiments, the Disclosure provides a kit comprising one or more nucleic acid constructs encoding various components of the Class 2, V-type effector and genome editing system described herein, for example, a nucleotide sequence encoding a component of the Class 2, V-type effector and genome editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequence includes a heterologous promoter that drives the expression of the RNA genome editing system component.

[0228] In some embodiments, gene editing systems comprising class 2, type V effectors, gRNAs, or any combination thereof disclosed herein are incorporated into pharmaceutical, diagnostic, or research kits to facilitate their use in therapeutic, diagnostic, or research applications. The kit may comprise one or more containers containing any of the vectors disclosed herein, and instructions for use.

[0229] The kit may be designed to facilitate researchers' use of the methods described herein and may take various forms. Each of the components of the kit may be provided in liquid form (e.g., in solution) or solid form (e.g., dry powder), where applicable. In certain cases, some of the components may be configurable (e.g., into an active form) or otherwise treatable by the addition of a suitable solvent or other type (e.g., water or cell culture medium), which may or may not be provided with the kit. Where used herein, “instructions” defines the components of the instructions and / or promotional materials and may typically be accompanied by written instructions on or associated with the packaging of the herein. Instructions may also consist of any oral or electronic instructions provided in any format that clearly indicates to the user that the instructions relate to the kit, such as audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communications. In some embodiments, written instructions may be in the form prescribed by a government agency that regulates the manufacture, use, or sale of a pharmaceutical or biological product, and such instructions may also reflect approval by an agency for manufacture, use, or sale for animal administration. [Examples]

[0230] Example 1 - A method for metagenomic analysis of a novel protein. Metagenomic samples were collected from sediments, soil, and animals. Deoxyribonucleic acid (DNA) was extracted using a DNA miniprep kit and sequenced. Samples were collected with the consent of the owners. Further raw sequence data from public sources included sequences of animal microbiomes, sediments, soil, hot springs, hydrothermal vents, oceans, peat bogs, permafrost, and sewage. To identify new Cas effectors, metagenomic sequence data was searched using a hidden Markov model generated based on known Cas protein sequences that included class 2, type V Cas effector proteins. Effector proteins identified by the search were aligned against known proteins to identify potential active sites. This metagenomic workflow led to the description of the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, and MG126 families described herein.

[0231] Example 2 - Discovery of the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, and MG126 Families of the CRISPR System <000(https: / / patents.google.com / patent / US20190276732A1 / en?oq=US20190276732A1)0967>Analysis of the data from the metagenomic analysis of Example 1 revealed a new cluster of putative CRISPR systems that included nine families (MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, and MG126). The corresponding protein sequences and nucleic acid sequences of these new enzymes and their exemplary subdomains are presented as SEQ ID NOs: 1 - 325, 420 - 431, 476 - 624, or 629. )0967>

[0232] Example 3 - Template DNA for Transcription and Translation E coli codon-optimized sequences for all MG VU and CasPhi nucleases were prepared using plasmids containing the T7 promoter. Linear templates were amplified from the plasmids by PCR to include the T7 and nuclease sequences. Minimal array linear templates were amplified from sequences consisting of the T7 promoter, native repeats, universal spacers, and native repeats, flanked by adapter sequences for amplification. The universal spacers correspond to spacers in an 8N target library, which have 8N mixed bases flanked by spacers for PAM determination. Three intergene sequences near ORF or CRISPR arrays were identified from metagenomic contigs and ordered as gBlocks with flanking adapter sequences for amplification.

[0233] Example 4 - In vitro transcription of crRNA, minimal array, and sgRNA. RNA was produced by in vitro transcription using an RNA Synthesis Kit and purified using an RNA Cleanup Kit. Various templates were used for T7 transcription. For crRNA, DNA oligos were designed using the T7 promoter, trimmed native repeats, and a universal spacer. The same templates were used for minimal arrays. For sgRNA, DNA ultramers were designed using the T7 promoter, trimmed tracrRNA, GAAA tetraloops, trimmed native repeats, and a universal spacer. Minimal array templates were amplified with adapter primers. The crRNA and sgRNA templates were ordered as reverse complements and annealed in 1× double-stranded buffer with primers containing the T7 promoter sequence at 95°C for 2 minutes, followed by cooling to 22°C at 0.1°C / sec to produce hybrid ds / ssDNA substrates suitable for transcription. After transcription, but before washing, each reactant was treated with DNAse I and incubated at 37°C for 15 minutes. All transcripts were validated for yield and purity via RNA electrophoresis or denatured urea PAGE gel.

[0234] Example 5 - TXTL expression Nucleases, intergenetic sequences, and minimal arrays were expressed in a transcription-translation reaction mixture. The final reaction mixture contained 5 nM nuclease DNA template, 12 nM intergenetic DNA template, 15 nM minimal array DNA template, 0.1 nM pTXTL-P70a-T7rnap, and 1X Master Mix. The reaction mixture was incubated at 29°C for 16 hours and then stored at 4°C.

[0235] Example 6 - Protein Expression A 10 nM nuclease PCR template was expressed at 37°C for 3 hours using an in vitro protein synthesis kit for cleavage with in vitro transcribed RNA. These reaction products were then used to test in vitro cleavage with 50 nM sgRNA or minimal array RNA, following the same procedure as described in the cleavage reaction section.

[0236] Example 7 - E. coli expression Plasmids encoding effectors, intergenetic sequences from genomic contigs, native repeats, and a universal spacer sequence with a T7 promoter were transformed into BL21 DE3 or T7 Express lysY / Iq and cultured at 37°C in 60 mL of terrific broth medium supplemented with 100 μg / mL ampicillin. After the culture reached 0.5 OD600 nm, expression was induced with 0.4 mM IPTG and incubated overnight at 16°C. 25 mL of cells were pelleted by centrifugation and resuspended in 1.5 mL of lysis buffer (20 mM Tris-HCl, 500 mM NaCl, 1 mM TCEP, 5% glycerol, and 10 mM MgCl2 pH 7.5 containing a protease inhibitor). Cells were then lysed by sonication. The supernatant and cell debris were separated by centrifugation.

[0237] Example 8 - Cutting reaction Plasmid library DNA cleavage was performed by mixing a 5 nM target library, a 5-fold dilution of protein expression, 10 nM Tris-HCl, 10 nM MgCl2, and 100 mM NaCl at 37°C for 2 hours. For the reaction with E. coli expression, 10 μL of clarified lysate was added. The reaction was stopped, cleaned with PCR cleanup beads, and eluted in pH 8.0 Tris-EDTA buffer. The 3 nM cleavage product ends were blunted at 25°C for 15 minutes using 3.33 μM dNTPs, 1×T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment. The 1.5 nM cleavage product was ligated at room temperature for 20 minutes using a 150 nM adapter, 1×T4 DNA ligase buffer, and 20 U / μL T4 DNA ligase. The ligated product was amplified by PCR using NGS primers, sequenced by NGS, and PAM was obtained. The in vitro activity of MG119-2 is shown in Figure 9, while the PAM determination for MG119-2 is shown in Figure 10.

[0238] Example 9 - RNA-seq library preparation for intergenetic enrichment from TXTL and E. coli lysates RNA was extracted according to the RNA Miniprep Kit, and the expression product from the cell lysates was eluted in 30-50 μL of water. The total concentration of the transcript was measured.

[0239] Total RNA (100 ng–1 μg) from each sample was prepared for RNA sequencing. Amplicons of 150–300 bp were quantified and pooled to a final concentration of 4 nM. The final concentration of 12.5 pM was loaded into a sequencing kit and sequenced over 176 total cycles. The tracr sequences of the genes were identified using RNAseq reads.

[0240] Example 10 - Predicted RNA folding The predicted RNA folding of an active single RNA sequence was calculated at 37°C. The shading of a base corresponds to the probability of base pairing for that base.

[0241] Example 11 - In vitro cutting efficiency In this case, the protein is expressed in E. coli protease-deficient strain B under a T7-inducible promoter, the cells are lysed using sonication, and the His-tagged protein of interest is purified using Ni-NTA affinity chromatography on FPLC. Purity is determined by SDS-PAGE and densitometry of the separated protein bands on Coomassie-stained acrylamide gel. The protein is desalted in a storage buffer consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol, pH 7.5, and stored at -80°C.

[0242] Construct target DNA containing a spacer sequence and a PAM determined via NGS. If degenerate bases are present within the PAM, select a single representative PAM for testing. The target DNA is a 2200 bp linear DNA derived from a plasmid via PCR amplification. The PAM and spacer are located 700 bp from one end. Successful cleavage yields 700 and 1500 bp fragments.

[0243] The target DNA, in vitro transcribed single RNA, and purified recombinant protein are combined in a cleavage buffer (10 mM Tris, 100 mM NaCl, 10 mM MgCl2) containing excess protein and RNA, and incubated for 5 minutes to 3 hours, usually 1 hour. The reaction is stopped by the addition of RNAse A and incubation at 60°. The reaction products are separated on a 1.2% TAE agarose gel, and the fraction of cleaved target DNA is quantified using imaging software.

[0244] Example 12 - Activity in E. coli To test nuclease activity in bacterial cells, a strain is constructed with a genome sequence containing a target spacer and corresponding PAM sequence specific to the enzyme of interest. The engineered strain is then transformed with the nuclease of interest, and the transformants are subsequently made chemocompetent and transformed with a 50 ng single guide that is either on-target or off-target. After heat shock, the transformants are recovered at 37°C for 2 hours in state of care (SOC), and nuclease efficiency is determined by a 5-fold dilution series grown on induction medium. Colonies are quantified from the dilution series in triplicates.

[0245] Example 13 - Activity in mammalian cells To demonstrate targeting and cleavage activity in mammalian cells, the protein sequence is cloned into two mammalian expression vectors: one with a C-terminal SV40 NLS and 2A-GFP tag, and the other without a GFP tag and with two NLS sequences (one at the N-terminus and one at the C-terminus). Alternative NLS sequences may also be used. The protein's DNA sequence may be the native sequence, an E. coli codon-optimized sequence, or a mammalian codon-optimized sequence. A single guide polynucleotide sequence with the target gene of interest is also cloned into a mammalian expression vector. The two plasmids are co-transfected into HEK293T cells. 72 hours after co-transfection of the expression plasmid and the sgRNA targeting plasmid into HEK293T cells, DNA is extracted and used for NGS library preparation. The NHEJ percentage is measured via indels in the sequencing of target sites to indicate the enzyme's targeting efficiency in mammalian cells. At least 10 different target sites are selected to test the activity of each protein.

[0246] Example 14 - Characterization of compact V-type nucleases in the MG119 family In silico identification of compact V-type nucleases in the MG119 family. The discovery of predictive proteins associated with nuclease sequences in the MG119 family of compact V-type nucleases was based on homology searches. The search was performed, and hits of V-type nuclease sequences were retained if they met the following criteria: (i) the hmmsearch e value was ≤10. -5 (ii) The gene encoding the nuclease was within 1 kb of the CRISPR array, and (iii) the amino acid sequence length was in the range of 350–700 aa. Sequences were clustered with 100% amino acid identity using coverage mode 1 and 80% coverage of the target sequence. Sequence representatives were selected to construct a phylogenetic tree for multiple sequence alignment. Careful investigation of individual clades on the phylogenetic tree, including the genomic context of the nuclease gene, led to the identification of several compact V-type nuclease sequences in the MG119 family (SEQ ID NOs. 476–624 and 629).

[0247] In vitro characterization for identifying predicted tracrRNA For example, to identify the putative tracrRNA sequence of nuclease MG119-2, the contiguous intergene sequence and minimal array were expressed in a transcription-translation reaction mixture using cell-free expression reactions (TXTL). The final reaction mixture contained 5 nM nuclease DNA template, 12 nM intergene DNA template, 15 nM minimal array DNA template, 0.1 nM pTXTL-P70a-T7rnap, and 1X Master Mix. The reaction mixture was incubated at 29°C for 16 hours and then stored at 4°C.

[0248] Ribonucleoprotein complexes were tested via in vitro cleavage reactions. Plasmid DNA library cleavage was performed by mixing a 5nM target plasmid DNA library representing all possible 8N PAMs, a 5-fold dilution of TXTL expression, 10nM Tris-HCl, 10nM MgCl2, and 100mM NaCl at 37°C for 2 hours. The reaction was stopped, cleaned with PCR cleanup beads, and eluted in Tris-EDTA buffer at pH 8.0.

[0249] To obtain the PAM sequence, the ends of 3 nM cleavage products were blunted at 25°C for 15 minutes using 3.33 μM dNTPs, 1 × T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment. 1.5 nM cleavage products were ligated at room temperature for 20 minutes using a 150 nM adapter, 1 × T4 DNA ligase buffer, and 20 U / μL T4 DNA ligase. The ligated products were amplified by PCR using NGS primers and sequenced by NGS.

[0250] To obtain tracrRNA and crRNA sequences, RNA was extracted from TXTL lysates according to the RNA Miniprep Kit and eluted in 30–50 μL of water. 100 ng–1 μg of total RNA from each sample was prepared for RNA sequencing using the Small RNA Library Prep Set. Amplicons of 150–300 bp were quantified and pooled to a final concentration of 4 nM. The 12.5 pM final concentration was loaded into a sequencing kit and sequenced over 176 total cycles. The tracr sequences of the genes were identified by mapping back to the original sequences using RNAseq reads.

[0251] In silico identification of tracrRNA sequences To identify additional non-coding regions containing potential tracrRNAs, the active tracrRNA sequence was mapped to other contigs containing nucleases from the same nuclease family (e.g., MG119-1 and MG119-3). A covariance model was generated using the newly identified sequences to predict additional tracrRNAs. The covariance model was constructed from multiple sequence alignments (MSAs) of the active tracrRNA sequence and the predicted tracrRNA sequence. Secondary structures of the MSAs were obtained, and the covariance model was constructed. Other contigs containing candidate nucleases were searched using the covariance model. TracrRNA candidates were tested in vitro (see below), and in iterative processes, the covariance model was improved using sequences from the active candidate to search for additional tracrRNAs in intergeneric regions associated with other nuclease candidates.

[0252] sgRNA design The predicted tracrRNAs and their associated CRISPR repeat sequences obtained from the covariance model were modified as follows to generate sgRNAs (Figure 11A): the 3' end of the predicted tracrRNA sequence and the 5' end of the repeat sequence were trimmed and then connected to a GAAA tetraloop.

[0253] In vitro cleavage reaction to confirm nuclease activity and enable PAM determination. A 5 nM nuclease-amplified DNA template and a 25 nM sgRNA-amplified DNA template (containing one of the spacer sequences listed in Table 2) were expressed in an in vitro protein expression system at 37°C for 3 hours. Plasmid library DNA cleavage was performed by mixing a 5 nM target library representing all possible 8N PAMs, a 5-fold dilution of the in vitro protein expression solution, 10 mM Tris-HCl pH 7.9, 10 mM MgCl2, 100 μg / mL BSA, and 50 mM NaCl at 37°C for 2 hours. The reaction was stopped, cleaned with PCR cleanup beads, and eluted in Tris-EDTA buffer at pH 8.0. The 3 nM cleavage product ends were blunted at 25°C for 15 minutes with 3.33 μM dNTPs, 1 × T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment. The 1.5 nM cleavage product was ligated at room temperature for 20 minutes using a 150 nM adapter, 1 × T4 DNA ligase buffer, and 20 U / μL T4 DNA ligase. The ligated product was amplified by PCR using NGS primers and sequenced by NGS to obtain the PAM. The active protein successfully cleaved in the PAM library produced bands of approximately 188 or 205 bp on the agarose gel, depending on which target site was encoded by the sgRNA (Figure 11B).

[0254] [Table 2]

[0255] The PAM recognized by the MG119 nuclease is shown as a sequence logo produced by the Seqlog maker (Figure 12). Table 3 lists preferred cleavage sites on the target strand for the protospacer sequence complementary to the U40 spacer.

[0256] [Table 3-1]

[0257] [Table 3-2]

[0258] Protein expression and purification Isolating pure and functional proteins is essential for extensive in vitro analysis of biochemical properties and mechanistic studies. Expression and purification of MG119 candidates were optimized to obtain sufficient quantities and quality of protein for such characterization. All constructs were expressed in E. coli. The constructs were expressed using either a pMGB expression vector (MBP fusion), a pMGBΔ expression vector (without fusion protein), or both.

[0259] Protein expression The protein expression protocols for the pMGB construct and the pMGBΔ construct were identical. Cultures were grown at 37°C in 2×YT medium (1.6% tryptone, 1% yeast extract, 0.5% NaCl) or TB medium with 100 μg / L carbenicillin. When the OD600 was approximately equal to 0.8–1.2, the cultures were induced with 0.5 mM IPTG and incubated overnight at 18°C ​​or for 4–6 hours at 24°C, depending on the construct. Next, the culture was collected by centrifugation at 6,000 × g for 10 minutes, and the pellet was resuspended in nickel-A buffer (50 mM Tris pH 7.5, 750 mM NaCl, 10 mM MgCl2, 20 mM imidazole, 0.5 mM EDTA, 5% glycerol, 0.5 mM TCEP) + protease inhibitor (protease inhibitor tablet, EDTA-free) and stored at -80°C.

[0260] Protein purification - pMGBΔ expression vector The protein expressed by this vector has the following sequence structure: 6xHis-(GS)2-PSP-nucleoplasmin binocular NLS-(GGS)1-(GS)1-MG119-X-(GGS)3-SV40 NLS (Table 5). The protein expressed by this vector is denoted as MG119-X△. The cell pellet was thawed and refilled to 120 mL with a Cf=0.5% n-octyl-β-D-glucoside surfactant. The sample was sonicated in an ice bath at 75% amplitude for a total processing time of 3 minutes using a 15-second on / 45-second off cycle. The lysate was clarified by centrifugation at 30,000 × g for 25 minutes, and the supernatant batch was conjugated to 5 mL of Ni-NTA resin for ≥20 minutes. The samples were loaded onto a gravity column, washed with 30 CV of nickel_A buffer, then eluted in 4 CV of nickel_B buffer (nickel_A buffer + 250 mM imidazole), and concentrated in a 50 kDa MWCO concentrator. Samples were collected throughout the entire purification process and electrophoresed on an SDS-PAGE protein gel, which was then imaged in an imaging system in an unstained channel after 5 minutes of UV activation (Figure 13A). The ΔMBP constructs were then loaded onto an S200i 10 / 300 GL column and flushed into nickel_A buffer (Figure 13B). The peak fractions were pooled and concentrated in a 50 kDa MWCO concentrator. Purification of proteins expressed with the pMGBΔ vector typically yielded 25–125 nmol of protein per L-expression culture (Figure 13F).

[0261] Protein purification - pMGB expression vector The protein expressed by this vector has the following sequence structure: 6xHis-(GS)1-MBP-(GS)1-TEV-nucleoplasmin binocular NLS-(GGGGS)3-(GS)1-MG119-X-(GGS)3-SV40 NLS (Table 5). The MBP fusion construct was purified to the same standard as the pMGBΔ protein by lysis, clarification, affinity purification, and elution with nickel B (Figure 13C). After protein concentration in a 50 kDa MWCO concentrator, TEV protease was added to each sample (Cf=1 UI / μL), and the samples were incubated overnight at 4°C with gentle rotation (rotating end-over-end). The samples were centrifuged (21,000 × g, 4°C, 10 min) to pelletize aggregates, and the supernatant was batch-conjugated to 3 mL of amylose resin at 4°C for 30 minutes, and then loaded onto a gravity column. The flow-through was collected and concentrated in a 50 kDa MWCO concentrator (Figure 13D). The sample was again centrifuged (21,000 × g, 4°C, 10 min) to pelletize aggregates, which were then loaded onto an S200i 10 / 300 GL column and flowed into nickel-A buffer (Figure 13E). The peak fraction was pooled and concentrated in a 50 kDa MWCO concentrator. The sample was collected throughout the entire purification process, run on an SDS-PAGE protein gel, and imaged in an imaging system in an unstained channel after 5 minutes of UV activation (Figure 13D).

[0262] Several selected MG119 candidates were purified from both pMGB and pMGB△ expression vectors. A comparison of the final protein yield, normalized to the initial expression culture volume, showed a trend towards higher expression yields from the pMGB△ vector (Figure 13E). Purification of the protein expressed with the pMGB△ vector typically yielded 2–15 nmol of protein per L expression culture (Figure 13E). The purified protein yields are shown in Table 4.

[0263] [Table 4]

[0264] [Table 5]

[0265] In vitro cleavage efficiency using purified proteins The active fraction of the protein aliquots was determined by a linear DNA substrate cleavage assay. Effector proteins were pre-incubated with 2x molar excess sgRNA at room temperature for 20 minutes to form ribonucleoprotein complexes (RNPs). The reaction was set up using titration of 25 nM DNA substrate and 0.25 × ~10 × molar excess RNP relative to the substrate. The reaction buffer composition was 10 mM Tris pH 7.5, 10 mM MgCl2, and 100 mM NaCl. The DNA substrate was 522 bp long. Successful cleavage yielded fragments of 172 and 350 bp. The reaction products were incubated at 37°C for 60 minutes, then at 75°C for 10 minutes. RNase was added to each reaction product (Cf = 0.33 μg / μL), and the samples were incubated at 37°C for 10 minutes. Proteinase K was added to each reactant (Cf = 60 units / mL), and the samples were incubated at 55°C for 15 minutes. The entire reactant was then electrophoresed on a 1.5% agarose gel containing GelGreen dye (Figure 14A), and imaged using an imaging system in the GelGreen channel. The percentage of cleaved substrate was calculated for each lane via densitometry analysis using imaging software. The active fraction was determined by the slope of the linear range of cleavage (Figure 14B).

[0266] In vitro cleavage of purified Hepa1-6 genomic DNA using purified protein To evaluate the cleavage of purified mouse Hepa1-6 genomic DNA (gDNA), the mouse albumin gene was targeted at intron 1 (Table 6). Using the Genomic DNA Mini kit, gDNA was extracted from a Hepa1-6 cell pellet containing 8 million cells and eluted in 10 mM Tris HCl at pH 8. sgRNA was prepared at 2 nmol and then resuspended at 20 μM in 10 mM Tris EDTA buffer (Table 6). Ribonucleoprotein (RNP) was prepared by pre-incubating 1× Effector Buffer (100 mM NaCl, 10 mM MgCl2, 10 mM Tris HCl, pH 7.5) at room temperature for 30 minutes with a nuclease in a 1:2 molar ratio with either a targeted or untargeted guide. All reactions were performed with three replicates, including a negative control without sgRNA. After RNP formation, the RNPs were added to a digest reaction mixture containing 20 ng / μL of purified gDNA in 1X effector buffer and incubated at 37°C for 1 hour. The nucleases were tested at two final concentrations: 7.8 and 15.6 nM. These concentrations were normalized by dividing the target concentration by the activity fraction of each nuclease. After incubation, these reaction mixtures were immediately transferred to 4°C, diluted 30× with water, and then prepared for qPCR containing 1× Master Mix, 10 μM forward primers, 10 μM reverse primers, and 5 μM 5'-FAM and ZEN / Iowa Black fluorescent quencher Taqman probe (Table 7). The real-time PCR system was used in the following cycles: 1) 95°C for 15 minutes, 2) 95°C for 5 seconds, and 3) 60°C for 1 minute, and steps 2-3 were repeated 40 times. Using the Cq value, the gDNA cleavage percentage for each reaction was calculated according to the cleavage percentage formula (see below). All were normalized against the untargeted control reaction. Figure 15A shows examples of average gDNA cleavage of 60% with MG119-28 and sgRNA3, and 21% with sgRNA2, at higher concentrations of the protein used. Percentage cutting formula Cut % = 100 - (2 -(Cq(実験的)-Cq(非標的化対照)) (x100)

[0267] [Table 6-1]

[0268] [Table 6-2]

[0269] [Table 7]

[0270] In vivo cleavage of genomic DNA in Hepa 1-6 cells using purified protein. Intracellular editing was demonstrated using an RNP complex of a guide and nuclease targeting the mouse albumin gene in intron 1 (Table 6). Hepa1-6 cells were thawed, washed, and resuspended in Dulbecco's modified Eagle medium (DMEM, 10% FBS, and 1% Pen-strep). Cells were then refrigerated in 30 mL of medium at 37°C, 4 × 10⁶ cells per 15 cm dish. 6 Cells were seeded at a specific density. After two days, when the cells reached 70-80% confluence, they were divided. The cells were trypsinized with 0.25% trypsin and incubated at 37°C for 30 seconds. DMEM was added, then the cells were divided into 3 mL portions and further diluted with 27 mL of culture medium. The divided cells were incubated for a further 2 days. Before nucleofection, the medium was aspirated from the plate, the cells were washed with 1× phosphate-buffered saline pH 7.2, and then trypsinized. The trypsin was neutralized, and the cells were resuspended in DMEM. The cells in the cell suspension were counted to calculate the volume of cells to be pelleted. A total of 100,000 cells were required for each downstream treatment. The cells were centrifuged at 300×g for 7 minutes, washed with PBS pH 7.2, and resuspended in the nucleofection solution.

[0271] RNP complexes were individually prepared by incubating 120 pmol of nuclease with 120 pmol of guide at room temperature for 90 minutes. 20 μL of the prepared cells were added to the RNP. Nucleofection was performed using a Nucleofector. Nucleofected cells were transferred from the nucleofection cassette to a 24-well plate containing 500 μL of medium in each well. After 2 days of incubation, gDNA from all treatments was extracted using the following cycle: 1) 65°C for 15 minutes, 2) 68°C for 15 minutes, and 3) 98°C for 10 minutes, and then kept at 4°C until use. The extracted gDNA was amplified using the following cycles for 30 cycles, targeting a 317 bp window: 1) 10 seconds at 98°C, 2) 1 second at 98°C, 3) 5 seconds at 63°C, 4) 15 seconds at 72°C, and 5) 1 minute at 72°C. Steps 2-5 were repeated, and the mixture was then held at 4°C. The amplicons were visualized on a 2% agarose gel, washed, and concentrated with magnetic beads having a bead volume of 1.8 × relative to the sample. The sample was eluted with water. Indels were sequenced by NGS (600 cycles, Table 8) using 5% phiX for 2 × 301 bp paired-end reads (minimum 20,000 reads per sample). Indel analysis was performed. The results are shown in Table 9 and Figure 15B.

[0272] [Table 8]

[0273] [Table 9]

[0274] Example 15 - Optimization of buffer solution for MG119 protein purification To date, the MG119 protein has been purified using Nickel A buffer. Nickel A buffer is unsuitable for downstream in vivo assays due to its high salt concentration, and rapid dilution to low-salt solutions induces protein precipitation. To optimize the buffer for protein stability and downstream assay compatibility, the MG119 nuclease is initially purified in a high-salt buffer (750 mM NaCl) and gradually washed to a Nickel A buffer variant containing 200 mM NaCl and the zwitterionic amino acids L-arginine (50 mM) and L-glutamate (50 mM). Empirically, various stabilizing sugars (ribose, sorbitol, mannitol, xylitol) have also been added to the buffer to enhance protein stability in low-salt buffers.

[0275] Example 16 - Fluorescence-based measurement of nuclease activity Cell line manipulation Current assays used to measure in vivo (i.e., in mammalian cell lines) nuclease activity require extensive data analysis and a turnaround time of up to one week. To facilitate the assessment of in vivo nuclease activity, immortalized mammalian cell lines are engineered to provide immediate data on genomic DNA editing. K562 mammalian cells grown in IMDM + 10% FBS are used in this assay. K562 mammalian cells are transfected with 1200 ng of plasmid (pUC backbone) containing 12 pmol of Cas9 protein, 60 pmol of sgRNA, and the expression sequence for mMBP-(GGS)3-eGFP protein. Genomic integration of this construct results in constitutive expression under a synthetic MND promoter. Cells are left to grow for 6 days and passaged every 3 days. Single-gene cell lines are isolated from single cells by sorting individual GFP-expressing cells into 96-well plates using a cell sorter.

[0276] Fluorescence-based in vivonuclease activity screening Appropriate sgRNAs are designed to direct nuclease cleavage along the mMBP and eGFP genes so that indel formation leads to frameshift mutations and loss of fluorescence. The MG119 RNP complex is formed by combining 100 pmol of protein and 200 pmol of sgRNA and incubating at room temperature for ≥20 minutes in 5 μL final volume. K562 cells are washed in 1×PBS and resuspended in nucleofector solution containing approximately 200,000 cells per well. Cells are combined with RNP in 25 μL final volume in a 96-well nucleofection plate, nucleofected (K562 cells), and harvested in IMDM + 10% FBS medium. Cells are left at 37°C for 2-3 days until recovered. For analysis, cells are washed twice with 1×PBS and then stained with 1×PBS + LIVE / DEAD dye for 20 minutes at room temperature. The cells are washed again with 1×PBS, resuspended in 1×PBS, and loaded into a flow cytometer for fluorescence analysis. Positive and negative fluorescence gates are established using a positive unedited control (nucleofected without RNP) and a negative control (non-fluorescent K562 cells), and the cell population is analyzed for fluorescence loss within the GFP channel to evaluate in vivonuclease activity.

[0277] Example 17 - Use for epigenome editing Epigenome editing is a gene regulation technique that involves constitutively or transiently turning genes on or off. Such techniques can utilize catalytically dead Cas9 (dCas9) fused to three proteins: Dnmt3A, Dnmt3L, and KRAB. Dnmt3A and Dnmt3L are DNA methyltransferases. The KRAB domain mediates histone methylation. DNA and histone methylation in the promoter region mediates constitutive gene repression. dCas9 and guide RNA can recruit the DNA and histone methylation complex to the promoter region without requiring nuclease activity. Together, Dnmt3A, Dnmt3L, and KRAB have 579 aa, while dCas9 has 1,368 aa. The fusion protein consists of 1,947 aa or 5,841 nucleotides, exceeding the adeno-associated virus vector (AAV) packaging limit (4.7 Kb). Therefore, there is a need to create more compact epigenome editors. Compact V-type nucleases from the MG119 family represent excellent candidates for use as dead nuclease partners in epigenome editing technologies. When fused to DNA and histone methylation complexes, their small size, ranging from 350 to 700 aa, allows for fusion protein sizes ranging, for example, from approximately 929 to 1,279 aa, or from approximately 2787 to 3837 nucleotides, enabling easy packaging in AAVs.

[0278] To test the MG119 fusion protein as an epigenome editor, HEK293T cells expressing GFP under a chimeric promoter (GAPDH-Srnpn) are generated by lentiviral transduction. An MG119 family guide RNA targeting the chimeric promoter is designed. The guide was fabricated by modifying the 5' and 3' nucleotides with three 2'-O-methyl substituents and three phosphorothioate bonds for stability. A dead version of the MG119 nuclease is fused to a DNA and histone methylation complex (MG119 epigenome editor). The fusion protein is cloned in a mammalian expression plasmid under a CMV promoter. GFP-expressing HEK293T cells are transfected with plasmids expressing the MG119 epigenome editor and the chemically synthesized guide. The transfected cells are analyzed by flow cytometry. Successful MG119 epigenome editor is determined by the loss of GFP fluorescence in the transfected cells. Next, the MG119 epigenome editor is used to target the gene of interest for treatment.

[0279] Example 18 - In silico identification of a medium-sized V-type nuclease Homology searches for V-type Cas nucleases were performed using HMMER software. Hits of V-type nuclease sequences were retained if the following criteria were met: (i) the hmmsearch e value was ≤ 10. -5The criteria were (ii) the gene encoding the nuclease was within 1Kb of the CRISPR array, and (iii) the amino acid sequence length was in the range of 700–1100aa. Using MMSeqs2, sequences were clustered with 100% amino acid identity at coverage mode 1 and 80% coverage of the target sequence (parameters --cov-mode 1-c 0.8 --min-seq-id 1.0). Sequence representatives were selected to construct a phylogenetic tree by building a multiple sequence alignment using MAFFT with the Needleman-Wunsch algorithm for global alignment. Careful examination of the phylogenetic tree led to the identification of the V-type nuclease family MG191 (sequence codes 1065–1090 and 1114–1118) and their associated potential CRISPR RNAs (sequence codes 1091–1113 and 1119–1120). The genomic context of representative nuclease genes is shown in Figure 16A, along with their associated crRNAs (Figure 16B).

[0280] Example 19 - Optimization of the construct for improving the solubility of MG119 nuclease Some effector proteins are readily purified with stable fusion proteins (e.g., maltose-binding protein, MBP), but precipitate upon cleavage of the fusion protein using proteases. MBP is a large fusion protein (approximately 35 kDa), and if not removed, it can impair nucleofection efficiency. Therefore, the insoluble MG119 nuclease was expressed via N-terminal fusion with SUMO (small ubiquitin-like modifier, approximately 11 kDa, Tables 10 and 22), which is commonly used as an expression / soluble fusion protein. Due to its relatively small size, the SUMO domain remains fused to the effector protein, preventing effector precipitation. Proteins with N-terminal SUMO fusion were expressed and purified in the same way as proteins expressed with pMGB△ expression vectors, as described in Example 14. In some cases, the inclusion of the N-terminal SUMO domain increased protein expression and solubility (Figure 17B), while the same protein without SUMO fusion was observed to be less pure and readily precipitated over time (Figure 17A). Furthermore, SUMO-fused effector proteins are also more active, as measured by activity fraction assays (e.g., MG119-1, Table 13).

[0281] [Table 10]

[0282] Example 20 - Extended characterization of MG119 nuclease Determination of non-template chain break sites and confirmation of PAM sequences To identify PAM sequences and non-target strand cleavage sites, 5 nM nuclease-amplified DNA templates and 25 nM sgRNA-amplified DNA templates (including one of the spacer sequences listed in Table 11) were expressed at 37°C for 3 hours using an in vitro protein synthesis kit. Plasmid library DNA cleavage was performed by mixing a 5 nM target library representing all possible 8N PAMs, a 5-fold dilution of the in vitro expression, 10 nM Tris-HCl, 10 nM MgCl2, and 100 mM NaCl at 37°C for 2 hours. The reaction was stopped, cleaned with PCR cleanup beads, and eluted in Tris-EDTA buffer at pH 8.0. The 3 nM cleavage product ends were blunted in 0.167 U / μL Mung Bean Nuclease and 1× Mung Bean Nuclease buffer at 30°C for 15–30 minutes. The ligated products were amplified by PCR using NGS primers and sequenced by NGS. The active proteins that successfully cleaved the PAM library produced a band of approximately 195 bp in agarose gel electrophoresis. The bands were extracted from the agarose gel and sequenced by NGS to determine the PAM sequences.

[0283] [Table 11]

[0284] Sequence logos generated using Seqlogo maker demonstrated that the PAM sequence recognized by the active effector on the non-target strand (NTS) matched the PAM sequence on the target strand (TS). For example, for MG119-137, the PAM determined from the NTS matched the PAM obtained from the TS (Figure 18). The cleavage sites on the NTS were determined from the number of reads at each nucleotide position. Table 12 shows preferred cleavage sites on any of the protospacer sequences targeted by the spacers (Table 11), including the corresponding sgRNAs used in the experiment.

[0285] [Table 12]

[0286] Determining the spacer length The preference for spacer lengths for MG119-1, MG119-2, and MG119-3 nucleases was determined by testing the activity of purified enzymes using single guide RNAs (SEQ ID NOs. 432, 755, and 761) each carrying 16-24 nt spacers targeting plasmids with preferred PAM(TTG) and protospacer sequences. 250 nM effectors were complexed with 500 nM of each guide for 20 minutes at room temperature. The RNPs were then reacted with 5 nM of the target plasmid in 1 × buffer at 37°C for 1.5 hours. The reaction was terminated by denaturing at 75°C for 10 minutes to analyze the cleavage products. To remove residual guide RNA, 0.1 μg / μL of RNAse A was added to each reactant and incubated at 37°C for 10 minutes. To remove proteins, 0.03 U / μL of proteinase K was added to each reactant and incubated at 55°C for 15 minutes. The reactants were mixed with 1× Gel Loading Dye, Purple no SDS. 6 μL was loaded onto a 1% agarose gel with GelGreen dye and electrophoresed at 135 V for 35 minutes. The target plasmids were visualized using an imaging system within the GelGreen channel. The cleavage products included linearized plasmids from dsDNA cleavage and nicked plasmids from ssDNA cleavage. In MG119-1, double-stranded and single-stranded cleavage was observed at all spacer lengths. The 18nt spacer guide produced less nicked product compared to other spacer lengths. MG119-2 also produced double-stranded and single-stranded cleavage, but minimal nicking was observed with the 18nt spacer guide. MG119-3 preferred a spacer length of 16 nt and tended to produce single-strand breaks more often than double-strand breaks (Figure 19).

[0287] Nuclease activity evaluation using purified protein Not all protein purifications produce properly folded and catalytically active proteins. To determine the quality of purified proteins, activity fraction assays were performed to quantify the percentage of cleavable purified proteins. These assays were performed by titrating the RNP complex against a constant concentration of linear DNA substrate and measuring DNA cleavage. Effector proteins were purified and pre-incubated with 1.5- or 2-fold molar excess sgRNA at room temperature for 20 minutes to form ribonucleoprotein complexes (RNPs). The reaction was set up using titration of 25 nM DNA substrate and 0.25 × ~10 × molar excess RNP relative to the substrate. The reaction buffer composition was 10 mM Tris pH 7.5, 10 mM MgCl2, and 100 mM NaCl. The DNA substrate was 522 bp long. Successful cleavage resulted in fragments of 172 and 350 bp. The reaction mixture was incubated at 37°C for 60 minutes, then at 75°C for 10 minutes. RNase was added to each reactant (Cf = 0.33 μg / μL), and the samples were incubated at 37°C for 10 minutes. Proteinase K was added to each reactant (Cf = 60 units / mL), and the samples were incubated at 55°C for 15 minutes. Then, the entire reactant was electrophoresed on a 1.5% agarose gel containing GelGreen dye, and imaged using an imaging system in the GelGreen channel. The percentage of cleaved substrate was calculated for each lane through densitometry analysis using image software. The active fraction was determined by the slope of the linear range of cleavage (Table 13).

[0288] [Table 13]

[0289] Example 21 - Single Guide RNA Manipulation Primary sequence optimization To prepare for testing gene editing in mammalian cells, four consecutive U strings were replaced with single or paired mutations. These changes prevent premature transcriptional termination and immunogenicity in mammalian cells while preserving the secondary structure of the sgRNA. These modified guides were designed and tested in parallel with wild-type (WT) guides using the same in vitro cleavage reaction described above. The amplified cleavage products were visualized from agarose gels and compared with each other. Modified guides that produced equivalent PCR bands were selected for downstream testing.

[0290] Optimization of guide RNA length for improved guide RNA synthesis Standard commercially available guide synthesis quality degrades as RNA sequences lengthen. Therefore, truncations of either wild-type or optimized sgRNAs with reasonable design were explored, with the ultimate goal of identifying guide sequences ≤100 nt that do not contain spacer sequences. Truncations were designed using predicted RNA folds, and predicted RNAs were produced. sgRNAs were resuspended in 100 μM TE buffer and then aliquoted to 25 μM diH2O. Guide engineering activity testing was performed using RNP:substrate ratios such that WT (non-truncated) sgRNAs cleaved approximately 50% of the available substrate, based on activity fraction measurements for each different effector. All sgRNAs were tested for a given effector at this same RNP:substrate ratio. To form RNP complexes, sgRNAs were incubated with effector proteins at a 1.5-fold molar excess relative to the effector protein for 20 minutes at room temperature. From there, activity testing proceeded in the same manner as activity fraction measurement, starting with a 60-minute incubation at 37°C. All guide modifications were synthesized using U40 spacers (Table 11). The process of designing, ordering, and testing sgRNA truncations was repeated in further engineering rounds, and new truncations were expanded or designed based on data from previous engineering rounds. MG119-1, MG119-3, MG119-32, MG119-54, MG119-129, and MG119-136 underwent one round of sgRNA engineering, while MG119-2 and MG119-28 both underwent two rounds of sgRNA engineering.

[0291] Throughout each round of guide engineering, guides that exhibited ≥80% activity against the WT were considered to be of sufficient quality to progress. In the first engineering round for MG119-2 guide design, both sgRNA2_4 (124nt, SEQ ID NO: 659) and sgRNA2_6 (129nt, SEQ ID NO: 434) met this threshold, but other truncated guide scaffolds did not (Figure 20A). During the second round of guide manipulation, sgRNA2_4 and sgRNA2_6 truncations were combined to produce the sgRNA2_4.6 guide (115nt, SEQ ID NO: 663), which was further combined with additional truncations (Figure 20B). All variations of the MG119-2 sgRNA2 guide tested in the second round of guide engineering met the ≥80% activity threshold. The shortest guide that met the threshold was MG119-2 sgRNA2_4.6.11 (96nt, SEQ ID NO: 668). Identifying guide structures that are short enough for high-yield, commercially viable synthesis, and yet enable efficient enzyme activity will help broaden the range of potential scientific and therapeutic applications for effector proteins.

[0292] Example 22 - Mammalian cell editing using MG119 nuclease Nuclease mRNA production The nuclease mRNA sequence was codon-optimized for human expression, then synthesized and cloned into a high-copy-number ampicillin plasmid. The synthesized construct encoding the T7 promoter, UTR, nuclease ORF, and NLS sequence was digested from the backbone with HindII and BamHI, and ligated to the pUC19 plasmid backbone using T4 DNA ligase and 1× reaction buffer. The completed nuclease mRNA plasmid consisted of a replication origin, an ampicillin-resistant cassette, the synthesized construct, and an encoded poly(A) tail. Nuclease mRNA was synthesized via in vitro transcription (IVT) using a linearized nuclease mRNA plasmid. This plasmid was linearized by incubation with the SapI enzyme at 37°C for 16 hours. The linearization reaction consisted of 50 μL of reaction mixture containing 10 μg of pDNA, 50 units of Sap I, and 1× reaction buffer. The linearized plasmid was purified with phenol:chloroform:isoamyl alcohol (25:24:1, v / v), precipitated in EtOH, and resuspended in nuclease-free water at a adjusted concentration of 500 ng / μL. An IVT reaction to produce nuclease mRNA was performed at 50°C for 1 hour under the following conditions: 1 μg linearized plasmid, 5 mM ATP, CTP, GTP, and N1-methylpseudo-UTP, 18750 U / mL Hi-T7 RNA polymerase, 4 mM CleanCap AG, 2.5 U / mL inorganic E. coli pyrophosphatase, 1000 U / mL mouse RNase inhibitor, and 1× transcription buffer. After 1 hour, the IVT was stopped, the plasmid DNA was digested with the addition of 250 U / mL DNase I, and incubated at 37°C for 10 minutes. The nuclease mRNA was purified. The transcript concentration was determined by UV light and further analyzed by capillary gel electrophoresis.

[0293] Cell culture, transfection, next-generation sequencing, and indel analysis in mammalian cells Experiments using MG119-28 were performed on Hepa1-6 cells grown and passaged at 37°C and 5% CO2 in Dulbecco's Modified Eagle Medium supplemented with 10% (v / v) fetal bovine serum and 1% Pen / Strep, with 1×NEAA added. Experiments using MG119-2, MG119-125, and MG119-129 were performed on K562 cells grown and passaged at 37°C and 5% CO2 in Dulbecco's Modified Eagle Medium supplemented with 10% (v / v) fetal bovine serum.

[0294] For experiments where the nuclease was delivered as mRNA, 500 ng of mRNA and 150 pmol of sgRNA were mixed together and incubated on ice until cells were prepared. For experiments where the nuclease and guide were delivered as an RNP complex, 100 pmol of purified nuclease was incubated with 200 pmol of sgRNA at room temperature for approximately 30 minutes prior to transfection. The sgRNA sequences used with MG119-28, which targets mouse albumin intron 1, are SEQ ID NOs: 625-628. The sgRNA sequences used with MG119-2, which targets human TRAC exon 3, are SEQ ID NOs: 767-798. The sgRNA sequences used with MG119-129, which targets human APOA1 exon 3, are SEQ ID NOs: 979-1021. The sgRNA sequences used with MG119-125, which targets human APOA1 exon 3, are SEQ ID NOs: 831-904.

[0295] Approximately 1×10 5Cells were transfected with RNP complexes or mRNA+sgRNA in a nucleofector. Transfected cells were grown for 3 days, harvested, and gDNA was extracted. The target regions of the indels were amplified using Q5 high-fidelity DNA polymerase with primers, DNA was extracted as a template, and the PCR product was purified. PCR primers suitable for use in NGS-based DNA sequencing were generated and optimized and used to amplify individual target sequences of each guide RNA. The amplicons were sequenced and analyzed using a proprietary Python script, and indel frequencies were measured.

[0296] Table 14 shows the percentage of amplicons from NGS amplicon sequencing containing insertions or deletions obtained using the MG119-28 guide targeting intron 1 of the mouse albumin gene. Compared to previous results using MG119-28 with RNP delivery, mRNA delivery enhanced editing in mouse albumin by up to >90% at one target site (mALB sgRNA3) and 43% at another target site (mALB sgRNA4).

[0297] [Table 14]

[0298] Figure 21A shows the percentage of amplicons obtained from NGS amplicon sequencing containing insertions or deletions obtained using an MG119-2 guide targeting exon 3 of the human TRAC gene. Five target sites in TRAC exhibited >10% editing by MG119-2 upon delivery by both mRNA and RNP, with the highest editing rate being 57%.

[0299] MG119-125 was delivered as mRNA to K562 cells targeting human APOA1 exon 3. The percentage of indel amplicons from NGS amplicon sequencing is shown in Figure 21B. Four target sites showed detectable editing, with one target showing up to 17% editing.

[0300] MG119-129 was delivered to K562 cells targeting human APOA1 exon 3 as both mRNA and RNP. The percentage of indel amplicons from NGS amplicon sequencing is shown in Figure 21C. At one target site, 33% editing was presented using RNP delivery and 39% editing was presented using mRNA delivery.

[0301] Example 23 - Mammalian cell editing using MG119 nuclease Nuclease mRNA production The MG119 nuclease mRNA sequence was codon-optimized for human expression, then synthesized and cloned into a high-copy-number ampicillin plasmid. The synthesized construct encoding the T7 promoter, UTR, nuclease ORF, and NLS sequence was digested from the backbone with HindII and BamHI enzymes and ligated to the pUC19 plasmid backbone using T4 DNA ligase and 1× reaction buffer. The completed nuclease mRNA plasmid consists of a replication origin, an ampicillin-resistant cassette, the synthesized construct, and an encoded poly(A) tail. Nuclease mRNA was synthesized via in vitro transcription (IVT) using a linearized nuclease mRNA plasmid. This plasmid was linearized by incubation with the SapI enzyme at 37°C for 16 hours. The linearization reaction consisted of 50 μL of reaction mixture containing 10 μg of pDNA, 50 units of Sap I, and 1× reaction buffer. The linearized plasmid was purified with phenol:chloroform:isoamyl alcohol (25:24:1, v / v), precipitated in EtOH, and resuspended in nuclease-free water at a adjusted concentration of 500 ng / μL. An IVT reaction to generate nuclease mRNA was performed at 50°C for 1 hour under the following conditions: 1 μg linearized plasmid, 5 mM ATP, CTP, GTP, and N1-methylpseudo-UTP, 18750 U / mL Hi-T7 RNA polymerase, 4 mM CleanCap AG, 2.5 U / mL inorganic E. coli pyrophosphatase, 1000 U / mL mouse RNase inhibitor, and 1× transcription buffer. After 1 hour, the IVT was stopped, the plasmid DNA was digested with the addition of 250 U / mL DNase I, and incubated at 37°C for 10 minutes. The nuclease mRNA was purified. The transcript concentration was determined by UV light and further analyzed by capillary gel electrophoresis using a Fragment Analyzer.

[0302] Cell culture, transfection, next-generation sequencing, and indel analysis in mammalian cells The experiment was conducted using K562 cells that were grown and passaged in Iskov-modified Dulbecco medium supplemented with 10% (v / v) fetal bovine serum at 37°C and 5% CO2.

[0303] For mRNA-based nucleofection (MG119-2 and MG119-28), 500 ng of mRNA and 200 pmol of sgRNA were mixed together and incubated on ice until cells were prepared. For RNP-based nucleofection (MG119-32), 100 pmol of nuclease was complexed with 200 pmol of sgRNA at room temperature for 30 minutes, and then the mixture was placed on ice until cells were prepared. The sgRNA sequences used with MG119-2, MG119-28, and MG119-32, which target human albumin intron 1, human APOA1 exon 3, and AAVS1, are SEQ ID NOs: 1121-1169, 1219-1237, 1257-1324, 1393-1441, 1491-1506, 1523-1562, 1603-1632, 1663-1669, and 1677-1686.

[0304] Approximately 1×10 5 Cells were transfected with mRNA+sgRNA. Transfected cells were grown for 3 days, harvested, and gDNA was extracted. The target regions of the indels were amplified using high-fidelity DNA polymerase with primers, and DNA was extracted as a template. The PCR products were purified. PCR primers suitable for use in NGS-based DNA sequencing were generated and optimized and used to amplify individual target sequences of each guide RNA. The amplicons were sequenced by next-generation sequencing and analyzed using a custom Python script to measure indel frequencies.

[0305] Figure 22 shows the percentage of amplicons obtained from NGS amplicon sequencing, including insertions or deletions, using the MG119-2 guide targeting intron 1 of the human ALB gene, exon 3 of the human APOA1 gene, and AAVS1. Sixteen target sites across the three loci showed >10% editing, with the highest editing being 43% in AAVS1-gD2. Figure 23 shows the percentage of amplicons obtained from NGS amplicon sequencing, including insertions or deletions, using the MG119-28 guide targeting intron 1 of the human ALB gene, exon 3 of the human APOA1 gene, and AAVS1. Twenty-five target sites across the three loci showed >10% editing, with 12 of them showing >50% editing. In AAVS1, two target sites achieved >90% editing. Figure 24 shows the percentage of amplicons obtained from NGS amplicon sequencing, including insertions or deletions, using the MG119-32 guide targeting intron 1 of the human albumin (ALB) gene, exon 3 of the human APOA1 gene, and AAVS1. The two target sites achieved >20% editing.

[0306] Example 24 - Guide RNA Engineering of the MG119 Family Split Guide Engineering Most MG119 effectors have a single guide scaffold in the range of 130–160 nt (without spacers). Minimizing the guide via truncation of unwanted regions is advantageous, but this is not always possible. A viable alternative is to split the guide RNA into two halves and anneal them before RNP complex formation. To test this idea, four or five split guide designs were evaluated for MG119-2_sg2_WT (Figure 25A), MG119-2_sg2_4.6.11 (Figure 26A), MG119-28_sg1_WT (Figure 27A), and MG119-28_sg1_8.5 (Figure 28A), respectively. Prior to RNP complexation, two halves of each splitting guide were heated to 80°C for 2 minutes in unfolding buffer (20 mM HEPES pH 7.0, 100 mM NaCl, 0.1 mM EDTA). As a control, each parental sgRNA was also unfolded and re-annealed together with the splitting guide. Then, an equal volume of pre-heated (80°C) 2× annealing buffer (40 mM HEPES pH 7.0, 200 mM NaCl, 2 mM MgCl2) was added to the guide, and the reaction mixture was slowly cooled to 4°C at 0.1°C / second. From this point, the in vitro cleavage reaction was carried out as described in Example 20. These assays included both unfolded parental sgRNA (N, native) and refolded parental sgRNA (R, refolded). For MG119-2, the reaction was set up with 25 nM of DNA substrate, 50 nM of MG119-2 effector protein, and 75 nM of guide RNA. For MG119-28, the reaction was set up with 20 nM of DNA substrate, 400 nM of MG119-28 effector protein, and 600 nM of guide RNA.

[0307] For MG119-2_sg2 (sequence number 656), all splitting guide designs (sequence numbers 1697-1704) were active to varying degrees (Figure 25B). For MG119-2_sg2_4.6.11 (sequence number 668), splitting guides v1 and v5 were substantially inactive, and the v2 splitting guide was only about half as active as the parental sgRNA (Figure 26B; sequences 1705-1712). For both MG119-28_sg1 (sequence number 686, Figure 27B) and MG119-28_sg1_8.5 (sequence number 704, Figure 28B), all versions of the splitting guide were highly active except for v2 (sequence numbers 1713-1720 and 1721-1728, respectively). These data suggest that splitting guide engineering is feasible for systems with large guides and that multiple splitting points may be active.

[0308] Minimizing a single guide RNA for improved synthesis The length of the sgRNA for the MG119-32 effector was optimized in the same manner as described in Example 21. For MG119-32, two WT variants of sgRNA, sg1 (SEQ ID NO: 448) and sg2 (SEQ ID NO: 720), were initially tested. Guide trimming was performed for both sg1 and sg2 WT guides. For the first round of guide engineering, similar truncations were designed for both sg1 (Figure 29A, SEQ ID NOs: 712-719) and sg2 (Figure 29C; SEQ ID NOs: 721-728). The cleavage test was performed in the same manner as described in Example 20, using 20 nM substrate, 200 nM MG119-32 effector protein, and 300 nM sgRNA.

[0309] Similar trends were observed in both sgRNA panels, with only truncations 5, 6, 7, and 8 showing any significant activity compared to WT sgRNA (Figures 29B and 29D). The percentage of cleaved substrates was calculated for each lane via densitometry analysis. Since cleavage appeared generally high in the sg1 sgRNA panel, only sg1 was advanced to round 2 of guide engineering. In the second round of guide engineering, truncation 5 was combined with truncation 6 (MG119-32_sg1_5.6, SEQ ID NO: 1729), truncation 7 (MG119-32_sg1_5.7, SEQ ID NO: 1730), or truncation 8 (MG119-32_sg1_5.8, SEQ ID NO: 1731) to create three different dual truncations (Figure 30A). The cleavage test was performed using 20 nM substrate, 500 nM MG119-32 effector protein, and 750 nM sgRNA, similar to the first round of guide engineering.

[0310] Of the new dual truncation guides, 119~32_sg1_5.6 and 119~32_sg1_5.7 were cleaved with high activity comparable to WT sg1 sgRNA. Guides contributing at least 80% of cleavage compared to WT are considered successful and used as the starting point for the second round of guide engineering. In contrast, 119~32_sg1_5.8 showed impaired activity compared to WT sg1 sgRNA. These data suggest that large sgRNAs of the MG119 family effector proteins are suitable for inducing trimming without significant loss of activity.

[0311] Characterization of trimmed sgRNA During guide length optimization, guide scaffold truncation progresses through engineering rounds primarily based on dsDNA cleavage activity. However, guide length optimization may unintentionally alter tangential enzymatic activity, such as nicking, incidental ssDNA cleavage, and incidental ssRNA cleavage.

[0312] To determine whether guide engineering of the MG119-2 and MG119-28 guides altered secondary nicking activity, a cleavage assay was performed using circular plasmid DNA as a substrate. In this assay, the uncleaved substrate remains supercoiled. The completely cleaved substrate becomes linear, and its migration on the agarose gel is therefore slower compared to the supercoiled uncleaved reactant. Furthermore, nicked plasmids are not linearized, but their supercoils relax, and this species migrates even slower on the agarose gel than the completely cleaved linear product, making it distinguishable. Therefore, tracking the intensity of this relaxed product leads to reading the prevalence of nicking activity. Moreover, performing this reaction with the WT guide, the completely trimmed guide, and each modification that contributed to the completely trimmed guide provides insight into which parts of the guide may mediate such activity (if differences are observed).

[0313] The MG119-2 cleavage reaction was performed using 5 nM plasmid substrate, 10 nM MG119-2 effector protein, and 15 nM sgRNA, as described in Example 20. The data suggest that the nicking activity remained largely unchanged for all guides tested (Figure 31A). The MG119-28 cleavage reaction was performed using 5 nM plasmid substrate, 35 nM MG119-28 effector protein, and 52.5 nM sgRNA, as described in Example 20. The data suggest that the nicking activity remained largely unchanged for all guides tested (Figure 31B).

[0314] Example 25 - Discovery of a medium-sized V-type nuclease and in vitro activity assay Discovery of a novel medium-sized V-type nuclease in silico. Nucleases of the MG191 family (750 - 800 aa) were identified with the crRNAs encoded adjacent to them. By a similar method, the discovery of four other families of type V nucleases of length 700 - 1100 aa (MG185, MG186, MG187, and MG188; SEQ ID NOs: 1746 - 1752) was brought about. Briefly, Cas12 hmm hits that were within 1 Kb from the CRISPR array and of length 700 - 1100 aa with an e-value ≤ 10 -5 were clustered with 100% amino acid identity (--cov-mode 1 -c 0.8 -min-seq-id 1.0). Representative sequences were aligned using MAFFT that employs the Needleman - Wunsch algorithm for global alignment, and a phylogenetic tree was constructed using FastTree (Figure 32).

[0315] MG191 DNA template for in vitro transcription and translation The minimal array (SEQ ID NOs: 1732 - 1738) was designed to contain a T7 promoter, a predicted repeat, a U40 spacer sequence (TGGAGATATCTTGAACCTTGCATCCCCGGA, SEQ ID NO: 19,01), a second identical repeat, followed by a primer binding sequence in this order. In a second design, while keeping the spacer constant, the repeat orientation of the minimal array was reversed to test which orientation of the repeat was active. The reverse complements of these sequences were ordered as single-stranded ultramers that anneal with a complementary T7 promoter oligo for transcription. The transcription template was prepared by mixing and incubating 20 μM of the T7 promoter oligo, 20 μM of the reverse complement minimal array, and 0.7× IDT duplex buffer, then heating to 95°C for 2 minutes and then annealing by cooling to room temperature at 0.1°C / second.

[0316] E. coli codon-optimized nuclease plasmids were generated using a pTAC-driven expression vector (pMGD expression vector described in Example 14) that had NLS sequences flanking each end, an N-terminal His tag, and MBPs with a precision protease region. Linear nuclease templates for in vitro transcription / translation (IVTT) were amplified by PCR, and a T7 promoter was added for simultaneous expression. The mixture was then purified with 10 mM Tris HCl pH 8.0 and eluted. The PCR templates were validated for yield and purity.

[0317] Minimal array MG191 transcription (precrRNA) The aforementioned dsDNA minimal array template was used for the synthesis of precrRNA.

[0318] RNA was synthesized and purified. The transcript was verified for yield and purity. The RNA produced was used to test the activity of the purified protein.

[0319] MG191 in vitro cleavage assay and PAM sequencing As previously mentioned, a 5 nM nuclease-amplified DNA template and a 25 nM promoter-annealed minimal array DNA template (including the U40 spacer sequence TGGAGATATCTTGAACCTTGCATCCCCGGA (SEQ ID NO: 1901)) were expressed in an in vitro protein expression system at 37°C for 2 hours. Plasmid library DNA cleavage was performed by mixing a 5 nM target library representing all possible 8N PAMs with a 5-fold dilution of the in vitro expression in 10 mM Tris-HCl pH 7.9, 10 mM MgCl2, 100 μg / mL BSA, and 50 mM NaCl at 37°C for 1 hour. The reaction was stopped, cleaned using PCR cleanup beads, and eluted in Tris-EDTA buffer at pH 8.0. The ends of 3 nM cleavage products were blunted at 25°C for 15 minutes using 3.33 μM dNTPs, 1 × T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment. 1.5 nM cleavage products were ligated at room temperature for 20 minutes using a 150 nM adapter, 1 × T4 DNA ligase buffer, and 20 U / μL T4 DNA ligase. The ligated products were amplified by PCR using NGS primers, sequenced by NGS, and PAMs were obtained.

[0320] Successfully cleaved active proteins from the PAM library produced a band of approximately 205 bp in agarose gel electrophoresis (Figure 33). Identified PAMs (SEQ ID NOs. 1739-1745) are shown as sequence logos (Figure 34). Table 15 lists preferred cleavage sites on the target strand of protospacer sequences complementary to the U40 spacer.

[0321] [Table 15]

[0322] MG191 expression and purification Isolating pure and functional proteins is essential for extensive in vitro analysis of biochemical properties and mechanistic studies. The MG191 candidate was expressed and purified to obtain sufficient quantities and quality of protein for such characterization. All constructs were expressed in E. coli. The constructs were expressed together with the N-terminal MBP fusion protein in a pMGD expression vector.

[0323] Protein expression Protein expression plasmids were used to transform competent cells, which were then cultured overnight at 37°C in 25 mL of growth medium containing 100 μg / L carbenicillin (1.6% tryptone, 1% yeast extract, 0.5% NaCl). The following day, 10 mL from each overnight culture was inoculated into 1000 mL of growth medium containing 100 μg / L carbenicillin, and the cultures were grown at 37°C with shaking. When the OD600 was approximately equal to 0.6, the cultures were cooled on ice, induced with 0.5 mM IPTG, and further incubated overnight at 16°C for approximately 18 hours with shaking. Next, the culture was collected by centrifugation at 6,000 × g for 10 minutes, and the pellet was resuspended in nickel-A buffer (50 mM HEPES, 500 mM NaCl, 10 mM MgCl2, 0.5 mM EDTA, 20 mM imidazole, 5% glycerol, pH 7.5) + EDTA-free protease inhibitor and stored at -80°C. Culture samples were collected before and after induction, and the cells were pelletized by centrifugation (15,000 × g, 1.5 min) and resuspended in 100 μL of 2 × Laemmli buffer per OD cell.

[0324] Protein purification All MG191 candidates were purified using the same method. MG191-15 is shown here as an example. The protein expressed in the pMGD vector has the following sequence structure: 6xHis-(GS)1-MBP-GSGSGGGSGS-PSP-nucleoplasmin binocular NLS-GGSGSGGS-MG119-X-GGSGGSG-SV40 NLS (Table 16).

[0325] [Table 16]

[0326] The cell pellet was thawed and refilled to 80 mL with nickel-A buffer containing 0.5% β-octyl glucoside. The sample was sonicated in an ice bath at 75% amplitude for a total processing time of 2 minutes using a 5-second on / 15-second off cycle. The lysate was clarified by centrifugation at 30,000 × g for 15 minutes, and the supernatant batch was conjugated to 2.5 mL of Ni-NTA resin for ≥15 minutes. The sample was loaded onto a gravity column, washed with 30 CV of nickel washing buffer, then eluted in 4 CV of nickel elution buffer (nickel washing buffer + 250 mM imidazole), and concentrated in a 50 kDa MWCO concentrator. Samples were collected throughout the purification process, run on SDS-PAGE protein gels, activated with UV for 5 minutes, and imaged with a UV transilluminator. These gels were used to track the progress of purification throughout the protocol (Figure 35A). Finally, the protein samples were filtered through a 0.22 μm cellulose acetate membrane, then placed on an S200i 10 / 300 GL column and further purified by flowing through SEC buffer (20 mM HEPES, 500 mM NaCl, 10 mM MgCl2, 0.5 mM EDTA, 5% glycerol, 0.5 mM TCEP, pH 7.5) (Figure 35B). The peak fractions were pooled and concentrated in a 50 kDa MWCO concentrator. The absorbance of each sample at 280 nm was measured, the protein concentration was calculated, and the yield was estimated using the sample volume and concentration (nmol / L, Table 17). To estimate purity, 2 μg of purified protein was run on an SDS-PAGE protein gel, activated with UV for 5 minutes, and imaged with a UV transilluminator (Figure 35C). The purity of each sample was estimated using densitometry (Table 17). Purified proteins were obtained for each candidate MG191 through this protocol.

[0327] [Table 17]

[0328] Nuclease activity evaluation using purified protein Activity assays were performed to determine whether purified MG191 protein could be cleaved. This assay was performed to test the activity of MG191-15, MG191-17, and MG191-23. The results of the MG191-15 assay are shown as an example. These assays are performed by titrating the RNP complex against a constant concentration of linear DNA substrate and measuring DNA cleavage. Effector proteins were pre-incubated with 1.5-fold molar excess pre-crRNA (repeat-spacer-repeat-primer binding site) at room temperature for 20 minutes to form ribonucleoprotein complexes (RNPs). The reaction was set up using 20 nM DNA substrate and dose settings of 5 × (100 nM) and 25 × (500 mM) molar excess RNP relative to the substrate. The reaction buffer composition was 10 mM Tris pH 7.5, 10 mM MgCl2, and 100 mM NaCl. The DNA substrate was 521 bp long and contained PAM determined via NGS and a spacer specified by a pre-crRNA array. Successful cleavage yielded fragments of approximately 171 and 350 bp. The reaction products were incubated at 37°C for 60 minutes, then at 75°C for 10 minutes. RNase A was added to each reaction product (Cf = 0.33 μg / μL), and the samples were incubated at 37°C for 10 minutes. Proteinase K was added to each reaction product (Cf = 60 units / mL), and the samples were incubated at 55°C for 15 minutes. The entire reaction product was then electrophoresed on a 1.5% agarose gel with nucleic acid dye and imaged with a transilluminator. The active protein produced a product band that moved faster than the uncleaved reaction product on the agarose gel (Figure 35D). The percentage of cleaved substrate was calculated for each lane via densitometry analysis. These data indicate that all candidates except MG191-17 are active as purified proteins.

[0329] Activity testing of purified MG191 protein using in vitro library cleavage and ligation assays. The same in vitro assay described above was used to test the activity of the purified preparations of MG191-18, MG191-25, and MG191-29. In this assay, plasmid library DNA cleavage was performed by mixing a 5 nM target library representing all possible 8N PAMs, 50 nM effector protein, 50 nM minimal array RNA, 10 mM Tris-HCl pH 7.9, 10 mM MgCl2, 100 μg / ml BSA, and 50 mM NaCl at 37°C for 30 minutes. The reaction was stopped, cleaned using PCR cleaning beads, and eluted in water. The 3 nM cleavage product ends were blunted at 25°C for 15 minutes with 3.33 μM dNTPs, 1 × T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment. The 1.5 nM cleavage product was ligated at room temperature for 20 minutes using a 150 nM adapter, 1 × T4 DNA ligase buffer, and 20 U / μL T4 DNA ligase. The ligated product was amplified by PCR. The active protein successfully cleaved from the PAM library produced a band of approximately 205 bp on agarose gel electrophoresis (Figure 35E).

[0330] Example 26-MG119-28 sgRNA successfully truncated without reducing its activity in cells. Three of the truncated guides that showed sustained high levels of cleavage in vitro were further tested in K562 cells targeting nine different sites and exhibited diverse activities according to the protocol described in Example 22 (Table 18). Using the guide scaffolds MG119-28_sg1_8 (SEQ ID NO: 703), MG119-28_sg1_5 (SEQ ID NO: 691), and MG119-28_sg1_8.5 (SEQ ID NO: 704), chemically modified guides targeting the spacers listed in Table 18 were designed to generate 27 sgRNAXs.

[0331] [Table 18]

[0332] NGS-edited data showed that all designs were active in cells, but guide scaffold 1_8 had activity equal to or higher than the WT guide. Guide designs 1_5 and 1_8.5 had, on average, lower activity than the WT guide (Figure 36).

[0333] Further analysis of guide 1_8 was performed by dose titration of mRNA and guide 1_8 or WT to K562 cells (Figure 37). These results showed that sg1_8 generated equivalent or better levels of indels even at doses below the saturation dose of the guide, indicating that this truncation to sgRNA is not detrimental to activity. This truncation to WT sgRNA reduced the guide length by 15 nt.

[0334] Example 27-MG119-28: Sequential truncation of sgRNA and nuclease activity testing in cells The activity of the guide with escalation truncation was evaluated across the entire repeat-anti-repeat region of the sgRNA. The guide design included 20nt spacers targeting five AAVS1 target sequences. The sequence of all additional escalation truncations was SEQ ID NO: 1876, which is a modified sgRNA sequence from SEQ ID NO: 703. K562 cells were grown and passaged at 37°C and 5% CO2 in Iskov-modified Dulbecco medium supplemented with 10% (v / v) fetal bovine serum. 120K cells were transfected with 200 pmol of each guide and 500 ng of nuclease mRNA in a 96-well plate format, as recommended in the Amaxa® 4D-Nucleofector® protocol for the 4D-Nucleofector® system (Lonza). Transfected cells were grown for 3 days, harvested, and gDNA was extracted using QuickExtract (Lucigen) according to the manufacturer's instructions. PCR primers suitable for use in NGS-based DNA sequencing were generated and optimized, and each guide RNA and extracted DNA were used as templates to amplify individual target sequences. The PCR products were purified. The amplicons were sequenced and analyzed.

[0335] Data analysis shows that truncation to the target stem-loop reduces activity compared to the wild-type guide. Each guide is plotted against its corresponding indel and compared to wild-type sgRNA activity (Figure 38). Detectable activity was observed for each guide at all five AAVS1 sites. Although activity is reduced compared to the WT guide, these alternative guides can be used when guide size or sequence is constrained.

[0336] Example 28 - Reconstruction of the ancestral sequence of MG191 nuclease, and in vitro activity assay of the modern and ancestral sequences. Reconstruction of the ancestral sequence of MG191 nuclease To generate further diversity within the MG191 family of medium-sized V-type nucleases, we used the Ancestry Sequence Reconstruction (ASR) algorithm. ASR is a computational technique that reconstructs potential ancestral sequences of ancient organisms using existing protein sequences and the relationships inferred between them. Briefly, sequences were aligned using MAFFT with the Needleman-Wunsch algorithm for global alignment (Katoh and Standley 2013). Two phylogenetic trees were constructed using FastTree (Price, Dehal and Arkin 2010) with different sequence sets, and both trees were rooted using known sequences as outgroups. Sequence reconstruction was performed using the codeml package in PAML 4.8 (Yang 2007). Insertions and deletions were manually identified for each reconstructed node. Thirty-one ancestral sequences were reconstructed with high confidence and have the amino acid sequences shown in SEQ ID NOs. 1842–1872.

[0337] In vitro cleavage assays for modern and ancestral candidates In vitro cutting for PAM and cutting site determination. Nucleases of the MG191 family can process their own crRNA. To investigate nuclease activity, modern MG191 nucleases MG191-1, -2, -5, and -28 (SEQ ID NOs. 1065, 1066, 1069, and 1115) were tested using corresponding repeat sequences (SEQ ID NOs. 1873 or SEQ ID NOs. 1095, 1098, and 1119) and the U40 spacers listed in Table 19. The ancestral nucleases MG191-37, -38, -40, -41, -42, -48, -51, -53, and -62 (SEQ ID NOs. 1847, 1848, 1850-1852, 1858, 1861, 1863, and 1872) were tested with the smallest array of their closest modern homologues (SEQ ID NOs. 1732-1734, 1738, and 1874).

[0338] Table 20 shows the combinations of active nucleases and minimal arrays listed in this section.

[0339] A 5 nM nuclease-amplified DNA template and a 25 nM promoter-annealed minimal array DNA template were expressed at 37°C for 2 hours using the PURExpress® In Vitro Protein Synthesis Kit (New England Biolabs Inc.). Plasmid library DNA cleavage was performed by mixing a 5 nM target library representing all possible 8N PAMs with a 5-fold dilution of the PURExpress expression in 10 mM Tris-HCl pH 7.9, 10 mM MgCl2, 100 μg / mL BSA, and 50 mM NaCl (NEB 2.1 Buffer, NEB Inc.) at 37°C for 2 hours. The reaction was stopped, the mixture was cleaned, and eluted in Tris-EDTA buffer at pH 8.0. The 3 nM cleavage product ends were blunted at 25°C for 15 minutes using 3.33 μM dNTPs, 1X T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment. The 1.5 nM cleavage product was ligated at room temperature for 20 minutes using a 150 nM adapter, 1X T4 DNA ligase buffer, and 20 U / μL T4 DNA ligase. The ligated product was amplified by PCR using NGS primers, sequenced by NGS, and PAM was obtained.

[0340] [Table 19]

[0341] The active proteins successfully cleaved from the PAM library produced bands of approximately 200–300 bp on a TapeStation D1000 gel (Agilent Technologies Inc.) (Figure 39). PAMs recognized by the active modern and ancestral MG191 nuclease are provided in sequence numbers 1837–1841. An example of a sequence logo created with Logo Maker is shown in Figure 40.

[0342] Modern nucleases MG191-1, MG191-2, and MG191-5 are active in the minimal arrays and repeat orientations tested, as listed in Table 20 (SEQ ID NOs: 1867, 1095, and 1098). The ancestral nucleases MG191-51 and MG191-53 were activated in multiple minimal arrays (SEQ ID NOs: 1874, 1732, 1734, and 1738 in the current sequence listing). Preferred cleavage sites on the target strand of protospacer sequences complementary to the U40 spacer are listed in Table 20.

[0343] [Table 20]

[0344] MG191 crRNA exchange capability To optimize the execution of downstream studies in human cells, the inventors tested repeat compatibility within the MG191 nuclease family using the same in vitro cleavage assay described above. MG191-12, -15, -17, -18, -25, and -29 nucleases (SEQ ID NOs. 1076, 1079, 1081, 1082, 1089, and 1116) were tested in combinations of their corresponding repeats in minimal array form using similar PAM sequences (SEQ ID NOs. 1739, 1741, 1744, and 1745). With the exception of MG191-29, the MG191 nucleases listed in Table 21 can use other crRNAs for in vitro cleavage (Figure 41).

[0345] Purification of SUMO-fused MG191-12 and MG191-25, and evaluation of their nuclease activity.

[0346] [Table 21]

[0347] MG191 candidates previously purified with MBP require removal of the MBP tag before any intracellular testing due to the large size of the MBP tag. To mitigate potential protein precipitation and degradation after MBP cleavage, MG191 candidates were expressed and purified using a SUMO-soluble tag that does not require cleavage due to its minimal size. The protein expressed with the SUMO tag has the following structure: 6xHis-(GS)1-SUMO-(GGSGS)2-PSP-nucleoplasmin binocular NLS-GGS-MG119-X-GP-SV40 NLS (Table 22). Densitometry was used to estimate the purity of the final sample of MG191-25 expressed with the SUMO construct compared to the same protein expressed with the MBP construct. The mean purity of all recovered fractions of MG191-25 expressed with the SUMO tag in the final purification step was 57%, compared to 46% observed for MG191-25 expressed with MBP after tag cleavage.

[0348] [Table 22]

[0349] Competent cells were transformed with SUMO-fused MG191-12 (SEQ ID NO: 1076) and MG191-25 protein expression plasmids and cultured overnight at 37°C in 25 mL of 2×YT medium (1.6% tryptone, 1% yeast extract, 0.5% NaCl) containing 100 μg / L carbenicillin. The following day, 10 mL from each overnight culture was inoculated into 1000 mL of 2×YT medium containing 100 μg / L carbenicill...

Claims

1. A manipulated nuclease system, a) An endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 57, 31, 1-30, 32-56, 58-325, 420-431, 476-624, 629, 1065-1090, 1114-1118, 1746-1752, and 1842-1872, b) An engineered nuclease system comprising an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence.

2. The manipulated nuclease system according to claim 1, wherein the endonuclease contains a sequence having at least 90% or 100% sequence identity with any one of sequence numbers 57, 31, 1-30, 32-56, 58-325, 420-431, 476-624, 629, 1065-1090, 1114-1118, 1746-1752, and 1842-1872.

3. The manipulated nuclease system according to claim 1, wherein the manipulated guide polynucleotide comprises crRNA and tracrRNA.

4. The manipulated guide polynucleotide is sequence numbers 333-335, 355-357, 410-411, 346-347, 368-369, 412-413, 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456 The manipulated nuclease system according to claim 1, comprising a sequence having at least 90% or 100% sequence identity with any one of 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, 1091-1113, 1119-1120, 1697-1731, 1807-1836, 1876, and 1887-1894.

5. An engineered nuclease system, (i) a) an endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 6 to 14, and b) An engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide includes a sequence having at least 80% sequence identity with any one of sequence numbers 333-335 and 355-357. (ii) a) an endonuclease comprising a sequence having at least 80% sequence identity with sequence number 15, and b) An engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide includes a sequence having at least 80% sequence identity with any one of sequence numbers 410 to 411. (iii) a) an endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 16 to 29, and b) An engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 346-347, 368-369, and 412-413. (iv) a) an endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 57, 31, 1-30, 32-56, 58-150, 420-431, 476-624, and 629, and b) An engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 326-332, 336-345, 348-354, 358-367, 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, 474, 647-766, and 1697-1731, or (v) a) Endonucleases comprising a sequence having at least 80% sequence identity with any one of sequence numbers 1065-1090, 1114-1118, and 1746-1752, and b) An engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1091-1113, 1119-1120, and 1876. A modified nuclease system, including [specific component].

6. a) The manipulated guide polynucleotide is a single guide nucleic acid, a dual guide nucleic acid, or RNA, and / or b) The endonuclease is non-covalently bonded to the manipulated guide polynucleotide, or the endonuclease is covalently bonded to the manipulated guide polynucleotide, or the endonuclease is fused to the manipulated guide polynucleotide. The manipulated nuclease system according to claim 1.

7. The modified nuclease system according to claim 1, further comprising DNA methyltransferase.

8. a) The DNA methyltransferase is non-covalently bonded to the endonuclease, or the DNA methyltransferase is fused to the endonuclease in a single polypeptide, and / or b) The DNA methyltransferase comprises Dmnt3A or Dnmt3L, The manipulated nuclease system according to claim 7.

9. A method for modifying a target nucleic acid sequence, comprising contacting the target nucleic acid sequence with an engineered nuclease system according to any one of claims 1 to 8.

10. The method according to claim 9, wherein the target nucleic acid sequence includes genomic DNA, viral DNA, viral RNA, or bacterial DNA.

11. The method according to claim 10, wherein the modification is in vitro, in vivo, or ex vivo.

12. Use of the engineered nuclease system according to any one of claims 1 to 8 for modifying a target nucleic acid sequence in a mammalian cell, wherein the use includes contact with the mammalian cell.

13. A method for modifying TRAC, comprising contacting TRAC with an engineered nuclease system, wherein the engineered nuclease system is a) An endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 57, 31, 1-30, 32-56, 58-150, 420-431, 476-624, and 629, b) A method comprising: an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity with any one of sequence numbers 767 to 798.

14. A method for modifying APOA1, comprising contacting APOA1 with an operated nuclease system, wherein the operated nuclease system is a) An endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 57, 31, 1-30, 32-56, 58-150, 420-431, 476-624, and 629, b) A method comprising: an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide includes a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 831-904, 979-1021, 1219-1237, 1491-1506, and 1663-1669.

15. A method for modifying AAVS1, comprising contacting APOA1 with an operated nuclease system, wherein the operated nuclease system is a) An endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 57, 31, 1-30, 32-56, 58-150, 420-431, 476-624, and 629, b) A method comprising: an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1257-1324, 1523-1562, 1677-1686, 1753-1779, and 1807-1836.

16. A method for modifying an albumin gene, comprising contacting the albumin gene with an engineered nuclease system, wherein the engineered nuclease system is a) An endonuclease comprising a sequence having at least 80% sequence identity with any one of sequence numbers 57, 31, 1-30, 32-56, 58-150, 420-431, 476-624, and 629, b) A method comprising: an engineered guide polynucleotide configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence, wherein the engineered guide polynucleotide contains a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 1121-1169, 1393-1441, 1603-1632, 1887, 1889, 1891, and 1892-1894.

17. A nucleic acid encoding an engineered nuclease system according to any one of claims 1 to 8.

18. A cell comprising the manipulated nuclease system according to any one of claims 1 to 8.

19. The aforementioned cells, a) eukaryotic cells, b) mammalian cells; c) immortalized cells; d) insect cells; e) yeast cells, f) plant cells; g) fungal cells; h) prokaryotic cells; i) A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or their derivatives, j) Manipulated cells, k) Stable cells, l) primary cells, m) T cells, or n) Hematopoietic stem cells (HSC) The cell according to claim 18.