Class 2, type v crispr systems

WO2026122844A9PCT designated stage Publication Date: 2026-08-13METAGENOMI THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-08-13

AI Technical Summary

Technical Problem

Existing CRISPR systems face challenges in achieving high specificity and efficiency in targeted nucleic acid editing, particularly in identifying off-target editing and ensuring minimal mismatch tolerance, which can lead to unintended genomic modifications.

Method used

Development of engineered nuclease systems comprising engineered endonucleases and guide polynucleotides with specific sequence identities, used in conjunction with lipid nanoparticles, to form complexes that target nucleic acid sequences within genes like VCP, B2M, TIM-3, and TRAC, with methods to assess and minimize off-target editing through barcode-based sequencing and linker optimization.

Benefits of technology

Enhances the specificity and efficiency of nucleic acid editing by reducing off-target effects, allowing precise modification of target sequences while maintaining low off-target activity, thereby improving the safety and accuracy of genetic interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025058191_13082026_PF_FP_ABST
    Figure US2025058191_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Described herein are engineered endonuclease systems for modifying target genes (e.g., valosin-containing protein (VCP) gene, Beta-2-Microglobulin (B2M) gene, T cell immunoglobulin and mucin domain 3 (TIM-3) gene, T Cell Receptor Alpha Constant (TRAC) gene).
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 00010.041.1801CLASS 2, TYPE V CRISPR SYSTEMSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 729,212, filed on December 6, 2024, the entire content of which is hereby incorporated by reference herein in its entirety for all purposes.SEQUENCE LISTING

[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety7. Said XML copy, created on December 1, 2025. is named 00010_041_1801_SL.xml and is 14,209,840 bytes in size.SUMMARY

[0003] Described herein, in certain embodiments, are engineered nuclease systems, comprising: a) an engineered endonuclease having at least 70% sequence identity to any one of SEQ ID NOs: 21 -225 and 3471-3684; and b)an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence within a valosin- containing protein (VCP) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3721-3723, and 3725-3727. In some embodiments, the engineered endonuclease has at least 80% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered nuclease is encoded by a sequence having at least 70% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered nuclease is encoded by a sequence having at least 90% sequence identity' to SEQ ID NO: 3729. In some embodiments, the engineered nuclease is encoded by a sequence having 100% sequence identity' to SEQ ID NO: 3729. In some embodiments, the engineered guide polynucleotide is a single guide nucleic acid. In some embodiments, the engineered guide polynucleotide is a dual guide nucleic acid. In some embodiments, the engineered guide polynucleotide is RNA. In some embodiments, the engineered endonuclease binds non-covalently to the engineered guide polynucleotide. In some embodiments, the endonuclease is covalently linked to the engineered guide polynucleotide. In some embodiments, the endonuclease is fused to the engineered guide polynucleotide.1#507551Attorney Docket No. 00010.041.1801

[0004] Described herein, in certain embodiments, are engineered nuclease systems, comprising: a) an engineered endonuclease comprising a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence within a Beta-2-Microglobulin (B2M) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3731- 3737 and 3746- 3752. In some embodiments, the engineered endonuclease has at least 80% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered nuclease is encoded by a sequence having at least 70% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered nuclease is encoded by a sequence having at least 90% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered nuclease is encoded by a sequence having 100% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered guide polynucleotide is a single guide nucleic acid. In some embodiments, the engineered guide polynucleotide is a dual guide nucleic acid. In some embodiments, the engineered guide polynucleotide is RNA. In some embodiments, the engineered endonuclease binds non-covalently to the engineered guide polynucleotide. In some embodiments, the endonuclease is covalently linked to the engineered guide polynucleotide. In some embodiments, the endonuclease is fused to the engineered guide polynucleotide.

[0005] Described herein, in certain embodiments, are engineered nuclease systems, comprising: a) an engineered endonuclease comprising a sequence having at least 70% sequence identity' to any one of SEQ ID NOs: 215-225 and 3471-3684; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a T cell immunoglobulin and mucin domain 3 (TIM-3) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered endonuclease has at least 80% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered nuclease is encoded by a sequence having at least 70% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered nuclease is encoded by a sequence having at least 90% sequence identity to SEQ ID NO: 3729. In some embodiments, the2#507551Attorney Docket No. 00010.041.1801 engineered nuclease is encoded by a sequence having 100% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered guide polynucleotide is a single guide nucleic acid. In some embodiments, the engineered guide polynucleotide is a dual guide nucleic acid. In some embodiments, the engineered guide polynucleotide is RNA. In some embodiments, the engineered endonuclease binds non-covalently to the engineered guide polynucleotide. In some embodiments, the endonuclease is covalently linked to the engineered guide polynucleotide. In some embodiments, the endonuclease is fused to the engineered guide polynucleotide.

[0006] Described herein, in certain embodiments, are engineered nuclease systems, comprising: a) an engineered endonuclease comprising a sequence having at least 70% sequence identity' to any one of SEQ ID NOs: 215-225 and 3471-3684; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a T Cell Receptor Alpha Constant (TRAC) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered endonuclease has at least 80% sequence identity’ to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the engineered nuclease is encoded by a sequence having at least 70% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered nuclease is encoded by a sequence having at least 90% sequence identity- to SEQ ID NO: 3729. In some embodiments, the engineered nuclease is encoded by a sequence having 100% sequence identity to SEQ ID NO: 3729. In some embodiments, the engineered guide polynucleotide is a single guide nucleic acid. In some embodiments, the engineered guide polynucleotide is a dual guide nucleic acid. In some embodiments, the engineered guide polynucleotide is RNA. In some embodiments, the engineered endonuclease binds non- covalently to the engineered guide polynucleotide. In some embodiments, the endonuclease is covalently linked to the engineered guide polynucleotide. In some embodiments, the endonuclease is fused to the engineered guide polynucleotide.

[0007] Described herein, in certain embodiments, are methods of modifying a target nucleic acid sequence in a mammalian cell comprising contacting the mammalian cell using the engineered nuclease system described herein. In some embodiments, modifying the target nucleic acid sequence comprises binding, nicking, or cleaving the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence comprises genomic DNA, viral DNA, viral RNA,3#507551Attorney Docket No. 00010.041.1801 or bacterial DNA. In some embodiments, the modification is in vitro. In some embodiments, the modification is in vivo. In some embodiments, the modification is ex vivo. In some embodiments, the methods further comprise selecting cells comprising the modification.

[0008] Described herein, in certain embodiments, are cells comprising the engineered nuclease system described herein. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokary otic cell. In some embodiments, the cell is an A549, HEK-293, HEK-293T, BHK, CHO. HeLa, MRC5, Sf9. Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C 12, L cell, HT1080, HepG2, Huh7. K562, primary cell, or a derivative thereof. In some embodiments, the cell is an engineered cell. In some embodiments, the cell is a stable cell. In some embodiments, the cell is a T cell. In some embodiments, the cell is a hematopoietic cell.

[0009] Described herein, in certain embodiments, are lipid nanoparticles comprising: (a) the engineered nuclease system described herein; (b) a cationic lipid; (c) a sterol; (d) a neutral lipid; and (e) a PEG-modified lipid. In some embodiments, the cationic lipid comprises Cl 2-200, the sterol comprises cholesterol, the neutral lipid comprises DOPE, or the PEG-modified lipid comprises DMG-PEG2000. In some embodiments, the cationic lipid comprises 98N12-5 (TETA5- LAP), DLin DMA, DLin-K-DMA (2,2-Dilinoleyl-4-dimethylaminomethyl-[l,3]-dioxolane), DLin-KC2-DMA, DLin-MC3-DMA, or C 12-200.

[0010] Described herein, are methods for identifying off-target editing efficiency relative to a target DNA sequence, comprising: a) contacting a cell with a library of polynucleotides and an engineered nuclease system, the engineered nuclease system comprising: an engineered endonuclease having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471- 3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease, wherein each polynucleotide of the library comprises: (i) an on-target nucleic acid sequence, (ii) an off-target nucleic acid sequence containing a single nucleotide mismatch at a defined position within a protospacer adjacent motif (PAM) or a spacer relative to the on-target sequence, (iii) a first barcode associated with the off-target nucleic acid sequence and a second barcode associated with the on-target nucleic acid sequence, (iv) a linker sequence separating the off-target nucleic acid sequence and the on-target nucleic acid sequence, and (v) adapter primers flanking each of the on-target and off-target nucleic acid sequences; b) isolating genomic DNA from the cell; c)4#507551Attorney Docket No. 00010.041.1801 amplifying the target DNA sequence using the adapter primers; and d) sequencing the first and second barcodes and aligning the sequence reads assigned to each polynucleotide of the library to a target DNA sequence and quantifying an off-target indel frequency and an on-target indel frequency, and computing a mismatch tolerance value, thereby identifying off-target editing efficiency relative to the target DNA sequence for the engineered nuclease system.

[0011] Also described herein, are methods for determining specificity of a nuclease, the method comprising: a) contacting a cell with a li brary of polynucleotides and an engineered nuclease system, the engineered nuclease system comprising: an engineered endonuclease having at least 70% sequence identify to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease, wherein each polynucleotide of the library comprises: (i) an on-target nucleic acid sequence, (ii) an off-target nucleic acid sequence containing a single nucleotide mismatch at a defined position within a protospacer adjacent motif (PAM) or a spacer relative to the on-target sequence, (iii) a first barcode associated with the off-target nucleic acid sequence and a second barcode associated with the on- target nucleic acid sequence, (iv) a linker sequence separating the off-target nucleic acid sequence and the on-target nucleic acid sequence, and (v) adapter primers flanking each of the on-target and off-target nucleic acid sequences; b) isolating genomic DNA from the cell; c) amplifying the target DNA sequence using the adapter primers; and d) sequencing the first and second barcodes and aligning the sequence reads assigned to each polynucleotide of the library to a target DNA sequence and quantifying an off-target indel frequency and an on-target indel frequency, and computing a mismatch tolerance value, thereby determining specificity of a nuclease.

[0012] In certain embodiments, the linker is 10 to 40 nucleotides in length. For example, about 10 nucleotides to about 20 nucleotides, about 11 nucleotides to about 20 nucleotides, about 12 nucleotides to about 20 nucleotides, about 13 nucleotides to about 18 nucleotides, about 28 to about 38 nucleotides, etc. In certain embodiments, the linker is unique across at least a subset of the library members to reduce barcode recombination during amplification. In certain embodiments, identifying comprises recovering a distal segment of a barcode when a proximal segment is removed by a deletion event. In certain embodiments, computing the mismatch tolerance value comprises calculating a bound off-target to bound on-target ratio and designating a library member to be off-target editing when the ratio exceeds a predetermined threshold value of 0.5. In some embodiments, wherein a mismatch tolerance value of less than 0.5 indicates higher specificity7for the engineered nuclease sy stem and a mismatch tolerance value of more than 0.5 indicates lower5#507551Attorney Docket No. 00010.041.1801 specificity for the engineered nuclease system. In certain embodiments, contacting comprises delivering the complex to a cell by nucleofection or lipid nanoparticle formulation.

[0013] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings of which:

[0015] FIG. 1 depicts sequences of the human and mouse Valosin-Containing Protein (VCP) locus showing the wild-type (WT) and R155H mutation-bearing sequences for human IBMPFD patients (top) and the murine disease model cells (bottom), with corresponding amino acids. The nucleotide sequence of WT human VCP gene and WT mouse VCP gene are denoted by SEQ ID NO: 3717 and 3719, respectively. The nucleotide sequence of mutant human VCP gene and mutant mouse VCP gene are denoted by SEQ ID NO: 3718 and 3720, respectively. The amino acid sequence of WT human and mouse VCP protein is denoted by SEQ ID NO: 4083 and the amino acid sequence of mutant human and mouse VCP protein is denoted by SEQ ID NO: 4084.

[0016] FIG. 2 depicts MG29-1 guides targeting the mutant VCP allele in mouse VCPR155II / Ifibroblasts. Fibroblasts derived from VCPR135H / +mice were electroporated with MG29-1 mRNA (designated mr!26) and guides specific to the mutant allele (gr2133, gr2134, and gr2135). As a positive control, Cas9 mRNA and guide gr2155 were electroporated. Three (3) days postelectroporation, genomic DNA was extracted and the targeted genomic region was amplified with primers designed for NGS-based sequencing. Amplicons were sequenced and analyzed to quantify gene editing efficiency. The left graph in FIG. 2 shows percent total indels, while the right graph displays the percentage of out-of-frame (OOF) indels.

[0017] FIG. 3 depicts editing outcomes by MG29-1 guide gr2133 at the VCP locus in mouse VCPR155H +fibroblasts. FIG. 3 further shows alignment of sequencing reads produced by MG29-16#507551Attorney Docket No. 00010.041.1801 and guide gr2133 in fibroblasts derived from VCPR155H / +miCe. The mutated allele is shown as reference.

[0018] FIG. 4 depicts editing outcomes by MG29-1 guide gr2134 at the VCP locus in mouse VCPR155H / + fibroblasts, and alignment of sequencing reads produced by MG29-1 and guide gr2134 in fibroblasts derived from ycPRI 55nmice. The mutated allele is shown as reference.

[0019] FIG. 5 depicts editing outcomes by MG29-1 guide gr2135 at the VCP locus in mouse VCPR135H / +fibroblasts, and alignment of sequencing reads produced by MG29-1 and guide gr2135 in fibroblasts derived from VCPR155H / +mice.

[0020] FIG. 6 depicts editing outcomes by Cas9 positive control guide gr2155 at the VCP locus in mouse VCPR155H / +fibroblasts, and shows alignment of sequencing reads produced by Cas9 and guide gr2155 in fibroblasts derived from VCPR 15511mice.

[0021] FIG. 7 depicts MG29-1 guide gr2135 targeting of the mutant VCP allele in patient T cells. T cells were purified from IBMPFD patient blood were activated for three (3) days, rested for one day, and then electroporated with MG29-1 mRNA (designated mrl26) and mutant allele-specific guide gr2135. After three (3) additional days of culture, genomic DNA was extracted and the targeted genomic region was amplified with primers designed for NGS-based sequencing. Amplicons were sequenced on an Illumina MiSeq platform and analyzed using a proprietary Python script to quantify gene editing efficiency. The unedited, wild-type sequence is represented in 54.43% of sequencing reads, while the intact unedited R155H mutant sequence is represented in 16.46% of sequencing reads.

[0022] FIG. 8 depicts a schematic overview of the workflow showing single mismatches across every position of PAM and Spacer were generated (1MM Target) and used to create the dual-target library members. The library w as ordered and cloned into a lenti viral plasmid to create the plasmid library from which the lentiviral library’ was produced. K562 cells were infected with the dualtarget lentilibrary and following antibiotic selection the cells were edited via nucleofection and the gDNA harvested. The data was analyzed following NGS sequencing.

[0023] FIG. 9 depicts a range of GC content and Indels of MG29-1 library sgRNAs. The range of GC contents and percent indels for the guides from which the library was designed were plotted for MG29-1.

[0024] FIG. 10 depicts a reference structure of dual-target oligo library member adapter sequences used for amplification of library’ and sequencing. BC1 and BC2 oligonucleotides are dual barcoded with unique sequence identifiers. A PAM (protospacer adjacent motif) is found downstream of a7#507551Attorney Docket No. 00010.041.1801 target sequence. "On-target" refers to a PAM endowed spacer sequence corresponding to a single guide RNA. " I MM Target” differs only by a single mismatch (1MM) from the on-target, single mismatches are present in the spacer and the PAM sequence.

[0025] FIG. 11 depicts editing controls. Control oligonucleotides were generated to simulate four possible editing events: cleavage at the on-target, cleavage at the off-target, cleavage at both targets, and no cleavage. These control dual-target sequences possessed identical on- and off-target protospacers, with cleavage events regulated by their active or inactive PAM sequences. To generate inactive PAMs, bases were substituted in the active PAM consensus sequence with alternatives that are not recognized by the Cas system. Editing at the left or right end of the dualtarget controls, following nucleofection with guides for MG29-1, was plotted.

[0026] FIG. 12 depicts editing ranges observed at off and on-target sequences for MG29-1. The ranges of editing observed at the off-target (single mismatch containing target) or the on-target sequence were plotted for MG29-1.

[0027] FIG. 13 depicts generating a nuclease-specific single mismatch tolerance profile. Raw indel percentages at mismatched targets and their corresponding on-target were calculated for each library member. The raw indel percentages were normalized across each position by taking the ratio of observed off: on target editing, termed mismatch tolerance, for each guide. A nucleasespecific single mismatch tolerance profile was generated by plotting the mean and 95% confidence intervals of these single mismatch tolerances, by position, of all guides for a given nuclease.

[0028] FIG. 14 depicts single mismatch tolerance of 22 nucleotide spacers for MG29-1. 15 guides targeting human B2M, TIM3 and TRAC loci were used to design dual-target library members in which the off-target sequence differs by 1 mismatch between the off and on-target sequence. The raw indel percentages at mismatched targets compared to on-target library members were normalized by taking the ratio of offon target editing, termed mismatch tolerance, observed by position for each guide and a nuclease specific profile was generated by plotting the mean and 95% confidence intervals of single mismatch tolerance by position.

[0029] FIG. 15 depicts MG29-1 PAM preference. PAM preference scores were calculated for MG29-1 by converting the mismatch tolerance scores and their respective off-target PAMs into information-based representation matrices using logomaker.

[0030] FIG. 16 depicts MG29-1 protospacer base preference heatmap. Heatmaps were created to illustrate the protospacer base preference at each mismatch position for different nucleases. The heatmaps were generated by grouping the data points by nuclease, mismatch location, and8#507551Attorney Docket No. 00010.041.1801 mismatched protospacer base, and then calculating the mean off-on ratio and 95% confidence interval for each group.

[0031] FIG. 17 depicts a circos plot for MG29-1 complexed with ALB-83b showing a representative illustration of translocations between the chromosome targeted by guide (chr4) and all other chromosomes after applying a RPM filter of 83.53 (top) and an upset plot for showing results from the translocation assay (bottom).

[0032] FIG. 18 depicts a circos plot for MG29-1 complexed with a control guide showing a representative illustration of translocations between the chromosome targeted by guide (chr 18) and all other chromosomes after applying an RPM filter of 203.43 (top) and an upset plot showing results from the translocation assay (bottom).

[0033] FIG. 19 is a schematic overview of the workflow. Dual-target library members were designed by combining the on target with single mismatches across every position of PAM and spacer (1MM Target) as well as genetic variants generated by CRISPRme paired with the gnomAD variant database. The library was cloned into a lentiviral plasmid to create the plasmid library from which the lentiviral library was produced. K562 cells were infected with the dual-target lenti library and following antibiotic selection the cells were edited via LNP delivery and the gDNA was harvested. NGS data was analyzed to evaluate mismatch tolerance at each spacer position.

[0034] FIG. 20 is a reference structure of dual-target oligo library member for Type V nuclease. Adapter sequences were used for amplification of library’ and sequencing. Oligos were dual barcoded with unique sequence identifiers (BC1 and BC2). PAM- protospacer adjacent motif found downstream of target sequence. On-target - PAM endowed spacer sequence, corresponding to the single guide RNA. Off-target - includes: single mismatch (1MM) walk along the spacer and PAM sequences, off-targets nominated by CRISPRme and off-target nominated by biochemical and cellular assays. Oligos included: SEQ ID Nos: 4085-8126.

[0035] FIG. 21 is a graph showing on-target average inDei % at albumin genomic locus. LNP transfection was performed in 83b mismatch library’ K562 cells with 3 replicates of each following conditions: mock and 3* EC90. Saturating on-target editing at ALB was observed. Average InDel% for mock and 3* EC90 are 0.04% and 96.58% respectively.

[0036] FIG. 22 is a scatterplot graph showing off-target editing percentages across oligonucleotides in the 83b library' from two edited samples. Comparison of on-target versus off- target editing in the 83b mismatch library' indicates that all oligos show a high percentage of on- target editing, which serves as an internal control for normalizing off-target editing rates. Most9#507551Attorney Docket No. 00010.041.1801 oligonucleotides with elevated off-target activity are singletons (single-nucleotide mismatches) containing a single mismatch within the spacer region. CRISPRme and CRISPRme with variation library members represent in silico predicted off-target sites with up to four edit distance. Additional detected hits included the on-target oligonucleotide (Oligo_0 (denoted by SEQ ID NO: 8126; off / on ratio ~1) and a CRISPRme with variation oligo, Oligo_3884 (denoted by SEQ ID NO: 7968) (off / on ratios -0.44 and -0.32).

[0037] FIG. 23 is a graph showing mismatch tolerance profile through off / on ratio for single generated mismatch library.

[0038] FIG. 24 is a graph showing off / on ratio separated by library7categories of 83b mismatch library. Most active hits from screen are the singletons, on-target site, and a single library member in the CRISPRme with variation category (Oligo_3884 (denoted by SEQ ID NO: 7968) off / on ratios -0.44 and -0.32).

[0039] FIG. 25 is a graph providing comparison of Oligo_3884 (denoted by SEQ ID NO: 7968) and its hg38 reference counterpart, Oligo_97 (denoted by SEQ ID NO: 4181). Oligo_3884 (denoted by SEQ ID NO: 7968) contains a CRISPRme-predicted sequence variation from GnomAd and exhibits detectable off-target editing (off / on ratios -0.44 and -0.32), whereas the hg38 reference oligonucleotide (Oligo_97, denoted by SEQ ID NO: 4181) shows no measurable off-target activity7(off / on ratios -0.000305 and -0.000296).

[0040] FIGs. 26A-26B are graphs showing barcode detection rate. FIG. 26A shows the detected barcodes while using both the full length 18bp barcodes for identification and FIG. 26B show's how using full-length barcode 1 and a shortened portion (13bp) of barcode 2 leads to higher recovery7of the reads w ith both barcodes detected.

[0041] FIG. 27 show7exemplary library designs w ith linkers of vary ing lengths and identity7(from top to bottom, designs 1 to 4). Design 1 denotes pool with same 28 nt linker as previous 83B. Design 2 denotes pool w ith 33 nt constant linker. Design 3 denotes pool with unique 33nt linker for each oligo. Design 4 denotes pool with unique 38 nt linker for each oligo. Oligos included: SEQ ID Nos: 8127- 9946.

[0042] FIG. 28 depicts barcode detection of 83b mismatch optimization library. The new oligonucleotide design incorporating an extended linker and longer barcode was tested to better accommodate the broad indel window of the Type V nuclease MG29-1, the plasmid barcode detection results showed reduced barcode recombination by using unique linkers through decreased constant-region overlap.10#507551Attorney Docket No. 00010.041.1801BRIEF DESCRIPTION OF THE SEQUENCE LISTING

[0043] The Sequence Listing filed herewith provides exemplary7polynucleotide and polypeptide sequences for use in methods, compositions, and systems according to the disclosure. Below are exemplary descriptions of sequences therein. Abbreviations: mN: 2'-O-methyl modified base N; fN: 2'-Fluoro modified base N; *: phosphorothioate linkage; N: standard ribonucleotide base; AltRl and AltR2 refer to 5’ and 3’ AltR modifications.MG11

[0044] SEQ ID NOs: 1-37 show the full-length peptide sequences of MG11 nucleases.MG13

[0045] SEQ ID NOs: 38-118 show the full-length peptide sequences of MG13 nucleases.MG19

[0046] SEQ ID NOs: 119-124 show the full-length peptide sequences of MG19 nucleases.MG20

[0047] SEQ ID NO: 125 shows the full-length peptide sequence of a MG20 nuclease.MG26

[0048] SEQ ID NOs: 126-140 show the full-length peptide sequences of MG26 nucleases.MG28

[0049] SEQ ID NOs: 141-214 show the full-length peptide sequences of MG28 nucleases.MG29

[0050] SEQ ID NOs: 215-225 and 3471-3684 show the full-length peptide sequences of MG29 nucleases.

[0051] SEQ ID NO: 3697 shows the nucleotide sequence of an MG29-1 nuclease containing 5’ UTR, NLS, CDS, NLS, 3’ UTR, and polyA tail.

[0052] SEQ ID NOs: 3693-3694 show effector repeat motifs of MG29 nucleases.

[0053] SEQ ID NO: 3695 shows the nucleotide sequence of a sgRNA engineered to function with a MG29 nuclease.

[0054] Table 3 shows PAM sequences compatible with MG29 nucleases.

[0055] SEQ ID NO: 3696 shows an MG29-1 coding sequence used for the generation of rnRNA.

[0056] SEQ ID NOs: 3701-3709 show DNA sequences encoding MG29-1 mRNAs.

[0057] SEQ ID NOs: 3721-3723, and 3725-3727 show the nucleic acid sequences of MG29-1 sgRNA targeting mutant VCP locus.

[0058] SEQ ID NO: 3729 shows the nucleic acid sequence encoding MG29-1 mRNA.11#507551Attorney Docket No. 00010.041.1801

[0059] SEQ ID NOs: 3731-3737 and 3746- 3752 show the nucleotide sequences of MG29-1 sgRNA targeting human locus hB2M.

[0060] SEQ ID NOs: 3738- 3744 and 3753- 3759 show the nucleotide sequences of MG29-1 sgRNA targeting human locus TIM3.

[0061] SEQ ID NOs: 3745 and 3760 show the nucleotide sequence of MG29-1 sgRNA targeting human locus TRAC.MG30

[0062] SEQ ID NOs: 226-228 show the full-length peptide sequences of MG30 nucleases.MG31

[0063] SEQ ID NOs: 229-260 show the full-length peptide sequences of MG31 nucleases.MG32

[0064] SEQ ID NO: 261 shows the full-length peptide sequence of a MG32 nuclease.MG37

[0065] SEQ ID NOs: 262-426 show the full-length peptide sequences of MG37 nucleases.MG53

[0066] SEQ ID NOs: 427-428 show the full-length peptide sequences of MG53 nucleases.MG54

[0067] SEQ ID NOs: 429-430 show the full-length peptide sequences of MG54 nucleases.MG55

[0068] SEQ ID NOs: 431-688 show the full-length peptide sequences of MG55 nucleases.MG56

[0069] SEQ ID NOs: 689-690 show the full-length peptide sequences of MG56 nucleases.MG57

[0070] SEQ ID NOs: 691-721 show the full-length peptide sequences of MG57 nucleases.MG58

[0071] SEQ ID NOs: 722-779 show the full-length peptide sequences of MG58 nucleases.MG59

[0072] SEQ ID NOs: 780-792 show the full-length peptide sequences of MG59 nucleases.MG60

[0073] SEQ ID NOs: 793-1163 show the full-length peptide sequences of MG60 nucleases.MG61

[0074] SEQ ID NOs: 1164-1469 show the full-length peptide sequences of MG61 nucleases.12#507551Attorney Docket No. 00010.041.1801MG62

[0075] SEQ ID NOs: 1470-1472 show the full-length peptide sequences of MG62 nucleases.MG70

[0076] SEQ ID NOs: 1473-1514 show the full-length peptide sequences of MG70 nucleases.MG75

[0077] SEQ ID NOs: 1515-1710 show the full-length peptide sequences of MG75 nucleases.MG77

[0078] SEQ ID NOs: 1711-1712 show the full-length peptide sequences of MG77 nucleases.MG78

[0079] SEQ ID NOs: 1713-1717 show the full-length peptide sequences of MG78 nucleases.MG79

[0080] SEQ ID NOs: 1718-1722 show the full-length peptide sequences of MG79 nucleases.MG80

[0081] SEQ ID NO: 1723 shows the full-length peptide sequence of a MG80 nuclease.MG81

[0082] SEQ ID NOs: 1724-2654 show the full-length peptide sequences of MG81 nucleases.MG82

[0083] SEQ ID NOs: 2655-2657 show the full-length peptide sequences of MG82 nucleases.MG83

[0084] SEQ ID NOs: 2658-2659 show the full-length peptide sequences of MG83 nucleases.MG84

[0085] SEQ ID NOs: 2660-2677 show the full-length peptide sequences of MG84 nucleases.MG85

[0086] SEQ ID NOs: 2678-2680 show the full-length peptide sequences of MG85 nucleases.MG90

[0087] SEQ ID NOs: 2681-2809 show the full-length peptide sequences of MG90 nucleases.MG91

[0088] SEQ ID NOs: 2810-3470 and 3685-3692 show the full-length peptide sequences of MG91 nucleases.Spacer segments

[0089] SEQ ID NOs: 3761-3764 show the nucleotide sequences of spacer segments.NLS13#507551Attorney Docket No. 00010.041.1801

[0090] SEQ ID NOs: 3765-3780 show the sequences of example nuclear localization sequences (NLSs) that can be appended to nucleases according to the disclosure.B2M Targeting

[0091] SEQ ID NOs: 3731-3737, 3781-3856, and 3746-3752 show the nucleotide sequences of sgRNAs engineered to function with an MG29-1 nuclease in order to target B2M.

[0092] SEQ ID NOs: 3857-3932 show the DNA sequences of B2M target sites.TRAC Targeting

[0093] SEQ ID NOs: 3933-4004 show the nucleotide sequences of sgRNAs engineered to function with an MG29-1 nuclease in order to target TRAC.

[0094] SEQ ID NOs: 4005-4076 show the DNA sequences of TRAC target sites.Human VCP Targeting

[0095] SEQ ID NOs: 4077-4082 show the nucleotide sequences of sgRNAs engineered to function with an MG29-1 nuclease in order to target human VCP.

[0096] SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727 show the DNA sequences of human VCP target sites.VCP Locus

[0097] SEQ ID NOs: 3719 and 3720 show the DNA sequences of mouse VCP locus.

[0098] SEQ ID NOs: 3724 and 3728 show the nucleic acid sequence of Cas9 sgRNA targeting mutant VCP locus.Cas9 mRNA

[0099] SEQ ID NO: 3730 shows the nucleic acid sequence of Cas9 mRNA.DETAILED DESCRIPTION

[0100] While various embodiments of the disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed.

[0101] The practice of some methods disclosed herein employ, unless otherwise indicated, techniques of immunology7, biochemistry7, chemistry7, molecular biology, microbiology7, cell biology, genomics, and recombinant DNA. See for example Sambrook and Green, Molecular14#507551Attorney Docket No. 00010.041.1801Cloning: A Laboratory' Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F. M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0102] As used herein, the singular forms“an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms “including”, “includes”, “having”, “has”, “with”, or variants thereof are used in either the detailed description and / or the claims, such terms are intended to be inclusive in a manner similar to the term “comprising”.

[0103] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within one or more than one standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1 % of a given value.

[0104] The term “nucleotide,” as used herein, refers to a base-sugar-phosphate combination. Contemplated nucleotides include naturally occurring nucleotides and synthetic nucleotides. Nucleotides are monomeric units of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates such as dATP, dCTP, diTP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [aS]dATP, 7-deaza- dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance on the nucleic acid molecule containing them. The term nucleotide as used herein encompasses dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of ddNTPs include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. A nucleotide may be unlabeled or detectably labeled, such as using moieties comprising optically detectable moieties (e.g., fluorophores) or quantum dots. Detectable labels include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels of nucleotides include but are not limited fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine. 6-15#507551Attorney Docket No. 00010.041.1801 carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X- rhodamine (ROX), 4-(4'dimethylaminophenylazo) benzoic acid (DABCYL). Cascade Blue, Oregon Green, Texas Red, Cyanine and 5-(2'-aminoethyl)aminonaphthalene-l-sulfonic acid (EDANS). Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA] dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA] dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP. [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif; FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham, Arlington Heights, IL; Fluorescein- 15-dATP, Fluorescein- 12-dUTP. Tetramethyl-rodamine-6-dUTP, IR770- 9-dATP, Fluorescein- 12-ddUTP, Fluorescein- 12-UTP, and Fluorescein-15-2'-dATP available from Boehringer Mannheim, Indianapolis, Ind.; and Chromosome Labeled Nucleotides, BODIPY- FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY- TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein- 12-UTP, fluorescein- 12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5 -UTP, Rhodamine Green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP available from Molecular Probes, Eugene, Oreg. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically-modified nucleotide is biotin-dNTP. Non-limiting examples of biotinylated dNTPs include, biotin-dATP (e.g. bio-N6-ddATP, biotin- 14-d ATP), biotin-dCTP (e.g., biotin- 11-dCTP, biotin- 14-dCTP), and biotin-dUTP (e.g., biotin- 11-dUTP, biotin- 16-dUTP, biotin-20-dUTP).

[0105] The terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, either in single-, double-, or multistranded form. Contemplated polynucleotides include a gene or fragment thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA). ribosomal RNA (rRNA), short interfering RNA (siRNA), shorthairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. In a polynucleotide when referring to a T, a T means U (Uracil) in RNA and T (Thymine) in DNA. A polynucleotide can be16#507551Attorney Docket No. 00010.041.1801 exogenous or endogenous to a cell and / or exist in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., altered backbone, sugar, or nucleobase). If present, modifications to the nucleotide structure are imparted before or after assembly of the polymer. Non-limiting examples of modifications include: 5 -bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g.. rhodamine or fluorescein linked to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. The sequence of nucleotides may be interrupted by nonnucleotide components.

[0106] The terms ‘"transfection” or “transfected” refer to introduction of a polynucleotide into a cell by non-viral or viral-based methods. The polynucleotides may be gene sequences encoding complete proteins or functional portions thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.

[0107] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to refer to a polymer of at least two amino acid residues joined by peptide bond(s). This term does not connote a specific length of polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzy matic synthesis, or is naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some cases, the polymer is interrupted by non-amino acids. The terms include amino acid chains of any length, including full length proteins, and proteins with or without secondary7or tertiary structure (e.g., domains). The terms also encompass an amino acid polymer that has been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation such as conjugation with a labeling component. The terms “amino acid” and “amino acids,” as used herein, refer to natural and non-natural amino acids, including, but not limited to, modified amino acids. Modified amino acids include amino acids that have been chemically modified to include a group or a chemical moiety not naturally present on the amino acid. The term “amino acid” includes both D-amino acids and L-amino acids.

[0108] As used herein, the “non-native” refers to a nucleic acid or polypeptide sequence that is non-naturally occurring. Non-native refers to a non-naturally occurring nucleic acid or polypeptide sequence that comprises modifications such as mutations, insertions, or deletions. The term non-17#507551Attorney Docket No. 00010.041.1801 native encompasses fusion nucleic acids or polypeptides that encodes or exhibits an activity (e.g., enzymatic activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitinating activity, etc.) of the nucleic acid or polypeptide sequence to which the non-native sequence is fused. A non-native nucleic acid or polypeptide sequence includes those linked to a naturally -occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid or polypeptide sequence encoding a chimeric nucleic acid or polypeptide.

[0109] The term "promoter", as used herein, refers to the regulatory DNA region which controls transcription or expression of a polynucleotide (e.g., a gene) and which may be located adjacent to or overlapping a nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences which bind protein factors, often referred to as transcription factors, which facilitate binding of RNA polymerase to the DNA leading to gene transcription. Eukary otic basal promoters ty pically, though not necessarily, contain a TATA-box and / or a CAAT box.

[0110] The term "‘expression’; as used herein, refers to the process by which a nucleic acid sequence or a polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently7translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides may be collectively referred to as “gene product.” If the polynucleotide is derived from genomic DNA, the term expression includes splicing of the mRNA in a eukaryotic cell.[OHl] As used herein, “operably linked”, “operable linkage”, “operatively linked”, or grammatical equivalents thereof refer to an arrangement of genetic elements, e.g, a promoter, an enhancer, a polyadenylation sequence, etc., wherein an operation (e.g., movement or activation) of a first genetic element has some effect on the second genetic element. The effect on the second genetic element can be, but need not be, of the same type as operation of the first genetic element. For example, two genetic elements are operably linked if movement of the first element causes an activation of the second element. For instance, a regulatory element, which may comprise promoter and / or enhancer sequences, is operatively linked to a coding region if the regulatory7element helps initiate transcription of the coding sequence. There may be intervening residues between the regulatory element and coding region so long as this functional relationship is maintained.

[0112] A “vector” as used herein, refers to a macromolecule or association of macromolecules that comprises or associates with a polynucleotide and which mediates delivery of the polynucleotide18#507551Attorney Docket No. 00010.041.1801 to a cell. Examples of vectors include nucleic-based vectors (e.g., plasmids and viral vectors) and liposomes. An exemplary nucleic-acid based vector comprises genetic elements, e.g., regulatory elements, operatively linked to a gene to facilitate expression of the gene in a target.

[0113] As used herein, “expression cassette” and “nucleic acid cassette” are used interchangeably to refer to a component of a vector comprising a combination of nucleic acid sequences or elements (e.g., therapeutic gene, promoter, and a terminator) that are expressed together or are operably linked for expression. The terms encompass an expression cassette including a combination of regulatory elements and a gene or genes to which they are operably linked for expression.

[0114] A “functional fragment” of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) that is substantially similar to a biological activity of the full-length DNA or protein sequence. A biological activity of a DNA sequence includes its ability to influence expression in a manner attributed to the full-length sequence.

[0115] The terms “engineered,” “synthetic,” and “artificial” are used interchangeably herein to refer to an object that has been modified by human intervention. For example, the terms refer to a polynucleotide or polypeptide that is non-naturally occurring. An engineered peptide has. but does not require, low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to a naturally occurring human protein. Non-limiting examples include the following: a nucleic acid modified by changing its sequence to a sequence that does not occur in nature; a nucleic acid modified by ligating it to a nucleic acid that it does not associate with in nature such that the ligated product possesses a function not present in the original nucleic acid; an engineered nucleic acid synthesized in vitro with a sequence that does not exist in nature; a protein modified by changing its amino acid sequence to a sequence that does not exist in nature; an engineered protein acquiring a new function or property. An “engineered” system comprises at least one engineered component.

[0116] As used herein, the term “Casl2a” refers to a family of Cas endonucleases that are class 2, Type V-A Cas endonucleases and that (a) use a relatively small guide RNA (about 42-44 nucleotides) that is processed by the nuclease itself following transcription from the CRISPR array, and (b) cleave DNA to leave staggered cut sites.

[0117] As used herein, a “guide nucleic acid” or “guide polynucleotide” refers to a nucleic acid that may hybridize to a target nucleic acid and thereby directs an associated nuclease to the target nucleic acid. A guide nucleic acid is, but is not limited to, RNA (guide RNA or gRNA), DNA, or19#507551Attorney Docket No. 00010.041.1801 a mixture of RNA and DNA. A guide nucleic acid can include a crRNA or a tracrRNA or a combination of both. The term guide nucleic acid encompasses an engineered guide nucleic acid and a programmable guide nucleic acid to specifically bind to the target nucleic acid. A portion of the target nucleic acid may be complementary to a portion of the guide nucleic acid. The strand of a double-stranded target polynucleotide that is complementary to and hybridizes with the guide nucleic acid is the complementary strand. The strand of the double-stranded target polynucleotide that is complementary to the complementary strand, and therefore is not complementary to the guide nucleic acid is called noncomplementary strand. A guide nucleic acid having a polynucleotide chain is a “single guide nucleic acid.” A guide nucleic acid having two polynucleotide chains is a “double guide nucleic acid.” If not otherwise specified, the term “guide nucleic acid” is inclusive, referring to both single guide nucleic acids and double guide nucleic acids. A guide nucleic acid may comprise a segment referred to as a “nucleic acid-targeting segment” or a “nucleic acid-targeting sequence,” or a “spacer.” A nucleic acid-targeting segment can include a sub-segment referred to as a “protein binding segment” or “protein binding sequence” or “Cas protein binding segment.”

[0118] As used herein, the terms “gene editing” and “genome editing” can be used interchangeably. Gene editing or genome editing means to change the nucleic acid sequence of a gene or a genome. Genome editing can include, for example, insertions, deletions, and mutations.

[0119] As used herein, the term “complex” refers to ajoining of at least two components. The two components may each retain the properties / activities they had prior to forming the complex or gain properties as a result of forming the complex. The joining includes, but is not limited to, covalent bonding, non-covalent bonding (i.e., hydrogen bonding, ionic interactions, Van der Waals interactions, and hydrophobic bond), use of a linker, fusion, or any other suitable method. Contemplated components of the complex include polynucleotides, polypeptides, or combinations thereof. For example, a complex comprises an endonuclease and a guide polynucleotide.

[0120] The term “sequence identity” or “percent identity” in the context of two or more nucleic acids or polypeptide sequences, generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same, when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, e.g., BLASTP using parameters of a wordlength (W) of 3, an expectation (E) of 10, and20#507551Attorney Docket No. 00010.041.1801 the BLOSUM62 scoring matrix setting gap costs at existence of 11. extension of 1, and using a conditional compositional score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of a wordlength (W) of 2, an expectation (E) of 1000000, and the PAM30 scoring matrix setting gap costs at 9 to open gaps and 1 to extend gaps for sequences of less than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW with the Smith-Waterman homology search algorithm parameters with a match of 2, a mismatch of -1, and a gap of -1; MUSCLE with default parameters; MAFFT with parameters of a retree of 2 and max iterations of 1000; Novafold with default parameters; HMMER hmmalign with default parameters.

[0121] The term “optimally aligned” in the context of two or more nucleic acids or polypeptide sequences, generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that have been aligned to maximal correspondence of amino acids residues or nucleotides, for example, as determined by the alignment producing a highest or “optimized” percent identity score.

[0122] As used herein, the term “mismatch tolerance” is a ratio of percentage of indel formation at a mismatched off-target nucleic acid which denotes a percentage of off-target editing to the percentage of indel formation at an on-target nucleic acid sequence which denotes a percentage of on-target editing. A mismatch tolerance value of less than 0.5 indicates higher specificity and a lower mismatch tolerance for the engineered nuclease system, ^mismatch tolerance value of more than 0.5 indicates lower specificity and a higher tolerance for the engineered nuclease system. In some embodiments, the indel formation is quantified by sequencing.

[0124] Included in the current disclosure are variants of any of the enzymes described herein with one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be accomplished by substituting amino acids with similar hydrophobicity , polarity, and R chain length for one another. Additionally, or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by locating amino acid residues that have been mutated between species (e.g., non-conserved residues) without altering the basic functions of the encoded proteins. Such conservatively substituted variants may include variants with at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at21#507551Attorney Docket No. 00010.041.1801 least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%. at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% sequence identity to any one of the endonuclease protein sequences described herein (e.g. MG11, MG13, MG19, MG20, MG26,MG28, MG29, MG30, MG31, MG32, MG37, MG53, MG54, MG55, MG56, MG57, MG58,MG59, MG60, MG61, MG62, MG70, MG75, MG77, MG78, MG79, MG80, MG81, MG82,MG83, MG84, MG85, MG90, OR MG91 family endonucleases described herein, or any other family nuclease described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can encompass sequences with substitutions such that the activity of one or more critical active site residues or guide RNA binding residues of the endonuclease are not disrupted. In some embodiments, a functional variant of any of the proteins described herein lacks substitution of at least one conserved or functional residue. In some embodiments, a functional variant of any of the proteins described herein lacks substitution of all conserved or functional residues.

[0125] Also included in the current disclosure are variants of any of the enzymes described herein with substitution of one or more catalytic residues to decrease or eliminate activity of the enzy me (e.g. decreased-activity variants). In some embodiments, a decreased activity' variant as a protein described herein comprises a disrupting substitution of at least one, at least two, or all three catalytic residues.

[0126] Conservative substitution tables providing functionally similar amino acids are available from a variety7of references (see, for e.g., Creighton, Proteins: Structures and Molecular Properties (W H Freeman & Co.; 2nd edition (December 1993)). The following eight groups each contain amino acids that are conservative substitutions for one another:1) Alanine (A), Glycine (G);2) Aspartic acid (D), Glutamic acid (E);3) Asparagine (N), Glutamine (Q);4) Arginine (R), Lysine (K);5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);7) Serine (S), Threonine (T); and8) Cysteine (C), Methionine (M)22#507551Attorney Docket No. 00010.041.1801Overview

[0127] The discovery of new Cas enzymes with unique functionality and structure offers the potential to further gene editing technologies, improving speed, specificity, functionality, and ease of use. Relative to the predicted prevalence of Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) systems in microbes and the sheer diversity of microbial species, relatively few functionally characterized CRISPR / Cas enzymes exist in the literature. This is partly because a huge number of microbial species may not be readily cultivated in laboratory conditions. Metagenomic sequencing from natural environmental niches containing large numbers of microbial species may offer the potential to drastically increase the number of new CRISPR / Cas systems characterized and speed the discovery of new oligonucleotide editing functionalities. A recent example of the fruitfulness of such an approach is demonstrated by the 2016 discovery' of CasX / CasY CRISPR systems from metagenomic analysis of natural microbial communities.

[0128] CRISPR / Cas systems are RNA-directed nuclease complexes that function as an adaptive immune system in microbes. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally are made up of two parts: (i) an array of short repetitive sequences (30-40 bp) separated by short spacer sequences, which encode the RNA-based targeting element; and (ii) ORFs encoding the Cas nuclease. Efficient nuclease targeting of a particular target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target nucleic acid and a crRNA guide; and (ii) presence of a protospacer- adjacent motif (PAM) sequence within a certain vicinity of the target nucleic acid sequence depending on the specific Cas nuclease (the PAM usually being a sequence not commonly represented within the host genome). Depending on the exact function and organization of the system, CRISPR-Cas systems are commonly organized into 2 classes, 5 types and 16 subtypes based on shared functional characteristics and evolutionary' similarity.

[0129] Class 1 CRISPR-Cas systems have large, multi-subunit effector complexes, and include Types I, III, and IV Cas nucleases. Class 2 CRISPR-Cas systems generally have single-polypeptide multidomain nuclease effectors, and include Types II, V and VI Cas nucleases.

[0130] Type II CRISPR-Cas systems are considered the simplest in terms of components. In Type II CRISPR-Cas systems, the processing of the CRISPR array into mature crRNAs does not require the presence of a special endonuclease subunit, but rather a small trans-encoded crRNA (tracrRNA) with a region complementary to the array repeat sequence; the tracrRNA interacts with23#507551Attorney Docket No. 00010.041.1801 both its corresponding effector nuclease (e.g. Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. Cas II nucleases are identified as DNA nucleases. Type 2 effectors generally exhibit a structure comprising a RuvC-like endonuclease domain that adopts the RNase H fold with an unrelated HNH nuclease domain inserted within the folds of the RuvC- like nuclease domain. The RuvC-like domain is responsible for the cleavage of the target (e.g., crRNA complementary) DNA strand, while the HNH domain is responsible for cleavage of the displaced DNA strand.

[0131] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g. Casl2) structure similar to that of Type II effectors, comprising a RuvC-like domain. Similar to Type II, most (but not all) Type V CRISPR systems use a tracrRNA to process pre-crRNAs into mature crRNAs. However, unlike Type II systems which requires RNAse III to cleave the pre-crRNA into multiple crRNAs. Type V systems are capable of using the effector nuclease itself to cleave pre- crRNAs. Like Type-II CRISPR-Cas systems, Type V CRISPR-Cas systems are again identified as DNA nucleases. Unlike Type II CRISPR-Cas systems, some Type V enzymes (e.g.. Casl2a) appear to have a robust single-stranded nonspecific deoxyribonuclease activity that is activated by the first crRNA directed cleavage of a double-stranded target sequence.

[0132] CRISPR-Cas systems have emerged in recent years as the gene editing technology of choice due to their targetabi I i ty and ease of use. The most commonly used systems are the Class 2 Type II SpCas9 and the Class 2 Type V-A Casl2a. The Type V-A systems in particular are becoming more widely used since their reported specificity in cells is higher than other nucleases, with fewer or no off-target effects. The V-A systems are also advantageous in that the guide RNA is small (42-44 nucleotides compared with approximately 100 nt for SpCas9) and is processed by the nuclease itself following transcription from the CRISPR array, simplifying multiplexed applications with multiple gene edits. Furthermore, the V-A systems have staggered cut sites, which may facilitate directed repair pathways, such as microhomology-dependent targeted integration (MITI).

[0133] The most commonly used Type V-A enzymes require a 5’ protospacer adjacent motif (PAM) next to the chosen target site: 5’-TTTV-3’ for Lachnospiraceae bacterium ND2006 bCasl2a and Acidaminococcus sp. AsCasl2a; and 5’-TTV-3’ for Francisella novicida ^nCasl2a. Recent exploration of orthologs has revealed proteins with less restrictive PAM sequences that are also active in mammalian cell culture, for example YTV, YYN or TTN. However, these enzymes24#507551Attorney Docket No. 00010.041.1801 do not fully encompass V-A biodiversity and targetability, and may not represent all possible activities and PAM sequence requirements. Here, thousands of genomic fragments were identified from numerous metagenomes for Type V-A nucleases. The diversity of identified V-A enzymes may have been expanded and improved systems may have been developed into highly targetable, compact, and precise gene editing agents.MG Enzymes

[0134] Described herein, in certain embodiments, are endonucleases (e.g., Type V endonucleases).

[0135] In some embodiments, the endonuclease is a MG11 nuclease described herein (e.g., SEQ ID NOs: 1-37). In some embodiments, the endonuclease is a MG13 nuclease described herein (e.g., SEQ ID NOs: 38-118). In some embodiments, the endonuclease is a MG19 nuclease described herein (e.g., SEQ ID NOs: 119-124). In some embodiments, the endonuclease is a MG20 nuclease described herein (e.g., SEQ ID NO: 125). In some embodiments, the endonuclease is a MG26 nuclease described herein (e.g., SEQ ID NOs: 126-140). In some embodiments, the endonuclease is a MG28 nuclease described herein (e.g., SEQ ID NOs: 141-214). In some embodiments, the endonuclease is a MG29 nuclease described herein (e.g., SEQ ID NOs: 215-225 and 3471-3684). In some embodiments, the endonuclease is a MG30 nuclease described herein (e.g., SEQ ID NOs: 226-228). In some embodiments, the endonuclease is a MG31 nuclease described herein (e.g.. SEQ ID NOs: 229-260). In some embodiments, the endonuclease is a MG32 nuclease described herein (e.g., SEQ ID NO: 261). In some embodiments, the endonuclease is a MG37 nuclease described herein (e.g., SEQ ID NOs: 262-426). In some embodiments, the endonuclease is a MG53 nuclease described herein (e.g., SEQ ID NOs: 427-428). In some embodiments, the endonuclease is a MG54 nuclease described herein (e.g., SEQ ID NOs: 429-430). In some embodiments, the endonuclease is a MG55 nuclease described herein (e.g., SEQ ID NOs: 431-688). In some embodiments, the endonuclease is a MG56 nuclease described herein (e.g., SEQ ID NOs: 689-690). In some embodiments, the endonuclease is a MG57 nuclease described herein (e.g., SEQ ID NOs: 691- 721). In some embodiments, the endonuclease is a MG58 nuclease described herein (e.g., SEQ ID NOs: 722-779). In some embodiments, the endonuclease is a MG59 nuclease described herein (e.g., SEQ ID NOs: 780-792). In some embodiments, the endonuclease is a MG60 nuclease described herein (e.g., SEQ ID NOs: 793-1163). In some embodiments, the endonuclease is a MG61 nuclease described herein (e.g., SEQ ID NOs: 1164-1469). In some embodiments, the endonuclease is a MG62 nuclease described herein (e.g., SEQ ID NOs: 1470-1472). In some embodiments, the endonuclease is a MG70 nuclease described herein (e.g., SEQ ID NOs: 1473-25#507551Attorney Docket No. 00010.041.18011514). In some embodiments, the endonuclease is a MG75 nuclease described herein (e.g., SEQ ID NOs: 1515-1710). In some embodiments, the endonuclease is a MG77 nuclease described herein (e.g., SEQ ID NOs: 1711-1712). In some embodiments, the endonuclease is a MG78 nuclease described herein (e.g., SEQ ID NOs: 1713-1717). In some embodiments, the endonuclease is a MG79 nuclease described herein (e.g., SEQ ID NOs: 1718-1722). In some embodiments, the endonuclease is a MG80 nuclease described herein (e.g., SEQ ID NO: 1723). In some embodiments, the endonuclease is a MG81 nuclease described herein (e.g., SEQ ID NOs: 1724-2654). In some embodiments, the endonuclease is a MG82 nuclease described herein (e.g., SEQ ID NOs: 2655-2657). In some embodiments, the endonuclease is a MG83 nuclease described herein (e.g., SEQ ID NOs: 2658-2659). In some embodiments, the endonuclease is a MG84 nuclease described herein (e.g.. SEQ ID NOs: 2660-2677). In some embodiments, the endonuclease is a MG85 nuclease described herein (e.g., SEQ ID NOs: 2678-2680). In some embodiments, the endonuclease is a MG90 nuclease described herein (e.g., SEQ ID NOs: 2681- 2809). In some embodiments, the endonuclease is a MG91 nuclease described herein (e.g., SEQ ID NOs: 2810-3470 and 3685-3692).

[0136] In some embodiments, the endonucleases described herein are about 1000-1100 amino acids in length. In some embodiments, the endonucleases described herein are less than about 400 amino acids in length. In some embodiments, the endonucleases described herein are about 500- 700 amino acids in length. In some embodiments, the endonucleases described herein comprise RuvC and HTH DNA binding domains. In some embodiments, the endonuclease comprises a RuvCI, II, or III domain. In some embodiments, the RuvCI domain comprises a D catalytic residue. In some embodiments the RuvCII domain comprises an E catalytic residue. In some embodiments the RuvCIII domain comprises a D catalytic residue. In some embodiments, the RuvC domain does not have nuclease activity. In some embodiments, the endonuclease further comprises a WED II domain. In some embodiments, the endonuclease further comprises a zinc finger-like domain. In some embodiments, the endonucleases described herein do not require a tracrRNA.

[0137] In some embodiments, the endonuclease is a Cas endonuclease. In some embodiments, he endonuclease is a class 2, type V Cas endonuclease. In some embodiments, the endonuclease is a class 2, type V-A Cas endonuclease. In some embodiments, the endonuclease is not a Cpfl or Cmsl endonuclease. In some embodiments, the endonuclease further comprises a zinc finger-like domain.26#507551Attorney Docket No. 00010.041.1801

[0138] In some embodiments, the endonuclease has at least about 70% sequence identity to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%. at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1- 3692. In some embodiments, the endonuclease has at least about 75% identity to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 80% identity' to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 85% identity to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 90% identity to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 95% identity7to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 96% identity to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 97% identity to any one of SEQ ID NOs: 1- 3692. In some embodiments, the endonuclease has at least about 98% identity to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has at least about 99% identity7to any one of SEQ ID NOs: 1-3692. In some embodiments, the endonuclease has 100% identity to any one of SEQ ID NOs: 1-3692.

[0139] In some embodiments, the endonuclease has at least about 70% sequence identity to any one of SEQ ID NOs: 3685-3692, and 3471-3684. In some embodiments, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%. at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity7to any one of SEQ ID NOs: 3685-3692, and 3471-3684. In some embodiments, the endonuclease has at least about 75% identity to any one of SEQ ID NOs: 3685-3692, and 3471-3684. In some embodiments, the endonuclease has at least about 80% identity to any one of SEQ ID NOs: 3685-3692, and 3471- 3684. In some embodiments, the endonuclease has at least about 85% identity to any one of SEQ ID NOs: 3685-3692, and 3471-3684. In some embodiments, the endonuclease has at least about 90% identity to any one of SEQ ID NOs: 3685-3692. and 3471-3684. In some embodiments, the27#507551Attorney Docket No. 00010.041.1801 endonuclease has at least about 95% identity to any one of SEQ ID NOs: 3685-3692. and 3471- 3684. In some embodiments, the endonuclease has at least about 96% identity to any one of SEQ ID NOs: 3685-3692, and 3471-3684. In some embodiments, the endonuclease has at least about 97% identity to any one of SEQ ID NOs: 3685-3692, and 3471-3684. In some embodiments, the endonuclease has at least about 98% identity to any one of SEQ ID NOs: 3685-3692, and 3471- 3684. In some embodiments, the endonuclease has at least about 99% identity to any one of SEQ ID NOs: 3685-3692, and 3471-3684. In some embodiments, the endonuclease has 100% identity to any one of SEQ ID NOs: 3685-3692, and 3471-3684.

[0140] In some embodiments, the endonuclease is a MG91 endonuclease (e.g., SEQ ID NOs: 3685-3692). In some embodiments, the endonuclease has at least about 70% sequence identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%. at least about 92%. at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 75% identity' to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 80% identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 85% identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 90% identity' to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 95% identity to any one of SEQ ID NOs: 3685- 3692. In some embodiments, the endonuclease has at least about 96% identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 97% identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 98% identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has at least about 99% identity to any one of SEQ ID NOs: 3685-3692. In some embodiments, the endonuclease has 100% identity to any one of SEQ ID NOs: 3685-3692.

[0141] In some embodiments, the endonuclease is a MG29 endonuclease (e.g., SEQ ID NOs: 215- 225 and 3471-3684). In some embodiments, the endonuclease has at least about 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at28#507551Attorney Docket No. 00010.041.1801 least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%. at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 75% identity’ to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 80% identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 85% identity7to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 90% identity’ to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 95% identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 96% identity’ to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 97% identity' to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 98% identity to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has at least about 99% identity7to any one of SEQ ID NOs: 215-225 and 3471-3684. In some embodiments, the endonuclease has 100% identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

[0142] In some embodiments, endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence. In some embodiments, the PAM has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity’ to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 70% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 75% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 80% identity’ to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 85% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 90% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 95% identity' to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 96% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least29#507551Attorney Docket No. 00010.041.1801 about 97% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 98% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has at least about 99% identity to any one of the sequences listed in Table 3. In some embodiments, the PAM has 100% identity to any one of the sequences listed in Table 3.

[0143] In some embodiments, the PAM has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%. at least 80%. at least 81%. at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity sequence identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 70% identity to any one of SEQ ID NOs: 54 and 96. In some embodiments, the PAM has at least about 75% identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 80% identity' to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 85% identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 90% identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 95% identity to TtTYn. GnYYn, or wCCC. In some embodiments, the PAM has at least about 96% identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 97% identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 98% identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has at least about 99% identity to TtTYn, GnYYn, or wCCC. In some embodiments, the PAM has 100% identity to TtTYn, GnYYn, or wCCC.

[0144] In some embodiments, the PAM has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%. at least 94%. at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity’ sequence identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or ryCC. In some embodiments, the PAM has at least about 70% identity to any one of SEQ ID NOs: 54 and 96. In some embodiments, the PAM has at least about 75% identity to tnTYn, GnGYCn. TTTY, Cc. Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc. tngn, gnGY, mCm, or ryCC. In some embodiments, the PAM has at least about 80% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or ryCC. In some embodiments, the PAM has at least about 85% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt. yYy, TtGc, tngn, gnGY, mCm, or ryCC. In some embodiments,#507551Attorney Docket No. 00010.041.1801 the PAM has at least about 90% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn. yYt, yYy, TtGc, tngn, gnGY, mCm. or ryCC. In some embodiments, the PAM has at least about 95% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or ryCC. In some embodiments, the PAM has at least about 96% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or ryCC. In some embodiments, the PAM has at least about 97% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or ryCC. In some embodiments, the PAM has at least about 98% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or ryCC. In some embodiments, the PAM has at least about 99% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn. yYt, yYy, TtGc, tngn, gnGY. mCm, or ryCC. In some embodiments, the PAM has 100% identity to tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or ryCC.

[0145] In some embodiments, the endonuclease comprises a PI (PAM interacting) domain having at least about 80%. at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85% at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to a PI domain.

[0146] In some embodiments, the endonucleases are discovered through metagenomic sequencing. In some embodiments, the metagenomic sequencing is conducted on samples. In some embodiments, the samples are collected from a variety of environments. Such environments may be a human microbiome, an animal microbiome, environments with high temperatures, environments with low temperatures. Such environments may include sediment.

[0147] In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLSs) proximal to an N- or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence of any one of SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence of any one of SEQ ID NOs: 3765-3780, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least31#507551Attorney Docket No. 00010.041.1801 about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 80% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 85% identity- to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 90% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 91% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS compnses a sequence having at least about 92% identity- to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 93% identity' to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 94% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 95% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 96% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 97% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 98% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having at least about 99% identity to SEQ ID NOs: 3765-3780. In some embodiments, the NLS comprises a sequence having 100% identity to SEQ ID NOs: 3765-3780.Table 1: Example NLS Sequences that may be used with Cas Effectors according to the disclosure.#507551Attorney Docket No. 00010.041.1801Guide Polynucleotides

[0148] Disclosed herein, in certain embodiments, are engineered endonuclease systems comprising (a) an endonuclease disclosed herein, and (b) an engineered guide polynucleotide e.g., a guide ribonucleic acid (gRNA), a single gRNA, or a dual guide RNA. In a polynucleotide when referring to a T, a T means U (Uracil) in RNA and T (Thymine) in DNA.

[0149] In some embodiments, the engineered guide polynucleotide is configured to form a complex with the engineered endonuclease. In some embodiments, the engineered guide polynucleotide comprises a spacer sequence. In some embodiments, the spacer sequence is configured to hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence.

[0150] In some embodiments, the guide polynucleotide (e.g., gRNA) targets a gene or locus in a cell. In some embodiments, the guide polynucleotide targets a gene or locus in a mammalian cell. In some embodiments, the mammalian cell is a pig, a cow, a goat, a sheep, a rodent, a rat, a mouse, a non-human primate, or a human cell. In some embodiments, the target gene or target locus is albumin, B2M, TRAC, and VCP.

[0151] In some embodiments, the guide polynucleotide (e.g., gRNA) comprises a sequence having least 80% sequence identity to the first 19 nucleotides or the non-degenerate nucleotides of SEQ ID NO: 3695. In some embodiments, the guide polynucleotide comprises a sequence with at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%. at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least33#507551Attorney Docket No. 00010.041.1801 about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the first 19 nucleotides or the non-degenerate nucleotides of SEQ ID NO: 3695. In some embodiments, the guide polynucleotide comprises a sequence 100% identical to the first 19 nucleotides or the nondegenerate nucleotides of SEQ ID NO: 3695.

[0152] In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity7to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity7to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity7to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity to SEQ ID NO: 3695.

[0153] In some embodiments, the target gene is B2M. In some embodiments, the guide polynucleotide is encoded by any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 20%, at least about 25%. at least about 30%, at least about 35%, at least about 40%, at least about45%. at least about 50%. at least about 55%, at least about 60%, at least about 65%, at least about70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having34#507551Attorney Docket No. 00010.041.1801 at least about 90% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856.

[0154] In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a target nucleic acid sequence within the B2M gene (e.g. within an intron of the B2M gene). In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3731- 3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity to any one of SEQ ID NOs: 3731-3737. 3746-3752. and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 85% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 95% identity' to any one of SEQ ID NOs: 3731- 3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a35#507551Attorney Docket No. 00010.041.1801 sequence having at least about 98% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary' to a sequence having at least about 99% identity' to any one of SEQ ID NOs: 3731- 3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having 100% identity' to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856.

[0155] In some embodiments, the guide polynucleotide hybridizes or targets a sequence within the B2M gene (e.g., within the intron of the B2M gene). In some embodiments, the guide polynucleotide hybridizes or targets a sequence according to any one of SEQ ID NOs: 3857-3932 or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 80% identity to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 85% identity to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 90% identity to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 95% identity' to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 96% identity' to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 97% identity to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 98% identity' to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 99% identity' to any one of SEQ ID NOs: 3857-3932. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having 100% identity to any one of SEQ ID NOs: 3857-3932.

[0156] In some embodiments, the target gene is TRAC. In some embodiments, the guide polynucleotide is encoded by any one of SEQ ID NOs: 3745, 3760, and 3933-4004, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide compnses a sequence comprising at least about 46-80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%. at least about 60%, at least about 65%, at least about 70%, at least about36#507551Attorney Docket No. 00010.041.180175%, at least about 80%. at least about 85%, at least about 90%, at least about 91%, at least about 92%. at least about 93%. at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity to any one of SEQ ID NOs: 3745, 3760, and 3933- 4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity to any one of SEQ ID NOs: 3745. 3760, and 3933-4004.

[0157] In some embodiments, the target gene is TIM-3. In some embodiments, the guide polynucleotide is encoded by any one of SEQ ID NOs: 3738- 3744 and 3753- 3759, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about92%, at least about 93%. at least about 94%, at least about 95%, at least about 96%, at least about97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90%37#507551Attorney Docket No. 00010.041.1801 identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759.

[0158] In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a target nucleic acid sequence within the TIM-3 gene (e.g., within an intron of the TIM-3 gene). In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 85% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 95% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 98% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a38#507551Attorney Docket No. 00010.041.1801 sequence having at least about 99% identity to any one of SEQ ID NOs: 3738- 3744 and 3753-3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary' to a sequence having 100% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759.

[0159] In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary’ to a target nucleic acid sequence within the TRAC gene (e.g., within an intron of the TRAC gene). In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary' to any one of SEQ ID NOs: 3745, 3760, and 3933-4004, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity’ to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity to any one of SEQ ID NOs: 3745,3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary' to a sequence having at least about 85% identity’ to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary' to a sequence having at least about 95% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity’ to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary' to a sequence having at least about 98% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 99% identity to any one of SEQ ID NOs: 3745, 3760, and 3933- 4004. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary’ to a sequence having 100% identity to any one of SEQ ID NOs: 3745, 3760, and 3933-4004.

[0160] In some embodiments, the guide polynucleotide hybridizes or targets a sequence within the TRAC gene (e.g., within the intron of the TRAC gene). In some embodiments, the guide polynucleotide hybridizes or targets a sequence according to any one of SEQ ID NOs: 4005-407639#507551Attorney Docket No. 00010.041.1801 or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 80% identity to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 85% identity to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 90% identity to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 95% identity to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 96% identity7to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 97% identity to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 98% identity' to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 99% identity7to any one of SEQ ID NOs: 4005-4076. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having 100% identity to any one of SEQ ID NOs: 4005-4076.

[0161] In some embodiments, the target gene is VCP. In some embodiments, the guide polynucleotide is encoded by any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727 or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide comprises a sequence comprising at least about 46-80 consecutive nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about45%, at least about 50%. at least about 55%, at least about 60%, at least about 65%, at least about70%. at least about 75%. at least about 80%, at least about 85%, at least about 90%, at least about91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about96%, at least about 97%, at least about 98%, or at least about 99% identity7to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 80% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 85% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 90% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In40#507551Attorney Docket No. 00010.041.1801 some embodiments, the guide polynucleotide is encoded by a sequence having at least about 95% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 96% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 97% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 98% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having at least about 99% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide is encoded by a sequence having 100% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727.

[0162] In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a target nucleic acid sequence within the VCP gene. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727, or a sequence having at least 90%, 95%. 97%. 98%, or 99% sequence identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 80% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 85% identity to any one of SEQ ID NOs: 4077- 4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 90% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 95% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 96% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 97% identity to any one of SEQ ID NOs: 4077- 4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary' to a sequence having at least about 98% identity to any one of SEQ ID NOs: 4077-4082. 3721-3723, and 3725-3727. In some embodiments, the guide41#507551Attorney Docket No. 00010.041.1801 polynucleotide hybridizes or targets a sequence complementary to a sequence having at least about 99% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to a sequence having 100% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725- 3727.

[0163] In some embodiments, the guide polynucleotide hybridizes or targets a sequence within the VCP gene (e.g., within the intron of the VCP gene). In some embodiments, the guide polynucleotide hybridizes or targets a sequence according to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 80% identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 85% identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 90% identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 95% identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 96% identity to any one of SEQ ID NOs: 3710-3718. 3721-3723. and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 97% identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 98% identity to any one of SEQ ID NOs: 3710-3718. 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having at least about 99% identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence having 100% identity to any one of SEQ ID NOs: 3710-3718, 3721-3723, and 3725-3727.

[0164] In some embodiments, the guide polynucleotide is configured to form a complex with the endonuclease. In some embodiments, the guide polynucleotide binds to the endonuclease to form a complex. In some embodiments, the guide polynucleotide binds (e.g., non-covalently through42#507551Attorney Docket No. 00010.041.1801 electrostatic interactions or hydrogen bonds) to the endonuclease to form a complex. In some embodiments, the guide polynucleotide is fused to the endonuclease to form a complex.

[0165] In some embodiments, the guide polynucleotide comprises a spacer sequence. In some embodiments, the spacer sequence is configured to hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence.

[0166] In some embodiments, the guide polynucleotide (e.g.. gRNA) targets a gene or locus in a cell. In some embodiments, the guide polynucleotide targets a gene or locus in a mammalian cell. In some embodiments, the mammalian cell is a pig, a cow, a goat, a sheep, a rodent, a rat, a mouse, a non-human primate, or a human cell.

[0167] In some embodiments, the guide polynucleotides (e.g., guide RNAs) comprise various structural elements including but not limited to: a spacer sequence which binds to the protospacer sequence (target sequence), a crRNA, and an optional tracrRNA. In some embodiments, the genome editing system comprises a CRISPR guide RNA. In some embodiments, the guide RNA comprises a crRNA comprising a spacer sequence. In some embodiments, the guide RNA additionally comprises a tracrRNA or a modified tracrRNA.

[0168] In some embodiments, the systems provided herein comprise one or more guide polynucleotides. In some embodiments, the guide polynucleotide comprises a sense sequence. In some embodiments, the guide polynucleotide comprises an anti-sense sequence. In some embodiments, the guide polynucleotide comprises nucleotide sequences other than the region complementary to or substantially complementary to a region of a target sequence. For example, a crRNA is part or considered part of a guide polynucleotide, or is comprised in a guide polynucleotide, e.g., a crRNA: tracrRNA chimera.

[0169] In some embodiments, the guide polynucleotide comprises synthetic nucleotides or modified nucleotides. In some embodiments, the guide polynucleotide comprises one or more inter-nucleoside linkers modified from the natural phosphodiester. In some embodiments, all of the inter-nucleoside linkers of the guide polynucleotide, or contiguous nucleotide sequence thereof, are modified. For example, in some embodiments, the inter nucleoside linkage comprises Sulphur (S), such as a phosphorothioate inter-nucleoside linkage. In some embodiments, the guide polynucleotide comprises greater than about 10%, 25%, 50%, 75%, or 90% modified inter- nucleoside tinkers. In some embodiments, the guide polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8,43#507551Attorney Docket No. 00010.041.18019, 10, or more than 10 modified inter-nucleoside linkers (e.g., phosphorothioate inter-nucleoside linkage).

[0170] In some embodiments, the guide polynucleotide comprises modifications to a ribose sugar or nucleobase. In some embodiments, the guide polynucleotide comprises one or more nucleosides comprising a modified sugar moiety, wherein the modified sugar moiety is a modification of the sugar moiety when compared to the ribose sugar moiety found in deoxyribose nucleic acid (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acids (LNA)), or an unlinked ribose ring which typically lacks a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleosides comprise bicyclohexose nucleic acids or tricyclic nucleic acids. In some embodiments, the modified nucleosides comprise nucleosides where the sugar moiety is replaced with a non-sugar moiety, for example peptide nucleic acids (PNA) or morpholino nucleic acids.

[0171] In some embodiments, the guide polynucleotide comprises one or more modified sugars. In some embodiments, the sugar modifications comprise modifications made by altering the substituent groups on the ribose ring to groups other than hydrogen, or the 2’-OH group naturally found in DNA and RNA nucleosides. In some embodiments, substituents are introduced at the 2’, 3’, 4', 5‘ positions, or combinations thereof. In some embodiments, nucleosides with modified sugar moieties comprise 2’ modified nucleosides, e.g.. 2’ substituted nucleosides. A 2’ sugar modified nucleoside, in some embodiments, is a nucleoside that has a substituent other than H or -OH at the 2’ position (2’ substituted nucleoside) or comprises a 2’ linked biradical, and comprises 2’ substituted nucleosides and LNA (2’-4’ biradical bridged) nucleosides. Examples of 2’- substituted modified nucleosides comprise, but are not limited to, 2’-O-alkyl-RNA, 2’-O-methyl- RNA, 2’-alkoxy-RNA, 2’-O-methoxyethyl- RNA (MOE), 2’-amino-DNA, 2’-Fluoro-RNA, and 2’-F-ANA nucleoside. In some embodiments, the modification in the ribose group comprises a modification at the 2’ position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2’-O-methyl, 2’-fluoro, 2'- deoxy, and 2’-O-(2-methoxyethyl).

[0172] In some embodiments, the guide polynucleotide comprises one or more modified sugars. In some embodiments, the guide polynucleotide comprises only modified sugars. In some embodiments, the guide polynucleotide comprises greater than about 10%, 25%, 50%. 75%, or44#507551Attorney Docket No. 00010.041.180190% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises a 2’-O-methyl. In some embodiments, the modified sugar comprises a 2’ -fluoro. In some embodiments, the modified sugar comprises a 2 -0- methoxyethyl group. In some embodiments, the guide polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 modified sugars (e g., comprising a 2’-O-methyl or 2'-fluoro).

[0173] In some embodiments, the guide polynucleotide comprises both inter-nucleoside linker modifications and nucleoside modifications. In some embodiments, the guide polynucleotide comprises greater than about 10%, 25%, 50%, 75%, or 90% modified inter-nucleoside linkers and greater than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the guide polynucleotide comprises 1, 2, 3, 4. 5, 6, 7, 8, 9, 10, or more than 10 modified inter-nucleoside linkers (e.g., phosphorothioate inter-nucleoside linkage) and 1, 2, 3. 4, 5, 6. 7, 8, 9, 10, or more than 10 modified sugars (e.g, comprising a 2’-O-methyl or 2’-fluoro).

[0174] In some embodiments, the guide polynucleotide (e.g., gRNA) comprises at least one of the following modifications: (i) a 2’-0 methyl or a 2'-fluoro base modification of at least one nucleotide within the first 4 bases of the 5' end of the guide polynucleotide or the last 4 bases of a 3’ end of the guide polynucleotide; (ii) a thiophosphate (PS) linkage between at least 2 of the first five bases of a 5’ end of the guide polynucleotide, or a thiophosphate linkage between at least two of the last five bases of a 3’ end of the guide polynucleotide; (iii) a thiophosphate linkage within a 3’ stem or a 5’ stem of the guide polynucleotide; (iv) a 2'-0 methyl or 2’ base modification within a 3’ stem or a 5’ stem of the guide polynucleotide; (v) a 2’-fluoro base modification of at least 7 bases of a spacer region of the guide polynucleotide; and (vi) a thiophosphate linkage within a loop region of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises a 2’-0 methyl or a 2’-fluoro base modification of at least one nucleotide within the first 5 bases of a 5' end the guide polynucleotide or the last 5 bases of a 3' end of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises a 2’-0 methyl or a 2’-fluoro base modification at a 5’ end of the guide polynucleotide or a 3’ end of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises a thiophosphate (PS) linkage between at least 2 of the first five bases of a 5’ end of the guide polynucleotide, or a thiophosphate linkage between at least two of the last five bases of a 3’ end of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises a thiophosphate linkage within a 3’ stem or a 5’ stem of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises a 2'-0 methyl base modification within a 3‘ stem or a 5’ stem of the guide polynucleotide. In some embodiments,45#507551Attorney Docket No. 00010.041.1801 the guide polynucleotide comprises a 2’ -fluoro base modification of at least 7 bases of a spacer region of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises a thiophosphate linkage within a loop region of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least three 2’-0 methyl or 2’-fluoro bases at the 5’ end of the guide polynucleotide, two thiophosphate linkages between the first 3 bases of the 5’ end of the guide polynucleotide, at least 4 2 -0 methyl or 2’-fluoro bases at the 4’ end of the guide polynucleotide, and three thiophosphate linkages between the last three bases of the 3’ end of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least two 2’- O-methyl bases and at least two thiophosphate linkages at a 5‘ end of the guide polynucleotide and at least one 2'-O-methyl bases and at least one thiophosphate linkage at a 3’ end of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least one 2 -0- methyl base in both the 3’ stem or the 5’ stem region of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least one to at least fourteen 2’-fluoro bases in the spacer region excluding a seed region of the guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least one 2’-O-methyl base in the 5’ stem region of the guide polynucleotide and at least one to at least fourteen 2’ -fluoro bases in the spacer region excluding a seed region of the guide RNA. In some embodiments, the guide RNA comprises a spacer sequence having at least 80% identity to any one of SEQ ID NOs: 3746-3764. In some embodiments, the guide RNA comprises the nucleotides of any one of SEQ ID NOs: 3746-3764 comprising a chemical modification. In some embodiments, the RNA-guided nuclease is a Cas endonuclease. In some embodiments, the Cas endonuclease is a class 2, type V Cas endonuclease. In some embodiments, the class 2, type V Cas endonuclease comprises a RuvC domain comprising a RuvCI subdomain, a RuvCII subdomain, and a RuvCIII subdomain. In some embodiments, the class 2, tvpe V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3692 or a variant thereof. In some embodiments, the class 2, type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721. In some embodiments, the guide polynucleotide comprises a sequence with at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NOs: 3695.

[0175] In some embodiments, the guide polynucleotide comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence complementary to a eukaryotic46#507551Attorney Docket No. 00010.041.1801 genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence complementary’ to a fungal genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence complementary to a plant genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence complementary to a mammalian genomic polynucleotide sequence. In some embodiments, the guide polynucleotide comprises a sequence complementary to a human genomic polynucleotide sequence.

[0176] In some embodiments, the guide polynucleotide is 30-250 nucleotides in length. In some embodiments, the guide polynucleotide is more than 90 nucleotides in length. In some embodiments, the guide polynucleotide is less than 245 nucleotides in length. In some embodiments, the guide polynucleotide is 30, 40, 50, 60, 70, 80, 90, 100. 120, 140. 160, 180, 200, 220, 240, or more than 240 nucleotides in length. In some embodiments, the guide polynucleotide is about 30 to about 40, about 30 to about 50, about 30 to about 60, about 30 to about 70, about 30 to about 80, about 30 to about 90, about 30 to about 100, about 30 to about 120, about 30 to about 140, about 30 to about 160, about 30 to about 180, about 30 to about 200, about 30 to about 220, about 30 to about 240, about 50 to about 60, about 50 to about 70, about 50 to about 80, about 50 to about 90, about 50 to about 100, about 50 to about 120, about 50 to about 140, about 50 to about 160, about 50 to about 180, about 50 to about 200, about 50 to about 220, about 50 to about 240, about 100 to about 120, about 100 to about 140, about 100 to about 160, about 100 to about 180, about 100 to about 200, about 100 to about 220, about 100 to about 240, about 160 to about 180, about 160 to about 200, about 160 to about 220, or about 160 to about 240 nucleotides in length.

[0177] In some embodiments, the guide polynucleotide comprises a hairpin comprising at least 8 base-paired ribonucleotides. In some embodiments, the guide polynucleotide comprises a hairpin comprising at least 9 base-paired ribonucleotides. In some embodiments, the guide polynucleotide comprises a hairpin comprising at least 10 base-paired ribonucleotides. In some embodiments, the guide polynucleotide comprises a hairpin comprising at least 11 base-paired ribonucleotides. In some embodiments, the guide polynucleotide comprises a hairpin comprising at least 12 basepaired ribonucleotides.

[0178] In some embodiments, the guide polynucleotide comprises a DNA-targeting segment. In some embodiments, the DNA-targeting segment comprises a nucleotide sequence that is complementary’ to a target sequence. In some embodiments, the target sequence is in a target DNA molecule. In some embodiments, the guide polynucleotide comprises (b) a protein-binding47#507551Attorney Docket No. 00010.041.1801 segment. In some embodiments, the protein-binding segment comprises two complementary stretches of nucleotides. In some embodiments, the two complementary stretches of nucleotides hybridize to form a double-stranded RNA (dsRNA) duplex. In some embodiments, the two complementary' stretches of nucleotides are covalently linked to one another with intervening nucleotides.

[0179] In some embodiments, the DNA-targeting segment is positioned 3’ of both of the two complementary' stretches of nucleotides.

[0180] In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 8 ribonucleotides. In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 9 ribonucleotides. In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 10 ribonucleotides. In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 11 ribonucleotides. In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 12 ribonucleotides.MG Endonuclease Systems

[0181] Described herein, in certain embodiments, are engineered nuclease systems comprising an engineered endonuclease and an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence.

[0182] In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 70% identity' to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725-3727, 3731-3737, 3746-3752, 3781- 3856, 3738- 3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 80% identity' to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 80% identity to any one of SEQ48#507551Attorney Docket No. 00010.041.1801ID NOs: 3695. 3721-3723, 3725-3727, 3731-3737, 3746-3752. 3781-3856, 3738- 3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725-3727, 3731-3737, 3746-3752, 3781-3856, 3738- 3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725- 3727, 3731-3737, 3746-3752, 3781-3856, 3738- 3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725-3727, 3731-3737, 3746- 3752, 3781-3856, 3738- 3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease composing a sequence having at least about 96% identity' to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 3695. 3721-3723, 3725-3727. 3731-3737, 3746-3752, 3781-3856, 3738-3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725-3727, 3731-3737, 3746-3752, 3781-3856, 3738- 3744, 3753- 3759,3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 1-3692 and49#507551Attorney Docket No. 00010.041.1801 b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725- 3727, 3731-3737, 3746-3752, 3781-3856, 3738- 3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725-3727, 3731-3737, 3746- 3752, 3781-3856, 3738- 3744, 3753- 3759, 3745, and 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising 100% identity to any one of SEQ ID NOs: 1-3692 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising 100% identity to any one of SEQ ID NOs: 3695, 3721-3723, 3725- 3727, 3731-3737, 3746-3752. 3781-3856, 3738- 3744, 3753- 3759, 3745. and 3760. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to any one of SEQ ID NOs: 3695, 3721-3723, 3725-3727, 3731-3737, 3746-3752, 3781-3856, 3738- 3744, 3753- 3759, 3745, and 3760 or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3695. 3721-3723, 3725-3727, 3731-3737, 3746-3752, 3781- 3856, 3738- 3744, 3753- 3759, 3745, and 3760.

[0183] In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 70% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 75% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide50#507551Attorney Docket No. 00010.041.1801 polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 80% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 85% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 90% identity7to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 90% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 95% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 96% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 97% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 98% identity7to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a51#507551Attorney Docket No. 00010.041.1801 sequence having at least about 99% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising a sequence having at least about 99% identity to SEQ ID NO: 3695. In some embodiments, the engineered nuclease system comprises an endonuclease comprising 100% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, the engineered guide polynucleotide comprising 100% identity to SEQ ID NO: 3695. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to SEQ ID NO: 3695 or a sequence having at least 90%. 95%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 3695.

[0184] In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to fonn a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene (e.g., within an intron of the VCP gene), the engineered guide polynucleotide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide52#507551Attorney Docket No. 00010.041.1801 polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 98% identity' to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a53#507551Attorney Docket No. 00010.041.1801VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 21 -225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising a sequence having at least about 99% identity' to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the engineered nuclease system comprises an endonuclease comprising 100% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a VCP gene, the engineered guide polynucleotide comprising 100% identity to any one of SEQ ID NOs: 4077-4082, 3721-3723, and 3725-3727. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary’ to any one of SEQ ID NOs: 4077-4082. 3721-3723, and 3725-3727 or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 4077- 4082, 3721-3723, and 3725-3727.

[0185] In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene (e.g., within an intron of the B2M gene), the engineered guide polynucleotide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752. and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 75% identity' to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 75% identity' to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide54#507551Attorney Docket No. 00010.041.1801 polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 97% identity' to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a55#507551Attorney Docket No. 00010.041.1801B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 98% identity' to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the engineered nuclease system comprises an endonuclease comprising 100% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a B2M gene, the engineered guide polynucleotide comprising 100% identity to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to any one of SEQ ID NOs: 3731-3737, 3746-3752, and 3781-3856, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3731- 3737, 3746-3752, and 3781-3856.

[0186] In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene (e.g., within an intron of the TRAC gene), the engineered guide polynucleotide comprising a sequence having at least about 70% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 75% identity' to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to56#507551Attorney Docket No. 00010.041.1801 form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 75% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 80% identity' to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 80% identity' to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 85% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 90% identity7to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 95% identity' to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 95% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide57#507551Attorney Docket No. 00010.041.1801 polynucleotide comprising a sequence having at least about 96% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 97% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 98% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising a sequence having at least about 99% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the engineered nuclease system comprises an endonuclease comprising 100% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TRAC gene, the engineered guide polynucleotide comprising 100% identity to SEQ ID NO: 3745 or SEQ ID NO: 3760. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary to SEQ ID NO: 3745 or SEQ ID NO: 3760 or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 3745 or SEQ ID NO: 3760.

[0187] In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a58#507551Attorney Docket No. 00010.041.1801 target nucleic acid sequence within a TIM-3 gene (e.g., within an intron of the TIM-3 gene), the engineered guide polynucleotide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 75% identity7to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 90% identity7to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 95% identity to any59#507551Attorney Docket No. 00010.041.1801 one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 97% identity7to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the engineered nuclease system comprises an endonuclease comprising 100% identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a TIM-3 gene, the engineered guide polynucleotide comprising 100% identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759. In some embodiments, the guide polynucleotide hybridizes or targets a sequence complementary7to any one of SEQ ID NOs:60#507551Attorney Docket No. 00010.041.18013738- 3744 and 3753- 3759, or a sequence having at least 90%, 95%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759.

[0188] In some embodiments, the engineered nuclease system further comprises a single- or double stranded DNA repair template. In some embodiments, the engineered nuclease system further comprises a single-stranded DNA repair template. In some embodiments, the engineered nuclease system further comprises a double-stranded DNA repair template. In some embodiments, the single- or double-stranded DNA repair template comprises from 5’ to 3’: a first homology arm comprising a sequence of at least 20 nucleotides 5' to said target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homology' arm comprising a sequence of at least 20 nucleotides 3' to said target sequence.

[0189] In some embodiments, the first homology arm comprises a sequence of at least 40. at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides. In some embodiments, the second homology arm comprises a sequence of at least 40. at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 1 10, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides.

[0190] In some embodiments, the first and second homology arms are homologous to a genomic sequence of a prokaryote. In some embodiments, the first and second homology’ arms are homologous to a genomic sequence of a bacteria. In some embodiments, the first and second homology arms are homologous to a genomic sequence of a fungus. In some embodiments, the first and second homology' arms are homologous to a genomic sequence of a eukaryote.

[0191] In some embodiments, engineered nuclease system further comprises a DNA repair template. In some embodiments, the DNA repair template comprises a double-stranded DNA segment. The double-stranded DNA segment may be flanked by one single-stranded DNA segment. In some embodiments, the double-stranded DNA segment are flanked by two singlestranded DNA segments. In some embodiments, the single-stranded DNA segments are conjugated to the 5’ ends of the double-stranded DNA segment. In some embodiments, the single stranded DNA segments are conjugated to the 3’ ends of the double-stranded DNA segment.

[0192] In some embodiments, the single-stranded DNA segments have a length from 1 to 15 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length from 4 to 10 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length of61#507551Attorney Docket No. 00010.041.18014 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length of 5 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length of 6 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length of 7 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length of 8 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length of 9 nucleotide bases. In some embodiments, the single-stranded DNA segments have a length of 10 nucleotide bases.

[0193] In some embodiments, the single-stranded DNA segments have a nucleotide sequence complementary to a sequence within the spacer sequence. In some embodiments, the doublestranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein-coding sequence, a miRNA coding sequence, an RNA coding sequence, or a transgene.

[0194] In some embodiments, the engineered nuclease system further comprises a source of Mg2+.

[0195] In some embodiments, the engineered nuclease system comprises 20 pmoles or less of the endonuclease (e.g., class 2, ty pe V Cas endonuclease). In some embodiments, the engineered nuclease system comprises 1 pmol or less of the endonuclease.

[0196] In some embodiments, the engineered nuclease system comprises: (a) an endonuclease comprising a RuvC domain, wherein the endonuclease is derived from an uncultivated microorganism, and wherein the endonuclease is a Cas 12a endonuclease; and (b) an engineered guide RNA, wherein the guide polynucleotide is configured to form a complex with the endonuclease and the guide polynucleotide comprises a spacer sequence configured to hybridize to a target nucleic acid sequence. In some embodiments, the Cas 12a endonuclease comprises the sequence GWxxxK. In some embodiments, the engineered guide RNA comprises UCUAC[N3- s]GUAGAU (NQ (SEQ ID NO: 9947). In some embodiments, the engineered guide RNA comprises CCUGC[N4]GCAGG (N3-4) (SEQ ID NO: 9948).

[0197] In some embodiments, the engineered nuclease system comprises: (a) a class 2, Type V-A Cas endonuclease configured to bind a 3- or 4-nucleotide PAM sequence, wherein the endonuclease has increased cleavage activity' relative to sMbCasl2a; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the class 2, Type V-A Cas endonuclease and the engineered guide RNA comprises a spacer sequence configured to hybridize to a target nucleic acid comprising a target nucleic acid sequence. In some embodiments, the cleavage activity' is measured in vitro by introducing the endonucleases alongside compatible guide RNAs to cells comprising the target nucleic acid and detecting cleavage of the target nucleic62#507551Attorney Docket No. 00010.041.1801 acid sequence in the cells. In some embodiments, the class 2, Type V-A Cas endonuclease has at least 75% identity to any one of 215-225 or a variant thereof. In some embodiments, the target nucleic acid further comprises a YYN PAM sequence proximal to the target nucleic acid sequence. In some embodiments, the class 2, Type V-A Cas endonuclease has at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or 200%, or more increased activity relative to sMbCasl2a.Delivery and Vectors

[0198] Disclosed herein, in some embodiments, are nucleic acid sequences encoding an engineered nuclease system described herein or components thereof (e g., endonuclease, engineered guide polynucleotide).

[0199] In some embodiments, the nucleic acid encoding the engineered nuclease system described herein or components thereof is a DNA, for example a linear DNA, a plasmid DNA, or a minicircle DNA. In some embodiments, the nucleic acid encoding the engineered nuclease system described herein or components thereof is an RNA, for example a mRNA.

[0200] In some embodiments, the nucleic acid encoding the engineered nuclease system described herein or components thereof is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g, circular DNA molecules that can autonomously replicate inside a cell), cosmid (e.g, pWE or sCos vectors), artificial chromosome, human artificial chromosome (HAC), yeast artificial chromosomes (YAC), bacterial artificial chromosome (BAC), Pl -derived artificial chromosomes (PAC), phagemid, phage derivative, bacmid, or virus. In some embodiments, the nucleic acid-based vector is selected from the list consisting of: pSF-CMV- NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST- TEV. pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF- CMV-FMDV-daGFP, pEFla-mCherry-Nl vector, pEFla-tdTomato vector, pSF-CMV-FMDV- Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20- BetaGal,pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, pRI 10I-AN DNA, pCambia2301, pTYB21, pKLAC2, pAc5. 1 / V5-His A, and pDEST8.

[0201] In some embodiments, the nucleic acid-based vector comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a mini promoter, an inducible promoter, a constitutive promoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EFla, CAG, PGK. TRE, U6, UAS, T7, Sp6,63#507551Attorney Docket No. 00010.041.1801 lac. araBad, trp, Ptac, p5, pl 9, p40, Synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments the promoter is a U6 promoter. In some embodiments, the promoter is a CAG promoter.

[0202] In some embodiments, the nucleic acid-based vector is a virus. In some embodiments, the virus is an alphavirus, a parvovirus, an adenovirus, an AAV, a baculovirus, a Dengue virus, a lentivirus, a herpesvirus, a poxvirus, an anellovirus, a bocavirus, a vaccinia virus, or a retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the vims is a Dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the vims is or a retrovirus.

[0203] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8. AAV9. AAV10, AAV11, AAV12, AAV13, AAV14. AAV15, AAV16, AAV-rh8. AAV- rhlO, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-l, AAV-hu37, AAV-Anc80, AAV- Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV- LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV- HSC7, AAV-HSC8. AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT. AAV-DJ / 8. AAV-Myo, AAV-NP40, AAV-NP59. AAV- NP22, AAV-NP66, AAV-HSC16, or a derivative thereof. In some embodiments, the herpesvims is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.

[0204] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the vims is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the vims is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the vims is AAV9 or a derivative thereof. In some embodiments, the virus is AAV 10 or a derivative thereof. In some embodiments, the virus is AAV 11 or a derivative thereof. In some embodiments, the virus is AAV 12 or a derivative thereof. In some embodiments, the virus is AAV 13 or a derivative thereof. In some embodiments, the virus is AAV 14 or a derivative thereof. In some embodiments, the virus is64#507551Attorney Docket No. 00010.041.1801AAV 15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rhlO or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-l or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV -PHP -B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof. In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the vims is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the vims is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the vims is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV -HSC 12 or a derivative thereof. In some embodiments, the vims is AAV-HSC13 or a derivative thereof. In some embodiments, the vims is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the vims is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV -NP66 or a derivative thereof. In some embodiments, the virus is AAV -HSC 16 or a derivative thereof.65#507551Attorney Docket No. 00010.041.1801

[0205] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.

[0206] In some embodiments, the nucleic acid encoding the engineered nuclease system described herein or components thereof is delivered by anon-nucleic acid-based delivery system (e.g., anon- viral delivery' system). In some embodiments, the non-viral delivery' system is a liposome. In some embodiments, the nucleic acid is associated with a lipid. The nucleic acid associated with a lipid, in some embodiments, is encapsulated in the aqueous interior of a liposome, interspersed within the lipid bilayer of a liposome, attached to a liposome via a linking molecule that is associated with both the liposome and the nucleic acid, entrapped in a liposome, complexed with a liposome, dispersed in a solution containing a lipid, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, contained or complexed with a micelle, or otherwise associated with a lipid. In some embodiments, the nucleic acid is comprised in a lipid nanoparticle (LNP).

[0207] In some embodiments, the engineered nuclease system described herein or components thereof is introduced into the cell in any suitable way, either stably or transiently. In some embodiments, the engineered nuclease system described herein or components thereof is transfected into the cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct that encodes the engineered nuclease system described herein or components thereof. For example, a cell is transduced (e.g., with a virus encoding the engineered nuclease system described herein or components thereof), or transfected (e.g., with a plasmid encoding the engineered nuclease system described herein or components thereof) with a nucleic acid that encodes the engineered nuclease system described herein or components thereof, or the translated the engineered nuclease system described herein or components thereof. In some embodiments, the transduction is a stable or transient transduction. In some embodiments, cells expressing the engineered nuclease system described herein or components thereof or containing the engineered nuclease system described herein or components thereof are transduced or transfected with one or more gRNA molecules, for example, when the engineered nuclease system described herein or components thereof comprises a CRISPR nuclease. In some embodiments, a plasmid expressing the engineered nuclease system described herein or components thereof is introduced into cells66#507551Attorney Docket No. 00010.041.1801 through electroporation, transient (e.g., lipofection) and stable genome integration (e.g., piggybac) and viral transduction (for example lentivirus or AAV) or other methods known to those of skill in the art. In some embodiments, the gene editing system is introduced into the cell as one or more polypeptides. In some embodiments, delivery is achieved through the use of RNP complexes. Delivery methods to cells for polypeptides and / or RNPs are known in the art, for example by electroporation or by cell squeezing.

[0208] Exemplary methods of delivery of nucleic acids include lipofection, nucleofection, electroporation, stable genome integration (e.g., piggybac), microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid nucleic acid conjugates, naked DNA, artificial virions, and agent-enhanced uptake of DNA. Lipofection is described in e.g., U.S. Pat. Nos. 5,049,386; 4,946,787; and 4,897,355) and lipofection reagents are sold commercially (e.g., Transfectam™, Lipofectin™ and SF Cell Line 4D-Nucleofector X Kit™ (Lonza)). Cationic and neutral lipids that are suitable for efficient receptor-recognition lipofection of polynucleotides include those of WO 91 / 17424 and WO 91 / 16024. In some embodiments, the delivery' is to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g.. in vivo administration). In some embodiments, the nucleic acid is comprised in a liposome or a nanoparticle that specifically targets a host cell.

[0209] Additional methods for the delivery of nucleic acids to cells are known to those skilled in the art. See, for example, US 2003 / 0087817.

[0210] In some embodiments, the present disclosure provides a cell comprising a vector or a nucleic acid described herein. In some embodiments, the cell expresses a gene editing system or parts thereof. In some embodiments, the cell is a human cell. In some embodiments, the cell is genome edited ex vivo. In some embodiments, the cell is genome edited in vivo.Lipid nanoparticles

[0211] Lipid nanoparticles as described herein can be 4-component lipid nanoparticles. Such nanoparticles can be configured for delivery' of RNA or other nucleic acids (e.g. synthetic RNA, mRNA, or in v / 7ro-synthesized mRNA) and can be generally formulated as described in WO2012135805A2. Such nanoparticles can generally comprise: (a) a cationic lipid (e.g. 98N12-5 (TETA5-LAP), DLin DMA, DLin-K-DMA (2,2-Dilinoleyl-4-dimethylaminomethyl-[l,3]- dioxolane), DLin-KC2-DMA, DLin-MC3-DMA, or C12-200), (b) a neutral lipid (e.g. DSPC or67#507551Attorney Docket No. 00010.041.1801DOPE), (c) a sterol (e.g. cholesterol or a cholesterol analog), and (d) a PEG-modified lipid (e.g. PEG-DMG).

[0212] The cationic lipid referred to herein as “Cl 2-200” is disclosed by Love et al., Proc Natl Acad Sci USA. 2010 107: 1864-1869 and Liu and Huang, Molecular Therapy. 2010 669-670. Cationic lipid formulations can include particles comprising either 3 or 4 or more components in addition to polynucleotide, primary construct, or RNA (e.g. mRNA). As an example, formulations with certain cationic lipids include, but are not limited to, 98N 12-5, and may contain 42% lipidoid, 48% cholesterol, and 10% PEG (Cl 4 or greater alky 1 chain length). As another example, formulations with certain lipidoids include, but are not limited to, Cl 2-200 and may contain 50% cationic lipid, 10% disteroylphosphatidyl choline, 38.5% cholesterol, and 1.5% PEG-DMG.

[0213] In some embodiments, the cationic lipid nanoparticle comprises a cationic lipid, a PEG- modified lipid, a sterol, and anon-cationic lipid. In some embodiments, the cationic lipid is selected from the group consisting of 98N12-5 (TETA5-LAP), DLin DMA, DLin-K-DMA (2,2-Dilinoleyl- 4-dimethylaminomethyl-[l,3]-dioxolane), DLin-KC2-DMA, DLin-MC3-DMA, and C 12-200. In some embodiments, the cationic lipid nanoparticle has a molar ratio of about 20-60% cationic lipid, about 5-25% non-cationic lipid, about 25-55% sterol, and about 0.5-15% PEG-modified lipid. In some embodiments, the cationic lipid nanoparticle comprises a molar ratio of about 50% cationic lipid, about 1.5% PEG-modified lipid, about 38.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid nanoparticle comprises a molar ratio of about 55% cationic lipid, about 2.5% PEG-modified lipid, about 32.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid is an ionizable cationic lipid, the non-cationic lipid is a neutral lipid, and the sterol is a cholesterol. In some embodiments, the cationic lipid nanoparticle has a molar ratio of 50:38.5: 10: 1.5 of cationic lipid: cholesterol: PEG2000-DMG:DSPC or DMG:DOPE. In some embodiments, lipid nanoparticles as described herein can comprise cholesterol, l,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), l,l ‘-((2-(4-(2-((2-(bis(2- hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-l- yl)ethyl)azanediyl)bis(dodecan-2-ol) (C 12-200), and DMG-PEG-2000 at molar ratios of 47.5: 16:35: 1.5.Cells

[0214] Described herein, in certain embodiments, is a cell comprising the engineered nuclease system described herein.68#507551Attorney Docket No. 00010.041.1801

[0215] In some embodiments, the cell is a eukary otic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungi cell), a mammalian cell (a Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embry o kidney (HEK), mouse myeloma (NSO), or human retinal cells), an immortalized cell (e.g., a HeLa cell, a COS cell, a HEK-293T cell, a MDCK cell, a 3T3 cell, a PC 12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, aN2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, a S2 cell, or aHeliothis virescens cell), a yeast cell (e.g., a Saccharomyces cerevisiae cell, a Cryptococcus cell, or a Candida cell), a plant cell (e.g. , a parenchyma cell, a collenchyma cell, or a sclerenchyma cell), a fungal cell (e.g., a Saccharomyces cerevisiae cell, a Cryptococcus cell, or a Candida cell), or a prokaryotic cell (e.g., a E. coli cell, a streptococcus bacterium cell, a streptomyces soil bacteria cell, or an archaea cell). In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.

[0216] In some embodiments, the cell is an A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, a primary7cell, or derivative thereof.

[0217] In some embodiments, the cell is an E. coli cell or a mammalian cell. In some embodiments, the cell is an E. coli cell, wherein the E. coli cell is a XDE3 lysogen or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT Ion genotype.

[0218] In some embodiments, the cell is aT cell. In some embodiments, the cell is a hematopoietic cell.Methods of Use

[0219] Systems of the present disclosure may be used for various applications, such as, for example, nucleic acid editing (e g., gene editing), binding to a nucleic acid molecule (e.g., sequence-specific binding). Such systems may be used, for example, for addressing (e.g., removing or replacing) a genetically inherited mutation that may cause a disease in a subject, inactivating a gene in order to ascertain its function in a cell, as a diagnostic tool to detect disease-causing genetic elements (e.g. via cleavage of reverse-transcribed viral RNA or an amplified DNA sequence encoding a disease-causing mutation), as deactivated enzymes in combination with a probe to69#507551Attorney Docket No. 00010.041.1801 target and detect a specific nucleotide sequence (e.g. sequence encoding antibiotic resistance int bacteria), to render viruses inactive or incapable of infecting host cells by targeting viral genomes, to add genes or amend metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish a gene drive element for evolutionary selection, to detect cell perturbations by foreign small molecules and nucleotides as a biosensor.

[0220] Described herein, in certain embodiments, are methods of modifying a target nucleic acid sequence comprising contacting the target nucleic acid sequence using the engineered nuclease systems described herein. In some embodiments, modifying the target nucleic acid sequence comprises binding, nicking, marking, or cleaving the target nucleic acid sequence. In some embodiments, modifying the target nucleic acid sequence comprises binding, nicking, or cleaving the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence is within a B2M, TRAC, or VCP.

[0221] In some embodiments, the target nucleic acid sequence comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid sequence comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid sequence is modified in vitro. In some embodiments, the target nucleic acid sequence is within a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell.

[0222] In some embodiments, the endonuclease is a Cas endonuclease. In some embodiments, the endonuclease is a class 2 Cas endonuclease. In some embodiments, the endonuclease is a class 2, type V Cas endonuclease. In some embodiments, the endonuclease is a class 2, type V-A Cas endonuclease. In some embodiments, the endonuclease is in complex with a guide polynucleotide. In some embodiments, the guide polynucleotide is configured to bind to the endonuclease. In some embodiments, the guide polynucleotide is configured to bind to the double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the guide polynucleotide is configured to bind to the endonuclease and to the double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM). In some embodiments, the PAM comprises a sequence comprising any one of the sequences listed in Table 3. In some embodiments, the PAM comprises a sequence comprising TtTYn, GnYYn, or wCCC. In some embodiments, the PAM70#507551Attorney Docket No. 00010.041.1801 comprises a sequence comprising tnTYn, GnGYCn, TTTY, Cc, Gnkynn, tTYnAA, nnnCn, yYt, yYy, TtGc. tngn. gnGY, mCm, or ryCC.

[0223] In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of the guide polynucleotide and a second strand comprising the PAM. In some embodiments, the PAM is directly adjacent to the 5' end of the sequence complementary to the sequence of the guide polynucleotide. In some embodiments, the endonuclease is not a Cpfl endonuclease or a Cmsl endonuclease. In some embodiments, the endonuclease is derived from an uncultivated microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the PAM comprises any one of the sequences listed in Table 3. In some embodiments, the PAM comprises a sequence comprising TtTYn, GnYYn, or wCCC. In some embodiments, the PAM comprises a sequence comprising tnTYn, GnGYCn, TTTY, Cc, Gnkynn, ttTYnAA, nnnCn, yYt, yYy, TtGc, tngn, gnGY, mCm, or r CC

[0224] In some embodiments, methods described herein comprise modifying a target nucleic acid sequence. In some embodiments, the method comprises delivering to the target nucleic acid sequence the engineered nuclease system described herein. In some embodiments, the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure. In some embodiments, the complex is configured such that upon binding of the complex to the target nucleic acid sequence, the complex modifies the target nucleic acid sequence.

[0225] In some embodiments, delivery of the engineered nuclease system to the target nucleic acid sequence comprises delivering the nucleic acid described herein or the vector described herein. In some embodiments, delivery' of engineered nuclease system to the target nucleic acid sequence comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease. In some embodiments, the nucleic acid comprises a promoter. In some embodiments, the open reading frame encoding the endonuclease is operably linked to the promoter.

[0226] In some embodiments, delivery' of the engineered nuclease system to the target nucleic acid sequence comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease. In some embodiments, delivery of the engineered nuclease system to the target nucleic acid sequence comprises delivering a translated polypeptide. In some embodiments, delivery' of the engineered nuclease system to the target nucleic acid sequence comprises delivering71#507551Attorney Docket No. 00010.041.1801 a deoxyribonucleic acid (DNA) encoding the guide polynucleotide operably linked to a ribonucleic acid (RNA) pol III promoter.

[0227] In some embodiments, the open reading frame is operably linked to a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP\>\v) promoter, a strong leftward promoter from phage lambda (pL promoter), or any combination thereof. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to a sequence encoding the endonuclease. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the endonuclease via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a Thrombin cleavage site, a Factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof.

[0228] In some embodiments, the open reading frame is codon-optimized for expression in the host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into a genome of the host cell.

[0229] In some embodiments, the endonuclease induces a single-stranded break or a doublestranded break at or proximal to the target sequence. In some embodiments, the endonuclease induces a staggered single stranded break within or 3' to said target sequence.

[0230] In some embodiments, effector repeat motifs are used to inform guide design of MG nucleases. For example, the processed gRNA in Type V-A systems comprises the last 20-22 nucleotides of a CRISPR repeat. This sequence may be synthesized into a crRNA (along with a spacer) and tested in vitro, along with the synthesized nucleases, for cleavage on a library of possible targets. Using this method, the PAM may be determined. In some embodiments. Type V- A enzymes may use a ’‘universal” gRNA. In some embodiments, Type V enzymes may utilize a unique gRNA.

[0231] In some embodiments, the present disclosure provides for a method of producing an endonuclease, comprising cultivating any of the host cells described herein in compatible growth72#507551Attorney Docket No. 00010.041.1801 medium. In some embodiments, the method further comprises inducing expression of the endonuclease. In some embodiments, the inducing expression of the nuclease is by addition of an additional chemical agent or an increased amount of a nutrient, or by temperature increase or decrease. In some embodiments, an additional chemical agent or an increased amount of a nutrient comprises Isopropyl 0-D-1 -thiogalactopyranoside (IPTG) or additional amounts of lactose. In some embodiments, the method further comprises isolating the host cell after the cultivation and lysing the host cell to produce a protein extract. In some embodiments, the method further comprises isolating the endonuclease. In some embodiments, the isolating comprises subjecting the protein extract to IMAC, ion-exchange chromatography, anion exchange chromatography, or cation exchange chromatography. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to a sequence encoding the endonuclease. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the endonuclease via a linker sequence encoding protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a Thrombin cleavage site, a Factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the method further comprises cleaving the affinity tag by contacting a protease corresponding to the protease cleavage site to the endonuclease. In some embodiments, the affinity7tag is an IMAC affinity7tag. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from a composition comprising the endonuclease.Off-Target Editing Efficiency

[0232] The present disclosure provides methods that deploy dual-target libraries and nextgeneration sequencing to measure editing at off-target sequences alongside the matched on-target sequence which in turn provide information on nuclease specificity. Employing both off and on target nucleic acid sequences in the same construct provides a built-in, same-well normalization of editing rate, reducing confounders from delivery, expression, and cell state. Off-target “hits” are identified when the candidate sequence shows measurable indels relative to the on-target control in the same construct / sample. The assays disclosed herein map how single-base mismatches across the spacer and / or PAM affect off-target editing. In certain embodiments, the assays measure editing as a functional proxy for off-target activity7wherein a site is called off-target if it is actually cleaved / edited relative to the on-target control within the same dual-target construct.73#507551Attorney Docket No. 00010.041.1801

[0233] Described herein, are methods for identifying off-target editing efficiency relative to a target DNA sequence, comprising: a) contacting a cell with a library of polynucleotides and an engineered nuclease system, the engineered nuclease system comprising: an engineered endonuclease having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471- 3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease, wherein each polynucleotide of the library comprises: (i) an on-target nucleic acid sequence, (ii) an off-target nucleic acid sequence containing a single nucleotide mismatch at a defined position within a protospacer adjacent motif (PAM) or a spacer relative to the on-target sequence, (iii) a first barcode associated with the off-target nucleic acid sequence and a second barcode associated with the on-target nucleic acid sequence, (iv) a linker sequence separating the off-target nucleic acid sequence and the on-target nucleic acid sequence, and (v) adapter primers flanking each of the on-target and off-target nucleic acid sequences; b) isolating genomic DNA from the cell: c) amplifying the target DNA sequence using the adapter primers; and d) sequencing the first and second barcodes and aligning the sequence reads assigned to each polynucleotide of the library' to a target DNA sequence and quantifying an off-target indel frequency and an on-target indel frequency, and computing a mismatch tolerance value, thereby identifying off-target editing efficiency relative to the target DNA sequence for the engineered nuclease-engineered guide polynucleotide complex.

[0234] Also described herein, are methods for determining specificity of a nuclease, the method comprising: a) contacting a cell with a library of polynucleotides and an engineered nuclease system, the engineered nuclease system comprising: an engineered endonuclease having at least 70% sequence identify to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease, wherein each polynucleotide of the library’ comprises: (i) an on-target nucleic acid sequence, (ii) an off-target nucleic acid sequence containing a single nucleotide mismatch at a defined position within a protospacer adjacent motif (PAM) or a spacer relative to the on-target sequence, (iii) a first barcode associated with the off-target nucleic acid sequence and a second barcode associated with the on- target nucleic acid sequence, (iv) a linker sequence separating the off-target nucleic acid sequence and the on-target nucleic acid sequence, and (v) adapter primers flanking each of the on-target and off-target nucleic acid sequences; b) isolating genomic DNA from the cell; c) amplifying the target DNA sequence using the adapter primers; and d) sequencing the first and second barcodes and aligning the sequence reads assigned to each polynucleotide of the library to a target DNA74#507551Attorney Docket No. 00010.041.1801 sequence and quantifying an off-target indel frequency and an on-target indel frequency, and computing a mismatch tolerance value, thereby determining specificity of a nuclease.

[0235] In certain embodiments, the linker is 10 to 40 nucleotides in length. For example, about 10 nucleotides to about 20 nucleotides, about 11 nucleotides to about 20 nucleotides, about 12 nucleotides to about 20 nucleotides, about 13 nucleotides to about 18 nucleotides, about 28 to about 38 nucleotides, etc. In certain embodiments, the linker is unique across at least a subset of the library members to reduce barcode recombination during amplification. In certain embodiments, identifying comprises recovering a distal segment of a barcode when a proximal segment is removed by a deletion event. In certain embodiments, computing the mismatch tolerance value comprises calculating a bound off-target to bound on-target ratio and designating a library member to be off-target editing when the ratio exceeds a predetermined threshold value of 0.5. In certain embodiments, contacting comprises delivering the complex to a cell by nucleofection or lipid nanoparticle formulation.EXAMPLESExample 1 - Sequences of the wild-type and mutant VCP gene present in IBMPFD patients and the VCPR155H / +mouse model of the disease used in this study

[0236] Inclusion body myopathy with Paget disease of bone and frontotemporal dementia (IBMPFD) is a multisystem proteinopathy. This is a progressive dominant disorder that is adultonset and ultimately lethal. The disease is caused by heterozygous missense mutations in valosin- containing protein (VCP), an ATPase involved in protein degradation and autophagy . Arginine 155 at the N-terminal domain of VCP / p97 is the most common amino acid affected, and the substitution of this residue to histidine, R155H, is the most common VCP mutation linked to IBMPFD. Transgenic mice expressing VCP / p97 harboring this mutation recapitulate the full spectrum of IBMPFD, including degeneration in muscle, brain, and bone, serving as a useful model for studying the disease. Importantly, mice carry ing only one allele of VCP / p97 were reported to be indistinguishable from their wild-type littermates.

[0237] The Example describes targeting the mutant VCP allele with CRISPR nucleases described herein. Here and in other examples as provided herein, MG29-1 with VCP mutation-specific guides gr2133, gr2134, and gr2135 were tested on fibroblasts derived from VCPR155H / +disease model mice and found to primarily edit the mutant allele while preserving the wild-type allele. Subsequently, guide gr2135 was tested in IBMPFD patient T cells and found to edit the75#507551Attorney Docket No. 00010.041.1801 mutant allele but not the wild-type allele. In addition, an increase in the wild-type allele percentage above the original 50% was observed in patient T cells.

[0238] FIG. 1 shows the DNA sequence of the codon implicated in IBMPFD with flanking sequences for both humans and the disease model mice used in this example. Arginine 155 is encoded by CGT on one allele (the wild-ty pe allele) in human patients. The corresponding codon on the mutant allele is CAT encoding histidine. Arginine 155 is encoded by CGG on one allele (the wild-type allele) in mouse models of the disease. The corresponding codon on the mutant allele in disease model mice is the same as in humans, CAT encoding histidine. FIG. 1 further shows SEQ ID NOs: 3717-3720.Example 2 - Evaluation of allele-specific editing across MG29-1 guides designed to target the mutant VCP allele in mouse VCPR15:,H / +fibroblasts

[0239] In this Example, mRNA from MG29-1 and chemi cally-modified, synthetic guides targeting the R155H mutant allele of the mouse VCP gene were electroporated into fibroblasts from VCPR155H +mice. The frequency of total indels and out of frame (OOF) indels was measured at the locus by NGS.

[0240] Methods

[0241] Cell culture, transfections, next generation sequencing, and editing analysis

[0242] Guide spacer SEQ ID NOs: 3721-3724. full-length guide SEQ ID NOs: 3725-3728, and mRNA sequence SEQ ID NOs: 3729 and 3730 were evaluated.

[0243] Mouse fibroblasts were cultured in high glucose DMEM supplemented with 15% FBS, lx GlutaMAX, 1 mM Pyruvate, 10 ug / mL gentamycin, 1% Antibiotic-Antimycotic, 1% Pen / Strep, and 2.5 ug / mL plasmocin prophylactic. Per electroporation reaction, 50,000 cells were trypsinized and washed twice in PBS before resuspension in 20 uL of SF buffer with 150 pmols of guide RNA and 500 ng of nuclease mRNA. The mixture was transferred to a nucleocuvette and pulsed with code CD-137. Cells were rested for 1-3 minutes in the cuvette and then transferred to a 96-well plate containing prewarmed media for three (3) days.

[0244] Genomic DNA (gDNA) was extracted. The target genomic regions were amplified from the extracted gDNA using DNA polymerase with primers designed for NGS-based sequencing. PCR products were purified. The amplicons were sequenced and analyzed using a script to quantify gene editing efficiency.

[0245] Results76#507551Attorney Docket No. 00010.041.1801

[0246] The average indels across 3 replicates were 47% for gr2133, 47% for gr2134, 52% for gr2135, and 52% for gr2155 (FIG. 2). The average out-of-frame indels across 3 replicates were 34% for gr2133, 33% for gr2134, 37% for gr2135, and 50% for gr2155.Example 3 - Sequencing outcomes show mutant allele-specific editing by MG29-1 guide gr2133 in mouse VCPR155H / +fibroblasts

[0247] The Example descnbes the sequencing outcomes specific to MG29-1 and guide gr2133 targeting the R155H mutant allele of the mouse VCP gene as presented in Example 2 above.

[0248] Methods

[0249] Methods adopted in this Example are the same as methods outlined in Example 2 above. Elere, guide spacer SEQ ID NO: 3721, full-length guide SEQ ID NO: 3725, and mRNA SEQ ID NO: 3729 were evaluated.

[0250] Results

[0251] The unedited, wild-type VCP sequence is represented in 49.41% of sequencing reads, while the intact unedited R155H mutant sequence is absent, demonstrating mutant-allele specific editing by MG29-1 guide gr2133 in mouse disease model cells (FIG. 3).Example 4 - Sequencing outcomes show mutant allele-specific editing by MG29-1 guide gr2134 in mouse VCPR155H / +fibroblasts.

[0252] The Example describes the sequencing outcomes specific to MG29-1 and guide gr2134 targeting the R155H mutant allele of the mouse VCP gene as presented in Example 2.

[0253] Methods

[0254] Methods adopted in this Example are the same as methods outlined in Example 2 above. Here, guide spacer SEQ ID NO: 3722, full-length guide SEQ ID NO: 3726, and mRNA SEQ ID NO: 3729 were evaluated.

[0255] Results

[0256] The unedited, wild-type sequence is represented in 48.48% of sequencing reads, while the intact unedited R155H mutant sequence is absent, demonstrating mutant-allele specific editing by MG29-1 guide gr2134 in mouse disease model cells (FIG. 4).Example 5 - Sequencing outcomes show mutant allele-specific editing by MG29-1 guide gr2135 in mouse VCPR155H / +fibroblasts77#507551Attorney Docket No. 00010.041.1801

[0257] The Example describes the sequencing outcomes specific to MG29-1 and guide gr2135 targeting the R155H mutant allele of the mouse VCP gene as presented in Example 2 above are shown.

[0258] Methods

[0259] Methods adopted in this Example are the same as methods outlined in Example 2 above. In this example, guide spacer SEQ ID NO: 3723, full-length guide SEQ ID NO: 3727, and mRNA S EQ ID NO: 3729 were evaluated.

[0260] Results

[0261] The unedited, wild-type sequence is represented in 45.33% of sequencing reads, while the intact unedited R155H mutant sequence is absent, demonstrating mutant-allele specific editing by MG29-1 guide gr2135 in mouse disease model cells (FIG. 5).Example 6 - Sequencing outcomes show mutant allele-specific editing by Cas9 guide gr2155 in mouse VCPR155H / +fibroblasts.

[0262] The Example describes the sequencing outcomes specific to Cas9 and guide gr2155 targeting the R155H mutant allele of the mouse VCP gene as presented in Example 2 above are shown.

[0263] Methods

[0264] Methods adopted in this Example are the same as methods outlined in Example 2 above. In this example, guide spacer SEQ ID NO: 3724, full-length guide SEQ ID NO: 3728, and mRNA SEQ ID NO: 3730 were evaluated.

[0265] Results

[0266] The unedited, wild-type sequence is represented in 45.51% of sequencing reads, while the intact unedited R155H mutant sequence is present in only 1.00% of reads, demonstrating mutantallele specific editing by Cas9 guide gr2155 in mouse disease model cells (FIG. 6).Example 7 - MG29-1 guide gr2135 targeting of the mutant VCP allele in patient T cells

[0267] The Example describes MG29-1 mRNA mr!26 and a chemically-modified, synthetic guide targeting the R155H mutant allele of the human VCP gene denoted gr2135 in IBMPFD patient T cells.

[0268] Methods

[0269] Cell culture, transfections, next generation sequencing, and editing analysis78#507551Attorney Docket No. 00010.041.1801

[0270] Guide spacer SEQ ID NO: 3723 and full-length guide SEQ ID NO: 3727 were evaluated with mRNA SEQ ID NO: 3729.

[0271] T cells were purified from IBMPFD patient blood. Cells were plated at a density of 0.5xl06 / mL in prewarmed T cell expansion medium supplemented with 5% CTS Immune Cell Serum Replacement and 100 lU / mL of IL-2. Cells were then activated with beads from the T Cell Activation and Expansion kit at a 1 :1 bead to cell ratio. After three (3) days, the activation beads were removed and the cells replated in the supplemented media above with 100 lU / mL IL-2 at IxlO6cells / mL. The cells were harvested the next day, washed twice in PBS, and resuspended at 2xl05cells per 20 uL of supplemented P3 buffer. Per electroporation reaction, 150 pmols guide RNA and 500 ng of nuclease mRNA were added. Cells were electroporated using code DS 120 and then grown in the supplemented media above with IL-2 for 3 days.

[0272] Genomic DNA (gDNA) was extracted. The target genomic regions were amplified from the extracted gDNA using the DNA polymerase with primers designed forNGS-based sequencing. PCR products were purified. The amplicons were sequenced and analyzed to quantify gene editing efficiency.

[0273] Results

[0274] Frequency of the mutant allele was reduced from 50% to 16.46% (FIG. 7). Frequency of the wild-type allele increased from 50% to 54.43%. This indicates that mutant allele-specific editing occurred with MG29-1 and guide gr2135 in IBMPFD patient T cells, accompanied by a 4.43% increase in the wild-type allele frequency.Example 8 - Mismatch tolerance evaluation of Type II nucleases

[0275] The Example describes evaluation of mismatch tolerance by nucleases described herein.

[0276] Lentivirus was used to deliver a library of dual-targets oligos, one containing a single mismatch in the spacer or PAM compared to the on-target, which are then edited in 96-well format by nucleofection with mRNA or RNP with chemically modified guides. The mismatch tolerance, obtained from the ratio of OFF : ON target editing for each dual target oligo, was plotted by guide and highlights the guide-intrinsic mismatch tolerance observed. The mean and 95% confidence intervals were plotted across all guides for a given nuclease, specifically MG29-1, by position and by mismatch ty pe, a nuclease-specific single mismatch tolerance profile was observed.

[0277] Methods79#507551Attorney Docket No. 00010.041.1801

[0278] A 96-well mismatch tolerance assay was developed (FIG. 8) in which cells were transduced with a library of dual-targets, one containing a single mismatch in the spacer or PAM compared to the on-target, and edited by nucleofection to directly compare the editing efficiencies observed at both mismatched and on-target sequences using chemically modified guides.

[0279] Library design

[0280] Single guide RNAs that represented a range of GC contents and percent indels were chosen for use in this assay (FIG. 9). An oligonucleotide library was generated from 15 guides for MG29- 1 (SEQ ID NOs: 3731-3745) to investigate their ON- and OFF-target effects. Each library member contained two targets: an ON-target sequence and an OFF-target sequence with mismatches systematically introduced at every position within both the protospacer adjacent motif (PAM) and the spacer region.

[0281] In addition, control oligonucleotides, were generated to simulate four possible editing events: cleavage at the ON-target, cleavage at the OFF-target, cleavage at both targets, and no cleavage. The library construct, reference structure shown in (FIG. 10), contains of the following components ordered left to right: a 20-bp left primer adapter sequence, an 18-bp barcode 1, a 24- 32 bp OFF-target sequence, a 28 bp linker segregating the two target sequences, a 24-32 bp on- target sequence, an 18-bp barcode 2, and an 20-bp right primer adapter sequence.

[0282] Editing efficiency was computed based on two quantification windows: one around the OFF-target sequence and one around the ON-target sequence. Groups with fewer than 100 reads, less than 1% on-target editing, or no editing at either target were discarded from downstream analysis. A final single mismatch tolerance score was calculated for each group as the ratio of OFF- target to ON-target editing. Single mismatch tolerance scores above 0.5 were considered to be highly tolerant, while those below 0.5 were considered to have low tolerance. PAM preference scores were also calculated by converting these mismatch tolerance scores and their respective OFF-target PAMs into information-based representation matrices using logomaker.

[0283] To further visualize the data, heatmaps were created to illustrate the protospacer base preference at each mismatch position for different nucleases. The heatmaps were generated bygrouping the data points by nuclease, mismatch location, and mismatch base, and then calculating the mean OFF-ON ratio and 95% confidence interval for each group. The resulting data was pivoted to create a matrix format suitable for heatmap visualization via seaborn, with color coding to indicate the mismatch tolerance at each base and position.

[0284] Results80#507551Attorney Docket No. 00010.041.1801

[0285] Editing controls were generated to characterize the editing outcomes of a dual-target library. The editing at left or right end was turned “OFF” by modifying the native PAM sequence (“ON”) associated with each on-target sequence. Editing observed across the four types of editing controls, no editing (OFF / OFF), left-end editing (OFF / ON, right-end editing, and both-end editing (ON / ON) were plotted in (FIG. 11) and show editing can be controlled through PAM inactivation and indels correctly assigned to the left or right end target.

[0286] Raw indel percentages at mismatched targets compared to on-target library members were plotted by guide to compare the ranges of editing levels observed and plotted in (FIG. 12). The raw indel percentages were normalized by taking the ratio of OFF: ON target editing, termed mismatch tolerance, observed by position for each guide. A nuclease-specific single mismatch tolerance profile was generated by plotting the mean and 95% confidence intervals of single mismatch tolerance, by position, of all guides for a given nuclease (FIG. 13). FIG. 14 shows the single mismatch tolerance profile observed with 15 MG29-1 guides of 22 nucleotide length targeting human B2M, TIM3, and TRAC loci. MG29-1 demonstrated lowest tolerance for single nucleotide mismatches at PAM proximal sites, increasing in tolerance moving distally from PAM.

[0287] PAM preference scores were calculated by converting mismatch tolerance scores and their respective OFF-target PAMs into information-based representation matrices using logomaker. The PAM preferences for MG29-1 observed with 22 nucleotide spacers are represented as a seqlogo in (FIG. 15) depicting a consensus PAM of TTTN.

[0288] Heatmaps were created to illustrate the protospacer base preference at each mismatched position for different nucleases. Due to the fixed nucleotide nature of certain nuclease PAMs, not all mismatches are represented and are thus represented by white blank squares. For example, if all PAM sequences for a given position are a C, there would be no PAM mismatches to a C (and thus because all protospacers contain a G, only A, C, and T mismatches would be displayed). FIG. 16 depicts the protospacer base preferences for MG29-1.Example 9 - Albumin locus targeting gRNA in primary human hepatocytes

[0289] Methods

[0290] Structural variants were identified after editing with MG29-1 complexed with either ALB- 83b or positive control guide. Structural variants encompass translocations, large deletions, as well as inversions. Edited samples in 5 replicates for each target were used in an assay. Mock edited81#507551Attorney Docket No. 00010.041.1801 samples were included in singlicate for each assay. Circos plots are used to visualize structural variants while Upset plots are used to show the agreement between biological replicates.

[0291] Results

[0292] The ALB MG29-1 assay discovered no reproducible hits within a 200bp breakpoint window of each other. A total of 38 non-reproducible hits unique to the edited sample (vs mock) were identified with a range of 2-14 non-reproducible hits found in each replicate (FIG. 17). The range of reads per million at each discovered breakpoint was 20-655 in the ALB MG29-1 assay. On the other hand, the positive control assay discovered a much elevated number of both reproducible and non-reproducible hits (FIG. 18). Out of the as many as 962 total hits found, 78 reproduced within a 200bp breakpoint window of each other in at least 2 out of 5 biological replicates resulting in 59 reproducible called hits for the positive control assay. For reproducible hits, the number of called hits ranged from 2-47 depending on the number of replicates called hits were present in with the numbers dropping as the number of replicates increased. For non- reproducible hits, the number ranged from 92-235 in each replicate. The range of reads per million at each discovered breakpoint was 20-419244 for the positive control assay. All 5 hits in the positive control assay that reproduced in 4 or 5 biological replicates were hits from an orthogonal cell based OT discovery assay. Additionally, 6 previously validated OT sites were re-discovered by the positive control assay with InDei rates as low as 1 %.Example 10 - Evaluation of nuclease activity on genomic variants

[0293] In order to further characterize editing in cellular context, a variant-aware cell-based assay was developed to query the impact of genetic variation at on- and potential off-target sites. This assay utilizes pooled, lentivirally delivered vector genome containing dual target sequences: the wild-type on-target sequence and one of many nominated off-target sites, their variants, or variation at the on-target site (FIG. 19). This assay is an extension of the assay described in Example 8, where in this case the off-targets are derived from rare genotypes present in public databases (e.g.. GnomAD) rather than solely designed mismatch sequences. These variants would be difficult to source from donor material and thus the integrated library offers a proxy for evaluating the editing propensity.

[0294] Library design

[0295] To query' potential off-target editing due to genetic variants of the 83b guide, an 83b mismatch library was designed and the mismatch assay was performed according to the illustrated82#507551Attorney Docket No. 00010.041.1801 workflow as exemplified in FIG. 19. The 83b variant library (total 4,042 spacers; SEQ ID NOs: 4085-8126) included the on-target sequence, single nucleotide mismatches (n=72), off-target loci hits obtained from a genomic method that identifies and maps cutting sites of Cas9 enzyme (n=476), CRISPRme hits up to edit distance 4 (including variants from gnomAD that modify off- target to be edit distance <4) (n=3,081) and control sequences: editing controls (n=4), and “truncated” spacer length where the spacer is incrementally shortened (n=40) (FIG. 20).

[0296] To determine which linker design gives the lowest barcode recombination rate, an 83b mismatch optimization 1 i brarx with different linker designs was designed to ensure that valid library7members with the correct matching barcodes at the plasmid stage can be identified and retained. The 83b mismatch linker-optimization library (total 1820 members) includes the singletons (72 singletons), all the truncation combinations up to 5bp (363 truncates), and 4 controls replicated 5 times, each with aunique barcode (20 controls) (FIG. 27) (SEQ ID Nos: 8127- 9946). As shown in FIG. 27, the library' has four different linker designs: (1) a constant 28 nt linker, (2) a constant 33 nt linker, (3) a unique 33 nt linker for each oligo, and (4) a unique 38 nt linker for each oligo.

[0297] Assembly of library oligo into lentiviral vectors and QC

[0298] A lyophilized oligo library pool was first resuspended with molecular grade water. The oligo DNA was then subjected to seven cycles of PCR amplification to reach enough yield for cloning into a lentiviral vector. The lentiviral plasmid. pRSG16-U6-sg-UbiC-TagRFP-2A-Puro, was used as a backbone with the molar ratio of insert: vector = 5: 1 at 50 °C for Ihr. The reactions were then subjected to ethanol precipitation and resuspended in molecular grade water. Electrocompetent cells w ere thawed on ice for about 10 minutes to about 20 minutes and 8 pl of DNA is added to 50pl of cells, mixed by swirling the tip in a tube. 29pl of DNA / cell mixture was added to each cuvette. Cells were electroporated with ECI program (1800 Volts, 10 uF, 600 Ohms) in 0. 1 cm cuvette and 2ml of recovery medium at room temperature was immediately added to the cells. Electroporated cells w ere then transferred to a 50ml falcon tube and shaken at 37°C for Ihr to recover. After the recovery , the cells were diluted at 1: 10, 1:30, 1 :90, 1 :270, plated in duplicate, and incubated at 37°C overnight for colony count and library diversity assessment. The rest of the cells were cultured with TB broth and incubated at 30°C overnight for plasmid maxi-prep. Initial plasmid library' diversity' was determined via amplifying library' oligos from 125pg plasmid input followed by NextSeqlOOO sequencing using a lx300nt kit according to the manufacturer’s83#507551Attorney Docket No. 00010.041.1801 instructions. Minimum 5M reads of oligo library' were analyzed for barcode recombination rate as well as library oligo representation.

[0299] Lentiviral production and titering of 83b mismatch library

[0300] Lentiviral packaging procedure was performed as outlined below. Briefly, the lentivirus production cells were cultured and seeded to aim for -60-80% confluency on the day of transfection. For lentivirus production, the transfer vector and packaging vector were combined at a 1 : 1 ratio in serum-free medium along with the transfection reagent. The mixture was gently mixed to ensure homogeneity before being added dropwise to the producer cells. Following transfection, the cells were incubated for 5-6 hours at 37 °C in a 5% CO2 incubator to allow complex formation and uptake. After this initial incubation period, transfection enhancers were added directly to the culture medium to boost viral yield.

[0301] At 48 hours post-transfection, the culture medium containing the lentiviral particles was harvested by centrifugation at 1,300 x g for 15 minutes to remove cells and large debris. The clarified viral supernatant was then passed through a 0.22 pm low-protein-binding filter to further eliminate residual particulates.

[0302] To concentrate the virus, every’ 3 parts of filtered viral supernatant were mixed with 1 part of a concentration reagent (e.g., 40 mL of concentrator for 120 mL of supernatant). The mixture was gently inverted 5-10 times to ensure uniform blending and then incubated at 4 °C for 1-2 hours. After incubation, the samples were centrifuged at 1,500 x g for 45 minutes in a pre-chilled centrifuge. The supernatant was carefully discarded, and the viral pellet was resuspended in 1 / 100th of the original supernatant volume using additive-free medium. The concentrated virus was aliquoted into pre-labeled cryogenic tubes and stored at -80 °C until further use.

[0303] To determine the functional viral titer, K562 cells were transduced with serial dilutions of the viral stock (2x, l x, o.5x, 0.25x, 0.125x, 0.063x, 0.031x, and Ox) in the presence of 8 pg / mL polybrene. At 3 days post-transduction, cells were analyzed for tagRFP expression by flow cytometry, and transducing units per milliliter (TU / mL) yvere estimated based on the percentage of tagRFP-positive cells.

[0304] Generation of 83b mismatch lentiviral library in K562 cell line

[0305] K562 cells are transduced with 0.167x viral dose in a 15cm dish at 1 e6 / mL in 30mL (with 8pg / mL polybrene) and puromycin (2pg / ml) selection started at day3 post transduction with an MOI <0.3. The percent of cells expressing tagRFP was determined on day 3, 6, 8, 10, 13 to ensure minimum lOOOx library coverage before LNP editing.84#507551Attorney Docket No. 00010.041.1801

[0306] LNP editing of 83b mismatch lentiviral library

[0307] For LNP delivery, the engineered K562 cells were seeded at 10 e7 cells at 1 e6 / mL per T75 flask and cultured at 37°C with 5% CCL in IMDM media. The following day, cells were edited with 3x EC90 concentration of the relevant LNP packaged mRNA and gRNA or left untreated and cultured for 3 days. After editing, cells are pelleted for gDNA extraction from which point compatible indexed libraries of amplified target sequences representing the treatment outcome were made. The libraries were then sequenced with 1x300 cycle kit using sequencers targeting a read coverage of at least 10,000 reads per site per sample.

[0308] In-cell library diversity' was determined via amplifying library' oligos from 3.75pg genomic DNA input followed by deep sequencing using a lx300nt kit according to the manufacturer’s instructions. Minimum 10k reads per oligo was required for barcode recombination rate, library oligo representation, and indel analysis.

[0309] Bioformatic analysis

[0310] After NGS sequencing, reads that retained both barcodes were used to identify oligos library members and determine editing efficiencies. For the 83b library’, the analysis method was modified to allow the use of a shorter portion of the barcode for identification; specifically, the 13bp segment distal to the target sites and adjacent to the adapter sequence. This adjustment ensures barcode recovery' in cases where large deletions extend into the barcode region, potentially removing part of the barcode sequence.

[0311] General quality control metrics, including library diversity, recombination rate, and barcode distribution, were calculated and visualized for each sample. Onfy sequencing reads containing valid barcode pairs were retained for further analysis. Reads were grouped by barcode pairs to represent individual library members and aligned to their corresponding reference oligonucleotides using CRISPResso2.

[0312] Editing efficiency was quantified within two defined windows: one encompassing the on- target site and one encompassing the off-target site. The on-target site serves as an internal control for normalizing off-target editing rates. For each library member, a mismatch tolerance score was calculated as the ratio of off- target to on-target editing efficiency. When evaluating off-target oligonucleotides in this library, threshold parameters for hit identification were intentionally lowered to ensure that all potential off-target events were captured.85#507551Attorney Docket No. 00010.041.1801

[0313] Results

[0314] 83b mismatch library barcode identification & variant assessment

[0315] Two untreated and two edited 83b variant library samples (On-target indel% -90%) were subjected to mismatch tolerance analysis (FIG. 21). Decreased detection of barcode 2 was observed in the edited samples, likely due to the super saturating editing at on-target spacer (FIGs. 26A-26B). In order to recover the maximum amount of reads for calculating editing efficiency, a portion of barcode 2 sequence distal to the target cut sites was used. The original barcodes were 18 base pairs in length and it was verified that the shorter segments retained sufficient uniqueness for accurate library' member identification.

[0316] Comparison between on- vs. off-target editing for 83b mm library (FIG. 22) indicates that most oligos with higher off-target editing are singletons (FIGs. 22-24). which contain single mismatches along the spacer. Additional detected hits included the on-target oligonucleotide control (Oligo O (SEQ ID NO: 8126); off / on ratio -1) and Oligo_3884 (SEQ ID NO: 7968) (off / on ratios -0.44 and -0.32) (FIG. 25). Oligo_3884 corresponds to a genomic variant predicted by CRISPRme, a variant-aware in silico off-target prediction tool. The oligo_3884 loci contains sequence variation from GnomAD relative to hg38, resulting in an edit distance of four, compared to five for the hg38 reference sequence. Consistent with this difference, Oligo_3884 exhibited 32- 39% off-target editing, whereas its hg38 reference counterpart (Oligo_97, SEQ ID NO: 4181) showed no detectable off-target activity (Table 2).

[0317] Table 2: 83b mismatch assay results for the CRISPRme variation hit Oligo_3884 and its corresponding hg38 pair, oligo_97#507551Attorney Docket No. 00010.041.1801

[0318] 83b mismatch optimization library

[0319] To further optimize the assay, a new library (SEQ ID Nos: 8127-9946) was designed with longer and unique linker sequences and longer barcodes to reduce barcode recombination while retaining intact barcodes upon editing. The conditions are described above. A longer linker and longer barcode could accommodate the wide indel window of type V nuclease / MG29-l, providing a more accurate editing readout at off-target spacers. In addition, the variable linker design helped reduce barcode recombination due to reduction of constant regions in the oligo structure.

[0320] The four oligo designs are illustrated in FIG. 27 including the original version (top design in FIG. 27). The barcode recombination rate of the plasmid library was evaluated and as shown in FIG. 28, all conditions had -70% of both barcodes detected and the presence of unique linkers reduced PCR recombination (z.e., differing barcodes detected group).Table 3 - Listing of PAMs referred to herein not included in the sequence listing87#507551Attorney Docket No. 00010.041.180188#507551Attorney Docket No. 00010.041.180189#507551Attorney Docket No. 00010.041.180190#507551Attorney Docket No. 00010.041.180191#507551Attorney Docket No. 00010.041.1801

[0321] While preferred embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the disclosure be limited by the specific examples provided within the specification. While the disclosure has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the disclosure. Furthermore, it shall be understood that all aspects of the disclosure are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed in practicing the disclosure. It is therefore contemplated that the disclosure shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the disclosure and that methods and structures within the scope of these claims and their equivalents be covered thereby.#507551

Claims

Attorney DocketNo. 00010.041.1801CLAIMSWHAT IS CLAIMED IS:

1. An engineered nuclease system, comprising: a) an engineered endonuclease having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence within a valosin-containing protein (VCP) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3721-3723, and 3725-3727.

2. The engineered nuclease system of claim 1, wherein the engineered endonuclease has at least 80% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

3. The engineered nuclease system of claim 1, wherein the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

4. The engineered nuclease system of any one of claims 1-3, wherein the engineered nuclease is encoded by a sequence having at least 70% sequence identity7to SEQ ID NO: 3729.

5. The engineered nuclease system of any one of claims 1-3, wherein the engineered nuclease is encoded by a sequence having at least 90% sequence identity to SEQ ID NO: 3729.

6. The engineered nuclease system of any one of claims 1-3, wherein the engineered nuclease is encoded by a sequence having 100% sequence identity to SEQ ID NO: 3729.

7. The engineered nuclease system of any one of claims 1-6, wherein the engineered guide polynucleotide is a single guide nucleic acid.

8. The engineered nuclease system of any one of claims 1 -6, wherein the engineered guide polynucleotide is a dual guide nucleic acid.

9. The engineered nuclease system of any one of claims 1-6, wherein the engineered guide polynucleotide is RNA.

10. The engineered nuclease system of any one of claims 1-9, wherein the engineered endonuclease binds non-covalently to the engineered guide polynucleotide.

11. The engineered nuclease system of any one of claims 1 -9, wherein the endonuclease is covalently linked to the engineered guide polynucleotide.

12. The engineered nuclease system of any one of claims 1-9. wherein the endonuclease is fused to the engineered guide polynucleotide.

13. An engineered nuclease system, comprising:Attorney DocketNo. 00010.041.1801 a) an engineered endonuclease comprising a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence within a Beta-2 -Microglobulin (B2M) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3731-3737. 3746-3752, and 3781-3856.

14. The engineered nuclease system of claim 13, wherein the engineered endonuclease has at least 80% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

15. The engineered nuclease system of claim 13, wherein the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

16. The engineered nuclease system of any one of claims 13-15, wherein the engineered nuclease is encoded by a sequence having at least 70% sequence identity to SEQ ID NO: 3729.

17. The engineered nuclease system of any one of claims 13-15, wherein the engineered nuclease is encoded by a sequence having at least 90% sequence identity to SEQ ID NO: 3729.

18. The engineered nuclease system of any one of claims 13-15, wherein the engineered nuclease is encoded by a sequence having 100% sequence identity to SEQ ID NO: 3729.

19. The engineered nuclease system of any one of claims 13-18, wherein the engineered guide polynucleotide is a single guide nucleic acid.

20. The engineered nuclease system of any one of claims 13-18, wherein the engineered guide polynucleotide is a dual guide nucleic acid.

21. The engineered nuclease system of any one of claims 13-18, wherein the engineered guide polynucleotide is RNA.

22. The engineered nuclease system of any one of claims 13-21, wherein the engineered endonuclease binds non-covalently to the engineered guide polynucleotide.

23. The engineered nuclease system of any one of claims 13-21, wherein the endonuclease is covalently linked to the engineered guide polynucleotide.

24. The engineered nuclease system of any one of claims 13-21, wherein the endonuclease is fused to the engineered guide polynucleotide.

25. An engineered nuclease system, comprising: a) an engineered endonuclease comprising a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion ofAttorney DocketNo. 00010.041.1801 a target nucleic acid sequence within a T cell immunoglobulin and mucin domain 3 (TIM-3) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3738- 3744 and 3753- 3759.

26. The engineered nuclease system of claim 25, wherein the engineered endonuclease has at least 80% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

27. The engineered nuclease system of claim 25, wherein the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

28. The engineered nuclease system of any one of claims 25-27, wherein the engineered nuclease is encoded by a sequence having at least 70% sequence identity to SEQ ID NO: 3729.

29. The engineered nuclease system of any one of claims 25-27, wherein the engineered nuclease is encoded by a sequence having at least 90% sequence identity to SEQ ID NO: 3729.

30. The engineered nuclease system of any one of claims 25-27, wherein the engineered nuclease is encoded by a sequence having 100% sequence identity to SEQ ID NO: 3729.

31. The engineered nuclease system of any one of claims 25-30, wherein the engineered guide polynucleotide is a single guide nucleic acid.

32. The engineered nuclease system of any one of claims 25-30, wherein the engineered guide polynucleotide is a dual guide nucleic acid.

33. The engineered nuclease system of any one of claims 25-30, wherein the engineered guide polynucleotide is RNA.

34. The engineered nuclease system of any one of claims 25-33, wherein the engineered endonuclease binds non-covalently to the engineered guide polynucleotide.

35. The engineered nuclease system of any one of claims 25-33, wherein the endonuclease is covalently linked to the engineered guide polynucleotide.

36. The engineered nuclease system of any one of claims 25-33, wherein the endonuclease is fused to the engineered guide polynucleotide.

37. An engineered nuclease system, comprising: a) an engineered endonuclease comprising a sequence having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684; and b) an engineered guide polynucleotide configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to at least a portion of a target nucleic acid sequence within a T Cell Receptor Alpha Constant (TRAC) gene, the engineered guide polynucleotide comprising a sequence having at least 90% sequence identity to SEQ ID NO: 3745 or SEQ ID NO: 3760.Attorney DocketNo. 00010.041.180138. The engineered nuclease system of claim 37, wherein the engineered endonuclease has at least 80% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

39. The engineered nuclease system of claim 37, wherein the engineered endonuclease has 100% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684.

40. The engineered nuclease system of any one of claims 37-39, wherein the engineered nuclease is encoded by a sequence having at least 70% sequence identity to SEQ ID NO: 3729.

41. The engineered nuclease system of any one of claims 37-39, wherein the engineered nuclease is encoded by a sequence having at least 90% sequence identity to SEQ ID NO: 3729.

42. The engineered nuclease system of any one of claims 37-39, wherein the engineered nuclease is encoded by a sequence having 100% sequence identity to SEQ ID NO: 3729.

43. The engineered nuclease system of any one of claims 37-42, wherein the engineered guide polynucleotide is a single guide nucleic acid.

44. The engineered nuclease system of any one of claims 37-42, wherein the engineered guide polynucleotide is a dual guide nucleic acid.

45. The engineered nuclease system of any one of claims 37-42, wherein the engineered guide polynucleotide is RNA.

46. The engineered nuclease system of any one of claims 37-45, wherein the engineered endonuclease binds non-covalently to the engineered guide polynucleotide.

47. The engineered nuclease system of any one of claims 37-45, wherein the endonuclease is covalently linked to the engineered guide polynucleotide.

48. The engineered nuclease system of any one of claims 37-45, wherein the endonuclease is fused to the engineered guide polynucleotide.

49. A method of modifying a target nucleic acid sequence in a mammalian cell comprising contacting the mammalian cell using the engineered nuclease system of any one of claims 1- 48.

50. The method of claim 49, wherein modifying the target nucleic acid sequence comprises binding, nicking, or cleaving the target nucleic acid sequence.

51. The method of any one of claims 49-50, wherein the target nucleic acid sequence comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA.

52. The method of any one of claims 49-51, wherein the modification is in vitro.

53. The method of any one of claims 49-51, wherein the modification is in vivo.

54. The method of any one of claims 49-51, wherein the modification is ex vivo.Attorney DocketNo. 00010.041.180155. The method of any one of claims 49-54. further comprising selecting cells comprising the modification.

56. A cell comprising the engineered nuclease system of any one of claims 1-48.

57. The cell of claim 56, wherein the cell is a eukaryotic cell.

58. The cell of claim 56, wherein the cell is a mammalian cell.

59. The cell of claim 56. wherein the cell is an immortalized cell.

60. The cell of claim 56, wherein the cell is an insect cell.

61. The cell of claim 56, wherein the cell is a yeast cell.

62. The cell of claim 56, wherein the cell is a plant cell.

63. The cell of claim 56, wherein the cell is a fungal cell.

64. The cell of claim 56, wherein the cell is a prokaryotic cell.

65. The cell of claim 56, wherein the cell is an A549, HEK-293, HEK-293T, BHK, CHO,HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cell, HT1080, HepG2, Huh7, K562, primary cell, or a derivative thereof.

66. The cell of claim 56, wherein the cell is an engineered cell.

67. The cell of claim 56, wherein the cell is a stable cell.

68. The cell of claim 56, wherein the cell is a T cell.

69. The cell of claim 56, wherein the cell is a hematopoietic cell.

70. A lipid nanoparticle comprising:(a) the engineered nuclease system of any one of claims 1-48;(b) a cationic lipid;(c) a sterol;(d) a neutral lipid; and(e) a PEG-modified lipid.

71. The lipid nanoparticle of claim 70, wherein the cationic lipid comprises C 12-200, the sterol comprises cholesterol, the neutral lipid comprises DOPE, or the PEG-modified lipid comprises DMG-PEG2000.

72. The lipid nanoparticle of claim 70, wherein the cationic lipid comprises 98N12-5 (TETA5-LAP), DLin DMA, DLin-K-DMA (2,2-Dilinoleyl-4-dimethylaminomethyl-[l,3]- dioxolane), DLin-KC2-DMA, DLin-MC3-DMA, or C 12-200.

73. A method for identifying off-target editing efficiency relative to a target DNA sequence, comprising:Attorney DocketNo. 00010.041.1801 a) contacting a cell with a library of polynucleotides and an engineered nuclease system, the engineered nuclease system comprising: an engineered endonuclease having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease, wherein each polynucleotide of the library comprises:(i) an on-target nucleic acid sequence,(ii) an off-target nucleic acid sequence containing a single nucleotide mismatch at a defined position within a protospacer adjacent motif (PAM) or a spacer relative to the on-target sequence,(iii) a first barcode associated with the off-target nucleic acid sequence and a second barcode associated with the on-target nucleic acid sequence,(iv) a linker sequence separating the off-target nucleic acid sequence and the on-target nucleic acid sequence, and(v) adapter primers flanking each of the on-target and off-target nucleic acid sequences; b) isolating genomic DNA from the cell; c) amplifying the target DNA sequence using the adapter primers; and d) sequencing the first and second barcodes and aligning the sequence reads assigned to each polynucleotide of the library to a target DNA sequence and quantifying an off-target indel frequency and an on-target indel frequency, and computing a mismatch tolerance value, thereby identifying off-target editing efficiency relative to the target DNA sequence for the engineered nuclease-engineered guide polynucleotide complex.

74. A method of determining specificity of a nuclease, the method comprising: a) contacting a cell with a library of polynucleotides and an engineered nuclease system, the engineered nuclease system comprising: an engineered endonuclease having at least 70% sequence identity to any one of SEQ ID NOs: 215-225 and 3471-3684 and an engineered guide polynucleotide configured to form a complex with the endonuclease, wherein each polynucleotide of the library comprises:(i) an on-target nucleic acid sequence,(ii) an off-target nucleic acid sequence containing a single nucleotide mismatch at a defined position within a protospacer adjacent motif (PAM) or a spacer relative to the on-target sequence,Attorney DocketNo. 00010.041.1801(iii) a first barcode associated with the off-target nucleic acid sequence and a second barcode associated with the on-target nucleic acid sequence,(iv) a linker sequence separating the off-target nucleic acid sequence and the on-target nucleic acid sequence, and(v) adapter primers flanking each of the on-target and off-target nucleic acid sequences; b) isolating genomic DNA from the cell; c) amplifying the target DNA sequence using the adapter primers; and d) sequencing the first and second barcodes and aligning the sequence reads assigned to each polynucleotide of the library to a target DNA sequence and quantifying an off-target indel frequency and an on-target indel frequency, and computing a mismatch tolerance value, thereby determining specificity of a nuclease.

75. The method of claim 73 or 74, wherein a mismatch tolerance value of less than 0.5 indicates higher specificity for the engineered nuclease system and a mismatch tolerance value of more than 0.5 indicates lower specificity for the engineered nuclease system.