Class II and V CRISPR systems

JP2024522086A5Pending Publication Date: 2025-06-03METAGENOMI INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023572062
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-14
Filing Date
2022-06-01
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing CRISPR/Cas systems, particularly Cas9, exhibit high immunogenicity and limited specificity when used in human subjects, necessitating improved endonucleases with enhanced target recognition and reduced immune response.

Method used

Development of a Class II, Type V Cas endonuclease derived from uncultured microorganisms, combined with engineered guide RNAs, to form a complex that targets specific nucleic acid sequences with high specificity and reduced immunogenicity.

Benefits of technology

The new endonuclease system achieves precise and efficient gene editing with reduced antibody immunogenicity, enabling targeted nucleic acid cleavage in human cells.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Described herein are methods, compositions, and systems derived from uncultured microorganisms that are useful for gene editing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application is a continuation of U.S. Provisional Patent Application Nos. 63 / 196,127, filed June 2, 2021, 63 / 233,653, filed August 16, 2021, 63 / 261,436, filed September 21, 2021, 63 / 262,169, filed October 6, 2021, 63 / 280,026, filed November 16, 2021, and 63 / 280,026, filed November 16, 2021. This application claims the benefit of PCT Application No. 63 / 299,664, filed January 14, 2022, No. 63 / 308,766, filed February 10, 2022, No. 63 / 323,014, filed March 23, 2022, and No. 63 / 331,076, filed April 14, 2022, each of which is incorporated by reference herein in its entirety. This application is related to PCT Application No. PCT / US21 / 21259, which is incorporated by reference herein in its entirety. [Background technology]

[0002] Cas enzymes, along with their associated clustered regularly interspaced short palindromic repeats (CRISPR)-guided ribonucleic acid (RNA), appear to be widespread components of prokaryotic immune systems (approximately 45% of bacteria and 84% of archaea), helping to protect these microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, through CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. While CRISPR DNA elements were observed as early as 1987, the programmable endonuclease cleavage capabilities of CRISPR / Cas complexes have only recently been recognized, leading to the use of recombinant CRISPR / Cas systems in a variety of DNA manipulation and gene editing applications.

[0003] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated by reference herein in its entirety. The ASCII copy, created on March 5, 2021, is named 55921-728_601_SL.txt, and is 23,886,645 bytes in size. Summary of the Invention

[0004] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC domain, wherein the endonuclease is derived from an uncultured microorganism, and the endonuclease is a Cas12a endonuclease; and (b) an engineered guide RNA comprising a spacer sequence configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence. In some embodiments, the Cas12a endonuclease comprises the sequence GWxxxK. In some embodiments, the engineered guide RNA comprises the sequence UCUAC[N 3-5 ]GUAGAU(N4). In some embodiments, the engineered guide RNA comprises CCUGC[N4]GCAGG(N 3-4In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof; and (b) an engineered guide RNA comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease comprises a RuvCI, II, or III domain. In some embodiments, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the RuvCI, II, or III domain of any one of SEQ ID NOs: 1-3470 or variants thereof. In some embodiments, the RuvCI domain comprises a D catalytic residue. In some embodiments, the RuvCII domain comprises an E catalytic residue. In some embodiments, the RuvCIII domain comprises a D catalytic residue. In some embodiments, the RuvC domain does not have nuclease activity.In some embodiments, the endonuclease further comprises a WED II domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the WED II domain of any one of SEQ ID NOs: 1-3470 or variants thereof. In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) comprising any one of SEQ ID NOs: 3862-3913, wherein the endonuclease is a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA comprising a spacer sequence configured to form a complex with the endonuclease and hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease further comprises a zinc finger-like domain. In some embodiments, the guide RNA is selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-355MG11, MG13, MG19, MG20, MG26, MG28, MG29, MG30, MG31, MG32, MG37, MG39, MG40, MG41, MG42, MG43, MG44, MG45, MG46, MG47, MG48, MG49, MG50, MG51, MG52, MG53, MG54, MG55, MG56, MG57, MG58, MG59, MG60, MG61, MG62, MG63, MG64, MG65, MG66, MG67, MG68, MG69, MG70, MG71, MG72, MG73, MG74, MG75, MG76, MG77, MG78, MG79, MG80, MG81, MG82, MG83, MG84, MG85, MG86, MG87, MG88, MG89, MG90, MG91, MG92, MG93, MG94, MG95, MG96, MG97, MG98, MG99, MG10, MG100, MG101, MG102, MG103, MG104, MG105, MG106, MG107, MG108, MG110, MG111, MG112, MG113, MG114, MG115, MG116, MG117, MG11 9, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, and 3851-3857.In some aspects, the present disclosure provides a method for the preparation of a nucleotide sequence selected from the group consisting of (a) nucleotides selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695- The present invention provides an engineered nuclease system comprising: (a) an engineered guide RNA comprising a sequence having at least 80% sequence identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 3696, 3729-3730, 3734-3735, or 3851-3857; and (b) a Class 2, V-type Cas endonuclease configured to bind to the engineered guide RNA. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence comprising any one of SEQ ID NOs: 3863-3913. In some embodiments, the guide RNA comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the guide RNA is 30-250 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 3938-3953. In some embodiments, the endonuclease comprises at least one of the following mutations when the sequence of the endonuclease is optimally aligned with SEQ ID NO: 215: S168R, E172R, N577R, or Y170R. In some embodiments, the endonuclease comprises mutations S168R and E172R when the sequence of the endonuclease is optimally aligned with SEQ ID NO: 215. In some embodiments, the endonuclease comprises mutations N577R or Y170R when the sequence of the endonuclease is optimally aligned with SEQ ID NO: 215. In some embodiments, the endonuclease comprises mutation S168R when the sequence of the endonuclease is optimally aligned with SEQ ID NO: 215.In some embodiments, the endonuclease does not contain a mutation at E172, N577, or Y170. In some embodiments, the engineered nuclease system further comprises a single-stranded or double-stranded DNA repair template comprising, from 5' to 3', a first homologous arm comprising a sequence of at least 20 nucleotides 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the first or second homologous arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the first and second homologous arms are homologous to a genomic sequence of a prokaryote, bacterium, fungus, or eukaryote. In some embodiments, the single-stranded or double-stranded DNA repair template comprises a transgene donor. In some embodiments, the engineered nuclease system further comprises a DNA repair template comprising a double-stranded DNA segment adjacent to one or two single-stranded DNA segments. In some embodiments, the single-stranded DNA segment is conjugated to the 5' end of the double-stranded DNA segment. In some embodiments, the single-stranded DNA segment is conjugated to the 3' end of the double-stranded DNA segment. In some embodiments, the single-stranded DNA segment has a length of 4 to 10 nucleotide bases. In some embodiments, the single-stranded DNA segment has a nucleotide sequence that is complementary to a sequence within a spacer sequence. In some embodiments, the double-stranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein-coding sequence, an miRNA-coding sequence, an RNA-coding sequence, or a transgene. In some embodiments, the double-stranded DNA sequence is adjacent to a nuclease cleavage site. In some embodiments, the nuclease cleavage site comprises a spacer and a PAM sequence. In some embodiments, the system comprises Mg. 2+In some embodiments, the guide RNA comprises a hairpin comprising at least 8, at least 10, or at least 12 base pair ribonucleotides. In some embodiments, the hairpin comprises 10 base pair ribonucleotides. In some embodiments, (a) the endonuclease comprises a sequence at least 75%, 80%, or 90% identical to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof, and (b) the guide RNA structure comprises a sequence at least 80% or 90% identical to a non-degenerate nucleotide of any one of SEQ ID NOs: 3608-3609, 3853, or 3851-3857. In some embodiments, the endonuclease is configured to bind to a PAM comprising any one of SEQ ID NOs: 3863-3913. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 3871. In some embodiments, sequence identity is determined by the BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW algorithm using Smith-Waterman homology search algorithm parameters. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.

[0005] In some aspects, the present disclosure provides an engineered guide RNA comprising: (a) a DNA-targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide; wherein the engineered guide ribonucleic acid polynucleotide is capable of forming a complex with an endonuclease having at least 75% sequence identity to any one of SEQ ID NOS: 1-3470 and targeting the complex to a target sequence in the target DNA molecule. In some embodiments, the DNA-targeting segment is positioned 3' to both of the two complementary stretches of nucleotides. In some embodiments, the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to non-degenerate nucleotides of SEQ ID NOS: 3608-3609. In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 5, at least 8, at least 10, or at least 12 ribonucleotides.

[0006] In some aspects, the present disclosure provides deoxyribonucleic acid polynucleotides that encode the engineered guide ribonucleic acid polynucleotides described herein.

[0007] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a Class 2, Type V Cas endonuclease, wherein the endonuclease is derived from an uncultured microorganism, and wherein the organism is not an uncultured organism. In some embodiments, the endonuclease comprises a variant having at least 70% or at least 80% sequence identity to any one of SEQ ID NOs: 1-3470. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N- or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 3938-3953. In some embodiments, the NLS comprises SEQ ID NO: 3939. In some embodiments, the NLS is proximal to the N-terminus of the endonuclease. In some embodiments, the NLS comprises SEQ ID NO: 3938. In some embodiments, the NLS is proximal to the C-terminus of the endonuclease. In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human.

[0008] In some aspects, the disclosure provides engineered vectors comprising a nucleic acid sequence encoding a Class 2, Type V Cas endonuclease, or a Cas12a endonuclease, wherein the endonuclease is derived from an uncultured microorganism.

[0009] In some aspects, the present disclosure provides engineered vectors comprising the nucleic acids described herein.

[0010] In some aspects, the present disclosure provides engineered vectors comprising the deoxyribonucleic acid polynucleotides described herein. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, a lentivirus, or an adenovirus.

[0011] In some aspects, the present disclosure provides a cell comprising a vector described herein.

[0012] In some aspects, the disclosure provides a method of producing an endonuclease, comprising culturing any of the host cells described herein.

[0013] In some aspects, the disclosure provides methods of binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with a Class 2, Type V Cas endonuclease complexed with an endonuclease and an engineered guide RNA configured to bind to the double-stranded deoxyribonucleic acid polynucleotide; (b) the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM); and (c) the PAM comprises a sequence comprising any one of SEQ ID NOs: 3863-3913. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of the engineered guide RNA and a second strand comprising a PAM. In some embodiments, the PAM is immediately adjacent to the 5' end of the sequence complementary to a sequence of the engineered guide RNA. In some embodiments, the PAM comprises SEQ ID NO: 3871. In some embodiments, the Class 2, Type V Cas endonuclease is derived from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the method comprises delivering an engineered nuclease system described herein to a target nucleic acid locus, wherein the endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure, and the complex is configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is within a cell.In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a human cell, or a primary cell. In some embodiments, the cell is a primary cell. In some embodiments, the primary cell is a T cell. In some embodiments, the primary cell is a hematopoietic stem cell (HSC). In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding the endonuclease is operably linked. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA comprising an open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering deoxyribonucleic acid (DNA) encoding an engineered guide RNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease induces a single-stranded or double-stranded break at or proximal to the target locus. In some embodiments, the endonuclease induces a staggered single-stranded break within or 3' of the target locus.

[0014] In some aspects, the disclosure provides a method of editing a TRAC locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the TRAC locus, and the engineered guide RNA comprises a targeting sequence having at least 85% identity to at least 18 contiguous nucleotides of any one of SEQ ID NOs: 4316-4369. In some embodiments, the RNA-guided endonuclease is a Cas endonuclease. In some embodiments, the Cas endonuclease is a Class 2, Type V Cas endonuclease. In some embodiments, the Class 2, Type V Cas endonuclease comprises a RuvC domain comprising a RuvCI subdomain, a RuvCII subdomain, and a RuvCIII subdomain. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof. In some embodiments, the engineered guide RNA further comprises a sequence having at least 80% sequence identity to at least 19 non-degenerate nucleotides of any one of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, and 3851-3857. In some embodiments, the endonuclease comprises a sequence that is at least 75%, 80%, or 90% identical to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof.In some embodiments, the guide RNA structure comprises a sequence that is at least 80% or at least 90% identical to at least 19 non-degenerate nucleotides of any one of SEQ ID NOs: 3608-3609, 3853, or 3851-3857. In some embodiments, the method further comprises contacting or introducing into the cell a donor nucleic acid comprising a cargo sequence flanked on the 3' or 5' end by a sequence having at least 80% identity to any one of SEQ ID NOs: 4424 or 4425. In some embodiments, the cell is a peripheral blood mononuclear cell (PBMC). In some embodiments, the cell is a T cell or a precursor thereof, or a hematopoietic stem cell (HSC). In some embodiments, the cargo sequence comprises a sequence encoding a T cell receptor polypeptide, a CAR-T polypeptide, or a fragment or derivative thereof. In some embodiments, the engineered guide RNA comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 4370-4423. In some embodiments, the engineered guide RNA comprises the nucleotide sequence of sgRNA1-54 from Table 5A, including the corresponding chemical modifications listed in Table 5A. In some embodiments, the engineered guide RNA comprises a targeting sequence having at least 80% sequence identity to any one of SEQ ID NOs: 4334, 4350, or 4324. In some embodiments, the engineered guide RNA comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 4388, 4404, or 4378. In some embodiments, the engineered guide RNA comprises the nucleotide sequence of sgRNA9, 35, or 19 from Table 5A.

[0015] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an RNA-guided endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a target nucleic acid sequence, and wherein the engineered guide RNA has one of the following modifications: (i) a 2'-O methyl or 2'-fluoro base modification of at least one nucleotide within the first four bases of the 5' end of the engineered guide RNA or within the last four bases of the 3' end of the engineered guide RNA; (ii) a 2'-O methyl or 2'-fluoro base modification of at least one nucleotide within the first four bases of the 5' end of the engineered guide RNA or within the last four bases of the 3' end of the engineered guide RNA; The engineered guide RNA may comprise at least one of the following: (i) a thiophosphate (PS) linkage between at least two of the first five bases at the 5' end of the engineered guide RNA or a thiophosphate linkage between at least two of the last five bases at the 3' end of the engineered guide RNA; (ii) a thiophosphate linkage within the 3' or 5' stem of the engineered guide RNA; (iii) a thiophosphate linkage within the 3' or 5' stem of the engineered guide RNA; (iv) a 2'-O methyl or 2' base modification within the 3' or 5' stem of the engineered guide RNA; (v) a 2'-fluoro base modification of at least seven bases in the spacer region of the engineered guide RNA; and (vi) a thiophosphate linkage within the loop region of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises a 2'-O methyl or 2'-fluoro base modification of at least one nucleotide within the first five bases at the 5' end of the engineered guide RNA or the last five bases at the 3' end of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises a 2'-O methyl or 2'-fluoro base modification at the 5' end of the engineered guide RNA or the 3' end of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises a thiophosphate (PS) linkage between at least two of the first five bases at the 5' end of the engineered guide RNA or a thiophosphate linkage between at least two of the last five bases at the 3' end of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises a thiophosphate linkage within the 3' stem or the 5' stem of the engineered guide RNA.In some embodiments, the engineered guide RNA comprises 2'-O methyl base modifications within the 3' stem or 5' stem of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises 2'-fluoro base modifications in at least 7 bases of the spacer region of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises thiophosphate linkages within the loop region of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises at least three 2'-O methyl or 2'-fluoro bases at the 5' end of the engineered guide RNA, two thiophosphate linkages within the first three bases of the 5' end of the engineered guide RNA, at least four 2'-O methyl or 2'-fluoro bases at the 4' end of the engineered guide RNA, and three thiophosphate linkages within the last three bases of the 3' end of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises at least two 2'-O-methyl bases and at least two thiophosphate linkages at the 5' end of the engineered guide RNA, and at least one 2'-O-methyl base and at least one thiophosphate linkage at the 3' end of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises at least one 2'-O-methyl base in both the 3' stem or the 5' stem region of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises at least one to at least 14 2'-fluoro bases in the spacer region, excluding the seed region, of the engineered guide RNA. In some embodiments, the engineered guide RNA comprises at least one 2'-O-methyl base in the 5' stem region of the engineered guide RNA, and at least one to at least 14 2'-fluoro bases in the spacer region, excluding the seed region, of the guide RNA. In some embodiments, the guide RNA comprises a spacer sequence that targets the VEGF-A gene. In some embodiments, the guide RNA comprises a spacer sequence having at least 80% identity to SEQ ID NO: 3985. In some embodiments, the guide RNA comprises nucleotides of guide RNAs 1-7 from Table 7, comprising the chemical modifications listed in Table 7.In some embodiments, the RNA-guided endonuclease is a Cas endonuclease. In some embodiments, the Cas endonuclease is a Class 2, Type V Cas endonuclease. In some embodiments, the Class 2, Type V Cas endonuclease comprises a RuvC domain comprising a RuvCI subdomain, a RuvCII subdomain, and a RuvCIII subdomain. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some embodiments, the engineered guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, and 3851-3857. In some embodiments, the engineered guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 3608-3609, 3853, or 3851-3857.

[0016] In some aspects, the disclosure provides a host cell comprising an open reading frame encoding a heterologous endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof. In some embodiments, the endonuclease has at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some embodiments, the host cell is an E. coli cell or a mammalian cell. In some embodiments, the host cell is an E. coli cell, wherein the E. coli cell is a λDE3 lysogen, or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the open reading frame is selected from the group consisting of a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, a BADIn some embodiments, the open reading frame is operably linked to a promoter, a strong leftward promoter from phage lambda (pL promoter), or any combination thereof. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to the sequence encoding the endonuclease. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose-binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the endonuclease via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the open reading frame is codon-optimized for expression in a host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of a host cell.

[0017] In some aspects, the disclosure provides a culture comprising any of the host cells described herein in a compatible liquid medium.

[0018] In some aspects, the disclosure provides methods for producing an endonuclease, comprising culturing any of the host cells described herein in a suitable growth medium. In some embodiments, the method further comprises inducing expression of the endonuclease. In some embodiments, the induction of nuclease expression is by adding an additional chemical agent or increasing the amount of a nutrient, or by increasing or decreasing the temperature. In some embodiments, the increased amount of the additional chemical agent or nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises isolating the endonuclease. In some embodiments, the isolating comprises subjecting the protein extract to IMAC, ion exchange chromatography, anion exchange chromatography, or cation exchange chromatography. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to the sequence encoding the endonuclease. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the endonuclease via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the method further comprises cleaving the affinity tag by contacting the endonuclease with a protease corresponding to the protease cleavage site. In some embodiments, the affinity tag is an IMAC affinity tag. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from the composition comprising the endonuclease.

[0019] In some aspects, the disclosure provides a system comprising: (a) a Class 2, VA-type Cas endonuclease configured to bind to a 3- or 4-nucleotide PAM sequence, wherein the endonuclease has increased cleavage activity compared to sMbCas12a; and (b) an engineered guide RNA configured to form a complex with the Class 2, VA-type Cas endonuclease and comprising a spacer sequence configured to hybridize to a target nucleic acid comprising the target nucleic acid sequence. In some embodiments, the cleavage activity is measured in vitro by introducing the endonuclease along with a compatible guide RNA into a cell containing the target nucleic acid and detecting cleavage of the target nucleic acid sequence in the cell. In some embodiments, the Class 2, VA-type Cas endonuclease comprises a sequence having at least 75% identity to 215-225 or a variant thereof. In some embodiments, the engineered guide RNA comprises a sequence having at least 80% identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the target nucleic acid further comprises a YYN PAM sequence proximal to the target nucleic acid sequence. In some embodiments, the Class 2, Type VA Cas endonuclease has at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or 200% or more increased activity compared to sMbCas12a.

[0020] In some aspects, the disclosure provides a system comprising: (a) a Class 2, Type V-A' Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA comprises a sequence having at least 80% identity to about 19 to about 25 or about 19 to about 31 contiguous nucleotides of a naturally occurring effector repeat sequence of the Class 2, Type V Cas endonuclease. In some embodiments, the naturally occurring effector repeat sequence is any one of SEQ ID NOs: 3560-3572. In some embodiments, the Class 2, Type V-A' Cas endonuclease has at least 75% identity to SEQ ID NO: 126.

[0021] In some aspects, the disclosure provides a system comprising: (a) a Class 2, VL-type Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA comprises a sequence having at least 80% identity to about 19 to about 25 or about 19 to about 31 contiguous nucleotides of a naturally occurring effector repeat sequence of the Class 2, V-type Cas endonuclease. In some embodiments, the Class 2, VL-type endonuclease has at least 75% sequence identity to any one of SEQ ID NOs: 793-1163.

[0022] In some aspects, the disclosure provides a method of disrupting a VEGF-A locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the VEGF-A locus, and wherein the engineered guide RNA comprises a targeting sequence having at least 80% identity to SEQ ID NO: 3985, or wherein the engineered guide RNA comprises the nucleotide sequence of any one of guide RNAs 1-7 from Table 7. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some embodiments, the engineered guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, and 3851-3857. In some embodiments, the engineered guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 3608-3609, 3853, or 3851-3857.

[0023] In some aspects, the disclosure provides a method for disrupting a genetic locus in a cell, the method comprising contacting the cell with a composition comprising: (a) a Class 2, type V Cas endonuclease having at least 75% identity to any one of SEQ ID NOS: 215-225 or a variant thereof; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the genetic locus, and the Class 2, type V Cas endonuclease has cleavage activity at least equivalent to that of spCas9 in the cell. In some embodiments, the cleavage activity is measured in vitro by introducing the endonuclease along with the compatible guide RNA into a cell containing a target nucleic acid and detecting cleavage of the target nucleic acid sequence in the cell. In some embodiments, the composition comprises 20 pmoles or less of the Class 2, type V Cas endonuclease. In some embodiments, the composition comprises 1 pmol or less of the Class 2, type V Cas endonuclease.

[0024] In some aspects, the disclosure provides a method of disrupting the CD38 locus in a cell, the method comprising introducing into the cell (a) a Class 2, V-type Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the CD38 locus, wherein the engineered guide RNA hybridizes to any one of SEQ ID NOs: 4466-4503 and 5686 by at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at Guide RNAs configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides with about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity include a nucleotide sequence with at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4428-4465 and 5685. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease with at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, and 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 4466, 4467, 4468, 4479, 4484, 4490, 4492, 4493, 4495, 4498. In some embodiments, the engineered guide RNA comprises a nucleotide sequence having at least 80% identity to any one of SEQ ID NOs: 4428, 4429, 4430, 4436, 4441, 4446, 4452, 4454, 4455, 4460, or 4461. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0025] In some aspects, the disclosure provides a method of disrupting a TIGIT locus in a cell, the method comprising introducing into the cell (a) a Class 2, V-type Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the TIGIT locus, wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, or at least about 94% of any one of SEQ ID NOs: 4521-4537. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides with at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4504-4520. The guide RNA comprises a nucleotide sequence with at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4504-4520. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease with at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, and 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 4521, 4527, 4528, 4535, or 4536. In some embodiments, the engineered guide RNA comprises a nucleotide sequence having at least 80% identity to any one of SEQ ID NOs: 4504, 4510, 4511, 4518, or 4519. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0026] In some aspects, the disclosure provides a method of disrupting the AAVS1 locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the AAVS1 locus, wherein the engineered guide RNA hybridizes to any one of SEQ ID NOs: 4569-4599 by at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% of any one of SEQ ID NOs: 4569-4599. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4538-4568. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNA is selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734 In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3609, -3735, 3851-3857, and 6033 to 6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides that have at least 80% identity to any one of SEQ ID NOs: 4574, 4577, 4578, 4579, 4582, 4584, 4585, 4586, 4587, 4589, 4590, 4591, 4592, 4593, 4595, 4596, or 4598. In some embodiments, the engineered guide RNA comprises a nucleotide sequence that has at least 80% identity to any one of SEQ ID NOs: 4543, 4546, 4547, 4548, 4551, 4553, 4554, 4555, 4556, 4558, 4559, 4560, 4561, 4562, 4565, or 4567. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, a hepatocyte, or a precursor thereof.

[0027] In some aspects, the disclosure provides a method of disrupting a B2M locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the B2M locus, wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% of any one of SEQ ID NOs: 4676-4751. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4600-4675. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNA is selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734 In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, and 6033 to 6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides that have at least 80% identity to any one of SEQ ID NOs: 4676, 4678-4687, 4690, 4692, 4698-4707, 4720-4723, 4725-4726, 4732-4733, 4736-4737, 4741, or 4750-4751. In some embodiments, the engineered guide RNA comprises a nucleotide sequence having at least 80% identity to any one of SEQ ID NOs: 4600, 4602-4611, 4614, 4616, 4622-4631, 4644-4647, 4649-4650, 4656-4657, 4660-4661, 4665, or 4674-4675. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0028] In some aspects, the disclosure provides a method of disrupting the CD2 locus in a cell, the method comprising introducing into the cell (a) a Class 2, V-type Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the CD2 locus, wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% of any one of SEQ ID NOs: 4837-4921. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4752-4836. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, and 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides that have at least 80% identity to any one of SEQ ID NOs: 4837, 4844, 4845, 4848, 4857-4858, 4883, 4887, 4892-4893, 4904-4909, 4914, 4916, or 4918. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 14E that target any one of SEQ ID NOs: 4837, 4844, 4845, 4848, 4857-4858, 4883, 4887, 4892-4893, 4904-4909, 4914, 4916, or 4918. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0029] In some aspects, the disclosure provides a method of disrupting the CD5 locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the CD5 locus, wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% of any one of SEQ ID NOs: 4946-4969. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4922-4945. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, and 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 4946-4947, 4949, 4951, 4957-4960, 4963, 4967, or 4969. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 14F that target any one of SEQ ID NOs: 4946-4947, 4949, 4951, 4957-4960, 4963, 4967, or 4969. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0030] In some aspects, the disclosure provides a method of disrupting a mouse TRAC locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the mouse TRAC locus, wherein the engineered guide RNA hybridizes to any one of SEQ ID NOs: 5126-5195, 5682, or 5684 by at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5056-5125, 5681, or 5683. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3731-3732, 3733-3734, 3735-3736, 3737-3738, 3740-3741, 3742-3743, 3744-3745, 3746-3747, 3748-3749, 3751-3752, 3753-3754, 3755-3756, 3757-3758, 3759-3760, 3761-3762, 3763-3764, 3765-3766, 3767-3768, 3771-3772, 3773-3774, 3775-3776, 3776-3778, 3779-3780, 3781-3782, 3783-3784, 3785-3786, 3787-3788, 3789-3790, 3791-37 In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 34-3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides that have at least 80% identity to any one of SEQ ID NOs: 5126-5130, 5133-5143, 5147-5150, 5172-5173, 5184-5189, or 5192-5194. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 14G that target any one of SEQ ID NOs: 5126-5130, 5133-5143, 5147-5150, 5172-5173, 5184-5189, or 5192-5194. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0031] In some aspects, the disclosure provides a method for disrupting a mouse TRBC1 or TRBC2 locus in a cell, the method comprising introducing into the cell (a) a Class 2, V-type Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the mouse TRBC1 or TRBC2 locus, and the engineered guide RNA has at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, or at least about 94% affinity to any one of SEQ ID NOs: 5211-5225 or 5247-5267. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5196-5210 or 5226-5246. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3731-3732, 3733-3734, 3735-3736, 3737-3738, 3740-3741, 3742-3743, 3744-3745, 3746-3747, 3748-3749, 3751-3752, 3753-3754, 3755-3756, 3757-3758, 3759-3760, 3761-3762, 3763-3764, 3765-3766, 3767-3768, 3771-3772, 3773-3774, 3775-3776, 3776-3778, 3777-3779, 3780-3781, 3782-3782, 3783-3784, 3785-3785, 3786-3786, 3787-37 In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 34-3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides that have at least 80% identity to any one of SEQ ID NOs: 5211, 5213-5215, 5217, 5221, 5223, 5247, 5249-5250, 5252-5253, 5258-5259, or 5264. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 14H that target any one of SEQ ID NOs: 5211, 5213-5215, 5217, 5221, 5223, 5247, 5249-5250, 5252-5253, 5258-5259, or 5264. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0032] In some aspects, the disclosure provides a method of disrupting the human TRBC1 or TRBC2 locus in a cell, the method comprising introducing into the cell (a) a Class 2, V-type Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the human TRBC1 or TRBC2 locus, and wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, or at least about 94% of any one of SEQ ID NOs: 5661-5679. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5642-5660. The guide RNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5642-5660. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5661-5663, 5672-5675, or 5678. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 14I that target any one of SEQ ID NOs: 5661-5663, 5672-5675, or 5678. In some embodiments, the cell is a eukaryotic cell, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0033] In some aspects, the disclosure provides methods of disrupting the HPRT locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the HPRT locus, wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94% of any one of SEQ ID NOs: 5562-5641. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5482-5561. The guide RNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5482-5561. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5562-5564 or 5568. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 14J that target any one of SEQ ID NOs: 5562-5564 or 5568. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0034] In some aspects, the disclosure provides a method of disrupting an APO-A1 locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the APO-A1 locus, and wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at least about 136%, at least about 137%, at least about 138 The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having 4%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5847-5860. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NOs: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5861-5866 or 5868-5869. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 43A that target any one of SEQ ID NOs: 5861-5866 or 5868-5869. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0035] In some aspects, the disclosure provides a method of disrupting the ANGPTL3 locus in a cell, the method comprising introducing into the cell (a) a Class 2, V-type Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the ANGPTL3 locus, and wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at least about 136%, at least about 137%, at least about 1 The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5875-5952. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the engineered guide RNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides that have at least 80% identity to any one of SEQ ID NOs: 5955-5963, 5968-5975, 5979-5987, 5989-5993, 5997, 5999, 6003-6010, 6014-6016, 6024-6025, or 6027-6030. In some embodiments, the engineered guide RNA has at least 80% sequence identity to any of the guide RNAs from Table 43B that target any one of SEQ ID NOs: 5955-5963, 5968-5975, 5979-5987, 5989-5993, 5997, 5999, 6003-6010, 6014-6016, 6024-6025, or 6027-6030. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0036] In some aspects, the disclosure provides a method of disrupting the human Rosa26 locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the human Rosa26 locus, wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, or at least about 94% of any one of SEQ ID NOs: 5013-5055. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4970-5012. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0037] In some aspects, the disclosure provides a method of disrupting a FAS locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the FAS locus, wherein the engineered guide RNA hybridizes to at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% of any one of SEQ ID NOs: 5367-5465. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity, and comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5268-5366. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0038] In some aspects, the disclosure provides a method of disrupting the PD-1 locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the PD-1 locus, and wherein the engineered guide RNA hybridizes to any one of SEQ ID NOs: 5474-5481 by at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, or at least about 95% of any one of SEQ ID NOs: 5474-5481. The guide RNA is configured or engineered to hybridize to a sequence having at least 20-22 contiguous nucleotides complementary to a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5466-5473. The guide RNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5466-5473. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0039] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease having at least 75% sequence identity to any one of SEQ ID NO: 215 or a variant thereof; and (b) an engineered guide RNA configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the system has reduced immunogenicity when administered to a human subject compared to a comparable system comprising a Cas9 enzyme. In some embodiments, the Cas9 enzyme is a SpCas9 enzyme. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the immunogenicity is antibody immunogenicity.

[0040] Aspects of the present disclosure provide methods for disrupting the mouse HAO-1 locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the mouse HAO-1 locus, wherein the engineered guide RNA comprises nucleotides of guide RNA mH29-1_37, mH29-15_37, mH29-29_37 from Table 25, comprising a nucleotide modification as set forth in Table 25, or wherein the engineered guide RNA comprises any one of SEQ ID NOs: 4184-4225. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof. In some embodiments, the engineered guide RNA is set forth in SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609.In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof. In some embodiments, the engineered guide RNA comprises nucleotides of guide RNA mH29-15_37 or mH29-29_37 from Table 25, wherein the nucleotide modification is set forth in Table 25. In some embodiments, the method further comprises disrupting expression of glycolate oxidase from the HAO-1 locus.

[0041] In some aspects, the disclosure provides a method of disrupting the human TRAC locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the human TRAC locus, and wherein the engineered guide RNA comprises the nucleotides of MG29-1-TRAC-sgRNA-35 from Table 28B, wherein the nucleotide modifications comprise those set forth in Table 28B. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof. In some embodiments, the engineered guide RNA is set forth in SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof.

[0042] Aspects of the present disclosure provide a method of disrupting an albumin locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the albumin locus, wherein the engineered guide RNA is a nucleoside selected from the group consisting of a nucleotide sequence listed in Table 29. or the engineered guide nucleotide RNA comprises nucleotides mAlb298-37, mAlb2912-37, mAlb2918-37, or mAlb298-34 from Table 29 containing a nucleotide modification, or the engineered guide nucleotide RNA comprises nucleotides mAlb29-8-44, mAlb29-8-50, mAlb29-8-50b, mAlb29-8-51b, mAlb29-8-52b, mAlb29-8-53b, or mAlb29-8-54b containing a nucleotide modification as set forth in Table 46. In some embodiments, the Class 2, Type V Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470, or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof.In some embodiments, the engineered guide RNAs are selected from the group consisting of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734- In some embodiments, the guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the non-degenerate nucleotides of any one of SEQ ID NO: 3735, 3851-3857, or 6033-6036. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO: 3609. In some embodiments, the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof. In some embodiments, the engineered guide RNA comprises nucleotides mAlb298-37, mAlb2912-37, mAlb2918-37, or mAlb298-34 from Table 29, comprising a nucleotide modification as set forth in Table 29.

[0043] In some aspects, the disclosure provides an engineered guide RNA comprising: (a) a DNA-targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment configured to bind to a Class 2, Type V Cas endonuclease, wherein the guide RNA comprises a nucleotide modification pattern set forth in any one of SEQ ID NOs:5695-5701 in Table 34. In some embodiments, the guide RNA comprises mAlb29-8-44, mAlb29-8-50, mAlb29-8-37, or mAlb29-12-44. In some embodiments, the guide RNA comprises hH29-4_50, hH29-21_50, hH29-23_50, hH29-41_50, hH29-4_50b, hH29-21_50b, hH29-23_50b, or hH29-41_50b, mH29-1-50, mH29-15-50, mH29-29-50, mH29-1-50b, mH29-15-50b, or mH29-29-50b. In some embodiments, the DNA-targeting segment is configured to hybridize to the HAO-1 gene or the albumin gene. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof. In some embodiments, the Class 2, V-type Cas endonuclease comprises an endonuclease having at least 75% sequence identity to SEQ ID NO: 215.

[0044] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease having at least 75% sequence identity to any one of SEQ ID NOS: 1-3470 or a variant thereof; and (b) a polynucleotide sequence encoding a CRISPR array, wherein the CRISPR array is configured to be processed by the endonuclease into an engineered guide RNA, the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a target nucleic acid sequence, the spacer sequence configured to hybridize to an albumin gene. In some embodiments, the polynucleotide sequence comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 5712. In some embodiments, the endonuclease comprises an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722, or a variant thereof. In some embodiments, the endonuclease comprises an endonuclease having at least 75% sequence identity to SEQ ID NO: 215.

[0045] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease having at least 75% sequence identity to SEQ ID NO:470 or a variant thereof; and (b) an engineered guide RNA configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize to a target nucleic acid sequence. In some embodiments, the engineered guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO:6031. In some embodiments, the endonuclease is configured to be selective for a 5' PAM sequence comprising SEQ ID NO:6032.

[0046] In some aspects, the disclosure provides a method for detecting a denatured endonuclease comprising: (a) an endonuclease having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 2824, 2841, or 2896 or a variant thereof; and (b) a nucleic acid molecule capable of forming a complex with the endonuclease. and an engineered guide RNA comprising a spacer sequence configured to hybridize to a target nucleic acid sequence, wherein the engineered guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 6033, 6034, or 6035. In some embodiments, the endonuclease has at least 80% sequence identity to SEQ ID NO: 2824 and the engineered guide RNA has at least 80% sequence identity to SEQ ID NO: 6033. In some embodiments, the endonuclease has at least 80% sequence identity to SEQ ID NO: 2841 and the engineered guide RNA has at least 80% sequence identity to SEQ ID NO: 6034. In some embodiments, the endonuclease has at least 80% sequence identity to SEQ ID NO: 2896 and the engineered guide RNA has at least 80% sequence identity to SEQ ID NO: 6035. In some embodiments, the endonuclease is configured to be selective for a 5' PAM sequence comprising any one of SEQ ID NOs: 6037-6039.

[0047] In some aspects, the present disclosure provides lipid nanoparticles comprising (a) any of the endonucleases described herein, (b) any of the engineered guide RNAs described herein, (c) a cationic lipid, (d) a sterol, (e) a neutral lipid, and (f) a PEG-modified lipid. In some embodiments, the cationic lipid comprises C12-200, the sterol comprises cholesterol, the neutral lipid comprises DOPE, or the PEG-modified lipid comprises DMG-PEG2000. In some embodiments, the cationic lipid comprises any of the cationic lipids shown in Figure 109.

[0048] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.

[0049] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]

[0050] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings.

[0051] [Figure 1] 1 shows exemplary organizations of different classes and types of CRISPR / Cas loci previously documented prior to this disclosure. [Figure 2]1 shows the environmental distribution of the MG nucleases described herein. Protein lengths are shown for representatives of the MG29 protein family. The shading of the circles indicates the environment or type of environment in which each protein was identified (dark gray circles indicate high-temperature sources, light gray circles indicate non-high-temperature sources). N / A indicates that the type of environment from which the sample was collected is unknown. [Figure 3] The number of predicted catalytic residues present in MG nucleases detected from the sample types described herein is shown (e.g., Figure 1). Protein lengths are shown for representatives of the MG29 protein family. The number of predicted catalytic residues for each protein is indicated in the figure legend (3.0 residues). The first, second, and third catalytic residues are located in the RuvCI, RuvCII, and RuvCIII domains, respectively. [Figure 4A] Diversity of CRISPR type VA effectors is shown. Figure 4A shows the family-by-family distribution of taxonomic classifications of contigs encoding novel type VA effectors. MG families are indicated in parentheses. PAM requirements for active nucleases are outlined in boxes associated with the families. Non-type VA reference sequences were used to root the gene tree (*MG61 family requires crRNAs with alternative stem-loop sequences). [Figure 4B] Figure 4B shows the diversity of CRISPR type VA effectors. Figure 4B shows a gene tree inferred from the alignment of 119 novel and 89 reference type V effector sequences. MG families are indicated in parentheses. The PAM requirements for active nucleases are outlined in the boxes associated with the families. Non-type VA reference sequences were used to root the gene tree (*MG61 family requires crRNAs with alternative stem-loop sequences). [Figure 5A] Various characteristic information about the nucleases described herein is provided: Figure 5A shows the distribution of effector protein lengths by family and sample type; [Figure 5B]Various characteristic information regarding the nucleases described herein is provided: Figure 5B shows the presence of RuvC catalytic residues. [Figure 5C] Figure 5C provides various characteristic information about the nucleases described herein. Figure 5C shows the number of CRISPR arrays with various repeat motifs. [Figure 5D] Various characteristic information about the nucleases described herein is provided. Figure 5D shows the family-wise distribution of repeat motifs. [Figure 6A] Figure 6A shows a multiple sequence alignment of the catalytic and PAM-interacting regions in VA-type sequences. Francisella novicida Cas12a (FnCas12) is the reference sequence. Other reference sequences are Acidaminococcus sp. (AsCas12a), Moraxella bovoculi (MbCas12a), and Lachnospiraceae bacterium ND2006 (LbCas12a). Figure 6A shows the blocks of conservation around the DED catalytic residues in the RuvC-I (left), RuvC-II (center), and RuvC-III (right) regions. The gray boxes below the FnCas12a sequence identify the domains. [Figure 6B] Figure 6B shows a multiple sequence alignment of the catalytic and PAM-interacting regions in a VA-type sequence. Francisella novicida Cas12a (FnCas12) is the reference sequence. Other reference sequences are Acidaminococcus sp. (AsCas12a), Moraxella bovoculi (MbCas12a), and Lachnospiraceae bacterium ND2006 (LbCas12a). Figure 6B shows the WED-II and PAM-interacting regions, including residues involved in PAM recognition and interaction. Gray boxes below the FnCas12a sequence identify domains. Dark boxes within the alignment indicate increasing sequence identity. Black boxes above the FnCas12a sequence indicate catalytic residues (and positions) in the reference sequence. Gray boxes indicate domains in the reference sequence at the top of the alignment (FnCas12a). Black boxes indicate catalytic residues (and positions) in the reference sequence. [Figure 7A] Figure 7A shows the V-A type and the associated V-A' effectors. Figure 7A shows the V-A type (MG26-1) and V-A' type (MG26-2) indicated by arrows showing the direction of transcription. The CRISPR array is shown as a gray bar. The predicted domains of each protein within the contig are indicated by boxes. [Figure 7B] Figure 7B shows a sequence alignment of type V-A' MG26-2 and AsCas12a reference sequences. Top: RuvC-I domain. Middle: Region containing RuvC-I and RuvC-II catalytic residues. Bottom: Region containing RuvC-III catalytic residues. Catalytic residues are indicated by boxes. [Figure 8] A schematic diagram of the structure of the sgRNA and target DNA in a ternary complex with AacC2C1 is shown (see Yang, Hui, Pu Gao, Kanagalaghatta R. Rajashankar, and Dinshaw J. Patel. 2016. "PAM-Dependent Target DNA Recognition and Cleavage by C2c1 CRISPR-Cas Endonuclease." Cell 167(7):1814-28. e12, which is incorporated by reference in its entirety). [Figure 9]Figure 1 shows the effect of mutations or truncations in the R-AR domain of sgRNA on AacC2c1-mediated cleavage of linear plasmid DNA; WT, wild-type sgRNA. Mutant nucleotides within the sgRNA (lanes 1-5) are highlighted in the left panel. Δ15: 15 nt deleted from the sgRNA R-AR1 region. Δ12: 12 nt removed from the sgRNA J2 / 4 R-AR1 region (see Liu, Liang, Peng Chen, Min Wang, Xueyan Li, Jiuyu Wang, Maolu Yin, and Yanli Wang. 2017. "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism." Molecular Cell 65(2):310-22, which is incorporated herein by reference in its entirety). [Figure 10A] Figure 10A shows the structural fold-up of the reference crRNA sequence in the LbCpf1 line, demonstrating that the CRISPR RNA (crRNA) structure is conserved among VA-type lines. [Figure 10B] Figure 10B shows that the CRISPR RNA (crRNA) structure is conserved among VA-type systems. Figure 10B shows a multiple sequence alignment of CRISPR repeats associated with novel VA-type systems. The LbCpf1 processing site is indicated by a black bar. [Figure 10C] Figure 10C shows that CRISPR RNA (crRNA) structure is conserved among VA-type strains. Figure 10C shows the structural fold-up of the MG61-2 putative crRNA, which has an alternative stem-loop motif, CCUGC[N3-4]GCAGG. [Figure 10D] Figure 10D shows that the CRISPR RNA (crRNA) structure is conserved among VA-type systems. Figure 10D shows multiple sequence alignments of CRISPR repeats with alternative repeat motif sequences. Processing sites and loops are indicated. [Figure 11] The predicted structure of the guide RNA utilized herein is shown (SEQ ID NO: 3608). [Figure 12-1]1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3636, 3637, 3641, 3640). [Figure 12-2] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3636, 3637, 3641, 3640). [Figure 13-1] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3644, 3645, 3649, 3648). [Figure 13-2] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3644, 3645, 3649, 3648). [Figure 14-1] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3652, 3653, 3657, 3656). [Figure 14-2] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3652, 3653, 3657, 3656). [Figure 15-1] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3660, 3661, 3665, 3664). [Figure 15-2] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3660, 3661, 3665, 3664). [Figure 16-1] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3666, 3667, 3672, 3671). [Figure 16-2] 1 shows the predicted structures of the corresponding sgRNAs of the MG enzymes described herein (clockwise, SEQ ID NOs: 3666, 3667, 3672, 3671). [Figure 17A]Figure 17A shows an agarose gel showing the results of PAM vector library cleavage in the presence of TXTL extracts containing various MG family nucleases and their corresponding tracrRNAs or sgRNAs (described in Example 12). Figure 17A shows lane 1: ladder. The bands, from top to bottom, are 766, 500, 350, 300, 350, 200, 150, 100, 75, and 50; lane 2: 28-1 + MGcrRNA spacer 1 (SEQ ID NO: 141 + 3860); lane 3: 29-1 + MGcrRNA spacer 1 (SEQ ID NO: 215 + 3860); lane 4: 30-1 + MGcrRNA spacer 1 (SEQ ID NO: 226 + 3860); lane 5: 31-1 + MGcrRNA spacer 1 (SEQ ID NO: 229 + 3860); lane 6: 32-1 + MGcrRNA spacer 1 (SEQ ID NO: 261 + 3860); lane 7: ladder. [Figure 17B] Figure 17B shows an agarose gel showing the results of PAM vector library cleavage in the presence of TXTL extracts containing various MG family nucleases and their corresponding tracrRNAs or sgRNAs (described in Example 12). Figure 17B shows lane 1: ladder; lane 2: LbaCas12a + LbaCas12a crRNA spacer 2; lane 3: LbaCas12a + MGcrRNA spacer 2; lane 4: Apo 13-1; lane 5: 28-1 + MGcrRNA spacer 2 (SEQ ID NO: 141 + 3861); lane 6: 29-1 + MGcrRNA spacer 2 (SEQ ID NO: 215 + 3861); lane 7: 30-1 + MGcrRNA spacer 2 (SEQ ID NO: 226 + 3861); lane 8: 31-1 + MGcrRNA spacer 2 (SEQ ID NO: 229 + 3861); lane 9: 32-1 + MGcrRNA spacer 2 (SEQ ID NO: 261 + 3861). [Figure 18A-1] We provide data demonstrating that the VA-type effectors described herein are active nucleases. Figure 18A-1 shows the seqLogo representation of the PAM sequences determined for the three nucleases described herein. [Figure 18A-2]We provide data demonstrating that the VA-type effectors described herein are active nucleases. Figure 18A-2 shows the seqLogo representation of the PAM sequences determined for the three nucleases described herein. [Figure 18B] Figure 18B provides data showing that the VA-type effectors described herein are active nucleases. Figure 18B shows boxplots of plasmid transfection activity assays estimated from indel editing frequencies for active nucleases. Boxplot boundaries indicate the first and third quartiles. The mean is indicated by an "x", and the median is represented by the midline within each box. [Figure 18C] This paper provides data demonstrating that the VA-type effectors described herein are active nucleases. Figure 18C shows the plasmid transfection editing frequencies of MG29-1 and AsCas12a at four target sites. One comparative experiment was performed using AsCas12a. Figures 18C and 18D show bar plots showing the average editing frequency with one standard deviation error bar. [Figure 18D] Figure 18C and Figure 18D show data demonstrating that the VA-type effectors described herein are active nucleases. Figure 18D shows the plasmid and RNP editing activity of nuclease MG29-1 at 14 target loci with either TTN or CCN PAMs. Figures 18C and 18D show bar plots showing the average editing frequency with one standard deviation error bars. [Figure 18E] This paper provides data demonstrating that the VA-type effectors described herein are active nucleases. Figure 18E shows the editing profile of nuclease MG29-1 from an RNP transfection assay. A comparative experiment was performed using AsCas12a. MG29-1 editing frequency and profile experiments were performed in duplicate. [Figure 19] FIG. 10 shows cellular indel formation produced by transfection of HEK cells with the MG29-1 constructs described in Example 12, along with corresponding sgRNAs containing a variety of different targeting sequences targeting various locations in the human genome. [Figure 20] 1 shows a seqLogo representation of the PAM sequences of certain MG family enzymes derived via NGS as described herein (as described in Example 13). [Figure 21] SEQ ID NOs: 3865, 3867, 3872. SeqLogo representations of PAM sequences for specific MG family enzymes derived via NGS as described herein are shown (top to bottom: 3865, 3867, 3872). [Figure 22] 1 shows a seqLogo representation of the PAM sequences derived via NGS as described herein (from top to bottom, SEQ ID NOs: 3878, 3879, 3880, 3881). [Figure 23] 1 shows a seqLogo representation of the PAM sequences derived via NGS as described herein (from top to bottom, SEQ ID NOs: 3883, 3884, 3885). [Figure 24]

[0023] Figure 1 shows a seqLogo representation of the PAM sequence derived via NGS as described herein (sequence number 3882). [Figure 25] FIG. 10 shows cellular indel formation produced by transfection of HEK cells with the MG31-1 constructs described in Example 14, along with corresponding sgRNAs containing a variety of different targeting sequences targeting various locations in the human genome. [Figure 26A] Biochemical characterization of type VA nucleases is shown. Figure 26A shows PCR of cleavage products with adapters ligated to their ends, demonstrating the activity of the nucleases described herein and Cpfl (positive control) when bound to universal crRNA. The expected cleavage product bands are labeled with arrows. [Figure 26B] Figure 26B shows the biochemical characterization of type VA nucleases. Figure 26B shows PCR of the cleavage products with adapters ligated to their ends, demonstrating the activity of the nucleases described herein when bound to native crRNA. The cleavage product bands are indicated by arrows. [Figure 26C]Figure 26C shows biochemical characterization of type VA nucleases. NGS cleavage site analysis shows cleavage on the target strand at position 22, with occasional low frequency cleavage after 21 or 23 nt. [Figure 27A] (Figure 27A) and (Figure 27B) show multiple sequence alignments of the VL-type nucleases described herein, showing exemplary locus organizations of VL-type nucleases. The region containing the putative RuvC-III domain is shown as a light gray rectangle. Putative RuvC catalytic residues are shown as small dark gray rectangles above each sequence. Putative single-stranded RNA binding sequences are small white rectangles, putative cleavable phosphate binding sites are shown as black rectangles above the sequences, and residues predicted to disrupt base stacking near the cleavable phosphate in the target sequence are shown as small gray rectangles above the sequences. [Figure 27B] (Figure 27A) and (Figure 27B) show multiple sequence alignments of the VL-type nucleases described herein, showing exemplary locus organizations of VL-type nucleases. The region containing the putative RuvC-III domain is shown as a light gray rectangle. Putative RuvC catalytic residues are shown as small dark gray rectangles above each sequence. Putative single-stranded RNA binding sequences are small white rectangles, putative cleavable phosphate binding sites are shown as black rectangles above the sequences, and residues predicted to disrupt base stacking near the cleavable phosphate in the target sequence are shown as small gray rectangles above the sequences. [Figure 28] Shown is a VL-type candidate labeled MG60 as an exemplary locus organization along with the effector repeat structure and a phylogenetic tree showing the position of the enzyme in the V-type family. [Figure 29] Examples of smaller V-type effectors are shown, one of which can be labeled as MG70. [Figure 30-1]

[0023] Figure 1 shows the characterization information for MG70 described herein. Exemplary locus organization is illustrated along with a phylogenetic tree showing the position of these enzymes in the V-type family. [Figure 30-2]

[0023] Figure 1 shows the characterization information for MG70 described herein. Exemplary locus organization is illustrated along with a phylogenetic tree showing the position of these enzymes in the V-type family. [Figure 31-1] 1 shows another example of the small type V effector MG81 described herein. An exemplary locus organization is shown along with a phylogenetic tree showing the position of these enzymes in the type V family. [Figure 31-2] 1 shows another example of the small type V effector MG81 described herein. An exemplary locus organization is shown along with a phylogenetic tree showing the position of these enzymes in the type V family. [Figure 32] The activity of individual enzymes in the V-type effector family identified herein (e.g., MG20, MG60, MG70, and others) is maintained across a range of different enzyme lengths (e.g., 400-1200 AA). Light dots (true) indicate active enzymes, while dark dots (unknown) indicate untested enzymes. [Figure 33] Figure 1 shows the sequence conservation of the MG nucleases described herein. Black bars indicate putative RuvC catalytic residues. [Figure 34] Shown is an expanded version of the multiple sequence alignment in Figure 33 of a region of the MG nuclease described herein, including the putative RuvC catalytic residues (dark grey rectangle), the scissile phosphate binding residues (black rectangle), and residues predicted to disrupt base stacking adjacent to the scissile phosphate (light grey rectangle). [Figure 35] Shown is an expanded version of the multiple sequence alignment in Figure 33 of a region of the MG nuclease described herein, including the putative RuvC catalytic residues (dark grey rectangle), the scissile phosphate binding residues (black rectangle), and residues predicted to disrupt base stacking adjacent to the scissile phosphate (light grey rectangle). [Figure 36] 1 shows the region of the MG nuclease described herein, including the putative RuvC-III domain and catalytic residues. [Figure 37]The region of the MG nuclease containing the putative single-stranded RNA binding residues (white rectangle above the sequence) is indicated. [Figure 38] Multiple protein sequence alignment of representatives from several MG type V families is shown. Conserved regions, including parts of the RuvC domain predicted to be involved in nuclease activity, are indicated. Predicted catalytic residues are highlighted. [Figure 39]

[0023] Figure 1 shows screening of the TRAC locus for MG29-1 gene editing. The bar graph shows indel generation resulting from transfection of MG29-1 with 54 distinct guide RNAs targeting the TRAC locus in primary human T cells. The corresponding guide RNAs shown in the figure are identified by SEQ ID NOs: 4316-4423. [Figure 40]

[0049] Figure 39 shows optimization of MG29-1 editing in TRAC. The bar graph shows indel generation resulting from transfection of MG29-1 with the four best 22-nt guide RNAs from Figure 39 (9, 19, 25, and 35) (at the concentrations shown). Legend: MG29-1 9 is MG29-1 effector (SEQ ID NO:215) and guide 9 (SEQ ID NO:4378), MG29-1 19 is MG29-1 effector (SEQ ID NO:215) and guide 19 (SEQ ID NO:4388), MG29-1 25 is MG29-1 effector (SEQ ID NO:215) and guide 25 (SEQ ID NO:4394), and MG29-1 35 is MG29-1 effector (SEQ ID NO:215) and guide 35 (SEQ ID NO:4404). [Figure 41] Figure 39 shows the optimization of dose and guide length for MG29-1 editing in TRAC. The line graph shows indel generation resulting from transfection of MG29-1 and either guide RNA #19 (SEQ ID NO: 4388) or guide RNA #35 (SEQ ID NO: 4404). Three different doses of nuclease / guide RNA were used. For each dose, six different guide lengths were tested: consecutive single nucleotide 3' truncations of SEQ ID NOs: 4388 and 4404. The guide used in Figures 39 and 40 is a 22-nt long spacer-containing guide in this case. [Figure 42] 1 shows the correlation between indel generation in TRAC and loss of T cell receptor expression in the experiments of Example 22. [Figure 43] Figure 1 shows targeted transgene integration in TRAC stimulated by MG29-1 cleavage. Cells receiving only the transgene donor by AAV infection retain TCR expression and lack CAR expression. Cells transfected with MG29-1 RNP and infected with 100,000 vg (vector genome) of the CAR transgene donor lose TCR expression and gain CAR expression. FACS plots of CAR antigen binding versus TCR expression are shown for cells transfected with AAV alone containing the CAR-T-containing donor sequence (AAV), AAV containing the CAR-T-containing donor sequence containing the MG29-1 enzyme and sgRNA 19 (SEQ ID NO: 4388) ("AAV+MG29-1-19-22" containing a 22-nucleotide spacer), or AAV containing the CAR-T-containing donor sequence containing the MG29-1 enzyme and sgRNA 35 (SEQ ID NO: 4404) ("AAV+MG29-1-35-22" containing a 22-nucleotide spacer). [Figure 44] MG29-1 gene editing in TRAC in hematopoietic stem cells. The bar graph shows the extent of indel generation in TRAC after transfection with MG29-1-9-22 ("MG29-1 9"; MG29-1 + guide RNA #19) and MG29-1-35-22 ("MG29-1 35"; MG29-1 + guide RNA #35) compared to mock-transfected cells. [Figure 45] Figure 1 shows the refinement of the MG29-1 PAM based on analysis of gene editing results in cells. Guide RNAs were designed using the 5'-NTTN-3' PAM sequence and then classified according to observed gene editing activity. The identity of the underlined base (5'-proximal N) is shown for each bin. All guides with activity greater than 10% have a T at this position in the genomic DNA, indicating that the MG29-1 PAM can best be described as 5'-TTTN-3'. The statistical significance of the overrepresentation of T at this position is shown for each bin. [Figure 46-1]Analysis of gene editing activity for the base composition of the MG29-1 spacer sequence. The bar graph shows experimental data demonstrating the relationship between GC content (%) and indel frequency ("High" means >50% indels (N=4), "Medium" means 10-50% indels (N=15), ">1%" means 1-5% indels (N=12), and "<1%" means less than 1% indels (N=82)). [Figure 46-2] Analysis of gene editing activity for the base composition of the MG29-1 spacer sequence. The bar graph shows experimental data demonstrating the relationship between GC content (%) and indel frequency ("High" means >50% indels (N=4), "Medium" means 10-50% indels (N=15), ">1%" means 1-5% indels (N=12), and "<1%" means less than 1% indels (N=82)). [Figure 47] MG29-1 guide RNA chemical modifications. The bar graph shows the results of the modifications from Table 7 on VEGF-A editing activity compared to the unmodified guide RNA (Sample #1). [Figure 48-1] Dose titration of various chemically modified MG29-1 RNAs is shown. Bar graphs show indel generation after transfection of RNPs with guides using modification patterns 1, 4, 5, 7, and 8. RNP doses were 126 pmol MG29-1 and 160 pmol guide RNA, or as indicated: full dose (A), 1 / 4 (B), 1 / 8 (C), 1 / 16 (D), and 1 / 32 (E). [Figure 48-2] Dose titration of various chemically modified MG29-1 RNAs is shown. Bar graphs show indel generation after transfection of RNPs with guides using modification patterns 1, 4, 5, 7, and 8. RNP doses were 126 pmol MG29-1 and 160 pmol guide RNA, or as indicated: full dose (A), 1 / 4 (B), 1 / 8 (C), 1 / 16 (D), and 1 / 32 (E). [Figure 49]1 shows the plasmid map of pMG450 (MG29-1 nuclease protein in a lac-inducible tac promoter E. coli BL21 expression vector). [Figure 50] 1 shows the indel profile of MG29-1 with spacer mALb29-1-8 (sequence number 3999) compared to spCas9 with a guide targeting mouse albumin intron 1. [Figure 51-1] Representative indel profile of MG29-1 with a guide targeting mouse albumin intron 1 as determined by next generation sequencing (approximately 15,000 total reads analyzed) as in Example 29. [Figure 51-2] Representative indel profile of MG29-1 with a guide targeting mouse albumin intron 1 as determined by next generation sequencing (approximately 15,000 total reads analyzed) as in Example 29. [Figure 52] 1 shows the editing efficiency of MG29-1 compared to spCas9 in the mouse liver cell line Hepa1-6 nucleofected with RNP as in Example 29. [Figure 53A] Figure 53A shows the editing efficiency in mammalian cells of MG29-1 variants with single and double amino acid substitutions compared to wild-type MG29-1. Figure 53B shows the editing efficiency in Hepa1-6 cells transfected with plasmids encoding MG29-1 WT or mutant forms. [Figure 53B] Figure 53B shows the editing efficiency in mammalian cells of MG29-1 variants with single and double amino acid substitutions compared to wild-type MG29-1. Figure 53B shows the editing efficiency in Hepa1-6 cells transfected with various concentrations of mRNA encoding WT or S168R. [Figure 53C] Figure 53C shows the editing efficiency in mammalian cells of MG29-1 variants with single and double amino acid substitutions compared to wild-type MG29-1. Figure 53D shows the editing efficiency in Hepa 1-6 cells transfected with mRNA-encoded versions of MG29-1 with single or double amino acid substitutions. [Figure 53D] Figure 53D shows the editing efficiency in mammalian cells of MG29-1 variants with single and double amino acid substitutions compared to wild-type MG29-1. Figure 53D shows the editing efficiency in Hepa1-6 cells and HEK293T cells transfected with MG29-1 WT vs. S168R in combination with 13 guides. Twelve guides correspond to the guides in Table 7. Guide "35(TRAC)" is a guide targeting the human locus TRAC. [Figure 54] The predicted secondary structure of MG29-1 guide mAlb29-1-8 is shown. [Figure 55] Figure 1 shows the effect of chemical modifications of the MG29-1 sgRNA sequence on the stability of the sgRNA in whole cell extracts of mammalian cells. [Figure 56A] Figure 56A shows the use of sequencing to identify cleavage sites on the target strand in an in vitro reaction performed with MG29-1 protein, guide RNA, and an appropriate template. Figure 56A shows the distance of the cleavage position in nucleotides from the PAM as determined by next-generation sequencing. [Figure 56B] Figure 56B shows the use of sequencing to identify cleavage sites on the target strand in an in vitro reaction performed with MG29-1 protein, guide RNA, and an appropriate template. Figure 56B shows the use of Sanger sequencing to define the MG29-1 cleavage site on the target strand. [Figure 56C]Figure 56C shows the use of sequencing to identify the cleavage site on the target strand in an in vitro reaction performed with MG29-1 protein, guide RNA, and appropriate template. Figure 56C shows the use of Sanger sequencing to define the MG29-1 cleavage site on the non-target strand. Runoff Sanger sequencing was performed on an in vitro reaction containing MG29-1, guide, and appropriate template to assess cleavage of both strands. The cleavage site on the target strand is at position 23, consistent with the NGS data in Figure 56A, which indicates cleavage at bases 21–23. The "A" peak at the end of the sequence is due to polymerase runoff and is expected. The cleavage site on the non-target strand can be seen in the reverse readout, where the expected terminating base is a "T." The marked location (line) indicates cleavage from position 17 from the PAM, followed by the terminal T. However, there are mixed T signals at positions 18, 19, and 20 from the PAM, suggesting variable cleavage of this strand at positions 17, 18, and 19. [Figure 57] The results of gene editing at the DNA level of CD38 are shown. The S. pyogenes (Spy) Cas9 guides for CD38 and TRAC are shown on the right. [Figure 58] The results of gene editing at the CD38 phenotypic level are shown. [Figure 59] This shows the results of gene editing at the DNA level using TIGIT. [Figure 60] The results of gene editing at the DNA level of AAVS1 are shown. [Figure 61] This shows the results of gene editing at the DNA level by B2M. [Figure 62] This shows the results of gene editing at the CD2 DNA level. [Figure 63] This shows the results of gene editing at the CD5 DNA level. [Figure 64A] This shows the results of gene editing at the DNA level of mouse TRAC. [Figure 64B] Flow cytometry results of gene editing of mouse TRAC are shown. [Figure 65] The percentage of TRAC knockout relative to the percentage of indels is shown. [Figure 66A] This shows the results of gene editing of mouse TRBC1 at the DNA level. [Figure 66B] This shows the results of gene editing of mouse TRBC2 at the DNA level. [Figure 66C] 1 shows the flow cytometry results of gene editing of human TRBC1 / 2. [Figure 67] This shows the results of gene editing at the DNA level of HPRT. [Figure 68] FIG. 1 shows the activity of chemically modified guides in Hepa1-6 cells when delivered as mRNA and gRNA using Lipofectamine Messenger Max. [Figure 69] The stability of guides modified with modification 44 relative to terminally modified or unmodified guides is shown. [Figure 70A] Figure 70A shows a comparison of guide stability between Type II and Type V systems. Figure 70A shows stability data for unmodified guides. [Figure 70B] Figure 70B shows a comparison of guide stability between Type II and Type V systems. Figure 70B shows stability data for guides with 5' and 3' end modifications. [Figure 71] Figure 1 shows the predicted secondary structures of MG29-1 (type V) and MG3-6 / 3-4 (type II) guide RNAs. The backbone (tracr) portion is indicated. [Figure 72] 1 shows the stability of guide mAlb298-34 compared to mAlb298-37 in cell lysates from Hepa1-6 cells. [Figure 73] 1 shows the editing efficiency of MG29-1 in mouse liver after in vivo delivery. [Figure 74] 1 shows analysis of gene editing results by NGS for mRNA electroporation in T cells. [Figure 75] 1 shows analysis of gene editing results by NGS for chemically modified guides. [Figure 76]ELISA results from screening performed at a 1:50 serum dilution to detect antibodies to MG29-1 are shown (n=50). Tetanus toxoid was used as a positive control due to widespread vaccination against this antigen. Serum samples above the dashed line were considered antibody positive; this line represents the mean absorbance of the negative control (human albumin) plus two standard deviations from the mean. *P<0.05, **P<0.01, ****P<0.0001 as determined by unpaired Student's t-test; ns, not significant. [Figure 77]

[0033] Figure 1 shows HAO-1 editing efficiency in mouse liver as measured by NGS. Each point represents an individual mouse. [Figure 78A] Figure 1 shows the effect of HAO-1 editing on glycolate oxidase (GO) protein levels in mouse liver as assessed by Western blot. 10 μg of total protein was loaded for each sample. [Figure 78B] Figure 1 shows the effect of HAO-1 editing on glycolate oxidase (GO) protein levels in mouse liver as assessed by Western blot. 10 μg of total protein was loaded for each sample. [Figure 79] Western blot analysis of glycolate oxidase (GO) protein levels in untreated mice compared with two individual mice treated with lipid nanoparticles (LNPs) encapsulating MG29-1 mRNA and either guide mH29-1_37 or mH29-5_37. Three different amounts of total liver protein (40 μg, 20 μg, and 10 μg) from each mouse were loaded onto a gel and then processed for detection of mouse glycolate oxidase protein. [Figure 80] An example of the indel profile of MG29-1 and sgRNA targeting the HAO-1 gene in mouse liver is shown. Samples were taken from mouse #17 (treated with lipid nanoparticles encapsulating mH29-29_37 and MG29-1 WT mRNA). [Figure 81] This shows the results of gene editing at the DNA level of TRAC in human peripheral blood B cells. [Figure 82] This shows the results of gene editing at the DNA level of TRAC in hematopoietic stem cells. [Figure 83] This shows the results of gene editing at the DNA level using TRAC in induced pluripotent stem cells (iPSCs). [Figure 84] 1 shows the results of in vivo genome editing with MG29-1, as quantified by next-generation sequencing (NGS). [Figure 85] 1 shows an exemplary indel profile generated by MG29-1 nuclease and guide 298-37 as measured by next-generation sequencing (NGS). [Figure 86] Optimization of spacer length for MG29-1 guide targeting two loci. [Figure 87] 1 shows the in vitro stability of sgRNAs for MG29-1 and MG3-6 / 3-4. [Figure 88] The predicted secondary structures of the backbone portions of the guide RNAs for MG29-1 and MG3-6 / 3-4 are shown. [Figure 89] 1 shows the predicted secondary structure of the MG3-6 / 3-4 guide with a spacer targeting mouse albumin. [Figure 90] 1 shows the predicted secondary structure of the MG29-1 guide with stem-loop 1 from MG3-6 / 3-4 added to the 5′ end. [Figure 91] Figure 1 shows the editing efficiency of MG29-1 using mouse albumin guide 8 with chemical 44 or 50 in Hepa1-6 cells by mRNA transfection or RNP nucleofection. [Figure 92] 1 shows editing in the liver of mice after administration with LNPs encapsulating MG29-1 mRNA and one of four different guide RNAs. [Figure 93] 1 shows the predicted secondary structure of the RNA molecule mAlb29-g8-37-array. [Figure 94]A plot showing editing efficiency in whole mouse livers 5 days after intravenous injection of LNPs encapsulating one of the following: (1) MG29-1 mRNA and guide mAlb29-8-50 (mA29-8-50) at three different doses, (2) spCas9 mRNA and guide mAlbR2 at three different doses, or (3) PBS buffer (control). Each circle represents a single mouse, and the bars show the mean and standard deviation. [Figure 95] Shows editing activity in Hep3B cells transfected with MG29-1 mRNA and 6sgRNA targeting human HAO-1. [Figure 96] 1 shows the editing activity of 4MG29-1 sgRNAs targeting human HAO-1 in HuH7 and Hep3B cells transfected with the ribonucleoprotein complex. [Figure 97] 1 shows the editing activity of MG29-1 with sgRNA targeting the human HAO-1 gene in primary human hepatocytes. [Figure 98A] Representative indel profiles of MG29-1 sgRNAs hH29-4-37 and hH29-21-37 in primary human hepatocytes are shown. [Figure 98B] Representative indel profiles of MG29-1 sgRNAs hH29-23-37 and hH29-41-37 in primary human hepatocytes are shown. [Figure 99] Shows the activity of MG29-1 guide RNAs with a 22 nucleotide or 20 nucleotide spacer targeting mouse HAO-1 in mouse liver. [Figure 100] 1 shows the effect of mRNA / guide RNA ratio and separate or co-formulation on editing efficiency in mouse liver. [Figure 101] 1 shows evaluation of MG29-1 guide chemistry for editing activity in mouse liver after in vivo delivery in LNPs. [Figure 102] This shows the results of gene editing at the DNA level of APO-A1 in Hepa1-6 cells. [Figure 103]This shows the results of gene editing at the DNA level of ANGPTL3 in Hepa1-6 cells. [Figure 104A] Figure 104A shows the in vitro characterization of MG55-43. Figure 104A shows the genomic region near the MG55-43 nuclease. The gene is represented by an orange arrow. The gene encoding the candidate nuclease contains a "putative transpotase DNA-binding domain." The CRISPR array is represented by repeats and spacers. The predicted tracrRNA is shown as an arrow between the array and the nuclease and is labeled "Predicted Trimmed TracrRNA-CM2." [Figure 104B] Figure 104B shows the in vitro characterization of MG55-43. Figure 104B shows the active single guide RNA design (tracrRNA and repeats connected by a tetraloop). The color of the nitrogenated base corresponds to the base pairing probability of that base, with red being high probability and blue being low probability. [Figure 104C] Figure 104C shows in vitro characterization of MG55-43. Figure 104C shows an agarose gel demonstrating in vitro cleavage of a plasmid target DNA library using sgRNA and two different spacers (U67 and U40). Lanes not associated with MG55-43 nuclease are not shown. [Figure 104D] Figure 104D shows the in vitro characterization of MG55-43. Figure 104D shows the sequence logo showing the MG55-43 PAM sequence. [Figure 104E] Figure 104E shows in vitro characterization of MG55-43. Figure 104E shows a histogram showing the cleavage site locations of spacer sequences tested in MG55-43. [Figure 105A] An example of a genomic region encoding the MG91 nuclease is shown. Genes are represented by arrows with candidate nuclease-encoding genes labeled as such. CRISPR arrays are represented by repeats and spacers. Intergenic regions potentially encoding active tracrRNAs are highlighted as bars labeled with IG numbers or intergenic region numbers. [Figure 105B]An example of a genomic region encoding the MG91 nuclease is shown. Genes are represented by arrows with candidate nuclease-encoding genes labeled as such. CRISPR arrays are represented by repeats and spacers. Intergenic regions potentially encoding active tracrRNAs are highlighted as bars labeled with IG numbers or intergenic region numbers. [Figure 105C] An example of a genomic region encoding the MG91 nuclease is shown. Genes are represented by arrows with candidate nuclease-encoding genes labeled as such. CRISPR arrays are represented by repeats and spacers. Intergenic regions potentially encoding active tracrRNAs are highlighted as bars labeled with IG numbers or intergenic region numbers. [Figure 106A] Figure 106A shows a multiple sequence alignment of the intergenic region nucleotide sequences potentially containing tracrRNA. The upper green bar indicates a high degree of similarity between the sequences in the intergenic region. Figure 106B shows the intergenic region 2 near the MG91-15 nuclease and its analogs. [Figure 106B] Figure 106B shows a multiple sequence alignment of the intergenic region nucleotide sequences potentially containing tracrRNA. The upper green bar indicates a high degree of similarity between the sequences in the intergenic region. Figure 106C shows the intergenic region 2 near the MG91-32 nuclease and its analogs. [Figure 106C] Figure 106C shows a multiple sequence alignment of the intergenic region nucleotide sequences potentially containing tracrRNA. The upper green bar indicates a high degree of similarity between the sequences in the intergenic region. Figure 106D shows the intergenic region 2 near the MG91-87 nuclease and its analogs. [Figure 107A] Figure 107A shows the single guide RNA designs and the results of in vitro cleavage assays. For the single guide RNA designs (tracrRNA and repeat sequences), the color of the base corresponds to the base pairing probability of that base, with red representing a high probability and blue representing a low probability. Figure 107B shows the MG91-15 sgRNA1. [Figure 107B]Figure 107B shows the single guide RNA designs and the results of in vitro cleavage assays. For the single guide RNA designs (tracrRNA and repeat sequences), the color of the base corresponds to the base pairing probability of that base, with red representing a high probability and blue representing a low probability. Figure 107B shows MG91-32 sgRNA1. [Figure 107C] Figure 107C shows the single guide RNA designs and the results of in vitro cleavage assays. For the single guide RNA designs (tracrRNA and repeat sequences), the color of the base corresponds to the base pairing probability of that base, with red representing a high probability and blue representing a low probability. Figure 107D shows the MG91-87 sgRNA1. [Figure 107D] Figure 107D shows the results of single-guide RNA designs and in vitro cleavage assays. For the single-guide RNA designs (tracrRNA and repeat sequences), the color of the base corresponds to the base-pairing probability of that base, with red representing a high probability and blue representing a low probability. Figure 107D shows an agarose gel demonstrating the in vitro cleavage of a plasmid target DNA library with different sgRNA designs (sgRNA1 and sgRNA2) and two different spacers (U67 and U40). Lanes not associated with these nucleases are not shown. [Figure 108A] Figures 108A, 108B, and 108C show sequence logos representing the PAM sequences of MG91-15, MG91-32, and MG91-87, respectively. [Figure 108B] Figures 108A, 108B, and 108C show sequence logos representing the PAM sequences of MG91-15, MG91-32, and MG91-87, respectively. [Figure 108C] Figures 108A, 108B, and 108C show sequence logos representing the PAM sequences of MG91-15, MG91-32, and MG91-87, respectively. [Figure 108D]Figures 108D, 108E, and 108F show histograms indicating the cleavage site locations of the spacer sequences tested in MG91-15, MG91-32, and MG91-87, respectively. [Figure 108E] Figures 108D, 108E, and 108F show histograms indicating the cleavage site locations of the spacer sequences tested in MG91-15, MG91-32, and MG91-87, respectively. [Figure 108F] Figures 108D, 108E, and 108F show histograms indicating the cleavage site locations of the spacer sequences tested in MG91-15, MG91-32, and MG91-87, respectively. [Figure 109] 1 shows the structures of exemplary cationic lipids that can be used in the lipid nanoparticles described herein.

[0052] Brief Description of Sequence Listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.

[0053] MG11 SEQ ID NOs: 1 to 37 show the full-length peptide sequences of MG11 nuclease.

[0054] SEQ ID NO: 3471 shows the crRNA 5' direct repeat designed to function with MG11 nuclease.

[0055] SEQ ID NOs: 3472 to 3538 show the effector repeat motif of MG11 nuclease.

[0056] SEQ ID NOs: 38 to 118 show the full-length peptide sequences of MG13 nuclease.

[0057] SEQ ID NOs: 3540 to 3550 show the effector repeat motif of MG13 nuclease.

[0058] MG19 SEQ ID NOs: 119 to 124 show the full-length peptide sequences of MG19 nuclease.

[0059] SEQ ID NOs: 3551-3558 show the nucleotide sequences of sgRNAs engineered to function with MG19 nuclease.

[0060] SEQ ID NOs: 3863 to 3866 show PAM sequences compatible with MG19 nuclease.

[0061] MG20 SEQ ID NO: 125 shows the full-length peptide sequence of MG20 nuclease.

[0062] SEQ ID NO: 3559 shows the nucleotide sequence of an sgRNA engineered to function with MG20 nuclease.

[0063] SEQ ID NO: 3867 shows a PAM sequence compatible with MG20 nuclease.

[0064] MG26 SEQ ID NOs: 126 to 140 show the full-length peptide sequences of MG26 nuclease.

[0065] SEQ ID NOs: 3560 to 3572 show the effector repeat motif of MG26 nuclease.

[0066] MG28 SEQ ID NOs: 141 to 214 show the full-length peptide sequences of MG28 nuclease.

[0067] SEQ ID NOs: 3573 to 3607 show the effector repeat motif of MG28 nuclease.

[0068] SEQ ID NOs: 3608-3609 show crRNA 5' direct repeats designed to function with MG28 nuclease.

[0069] SEQ ID NOs: 3868 to 3869 show PAM sequences compatible with MG28 nuclease.

[0070] MG29 SEQ ID NOs: 215 to 225 show the full-length peptide sequences of MG29 nuclease.

[0071] SEQ ID NO: 5680 shows the nucleotide sequence of the MG29-1 nuclease, including the 5'UTR, NLS, CDS, NLS, 3'UTR, and polyA tail.

[0072] SEQ ID NOs: 3610-3611 show the effector repeat motif of MG29 nuclease.

[0073] SEQ ID NO: 3612 shows the nucleotide sequence of an sgRNA engineered to function with MG29 nuclease.

[0074] SEQ ID NOs: 3870 to 3872 show PAM sequences compatible with MG29 nuclease.

[0075] SEQ ID NO: 5687 shows the MG29-1 coding sequence used to produce mRNA.

[0076] SEQ ID NOs: 5830 and 5846 show the DNA sequence encoding MG29-1 mRNA.

[0077] MG30 SEQ ID NOs: 226 to 228 show the full-length peptide sequences of MG30 nuclease.

[0078] SEQ ID NOs: 3613 to 3615 show the effector repeat motif of MG30 nuclease.

[0079] SEQ ID NO: 3873 shows a PAM sequence compatible with MG30 nuclease.

[0080] MG31 SEQ ID NOs: 229 to 260 show the full-length peptide sequences of MG31 nuclease.

[0081] SEQ ID NOs: 3616 to 3632 show the effector repeat motifs of MG31 nuclease.

[0082] SEQ ID NOs: 3874 to 3876 show PAM sequences compatible with MG31 nuclease.

[0083] MG32 SEQ ID NO: 261 shows the full-length peptide sequence of MG32 nuclease.

[0084] SEQ ID NOs: 3633-3634 show the effector repeat motif of MG32 nuclease.

[0085] SEQ ID NO: 3876 shows a PAM sequence compatible with MG32 nuclease.

[0086] MG37 SEQ ID NOs: 262 to 426 show the full-length peptide sequences of MG37 nuclease.

[0087] SEQ ID NO: 3635 shows the effector repeat motif of MG37 nuclease.

[0088] SEQ ID NOs: 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, and 3660-3661 set forth the nucleotide sequences of sgRNAs engineered to function with MG37 nuclease.

[0089] SEQ ID NOs: 3638, 3642, 3646, 3650, 3654, 3658, and 3662 set forth the nucleotide sequence of the MG37 tracrRNA, which is derived from the same locus as the MG37 nuclease described above.

[0090] SEQ ID NOs: 3639, 3643, 3647, 3651, 3655, and 3659 show 5' direct repeat sequences derived from the native MG37 locus that function as crRNAs when placed 5' to a 3' targeting sequence or spacer sequence.

[0091] MG53 SEQ ID NOs: 427 to 428 show the full-length peptide sequences of MG53 nuclease.

[0092] SEQ ID NO: 3663 shows a 5' direct repeat sequence derived from the native MG53 locus that functions as a crRNA when placed 5' to a 3' targeting sequence or spacer sequence.

[0093] SEQ ID NOs: 3664-3667 show the nucleotide sequences of sgRNAs engineered to function with MG53 nuclease.

[0094] SEQ ID NOs: 3668-3669 show the nucleotide sequence of MG53 tracrRNA, which is derived from the same locus as the MG53 nuclease described above.

[0095] MG54 SEQ ID NOs: 429 to 430 show the full-length peptide sequences of MG54 nuclease.

[0096] SEQ ID NO: 3670 shows a 5' direct repeat sequence derived from the native MG54 locus that functions as a crRNA when placed 5' to a 3' targeting sequence or spacer sequence.

[0097] SEQ ID NOs: 3671-3672 show the nucleotide sequences of sgRNAs engineered to function with MG54 nuclease.

[0098] SEQ ID NOs: 3673-3676 show the nucleotide sequence of MG54 tracrRNA, which is derived from the same locus as the MG54 nuclease described above.

[0099] MG55 SEQ ID NOs: 431 to 688 show the full-length peptide sequences of MG55 nuclease.

[0100] SEQ ID NO: 6031 shows the nucleotide sequence of an sgRNA engineered to function with MG55 nuclease.

[0101] SEQ ID NO: 6032 shows a PAM sequence compatible with MG55 nuclease.

[0102] MG56 SEQ ID NOs: 689 to 690 show the full-length peptide sequences of MG56 nuclease.

[0103] SEQ ID NO: 3678 shows the crRNA 5' direct repeat designed to function with MG56 nuclease.

[0104] SEQ ID NOs: 3679 to 3680 show the effector repeat motif of MG56 nuclease.

[0105] MG57 SEQ ID NOs: 691 to 721 show the full-length peptide sequences of MG57 nuclease.

[0106] SEQ ID NOs: 3681 to 3694 show the effector repeat motif of MG57 nuclease.

[0107] SEQ ID NOs: 3695-3696 show the nucleotide sequences of sgRNAs engineered to function with MG57 nuclease.

[0108] SEQ ID NOs: 3879 to 3880 show PAM sequences compatible with MG57 nuclease.

[0109] MG58 SEQ ID NOs: 722 to 779 show the full-length peptide sequences of MG58 nuclease.

[0110] SEQ ID NOs: 3697 to 3711 show the effector repeat motif of MG58 nuclease.

[0111] MG59 SEQ ID NOs: 780 to 792 show the full-length peptide sequences of MG59 nuclease.

[0112] SEQ ID NOs: 3712 to 3728 show the effector repeat motifs of MG59 nuclease.

[0113] SEQ ID NOs: 3729-3730 show the nucleotide sequences of sgRNAs engineered to function with MG59 nuclease.

[0114] SEQ ID NOs: 3881-3882 show PAM sequences compatible with MG59 nuclease.

[0115] MG60 SEQ ID NOs: 793 to 1163 show the full-length peptide sequences of MG60 nuclease.

[0116] SEQ ID NOs: 3731 to 3733 show the effector repeat motif of MG60 nuclease.

[0117] MG61 SEQ ID NOs: 1164 to 1469 show the full-length peptide sequences of MG61 nuclease.

[0118] SEQ ID NOs: 3734-3735 show crRNA 5' direct repeats designed to function with MG61 nuclease.

[0119] SEQ ID NOs: 3736 to 3847 show the effector repeat motif of MG61 nuclease.

[0120] MG62 SEQ ID NOs: 1470 to 1472 show the full-length peptide sequences of MG62 nuclease.

[0121] SEQ ID NOs: 3848 to 3850 show the effector repeat motif of MG62 nuclease.

[0122] MG70 SEQ ID NOs: 1473 to 1514 show the full-length peptide sequences of MG70 nuclease.

[0123] MG75 SEQ ID NOs: 1515 to 1710 show the full-length peptide sequences of MG75 nuclease.

[0124] MG77 SEQ ID NOs: 1711 to 1712 show the full-length peptide sequences of MG77 nuclease.

[0125] SEQ ID NOs: 3851-3852 show the nucleotide sequences of sgRNAs engineered to function with MG77 nuclease.

[0126] SEQ ID NOs: 3883 to 3884 show PAM sequences compatible with MG77 nuclease.

[0127] MG78 SEQ ID NOs: 1713 to 1717 show the full-length peptide sequences of MG78 nuclease.

[0128] SEQ ID NO: 3853 shows the nucleotide sequence of an sgRNA engineered to function with MG78 nuclease.

[0129] SEQ ID NO: 3885 shows a PAM sequence compatible with MG78 nuclease.

[0130] MG79 SEQ ID NOs: 1718 to 1722 show the full-length peptide sequences of MG79 nuclease.

[0131] SEQ ID NOs: 3854-3857 show the nucleotide sequences of sgRNAs engineered to function with MG79 nuclease.

[0132] SEQ ID NOs: 3886 to 3889 show PAM sequences compatible with MG79 nuclease.

[0133] MG80 SEQ ID NO: 1723 shows the full-length peptide sequence of MG80 nuclease.

[0134] MG81 SEQ ID NOs: 1724 to 2654 show the full-length peptide sequences of MG81 nuclease.

[0135] MG82 SEQ ID NOs: 2655 to 2657 show the full-length peptide sequences of MG82 nuclease.

[0136] MG83 SEQ ID NOs: 2658 to 2659 show the full-length peptide sequence of MG83 nuclease.

[0137] MG84 SEQ ID NOs: 2660 to 2677 show the full-length peptide sequences of MG84 nuclease.

[0138] MG85 SEQ ID NOs: 2678 to 2680 show the full-length peptide sequences of MG85 nuclease.

[0139] MG90 SEQ ID NOs: 2681 to 2809 show the full-length peptide sequences of MG90 nuclease.

[0140] MG91 SEQ ID NOs: 2810 to 3470 show the full-length peptide sequences of MG91 nuclease.

[0141] SEQ ID NOs: 6033-6036 show the nucleotide sequences of sgRNAs engineered to function with MG91 nuclease.

[0142] SEQ ID NOs: 6037 to 6039 show PAM sequences compatible with MG91 nuclease.

[0143] SEQ ID NOs: 6040-6049 show the MG91 intergenic region potentially encoding tracrRNA.

[0144] SEQ ID NOs: 6050-6059 show MG91 CRISPR repeats. Spacer Segment

[0145] SEQ ID NOs: 3858 to 3861 show the nucleotide sequences of the spacer segments.

[0146] NLS SEQ ID NOs: 3938-3953 show the sequences of examples of nuclear localization sequences (NLS) that can be added to nucleases according to the present disclosure.

[0147] CD38 targeting SEQ ID NOs: 4428-4465 and 5685 set forth the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target CD38.

[0148] SEQ ID NOs: 4466 to 4503 and 5686 show the DNA sequences of the CD38 target site.

[0149] TIGIT targeting SEQ ID NOs: 4504-4520 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target TIGIT.

[0150] SEQ ID NOs: 4521 to 4537 show the DNA sequences of the TIGIT target sites.

[0151] AAVS1 targeting SEQ ID NOs: 4538-4568 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target AAVS1.

[0152] SEQ ID NOs: 4569 to 4599 show the DNA sequences of the AAVS1 target sites.

[0153] B2M targeting SEQ ID NOs: 4600-4675 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target B2M.

[0154] SEQ ID NOs: 4676 to 4751 show the DNA sequences of the B2M target sites.

[0155] CD2 targeting SEQ ID NOs: 4752-4836 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target CD2.

[0156] SEQ ID NOs: 4837 to 4921 show the DNA sequences of the CD2 target sites.

[0157] CD5 targeting SEQ ID NOs: 4922-4945 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target CD5.

[0158] SEQ ID NOs: 4946 to 4969 show the DNA sequences of the CD5 target sites.

[0159] hRosa26 targeting SEQ ID NOs: 4970-5012 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target hRosa26.

[0160] SEQ ID NOs: 5013 to 5055 show the DNA sequences of the hRosa26 target sites.

[0161] TRAC targeting SEQ ID NOs: 5056-5125, 5681, and 5683 set forth the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target TRAC.

[0162] SEQ ID NOs: 5126-5195, 5682, and 5684 show the DNA sequences of the TRAC target sites.

[0163] TRBC1 targeting SEQ ID NOs: 5196-5210 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target TRBC1.

[0164] SEQ ID NOs: 5211 to 5225 show the DNA sequences of the TRBC1 target sites.

[0165] TRBC2 targeting SEQ ID NOs: 5226-5246 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target TRBC2.

[0166] SEQ ID NOs: 5247 to 5267 show the DNA sequences of the TRBC2 target sites.

[0167] TRBC1 / 2 targeting SEQ ID NOs: 5642-5660 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target TRBCs.

[0168] SEQ ID NOs: 5661 to 5679 show the DNA sequences of the TRBC target sites.

[0169] FAS targeting SEQ ID NOs: 5268-5366 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target FAS.

[0170] SEQ ID NOs: 5367 to 5465 show the DNA sequences of the FAS target sites.

[0171] PD-1 targeting SEQ ID NOs: 5466-5473 set forth the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target PD-1.

[0172] SEQ ID NOs: 5474 to 5481 show the DNA sequences of the PD-1 target sites.

[0173] HPRT targeting SEQ ID NOs: 5482-5561 set forth the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target HPRT.

[0174] SEQ ID NOs: 5562 to 5641 show the DNA sequences of the HPRT target sites.

[0175] HAO-1 targeting SEQ ID NOs: 5788-5829 and 5831-5834 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target human HAO-1.

[0176] SEQ ID NOs: 5836-5845 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target mouse HAO-1.

[0177] APO-A1 targeting SEQ ID NOs: 5847-5860 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target mouse APO-A1.

[0178] SEQ ID NOs: 5861 to 5874 show the DNA sequences of the APO-A1 target sites.

[0179] ANGPTL3 targeting SEQ ID NOs: 5875-5952 show the nucleotide sequences of sgRNAs engineered to function with MG29-1 nuclease to target mouse ANGPTL3.

[0180] SEQ ID NOs: 5953 to 6030 show the DNA sequences of the ANGPTL3 target sites. DETAILED DESCRIPTION OF THE INVENTION

[0181] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0182] The practice of some methods disclosed herein employs, unless otherwise indicated, techniques in immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, e.g., Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M.A.usubel, et al. eds.); the series Methods in Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J.MacPherson, B.D.Hames, and G.R.Taylor eds. (1995)), Harlow and Lane, eds. (1988), Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)) (incorporated herein by reference in their entireties).

[0183] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent that the terms "comprise," "include," "have," "have," "having," or variants thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0184] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.

[0185] As used herein, "cell" generally refers to a biological cell. A cell can be the basic structural, functional, and / or biological unit of a living organism. A cell can originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-celled eukaryotic cells, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, bryophytes, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.). C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from vertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), etc. In some cases, the cells are not derived from a naturally occurring organism (e.g., the cells may be synthetically produced and sometimes referred to as artificial cells).

[0186] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term "nucleotide" may refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as by using a moiety containing an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4′dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer, Foster City, Calif.; fluoro-conjugated deoxynucleotides, fluoro-conjugated Cy3-dCTP, fluoro-conjugated Cy5-dCTP, fluoro-conjugated fluoroX-dCTP, fluoro-conjugated Cy3-dUTP, and fluoro-conjugated Cy5-dUTP, available from Amersham, Arlington Heights, Ill.; and fluoro-conjugated Cy5-dUTP, available from Boehringer Ingelheim. Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2′-dATP available from Mannheim, Indianapolis, Ind.; and Molecular Examples of chromosomal labeled nucleotides include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, all available from Probes, Eugene, Oreg. Nucleotides may also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0187] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-, double-, or multiple-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and may perform any function. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, heterologous nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci (locuses) defined by binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides may be interrupted by non-nucleotide components.

[0188] The terms "transfection" or "transfected" generally refer to the introduction of nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a genetic sequence encoding an entire protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88 (hereby incorporated by reference in its entirety).

[0189] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein and generally refer to a polymer of at least two amino acid residues joined by a peptide bond. The term does not denote a specific length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary and / or tertiary structure (e.g., domains). The term also encompasses amino acid polymers modified by any other manipulation, such as disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" generally refer to natural and unnatural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids may include natural amino acids and unnatural amino acids, which are chemically modified to include non-naturally occurring groups or chemical moieties on the amino acid. Amino acid analogs may refer to amino acid derivatives. The term "amino acid" includes both D- and L-amino acids.

[0190] As used herein, "non-naturally occurring" can generally refer to a nucleic acid or polypeptide sequence that is not present in a naturally occurring nucleic acid or protein. Non-naturally occurring can refer to an affinity tag. Non-naturally occurring can refer to a fusion. Non-naturally occurring can refer to a naturally occurring nucleic acid or polypeptide sequence that includes mutations, insertions, and / or deletions. A non-naturally occurring sequence can exhibit and / or encode an activity (e.g., an enzymatic activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) that can also be exhibited by the nucleic acid and / or polypeptide sequence to which the non-naturally occurring sequence is fused. A non-naturally occurring nucleic acid or polypeptide sequence can be linked to a naturally occurring nucleic acid and / or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid and / or polypeptide sequence that encodes a chimeric nucleic acid or polypeptide.

[0191] As used herein, the term "promoter" generally refers to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping the nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often called transcription factors, which promote the binding of RNA polymerase to DNA, thereby resulting in gene transcription. A "basal promoter," also referred to as a "core promoter," may generally refer to a promoter that contains all the basic elements required to promote the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters may contain a TATA box and / or a CAAT box.

[0192] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression includes splicing of the mRNA in a eukaryotic cell.

[0193] As used herein, "operably linked," "operably linked," "operably linked," or grammatical equivalents thereof generally refer to the juxtaposition of genetic elements, e.g., promoters, enhancers, polyadenylation sequences, etc., where the elements are in a relationship permitting them to operate in an expected manner. For example, a regulatory element, which may include a promoter sequence and / or an enhancer sequence, is operably linked to a coding region if the regulatory element helps initiate transcription of the coding sequence. There can be intervening residues between the regulatory element and the coding region so long as this functional relationship is maintained.

[0194] As used herein, a "vector" generally refers to a polymer or an association of polymers that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally contains genetic elements, e.g., regulatory elements, operably linked to a gene to facilitate expression of the gene in a target.

[0195] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably to refer to a nucleic acid sequence or combination of elements operably linked for expression. In some cases, an expression cassette refers to a combination of a gene or genes with regulatory elements operably linked for expression.

[0196] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be the ability to affect expression in a manner attributable to the full-length sequence.

[0197] As used herein, an "engineered" subject generally indicates that the subject has been modified by human intervention. By way of non-limiting examples, a nucleic acid may be modified by altering its sequence to a sequence that does not occur in nature, a nucleic acid may be modified by ligating to a nucleic acid with which it is not naturally associated such that the ligated product has a function not present in the original nucleic acid, an engineered nucleic acid may be synthesized in vitro using a sequence that does not occur in nature, a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature, and an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.

[0198] As used herein, "synthetic" and "artificial" may generally be used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains.

[0199] As used herein, the term "Cas12a" generally refers to a family of Cas endonucleases that are class 2, type VA Cas endonucleases that (a) use relatively small guide RNAs (approximately 42-44 nucleotides) that are processed by the nuclease itself after transcription from the CRISPR array, and (b) cleave DNA to leave staggered cleavage sites. Further characteristics of this enzyme family can be found, for example, in Zetsche B, Heidenreich M, Mohanraju P, et al. Nat Biotechnol 2017;35:31-34, and Zetsche B, Gootenberg JS, Abudayyeh OO, et al. Cell 2015;163:759-771, which are incorporated herein by reference.

[0200] As used herein, "guide nucleic acid" can generally refer to a nucleic acid that can hybridize to another nucleic acid. A guide nucleic acid can be RNA. A guide nucleic acid can be DNA. A guide nucleic acid can be programmed to site-specifically bind to a nucleic acid sequence. The nucleic acid to be targeted, or the target nucleic acid, can comprise nucleotides. A guide nucleic acid can comprise nucleotides. A portion of the target nucleic acid can be complementary to a portion of the guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to and hybridizes with a guide nucleic acid can be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore not complementary to the guide nucleic acid can be referred to as a non-complementary strand. A guide nucleic acid can comprise a polynucleotide strand and can be referred to as a "single guide nucleic acid." A guide nucleic acid can comprise two polynucleotide strands and can be referred to as a "dual guide nucleic acid." Otherwise, the term "guide nucleic acid" can be inclusive, referring to both single and double guide nucleic acids. A guide nucleic acid may comprise a segment that may be referred to as a "nucleic acid targeting segment" or a "nucleic acid targeting sequence" or a "spacer sequence." The nucleic acid targeting segment may comprise a subsegment that may be referred to as a "protein binding segment" or a "protein binding sequence" or a "Cas protein binding segment."

[0201] The terms "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refer to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a certain percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and presence of 11, gap cost at extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of word length (W) of 2, expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (default parameters for BLASTP are available at BLAST, available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE using default parameters; MAFFT using parameters of retries of 2 and maximum repeats of 1,000; Novafold using default parameters; and HMMER hmmalign using default parameters.

[0202] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences aligned for maximum amino acid residue or nucleotide correspondence, as determined, for example, by the alignment producing the highest or "optimized" percent identity score.

[0203] The present disclosure also encompasses variants of any of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues (e.g., non-conserved residues) that vary between species without altering the basic function of the encoded protein. Such conservatively substituted variants have at least about 20% or less similarity to any one of the endonuclease protein sequences described herein (e.g., an MG11, MG13, MG26, MG28, MG29, MG30, MG31, MG32, MG37, MG53, MG54, MG55, MG56, MG57, MG58, MG59, MG60, MG61, MG62, MG70, MG82, MG83, MG84, or MG85 family endonuclease described herein, or any family nuclease described herein). Conservatively substituted variants may also include variants having at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity. In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions such that the activity of one or more critical active site residues or guide RNA binding residues of the endonuclease are not destroyed.In some embodiments, a functional variant of any of the proteins described herein lacks at least one of the conserved or functional residues called out in Figures 17, 18, 10, 20, or 25 or the residues listed in Table 1B. In some embodiments, a functional variant of any of the proteins described herein lacks all of the conserved or functional residues called out in Figures 17, 18, 10, 20, or 25 or the residues listed in Table 1B.

[0204] The present disclosure also includes variants (e.g., reduced activity variants) of any of the enzymes described herein that have substitutions of one or more catalytic residues to reduce or eliminate activity of the enzyme. In some embodiments, reduced activity variants of the proteins described herein include disruptive substitutions of at least one, at least two, or all three catalytic residues identified in Table 1B.

[0205] Conservative substitution tables providing functionally similar amino acids are available in various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993))). The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), Glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V), 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W), 7) Serine (S), Threonine (T), and 8) Cysteine ​​(C), methionine (M).

[0206] overview The discovery of new Cas enzymes with unique functionality and structure could further disrupt deoxyribonucleic acid (DNA) editing technologies, offering the potential to improve their speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microorganisms and the sheer diversity of microbial species, relatively few functionally characterized CRISPR / Cas enzymes exist in the literature. This is in part due to the inability to easily cultivate vast numbers of microbial species under laboratory conditions. Metagenomic sequencing from natural environmental niches containing large numbers of microbial species could dramatically increase the number of characterized new CRISPR / Cas systems and potentially expedite the discovery of novel oligonucleotide editing functions. A recent example of the fruitfulness of such an approach is demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.

[0207] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode RNA-based targeting elements; and (ii) an ORF encoding a Cas, which encodes a nuclease polypeptide directed by the RNA-based targeting element flanked by accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (PAM is typically a sequence not commonly represented in the host genome). Depending on the exact function and composition of the system, CRISPR-Cas systems are generally organized into two classes, five types, and 16 subtypes based on shared functional characteristics and evolutionary similarities (see figure).

[0208] Class I CRISPR-Cas systems have large, multi-subunit effector complexes and include types I, III, and IV. Class II CRISPR-Cas systems generally have single-polypeptide, multi-domain nuclease effectors and include types II, V, and VI.

[0209] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, processing of the CRISPR array into mature crRNA does not require the presence of a specialized endonuclease subunit, but rather a small transcoding crRNA (tracrRNA) with a region complementary to the array repeat sequence. The tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate a mature effector enzyme loaded with both the tracrRNA and crRNA. Cas II nuclease is specified as a DNA nuclease. Type II effectors generally exhibit a structure containing a RuvC-like endonuclease domain that fits into an RNase H fold with an unrelated HNH nuclease domain inserted into the RuvC-like nuclease domain fold. The RuvC-like domain is responsible for cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleavage of the displacement DNA strand.

[0210] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g., Cas12) structure similar to that of type II effectors, including a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use a tracrRNA to process pre-crRNA into mature crRNA. However, unlike type II systems, which require RNAse III to cleave the pre-crRNA into multiple crRNAs, type V systems can cleave the pre-crRNA using the effector nuclease itself. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems have also been identified as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to have robust single-strand nonspecific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.

[0211] CRISPR-Cas systems have recently emerged as a gene editing technology due to their targeting and ease of use. The most commonly used systems are class 2, type II SpCas9 and class 2, type VA Cas12a. Specifically, type VA systems have become more widely used due to their reported higher specificity in cells than other nucleases and fewer or no off-target effects. VA systems are also advantageous in that the guide RNA is small (42-44 nucleotides compared to approximately 100 nt for SpCas9) and is processed by the nuclease itself after transcription from the CRISPR array, simplifying multiplexed applications involving multiple gene editing. Furthermore, VA systems have staggered cut sites, which may facilitate directed repair pathways such as microhomology-mediated targeted integration (MITI).

[0212] The most commonly used type VA enzymes require a 5' protospacer adjacent motif (PAM) next to the selected target site: 5'-TTTV-3' for Lachnospiraceae bacterium ND2006 LbCas12a and Acidaminococcus species AsCas12a; and 5'-TTV-3' for Francisella novicida FnCas12a. Recent surveys of orthologs have revealed proteins with less restrictive PAM sequences, such as YTV, YYN, or TTN, that are also active in mammalian cell culture. However, these enzymes do not fully encompass the biodiversity and targeting of VA enzymes and may not represent all possible activities and PAM sequence requirements. Here, we extracted thousands of genome fragments from multiple metagenomes for type VA nucleases. The diversity of identified VA enzymes may be expanded, and novel systems may be developed into highly targetable, compact, and precise gene editing agents.

[0213] MG enzyme Type V CRISPR systems are rapidly being adopted for use in a variety of genome editing applications. These programmable nucleases are part of adaptive microbial immune systems, and their natural diversity has remained largely unexplored. A novel family of type V CRISPR enzymes was identified through large-scale analysis of metagenomes collected from various complex environments, and representatives of these systems were developed into gene editing platforms. The nucleases are phylogenetically diverse (Figure 4A) and recognize single guide RNAs with specific motifs. The majority of these systems are derived from uncultivated organisms, some of which encode divergent type V effectors within the same CRISPR operon. Biochemical analysis revealed unexpected PAM diversity (see Figure 4B), indicating that these systems will facilitate a variety of genome engineering applications. The simplicity of the guide sequences and activity in human cell lines suggests their utility in gene and cell therapy.

[0214] In some embodiments, the present disclosure provides novel VL-type candidates (Figures 27A and B). The VL-types may be novel subtypes, and some subfamilies may have been identified. These nucleases are approximately 1000-1100 amino acids in length. The VL-types may be found in the same CRISPR loci as VA-type effectors. RuvC catalytic residues may have been identified for the VL-type candidates, and these VL-type candidates may not require tracrRNA. An example of a VL-type is the MG60 nuclease described herein (see Figures 28 and 32).

[0215] In some embodiments, the present disclosure provides smaller V-shaped effectors (see Figure 30). Such effectors may be small putative effectors. These effectors may simplify delivery and broaden therapeutic applications.

[0216] In some embodiments, the present disclosure provides novel V-type effectors. Such effectors can be MG70, as described herein (FIG. 29). MG70 can be an ultra-small enzyme approximately 373 amino acids in length. MG70 can have a single N-terminal transportase domain and a predicted tracrRNA (see FIGS. 30 and 32).

[0217] In some embodiments, the present disclosure provides smaller V-type effectors (see Figure 31). Such effectors can be MG81, as described herein. MG81 can be approximately 500-700 amino acids in length and can contain RuvC and HTH DNA-binding domains.

[0218] In one aspect, the present disclosure provides engineered nuclease systems discovered through metagenomics sequencing. In some cases, metagenomics sequencing is performed on a sample. In some cases, the sample can be collected from a variety of environments. Such environments can be human microbiomes, animal microbiomes, high temperature environments, low temperature environments. Such environments can include sediments. Examples of the types of environments for the engineered nuclease systems described herein can be found in the figures.

[0219] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a Class 2, Type V Cas endonuclease. In some cases, the endonuclease is a Class 2, Type VA Cas endonuclease. In some cases, the endonuclease is derived from an uncultured microorganism. The endonuclease may comprise a RuvC domain. In some cases, the engineered nuclease system comprises: (b) an engineered guide RNA. In some cases, the engineered guide RNA is configured to form a complex with the endonuclease. In some cases, the engineered guide RNA comprises a spacer sequence. In some cases, the spacer sequence is configured to hybridize to a target nucleic acid sequence.

[0220] In one aspect, the disclosure provides (a) an engineered nuclease system comprising an endonuclease. In some embodiments, the endonuclease has at least 70% sequence identity to any one of SEQ ID NOs: 1-3470. In some cases, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-3470.

[0221] In some cases, the endonuclease comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-3470. In some cases, the endonuclease can be substantially identical to any one of SEQ ID NOs: 1-3470.

[0222] In some cases, the engineered nuclease system comprises an engineered guide RNA. In some cases, the engineered guide RNA is configured to form a complex with an endonuclease. In some cases, the engineered guide RNA comprises a spacer sequence. In some cases, the spacer sequence is configured to hybridize to a target nucleic acid sequence.

[0223] In one aspect, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease. In some cases, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence. In some cases, the PAM sequence is substantially identical to any one of SEQ ID NOs: 3863-3913. In some cases, the PAM sequence is any one of SEQ ID NOs: 3863-3913. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a Class 2 Cas endonuclease. In some cases, the endonuclease is a Class 2, Type V Cas endonuclease. In some cases, the endonuclease is a Class 2, Type VA Cas endonuclease. In some cases, the engineered nuclease system comprises: (b) an engineered guide RNA. In some cases, the engineered guide RNA is configured to form a complex with the endonuclease. In some cases, the engineered guide RNA comprises a spacer sequence. In some cases, the spacer sequence is configured to hybridize to a target nucleic acid sequence.

[0224] In some cases, the endonuclease is not a Cpf1 or Cms1 endonuclease. In some cases, the endonuclease further comprises a zinc finger-like domain.

[0225] In some cases, the guide RNA comprises the first 19 nucleotides or a sequence having at least 80% sequence identity to non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, or 3851-3857. In some cases, the guide RNA has at least about 20% or at least about 20% identical non-degenerate nucleotides in the first 19 nucleotides or sequences of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, or 3851-3857. Also includes sequences having about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity.In some cases, the guide RNA has at least about 20% or at least about 20% non-degenerate nucleotides in the first 19 nucleotides or sequences of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, or 3851-3857. Includes variants having about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity. In some cases, the guide RNA comprises the first 19 nucleotides or a sequence substantially identical to the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, or 3851-3857.

[0226] In some cases, the guide RNA has at least about 20% or at least about 20% identical non-degenerate nucleotides in the first 19 nucleotides or sequences of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, or 3851-3857. or a sequence having at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity. In some cases, the endonuclease is configured to bind to an engineered guide RNA. In some cases, the Cas endonuclease is configured to bind to an engineered guide RNA. In some cases, the Class 2, Cas endonuclease is configured to bind to an engineered guide RNA. In some cases, the Class 2, V-type Cas endonuclease is configured to bind to an engineered guide RNA. In some cases, a class 2, VA-type Cas endonuclease is configured to bind to an engineered guide RNA.

[0227] In some cases, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence comprising any one of SEQ ID NOs: 3863-3913.

[0228] In some cases, the guide RNA comprises a sequence that is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a eukaryotic genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a fungal genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a plant genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a mammalian genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a human genomic polynucleotide sequence.

[0229] In some cases, the guide RNA is 30-250 nucleotides in length. In some cases, the guide RNA is 42-44 nucleotides in length. In some cases, the guide RNA is 42 nucleotides in length. In some cases, the guide RNA is 43 nucleotides in length. In some cases, the guide RNA is 44 nucleotides in length. In some cases, the guide RNA is 85-245 nucleotides in length. In some cases, the guide RNA is more than 90 nucleotides in length. In some cases, the guide RNA is less than 245 nucleotides in length.

[0230] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus of any one of SEQ ID NOs: 3938-3953 or a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3938-3953. In some cases, the NLS can comprise a sequence substantially identical to any one of SEQ ID NOs: 3938-3953.

[0231] [Table 1]

[0232] In some cases, the engineered nuclease system further comprises a single-stranded or double-stranded DNA repair template. In some cases, the engineered nuclease system further comprises a single-stranded DNA repair template. In some cases, the engineered nuclease system further comprises a double-stranded DNA repair template. In some cases, the single-stranded or double-stranded DNA repair template can comprise, from 5' to 3', a first homologous arm comprising a sequence of at least 20 nucleotides 5' to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence.

[0233] In some cases, the first homology arm comprises a sequence of at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides. In some cases, the second homology arm comprises a sequence of at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides.

[0234] In some cases, the first and second homologous arms are homologous to a prokaryotic genomic sequence. In some cases, the first and second homologous arms are homologous to a bacterial genomic sequence. In some cases, the first and second homologous arms are homologous to a fungal genomic sequence. In some cases, the first and second homologous arms are homologous to a eukaryotic genomic sequence.

[0235] In some cases, the engineered nuclease system further comprises a DNA repair template. The DNA repair template may comprise a double-stranded DNA segment. The double-stranded DNA segment may be adjacent to one single-stranded DNA segment. The double-stranded DNA segment may be adjacent to two single-stranded DNA segments. In some cases, the single-stranded DNA segment is conjugated to the 5' end of the double-stranded DNA segment. In some cases, the single-stranded DNA segment is conjugated to the 3' end of the double-stranded DNA segment.

[0236] In some cases, a single-stranded DNA segment has a length of 1 to 15 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 4 to 10 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 4 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 5 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 6 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 7 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 8 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 9 nucleotide bases. In some cases, a single-stranded DNA segment has a length of 10 nucleotide bases.

[0237] In some cases, the single-stranded DNA segment has a nucleotide sequence that is complementary to a sequence within the spacer sequence. In some cases, the double-stranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein-coding sequence, an miRNA-coding sequence, an RNA-coding sequence, or a transgene.

[0238] In some cases, the engineered nuclease system 2+ The present invention further comprises a source of

[0239] In some cases, the guide RNA comprises a hairpin comprising at least 8 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin comprising at least 9 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin comprising at least 10 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin comprising at least 11 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin comprising at least 12 base pair ribonucleotides.

[0240] In some cases, the endonuclease comprises a sequence that is at least 70% identical to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some cases, the endonuclease comprises a sequence that is at least 75% identical to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some cases, the endonuclease comprises a sequence that is at least 80% identical to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some cases, the endonuclease comprises a sequence that is at least 85% identical to a variant of any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some cases, the endonuclease comprises a sequence that is at least 90% identical to a variant of any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof. In some cases, the endonuclease comprises a sequence that is at least 95% identical to a variant of any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1721, or a variant thereof.

[0241] In some cases, the guide RNA structure comprises a sequence at least 70% identical to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 3608. In some cases, the guide RNA structure comprises a sequence at least 75% identical to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 3608. In some cases, the guide RNA structure comprises a sequence at least 80% identical to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 3608. In some cases, the guide RNA structure comprises a sequence at least 85% identical to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 3608. In some cases, the guide RNA structure comprises a sequence at least 90% identical to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 3608. In some cases, the guide RNA structure comprises a sequence at least 95% identical to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 3608. In some cases, the endonuclease is configured to bind to a PAM comprising any one of SEQ ID NOs: 3863-3913.

[0242] In some cases, sequences may be determined by the BLASTP, CLUSTALW, MUSCLE, or MAFFT algorithm, or the CLUSTALW algorithm using Smith-Waterman homology search algorithm parameters. Sequence identity may be determined by the BLASTP homology search algorithm using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at 11 presence and 1 extension, with a conditional composition score matrix adjustment.

[0243] In one aspect, the present disclosure provides an engineered guide RNA comprising (a) a DNA-targeting segment. In some cases, the DNA-targeting segment comprises a nucleotide sequence complementary to a target sequence. In some cases, the target sequence is in a target DNA molecule. In some cases, the engineered guide RNA comprises (b) a protein-binding segment. In some cases, the protein-binding segment comprises two complementary stretches of nucleotides. In some cases, the two complementary stretches of nucleotides hybridize to form a double-stranded RNA (dsRNA) duplex. In some cases, the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide. In some cases, the engineered guide ribonucleic acid polynucleotide can form a complex with an endonuclease. In some cases, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-3470. In some cases, the complex targets a target sequence of a target DNA molecule.

[0244] In some embodiments, the DNA-targeting segment is positioned 3' to both of the two complementary stretches of nucleotides. In some cases, the protein-binding segment comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 3608.

[0245] In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 8 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 9 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 10 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 11 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 12 ribonucleotides.

[0246] In some cases, the deoxyribonucleic acid polynucleotide encodes an engineered guide ribonucleic acid polynucleotide.

[0247] In one aspect, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence. In some cases, the engineered nucleic acid sequence is optimized for expression in an organism. In some cases, the nucleic acid encodes an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a Class 2 endonuclease. In some cases, the endonuclease is a Class 2, Type V Cas endonuclease. In some cases, the endonuclease is a Class 2, Type VA Cas endonuclease. In some cases, the endonuclease is derived from an uncultured microorganism. In some cases, the organism is not an uncultured organism.

[0248] In some cases, the endonuclease comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-3470.

[0249] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLSs). The NLS may be located proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus of any one of SEQ ID NOs: 3938-3953 or a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 3938-3953.

[0250] In some cases, the organism is a prokaryote. In some cases, the organism is a bacterium. In some cases, the organism is a eukaryote. In some cases, the organism is a fungus. In some cases, the organism is a plant. In some cases, the organism is a mammal. In some cases, the organism is a rodent. In some cases, the organism is a human.

[0251] In one aspect, the disclosure provides an engineered vector. In some cases, the engineered vector comprises a nucleic acid sequence encoding an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a Class 2 Cas endonuclease. In some cases, the endonuclease is a Class 2, Type V Cas endonuclease. In some cases, the endonuclease is a Class 2, Type VA Cas endonuclease. In some cases, the endonuclease is derived from an uncultured microorganism.

[0252] In some cases, the engineered vector comprises a nucleic acid described herein. In some cases, the nucleic acid described herein is a deoxyribonucleic acid polynucleotide described herein. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.

[0253] In one aspect, the present disclosure provides a cell comprising a vector described herein.

[0254] In one aspect, the present disclosure provides a method for producing an endonuclease. In some cases, the method includes culturing a cell.

[0255] In one aspect, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide. The method can include contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a Class 2 Cas endonuclease. In some cases, the endonuclease is a Class 2, Type V Cas endonuclease. In some cases, the endonuclease is a Class 2, Type VA Cas endonuclease. In some cases, the endonuclease is complexed with an engineered guide RNA. In some cases, the engineered guide RNA is configured to bind to the endonuclease. In some cases, the engineered guide RNA is configured to bind to the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the engineered guide RNA is configured to bind to the endonuclease and to the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM). In some cases, the PAM sequence is a sequence comprising any one of SEQ ID NOs: 3863-3913.

[0256] In some cases, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of an engineered guide RNA and a second strand comprising a PAM. In some cases, the PAM is immediately adjacent to the 5' end of the sequence complementary to a sequence of an engineered guide RNA. In some cases, the endonuclease is not Cpf1 endonuclease or Cms1 endonuclease. In some cases, the endonuclease is derived from an uncultured microorganism. In some cases, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some cases, the PAM comprises any one of SEQ ID NOs: 3863-3913.

[0257] In one aspect, the present disclosure provides a method for modifying a target nucleic acid locus. The method can include delivering an engineered nuclease system described herein to the target nucleic acid locus. In some cases, the endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure. In some cases, the complex is configured such that, upon binding of the complex to the target nucleic acid locus, the complex modifies the target nucleic acid locus.

[0258] In some cases, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some cases, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some cases, the target nucleic acid locus is in vitro. In some cases, the target nucleic acid locus is in a cell. In some cases, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a primary cell, or a human cell.

[0259] In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some cases, the nucleic acid comprises a promoter. In some cases, the open reading frame encoding the endonuclease is operably linked to a promoter.

[0260] In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA comprising an open reading frame encoding the endonuclease. In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide. In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide RNA operably linked to a ribonucleic acid (RNA) pol III promoter.

[0261] In some cases, the endonuclease induces a single-stranded or double-stranded break at or near the target locus, hi some cases, the endonuclease induces a staggered single-stranded break within the target locus or 3' to the target locus.

[0262] In some cases, the effector repeat motif is used to inform the guide design of the MG nuclease. For example, the processed gRNA in a type VA system includes the last 20-22 nucleotides of the CRISPR repeat. This sequence can be synthesized into a crRNA (along with a spacer) and tested in vitro with the synthetic nuclease for cleavage on a library of potential targets. Using this method, the PAM can be determined. In some cases, type VA enzymes may use a "universal" gRNA. In some cases, type V enzymes may utilize unique gRNAs.

[0263] lipid nanoparticles The lipid nanoparticles described herein may be four-component lipid nanoparticles. These nanoparticles may be configured for delivery of RNA or other nucleic acids (e.g., synthetic RNA, mRNA, or in vitro synthesized mRNA) and may be formulated generally as described in WO2012 / 135805(A2), which is incorporated herein by reference for all purposes. These nanoparticles may generally comprise (a) a cationic lipid (e.g., any of the lipids described in Figure 109), (b) a neutral lipid (e.g., DSPC or DOPE), (c) a sterol (e.g., cholesterol or a cholesterol analog), and (d) a PEG-modified lipid (e.g., PEG-DMG).

[0264] The cationic lipid referred to herein as "C12-200" is disclosed by Love et al., Proc Natl Acad Sci USA. 2010 107:1864-1869 and Liu and Huang, Molecular Therapy. 2010 669-670, both of which are incorporated herein by reference in their entirety. Cationic lipid formulations may include particles containing any three, four, or more components in addition to polynucleotides, primary constructs, or RNA (e.g., mRNA). By way of example, a formulation containing a particular cationic lipid, including, but not limited to, 98N12-5 (or any of the other structures depicted in Figure 109), may contain 42% lipidoid, 48% cholesterol, and 10% PEG (alkyl chain length of C14 or greater). As another example, a formulation with certain lipidoids, including but not limited to C12-200, may contain 50% cationic lipid, 10% disteroylphosphatidylcholine, 38.5% cholesterol, and 1.5% PEG-DMG.

[0265] In some embodiments, the lipid nanoparticles are formulated as described in US 10709779 (B2), which is incorporated herein by reference in its entirety. In some embodiments, the cationic lipid nanoparticles comprise a cationic lipid, a PEG-modified lipid, a sterol, and a non-cationic lipid. In some embodiments, the cationic lipid is selected from the group consisting of any of the cationic lipids shown in Figure 109. In some embodiments, the cationic lipid nanoparticles have a molar ratio of about 20-60% cationic lipid, about 5-25% non-cationic lipid, about 25-55% sterol, and about 0.5-15% PEG-modified lipid. In some embodiments, the cationic lipid nanoparticles comprise a molar ratio of about 50% cationic lipid, about 1.5% PEG-modified lipid, about 38.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid nanoparticles comprise a molar ratio of about 55% cationic lipid, about 2.5% PEG-modified lipid, about 32.5% cholesterol, and about 10% non-cationic lipid. In some embodiments, the cationic lipid is an ionic cationic lipid, the non-cationic lipid is a neutral lipid, and the sterol is cholesterol. In some embodiments, the cationic lipid nanoparticles have a molar ratio of cationic lipid:cholesterol:PEG2000-DMG:DSPC or DMG:DOPE of 50:38.5:10:1.5. In some embodiments, the lipid nanoparticles described herein can comprise cholesterol, 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), 1,1'-((2-(4-(2-((2-(bis(2-hydroxydodecyl)amino)ethyl)(2-hydroxydodecyl)amino)ethyl)piperazin-1-yl)ethyl)azanediyl)bis(dodecan-2-ol) (C12-200), and DMG-PEG-2000 in a molar ratio of 47.5:16:35:1.5.

[0266] The disclosed system can be used for a variety of applications, such as binding to nucleic acid molecules (e.g., sequence-specific binding), nucleic acid editing (e.g., gene editing), etc. Such systems can be used, for example, to address (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivated enzymes combined with probes to target specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to prevent them from infecting host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and as biosensors to detect cellular perturbations by exogenous small molecules and nucleotides. [Example]

[0267] Example 1 - New protein metagenomics analysis method Metagenomics samples were collected from sediment, soil, and animals. Deoxyribonucleic acid (DNA) was extracted using a Zymobiomics DNA miniprep kit and sequenced on an Illumina HiSeq® 2500. Samples were collected with the owner's consent. A hidden Markov model generated based on identified Cas protein sequences, including class II and type V Cas effector proteins, was used to search metagenomics sequence data to identify novel Cas effectors (see the figure, which shows the distribution of proteins detected in one family, MG29, identified from sample types such as high-temperature samples). Novel effector proteins identified by the search were aligned to the identified proteins to identify potential active sites (see, for example, the figure, which shows that all MG29 family effectors identified from various samples possess three catalytic residues from the RuvCI, RuvCII, and RuvCIII catalytic domains and are predicted to be active). This metagenomics workflow resulted in the description of the following families: 3471, 3539, 3551-355, MG11, MG13, MG19, MG20, MG26, MG28, MG29, MG30, MG31, MG32, MG37, MG9, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, and MG91. Putative spacer sequences were identified by their location adjacent to genomic loci encoding effector proteins.

[0268] Example 2 - Methods for metagenomic analysis of novel proteins Thirteen animal microbiome, high-temperature biofilm, and sediment samples were collected and stored on ice or in Zymo DNA / RNA Shield after collection. DNA was extracted from samples using either the Qiagen DNeasy PowerSoil Kit or the ZymoBIOMICS DNA Miniprep Kit. DNA sequencing libraries were constructed and sequenced on an Illumina HiSeq 4000 or Novaseq machine at the Vincent J. Coates Genomics Sequencing Laboratory (UC Berkeley) using paired 150-bp reads with a target insert size of 400–800 bp (targeting 10 GB of sequencing per sample). Publicly available metagenomics sequencing data was downloaded from NCBI SRA. Sequencing reads were trimmed and assembled with Megahit 11 using BBMap (Bushnell B., sourceforge.net / projects / bbmap / ). Open reading frames and protein sequences were predicted using Prodigal. HMM profiles of identified VA-type CRISPR nucleases were constructed and searched against all predicted proteins using HMMER3 (hmmer.org) to identify potential effectors. CRISPR arrays on assembled contigs were predicted using Minced (https: / / github.com / ctSkennerton / minced). Taxonomy was assigned to proteins with Kaiju, and contig taxonomy was determined by finding a consensus of all encoded proteins.

[0269] Predicted and reference (e.g., LbCas12a, AsCas12a, FnCas12a) type V effector proteins were aligned with MAFFT, and phylogenetic trees were inferred using FasTree2. Novel families were delineated by identifying clades composed of sequences recovered from this study. Within families, candidates were selected if they contained components for laboratory analysis (i.e., they were found in fully assembled and annotated contigs using CRISPR arrays) in a manner that sampled as much phylogenetic diversity as possible. Small effectors from diverse families (i.e., families with representatives that share a broader range of protein sequences) were prioritized. Selected representative sequences and reference sequences were aligned using MUSCLE and Clustal W to identify catalytic and PAM-interacting residues. CRISPR array repeats were searched for motifs related to type V-A systems, TCTAC-N-GTAGA (containing 1 to 8 N residues). From this analysis, if a representative CRISPR array contained one of these motif sequences, the family was putatively classified as VA. Using this dataset, we identified HMM profiles associated with the VA family, which were then used to classify additional families (Figures 33-37). Conventional wisdom is to name novel Cas12 nucleases based on the organisms that encode them, but this is not possible for the nucleases described herein. Therefore, to best adhere to convention, the systems described herein are named with the prefix MG to indicate that they are derived from assembled metagenomics fragments.

[0270] We mined 140,867 Mbp of assembled metagenomics sequencing data from diverse environments (soil, thermophiles, sediments, human and non-human microbiomes). A total of 119 genomic fragments encoded CRISPR effectors distally associated with type VA nucleases flanking the CRISPR array (Figure 4B). The type VA effectors were classified into 14 novel families that shared an average pairwise amino acid identity of less than 30% with each other and with reference sequences (e.g., LbCas12a, AsCas12a, FnCas12a). Some effectors contained conserved DED nuclease catalytic residues from the RuvC and alpha-helical recognition domains, as well as the RuvCI / CII / CIII domains (identified in multiple sequence alignments; see, for example, Table 1A below), suggesting that these effectors were active nucleases (Figures 5-7). The novel VA-type nucleases range in size from <800 to 1,400 amino acids in length (see Figure 5A), and their classification spans a diverse array of phyla (Figure 4A), suggesting the possibility of horizontal transfer.

[0271] Some genomic fragments carrying the VA-type CRISPR system also encoded a second effector, termed the VA-type prime (V-A', Figure 7A). For example, VA-type MG26-1, which shares 16.6% amino acid identity with VA-type MG26-2, may be encoded by the same CRISPR Cas operon and share the same crRNA with MG26-1 (Figure 7B). Although no nuclease domain was predicted, MG26-2 contained three RuvC catalytic residues identified from multiple sequence alignments (Figure 7B).

[0272] [Table 2]

[0273] Example 3 - (General Protocol) Identification / Confirmation of PAM Sequence PAM sequences that can be cleaved in vitro by CRISPR effectors were identified by incubating the effectors with crRNA and a plasmid library containing eight randomized nucleotides located adjacent to the 5' end of a sequence complementary to the spacer of the crRNA. The plasmids were configured so that the plasmids were cleaved when the eight randomized nucleotides formed a functional PAM sequence. Adapters were then ligated to the ends of the cleaved plasmids, and functional PAM sequences were identified by sequencing the adapter-containing DNA fragments. Putative endonucleases were expressed in an E. coli lysate-based expression system (myTXTL, Arbor Biosciences). E. coli codon-optimized nucleotide sequences encoding the putative nucleases were transcribed and translated in vitro from PCR fragments under the control of a T7 promoter. A second PCR fragment containing a minimal CRISPR array consisting of a T7 promoter followed by a repeat-spacer repeat sequence was transcribed in the same reaction. Successful expression of the endonuclease sequence and repeat-spacer repeat sequence, followed by CRISPR array processing, yielded active in vitro CRISPR nuclease complexes.

[0274] A library of target plasmids containing a spacer sequence matching that in the minimal array, preceded by 8N (degenerate) bases (the putative PAM sequence), was incubated with the products of the TXTL reaction. After 1–3 hours, the reaction was stopped, and the DNA was recovered using a DNA cleanup kit, such as Zymo DCC, AMPure XP beads, or QiaQuick. Adapter sequences were blunt-end ligated to DNA fragments containing active PAM sequences cleaved by the endonuclease; uncleaved DNA was inaccessible to ligation. DNA segments containing active PAM sequences were then amplified by PCR using primers specific for the library and adapter sequences. PCR amplification products were resolved on a gel to identify amplicons corresponding to cleavage events. Amplified segments from the cleavage reaction were also used as templates for NGS library preparation or as substrates for Sanger sequencing. Sequencing of this resulting library, a subset of the starting 8N library, revealed sequences with PAM activity compatible with CRISPR complexes. For PAM studies using treated RNA constructs, the same procedure was repeated except that in vitro transcribed RNA was added along with the plasmid library and the minimal CRISPR array template was omitted. The following sequences were used as targets in these assays: CGTGAGCCACCACGTCGCAAGCCT (SEQ ID NO: 3860); GTCGAGGCTTGCGACGTGGTGGCT (SEQ ID NO: 3861); GTCGAGGCTTGCGACGTGGTGGCT (SEQ ID NO: 3858); and TGGAGATATCTTGAACCTTGCATC (SEQ ID NO: 3859).

[0275] Example 4 - PAM sequence identification / confirmation of the endonucleases described herein PAM requirements were determined with modifications via an E. coli lysate-based expression system (myTXTL, Arbor Biosciences). Briefly, E. coli codon-optimized effector protein sequences were expressed under the control of a T7 promoter for 16 hours at 29°C. This crude protein stock was then used in an in vitro digestion reaction at a concentration of 20% of the total reaction volume. The reaction was incubated for 3 hours at 37°C with 5 nM of a plasmid library containing a constant target sequence preceded by 8N mixed bases and 50 nM of in vitro-transcribed crRNA derived from the same CRISPR locus as the effector linked to a sequence complementary to the target sequence in NEB Buffer 2.1 (New England Biolabs; NEB Buffer 2.1 was chosen to compare candidates with commercially available proteins). Protein concentrations were not normalized in the PAM discovery assay (PCR amplification signals provide high sensitivity to low expression or activity). Cleavage products from the TXTL reaction were recovered via cleanup using AMPure SPRI beads (Beckman Coulter). The DNA was blunted via the addition of Klenow fragment and dNTPs (New England Biolabs). The blunt-ended product was ligated with a 100-fold excess of double-stranded adapter sequences and used as a template for the preparation of an NGS library, from which the PAM requirements were determined by sequence analysis.

[0276] Raw NGS reads were filtered by Phred quality scores >20. Using 28 bp representing the DNA sequence identified from the scaffold adjacent to the PAM as a reference, the PAM-proximal region was located, and the adjacent 8 bp were identified as putative PAMs. The distance between the PAM and the ligated adapter was also measured for each read. Reads that did not exactly match the reference sequence or adapter sequence were excluded. PAM sequences were filtered by cleavage site frequency so that PAMs with ±2 bp of the most frequent cleavage site were included in the analysis. This correction removed low-level background cleavage that could occur at random positions due to the use of crude E. coli lysates. This filtering step can remove 2%–40% of reads, depending on the signal-to-noise ratio of the candidate protein; less active proteins have more background signal. For reference MG29-1, 2% of reads were filtered at this stage. The filtered list of PAMs was used to generate sequence logos using Logomaker. Depictions of these sequence logos for PAMs are shown in Figures 20–24.

[0277] Example 5 - tracrRNA prediction and guide design The crystal structure of the ternary complex of AacC2c1 (Cas12b) bound to sgRNA and target DNA reveals two distinct repeat-anti-repeat (R-AR) motifs in the bound sgRNA, designated R-AR duplex 1 and R-AR duplex 2 (see Figures 8 and 9 herein, and Yang, Hui, Pu Gao, Kanagalaghatta R. Rajashankar, and Dinshaw J. Patel. 2016. "PAM-Dependent Target DNA Recognition and Cleavage by C2c1 CRISPR-Cas Endonuclease." Cell 167(7):1814-28. e12 and Liu, Liang, Peng Chen, Min Wang, Xueyan Li, Jiuyu Wang, Maolu Yin, and Yanli Wang. 2017. "C2c1-sgRNA Complex Structure Reveals RNA-Guided DNA Cleavage Mechanism." Molecular Cell 65(2):310-22, each of which is incorporated herein by reference in its entirety. The putative tracrRNA sequences for the CRISPR effectors disclosed herein were identified by searching for anti-repeat sequences in the genomic context surrounding natural CRISPR arrays, with the R-AR duplex 2 anti-repeat sequence occurring approximately 20-90 nucleotides upstream of (closer to the 5' end of) the R-AR duplex 1 anti-repeat sequence. After identifying the tracrRNA sequences, two guide sequences were designed for each enzyme. The first included both R-AR duplex 1 and 2 (see, e.g., SEQ ID NOs: 3636, 3640, 3644, 3648, 3652, 3656, 3660, 3671, and 3672), and the second was a short guide sequence that deleted the R-AR duplex 1 region because this region may not be essential for cleavage (see, e.g., SEQ ID NOs: 3637, 3641, 3645, 3649, 3653, 3657, and 3661).

[0278] Example 6 - Predicted RNA folding protocol The predicted RNA folding of the RNA sequences at 37° C. was calculated using the method of Andronescu 2007, which is incorporated herein by reference in its entirety.

[0279] Example 7 - Identification of RNA guides For contigs encoding type VA effectors and CRISPR arrays, secondary structure folding of the repeats indicated that the novel type VA system required a single guide crRNA (sgRNA, Figure 10A-D). The tracrRNA sequence was not identified. The sgRNA contained approximately 19-22 nt from the 3' end of the CRISPR repeat. Multiple sequence alignment of CRISPR repeats from six of the type VA candidates tested for in vitro activity revealed a highly conserved motif at the 3' end of the repeat, which formed a stem-loop structure for the sgRNA (Figure 10C). The motif, UCUAC[N3-5]GUAGAU, contained short palindromic repeats (stem) separated by 3-5 nucleotides (loop).

[0280] Conservation of sgRNA motifs was used to reveal novel effectors that may not show similarity to classified VA-type nucleases. Motifs were searched for in repeats from 69,117 CRISPR arrays. The most common motif contained a 4-nucleotide loop, while 3- and 5-nucleotide loops were less common (Figures 12, 13, 14, 15, and 16). Examination of the genomic context surrounding CRISPR arrays containing repeat motifs revealed numerous effectors of varying lengths. For example, effectors from family MG57 were the largest of the identified VA-type nucleases (average approximately 1400 aa) and encoded repeats with a 4-bp loop. Another family identified from HMM analysis contained a distinct repeat motif, CCUGC[N 3-4 ]GCAGG (see Figures 5C and 5D). Although the sequences differ, the structures were predicted to fold into very similar stem-loop structures.

[0281] Example 8 - In vitro cleavage efficiency of MG CRISPR complexes The endonuclease is expressed as a His-tagged fusion protein from an inducible T7 promoter in a protease-deficient E. coli B strain. Cells expressing the His-tagged protein are lysed by sonication, and the His-tagged protein is purified by Ni-NTA affinity chromatography on a HisTrap FF column (GE Lifescience) on an AKTA Avant FPLC (GE Lifescience). The eluate is resolved by SDS-PAGE on an acrylamide gel (Bio-Rad) and stained with InstantBlue Ultrafast coomassie (Sigma-Aldrich). Purity is determined using densitometry of the protein bands with ImageLab software (Bio-Rad). The purified endonuclease is dialyzed into a storage buffer consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol, pH 7.5, and stored at -80°C. Target DNA containing a pacer sequence and a PAM sequence (e.g., as determined in either Example 3 or Example 4) is constructed by DNA synthesis. A single representative PAM is selected for testing when the PAM has degenerate bases. The target DNA consists of a 2200-bp linear DNA derived from a plasmid via PCR amplification using the PAM and a spacer located 700 bp from one end. Successful cleavage results in 700- and 1500-bp fragments. The target DNA, in vitro-transcribed single RNA, and purified recombinant protein are combined in cleavage buffer (10 mM Tris, 100 mM NaCl, 10 mM MgCl2) containing excess protein and RNA and incubated for 5 minutes to 3 hours, typically 1 hour. The reaction is stopped by the addition of RNAse A and a 60-minute incubation. The reaction is then resolved on a 1.2% TAE agarose gel, and the fraction of cleaved target DNA is quantified using ImageLab software.

[0282] Example 9 - Testing the genome cleavage activity of MG CRISPR complexes in E. coli E. coli lacks the ability to efficiently repair double-stranded DNA breaks. Therefore, cleavage of genomic DNA can be a lethal event. Taking advantage of this phenomenon, endonuclease activity is tested in E. coli by recombinantly expressing the endonuclease and guide RNA (e.g., as determined in Example 6) in a target strain with a spacer / target PAM sequence integrated into its genomic DNA (e.g., as determined in Example 4), and transforming the endonuclease-encoding DNA. The transformants are then chemically competent and transformed with 50 ng of guide RNA (e.g., crRNA) either specific for the target sequence ("on-target") or non-specific for the target ("non-target"). After heat shock, the transformants were allowed to recover in SOC at 37°C for 2 hours. Nuclease efficiency was then determined by a 5-fold dilution series grown on induction medium. Colonies were quantified from the dilution series in triplicate. A reduction in the number of colonies transformed with the on-target guide RNA compared to the number of colonies transformed with the off-target guide RNA indicates specific genomic cleavage by the endonuclease.

[0283] Example 10 - General Procedure: Testing the Genome Cleavage Activity of MG CRISPR Complexes in Mammalian Cells Two types of mammalian expression vectors are used to detect targeting and cleavage activity in mammalian cells. First, the MG Cas effector is fused to a viral 2A consensus cleavable peptide sequence linked to a C-terminal SV40 NLS and a GFP tag (2A-GFP tag for monitoring protein expression). Second, the MG Cas effector is fused to two SV40 NLS sequences, one at the N-terminus and the other at the C-terminus. The NLS sequences include any of the NLS sequences described herein (e.g., SEQ ID NOS: 3938-3953). In some cases, the nucleotide sequence encoding the endonuclease is codon-optimized for expression in mammalian cells.

[0284] A single guide RNA carrying a crRNA sequence fused to a sequence complementary to the mammalian target DNA is cloned into a second mammalian expression vector. The two plasmids are co-transfected into HEK293T cells. 72 hours after co-transfection, DNA is extracted from the transformed HEK293T cells and used to prepare an NGS library. The percentage of NHEJ is measured by quantifying indels at the target site to indicate the targeting efficiency of the enzyme in mammalian cells. At least 10 different target sites are selected to test the activity of each protein.

[0285] Example 11 - Testing genome cleavage activity of MG CRISPR complexes in mammalian cells To demonstrate targeting and cleavage activity in mammalian cells, the MG Cas effector protein sequence was cloned C-terminally after the His tag into a mammalian expression vector with flanking N- and C-terminal SV40 NLS sequences, a C-terminal His tag, and a 2A-GFP (e.g., viral 2A consensus cleavable peptide sequence linked to GFP) tag (Scaffold 1). In some cases, the endonuclease-encoding nucleotide sequence was codon-optimized for expression in E. coli cells or was a native sequence codon-optimized for expression in mammalian cells.

[0286] A single guide RNA sequence (sgRNA) carrying the desired gene target was also cloned into a mammalian expression vector. The two plasmids were co-transfected into HEK293T cells. 72 hours after co-transfection of the expression plasmid and the sgRNA-targeting plasmid into HEK293T cells, DNA was extracted and used for NGS library preparation. The percentage of NHEJ was measured via indels in sequencing of the target site to demonstrate the targeting efficiency of the enzyme in mammalian cells. Seven to 12 different target sites were selected for testing the activity of each protein. An arbitrary threshold of 5% indels was used to identify active candidates. Genome editing efficiency in human cells was assessed from NGS reads using CRISPResso with the following parameters: cleavage offset = -4 and window = 10. All post-cleavage events from the CRISPResso output were summed for indels / mutations of ±1 bp and deletions, insertions, and mutations of ≥2 bp. All results were normalized to the total sequence aligned to the expected amplicon (Figure 18A-E).

[0287] Example 12 - Characterization of the MG29 family PAM specificity, tracrRNA / sgRNA validation The targeted endonuclease activity of the MG29 family endonuclease systems was confirmed using the myTXTL system described in Examples 3 and 8. In this assay, PCR amplification of the cleaved target plasmid yields a product migrating at approximately 170 bp in the gel, as shown in Figures 17A and 17B. For MG29-1, an amplification product was observed using crRNA corresponding to SEQ ID NO: 3609 (Figure 17A, lane 7). Sequencing of the PCR products revealed the active PAM sequences for these enzymes, as shown in Table 2 below.

[0288] [Table 3]

[0289] Targeted endonuclease activity in mammalian cells MG29-1 target loci were selected for testing locations within the genome using the PAM YYn (SEQ ID NO: 3871). Spacers corresponding to the selected target sites were cloned into the sgRNA scaffold of mammalian vector-based backbone 1 described in Example 9. The sites are listed in Table 3 below. The activity of MG29-1 at various target sites is shown in Table 2 and Figure 19.

[0290] [Table 4]

[0291] Example 13 - High replication PAM determination via NGS Type V endonucleases (e.g., MG28, MG29, MG30, and MG31 endonucleases) were tested for cleavage activity using E. coli lysate-based expression in the myTXTL kit, as described in Examples 3 and 8. Upon incubation of the crRNA with a plasmid library containing spacer sequencing matching the crRNA preceded by eight degenerate (N) bases (5' PAM library), they cleaved a subset of the plasmid library with a functional PAM. Ligation to this cleavage site and PCR amplification provided evidence of activity, indicated by a band observed on a gel at 170 bp (Figure 17B). Gel 1 (top panel, A) lanes are as follows: 1 (ladder; the darkest band corresponds to 200 bp); 2: positive control (previously validated library); 3 (n / a); 4 (n / a); 5 (MG28-1); 6 (MG29-1); 7 (MG30-1); 8 (MG31-1); 9 (MG32-1); and 10 (ladder). Gel 2 (bottom panel, B) lanes are as follows: 1 (ladder; the darkest band corresponds to 200 bp); 2 (LbCpf1 positive control); 3 (LbCpf1 positive control); 4 (negative control); 5 (n / a); 6 (n / a); 7 (MG28-1); 8 (MG29-1); 9 (MG30-1); 10 (MG31-1); and 11 (MG32-1).

[0292] The PCR products were further subjected to NGS sequencing, and the PAM was matched to the seqLogo (see, e.g., Huber et al. Nat Methods. 2015 Feb;12(2):115-21, which is incorporated herein by reference) representation (Figure 20). The seqLogo representation shows 8 bp upstream (5') of the spacer, labeled as positions 0-7. As shown in Figure 20, the PAM is pyrimidine-rich (C and T), with most of the sequence requirement being 2-4 bp upstream of the spacer (positions 4-6 in SeqLogo).

[0293] The PAMs of the MG candidates are shown in Table 4 below.

[0294] [Table 5]

[0295] In some cases, the position immediately adjacent to the spacer may have weaker specificity, for example, for "m" or "v" instead of "n".

[0296] Example 14 - Targeted endonuclease activity in mammalian cells with MG31 nuclease Targeted endonuclease activity in mammalian cells The MG31-1 target locus was selected for testing locations within the genome using the PAM TTTR (SEQ ID NO: 3875). Spacers corresponding to the selected target sites were cloned into the sgRNA scaffold of mammalian vector-based backbone 1 described in Example 11. The sites are listed in Table 5 below. The activity of MG31-1 at various target sites is shown in Table 5 and in Figure 25.

[0297] [Table 6]

[0298] Example 3 - In vitro activity Promising candidates from the bioinformatics analysis and preliminary screening were selected for further biochemical analysis, as described in this Example. Using the conserved 3' sgRNA structure, a "universal" sgRNA was designed, containing a 3' 20-nt CRISPR repeat and a 24-nt spacer (Figure 10A-D). Of the seven candidates tested, six showed activity in vitro against the 8N PAM library (Figure 26A). The remaining inactive candidate (30-1) showed activity when tested with its predicted endogenous trimming CRISPR repeat (SEQ ID NO: 3608, see Figure 26B), but not included in the NGS library assay (Figure 26C).

[0299] The majority of identified PAMs are thymine-rich sequences of 2-3 bases (Figure 18A). However, two enzymes, MG26-1 (PAM YYn) and MG29-1 (PAM YYn), possess PAM specificity for either pyrimidine bases, thymine, or cytosine, allowing for broader sequence targeting. Analysis of putative PAM-interacting residues showed that active type VA nucleases contain a conserved lysine and a GWxxxK motif, which are shown to be important for recognition and interaction with different PAMs in FnCas12a.

[0300] Because our PAM detection assay requires ligation to generate blunt-ended fragments before PAM enrichment, this suggested that these enzymes created staggered double-stranded DNA breaks, similar to reported VA-type nucleases. The cleavage site on the target strand could be identified by analysis of the NGS reads used for indel detection (Figure 18B), which showed cleavage after the 22nd PAM-distal base.

[0301] In vitro cleavage by MG29-1 was further investigated by sequencing the cleavage products. The cleavage position on the target strand was 22 nucleotides away from the PAM in most sequences, less frequently by 21 or 23 nucleotides (Figures 56A-C). The cleavage position on the non-target strand was 17-19 nucleotides from the PAM. Combined, these results indicate a 3-5 bp overhang.

[0302] Example 4 - Genome editing After PAM validation, the novel proteins described herein were tested for gene targeting activity in HEK293T cells. All candidates demonstrated greater than 5% NHEJ (background-corrected) activity against at least one of the 10 tested target loci. MG29-1 showed the highest overall activity in NHEJ modification results (Figure 18B) and was active against the highest number of targets. Therefore, this nuclease was selected for purified ribonucleoprotein complex (RNP) testing in HEK293 cells. RNP transfection of the MG29-1 holoenzyme demonstrated higher RNP-mediated editing levels than plasmid-based transfection in four of nine targets, with editing efficiencies of 80% in some cases (Figure 18C). Analysis of the editing profile of MG29-1 indicates that this nuclease produces deletions at its target sites by more than 2 bp more frequently than other types of editing (Figure 18D). For some targets (5 and 8), the indel frequency in MG29-1 was twice that of AsCpf1 (Figure 18E).

[0303] Example 17 - Discussion VA-type CRISPRs were identified from metagenomes collected from diverse complex environments and arranged into families. These novel VA-type nucleases possess diverse sequences and phylogenetic origins within and between families, as well as cleaved targets with diverse PAM sites. Similar to other VA-type nucleases (e.g., LbCas12a, AsCas12a, and FnCas12a), the effectors described herein utilize a single guide CRISPR RNA (sgRNA) to target staggered double-strand breaks in DNA, simplifying guide design and synthesis and facilitating multiplexed editing. Analysis of CRISPR repeat motifs that formed stem-loop structures in crRNAs suggested that the VA-type effectors described herein frequently possess 4-nt loop guides rather than shorter or longer loops. The sgRNA motif of LbCpf1 possesses a less common 5-nt loop, but 4-nt loops were also observed in 16 previously identified Cpf1 orthologs. An unusual stem-loop CRISPR repeat motif sequence, CCUGC[N 3-4 ]GCAGG was identified for the MG61 family of VA-type effectors. The high degree of conservation of sgRNAs with variable loop lengths in VA-type effectors may result in flexible levels of activity, as shown for the proteins described herein. Collectively, these effectors are not close homologs of previously studied enzymes, and significantly expand the diversity of VA-type-like sgRNA nucleases.

[0304] The additional type V effectors described herein may have evolved from duplications of type V-like nucleases, referred to herein as type V-A prime effectors (V-A'), which may be encoded adjacent to the Cas12a nuclease. Both type V and these type V-A' systems may share a CRISPR sgRNA, but type V-A' systems have diverged from Cas12a (Figure 4). The CRISPR repeats associated with these prime effectors also contain UCUAC[N 3-5] folds into a single-guide crRNA with a GUAGAU motif. One report identified a V-type cms1 effector encoded next to a VA-type nuclease, which required a single-guide crRNA for cleavage activity in plant cells. Although different CRISPR arrays were reported for each effector, the V-A' type system described herein suggested that both VA-type and V-A'-type effectors may require the same crRNA for DNA targeting and cleavage. As described in the Roizman bacterial genome (see, e.g., Chen et al. Front Microbiol. 2019 May 3;10:928), both VA-type and V-A'-type effectors are distantly related based on sequence homology and phylogenetic analysis. Therefore, prime effectors do not belong within the VA-type classification and deserve a separate V-type subclassification.

[0305] The PAMs determined for active type VA nucleases were generally thymine-rich and similar to those described for other type VA nucleases. In contrast, MG29-1 required a shorter YYN PAM sequence, which increases target flexibility compared to the four-nucleotide TTTV PAM of LbCpf1. Furthermore, RNPs containing MG29-1 had higher activity in HEK293 cells compared to sMbCas12a, which has a three-nucleotide PAM.

[0306] When novel nucleases were tested for in vitro editing activity, MG29-1 exhibited comparable or better activity than other reported enzymes in its class. Reports of plasmid transfection editing efficiency in mammalian cells using Cas12a orthologs showed indel frequencies of 21% to 26% for guides with T-rich PAMs, and one of 18 guides with a CCN PAM showed approximately 10% activity with Mb3Cas12a (Moraxella bovoculi AAX11_00205 Cas12a, see, e.g., Wang et al. Journal of Cell Science 2020 133:jcs240705). Notably, MG29-1 activity in plasmid transfection appears to be greater than that reported for Mb3Cas12a for targets with TTN and CCN PAMs (see, e.g., Figures 18A-E). Because the target sites for plasmid transfection had the same TTG PAM in all experiments, differences in editing efficiency could be attributed to differences in genome accessibility at different target genes. MG29-1 editing via RNP was much more efficient than via plasmid and more efficient than AsCas12a at two of the seven target loci. Therefore, MG29-1 may be a highly active and efficient gene-editing nuclease. These findings increase the diversity of identified single-guide VA-type CRISPR nucleases and demonstrate the genome editing potential of novel enzymes from uncultured microorganisms. Seven novel nucleases demonstrated in vitro activity with diverse PAM requirements, and RNP data demonstrated editing efficiencies exceeding 80% for therapeutically relevant targets in human cell lines. These novel nucleases expand the toolkit of CRISPR-associated enzymes, enabling diverse genome engineering applications.

[0307] Example 18 - MG29-1 induced editing of the TRAC locus in T cells Three exons of the T cell receptor alpha chain constant region (TRACA) were scanned for sequences matching the initial predicted 5'-TTN-3' PAM specificity of MG29-1, and single-stranded guide RNAs with the proprietary Alt-R modification were ordered from IDT. All guide spacer sequences were 22 nt in length. Guides (80 pmol) were mixed with purified MG29-1 protein (63 pmol) and incubated at room temperature for 15 minutes. T cells were purified from PBMCs by negative selection using Stemcell Technologies Human T Cell Isolation Kit #17951 and activated with CD2 / 3 / 28 beads (Miltenyi T Cell Activation / Proliferation Kit #130-091-441). After 4 days of cell expansion, each MG29-1 / guide RNA mixture was electroporated into 200,000 T cells using a Lonza 4-D Nucleofector with program EO-115 and P3 buffer. Cells were harvested 72 hours after transfection, and genomic DNA was isolated and PCR-amplified for analysis using high-throughput DNA sequencing with primers targeting the TRACA locus. The creation of insertions and deletions characteristic of NHEJ-based gene editing was quantified using a proprietary Python script (see Figure 39).

[0308] [Table 7-1]

[0309] [Table 7-2]

[0310] [Table 7-3]

[0311] [Table 7-4]

[0312] [Table 7-5]

[0313] [Table 7-6]

[0314] Example 19 - Retest of lead guide of MG29-1 An experiment was conducted to retest the MG29-1 lead guide. Three exons of the T cell receptor alpha chain constant region were scanned for sequences matching the 5'-TTN-3' and single-stranded guide RNAs ordered from IDT using Alt-R modifications. All guide spacer sequences were 22 nt in length. The guides were mixed with purified MG29-1 protein (80 pmol gRNA + 63 pmol MG29-1, or 160 pmol gRNA with 126 pmol MG29-1) and incubated for 15 minutes at room temperature. T cells were purified from PBMCs by negative selection using (Stemcell Technologies Human T Cell Isolation Kit #17951) and activated with CD2 / 3 / 28 beads (Miltenyi T Cell Activation / Proliferation Kit #130-091-441). After 4 days of cell growth, each MG29-1 / guide RNA mixture was electroporated into 200,000 T cells using a Lonza 4-D Nucleofector with program EO-115 and P3 buffer. 72 hours after transfection, genomic DNA was harvested and PCR-amplified for analysis using high-throughput DNA sequencing. The creation of insertions and deletions characteristic of NHEJ-based gene editing was quantified using a proprietary Python script (see Figure 40).

[0315] Example 20 - Testing the length of the guide spacer of MG29-1 Experiments were performed to determine the optimal guide spacer length. Three exons of the T cell receptor alpha chain constant region were scanned for sequences matching the 5'-TTN-3' and single-stranded guide RNAs ordered from IDT using Alt-R modifications. The guides were mixed with purified MG29-1 protein (80 pmol gRNA + 60 pmol effector, 160 pmol gRNA + 120 pmol effector, or 320 pmol gRNA + 240 pmol effector) and incubated at room temperature for 15 minutes. T cells were purified from PBMCs by negative selection using (Stemcell Technologies Human T Cell Isolation Kit #17951) and activated with CD2 / 3 / 28 beads (Miltenyi T Cell Activation / Proliferation Kit #130-091-441). After 4 days of cell growth, each MG29-1 / guide RNA mixture was electroporated into 200,000 T cells using a Lonza 4-D Nucleofector with program EO-115 and P3 buffer. Seventy-two hours after transfection, genomic DNA was harvested and PCR-amplified for analysis using high-throughput DNA sequencing. The creation of insertions and deletions characteristic of NHEJ-based gene editing was quantified using a proprietary Python script. The results are shown in Figure 41, which demonstrates that guide spacer lengths of 20-24 nt work well, with a drop-off at 19 nt.

[0316] Example 21 - Determination of MG29-1 indel generation on TCR expression The cells in Figure 41 were analyzed for TCR expression by flow cytometry using an APC-labeled anti-human TCR α / β Ab (Biolegend #306718, clone IP26) and an Attune NxT flow cytometer (Thermo Fisher). Indel data is taken from Figure 41.

[0317] Example 22 - Targeted CAR integration with MG29-1 Three exons of the T cell receptor alpha chain constant region were scanned for sequences matching the 5'-TTN-3' and single-stranded guide RNAs ordered from IDT using IDT's proprietary Alt-R modification. The guide (80 pmol) was mixed with purified MG29-1 protein (63 pmol) and incubated at room temperature for 15 minutes. T cells were purified from PBMCs by negative selection using Stemcell Technologies Human T Cell Isolation Kit #17951 and activated with CD2 / 3 / 28 beads (Miltenyi T Cell Activation / Proliferation Kit #130-091-441). After 4 days of cell expansion, each MG29-1 / guide RNA mixture was electroporated into 200,000 T cells using a Lonza 4-D Nucleofector with program EO-115 and P3 buffer. 100,000 vector genomes of serotype 6 adeno-associated virus (AAV-6) containing the coding sequence of a customized chimeric antigen receptor flanked by 5' and 3' homology arms targeting the TRAC gene (5' arm SEQ ID NO: 4424 is approximately 500 nt in length, and the 3' arm SEQ ID NO: 4425 is approximately 500 nt in length) were added to the cells immediately after transfection. Replicates were analyzed for TCR expression relative to TRAC indels (Figure 42), showing that indels in the TRAC gene correlated with loss of TCR expression. Cells were also simultaneously analyzed by flow cytometry, as in Example 21, for TCR expression (Figure 42) and for target antigen binding to the CAR (Figure 43, where plots are gated on single live cells). The flow analysis results in Figure 43 show that while guide RNA alone was effective in eliminating TCR expression ("RNP only"), the addition of guide RNA and AAV resulted in a new population of cells that bound to the CAR antigen (top left of the plot, "AAV+MG29-1-19-22" and "AAV+MG29-1-35-22"). sgRNA 35 (SEQ ID NO: 4404) was slightly more effective at inducing CAR integration than sgRNA 19 (SEQ ID NO: 4388). One possible explanation for the difference is that the predicted nuclease cleavage site of guide 19 is approximately 160 bp away from the end of the right homology arm.

[0318] [Table 8]

[0319] Example 5 - MG29-1 TRAC editing in HSCs Hematopoietic stem cells were purchased from Allcells, thawed according to the supplier's instructions, washed with DMEM + 10% FBS, and resuspended in Stemspan II medium + CC110 cytokines. One million cells were cultured in 4 mL of medium in a 6-well dish for 72 hours. MG29-1 RNPs were generated, transfected, and analyzed for gene editing as in Example 18, except for the use of the EO-100 nucleofection program. The results are shown in Figure 61, which shows gene editing at TRAC in hematopoietic stem cells using the #19 (SEQ ID NO: 4388) and #35 (SEQ ID NO: 4404) sgRNAs in Table 5B below. The results also show that the #35 sgRNA is highly effective at targeting the TRAC locus.

[0320] [Table 9]

[0321] Example 6 - Further analysis of PAM specificity relative to MG29-1 Further analysis was performed to more precisely determine the PAM specificity of MG29-1. N

[00130] Guides were designed using the TTN-3' PAM sequence and then classified according to observed gene editing activity (Figure 45, where the identity of the underlined base -5'-proximal N is shown for each bin). All guides with activity greater than 10% had a T at this position in the genomic DNA, indicating that the MG29-1 PAM can be better described as 5'-TTTN-3'. The statistical significance of the overrepresentation of T at this position is shown for each bin. In Figure 45, the various bins (high, medium, low, >1%, <1%) indicate the following: High: >50% indels (N=4) Medium: 10-50% indels (N=15) Low: 5-10% indels (N=5) >1%: 1-5% indels (N=12) <1% (N=82)

[0322] [Table 10]

[0323] Example 7 - Determination of MG29-1 indel-inducing capacity versus spacer-based composition Further analysis of gene editing activity against the base composition of the MG29-1 spacer sequence was performed. The correlation was moderate (R^2=0.23), but there was a trend toward better activity with higher GC content (see Figure 46, where the correlation between indels induced in cultured cells and the GC content of the spacer sequence is represented as a dot plot).

[0324] Example 26 - MG29-1 guided chemical modification Experiments to optimize chemical modifications for targeting the VEGF-A locus using MG29-1 were performed using the procedures of Example 18, but with the indicated guide RNAs targeting VEGF-A (see Table 7 below). The experiment used 126 pmol of MG29-1 and 160 pmol of guide RNA. The results are presented in Figure 47. Guides #4, 5, 6, 7, and 8 showed improved activity relative to unmodified guide #1, indicating that the corresponding modifications in these sequences improved the activity of these guide RNAs relative to the unmodified RNA sequence.

[0325] [Table 11]

[0326] Example 27 - Titration of modified MG29-1 guide from Example 26 Further experiments were performed to determine the dose-dependence of activity of the modified guide used in Example 26 and to identify possible dose-dependent toxic effects. Experiments were performed as in Example 26, but using 1 / 4 (B), 1 / 8 (C), 1 / 16 (D), and 1 / 32 (E) of the starting dose (A, 126 pmol MG29-1, and 160 pmol guide RNA). Results are presented in Table 48).

[0327] Example 28 - Large-scale synthesis of nucleases described herein Project Overview Production of Metagenomi's VA-type CRISPR nuclease, MG29-1, will be scaled up to an initial culture volume of 10 L. Expression screening, scale-up expression, downstream development, formulation studies, and delivery of ≥90% purified protein by SDS-PAGE will be performed.

[0328] Expression and purification screening Expression: Expression of MG29-1 from the pMG450 vector illustrated in Figure 49 is tested in a screen varying the following conditions: host strain, expression medium, inducer, induction time, and temperature.

[0329] After screening, E. coli is transformed with the appropriate expression plasmid, cultures are grown in shake flasks at the appropriate density, and the cultures are induced using materials and methods according to the optimal expression conditions identified during expression screening. Cell paste is harvested and expression is verified by SDS-PAGE. For these experiments, the cell culture volume is limited to 20 L. Up to 1 gram of protein is purified using the following method, formulated into a storage buffer, and assessed for yield and concentration by A280 and purity by SDS-PAGE.

[0330] purification: Total soluble protein extracted from E. coli cell paste is analyzed by SDS-PAGE for all conditions. Immobilized metal affinity chromatography (IMAC) pulldown followed by SDS-PAGE is performed for the top three expression conditions to estimate yield and purity and identify optimal expression conditions. A scale-up method is developed for lysis. Critical parameters are identified for purification by IMAC and subtractive IMAC (including tobacco etch virus protease (TEV) cleavage). Column fractions are tested using SDS-PAGE. The elution pool is tested using SDS-PAGE and optical absorbance at 280 nm (A280). A method is developed for buffer exchange and concentration by tangential flow filtration (TFF).

[0331] If purity is below 90%, additional chromatography steps are developed to achieve ≥90% purity. One chromatography mode is tested (e.g., ceramic hydroxyapatite chromatography). Up to eight unique conditions (e.g., two to six resins with two to three buffer systems each) are tested. Column fractions are tested using SDS-PAGE. The elution pool is tested using SDS-PAGE and A280. One condition is selected and a three-condition loading study is performed. Column fractions and elution pools are analyzed as described above. A method for buffer exchange can be developed, followed by concentration by TFF.

[0332] Formulation Research Formulation studies are performed using the purified protein to determine optimal storage conditions for the purified protein. Studies may explore concentration, storage buffer, storage temperature, maximum freeze / thaw cycles, storage time, or other conditions.

[0333] Example 29 - Demonstration of the ability of the nucleases described herein to edit intron regions in cultured mouse liver cells The intron region of an expressed gene is an attractive genomic target for integrating the coding sequence of a therapeutic protein of interest for the purpose of expressing that protein to treat or cure a disease. Integration of the protein-coding sequence can be achieved by creating a double-strand break within the intron using a sequence-specific nuclease in the presence of an exogenously supplied donor template. The donor template may be integrated into the double-strand break via one of two major cellular repair pathways, called homology-directed repair (HDR) and non-homologous end joining (NHEJ), resulting in targeted integration of the donor template. The NHEJ pathway is predominant in non-dividing cells, while the HDR pathway is primarily active in dividing cells. The liver is a particularly attractive tissue for targeted integration of protein-coding sequences due to the availability of in vivo delivery systems and the liver's ability to express and secrete proteins with high efficiency.

[0334] To evaluate the potential of MG29-1 to create double-stranded breaks in intronic regions, intron 1 of serum albumin was selected as the target locus. Single guide RNAs (sgRNAs) with a 22-nt spacer length targeting mouse albumin intron 1 were identified using the guide discovery algorithm in Geneious Prime nucleic acid analysis software (https: / / www.geneious.com / prime / ). A total of 112 potential sgRNAs were identified within mouse albumin intron 1 using a PAM of KTTG (sequence number 3870) located 5' of the spacer. Guides spanning intron / exon boundaries were excluded. Using Geneious Prime, the spacer sequences of these 112 guides were searched against the mouse genome, and the software assigned specificity scores based on alignment to additional sites within the genome. Spacer sequences with four or more consecutive bases of the same base were excluded due to concerns about specificity. A total of 12 spacers with the highest specificity scores were selected for testing. To generate the sgRNAs, the backbone sequence "TAATTTCTACTGTTGTAGAT" was added to the 3' end of the spacer sequence. The sgRNAs were chemically synthesized incorporating chemically modified bases identified to improve sgRNA performance for the cpf1 guide (AltR1 / AltR2 chemistry available from Integrated DNA Technologies). The spacer sequences for these guides are listed below in Table 8.

[0335] [Table 12]

[0336] Hepa1-6 cells, a transformed mouse liver cell line, were cultured under standard conditions (DMEM medium with 10% FBS in a 5% CO2 incubator) and nucleofected with ribonucleoproteins formed by mixing sgRNA and purified MG29-1 protein in PBS buffer. Hepa1-6 cells (1 × 10) in suspension in Complete SF Nucleofection Reagent (Lonza) were cultured under standard conditions (DMEM medium with 10% FBS in a 5% CO2 incubator).5 ) was nucleofected using a 4D nucleofection device (Lonza) with RNPs formed by mixing 50 pmol of MG29-1 protein and 100 pmol of sgRNA. After nucleofection, cells were seeded into 24-well plates in DMEM + 10% FBS and incubated in a 5% CO2 incubator for 48–72 hours. Genomic DNA was then extracted from the cells using a column-based purification kit (Purelink genomic DNA mini kit, ThermoFisher Scientific) and quantified by absorbance at 260 nm. The albumin intron 1 region was PCR amplified from 50 ng of genomic DNA in a reaction containing 0.5 micromoles each of primers mAlb90F (CTCCTCTTCGTCTCCGGC) (SEQ ID NO: 4031) and mAlb1073R (CTGCCACATTGCTCAGCAC) (SEQ ID NO: 4032) and 1x Pfusion Flash PCR Master Mix.

[0337] The resulting 984-bp PCR product spanning intron 1 of mouse albumin was purified using a column-based purification kit (DNA Clean and Concentrator, Zymo Research) and sequenced using primers located within 150–350 bp of the predicted target site of each sgRNA. PCR products generated using primers mAlb90F (SEQ ID NO: 4031) and mAlb1073R (SEQ ID NO: 4032) from untransfected Hepa1-6 cells were sequenced in parallel as a control. Sanger sequencing chromatograms were analyzed using Inference of CRISPR Edits (ICE) to determine indel frequencies and indel profiles (Hsiau et al., Inference of CRISPR Edits from Sanger Trace Data. BioArxiv. 2018 https: / / www.biorxiv.org / content / early / 2018 / 01 / 20 / 251082).

[0338] When nucleases create double-strand breaks (DSBs) in DNA within living cells, the DSBs are repaired by cellular DNA repair mechanisms. In actively dividing cells, such as transformed mammalian cells in culture, and in the absence of a repair template, this repair occurs via the NHEJ pathway. The NHEJ pathway is an error-prone process that introduces base insertions or deletions at the site of the double-strand break (Lieber, MR, Annu Rev Biochem. 2010;79:181-211). Therefore, these insertions and deletions are characteristic of the generated and subsequently repaired double-strand break and are widely used as a readout of the editing or cleavage efficiency of nucleases. The insertion and deletion profiles depend not only on the properties of the nuclease that created the double-strand break but also on the sequence context at the cleavage site. Based on in vitro assays, MG29-1 nuclease creates staggered cuts located 3' to the PAM. Staggered cuts often result in larger deletions due to trimming of single-stranded ends before end-joining. Table 8 lists the total indel frequencies generated by each of the 19 sgRNAs targeting mouse albumin intron 1 tested in Hepa1-6 cells. Eleven of the 18 sgRNAs resulted in detectable indels at the target site, with five sgRNAs resulting in indel frequencies greater than 50% and four sgRNAs resulting in indel frequencies greater than 75%. These data demonstrate that MG29-1 nuclease can edit the genome of cultured mouse liver cells at the predicted target site of the sgRNA with greater than 75% efficiency.

[0339] The editing efficiency of the same set of sgRNAs was assessed by cotransfection of sgRNAs and mRNA encoding MG29-1 nuclease using a commercially available lipid-based transfection reagent (Lipofectamine MessengerMAX, Invitrogen). The mRNA encoding MG29-1 was generated by in vitro transcription using T7 polymerase from a plasmid into which the MG29-1 coding sequence had been cloned. The MG29-1 coding sequence was codon-optimized using a human codon usage table and flanked by nuclear localization signals derived from SV40 at the N-terminus and nucleoplasmin at the C-terminus. Furthermore, to improve translation, a UTR was included at the 3' end of the coding sequence. To improve mRNA stability in vivo, a 3' UTR followed by a poly(A) tract of approximately 90–110 nucleotides was included at the 3' end of the coding sequence (see, for example, SEQ ID NO: 4426 for wild-type MG29-1 and SEQ ID NO: 3327 for the S168R variant). The in vitro transcription reaction included CleanCap® capping reagent (Trilink BioTechnologies), and the resulting RNA was purified using the MEGAClear™ Transcription Clean-Up Kit (Invitrogen). Purity was assessed using a TapeStation (Agilent) and found to consist of >90% full-length RNA. As seen in Table 1, editing efficiency after mRNA / sgRNA lipid transfection of Hepa1-6 cells was similar, but not identical, to that seen with RNP nucleofection, confirming, however, that MG29-1 nuclease is active in cultured liver cells when delivered in the form of mRNA.

[0340] Figure 50 shows a representative example of the indel profile of MG29-1, as determined by ICE analysis using mALb29-1-8 as a guide (SEQ ID NO: 3999), showing that 4-base deletions were the most frequent event (25% of all sequences), with 1-, 5-, 6-, or 7-base deletions each accounting for approximately 10-15% of sequences. Longer deletions of up to 13 bases were also detected, but insertions were undetectable. In contrast, spCas9 with a guide targeting mouse albumin intron 1 generated primarily 1-base insertions or deletions.

[0341] Figure 51 shows representative examples of indel profiles for MG29-1 and sgRNA mAlb29-1-8, as determined by next-generation sequencing (NGS) of PCR products from the mouse albumin intron 1 region. A total of approximately 15,000 sequence reads were obtained. A 4-base deletion was identified by NGS as the most frequent indel (approximately 20% of the total), with 1-, 5-, 6-, and 7-base deletions each accounting for approximately 10% of indels. Larger deletions of up to 19 bp were also detected. The profile observed by NGS analysis closely matches that measured by ICE. These results indicate that MG29-1 generates larger deletions at the target site, consistent with the staggered cleavage observed in vitro.

[0342] Example 8 - Demonstration of the ability of nucleases described herein to target intron regions in cultured human liver cells (HepG2) To evaluate the potential of MG29-1 to create double-stranded breaks in intronic regions in human cells, intron 1 of human serum albumin was selected as the target locus. Single guide RNAs (sgRNAs) with a 22-nt spacer length targeting human albumin intron 1 were identified using the guide discovery algorithm in Geneious Prime nucleic acid analysis software (https: / / www.geneious.com / prime / ). A total of 90 potential sgRNAs were identified within human albumin intron 1 using a PAM of KTTG (sequence number 3870) located 5' of the spacer. Guides spanning intron / exon boundaries were excluded. Using Geneious Prime, the spacer sequences of these guides were searched against the mouse genome, and the software assigned specificity scores based on alignment to additional sites within the genome. Spacer sequences with four or more consecutive bases of the same base were excluded due to concerns about specificity. A total of 23 spacers with the highest specificity scores were selected for testing. To generate the sgRNAs, the backbone sequence "TAATTTCTACTGTTGTAGAT" was added to the 3' end of the spacer sequence. The sgRNAs were chemically synthesized incorporating chemically modified bases identified to improve sgRNA performance for the cpf1 guide (AltR1 / AltR2 chemistry available from Integrated DNA Technologies). The spacer sequences for these guides are listed below in Table 9.

[0343] [Table 13]

[0344] HepG2 cells, a transformed human liver cell line, were cultured under standard conditions (MEM medium with 10% FBS in a 5% CO2 incubator) and nucleofected with a ribonucleoprotein formed by mixing sgRNA and purified MG29-1 protein in PBS buffer. A total of 1 e5 HepG2 cells in suspension in Complete SF Nucleofection Reagent (Lonza) were nucleofected using a 4D Nucleofection Device (Lonza) with an RNP formed by mixing 80 pmol of MG29-1 protein and 160 pmol of sgRNA. After nucleofection, cells were seeded into 24-well plates in DMEM + 10% FBS and incubated for 48–72 h in a 5% CO2 incubator. Genomic DNA was then extracted from the cells using a column-based purification kit (Purelink genomic DNA mini kit, ThermoFisher Scientific) and quantified by absorbance at 260 nm. The albumin intron 1 region was PCR amplified from 50 ng of genomic DNA in a reaction containing 0.5 micromoles each of primers hAlb 11F (TCTTCTGTCAACCCCACACGCC) (SEQ ID NO: 4079) and hAlb834R (CTTGTCTGGGCAAGGGAAGA) (SEQ ID NO: 4080) and 1x Pfusion Flash PCR Master Mix. The resulting 826-bp PCR product spanning intron 1 of mouse albumin was purified using a column-based purification kit (DNA Clean and Concentrator, Zymo Research) and sequenced using primers located within 150–350 bp of the predicted target site of the sgRNA.

[0345] PCR products generated from non-transfected HepG2 cells using primers hAlb 11F (TCTTCTGTCAACCCCACACGCC) (SEQ ID NO: 4079) and hAlb834R (CTTGTCTGGGCAAGGGAAGA) (SEQ ID NO: 4080) were sequenced in parallel as a control. Sanger sequencing chromatograms were analyzed using CRISPR editing inference (ICE), which determines indel frequency and indel profile. When nucleases create double-strand breaks (DSBs) in DNA in living cells, the DSBs are repaired by cellular DNA repair mechanisms. In actively dividing cells, such as transformed mammalian cells in culture, and in the absence of a repair template, this repair occurs via the NHEJ pathway. The NHEJ pathway is an error-prone process that introduces base insertions or deletions at the site of a double-strand break (Lieber, MR, Annu Rev Biochem. 2010;79:181-211).

[0346] These insertions and deletions are therefore characteristic of the double-strand breaks that have occurred and subsequently repaired, and are widely used as a readout of the editing or cleavage efficiency of nucleases. The insertion and deletion profiles depend not only on the properties of the nuclease that created the double-strand break, but also on the sequence context at the cleavage site. Based on in vitro assays, MG29-1 nuclease cleaves the target strand 22 nucleotides from the PAM (less frequently at 21 nucleotides from the PAM) and the non-target strand 18 nucleotides from the PAM, thus creating a 4-nucleotide staggered end located 3' from the PAM. Staggered cleavage often results in larger deletions due to trimming of single-stranded ends before end-joining.

[0347] Table 9 lists the total indel frequencies generated by each of the 23 sgRNAs targeting human albumin intron 1 tested in HepG2 cells. Sixteen of the 23 sgRNAs resulted in detectable indels at the target site, eight sgRNAs resulted in indels greater than 50%, and five sgRNAs resulted in indel frequencies greater than 90%. These data demonstrate that MG29-1 nuclease can edit the genome of cultured human liver cell lines at the predicted target site of the sgRNA with greater than 90% efficiency.

[0348] Example 9 - Demonstration of the ability of the nucleases described herein to edit exon regions in cultured mouse liver cells Sequence-specific nucleases can be used to disrupt the coding sequence of a gene, thereby creating a functional knockout of a target protein. This can be used for therapeutic purposes when knockdown of a specific protein has a beneficial effect in a specific disease. One way to disrupt the coding sequence of a gene is to use sequence-specific nucleases to create double-strand breaks within the exon region of the gene. These double-strand breaks are repaired through error-prone repair pathways, generating insertions or deletions that can result in frameshift mutations or changes in the amino acid sequence that disrupt the function of the protein.

[0349] To evaluate the potential of MG29-1 to create double-stranded breaks in exon regions of genes expressed in liver cells, we selected the gene encoding glycolate oxidase (hao-1) as the target locus. Single guide RNAs (sgRNAs) with a 22-nt spacer length targeting exons 1–4 of mouse hao-1 were identified using the guide discovery algorithm in Geneious Prime nucleic acid analysis software (https: / / www.geneious.com / prime / ). The first four exons of the hao-1 gene comprise approximately the N-terminal 50% of the hao-1 coding sequence. The first four exons were chosen because indels created toward the N-terminus of the gene's coding sequence are more likely to create frameshift or missense mutations that disrupt protein activity. A total of 45 potential sgRNAs were identified within mouse hao-1 exons 1–4 using a PAM of KTTG (sequence number 3870) located 5′ of the spacer. Guides spanning intron / exon boundaries were included because they can create indels that interfere with splicing. Using Geneious Prime, the spacer sequences of these 45 guides were searched against the mouse genome, and the software assigned specificity scores based on alignment to additional sites within the mouse genome. Spacer sequences with four or more consecutive bases of the same base were excluded due to concerns about specificity. A total of 45 spacers with the highest specificity scores were selected for testing.

[0350] To generate the sgRNA, the backbone sequence "TAATTTCTACTGTTGTAGAT" was added to the 3' end of the spacer sequence. The sgRNA was chemically synthesized incorporating chemically modified bases identified to improve sgRNA performance for the cpf1 guide (AltR1 / AltR2 chemistry available from Integrated DNA Technologies). The spacer sequences for these guides are listed below in Table 3.

[0351] [Table 14]

[0352] Hepa1-6 cells, a transformed mouse liver cell line, were cultured under standard conditions (DMEM medium with 10% FBS in a 5% CO incubator) and nucleofected with a ribonucleoprotein formed by mixing sgRNA and purified MG29-1 protein in PBS buffer for a total of 1 e in suspension in Complete SF Nucleofection Reagent (Lonza). 5Hepa1-6 cells were nucleofected using a 4D nucleofection device (Lonza) with RNPs formed by mixing 50 pmol of MG29-1 protein and 100 pmol of sgRNA. After nucleofection, cells were seeded into 24-well plates in DMEM + 10% FBS and incubated for 48–72 hours in a 5% CO2 incubator. Genomic DNA was then extracted from the cells using a column-based purification kit (Purelink genomic DNA mini kit, ThermoFisher Scientific) and quantified by absorbance at 260 nm. Exons 1–4 of the mouse hao-1 gene were PCR amplified from 40 ng of genomic DNA in reactions containing 0.5 μmol pairs of primers specific for each exon. The PCR primers used for exon 1 were PCR_mHE1_F_+233 (GTGACCAACCCTACCCGTTT) (SEQ ID NO: 4171) and PCR_mHE1_R_-553 (GCAAGCACCTACTGTCTCGT) (SEQ ID NO: 4172). The PCR primers used for exon 2 were HAO1_E2_F5721 (CAACGAAGGTTCCCTCCAGG) (SEQ ID NO: 4173) and HAO1_E2_R6271 (GGAAGGGTGTTCGAGAAGGA) (SEQ ID NO: 4174). The PCR primers used for exon 3 were HAO1_E3_F23198 (TGCCCTAGACAAGCTGACAC) (SEQ ID NO: 4175) and HAO1_E3_R23879 (CAGATTCTGGAAGTGGCCCA) (SEQ ID NO: 4176). The PCR primers used for exon 4 were HAO1_E4_F31087 (CCTGTAGGTGGCTGAGTACG) (SEQ ID NO: 4177) and HAO1_E4_R31650 (AGGTTTGGTTCCCCTCACCT) (SEQ ID NO: 4178).

[0353] In addition to primers and genomic DNA, the PCR reaction contained 1x Pfusion Flash PCR Master Mix (Thermo Fisher Scientific). The resulting PCR product contained a single band when analyzed on an agarose gel, indicating that the PCR reaction was specific, and was purified using a column-based purification kit (DNA Clean and Concentrator, Zymo Research). For sequencing, primers complementary to sequences at least 100 nt from each cleavage site were used. The primer for sequencing exon 1 was Seq_mHE1_F_+139 (GTCTAGGCATACAATGTTTGCTCA) (SEQ ID NO: 4179). The primer for sequencing exon 2 was 5938F Seq_HAO1_E2 (CTATGCAAGGAAAAGATTTGGCC) (SEQ ID NO: 4180). The primer for sequencing exon 3 was HAO1_E3_F23476 (TCTTCCCCCTTGAATGAAACACT) (SEQ ID NO: 4181) and the reverse PCR primer, HAO1_E3_R23879 (CAGATTCTGGAAGTGGCCCA) (SEQ ID NO: 4182). The primer for sequencing exon 4 was the reverse PCR primer, HAO1_E4_R31650 (AGGTTTGGTTCCCCTCACCT) (SEQ ID NO: 4183).

[0354] Sequencing of the PCR products showed that they contained the expected sequences of the hao-1 exons. PCR products from different RNP-nucleofected Hepa-16 cells or untreated controls were sequenced using primers located within 100–350 bp of the predicted target site of each sgRNA. Sanger sequencing chromatograms were analyzed using Inference of CRISPR Editing (ICE), which determines indel frequencies and indel profiles (Hsiau et al., Inference of CRISPR Edits from Sanger Trace Data. BioArxiv. 2018 https: / / www.biorxiv.org / content / early / 2018 / 01 / 20 / 251082). When nucleases create double-strand breaks (DSBs) in DNA in living cells, the DSBs are repaired by cellular DNA repair mechanisms. In actively dividing cells, such as transformed mammalian cells in culture, and in the absence of a repair template, this repair occurs via the NHEJ pathway. The NHEJ pathway is an error-prone process that introduces base insertions or deletions at the site of a double-strand break (Lieber, MR, Annu Rev Biochem. 2010;79:181-211). Therefore, these insertions and deletions are characteristic of the generated and subsequently repaired double-strand break and are widely used in the art as a readout of the editing or cleavage efficiency of nucleases. As shown in Table 3, 14 guides showed detectable editing at their predicted target sites. Four guides exhibited editing activity greater than 90%. All 14 active guides contained the TTTN PAM sequence, indicating that this PAM is more efficient in vivo. However, not all guides utilizing the TTTN PAM were active. These data demonstrate that MG29-1 nuclease can generate RNA guide sequence-specific double-strand breaks in exon regions in cultured liver cells with high efficiency.

[0355] Example 10 - Design of additional sgRNAs for disruption of the Hao-1 gene Additional sgRNAs were designed to target exon portions of the hao-1 gene. These were designed to target the first four exons because they contain approximately 50% of the coding sequence, and indels created toward the N-terminus of the gene's coding sequence are more likely to create frameshift or missense mutations that disrupt protein activity. Using the more restrictive PAM of KTTG (SEQ ID NO: 3870), shown in Example 9 to be more active in mammalian cells, a total of 42 potential sgRNAs were identified within human hao-1 exons 1-4 (Table 4).

[0356] [Table 15]

[0357] Guides spanning intron / exon boundaries were included because they can create indels that interfere with splicing. Using Geneious Prime, the spacer sequences of these 42 guides were searched against the human genome, and the software assigned specificity scores based on alignment to the human genome. A higher specificity score indicates a lower probability of the guide recognizing one or more sequences in the human genome other than the site for which the spacer was designed. Specificity scores ranged from 10% to 100%, with 25 guides having specificity scores above 90% and 33 guides having specificity scores above 80%. This analysis indicates that guides targeting exon regions of human genes with high specificity scores can be easily identified, and several highly active guides are expected to be identified.

[0358] Example 11 - Comparison of the editing capabilities of nucleases described herein with spCas9 in mouse liver cells The CRISPR Cas9 nuclease from the bacterial species Streptococcus pyogenes (spCas9) is widely used for genome editing and is the most active RNA-guided nuclease identified. The relative efficacy of MG29-1 compared to spCas9 was evaluated by nucleofection of different doses of RNP in the mouse liver cell line Hepa1-6. An sgRNA targeting intron 1 of mouse albumin was used for both nucleases. For MG29-1, the sgRNA mAlb29-1-8 identified in Example 29 was selected. The guide mAlb29-1-8 (see Example 29) was chemically synthesized incorporating chemical modifications called AltR1 / AltR2 (Integrated DNA Technologies) designed to improve the efficacy of the guide for the V-type nuclease cpf1, which has a similar sgRNA structure to MG29-1. For spCas9, sgRNAs that efficiently edited mouse albumin intron 1 were identified by testing three guides selected from in silico screening. The spCas9 protein used in these studies was obtained from a commercial supplier (Integrated DNA technologies AltR-sPCas9).

[0359] The sgRNA mAlbR1 (spacer sequence TTAGTATAGCATGGTCGAGC) was chemically synthesized and incorporated chemical modifications consisting of 2'O-methyl bases and phosphorothioate (PS) linkages at the three bases on both ends of the guide, which improved its efficacy in cells. The mAlbR1 sgRNA generated indels at a frequency of 90% when an RNP consisting of 20 pmol of spCas9 protein and 50 pmol of guide was nucleofected into Hepa1-6 cells, indicating that it is a highly active guide. RNPs formed with nuclease protein ranging from 20 pmoles to 1 pmole and a constant protein-to-sgRNA ratio of 1:2.5 were nucleofected into Hepa1-6 cells. Indels at the target site in mouse albumin intron 1 were quantified using Sanger sequencing and ICE analysis of PCR-amplified genomic DNA. The results, shown in Figure 52, indicate that MG29-1 generated a higher percentage of indels than spCas9 at lower RNP doses when editing was not saturated. These data indicate that MG29-1 is at least as active, and potentially more active, than spCas9 in liver-derived mammalian cells.

[0360] Example 34 - Engineering sequence variants of the nucleases described herein and evaluation in mouse liver cells To improve the editing efficiency of MG29-1, a set of mutations substituting one or two amino acids was introduced into the MG29-1 coding region. The set of amino acid substitutions was determined by alignment to Acidaminococcus sp. Cas12a (AsCas12a). Structure-guided manipulation (Kleinstiver, et al., Nat Biotechnol. 2019, 37, 276-282) was used to substitute different amino acids in AsCas12a to alter or improve PAM binding. Four amino acid substitutions in AsCas12a: S170R, E174R, N577R, and K583R, showed higher editing efficiency with standard and non-standard PAMs. The sites corresponding to these substitutions were identified by multiple alignments in MG29-1 and corresponded to the following in MG29-1: S168R, E172R, N577R, and K583R.

[0361] To test single amino acid substitutions, a two-plasmid delivery system was used. Expression plasmids encoding MG29-1 with single amino acid substitutions were constructed using standard molecular cloning techniques. One plasmid encoded MG29-1 under the CMV promoter, and the second plasmid contained the mAlb29-1-8 sgRNA (see Table 8), which has high editing efficiency in Hepa 1-6 cells. Transcription of the guide was driven by the human U6 promoter. Confirmation of initial results from testing single and double amino acid substitutions using the two-plasmid system was performed using in vitro transcribed (IVT) mRNA encoding MG29-1 (see Example 11 for details on how the IVT mRNA was generated) and a chemically synthesized guide (synthesized by Integrated DNA technologies) incorporating AltR1 / AltR2 chemical modifications optimized by Integrated DNA Technologies for Cpf1. For delivery of the two-plasmid system, 100 ng of the plasmid encoding MG29-1 and 400 ng of the plasmid encoding the guide were mixed with Lipofectamine 3000, added to Hepa1-6 cells, and incubated for 3 days before genomic DNA isolation.

[0362] For delivery of IVT mRNA and synthetic guides, 300 ng of mRNA and 120 ng of synthetic guide were mixed with Lipofectamine Messenger Max, added to cells, and incubated for 2 days before genomic DNA isolation. The synthetic guides screened using IVT mRNA correspond to the guides detailed in Table 8, but for simplicity, the guide names in Figure 53 are shortened so that guide "mAlb29-1-1" is represented as g1-1, "mAlb29-1-8" is represented as g1-8, and so on. One guide targeting the human T-cell receptor locus (TRAC) was also tested (35TRAC in Figure 53D). The guide 35 TRAC spacer is GAGTCTCTCAGCTGGTACACGG (SEQ ID NO: 4268) with a TTTG PAM. Guide 35 TRAC was ordered with the same modifications as previously described. For MG29-1 editing of mouse albumin intron 1, genomic DNA and PCR amplification were performed as described in the previous example. For guide35 TRAC, the human TRAC locus was amplified with primer F: TGCTTTGCTGGGCCTTTTTC (SEQ ID NO: 4269) and primer R: ACAGTCTGAGCAAAGGCAGG (SEQ ID NO: 4270). The resulting 957 bp PCR product was purified as described above. Editing was assessed by Sanger sequencing using primer ATCACGAGCAGCTGGTTTCT (SEQ ID NO: 4271).

[0363] Editing efficiencies of mouse albumin intron 1 and the human TRAC locus were quantified using Sanger sequencing of PCR products followed by CRISPR editing inference (ICE). Data representing up to four biological replicates are plotted in Figures 53A-D. The single amino acid substitution S168R showed improved editing efficiency when using the guide mAlb29-1-8 in a two-plasmid system (Figure 53A). The mutation E172R did not provide significant improvement with the guide mAlb29-1-8, while the mutation K583R completely prevented editing with the mAlb29-1-8 guide. Transfection with MG29-1 mRNA and the synthetic guide mAlb29-1-8 confirmed the results of the plasmid transfection (Figure 53B). The single amino acid substitution S168R conferred higher editing efficiency across the different concentrations of mRNA tested with the guide mAlb29-1-8 (Figure 53B). Double amino acid substitutions of S168R with E172R (a substitution that did not impair activity, as seen in Figure 53A), or N577R (a substitution that was not tested in MG29-1 plasmid transfection but conferred higher editing efficiency of cpfl) and Y170R (a substitution that was hypothesized to improve editing efficiency based on the predicted MG29-1 protein structure) were tested and compared with the single S168R mutant.

[0364] Neither double mutation conferred improved editing efficiency under the conditions tested (Figure 53C). The editing efficiencies of the S168R variants of MG29-1 and MG29-1 WT were compared in parallel with 12 guides targeting mouse albumin intron 1 and one guide targeting the human T-cell receptor locus (TRAC). The S168R variant of MG29-1 exhibited improved editing efficiency with all 13 guides, with some guides benefiting more than others (Figure 4d). Importantly, S168R did not impair mammalian editing efficiency with any of the guides tested. These results indicate that the S168R (serine at amino acid position 168 is changed to arginine) variant of MG29-1 improves editing activity and is advantageous for identifying highly active guides for therapeutic use.

[0365] Example 35 - Identification of chemical modifications of sgRNAs of nucleases described herein that improve guide stability and improve editing efficiency in mammalian cells RNA molecules are inherently unstable in biological systems due to their susceptibility to cleavage by nucleases. Modification of the natural chemical structure of RNA has been widely used to improve the stability of RNA molecules used in RNA interference (RNAi) in the context of therapeutic drug development (Corey, J Clin Invest. 2007 Dec 3; 117(12): 3615-3622, J.B. Bramsen, J.Kjems Frontiers in Genetics, 3 (2012), p. 154). The introduction of chemical modifications into the nucleobase or phosphodiester backbone of RNA is crucial for improving the stability, and therefore efficacy, of short RNA molecules in vivo. A wide range of chemical modifications have been developed with different properties in terms of stability against nucleases and affinity for complementary DNA or RNA.

[0366] Similar chemical modifications have been applied to guide RNAs for CRISPR Cas9 nucleases (Hendel et al, Nat Biotechnol. 2015 Sep;33(9):985-989, Ryan et al Nucleic Acids Res 2018 Jan 25;46(2):792-803, Mir et al Nature Communications volume 9, Article number:2641(2018), O'Reilly et al Nucleic Acids Res 2019 47,546-558, Yin et al Nature Biotechnology volume 35, pages 1179-1187(2017), each of which is incorporated herein by reference in its entirety).

[0367] MG29-1 nuclease is a novel nuclease with limited amino acid sequence similarity to identified Type V CRISPR enzymes such as cpf1. The sequences of the structural (backbone) components of the guide RNA identified for MG29-1 are similar to those of cpf1 chemical modifications to the MG29-1 guide, which allow for improved stability, but no retention activity. A series of chemical modifications of the MG29-1 sgRNA were designed to assess their effect on sgRNA activity in mammalian cells and stability in the presence of mammalian cell protein extracts.

[0368] We selected the sgRNA mAlb29-1-8, which was highly active in the mouse liver cell line Hepa1-6 when the guide contained a commercially available set of proprietary chemical modifications developed by IDT (called AltR1 / AltR2) designed to improve the activity of guide RNAs against cpf1. We chose to test two chemical modifications of the nucleobase: 2'-O-methyl, in which the 2' hydroxyl group is replaced with a methyl group, and 2'-fluoro, in which the 2' hydroxyl group is replaced with a fluorine. Both 2'-O-methyl and 2'-fluoro modifications improve resistance to nucleases. The 2'-O-methyl modification is a naturally occurring post-transcriptional modification of RNA that improves the binding affinity of RNA:RNA duplexes but has little effect on RNA:DNA stability. 2'-Fluoro modified bases have reduced immunostimulatory effects and increase the binding affinity of both RNA:RNA and RNA:DNA hybrids (see, e.g., Pallan et al. Nucleic Acids Res 2011 Apr;39(8):3482-95; Chen et al. Scientific Reports volume 9, Article number: 6078 (2019); Kawasaki, AM et al. J Med Chem 36, 831-841 (1993)).

[0369] The inclusion of phosphorothioate (PS) linkages instead of phosphodiester bonds between bases was also evaluated, as PS linkages improve resistance to nucleases (Monia et al., Nucleic Acids, Protein Synthesis, and Molecular Genetics, Volume 271, Issue 24, P14533-14540, June 14, 1996).

[0370] The predicted secondary structure of the MG29-1 sgRNA with a spacer targeting mouse albumin intron 1 (mAlb-1-8) is shown in Figure 54. Based on the sequence structure of other CRISPR-cas systems, the stem-loop of the guide scaffold was predicted to be important for interaction with the MG29-1 protein. Based on the secondary structure, a series of chemical modifications were designed at different structural and functional regions of the guide. A modular approach was taken to allow for the initial testing of guides with fewer chemical modifications, indicating that the structural and functional regions of the guide could tolerate different chemical modifications without significant loss of activity. The structural and functional regions were defined as follows: The 3' and 5' ends of the guide are targets for exonucleases and can be protected by various chemical modifications, including 2'-O-methyl and PS linkages, an approach that has been used to improve the stability of spCas9 guides (Hendel et al., Nat Biotechnol. 2015 Sep;33(9):985-989).

[0371] Sequences containing both half of the stem and loop in the guide scaffold region were selected for modification. The spacer was divided into a seed region (the first six nucleotides closest to the PAM) and the remaining 16 nucleotides of the spacer (referred to as the non-seed region). A total of 43 guides were designed, and 39 were synthesized. All 43 guides contained the same nucleotide sequence but had different chemical modifications. The editing activity of 39 of the guides was evaluated in Hepa1-6 cel...

Claims

1. A method for disrupting the CD38 locus in a cell, the method comprising introducing into the cell: (a) a class 2, type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a region of the CD38 locus, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4466-4503 and 5686, or wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4428-4465 and 5685.

2. The method according to claim 1, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof.

3. The method according to claim 1, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof.

4. The method according to claim 1, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, and 6033-6036.

5. The method according to claim 1, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotide of SEQ ID NO: 3609.

6. The method according to claim 1, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20-22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 4466, 4467, 4468, 4479, 4484, 4490, 4492, 4493, 4495, 4498.

7. The method according to claim 1, wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least 80% sequence identity with any one of SEQ ID NOs: 4428, 4429, 4430, 4436, 4441, 4446, 4452, 4454, 4455, 4460, or 4461.

8. A method for disrupting the TIGIT locus in a cell, comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease, and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a region of the TIGIT locus. The engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 4521 to 4537, or The engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 4504 to 4520, method. **Claim 9**: The method according to claim 8, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 10**: The method according to claim 8, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 11**: The method according to claim 8, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551 to 3559, 3608 to 3609, 3612, 3636 to 3637, 3640 to 3641, 3644 to 3645, 3648 to 3649, 3652 to 3653, 3656 to 3657, 3660 to 3661, 3664 to 3667, 3671 to 3672, 3678, 3695 to 3696, 3729 to 3730, 3734 to 3735, 3851 to 3857, and 6033 to 6036. **Claim 12**: The method according to claim 8, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 13**: The method according to claim 8, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NO: 4521, 4527, 4528, 4535, or 4536. **Claim 14**: The method according to claim 8, wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least 80% sequence identity with any one of SEQ ID NO: 4504, 4510, 4511, 4518, or 4519. **Claim 15**: A method of disrupting the AAVS1 locus in a cell, the method comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a region of the AAVS1 locus. The engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NO: 4569 - 4599, or the engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NO: 4538 - 4568. **Claim 16**: The method according to claim 15, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 17**: The method according to claim 15, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 18**: The method according to claim 15, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551 - 3559, 3608 - 3609, 3612, 3636 - 3637, 3640 - 3641, 3644 - 3645, 3648 - 3649, 3652 - 3653, 3656 - 3657, 3660 - 3661, 3664 - 3667, 3671 - 3672, 3678, 3695 - 3696, 3729 - 3730, 3734 - 3735, 3851 - 3857, 6033 - 6036. **Claim 19**: The method according to claim 15, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotide of SEQ ID NO: 3609. **Claim 20**: The method according to claim 15, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 - 22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 4574, 4577, 4578, 4579, 4582, 4584, 4585, 4586, 4587, 4589, 4590, 4591, 4592, 4593, 4595, 4596, or 4598. **Claim 21**: The method according to claim 15, wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least 80% sequence identity with any one of SEQ ID NOs: 4543, 4546, 4547, 4548, 4551, 4553, 4554, 4555, 4556, 4558, 4559, 4560, 4561, 4562, 4565, or 4567. **Claim 22**: A method for disrupting the B2M locus in a cell, the method comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a region of the B2M locus. The engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 4676 - 4751, or the engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 4600 - 4675. **Claim 23**: The method according to claim 22, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 - 3470 or a variant thereof. **Claim 24**: The method according to claim 22, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 - 1722 or a variant thereof. **Claim 25**: The method according to claim 22, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, and 6033-6036. **Claim 26**: The method according to claim 22, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 27**: The method according to claim 22, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20-22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 4676, 4678-4687, 4690, 4692, 4698-4707, 4720-4723, 4725-4726, 4732-4733, 4736-4737, 4741, or 4750-4751. **Claim 28**: The method according to claim 22, wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least 80% sequence identity with any one of SEQ ID NOs: 4600, 4602-4611, 4614, 4616, 4622-4631, 4644-4647, 4649-4650, 4656-4657, 4660-4661, 4665, or 4674-4675. **Claim 29**: A method for disrupting the CD2 locus in a cell, the method comprising introducing into the cell: **(a)** a Class 2, Type V Cas endonuclease, and introducing an engineered guide ribonucleic acid (gRNA) that is configured to form a complex with the endonuclease and contains a spacer sequence configured to hybridize to a region of the CD2 locus, wherein the engineered gRNA is configured to hybridize to a sequence having at least 20 to 22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4837 to 4921, or the engineered gRNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4752 to 4836. **Claim 30**: The method according to claim 29, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 31**: The method according to claim 29, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 32**: The method according to claim 29, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551 - 3559, 3608 - 3609, 3612, 3636 - 3637, 3640 - 3641, 3644 - 3645, 3648 - 3649, 3652 - 3653, 3656 - 3657, 3660 - 3661, 3664 - 3667, 3671 - 3672, 3678, 3695 - 3696, 3729 - 3730, 3734 - 3735, 3851 - 3857, and 6033 - 6036. **Claim 33**: The method according to claim 29, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 34**: The method according to claim 29, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 - 22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 4837, 4844, 4845, 4848, 4857 - 4858, 4883, 4887, 4892 - 4893, 4904 - 4909, 4914, 4916, or 4918. **Claim 35**: The method according to claim 29, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the engineered guide RNAs from Table 14E that targets any one of SEQ ID NOs: 4837, 4844, 4845, 4848, 4857 - 4858, 4883, 4887, 4892 - 4893, 4904 - 4909, 4914, 4916, or 4918. **Claim 36**: A method of disrupting the CD5 locus in a cell, the method comprising introducing into the cell **(a)** a Class 2, Type V Cas endonuclease, and introducing an engineered guide ribonucleic acid (gRNA) that is configured to form a complex with the endonuclease and contains a spacer sequence configured to hybridize to a region of the CD5 locus; wherein the engineered gRNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4946-4969, or the engineered gRNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4922-4945. **Claim 37**: The method according to claim 36, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. **Claim 38**: The method according to claim 36, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof. **Claim 39**: The method according to claim 36, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, and 6033-6036. **Claim 40**: The method according to claim 36, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 41**: The method according to claim 36, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20-22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 4946-4947, 4949, 4951, 4957-4960, 4963, 4967, or 4969. **Claim 42**: The method according to claim 36, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the engineered guide ribonucleic acids from Table 14F that targets any one of SEQ ID NOs: 4946-4947, 4949, 4951, 4957-4960, 4963, 4967, or 4969. **Claim 43**: A method of disrupting the mouse TRAC locus in a cell, the method comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a region of the mouse TRAC locus. The engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5126-5195, 5682, or 5684, or The engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5056-5125, 5681, or 5683, method. **Claim 44**: The method according to claim 43, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1-3470 or a variant thereof. **Claim 45**: The method according to claim 43, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof. **Claim 46**: The method according to claim 43, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, or 6033-6036. **Claim 47**: The method according to claim 43, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with a non-degenerate nucleotide of SEQ ID NO: 3609. **Claim 48**: The method according to claim 43, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20-22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 5126-5130, 5133-5143, 5147-5150, 5172-5173, 5184-5189, or 5192-5194. **Claim 49**: The method according to claim 43, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the engineered guide ribonucleic acids from Table 14G that targets any one of SEQ ID NOs: 5126-5130, 5133-5143, 5147-5150, 5172-5173, 5184-5189, or 5192-5194. **Claim 50**: A method of disrupting the mouse TRBC1 or TRBC2 locus in a cell, the method comprising introducing into the cell: **(a)** a Class 2, Type V Cas endonuclease, and (b) An engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a region of the mouse TRBC1 or TRBC2 locus, and introducing the engineered guide ribonucleic acid, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5211-5225 or 5247-5267, or wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5196-5210 or 5226-5246, a method. **Claim 51**: The method according to claim 50, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1-3470 or a variant thereof. **Claim 52**: The method according to claim 50, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof. **Claim 53**: The method according to claim 50, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3677-3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, or 6033-6036. **Claim 54**: The method according to claim 50, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 55**: The method according to claim 50, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20-22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 5211, 5213-5215, 5217, 5221, 5223, 5247, 5249-5250, 5252-5253, 5258-5259, or 5264. **Claim 56**: The method according to claim 50, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the guide ribonucleic acids from Table 14H that targets any one of SEQ ID NOs: 5211, 5213-5215, 5217, 5221, 5223, 5247, 5249-5250, 5252-5253, 5258-5259, or 5264. **Claim 57**: A method of disrupting the human TRBC1 or TRBC2 locus in a cell, the method comprising introducing into the cell: **(a)** a Class 2, Type V Cas endonuclease, and introducing an engineered guide ribonucleic acid (gRNA) that is configured to form a complex with the endonuclease and is configured to hybridize to a region of the human TRBC1 or TRBC2 locus, and the engineered gRNA is configured to hybridize to a sequence having at least 20 to 22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5661-5679, or the engineered gRNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5642-5660. **Claim 58**: The method according to claim 57, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. **Claim 59**: The method according to claim 57, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof. **Claim 60**: The method according to claim 57, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, or 6033-6036. **Claim 61**: The method according to claim 57, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 62**: The method according to claim 57, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20-22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 5661-5663, 5672-5675, or 5678. **Claim 63**: The method according to claim 57, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the engineered guide ribonucleic acids from Table 14I that targets any one of SEQ ID NOs: 5661-5663, 5672-5675, or 5678. **Claim 64**: A method for disrupting the HPRT locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease, and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a region of the HPRT locus. The engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5562 to 5641, or The engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5482 to 5561, method. **Claim 65**: The method according to claim 64, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 66**: The method according to claim 64, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 67**: The method according to claim 64, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551 to 3559, 3608 to 3609, 3612, 3636 to 3637, 3640 to 3641, 3644 to 3645, 3648 to 3649, 3652 to 3653, 3656 to 3657, 3660 to 3661, 3664 to 3667, 3671 to 3672, 3678, 3695 to 3696, 3729 to 3730, 3734 to 3735, 3851 to 3857, or 6033 to 6036. **Claim 68**: The method according to claim 64, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 69**: The method according to claim 64, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 5562-5564 or 5568. **Claim 70**: The method according to claim 64, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the engineered guide ribonucleic acids from Table 14J that targets any one of SEQ ID NOs: 5562-5564 or 5568. **Claim 71**: A method of disrupting the APO-A1 locus in a cell, the method comprising introducing into the cell: (a) a class 2, type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a region of the APO-A1 locus. The engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5861-5874, or The engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5847-5860. **Claim 72**: The method according to claim 71, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 73**: The method according to claim 71, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 74**: The method according to claim 71, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551 - 3559, 3608 - 3609, 3612, 3636 - 3637, 3640 - 3641, 3644 - 3645, 3648 - 3649, 3652 - 3653, 3656 - 3657, 3660 - 3661, 3664 - 3667, 3671 - 3672, 3678, 3695 - 3696, 3729 - 3730, 3734 - 3735, 3851 - 3857, or 6033 - 6036. **Claim 75**: The method according to claim 71, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotide of SEQ ID NO: 3609. **Claim 76**: The method according to claim 71, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 - 22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 5861 - 5866 or 5868 - 5869. **Claim 77**: The method according to claim 71, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the engineered guide ribonucleic acids from Table 43A that targets any one of SEQ ID NOs: 5861 - 5866 or 5868 - 5869. **Claim 78**: A method of disrupting the ANGPTL3 locus in a cell, the method comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease; and introducing an engineered guide ribonucleic acid (gRNA) that is configured to form a complex with the endonuclease and that includes a spacer sequence configured to hybridize to a region of the ANGPTL3 locus; wherein the engineered gRNA is configured to hybridize to a sequence having at least 20-22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5953-6030, or the engineered gRNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5875-5952. **Claim 79**: The method of claim 78, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. **Claim 80**: The method of claim 78, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof. **Claim 81**: The method according to claim 78, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551 - 3559, 3608 - 3609, 3612, 3636 - 3637, 3640 - 3641, 3644 - 3645, 3648 - 3649, 3652 - 3653, 3656 - 3657, 3660 - 3661, 3664 - 3667, 3671 - 3672, 3678, 3695 - 3696, 3729 - 3730, 3734 - 3735, 3851 - 3857, or 6033 - 6036. **Claim 82**: The method according to claim 78, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with a non-degenerate nucleotide of SEQ ID NO: 3609. **Claim 83**: The method according to claim 78, wherein the engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20 - 22 consecutive nucleotides having at least 80% sequence identity with any one of SEQ ID NOs: 5955 - 5963, 5968 - 5975, 5979 - 5987, 5989 - 5993, 5997, 5999, 6003 - 6010, 6014 - 6016, 6024 - 6025, or 6027 - 6030. **Claim 84**: The method according to claim 78, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with any one of the engineered guide ribonucleic acids from Table 43B that targets any one of SEQ ID NOs: 5955 - 5963, 5968 - 5975, 5979 - 5987, 5989 - 5993, 5997, 5999, 6003 - 6010, 6014 - 6016, 6024 - 6025, or 6027 - 6030. **Claim 85**: A method for disrupting the human Rosa26 locus in a cell, the method comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease, and introducing an engineered guide ribonucleic acid (gRNA) that is configured to form a complex with the endonuclease and that includes a spacer sequence configured to hybridize to a region of the human Rosa26 locus, wherein the engineered gRNA is configured to hybridize to a sequence having at least 20 to 22 consecutive nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5013 - 5055, or the engineered gRNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 4970 - 5012. **Claim 86** The method of claim 85, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1 - 3470 or a variant thereof. **Claim 87** The method of claim 85, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 - 1722 or a variant thereof. **Claim 88**: The method according to claim 85, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, or 6033-6036. **Claim 89**: The method according to claim 85, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 90**: A method for disrupting the FAS locus in a cell, the method comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a region of the FAS locus. The engineered guide ribonucleic acid is configured to hybridize to a sequence having at least 20-22 consecutive nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5367-5465, or A method, wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NOs: 5268 to 5366. **Claim 91**: The method according to claim 90, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 92**: The method according to claim 90, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 93**: The method according to claim 90, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with the non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551 to 3559, 3608 to 3609, 3612, 3636 to 3637, 3640 to 3641, 3644 to 3645, 3648 to 3649, 3652 to 3653, 3656 to 3657, 3660 to 3661, 3664 to 3667, 3671 to 3672, 3678, 3695 to 3696, 3729 to 3730, 3734 to 3735, 3851 to 3857, or 6033 to 6036. **Claim 94**: The method according to claim 90, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotide of SEQ ID NO: 3609. **Claim 95**: A method for disrupting the PD-1 locus in a cell, comprising introducing into the cell (a) a Class 2, Type V Cas endonuclease, and introducing an engineered guide ribonucleic acid (gRNA) that is configured to form a complex with the endonuclease and that includes a spacer sequence configured to hybridize to a region of the PD-1 locus, wherein the engineered gRNA is configured to hybridize to a sequence having at least 20 to 22 contiguous nucleotides having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5474-5481, or wherein the engineered gRNA comprises a nucleotide sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 5466-5473. **Claim 96**: The method according to claim 95, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-3470 or a variant thereof. **Claim 97**: The method according to claim 95, wherein the endonuclease comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof. **Claim 98**: The method according to claim 95, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551 - 3559, 3608 - 3609, 3612, 3636 - 3637, 3640 - 3641, 3644 - 3645, 3648 - 3649, 3652 - 3653, 3656 - 3657, 3660 - 3661, 3664 - 3667, 3671 - 3672, 3678, 3695 - 3696, 3729 - 3730, 3734 - 3735, 3851 - 3857, or 6033 - 6036. **Claim 99**: The method according to claim 95, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 100**: An engineered nuclease system comprising: (a) an endonuclease having at least 75% sequence identity with any one of SEQ ID NO: 215 or a variant thereof; (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a target nucleic acid sequence, wherein the engineered nuclease system has reduced immunogenicity when administered to a subject as compared to an equivalent system comprising a Cas9 enzyme. **Claim 101**: The engineered nuclease system according to claim 100, wherein the Cas9 enzyme is an SpCas9 enzyme. **Claim 102**: The engineered nuclease system according to claim 100 or 101, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 103**: The engineered nuclease system according to claim 100, wherein the immunogenicity is antibody immunogenicity. **Claim 104**: A method of disrupting the mouse HAO-1 locus in a cell, the method comprising introducing into the cell: (a) a class 2, type V Cas endonuclease; introducing an engineered guide ribonucleic acid (gRNA) that is configured to form a complex with the endonuclease and that hybridizes to a region of the mouse HAO-1 locus, and wherein the engineered gRNA comprises nucleotide modifications set forth in Table 25, and comprises nucleotides of gRNAs mH29-1_37, mH29-15_37, mH29-29_37 from Table 25, or wherein the engineered gRNA comprises any one of SEQ ID NOs: 4184 to 4225. **Claim 105**: The method of claim 104, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 106**: The method of claim 104, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 107**: The method of claim 104, wherein the engineered gRNA comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 3471, 3539, 3551 - 3559, 3608 - 3609, 3612, 3636 - 3637, 3640 - 3641, 3644 - 3645, 3648 - 3649, 3652 - 3653, 3656 - 3657, 3660 - 3661, 3664 - 3667, 3671 - 3672, 3678, 3695 - 3696, 3729 - 3730, 3734 - 3735, 3851 - 3857, or 6033 - 6036. **Claim 108**: The method of claim 104, wherein the engineered gRNA comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 109**: The method according to claim 104, wherein the engineered guide ribonucleic acid comprises nucleotides of guide ribonucleic acid mH29-15_37 or mH29-29_37 from Table 25, which comprise the nucleotide modifications described in Table 25. **Claim 110**: The method according to claim 104, further comprising disrupting the expression of glycolic acid oxidase from the HAO-1 locus. **Claim 111**: A method of disrupting the human TRAC locus in a cell, the method comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a region of the human TRAC locus, wherein the engineered guide ribonucleic acid comprises nucleotides of MG29-1-TRAC-sgRNA-35 from Table 28B, which comprise the nucleotide modifications described in Table 28B. **Claim 112**: The method according to claim 111, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1 to 3470 or a variant thereof. **Claim 113**: The method according to claim 111, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 to 1722 or a variant thereof. **Claim 114**: The method according to claim 111, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of the non-degenerate nucleotides of SEQ ID NOs: 3471, 3539, 3551-3559, 3608-3609, 3612, 3636-3637, 3640-3641, 3644-3645, 3648-3649, 3652-3653, 3656-3657, 3660-3661, 3664-3667, 3671-3672, 3678, 3695-3696, 3729-3730, 3734-3735, 3851-3857, or 6033-6036. **Claim 115**: The method according to claim 111, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non-degenerate nucleotides of SEQ ID NO: 3609. **Claim 116**: A method for disrupting the albumin locus in a cell, the method comprising introducing into the cell: (a) a Class 2, Type V Cas endonuclease; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and to hybridize to a region of the albumin locus. The engineered guide ribonucleic acid comprises nucleotides of mAlb298-37, mAlb2912-37, mAlb2918-37, or mAlb298-34 from Table 29, which comprise the nucleotide modifications described in Table 29, or the engineered guide ribonucleic acid comprises nucleotides of mAlb29-8-44, mAlb29-8-50, mAlb29-8-50b, mAlb29-8-51b, mAlb29-8-52b, mAlb29-8-53b, or mAlb29-8-54b from Table 46, which comprise the nucleotide modifications described in Table 46. **Claim 117**: The method according to claim 116, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 1-3470 or a variant thereof. **Claim 118**: The method according to claim 116, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 - 1722 or a variant thereof. **Claim 119**: The method according to claim 116, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with the non - degenerate nucleotides of any one of SEQ ID NOs: 3471, 3539, 3551 - 3559, 3608 - 3609, 3612, 3636 - 3637, 3640 - 3641, 3644 - 3645, 3648 - 3649, 3652 - 3653, 3656 - 3657, 3660 - 3661, 3664 - 3667, 3671 - 3672, 3678, 3695 - 3696, 3729 - 3730, 3734 - 3735, 3851 - 3857, or 6033 - 6036. **Claim 120**: The method according to claim 116, wherein the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with the non - degenerate nucleotides of SEQ ID NO: 3609. **Claim 121**: The method according to any one of claims 1 - 99 or claims 104 - 120, wherein the cell is a eukaryotic cell, a hepatocyte, a T cell, a hematopoietic stem cell, or a precursor thereof. **Claim 122**: The method according to any one of claims 116 - 120, wherein the engineered guide ribonucleic acid comprises nucleotide modifications described in Table 29 and comprises nucleotides of mAlb298 - 37, mAlb2912 - 37, mAlb2918 - 37, or mAlb298 - 34 from Table 29. **Claim 123**: An engineered guide ribonucleic acid, (a) comprising a DNA targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule, and (b) a protein - binding segment configured to bind to a class 2, type V Cas endonuclease, wherein the engineered guide ribonucleic acid comprises a nucleotide modification pattern shown in any one of SEQ ID NOs: 5695 - 5701 of Table 34.

124. The engineered guide ribonucleic acid according to claim 123, wherein the engineered guide ribonucleic acid comprises mAlb29-8-44, mAlb29-8-50, mAlb29-8-37, or mAlb29-12-44.

125. The engineered guide ribonucleic acid according to claim 123, wherein the engineered guide ribonucleic acid comprises hH29-4_50, hH29-21_50, hH29-23_50, hH29-41_50, hH29-4_50b, hH29-21_50b, hH29-23_50b, or hH29-41_50b, mH29-1-50, mH29-15-50, mH29-29-50, mH29-1-50b, mH29-15-50b, or mH29-29-50b.

126. The engineered guide ribonucleic acid according to claim 123, wherein the DNA targeting segment is configured to hybridize to the HAO-1 gene or the albumin gene.

127. The engineered guide ribonucleic acid according to claim 123, wherein the endonuclease comprises a sequence having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711-1722 or a variant thereof.

128. The engineered guide ribonucleic acid according to any one of claims 123-127, wherein the endonuclease comprises a sequence having at least 75% sequence identity with SEQ ID NO:

215.

129. An engineered nuclease system, (a) an endonuclease having at least 75% sequence identity with any one of SEQ ID NOs: 1-3470 or a variant thereof, or a nucleotide sequence encoding the endonuclease, and (b) a polynucleotide sequence encoding a CRISPR array, wherein the CRISPR array is configured to be processed by an engineered guide ribonucleic acid that is engineered by the endonuclease, the engineered guide ribonucleic acid is configured to form a complex with the endonuclease, and the engineered guide ribonucleic acid comprises a spacer sequence configured to hybridize to a target nucleic acid sequence, The engineered nuclease system, wherein the spacer sequence is configured to hybridize to the albumin gene.

130. The engineered nuclease system according to claim 129, wherein the polynucleotide sequence comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with SEQ ID NO: 5712.

131. The engineered nuclease system according to claim 129, wherein the endonuclease comprises an endonuclease having at least 75% sequence identity with any one of SEQ ID NOs: 141, 215, 229, 261, or 1711 - 1722 or a variant thereof.

132. The engineered nuclease system according to any one of claims 129 - 131, wherein the endonuclease comprises an endonuclease having at least 75% sequence identity with SEQ ID NO:

215.

133. An engineered nuclease system, comprising: (a) an endonuclease having at least 75% sequence identity with SEQ ID NO: 470 or a variant thereof; and (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a target nucleic acid sequence.

134. The engineered nuclease system according to claim 133, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with SEQ ID NO: 6031.

135. The engineered nuclease system according to claim 133 or 134, wherein the endonuclease is configured to be selective for a 5' PAM sequence comprising SEQ ID NO: 6032.

136. An engineered nuclease system, comprising: (a) an endonuclease having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NO: 2824, 2841, or 2896 or a variant thereof; (b) an engineered guide ribonucleic acid comprising a spacer sequence configured to form a complex with the endonuclease and configured to hybridize to a target nucleic acid sequence, and an engineered guide ribonucleic acid, wherein the engineered guide ribonucleic acid comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity with any one of SEQ ID NO: 6033, 6034, or 6035, an engineered nuclease system.

137. The engineered nuclease system according to claim 136, wherein the endonuclease comprises a sequence having at least 80% sequence identity with SEQ ID NO: 2824, and the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with SEQ ID NO: 6033.

138. The engineered nuclease system according to claim 136, wherein the endonuclease comprises a sequence having at least 80% sequence identity with SEQ ID NO: 2841, and the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with SEQ ID NO: 6034.

139. The engineered nuclease system according to claim 136, wherein the endonuclease comprises a sequence having at least 80% sequence identity with SEQ ID NO: 2896, and the engineered guide ribonucleic acid comprises a sequence having at least 80% sequence identity with SEQ ID NO: 6035.

140. The engineered nuclease system according to any one of claims 136 to 139, wherein the endonuclease is configured to be selective for a 5' PAM sequence containing any one of SEQ ID NOs: 6037 to 6039.

141. A lipid nanoparticle, comprising: (a) any one of the endonucleases described herein; (b) any one of the engineered guide ribonucleic acids described herein; (c) a cationic lipid; (d) a sterol; (e) a neutral lipid; (f) a PEGylated lipid.

142. The lipid nanoparticle according to claim 141, wherein the cationic lipid contains C12-200, the sterol contains cholesterol, the neutral lipid contains DOPE, or the PEGylated lipid contains DMG-PEG2000.

143. The lipid nanoparticle according to claim 141, wherein the cationic lipid contains any one of the cationic lipids shown in FIG. 109.