Class II and V CRISPR systems
Patent Information
- Application Number
- JP2024507845
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-08
- Filing Date
- 2022-09-06
- Publication Date
- 2025-09-16
AI Technical Summary
The existing CRISPR/Cas system has problems of insufficient specificity and efficiency in gene editing, especially the function of uncultured microbial-derived Cas proteins is not fully explored, which limits its application in gene editing technology.
A series of new category 2 V-CRISPR/Cas systems have been developed, including MG90, MG91, MG118, MG119, MG120, MG122 and MG126, extracted from the metagenome of uncultured microorganisms, and gene editing is used to design specific guide RNA and Cas protein complexes to achieve efficient identification and cleavage of target sequences.
It improves the specificity and efficiency of gene editing, expands the diversity of Cas proteins, simplifies gene editing operations, reduces non-target effects, and enhances the application potential in complex environments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications This application is related to PCT Patent Application No. PCT / US2021 / 021259 and PCT Patent Application No. PCT / US2022 / 031849, each of which is incorporated by reference in its entirety herein.
[0002] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 241,928, entitled "CLASS II, TYPE V CRISPR SYSTEMS," filed September 8, 2021, which is incorporated by reference in its entirety herein. [Background technology]
[0003] Cas enzymes, along with associated clustered regularly interspaced short palindromic repeats (CRISPR)-guided ribonucleic acid (RNA), appear to be widespread components of the immune system of prokaryotes (~45% of bacteria, ~84% of archaea), and serve to protect these microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, by CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. Although CRISPR DNA elements have been observed as early as 1987, the programmable endonuclease cleavage capabilities of CRISPR / Cas complexes have only been recognized relatively recently, leading to the use of recombinant CRISPR / Cas systems in a variety of DNA engineering and gene editing applications.
[0004] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, which is incorporated herein by reference in its entirety. The XML copy created on September 6, 2022 is named 55921-732601_revised_2.xml and is 1,114,268 bytes in size. Summary of the Invention
[0005] In some aspects, the disclosure provides an engineered nuclease system comprising an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof; and an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a target nucleic acid sequence.
[0006] In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 410-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, and 474. In some embodiments, the endonuclease has at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 30-33, 39, 48, 56, 57, 61, 83, 92, 100, 110, 124, 136, 145, 148, 424, 425, 429, 476, or 629. In some embodiments, the guide RNA comprises a sequence having at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, and 474.
[0007] In some aspects, the present disclosure provides an engineered nuclease system, comprising an engineered guide RNA comprising a sequence having at least 80% sequence identity with a non-degenerate nucleotide of any one of SEQ ID NOs: 410-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, and 474, and a class 2, V-type Cas endonuclease configured to bind to the engineered guide RNA. In some embodiments, the engineered nuclease system further comprises a DNA repair template comprising a double-stranded DNA segment adjacent to one or two single-stranded DNA segments. In some embodiments, the single-stranded DNA segment is conjugated to the 5' end of the double-stranded DNA segment. In some embodiments, the single-stranded DNA segment is conjugated to the 3' end of the double-stranded DNA segment. In some embodiments, the single-stranded DNA segment has a length of 4 to 10 nucleotide bases.
[0008] In some embodiments, the single stranded DNA segment has a nucleotide sequence that is complementary to a sequence within the spacer sequence. In some embodiments, the double stranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein coding sequence, an miRNA coding sequence, an RNA coding sequence, or a transgene. In some embodiments, the double stranded DNA sequence is flanked by nuclease cleavage sites. In some embodiments, the nuclease cleavage sites comprise a spacer and a PAM sequence. In some embodiments, the PAM comprises any one of SEQ ID NOs: 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, and 475. In some embodiments, the system comprises a Mg 2+In some embodiments, the guide RNA comprises a hairpin comprising at least 8, at least 10, or at least 12 base pair ribonucleotides. In some embodiments, the hairpin comprises 10 base pair ribonucleotides. In some embodiments, the endonuclease comprises a sequence that is at least 75%, 80%, or 90% identical to any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a variant thereof, and the guide RNA structure comprises a sequence that is at least 80% or 90% identical to a non-degenerate nucleotide of any one of SEQ ID NOs: 410-419. In some embodiments, the endonuclease has at least about 75%, at least about 80%, at least about 85%, at least about at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs:30-33, 39, 48, 56, 57, 61, 83, 92, 100, 110, 124, 136, 145, 148, 424, 425, 429, 476, or 629. wherein the guide RNA structure comprises a sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 414 to 419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, and 474.In some embodiments, the sequence identity is determined by the BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW algorithm with Smith-Waterman homology search algorithm parameters. In some embodiments, the sequence identity is determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and a BLOSUM62 scoring matrix setting gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.
[0009] In some aspects, the disclosure provides an engineered guide ribonucleic acid (RNA) polynucleotide comprising a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule, and a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide, wherein the engineered guide ribonucleic acid polynucleotide is capable of forming a complex with a type 2, class V Cas endonuclease. In some embodiments, the type 2, class V Cas endonuclease is derived from an uncultured organism. In some embodiments, the Cas endonuclease has at least 75% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, and targets the complex to the target sequence of the target DNA molecule. In some embodiments, the DNA-targeting segment is located 3' to both of the two complementary stretches of nucleotides. In some embodiments, the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to a non-degenerate nucleotide of SEQ ID NOs: 410-419. In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 5, at least 8, at least 10, or at least 12 ribonucleotides.
[0010] In some aspects, the present disclosure provides a deoxyribonucleic acid polynucleotide encoding any of the engineered guide RNAs disclosed herein.
[0011] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding a class 2, type V Cas endonuclease, the endonuclease being derived from an uncultured microorganism, and the organism not being the uncultured organism. In some embodiments, the endonuclease comprises a variant having at least 70% or at least 80% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 630-645. In some embodiments, the NLS comprises SEQ ID NO: 631. In some embodiments, the NLS is proximal to the N-terminus of the endonuclease. In some embodiments, the NLS comprises SEQ ID NO: 630. In some embodiments, the NLS is proximal to the C-terminus of the endonuclease, hi some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.
[0012] In some aspects, the disclosure provides an engineered vector comprising a nucleic acid sequence encoding a Class 2, V-type Cas endonuclease, wherein the endonuclease is derived from an uncultured microorganism.
[0013] In some aspects, the disclosure provides an engineered vector comprising any of the nucleic acids disclosed herein. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, a lentivirus, or an adenovirus.
[0014] In some aspects, the disclosure provides a cell comprising any of the engineered vectors disclosed herein.
[0015] In some aspects, the disclosure provides a method of producing an endonuclease comprising culturing any of the cells disclosed herein.
[0016] In some aspects, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising contacting the double-stranded deoxyribonucleic acid polynucleotide with a class 2, V-type Cas endonuclease complexed with the endonuclease and an engineered guide RNA configured to bind to the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the guide RNA structure comprises a sequence that is at least 80% or 90% identical to a non-degenerate nucleotide of any one of SEQ ID NOs: 410-419. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand that comprises a sequence that is complementary to a sequence of the engineered guide RNA and a second strand that comprises the PAM. In some embodiments, the PAM is immediately adjacent to the 5' end of the sequence that is complementary to the sequence of the engineered guide RNA. In some embodiments, the PAM comprises any one of SEQ ID NOs: 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, and 475. In some embodiments, the class 2, V-type Cas endonuclease is from an uncultured microorganism. In some embodiments, the class 2, V-type Cas endonuclease further comprises a PAM interaction domain. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0017] In some aspects, the disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering an engineered nuclease system according to any one of claims 1 to 29 to the target nucleic acid locus, the endonuclease being configured to form a complex with the engineered guide ribonucleic acid structure, the complex being configured to modify the target nucleic acid locus upon binding of the complex to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a human cell, or a primary cell. In some embodiments, the cell is a primary cell. In some embodiments, the primary cell is a T cell. In some embodiments, the primary cell is a hematopoietic stem cell (HSC). In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering any nucleic acid disclosed herein or any vector disclosed herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease. In some embodiments, the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease.In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding the engineered guide RNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease induces a single-stranded or double-stranded break at or proximal to the target locus. In some embodiments, the endonuclease induces a staggered single-stranded break within or 3' of the target locus.
[0018] In some aspects, the disclosure provides a host cell comprising an open reading frame encoding a heterologous endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof. In some embodiments, the endonuclease has at least 75% sequence identity to any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a variant thereof. In some embodiments, the endonuclease has at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 30-33, 39, 48, 56, 57, 61, 83, 92, 100, 110, 124, 136, 145, 148, 424, 425, 429, 476, or 629. In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cells are λDE3 lysogens or the E. coli cells are a BL21(DE3) strain. In some embodiments, the E. coli cells have an ompT lon genotype. In some embodiments, the open reading frame is operably linked to a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araPBAD promoter, a strong leftward promoter from phage lambda (pL promoter), or any combination thereof. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in frame to a sequence encoding the endonuclease.In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in frame to the sequence encoding the endonuclease via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease (PSP) cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the open reading frame is codon-optimized for expression in the host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of the host cell.
[0019] In some aspects, the disclosure provides a culture comprising any of the host cells disclosed herein in a compatible liquid medium.
[0020] In some aspects, the disclosure provides a method of producing an endonuclease comprising culturing any of the host cells disclosed herein in a suitable growth medium. In some embodiments, the method further comprises inducing expression of the endonuclease by adding additional chemical agents or increased amounts of nutrients. In some embodiments, the method further comprises isolating the host cells after the culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to IMAC or ion affinity chromatography. In some embodiments, the method further comprises cleaving the IMAC affinity tag by contacting the endonuclease with a protease corresponding to the protease cleavage site. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from a composition comprising the endonuclease.
[0021] In some aspects, the disclosure provides a method of disrupting a locus in a cell, comprising contacting the cell with a composition comprising a Class 2, V-type Cas endonuclease having at least 75% identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof, and an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the locus, wherein the Class 2, V-type Cas endonuclease has cleavage activity in the cell at least equivalent to spCas9. In some embodiments, the cleavage activity is measured in vitro by introducing the endonuclease, along with a compatible guide RNA, into a cell containing the target nucleic acid and detecting cleavage of the target nucleic acid sequence in the cell. In some embodiments, the composition comprises 20 picomoles (pmol) or less of the Class 2, V-type Cas endonuclease. In some embodiments, the composition comprises 1 pmol or less of the Class 2, V-type Cas endonuclease.
[0022] In some aspects, the disclosure provides a method of disrupting an albumin locus in a cell, comprising contacting the cell with a composition comprising an endonuclease having at least 75% identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof, and an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the locus, wherein the engineered guide RNA is configured to hybridize to any one of the target sequences in Table 6. In some embodiments, the engineered guide RNA comprises a sequence having at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to at least 18 non-degenerate nucleotides of any one of SEQ ID NOs: 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, and 474. In some embodiments, the engineered guide RNA comprises modified nucleotides of any of the single guide RNA (sgRNA) sequences in Table 6.In some embodiments, the endonuclease has at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 30-33, 39, 48, 56, 57, 61, 83, 92, 100, 110, 124, 136, 145, 148, 424, 425, 429, 476, or 629. In some embodiments, the endonuclease has at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to SEQ ID NO: 57. In some embodiments, the region is 5' to a PAM sequence comprising any one of SEQ ID NOs: 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, and 475.
[0023] In some aspects, the disclosure provides an isolated RNA molecule comprising a sequence of at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any of the sequences in Table 6. In some embodiments, the isolated RNA molecule further comprises a pattern of chemical modifications listed in any of the guide RNAs listed in Table 6.
[0024] In some aspects, the disclosure provides for the use of any of the RNA molecules disclosed herein to modify the albumin locus of a cell.
[0025] In some aspects, the disclosure provides an engineered nuclease system comprising an endonuclease configured to be selective for a protospacer adjacent motif (PAM) comprising any one of SEQ ID NOs: 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, and 475, and an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease is a class 2, type V Cas endonuclease. In some embodiments, the endonuclease is not a Cas12a nuclease. In some embodiments, the endonuclease is derived from an uncultured organism. In some embodiments, the endonuclease further comprises a PAM interaction domain configured to interact with the PAM. In some embodiments, the endonuclease has at least 75% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof. In some embodiments, the endonuclease has at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 30-33, 39, 48, 56, 57, 61, 83, 92, 100, 110, 124, 136, 145, 148, 424, 425, 429, 476, or 629.
[0026] In some aspects, the disclosure provides an engineered nuclease system comprising an endonuclease having at least 75% identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof, and a DNA methyltransferase. In some embodiments, the endonuclease has at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 30-33, 39, 48, 56, 57, 61, 83, 92, 100, 110, 124, 136, 145, 148, 424, 425, 429, 476, or 629. In some embodiments, the DNA methyltransferase is non-covalently bound to the endonuclease. In some embodiments, the DNA methyltransferase is fused to the endonuclease in a single polypeptide. In some embodiments, the DNA methyltransferase comprises Dmnt3A or Dnmt3L. In some embodiments, the KRAB domain is non-covalently linked to the endonuclease or the DNA methyltransferase.
[0027] In some embodiments, the KRAB domain is covalently linked to the endonuclease or the DNA methyltransferase. In some embodiments, the KRAB domain is fused to the endonuclease or the DNA methyltransferase in a single polypeptide. In some embodiments, the endonuclease is a nickase or is catalytically dead. In some embodiments, the engineered nuclease system further comprises an engineered guide RNA structure configured to form a complex with the endonuclease, the engineered guide RNA structure comprising a spacer sequence configured to hybridize to a target nucleic acid sequence. In some embodiments, the target nucleic acid sequence is comprised within or proximal to a promoter of the target genome. In some embodiments, the engineered guide RNA structure comprises one or more of: (a) 2'-O-methyl nucleotides, (b) 2'-fluoro nucleotides, or (c) phosphorothioate linkages. In some embodiments, the engineered guide RNA structure comprises a pattern of chemically modified nucleotides of any of the single guide RNAs in Table 6.
[0028] In some aspects, the disclosure provides methods of modifying a target nucleic acid locus, the method comprising delivering any of the engineered nuclease systems disclosed herein to the target nucleic acid locus, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, the complex being configured such that upon binding of the complex to the target nucleic acid locus, the DNA methyltransferase modifies the target nucleic acid locus.
[0029] In some aspects, the disclosure provides for the use of any of the engineered nuclease systems disclosed herein to modify a nucleic acid locus, in some embodiments, modifying the nucleic acid locus comprises methylating or demethylating a nucleotide at the nucleic acid locus.
[0030] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a RuvC domain, wherein the endonuclease is derived from an uncultured microorganism, and wherein the endonuclease is not a Cas12a endonuclease; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and wherein the engineered guide RNA comprises a spacer sequence configured to hybridize to a target nucleic acid sequence. In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof; and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, and the engineered guide RNA comprises a spacer sequence configured to hybridize to a target nucleic acid sequence. In some embodiments, the endonuclease comprises a RuvCI, II, or III domain. In some embodiments, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to the RuvCI, II, or III domain of any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629 or a variant thereof. In some embodiments, the RuvCI domain comprises a D catalytic residue. In some embodiments, the RuvCII domain comprises an E catalytic residue.In some embodiments, the RuvCIII domain comprises a D catalytic residue. In some embodiments, the RuvC domain does not have nuclease activity. In some embodiments, the endonuclease further comprises a WED II domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to a WED II domain of any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof. In some embodiments, the guide RNA comprises a sequence having at least 80% sequence identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 410-419. In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an engineered guide RNA comprising a sequence having at least 80% sequence identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 410-419; and (b) a class 2, V-type Cas endonuclease configured to bind to the engineered guide RNA. In some embodiments, the guide RNA comprises a sequence that is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the guide RNA is 30-250 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence that is at least 80% identical to a sequence from the group consisting of SEQ ID NOs: 630-645.
[0031] In some embodiments, the engineered nuclease system further comprises a single-stranded or double-stranded DNA repair template comprising, from 5' to 3', a first homologous arm comprising a sequence of at least 20 nucleotides 5' to a target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence. In some embodiments, the first or second homologous arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the first and second homologous arms are homologous to a prokaryotic, bacterial, fungal, or eukaryotic genomic sequence. In some embodiments, the single-stranded or double-stranded DNA repair template comprises a transgene donor. In some embodiments, the engineered nuclease system further comprises a DNA repair template comprising a double-stranded DNA segment adjacent to one or two single-stranded DNA segments. In some embodiments, the single stranded DNA segment is conjugated to the 5' end of the double stranded DNA segment. In some embodiments, the single stranded DNA segment is conjugated to the 3' end of the double stranded DNA segment. In some embodiments, the single stranded DNA segment has a length of 4-10 nucleotide bases. In some embodiments, the single stranded DNA segment has a nucleotide sequence that is complementary to a sequence within a spacer sequence. In some embodiments, the double stranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein coding sequence, an miRNA coding sequence, an RNA coding sequence, or a transgene. In some embodiments, the double stranded DNA sequence is adjacent to a nuclease cleavage site. In some embodiments, the nuclease cleavage site comprises a spacer and a PAM sequence. In some embodiments, the system comprises a Mg 2+In some embodiments, the guide RNA comprises a hairpin comprising at least 8, at least 10, or at least 12 base pair ribonucleotides. In some embodiments, the hairpin comprises 10 base pair ribonucleotides. In some embodiments, a) the endonuclease comprises a sequence that is at least 75%, 80%, or 90% identical to any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a variant thereof, and b) the guide RNA structure comprises a sequence that is at least 80% or 90% identical to a non-degenerate nucleotide of any one of SEQ ID NOs: 410-419. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT algorithms, or the CLUSTALW algorithm using Smith-Waterman homology search algorithm parameters. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.
[0032] In some aspects, the disclosure provides an engineered guide RNA comprising: a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; and b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide, and wherein the engineered guide ribonucleic acid polynucleotide is capable of forming a complex with an endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, and targeting the complex to a target sequence of the target DNA molecule. In some embodiments, the DNA-targeting segment is located 3' of both of the two complementary stretches of nucleotides. In some embodiments, the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to the non-degenerate nucleotides of SEQ ID NOs: 410-419. In some embodiments, the double-stranded RNA (dsRNA) duplex comprises at least 5, at least 8, at least 10, or at least 12 ribonucleotides.
[0033] In some aspects, the disclosure provides a deoxyribonucleic acid polynucleotide that encodes an engineered guide ribonucleic acid polynucleotide described herein.
[0034] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding a class 2, V-type Cas endonuclease, the endonuclease being derived from an uncultured microorganism, and the organism not being an uncultured organism. In some embodiments, the endonuclease comprises a variant having at least 70% or at least 80% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 630-645. In some embodiments, the NLS comprises SEQ ID NO: 631. In some embodiments, the NLS is proximal to the N-terminus of the endonuclease. In some embodiments, the NLS comprises SEQ ID NO: 630. In some embodiments, the NLS is proximal to the C-terminus of the endonuclease, hi some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.
[0035] In some aspects, the disclosure provides an engineered vector comprising a nucleic acid sequence encoding a Class 2, V-type Cas endonuclease, wherein the endonuclease is derived from an uncultured microorganism.
[0036] In some aspects, the disclosure provides engineered vectors comprising a nucleic acid described herein.
[0037] In some aspects, the present disclosure provides an engineered vector comprising a deoxyribonucleic acid polynucleotide described herein. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, a lentivirus, or an adenovirus.
[0038] In some aspects, the disclosure provides a cell comprising the vector described herein.
[0039] In some aspects, the disclosure provides a method of producing an endonuclease comprising culturing any of the host cells described herein.
[0040] In some aspects, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with a class 2, V-type Cas endonuclease complexed with an endonuclease and an engineered guide RNA configured to bind to the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the guide RNA structure comprises a sequence that is at least 80% or 90% identical to a non-degenerate nucleotide of any one of SEQ ID NOs: 410-419. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand that comprises a sequence that is complementary to a sequence of the engineered guide RNA, and a second strand that comprises a PAM. In some embodiments, the PAM is immediately adjacent to the 5' end of the sequence that is complementary to a sequence of the engineered guide RNA. In some embodiments, the class 2, V-type Cas endonuclease is derived from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0041] In some aspects, the disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering an engineered nuclease system described herein to a target nucleic acid locus, the endonuclease being configured to form a complex with an engineered guide ribonucleic acid structure, the complex being configured to modify the target nucleic acid locus upon binding of the complex to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a human cell, or a primary cell. In some embodiments, the cell is a primary cell. In some embodiments, the primary cell is a T cell. In some embodiments, the primary cell is a hematopoietic stem cell (HSC). In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding an endonuclease is operably linked. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding an endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide.In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide RNA operably linked to a ribonucleic acid (RNA) pol III promoter. In some embodiments, the endonuclease induces a single-stranded or double-stranded break at or proximal to the target locus. In some embodiments, the endonuclease induces a staggered single-stranded break within or 3' of the target locus.
[0042] In some aspects, the disclosure provides a host cell comprising an open reading frame encoding a heterologous endonuclease having at least 75% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof. In some embodiments, the endonuclease has at least 75% sequence identity to any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a variant thereof. In some embodiments, the host cell is an E. coli cell or a mammalian cell. In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is a λDE3 lysogen or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the open reading frame is selected from the group consisting of a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, BADIn some embodiments, the open reading frame is operably linked to a nucleotide sequence encoding an affinity tag linked in-frame to a sequence encoding an endonuclease. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the endonuclease via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the open reading frame is codon-optimized for expression in the host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of the host cell.
[0043] In some aspects, the disclosure provides a culture comprising any of the host cells described herein in a compatible liquid medium.
[0044] In some aspects, the disclosure provides a method of producing an endonuclease, comprising culturing any of the host cells described herein in a compatible growth medium. In some embodiments, the method further comprises inducing expression of the endonuclease by adding an additional chemical agent or an increased amount of a nutrient. In some embodiments, the additional chemical agent or the increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to IMAC, or ion affinity chromatography. In some embodiments, the open reading frame comprises a sequence encoding an IMAC affinity tag linked in frame to the sequence encoding the endonuclease. In some embodiments, the IMAC affinity tag is linked in frame to the sequence encoding the endonuclease via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the method further comprises cleaving the IMAC affinity tag by contacting a protease corresponding to the protease cleavage site with an endonuclease. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from the composition comprising the endonuclease.
[0045] In some aspects, the disclosure provides a method of disrupting a locus in a cell, comprising contacting the cell with a composition, the composition comprising: (a) a class 2, type V Cas endonuclease having at least 75% identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629, or a variant thereof; and (b) an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the locus, wherein the class 2, type V Cas endonuclease has cleavage activity at least equivalent to spCas9 in the cell. In some embodiments, the cleavage activity is measured in vitro by introducing the endonuclease, along with a compatible guide RNA, into a cell containing a target nucleic acid and detecting cleavage of the target nucleic acid sequence in the cell. In some embodiments, the composition comprises 20 pmoles or less of a class 2, type V Cas endonuclease. In some embodiments, the composition comprises 1 pmol or less of a Class 2, Type V Cas endonuclease.
[0046] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0047] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief description of the drawings]
[0048] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings.
[0049] [Figure 1] 1 shows a typical organization of the different classes and types of CRISPR / Cas loci previously described herein. [Figure 2A]
[0023] Figure 2 shows an overview of the MG119 family. Figure 2A shows a multiple alignment of MG119 effector representatives showing the domain composition and conservation of RuvC catalytic residues important for the function of double-stranded DNA cleavage activity. [Figure 2B] Figure 2B shows an overview of the MG119 family. Figure 2B shows a depiction of a CRISPR-containing contig with the genomic context surrounding the CRISPR array and Cas effector (example MG119-1). [Figure 2C]
[0023] Figure 2C shows the folding of the MG119-1 direct repeat. [Figure 2D]
[0023] Figure 2D shows an overview of the MG119 family. Figure 2D shows the single guide RNA designed for MG119-1. [Figure 3A]
[0036] Figure 3A shows an overview of the MG90 family. Figure 3A shows a multiple alignment of MG90 effector representatives showing the domain composition and conservation of RuvC catalytic residues important for the function of double-stranded DNA cleavage activity. [Figure 3B] An overview of the MG90 family is shown in Figure 3B. Figure 3B shows a depiction of a CRISPR-containing contig with the genomic context surrounding the CRISPR array and Cas effectors (example of MG90-5). [Figure 3C] MG90 family overview: Figure 3C shows the folding of the MG90-5 direct repeat. [Figure 4A]
[0036] Figure 4A shows an overview of the MG126 family. Figure 4A shows a multiple alignment of MG126 effector representatives showing the domain composition and conservation of RuvC catalytic residues important for the function of double-stranded DNA cleavage activity. [Figure 4B] Figure 4B shows an overview of the MG126 family. Figure 4B shows a depiction of a CRISPR-containing contig with the genomic context surrounding the CRISPR array and Cas effectors (example of MG126-4). [Figure 4C] 4A shows an overview of the MG126 family. Figure 4C shows the folding of the MG126-4 direct repeat. [Figure 5A]
[0036] Figure 5 shows an overview of the MG118 family. Figure 5A shows a multiple alignment of MG118 effector representatives showing the domain composition and conservation of RuvC catalytic residues important for the function of double-stranded DNA cleavage activity. [Figure 5B] Figure 5B shows an overview of the MG118 family. Figure 5B shows a depiction of a CRISPR-containing contig with the genomic context surrounding the CRISPR array and Cas effector (example MG118-1). [Figure 5C] 5A shows an overview of the MG118 family. Figure 5C shows the folding of the MG118-1 direct repeat. [Figure 6A]
[0036] Figure 6A shows an overview of the MG122 family. Figure 6A shows a multiple alignment of MG122 effector representatives showing the domain composition and conservation of RuvC catalytic residues important for the function of double-stranded DNA cleavage activity. [Figure 6B] Figure 6B shows an overview of the MG122 family. Figure 6B shows a depiction of a CRISPR-containing contig with the genomic context surrounding the CRISPR array and Cas effector (example of MG122-4). [Figure 6C] 6A shows an overview of the MG122 family. Figure 6C shows the folding of the MG122-4 direct repeat. [Figure 7A]
[0036] Figure 7 shows an overview of the MG120 family. Figure 7A shows a multiple alignment of MG120 effector representatives showing the domain composition and conservation of RuvC catalytic residues important for the function of double-stranded DNA cleavage activity. [Figure 7B] Figure 7B shows an overview of the MG120 family. Figure 7B shows a depiction of a CRISPR-containing contig with the genomic context surrounding the CRISPR array and Cas effectors (example MG120-1). [Figure 7C] 7 shows an overview of the MG120 family. Figure 7C shows the folding of the MG120-1 direct repeat. [Figure 8A] 8A shows an overview of the MG91 family. Figure 8A shows a depiction of a CRISPR-containing contig with the genomic context surrounding the CRISPR array and Cas effector (example of MG91B-24). [Figure 8B] 8A shows an overview of the MG91 family. FIG. 8B shows the folding of the MG91B-24 direct repeat. [Figure 8C] Figure 8C shows a representation of a CRISPR-containing contig with genomic context surrounding the CRISPR array and Cas effector (example of MG91C-10). [Figure 8D]
[0036] Figure 8D shows the folding of the MG91C-10 direct repeat. [Figure 9] Figure 1 shows the in vitro activity of MG119-2 using a TXTL assay. MG119-2 was tested for dsDNA cleavage with two intergenic sequences from the MG119-2 contig, a minimal array (MA) sequence containing repeats in either forward or reverse orientation, and a PAM library target plasmid. In lane 1, positive intergenic enrichment was observed as an amplified cleavage product with intergenic (IG) sequence 1 and the minimal array with repeats in forward orientation. Lanes 3 and 7 are negative controls where the IG was omitted, and lane 4 is a third negative control where both the array and IG were omitted. [Figure 10A]Figure 10A shows the SeqLogo of MG119-2 PAM (5'-nTnn-3') determined via next generation sequencing (NGS) of the cleavage products obtained from an in vitro cleavage assay. [Figure 10B] FIG. 10B shows a histogram of cleavage sites (23 bp away from the PAM). [Figure 11A] Examples of active MG119 nucleases and their sgRNA designs are shown. Figure 11A shows the predicted fold of a single guide RNA sequence without a spacer. The blue circle represents the first 5' nucleotide of the tracrRNA and the red circle represents the 3' nucleotide of the repeat. The tracrRNA and the repeat are looped with a GAAA tetraloop. The repeat-anti-repeat fold is at the 3' end of each structure. Three different RNA structures of active guides within the same family are shown. From left to right: the MG119-28 guide has a very long hairpin with four hairpins, three smaller hairpins at the 5' end, and two bulges next to the repeat-anti-repeat fold. The MG119-83 sgRNA has three small hairpins and the repeat-anti-repeat has two bulges. MG119-118 has four hairpins, the second of which branches into three hairpins from the 5' end, while the third hairpin and the repeat-anti-repeat have a bulge. This guide also has several paired nucleotides between the 5' end of tracr and the 3' end of the repeat. [Figure 11B]Figure 11B shows an example of active MG119 nucleases and their sgRNA designs. Figure 11B shows in vitro cleavage assay amplification products on a 2% agarose gel. Low molecular weight DNA ladder (NEB) is in lane 1, lane 7, and lane 11. Contents of other lanes from left to right: (2) MG119-28 nuclease only, MG119-28 nuclease plus (3) sgRNA1 with U67 spacer, (4) sgRNA1 with U40 spacer, (5) sgRNA2 with U67 spacer, and (6) sgRNA2 with U40 spacer, (8) MG119-83 nuclease only, MG119-83 nuclease plus (9) sgRNA1 with U67 spacer, and (10) sgRNA1 with U40 spacer, (12) MG119-118 nuclease only, MG119-118 nuclease plus (13) sgRNA1 with U67 spacer, and (14) sgRNA1 with U40 spacer. The resulting amplicon products are 188 bp with a U67 spacer carrying guide, or 205 bp with a U40 spacer carrying guide. [Figure 12] The sequence logo of the protospacer adjacent motif (PAM) for active MG119 nuclease is shown. [Figure 13A] Exemplary SDS-PAGE gels and size-exclusion chromatography (SEC) A280 traces of protein purification steps are shown in Figure 13A, which shows MG119-28Δ purification using (1) lysis after sonication, (2) centrifugation after clarification, (3) Ni-NTA gravity column flow-through, (4) elution from Ni-NTA resin, and (5) sample recovered from concentrated sample. [Figure 13B] Exemplary SDS-PAGE gels and size exclusion chromatography (SEC) A280 traces of the protein purification steps are shown. Figure 13B shows the S200i 10 / 300 GL column SEC A280 trace. Peak fractions were pooled and concentrated. [Figure 13C]Exemplary SDS-PAGE gels and size-exclusion chromatography (SEC) A280 traces of protein purification steps are shown. Figure 13C shows MBP-tagged / cleaved MG119-28Δ purification using samples collected from (1) lysis after sonication, (2) centrifugation after clarification, (3) Ni-NTA gravity column flow-through, (4) eluate from Ni-NTA resin, (5) concentrated protein, (6) concentrated protein cleaved overnight with TEV protease, (7) and centrifuged (21,000 x g, 4°C, 10 min) to pellet aggregates, (8) amylose column flow-through, (9) flow-through centrifuged (21,000 x g, 4°C, 10 min) to pellet aggregates, and (10) concentrated flow-through. [Figure 13D] Exemplary SDS-PAGE gels and size-exclusion chromatography (SEC) A280 traces of protein purification steps are shown. Figure 13D shows MBP-tagged / cleaved MG119-28Δ purification using samples collected from (1) lysis after sonication, (2) centrifugation after clarification, (3) Ni-NTA gravity column flow-through, (4) eluate from Ni-NTA resin, (5) concentrated protein, (6) concentrated protein cleaved overnight with TEV protease, (7) and centrifuged (21,000 x g, 4°C, 10 min) to pellet aggregates, (8) amylose column flow-through, (9) flow-through centrifuged (21,000 x g, 4°C, 10 min) to pellet aggregates, and (10) concentrated flow-through. [Figure 13E] Exemplary SDS-PAGE gels and size exclusion chromatography (SEC) A280 traces of the protein purification steps are shown. Figure 13E shows the S200i 10 / 300 GL column SEC A280 trace. [Figure 13F] An exemplary SDS-PAGE gel and size exclusion chromatography (SEC) A280 trace of the protein purification steps are shown. Figure 13F shows data demonstrating that of the five MG119 candidates expressed in both pMGB and pMGBΔ expression vectors, all showed higher yields in the pMGBΔ vector. [Figure 14A] An example of in vitro cleavage efficiency using purified proteins is shown in Figure 14A, which shows an agarose gel showing titration of RNP:substrate ratios and increased substrate cleavage at higher ratios. [Figure 14B] An example of in vitro cleavage efficiency using purified proteins is shown. Figure 14B shows the percent of substrate cleaved determined for each lane using densitometry. Cleavage fractions were plotted in Prism8 and the slope of the linear range of cleavage was used to calculate protein activity fraction. The assay used MG119-28 expressed in a pMGBΔ backbone. [Figure 15A] Examples of in vitro cleavage and editing efficiency of mouse Hepa1-6 cell DNA are shown. Figure 15A shows the percent cleavage of MG119-28 with four chemically modified guides targeting the mouse albumin gene in intron 1 (Table 6). Two concentrations of nuclease were tested: 15.6 nM (black bars) and 7.8 nM (white bars). Cleavage was normalized to a non-targeting control. MG119-28 can cleave Hepa 1-6 gDNA up to an average of 60% with sgRNA4 at 15.6 nM RNP and up to 33% at 7.8 nM RNP. [Figure 15B] Examples of in vitro cleavage and editing efficiency of mouse Hepa 1-6 cell DNA are shown. Figure 15B shows the percent of indels generated by MG119-28 in Hepa 1-6 cells normalized to the apo reaction. Each condition was performed in triplicate. An average of 25.12% of sequenced reads were edited with sgRNA3. sgRNA3 is consistently active in vitro and in cells as shown here. The next best guide in cells is sgRNA4 with an average of 4.11% editing. Edits observed are primarily 4-24 bp deletions.
[0050] Brief Description of the Sequence Listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein. MG122
[0051] SEQ ID NOs: 1 to 5 show the full-length peptide sequences of MG122 nuclease. MG120
[0052] SEQ ID NOs: 6 to 14 show the full-length peptide sequences of MG120 nuclease.
[0053] SEQ ID NOs: 333-335 and 355-357 show the nucleotide sequences of MG120 tracrRNA derived from the same locus as the MG120 Cas effector.
[0054] SEQ ID NOs:374-375 and 389-390 show the nucleotide sequences of the MG120 minimal array. MG118
[0055] SEQ ID NO: 15 shows the full-length peptide sequence of MG118 nuclease.
[0056] SEQ ID NO: 376 shows the nucleotide sequence of the MG118 minimal array.
[0057] SEQ ID NO: 391 shows the nucleotide sequence of the MG118 minimal array.
[0058] SEQ ID NOs: 400-401 show the nucleotide sequences of MG118-targeted CRISPR repeats.
[0059] SEQ ID NOs: 410-411 show the nucleotide sequence of MG118 crRNA. MG90
[0060] SEQ ID NOs: 16 to 29 show the full-length peptide sequences of MG90 nuclease.
[0061] SEQ ID NOs: 346-347 and 368-369 show the nucleotide sequences of MG90 tracrRNA derived from the same locus as the MG90 Cas effector.
[0062] SEQ ID NOs: 383-384 and 398-399 show the nucleotide sequences of the MG90 minimal array.
[0063] SEQ ID NOs: 402-403 show the nucleotide sequences of MG90-targeted CRISPR repeats.
[0064] SEQ ID NOs: 412-413 show the nucleotide sequences of MG90 sgRNA. MG119
[0065] SEQ ID NOs: 30 to 150, 420 to 431, 476 to 624, and 629 show full-length peptide sequences of MG119 nuclease.
[0066] SEQ ID NOs: 326-332, 336-345, 348-354, and 358-367 show the nucleotide sequences of MG119 tracrRNA derived from the same locus as the MG119 Cas effector.
[0067] SEQ ID NOs: 370-373, 377-382, 385-388, and 392-397 show the nucleotide sequences of the MG119 minimal array.
[0068] SEQ ID NOs: 404-409 show the nucleotide sequences of MG119-targeted CRISPR repeats.
[0069] SEQ ID NOs: 414-419, 432, 434, 436, 438, 440, 442, 444, 446, 448, 450, 452, 454, 456, 458, 460, 462, 464, 466, 468, 470, 472, and 474 show the nucleotide sequences of MG119 sgRNA.
[0070] SEQ ID NOs: 433, 435, 437, 439, 441, 443, 445, 447, 449, 451, 453, 455, 457, 459, 461, 463, 465, 467, 469, 471, 473, and 475 show the nucleotide sequence of MG119 PAM. MG91B
[0071] SEQ ID NOs: 151 to 291 show the full-length peptide sequences of MG91B nuclease. MG91C
[0072] SEQ ID NOs: 292 to 318 show the full-length peptide sequences of MG91C nuclease. MG91A
[0073] SEQ ID NO: 319 shows the full-length peptide sequence of MG91A nuclease. MG126
[0074] SEQ ID NOs: 320 to 325 show the full-length peptide sequences of MG126 nuclease. Detailed Description of the Invention
[0075] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used.
[0076] The practice of some of the methods disclosed herein may involve immunological, biochemical, chemical, molecular biology, microbiology, cell biology, genomics, and recombinant DNA techniques, unless otherwise indicated. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FMA Usubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)) (incorporated herein in its entirety by reference).
[0077] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, to the extent the terms "comprising," "including," "having," "having," "having," or variations thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0078] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0079] As used herein, "cell" generally refers to a biological cell. A cell may be the basic structural, functional, and / or biological unit of a living organism. A cell may originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-cell eukaryotic cells, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, bryophytes, mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.), and cells from other organisms (e.g., cereals, vegetables, fruits, and vegetables). C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from vertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), etc. In some cases, the cells are not derived from a naturally occurring organism (e.g., the cells may be synthetically produced and sometimes referred to as artificial cells).
[0080] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide may refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as by using a moiety that includes an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4′dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif.; fluoro-conjugated deoxynucleotides, fluoro-conjugated Cy3-dCTP, fluoro-conjugated Cy5-dCTP, fluoro-conjugated fluoroX-dCTP, fluoro-conjugated Cy3-dUTP, and fluoro-conjugated Cy5-dUTP available from Amersham, Arlington Heights, Ill.; Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2′-dATP available from Mannheim, Indianapolis, Ind.; and Molecular Examples of chromosomal labeling nucleotides available from Probes, Eugene, Oreg. include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides may also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0081] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-stranded, double-stranded, or multiple-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and may perform any function. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, heterologous nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci defined from binding analyses, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers.The sequence of nucleotides may be interrupted by non-nucleotide components.
[0082] The term "transfection" or "transfected" generally refers to the introduction of a nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a genetic sequence encoding a complete protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88, which is incorporated herein by reference in its entirety.
[0083] The terms "peptide", "polypeptide" and "protein" are used interchangeably herein and generally refer to a polymer of at least two amino acid residues linked by peptide bonds. The term does not refer to a particular length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers that include at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins (e.g., domains) with or without secondary and / or tertiary structure. The term also encompasses amino acid polymers that have been modified by any other manipulation, such as, for example, disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids may include natural amino acids and unnatural amino acids, which are chemically modified to include a non-naturally occurring group or chemical moiety on the amino acid. An amino acid analog may refer to an amino acid derivative. The term "amino acid" includes both D- and L-amino acids.
[0084] As used herein, "non-natural" may generally refer to a nucleic acid or polypeptide sequence that is not present in a naturally occurring nucleic acid or protein. Non-natural may refer to an affinity tag. Non-natural may refer to a fusion. Non-natural may refer to a naturally occurring nucleic acid or polypeptide sequence that includes mutations, insertions, and / or deletions. A non-natural sequence may exhibit and / or encode an activity (e.g., an enzyme activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, an ubiquitination activity, etc.) that may also be exhibited by the nucleic acid and / or polypeptide sequence to which the non-natural sequence is fused. A non-natural nucleic acid or polypeptide sequence may be linked to a naturally occurring nucleic acid and / or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid and / or polypeptide sequence that encodes a chimeric nucleic acid or polypeptide.
[0085] As used herein, the term "promoter" generally refers to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping the nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often called transcription factors, which promote the binding of RNA polymerase to DNA, thereby resulting in gene transcription. A "basal promoter", also called a "core promoter", may generally refer to a promoter that contains all the basic essential elements to promote the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters typically, but not necessarily, contain a TATA-box and / or a CAAT box.
[0086] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0087] As used herein, "operably linked," "operably linked," "operably linked," or grammatical equivalents thereof generally refer to the juxtaposition of genetic elements, such as promoters, enhancers, polyadenylation sequences, and the like, where the elements are in a relationship that allows them to operate in an expected manner. For example, a regulatory element, which may include a promoter sequence and / or an enhancer sequence, is operably linked to a coding region if the regulatory element helps to initiate transcription of the coding sequence. There may be intervening residues between the regulatory element and the coding region so long as this functional relationship is maintained.
[0088] As used herein, a "vector" generally refers to a polymer or an association of polymers that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally includes a genetic element, such as a regulatory element, operably linked to a gene to facilitate expression of the gene in a target.
[0089] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of a gene or genes with regulatory elements that are operably linked for expression.
[0090] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be the ability to affect expression in a manner known to be attributable to the full-length sequence.
[0091] As used herein, an "engineered" subject generally refers to a subject that has been modified by human intervention. By way of non-limiting examples, a nucleic acid may be modified by altering its sequence to a sequence that does not occur in nature, a nucleic acid may be modified by ligating to a nucleic acid with which it is not naturally associated such that the ligated product has a function not present in the original nucleic acid, an engineered nucleic acid may be synthesized in vitro with a sequence that does not occur in nature, a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature, and an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.
[0092] As used herein, "synthetic" and "artificial" may generally be used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) with naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains.
[0093] As used herein, the term "Cas12a" generally refers to a family of Cas endonucleases that are class 2, type VA Cas endonucleases that (a) use relatively small guide RNAs (about 42-44 nucleotides) that are processed by the nuclease itself after transcription from the CRISPR array, and (b) cleave DNA to leave staggered cleavage sites. Further characteristics of this enzyme family can be found, for example, in Zetsche B, Heidenreich M, Mohanraju P, et al. Nat Biotechnol 2017;35:31-34, and Zetsche B, Gootenberg JS, Abudayyeh OO, et al. Cell 2015;163:759-771, which are incorporated herein by reference.
[0094] As used herein, a "guide nucleic acid" can generally refer to a nucleic acid that can hybridize to another nucleic acid. A guide nucleic acid can be RNA. A guide nucleic acid can be DNA. A guide nucleic acid can be programmed to bind to a sequence of a nucleic acid in a site-specific manner. The nucleic acid to be targeted, or the target nucleic acid, can include nucleotides. A guide nucleic acid can include nucleotides. A portion of a target nucleic acid can be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to a guide nucleic acid and hybridizes with the guide nucleic acid can be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore not complementary to the guide nucleic acid can be referred to as a non-complementary strand. A guide nucleic acid can include a polynucleotide strand and can be referred to as a "single guide nucleic acid." A guide nucleic acid can include two polynucleotide strands and can be referred to as a "double guide nucleic acid." Unless otherwise specified, the term "guide nucleic acid" is inclusive and can refer to both single guide nucleic acid and double guide nucleic acid. A guide nucleic acid may include a segment that may be referred to as a "nucleic acid targeting segment" or a "nucleic acid targeting sequence" or a "spacer sequence." The nucleic acid targeting segment may include a sub-segment that may be referred to as a "protein binding segment" or a "protein binding sequence" or a "Cas protein binding segment."
[0095] The terms "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refer to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a certain percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of a word length (W) of 3, an expectation (E) of 10, and a gap cost of 11, an extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of a word length (W) of 2, an expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE using default parameters; MAFFT using parameters of 2 retrees and 1000 maximum repeats; Novafold using default parameters; HMMER hmmalign using default parameters.
[0096] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences aligned for maximum amino acid residue or nucleotide correspondence, e.g., as determined by the alignment producing the highest or "optimized" percent identity score.
[0097] The present disclosure includes any variant of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without destroying the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by replacing amino acids with similar hydrophobicity, polarity, and R chain length. Additionally or alternatively, by comparing the aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues (e.g., non-conserved residues) that are mutated between species without changing the basic function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of the endonuclease protein sequences described herein (e.g., the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, or MG126 family endonucleases, or any other family nuclease described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions such that the activity of one or more important active site residues or guide RNA binding residues of the endonuclease is not destroyed. In some embodiments, functional variants of any of the proteins described herein lack at least one substitution of the conserved or functional residues called out in Figure 2A, Figure 3A, Figure 4A, Figure 5A, or Figure 6A.In some embodiments, a functional variant of any of the proteins described herein lacks all substitutions of the conserved or functional residues called out in Figure 2A, Figure 3A, Figure 4A, Figure 5A, or Figure 6A.
[0098] The disclosure also includes variants (e.g., reduced activity variants) of any of the enzymes described herein having substitutions of one or more catalytic residues to reduce or eliminate activity of the enzyme. In some embodiments, reduced activity variants of the proteins described herein include disruptive substitutions of at least one, at least two, or all three catalytic residues called out in Figure 2A, Figure 3A, Figure 4A, Figure 5A, or Figure 6A.
[0099] Conservative substitution tables providing functionally similar amino acids are available in a variety of references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993)). Each of the following eight groups contains amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), Glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V), 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W), 7) Serine (S), Threonine (T), and 8) Cysteine (C), Methionine (M).
[0100] overview Discovery of new Cas enzymes with unique functionality and structure could further disrupt deoxyribonucleic acid (DNA) editing technologies, endowing them with the potential to improve speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microbes and the sheer diversity of microbial species, there are relatively few functionally characterized CRISPR / Cas enzymes in the literature. This is in part because the vast number of microbial species cannot be easily cultured in laboratory conditions. Metagenomic sequencing from natural environmental niches containing large numbers of microbial species could dramatically increase the number of known new CRISPR / Cas systems, endowing them with the potential to expedite the discovery of new oligonucleotide editing functions. A fruitful recent example of such an approach is demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.
[0101] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode RNA-based targeting elements; and (ii) an ORF encoding a Cas that encodes a nuclease polypeptide directed by the RNA-based targeting element flanked by accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (PAM is usually a sequence that is not commonly represented in the host genome). Depending on the exact function and organization of the system, CRISPR-Cas systems are commonly organized into two classes, five types, and 16 subtypes based on shared functional features and evolutionary similarities (see Figure 1).
[0102] Class I CRISPR-Cas systems have large, multi-subunit effector complexes and include Types I, III, and IV. Class II CRISPR-Cas systems generally have single polypeptide multi-domain nuclease effectors and include Types II, V, and VI.
[0103] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, processing of CRISPR arrays into mature crRNA does not require the presence of special endonuclease subunits, but rather a small transcoding crRNA (tracrRNA) with a region complementary to the array repeat sequence, which interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure that is cleaved by endogenous RNAse III to generate the mature effector enzyme loaded with both tracrRNA and crRNA. Cas II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure consisting of a RuvC-like endonuclease domain that fits into an RNase H fold with an unrelated HNH nuclease domain inserted into the fold of the RuvC-like nuclease domain. The RuvC-like domain is involved in cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is involved in cleavage of the replacement DNA strand.
[0104] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g., Cas12) structure similar to type II effectors, including a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use tracrRNA to process pre-crRNA into mature crRNA, but unlike type II systems that require RNAse III to cleave pre-crRNA into multiple crRNAs, type V systems can cleave pre-crRNA using the effector nuclease itself. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are also known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to have robust single-stranded non-specific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.
[0105] CRISPR-Cas systems have emerged as a gene editing technology in recent years due to their targeting and ease of use. The most commonly used systems are class 2 type II SpCas9 and class 2 type VA Cas12a (previously Cpf1). In particular, VA systems have become more widely used due to their reported higher specificity in cells than other nucleases and fewer or no off-target effects. VA systems are also advantageous in that the guide RNAs are small (42-44 nucleotides compared to approximately 100 nt for SpCas9) and processed by the nuclease itself after transcription from the CRISPR array, simplifying multiplexed applications involving multiple gene editing. Furthermore, VA systems have staggered cut sites, which may facilitate directed repair pathways such as microhomology-dependent targeted integration (MITI).
[0106] The most commonly used type VA enzymes require a 5' protospacer adjacent motif (PAM) next to the selected target site: 5'-TTTV-3' for Lachnospiraceae bacteria ND2006 LbCas12a and Acidaminococcus species AsCas12a; and 5'-TTV-3' for Francisella novicida FnCas12a. Recent surveys of orthologs have revealed proteins with less restrictive PAM sequences that are also active in mammalian cell culture, e.g., YTV, YYN, or TTN. However, these enzymes do not fully encompass the biodiversity and targetability of type V and may not represent all possible activities and PAM sequence requirements. Here, we retrieved thousands of genomic fragments from multiple metagenomes for type V nucleases. The known diversity of V enzymes may be expanded, and novel systems may be developed into highly targetable, compact, and precise gene editing agents.
[0107] MG enzyme Type V CRISPR systems are being rapidly adopted for use in a variety of genome editing applications. These programmable nucleases are part of the adaptive microbial immune system, and their natural diversity remains largely unexplored. Novel families of type V CRISPR enzymes were identified through large-scale analysis of metagenomics collected from various complex environments, and representatives of these were developed systems into gene editing platforms. The majority of these systems are derived from uncultured organisms, and some of them encode divergent type V effectors within the same CRISPR operon.
[0108] In some embodiments, the present disclosure provides novel V-type candidates. These candidates may represent one or more novel subtypes, and several subfamilies may be identified. These nucleases are less than about 900 amino acids long. These novel subtypes may be found in the same CRISPR locus as known V-type effectors. RuvC catalytic residues may be identified for novel V-type candidates, and these novel V-type candidates may not require tracrRNA.
[0109] In some embodiments, the present disclosure provides smaller V-type effectors. Such effectors can be small putative effectors. These effectors may simplify delivery and expand therapeutic applications.
[0110] In some embodiments, the present disclosure provides novel V-type effectors. Such effectors can be MG90 as described herein (see Figures 3A-3C). Such effectors can be MG91 as described herein (see Figures 8A-8B). Such effectors can be MG118 as described herein (see Figures 5A-5C). Such effectors can be MG119 as described herein (see Figures 2A-2D). Such effectors can be MG120 as described herein (see Figures 7A-7C). Such effectors can be MG122 as described herein (see Figures 6A-6C). Such effectors can be MG126 as described herein (see Figures 4A-4C).
[0111] In one aspect, the present disclosure provides engineered nuclease systems discovered through metagenomic sequencing. In some cases, metagenomic sequencing is performed on a sample. In some cases, the sample can be collected from a variety of environments. Such environments can be human microbiomes, animal microbiomes, hot environments, cold environments. Such environments can include sediments.
[0112] In one aspect, the disclosure provides an engineered nuclease system comprising an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2, type V Cas endonuclease. In some cases, the endonuclease is a novel subtype class 2, type V Cas endonuclease. In some cases, the endonuclease is derived from an uncultured microorganism. The endonuclease may comprise a RuvC domain. In some cases, the engineered nuclease system comprises an engineered guide RNA. In some cases, the engineered guide RNA is configured to form a complex with the endonuclease. In some cases, the engineered guide RNA comprises a spacer sequence. In some cases, the spacer sequence is configured to hybridize to a target nucleic acid sequence.
[0113] In one aspect, the disclosure provides an engineered nuclease system comprising an endonuclease. In some cases, the endonuclease has at least about 70% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629. In some cases, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629.
[0114] In some cases, the endonuclease comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629.
[0115] In some cases, the engineered nuclease system includes an engineered guide RNA. In some cases, the engineered guide RNA is configured to form a complex with an endonuclease. In some cases, the engineered guide RNA includes a spacer sequence. In some cases, the spacer sequence is configured to hybridize to a target nucleic acid sequence. In some cases, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence.
[0116] In some cases, the endonuclease is not a Cpf1 or Cms1 endonuclease.
[0117] In some cases, the guide RNA comprises a sequence having at least 80% sequence identity to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 410-419. In some cases, the guide RNA comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 410-419. In some cases, the guide RNA comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 410-419. In some cases, the guide RNA comprises a sequence that is substantially identical to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NO: 410-419.
[0118] In some cases, the guide RNA comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the first 19 nucleotides or non-degenerate nucleotides of SEQ ID NOs: 410-419. In some cases, the endonuclease is configured to bind to the engineered guide RNA. In some cases, the Cas endonuclease is configured to bind to the engineered guide RNA. In some cases, the class 2 Cas endonuclease is configured to bind to the engineered guide RNA. In some cases, a class 2, V-type Cas endonuclease is configured to bind to an engineered guide RNA. In some cases, a class 2, V-type, novel subtype Cas endonuclease is configured to bind to an engineered guide RNA.
[0119] In some cases, the guide RNA comprises a sequence that is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a eukaryotic genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a fungal genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a plant genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a mammalian genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence that is complementary to a human genomic polynucleotide sequence.
[0120] In some cases, the guide RNA is 30-250 nucleotides in length. In some cases, the guide RNA is 42-44 nucleotides in length. In some cases, the guide RNA is 42 nucleotides in length. In some cases, the guide RNA is 43 nucleotides in length. In some cases, the guide RNA is 44 nucleotides in length. In some cases, the guide RNA is 85-245 nucleotides in length. In some cases, the guide RNA is more than 90 nucleotides in length. In some cases, the guide RNA is less than 245 nucleotides in length.
[0121] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLS). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus of any one of SEQ ID NOs: 630-645, or to a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 630-645. In some cases, the NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 630-645.
[0122] [Table 1]
[0123] In some cases, the engineered nuclease system further comprises a single-stranded or double-stranded DNA repair template. In some cases, the engineered nuclease system further comprises a single-stranded DNA repair template. In some cases, the engineered nuclease system further comprises a double-stranded DNA repair template. In some cases, the single-stranded or double-stranded DNA repair template can comprise, from 5' to 3', a first homologous arm comprising a sequence of at least 20 nucleotides 5' to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homologous arm comprising a sequence of at least 20 nucleotides 3' to the target sequence.
[0124] In some cases, the first homology arm comprises a sequence of at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides. In some cases, the second homology arm comprises a sequence of at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 750, or at least 1000 nucleotides.
[0125] In some cases, the first and second homologous arms are homologous to a prokaryotic genomic sequence. In some cases, the first and second homologous arms are homologous to a bacterial genomic sequence. In some cases, the first and second homologous arms are homologous to a fungal genomic sequence. In some cases, the first and second homologous arms are homologous to a eukaryotic genomic sequence.
[0126] In some cases, the engineered nuclease system further comprises a DNA repair template. The DNA repair template may comprise a double-stranded DNA segment. The double-stranded DNA segment may be adjacent to one single-stranded DNA segment. The double-stranded DNA segment may be adjacent to two single-stranded DNA segments. In some cases, the single-stranded DNA segment is conjugated to the 5' end of the double-stranded DNA segment. In some cases, the single-stranded DNA segment is conjugated to the 3' end of the double-stranded DNA segment.
[0127] In some cases, the single stranded DNA segment has a length of 1-15 nucleotide bases. In some cases, the single stranded DNA segment has a length of 4-10 nucleotide bases. In some cases, the single stranded DNA segment has a length of 4 nucleotide bases. In some cases, the single stranded DNA segment has a length of 5 nucleotide bases. In some cases, the single stranded DNA segment has a length of 6 nucleotide bases. In some cases, the single stranded DNA segment has a length of 7 nucleotide bases. In some cases, the single stranded DNA segment has a length of 8 nucleotide bases. In some cases, the single stranded DNA segment has a length of 9 nucleotide bases. In some cases, the single stranded DNA segment has a length of 10 nucleotide bases.
[0128] In some cases, the single-stranded DNA segment has a nucleotide sequence that is complementary to a sequence within the spacer sequence. In some cases, the double-stranded DNA sequence comprises a barcode, an open reading frame, an enhancer, a promoter, a protein coding sequence, an miRNA coding sequence, an RNA coding sequence, or a transgene.
[0129] In some cases, the engineered nuclease system is 2+ The present invention further comprises a source of
[0130] In some cases, the guide RNA comprises a hairpin that comprises at least 8 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin that comprises at least 9 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin that comprises at least 10 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin that comprises at least 11 base pair ribonucleotides. In some cases, the guide RNA comprises a hairpin that comprises at least 12 base pair ribonucleotides.
[0131] In some cases, the endonuclease comprises a variant of any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a sequence that is at least 70% identical to a variant thereof. In some cases, the endonuclease comprises a variant of any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a sequence that is at least 75% identical to a variant thereof. In some cases, the endonuclease comprises a variant of any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a sequence that is at least 80% identical to a variant thereof. In some cases, the endonuclease comprises a variant of any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a sequence that is at least 85% identical to a variant of any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a variant thereof. In some cases, the endonuclease comprises a variant of any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a sequence that is at least 90% identical to a variant thereof. In some cases, the endonuclease comprises a variant of any one of SEQ ID NOs: 1, 6, 15, 30, 151, 292, or 319, or a sequence that is at least 95% identical to a variant thereof.
[0132] In some cases, sequences may be determined by the BLASTP, CLUSTALW, MUSCLE, or MAFFT algorithms, or the CLUSTALW algorithm using Smith-Waterman homology search algorithm parameters. Sequence identity may be determined by the BLASTP homology search algorithm using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11, extension of 1, and using a conditional composition score matrix adjustment.
[0133] In one aspect, the present disclosure provides an engineered guide RNA comprising a DNA targeting segment. In some cases, the DNA targeting segment comprises a nucleotide sequence that is complementary to a target sequence. In some cases, the target sequence is in a target DNA molecule. In some cases, the engineered guide RNA comprises a protein-binding segment. In some cases, the protein-binding segment comprises two complementary stretches of nucleotides. In some cases, the two complementary stretches of nucleotides hybridize to form a double-stranded RNA (dsRNA) duplex. In some cases, the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide. In some cases, the engineered guide ribonucleic acid polynucleotide can form a complex with an endonuclease. In some cases, the endonuclease has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629. In some cases, the complex targets a target sequence of a target DNA molecule. In some embodiments, the DNA targeting segment is located 3' of both of the two complementary stretches of nucleotides.
[0134] In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 8 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 9 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 10 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 11 ribonucleotides. In some cases, the double-stranded RNA (dsRNA) duplex comprises at least 12 ribonucleotides.
[0135] In some cases, the deoxyribonucleic acid polynucleotide encodes an engineered guide ribonucleic acid polynucleotide.
[0136] In one aspect, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence. In some cases, the engineered nucleic acid sequence is optimized for expression in an organism. In some cases, the nucleic acid encodes an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 endonuclease. In some cases, the endonuclease is a class 2, type V Cas endonuclease. In some cases, the endonuclease is a class 2, type V, novel subtype Cas endonuclease. In some cases, the endonuclease is from an uncultured microorganism. In some cases, the organism is not an uncultured organism.
[0137] In some cases, the endonuclease comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-325, 420-431, 476-624, or 629.
[0138] In some cases, the endonuclease may include a variant having one or more nuclear localization sequences (NLS). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus of any one of SEQ ID NOs: 630-645, or to a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 630-645.
[0139] In some cases, the organism is a prokaryote. In some cases, the organism is a bacterium. In some cases, the organism is a eukaryote. In some cases, the organism is a fungus. In some cases, the organism is a plant. In some cases, the organism is a mammal. In some cases, the organism is a rodent. In some cases, the organism is a human.
[0140] In one aspect, the disclosure provides an engineered vector. In some cases, the engineered vector comprises a nucleic acid sequence encoding an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 Cas endonuclease. In some cases, the endonuclease is a class 2, type V Cas endonuclease. In some cases, the endonuclease is a class 2, type V, novel subtype Cas endonuclease. In some cases, the endonuclease is derived from an uncultured microorganism.
[0141] In some cases, the engineered vector comprises a nucleic acid described herein. In some cases, the nucleic acid described herein is a deoxyribonucleic acid polynucleotide described herein. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus.
[0142] In one aspect, the disclosure provides a cell comprising the vector described herein.
[0143] In one aspect, the disclosure provides a method of producing an endonuclease. In some cases, the method includes culturing a cell.
[0144] In one aspect, the disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide. The method may include contacting the double-stranded deoxyribonucleic acid polynucleotide with an endonuclease. In some cases, the endonuclease is a Cas endonuclease. In some cases, the endonuclease is a class 2 Cas endonuclease. In some cases, the endonuclease is a class 2, type V Cas endonuclease. In some cases, the endonuclease is a class 2, type V, novel subtype Cas endonuclease. In some cases, the endonuclease is complexed with an engineered guide RNA. In some cases, the engineered guide RNA is configured to bind to the endonuclease. In some cases, the engineered guide RNA is configured to bind to the double-stranded deoxyribonucleic acid polynucleotide. In some cases, the engineered guide RNA is configured to bind to an endonuclease and to a double-stranded deoxyribonucleic acid polynucleotide. In some cases, the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM).
[0145] In some cases, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of an engineered guide RNA and a second strand comprising a PAM. In some cases, the PAM is immediately adjacent to the 5' end of the sequence complementary to a sequence of an engineered guide RNA. In some cases, the endonuclease is not Cpf1 endonuclease or Cms1 endonuclease. In some cases, the endonuclease is from an uncultured microorganism. In some cases, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0146] In one aspect, the present disclosure provides a method for modifying a target nucleic acid locus. The method may include delivering an engineered nuclease system as described herein to a target nucleic acid locus. In some cases, the endonuclease is configured to form a complex with an engineered guide ribonucleic acid structure. In some cases, the complex is configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies the target nucleic acid locus.
[0147] In some cases, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some cases, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some cases, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some cases, the target nucleic acid locus is in vitro. In some cases, the target nucleic acid locus is in a cell. In some cases, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a primary cell, or a human cell.
[0148] In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some cases, the nucleic acid comprises a promoter. In some cases, the open reading frame encoding the endonuclease is operably linked to a promoter.
[0149] In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding an endonuclease. In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide. In some cases, delivery of the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide RNA operably linked to a ribonucleic acid (RNA) pol III promoter.
[0150] In some cases, the endonuclease induces a single-stranded or double-stranded break at or proximal to the target locus, in some cases, the endonuclease induces a staggered single-stranded break within the target locus or 3' to the target locus.
[0151] In some cases, the effector repeat motif is used to inform the guide design of the MG nuclease. For example, the processed gRNA in a type V system consists of the last 20-22 nucleotides of the CRISPR repeat. This sequence can be synthesized into a crRNA (with a spacer) and tested in vitro with a synthetic nuclease for cleavage on a library of potential targets. Using this method, the PAM can be determined. In some cases, type V enzymes may use a "universal" gRNA. In some cases, type V enzymes may require a unique gRNA.
[0152] The disclosed system can be used for a variety of applications, such as, for example, nucleic acid editing (e.g., gene editing), binding to nucleic acid molecules (e.g., sequence-specific binding), etc. Such systems can be used, for example, to address (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivating enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to prevent them from infecting host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors. EXAMPLES
[0153] In accordance with IUPAC convention, the following abbreviations are used throughout the examples: A=Adenine C=Cytosine G=guanine T=Thymine R = adenine or guanine Y = cytosine or thymine S = guanine or cytosine W = adenine or thymine K = guanine or thymine M = adenine or cytosine B=C, G, or T D=A, G, or T H=A, C, or T V=A, C, or G
[0154] Example 1 - Methods for metagenomic analysis of new proteins Metagenomic samples were collected from sediments, soils, and animals. Deoxyribonucleic acid (DNA) was extracted using Zymobiomics DNA miniprep kits and sequenced on an Illumina HiSeq® 2500. Samples were collected with the consent of the owners. Additional raw sequence data from public sources included animal microbiome, sediment, soil, hot springs, hydrothermal vents, ocean, peat, Permafrost, and sewage sequences. To identify new Cas effectors, metagenomic sequence data was searched using hidden Markov models generated based on known Cas protein sequences, including class II type V Cas effector proteins. Novel effector proteins identified by the search were aligned against known proteins to identify potential active sites. This metagenomic workflow resulted in the delineation of the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, and MG126 families described herein.
[0155] Example 2 - Discovery of the MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, and MG126 families of CRISPR systems Analysis of the data from the metagenomic analysis of Example 1 revealed a new cluster of previously undescribed putative CRISPR systems, including nine families (MG90, MG91A, MG91B, MG91C, MG118, MG119, MG120, MG122, and MG126). The corresponding protein and nucleic acid sequences of these new enzymes and their exemplary subdomains are presented as SEQ ID NOs: 1-325, 420-431, 476-624, or 629.
[0156] Example 3 - Template DNA for transcription and translation All MG VU and CasPhi nuclease E. coli codon-optimized sequences were ordered in plasmids with a T7 promoter (Twist Biosciences). Linear templates were amplified from the plasmids by PCR to include the T7 and nuclease sequences. Minimal array linear templates were amplified from sequences consisting of the T7 promoter, natural repeats, universal spacer, and natural repeats flanked by adapter sequences for amplification. The universal spacer matches the spacer in the 8N targeted library, with 8N mixed bases flanking the spacer for PAM determination. Three intergenic sequences near the ORF or CRISPR array were identified from metagenomic contigs and ordered as gBlocks with flanking adapter sequences for amplification (Integrated DNA Technologies).
[0157] Example 4 - In vitro transcription of crRNA, minimal arrays, and sgRNA RNA was produced by in vitro transcription using the HiScribe™ T7 High Yield RNA Synthesis Kit and purified using the Monarch® RNA Cleanup Kit (New England Biolabs Inc.). Templates for T7 transcription were varied. For crRNA, DNA oligos were designed with a T7 promoter, trimmed native repeats, and a universal spacer. For minimal arrays, the same templates as above were used. For sgRNA, DNA ultramers were designed with a T7 promoter, trimmed tracrRNA, GAAA tetraloop, trimmed native repeats, and a universal spacer. Minimal array templates were amplified with adapter primers. Templates for crRNA and sgRNA were ordered as reverse complements and annealed with a primer carrying the T7 promoter sequence in 1X IDT duplex buffer at 95°C for 2 minutes, followed by cooling to 22°C at 0.1°C / s to produce hybrid ds / ssDNA substrates suitable for transcription. After transcription, but before washing, each reaction was treated with DNAse I and incubated for 15 minutes at 37° C. All transcription products were verified for yield and purity via RNA TapeStation or via denaturing urea PAGE gels.
[0158] Example 5 - TXTL Expression Nucleases, intergenic sequences, and minimal arrays were expressed in transcription-translation reaction mixtures using the myTXTL® Sigma 70 Master Mix Kit (Arbor Biosciences). The final reaction mixture contained 5 nM nuclease DNA template, 12 nM intergenic DNA template, 15 nM minimal array DNA template, 1.1 nM pTXTL-P70a-T7rnap, and 1× myTXTL® Sigma 70 Master Mix. Reactions were incubated at 29° C. for 16 hours and then stored at 4° C.
[0159] Example 6 - PURExpress Expression 10 nM nuclease PCR template was expressed for cleavage with in vitro transcribed RNA using the PURExpress® In Vitro Protein Synthesis Kit (New England Biolabs Inc.) for 3 hours at 37° C. These reactions were used to test in vitro cleavage with 50 nM sgRNA or minimal array RNA following the same procedure as described in the cleavage reaction section.
[0160] Example 7 - E. coli Expression Plasmids encoding the effectors, intergenic sequences from the genomic contigs, natural repeats, and universal spacer sequences with a T7 promoter were transformed into BL21 DE3 or T7 Express lysY / Iq and grown at 37°C in 60 mL of terrific broth medium supplemented with 100 μg / mL ampicillin. Expression was induced with 0.4 mM IPTG after the cultures reached an OD600nm of 0.5 and were incubated overnight at 16°C. 25 mL of cells were pelleted by centrifugation and resuspended in 1.5 mL of lysis buffer (20 mM Tris-HCl, 500 mM NaCl, 1 mM TCEP, 5% glycerol, 10 mM MgCl2 pH 7.5 with Pierce Protease Inhibitor (Thermo Scientific™). Cells were then lysed by sonication. Supernatant and cell debris were separated by centrifugation.
[0161] Example 8 - Cleavage Reaction Plasmid library DNA cleavage reactions contained 5 nM target library, 5-fold dilutions of TXTL or PURExpress expression, 10 nM Tris-HCl, 10 nM MgCl 2, and 100 mM NaCl for 2 hours at 37°C. For reactions with E. coli expression, 10 μL of clarified lysate was added. The reactions were stopped, washed with HighPrep™ PCR clean-up beads (MAGBIO Genomics, Inc.), and eluted in Tris-EDTA pH 8.0 buffer. The ends of the 3 nM cleavage products were blunted with 3.33 μM dNTPs, 1× T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment (New England Biolabs Inc.) for 15 minutes at 25°C. The 1.5 nM cleavage products were ligated with 150 nM adapters, 1× T4 DNA ligase buffer (New England Biolabs Inc.), 20 U / μL T4 DNA ligase (New England Biolabs Inc.) for 20 minutes at room temperature. The ligated product was amplified by PCR using NGS primers and sequenced by NGS to obtain PAM. The in vitro activity of MG119-2 is shown in Figure 9, while the PAM determination for MG119-2 is shown in Figure 10.
[0162] Example 9 - RNAseq library preparation of intergenic enrichment from TXTL and E. coli lysates RNA was extracted from TXTL and cell lysate expression according to the Quick-RNA™ Miniprep Kit (Zymo Research) and eluted in 30-50 μL of water. Total transcript concentration was measured on the Nanodrop, Tapestation, and Qubit.
[0163] Total RNA from 100ng to 1ug from each sample was prepared for RNA sequencing using NEBNext Small RNA Library Prep Set for Illumina (New England Biolabs Inc.). Amplicons of 150-300bp were quantified by Tapestation and Qubit and pooled to a final concentration of 4nM. A final concentration of 12.5pM was loaded into a MiSeq V3 kit and sequenced for 176 total cycles on a Miseq system (Illumina). RNAseq reads were used to identify tracr sequences of genes.
[0164] Example 10 - Predicted RNA folding The predicted RNA folding of the active single RNA sequences was calculated at 37° C. using the method of Andronescu 2007. The shading of a base corresponds to the base pairing probability of that base.
[0165] Example 11 - In vitro cleavage efficiency (predictive) Proteins are expressed in E. coli protease-deficient B strain under a T7 inducible promoter, cells are lysed using sonication, and His-tagged proteins of interest are purified using HisTrap FF (GE Lifescience) Ni-NTA affinity chromatography on an AKTA Avant FPLC (GE Lifescience). Purity is determined using SDS-PAGE and densitometry in ImageLab software (Bio-Rad) of protein bands resolved on InstantBlue Ultrafast (Sigma-Aldrich) Coomassie-stained acrylamide gels (Bio-Rad). Proteins are desalted in a storage buffer composed of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol, pH 7.5, and stored at -80°C.
[0166] A target DNA is constructed containing a spacer sequence and a PAM determined via NGS. If there are degenerate bases within the PAM, a single representative PAM is selected for testing. The target DNA is a 2200 bp linear DNA derived from a plasmid via PCR amplification. The PAM and spacer are located 700 bp from one end. Successful cleavage results in fragments of 700 and 1500 bp.
[0167] Target DNA, in vitro transcribed single RNA, and purified recombinant protein are combined in cleavage buffer (10 mM Tris, 100 mM NaCl, 10 mM MgCl2) containing excess protein and RNA and incubated for 5'-3 h, typically 1 h. Reactions are stopped via addition of RNAse A and incubation at 60°C. Reactions are resolved on a 1.2% TAE agarose gel and the fraction of cleaved target DNA is quantified with ImageLab software.
[0168] Example 12 - Activity in E. coli (predictive) To test for nuclease activity in bacterial cells, strains are constructed with genomic sequences containing the target spacer and corresponding PAM sequence specific for the enzyme of interest. The engineered strains are then transformed with the nuclease of interest, after which the transformants are made chemically competent and transformed with 50 ng of a single guide either specific for the target sequence (on-target) or non-specific for the target (off-target). After heat shock, transformants are allowed to recover in SOC at 37°C for 2 hours, and nuclease efficiency is determined by a 5-fold dilution series grown on induction medium. Colonies are quantified from the dilution series in triplicate.
[0169] Example 13 - Activity in mammalian cells (predictive) To demonstrate targeting and cleavage activity in mammalian cells, the protein sequence is cloned into two mammalian expression vectors, one with a C-terminal SV40 NLS and a 2A-GFP tag, and one without a GFP tag and with two NLS sequences (one on the N-terminus and one on the C-terminus). Alternative NLS sequences can also be used. The DNA sequence of the protein can be a native sequence, an E. coli codon-optimized sequence, or a mammalian codon-optimized sequence. A single guide RNA sequence with the gene target of interest is also cloned into a mammalian expression vector. The two plasmids are co-transfected into HEK293T cells. 72 hours after co-transfection of the expression plasmid and the sgRNA targeting plasmid into HEK293T cells, DNA is extracted and used for preparation of an NGS library. NHEJ percent is measured via indels in the sequencing of the target site to demonstrate the targeting efficiency of the enzyme in mammalian cells. At least 10 different target sites are selected to test the activity of each protein.
[0170] Example 14 - Characterization of compact V-type nucleases in the MG119 family In silico identification of a novel compact V-type nuclease in the MG119 family
[0171] The detection of predicted proteins related to nuclease sequences in the MG119 family of compact type V nucleases was based on homology searches. Searches were performed using HMMER software (http: / / hmmer.org / ). Type V nuclease sequence hits were retained if they fulfilled the following criteria: (i) hmmsearch e-value was ≤ 10 -5(ii) the nuclease-encoding gene was within 1 kb from the CRISPR array, and (iii) the amino acid sequence length ranged from 350 to 700 aa. Sequences were clustered with 100% amino acid identity at coverage mode 1 and 80% coverage of the target sequence using MMSeqs2 (https: / / github.com / soedinglab / MMseqs2) (parameters --cov-mode 1-c 0.8--min-seq-id 1.0). Sequence representatives were selected to construct multiple sequence alignments using MAFFT (https: / / mafft.cbrc.jp / alignment / software / ) with the Needleman-Wunsch algorithm for global alignment, and phylogenetic trees were constructed using FastTree (https: / / doi.org / 10.1371 / journal.pone.0009490). Careful investigation of individual clades on the phylogenetic tree, including the genomic context of the nuclease genes, led to the identification of several novel compact type V nuclease sequences in the MG119 family (sequence numbers 476-624 and 629).
[0172] In vitro characterization to identify putative tracrRNAs To identify putative tracrRNA sequences, for example for nuclease MG119-2, the flanking intergenic sequences and minimal array were expressed in a transcription-translation reaction mixture using myTXTL® Sigma 70 Master Mix Kit (Arbor Biosciences). The final reaction mixture contained 5 nM nuclease DNA template, 12 nM intergenic DNA template, 15 nM minimal array DNA template, 0.1 nM pTXTL-P70a-T7rnap, and 1× myTXTL® Sigma 70 Master Mix. The reaction was incubated at 29° C. for 16 hours and then stored at 4° C.
[0173] Ribonucleoprotein complexes were tested via in vitro cleavage reactions. Plasmid DNA library cleavage reactions were performed using all possible 8N PAMs, 5-fold dilutions of TXTL expression, 10 nM Tris-HCl, 10 nM MgCl. 2 , and 100 mM NaCl, for 2 hours at 37° C. The reaction was stopped and washed with HighPrep™ PCR Clean-up Beads (MAGBIO Genomics, Inc.) and eluted in Tris-EDTA pH 8.0 buffer.
[0174] To obtain the PAM sequence, the ends of 3 nM cleavage products were blunted with 3.33 μM dNTPs, 1× T4 DNA ligase buffer, and 0.167 U / μL Klenow Fragment (New England Biolabs Inc.) for 15 min at 25° C. 1.5 nM cleavage products were ligated with 150 nM adapters, 1× T4 DNA ligase buffer (New England Biolabs Inc.), and 20 U / μL T4 DNA ligase (New England Biolabs Inc.) for 20 min at room temperature The ligated products were amplified by PCR using NGS primers and sequenced by NGS.
[0175] To obtain tracrRNA and crRNA sequences, RNA was extracted from TXTL lysates according to the Quick-RNA™ Miniprep Kit (Zymo Research) and eluted in 30–50 μL of water. 100 ng–1 μg of total RNA from each sample was prepared for RNA sequencing using the NEBNext Small RNA Library Prep Set for Illumina (New England Biolabs Inc.). Amplicons of 150–300 bp were quantified by Tapestation and Qubit and pooled to a final concentration of 4 nM. A final concentration of 12.5 pM was loaded into a MiSeq V3 kit and sequenced on a Miseq system (Illumina) for 176 total cycles. RNAseq reads were used to identify the tracr sequences of genes by mapping back to the original sequences.
[0176] In silico search for novel tracrRNA sequences To identify additional non-coding regions containing potential tracrRNAs, the sequence of the active tracrRNA was mapped to other contigs containing nucleases of the same nuclease family (e.g., MG119-1 and MG119-3). The newly identified sequences were used to generate a covariance model and predict additional tracrRNAs. The covariance model was built from a multiple sequence alignment (MSA) of the active and predicted tracrRNA sequences. The secondary structure of the MSA was obtained with RNAalifold (Vienna Package) and a covariance model was built with the Infernal package (http: / / eddylab.org / infernal / ). Other contigs containing candidate nucleases were searched using the covariance model with the Infernal command "cmsearch". TracrRNA candidates were tested in vitro (see below) and, in an iterative process, sequences from the active candidates were used to improve the covariance model and search for additional tracrRNAs in intergenic regions associated with other nuclease candidates.
[0177] sgRNA design The predicted tracrRNAs and their associated CRISPR repeat sequences obtained from the covariance model were modified to generate sgRNAs as follows (Figure 11A): the 3' end of the predicted tracrRNA sequence and the 5' end of the repeat sequence were trimmed and then connected with a GAAA tetraloop.
[0178] In vitro cleavage reactions to confirm nuclease activity and enable PAM determination 5 nM of nuclease amplified DNA template and 25 nM of sgRNA amplified DNA template (containing one of the spacer sequences listed in Table 2) were expressed for 3 hours at 37° C. using the PURExpress® In Vitro Protein Synthesis Kit (New England Biolabs Inc.). Plasmid library DNA cleavage reactions were run in 10 mM Tris-HCl pH 7.9, 10 mM MgCl, 5-fold dilutions of all possible 8N PAMs, 10 mM MgCl ... 2The reaction was performed by mixing 5 nM of target library representing 100 μg / mL BSA and 50 mM NaCl (NEB 2.1 Buffer, NEB Inc.) for 2 hours at 37° C. The reaction was stopped, washed with HighPrep™ PCR clean-up beads (MAGBIO Genomics, Inc.) and eluted in Tris-EDTA pH 8.0 buffer. The ends of 3 nM cleavage products were blunted with 3.33 μM dNTPs, 1× T4 DNA ligase buffer and 0.167 U / μL Klenow Fragment (New England Biolabs Inc.) for 15 minutes at 25° C. The 1.5 nM cleavage products were ligated with 150 nM adapters, 1× T4 DNA ligase buffer (New England Biolabs Inc.) and 20 U / μL T4 DNA ligase (New England Biolabs Inc.) for 20 minutes at room temperature. The ligated products were amplified by PCR using NGS primers and sequenced by NGS to obtain the PAMs. Active proteins that successfully cleaved the PAM library produced bands of approximately 188 or 205 bp in agarose gels, depending on which target site was encoded by the sgRNA (Figure 11B).
[0179] [Table 2]
[0180] The PAMs recognized by MG119 nuclease are shown as sequence logos generated in Seqlog Maker (Figure 12). The preferred cleavage positions in the target strand of the protospacer sequence complementary to the U40 spacer are listed in Table 3.
[0181] [Table 3]
[0182] Protein expression and purification Isolation of pure and functional proteins is essential for extensive in vitro analysis of biochemical properties and mechanistic studies. Expression and purification of the MG119 candidates were optimized to obtain sufficient quantity and quality of protein for such characterization. All constructs were transformed into E. coli (NEBExpress I). q The constructs were expressed in either the pMGB expression vector (MBP fusion), the pMGBΔ expression vector (no fusion protein), or both.
[0183] Protein expression Protein expression protocols for pMGB and pMGBΔ constructs are identical. Cultures were grown at 37°C in TB medium (Teknova T0690) or 2xYT medium (1.6% tryptone, 1% yeast extract, 0.5% NaCl) with 100 μg / L carbenicillin. At an OD600 of 0.8-1.2, cultures were induced with 0.5 mM IPTG (GoldBio I2481) and incubated overnight at 18°C or for 4-6 hours at 24°C, depending on the construct. Cultures were then harvested by centrifugation at 6,000×g for 10 min and pellets were resuspended in Nickel_A buffer (50 mM Tris pH 7.5, 750 mM NaCl, 10 mM MgCl 2 , 20 mM imidazole, 0.5 mM EDTA, 5% glycerol, 0.5 mM TCEP) plus protease inhibitors (Pierce Protease Inhibitor Tablets, EDTA-free, ThermoFisher A32965) and stored at -80°C.
[0184] Protein purification-pMGBΔ expression vector The protein expressed by this vector has the following sequence structure: 6xHis-(GS)2-PSP-nucleoplasmin bipartite NLS-(GGS)1-(GS)1-MG119-X-(GGS)3-SV40 NLS (Table 5). The protein expressed by this vector is designated MG119-XΔ. The cell pellet was thawed and the volume was supplemented to 120 mL with n-octyl- -D-glucoside detergent (P212121, CI-00234) at Cf = 0.5%. The sample was sonicated in an ice-water bath at 75% amplitude using a cycle of 15 sec on / 45 sec off for a total processing time of 3 min. Lysates were clarified by centrifugation at 30,000×g for 25 min and supernatant batches were bound to 5 mL Ni-NTA resin (HisPur Ni-NTA resin, ThermoFisher 88223) for ≧20 min. Samples were loaded onto a gravity column, washed with 30 CV of Nickel_A buffer, then eluted in 4 CV of Nickel_B buffer (Nickel_A buffer+250 mM imidazole) before being concentrated in a 50 kDa MWCO concentrator (Amicon Ultra-15, MilliporeSigma UFC9050). Samples were taken throughout the purification process and run on an SDS-PAGE protein gel (BioRad #4568126), which was imaged on a ChemiDoc in the unstained channel after 5 min of UV activation ( FIG. 13A ). The ΔMBP construct was then loaded onto a S200i 10 / 300 GL column (Cytiva 28-9909-44) and run into Nickel_A buffer (Figure 13B). Peak fractions were pooled and concentrated in a 50 kDa MWCO concentrator. Purification of proteins expressed in the pMGBΔ vector typically yielded 25-125 nmol of protein per L expression culture (Figure 13F).
[0185] Protein purification - pMGB expression vector The protein expressed by this vector has the following sequence structure: 6xHis-(GS)1-MBP-(GS)1-TEV-nucleoplasmin bipartite NLS-(GGGGS)3-(GS)1-MG119-X-(GGS)3-SV40 NLS (Table 5). The MBP fusion construct was purified identically to the pMGBΔ protein through lysis, clarification, affinity purification, and elution with Nickel_B (Figure 13C). After protein concentration in 50 kDa MWCO concentrators, TEV protease (GenScript Z03030) was added to each sample (Cf=1UI / μL) and incubated overnight at 4°C with gentle rotating end-over-end. The sample was centrifuged (21,000×g, 4° C., 10 min) to pellet aggregates, then the supernatant was batch bound to 3 mL of amylose resin (NEB E8021L) at 4° C. for 30 min, then loaded onto a gravity column. The flow-through was collected and concentrated in a 50 kDa MWCO concentrator (FIG. 13D). Again, the sample was centrifuged (21,000×g, 4° C., 10 min) to pellet aggregates, then loaded onto a S200i 10 / 300 GL column and run into Nickel_A buffer (FIG. 13E). Peak fractions were pooled and concentrated in a 50 kDa MWCO concentrator. Samples were taken throughout the purification process and run on an SDS-PAGE protein gel (BioRad #4568126), which was imaged on a ChemiDoc in the unstained channel after 5 min of UV activation (FIG. 13D).
[0186] Several selected MG119 candidates were purified from both pMGB and pMGBΔ expression vectors. Comparison of final protein yields normalized to initial expression culture volume shows a trend for higher expression yields from the pMGBΔ vector (Figure 13E). Purification of proteins expressed in the pMGBΔ vector typically yielded 2-15 nmol protein per L expression culture (Figure 13E). Protein purification yields are shown in Table 4.
[0187] [Table 4]
[0188] [Table 5]
[0189] In vitro cleavage efficiency using purified proteins The active fraction of protein aliquots was determined in a linear DNA substrate cleavage assay. Effector proteins were pre-incubated with a 2-fold molar excess of sgRNA for 20 min at room temperature to form ribonucleoprotein complexes (RNPs). Reactions were set up using 25 nM DNA substrate and a titration of 0.25X to 10X molar excess of RNP over substrate. The reaction buffer composition was 10 mM Tris pH 7.5, 10 mM MgCl 2 , and 100 mM NaCl. The DNA substrate is 522 bp long. Successful cleavage results in fragments of 172 and 350 bp. Reactions were incubated at 37° C. for 60 min, then at 75° C. for 10 min. RNase (NEB T3018) was added to each reaction (Cf=0.33 μg / μL) and the samples were incubated at 37° C. for 10 min. Proteinase K (NEB P8107) was added to each reaction (Cf=60 units / mL) and the samples were incubated at 55° C. for 15 min. Each entire reaction was then run on a 1.5% agarose gel with GelGreen dye (Biotium, #41005) (FIG. 14A) and imaged on a ChemiDoc in the GelGreen channel. The percent of cleaved substrate was calculated for each lane through densitometric analysis using BioRad's Image Lab software (version 6.1.0 build 7). The active fraction was determined by the slope of the linear range of cleavage (Figure 14B).
[0190] In vitro cleavage of purified Hepa1-6 genomic DNA with purified proteins To evaluate cleavage of purified mouse Hepa1-6 genomic DNA (gDNA), the mouse albumin gene was targeted in intron 1 (Table 6). gDNA was extracted from Hepa1-6 cell pellets with 8 million cells according to the PurelinkTM Genomic DNA Mini kit (Invitrogen) and eluted in 10 mM Tris-HCl at pH 8. sgRNA was ordered from Integrated DNA technologies (IDT) at 2 nmol and then resuspended in 10 mM Tris-EDTA buffer at 20 μM (Table 6). Ribonucleoproteins (RNPs) were diluted in 1X effector buffer (100 mM NaCl, 10 mM MgCl 2Nucleases were generated by pre-incubating nucleases with targeted or non-targeted guides at a 1:2 molar ratio in 10 mM Tris-HCl, pH 7.5 for 30 min at room temperature. All reactions were performed in triplicate, including a negative control with no sgRNA. After RNP formation, RNPs were added to digestion reactions containing 20 ng / μL purified gDNA in 1× effector buffer and incubated at 37° C. for 1 h. Nucleases were tested at two final concentrations: 7.8 and 15.6 nM. These concentrations were normalized by dividing the target concentration by the fractional activity of each nuclease. After incubation, these reactions were immediately transferred to 4°C, diluted 30X with water, and then prepared for qPCR in a master mix containing 1X PrimeTime® Gene Expression Master Mix, 10 μM forward primer, 10 μM reverse primer, and 5 μM 5'-FAM and ZEN / Iowa Black fluorescent quencher Taqman probe (IDT) (Table 7). An AriaMx Real-Time PCR System (Agilent) was used with the following cycles: 1) 95°C for 15 minutes, 2) 95°C for 5 seconds, and 3) 60°C for 1 minute, with steps 2-3 repeated 40X. Cq values were used to calculate the percent gDNA cleavage for each reaction according to the percent cleavage formula (below). All were normalized to the non-targeting control reaction. Figure 15A shows an example of an average of 60% gDNA cleavage with MG119-28 and sgRNA3, and 21% cleavage with sgRNA2, at the higher concentrations of protein used. Cutoff Percentage Formula Cut%=100-(2 -(Cq(実験的)-Cq(非標的化対照)) ×100)
[0191] [Table 6-1]
[0192] [Table 6-2]
[0193] [Table 7]
[0194] In vivo cleavage of genomic DNA in Hepa 1-6 cells with purified proteins Intracellular editing was demonstrated using a guide and nuclease RNP complex targeting the mouse albumin gene in intron 1 (Table 6). Hepa1-6 cells were thawed, washed, and resuspended in Dulbecco's modified Eagle's medium (DMEM, 10% FBS, and 1% Pen-strep). Cells were plated at 4 × 10 per 15 cm dish in 30 mL of medium at 37 °C. 6 Cells were seeded at a density of 1000 x 1000 cells / well. After 2 days, when the cells reached 70-80% confluence, the cells were split. The cells were trypsinized with 0.25% trypsin and then incubated at 37°C for 30 seconds. DMEM was added, then split into 3 mL and further diluted with 27 mL of medium. The split cells were incubated for another 2 days. Prior to nucleofection, the medium was aspirated from the plate and the cells were washed with 1X phosphate buffered saline (PBS, Gibco™) pH 7.2 and then trypsinized. The trypsin was neutralized and the cells were resuspended in DMEM. The cells in the cell suspension were counted with a Countess 3 FL (Invitrogen) to calculate the volume of cells to be pelleted. A total of 100,000 cells were required for each downstream process. Cells were centrifuged at 300×g for 7 min in a sorvall X Pro Series Centrifuge (Thermo Fisher) and then washed with PBS pH 7.2 before being resuspended in Nucleofector™ solution from the Amaxa™ 4D-Nucleofector™ Kit (Lonza).
[0195] RNP complexes were prepared separately by incubating 120 pmol of nuclease with 120 pmol of guide for 90 min at room temperature. 20 μL of prepared cells were added to the RNP. Nucleofection was performed as recommended by the Amaxa™ 4D-Nucleofector™ Protocol for the 4D-Nucleofector™ System (Lonza). Nucleofected cells were transferred from the nucleofection cassette to a 24-well plate with each well containing 500 μL of medium. After 2 days of incubation, gDNA from all treatments was extracted with QuickExtract (Lucigen) using the following cycles: 1) 65°C for 15 min, 2) 68°C for 15 min, and 3) 98°C for 10 min, then kept at 4°C until use. A 317 bp targeting window was amplified from the resulting extracted gDNA with Phusion Flash High-Fidelity PCR Master Mix (Thermo Fisher) using the following cycles: 1) 98°C for 10 s, 2) 98°C for 1 s, 3) 63°C for 5 s, 4) 72°C for 15 s, and 5) 72°C for 1 min, repeating steps 2-5 for 30 cycles, then held at 4°C. Amplicons were visualized on a 2% agarose gel, then washed and enriched with HighPrep Magnetic Beads (MagBio Genomics Inc.) with 1.8X bead volume relative to the sample. Samples were eluted with water. Indels were sequenced by NGS on a MiSeq (minimum 20,000 reads per sample) using 5% phiX v3 reagent kit (600 cycles, Table 8) and 2 × 301 bp paired-end reads. Indel analysis was performed using a modified CRISPResso2 program (Clement et al., 2019; https: / / doi.org / 10.1038 / s41587-019-0032-3) and the results are shown in Table 9 and Figure 15B.
[0196] [Table 8]
[0197] [Table 9]
[0198] Example 15 - Buffer optimization for MG119 protein purification (predictive) So far, MG119 protein has been purified in Nickel_A buffer. Nickel_A buffer is incompatible with downstream in vivo assays due to its high salt concentration, and rapid dilution into low salt solutions induces protein precipitation. To optimize the buffer for protein stability and downstream assay compatibility, MG119 nuclease is first purified in high salt buffer (750 mM NaCl) and gradually washed into Nickel_A buffer variants with 200 mM NaCl and zwitterionic amino acids L-arginine (50 mM) and L-glutamate (50 mM). Empirically, various stabilizing sugars (ribose, sorbitol, mannitol, xylitol) are also added to the buffer to enhance protein stability in low salt buffers.
[0199] Example 16 - Fluorescence-based measurement of nuclease activity (predictive) Novel cell line engineering Current assays used to measure nuclease activity in vivo (i.e., in mammalian cell lines) require extensive data analysis and turnaround times of up to one week. To facilitate the evaluation of in vivo nuclease activity, immortalized mammalian cell lines are engineered to provide immediate data on genomic DNA editing. K562 mammalian cells grown in IMDM (Gibco #12440053) + 10% FBS (Corning™ Regular Fetal Bovine Serum, MT35011CV) are used for this assay. K562 mammalian cells are transfected with 12 pmol of Cas9 protein (IDT #1081058), 60 pmol of sgRNA (Mali et al. Science, 2013 Feb 15;339(6121):823-6.), and 1200 ng of plasmid (pUC backbone) containing the expression sequence for mMBP-(GGS)3-eGFP protein. Genomic integration of this construct results in constitutive expression under the synthetic MND promoter. Cells are left to grow for 6 days and passaged every 3 days. Single-gene cell lines are isolated from single cells by sorting individual GFP-expressing cells into 96-well plates using a Sony MA900 Cell Sorter.
[0200] Fluorescence-based in vivo nuclease activity screening The appropriate sgRNA is designed to direct nuclease cleavage along the mMBP and eGFP genes such that indel formation results in frameshift mutations leading to loss of fluorescence. MG119 RNP complexes are formed by combining 100 pmol of protein with 200 pmol of sgRNA and incubating at room temperature for >20 minutes in a final volume of 5 μL. K562 cells are washed in 1x PBS and resuspended in Nucleofector solution (SF Cell Line 96-well Nucleofector™ solution) with approximately 200,000 cells per well. Cells and RNPs are combined in a Lonza 96-well nucleofector™ Kit, V4SC-2096, in a final volume of 25 μL, nucleofected (K562 cells, FF-120), and harvested in IMDM+10% FBS medium. Cells are left to recover at 37 °C for 2-3 days. To analyze, cells are washed twice with 1x PBS and then stained with 1x PBS+LIVE / DEAD Fixable Near-IR Dead Cell Stain Kit dye (ThermoFisher L10119) for 20 min at room temperature. Cells are washed once more with 1x PBS, then resuspended in 1x PBS and loaded onto an Attune NxT, Acoustic Focusing Flow Cytometer (model AFC2) for fluorescence analysis. Positive unedited controls (nucleofected without RNP) and negative controls (non-fluorescent K562 cells) are used to establish positive and negative fluorescent gates, and cell populations are analyzed for fluorescence loss in the GFP channel to assess in vivo nuclease activity.
[0201] Example 17 - Use for epigenome editing (predictive) Epigenome editing is a gene regulation technique that involves turning genes on or off constitutively or transiently. Such techniques may use catalytically dead Cas9 (dCas9) fused to three proteins: Dnmt3A, Dnmt3L, and KRAB (e.g., as described in Nunez et al. Cell 2021, 184(9), 2503-2519, which is incorporated herein by reference in its entirety). Dnmt3A and Dnmt3L are DNA methyltransferases. The KRAB domain mediates histone methylation. DNA and histone methylation in promoter regions mediates constitutive gene repression. dCas9 and guide RNAs can recruit DNA and histone methylation complexes to promoter regions without the need for nuclease activity. Together, Dnmt3A, Dnmt3L, and KRAB are 579aa, and dCas9 is 1,368aa. The fusion protein consists of 1,947 aa or 5,841 nucleotides, exceeding the adeno-associated viral vector (AAV) packaging limit (4.7 Kb). Therefore, there is a need to create more compact epigenome editors. Compact V-type nucleases from the MG119 family represent excellent candidates for use as dead nuclease partners in technologies for epigenome editing. When fused to DNA and histone methylation complexes, due to their small size, ranging from 350 to 700 aa, the size of the fusion protein can range, for example, from about 929 to about 1,279 aa, or from about 2787 to about 3837 nucleotides, allowing easy packaging in AAV.
[0202] To test the MG119 fusion protein as an epigenome editor, HEK293T cells expressing GFP under a chimeric promoter (GAPDH-Srnpn) are generated by lentiviral transduction. MG119 family guide RNAs targeting the chimeric promoter are designed. Guides are ordered from IDT and modified at the 5' and 3' nucleotides with three 2'-O-methyl substituents and three phosphorothioate linkages for stability. A dead version of MG119 nuclease is fused to the DNA and histone methylation complex (MG119 epigenome editor). The fusion protein is cloned in a mammalian expression plasmid under a CMV promoter. GFP-expressing HEK293T cells are transfected with a plasmid expressing the MG119 epigenome editor and the chemically synthesized guide. The transfected cells are analyzed by flow cytometry. Successful MG119 epigenome editors are determined by loss of GFP fluorescence in transfected cells. The MG119 epigenome editor is then used to target the gene of therapeutic interest.
[0203] [Table 10-1]
[0204] [Table 10-2]
[0205] [Table 10-3]
[0206] [Table 10-4]
[0207] [Table 10-5]
[0208]
Table 10-6
[0209]
Table 10-7
[0210]
Table 10-8
[0211]
Table 10-9
[0212]
Table 10-10
[0213]
Table 10-11
[0214]
Table 10-12
[0215]
Table 10-13
[0216]
Table 10-14
[0217]
Table 10-15
[0218]
Table 10-16
[0219]
Table 10-17
[0220]
Table 10-18
[0221]
Table 10-19
[0222]
Table 10-20
[0223]
Table 10-21
[0224]
Table 10-22
[0225]
Table 10-23
[0226]
Table 10-24
[0227]
Table 10-25
[0228] [Table 11]
[0229] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided for illustrative purposes only. The present invention is not intended to be limited by the specific examples provided herein. Although the present invention has been described with reference to the foregoing description, the description and explanation of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in the practice of the present invention. It is therefore contemplated that the present invention will encompass any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. An engineered nuclease system, comprising: (a) an endonuclease or a nucleic acid encoding said endonuclease, wherein said endonuclease is a class 2, type V endonuclease, and said endonuclease comprises a RuvC domain comprising an amino acid sequence having at least 75% sequence identity to the RuvC domain of SEQ ID NO: 57; (b) an engineered guide ribonucleic acid or a nucleic acid encoding said engineered guide ribonucleic acid, wherein said engineered guide ribonucleic acid is configured to form a complex with said endonuclease, and wherein said engineered guide ribonucleic acid comprises a spacer sequence configured to hybridize to a target nucleic acid sequence.
2. The engineered nuclease system of claim 1, wherein the RuvC domain comprises a RuvCI domain, a RuvCII domain, or a RuvCIII domain.
3. The engineered nuclease system described in claim 2, wherein the RuvCI domain, the RuvCII domain, or the RuvCIII domain respectively comprises the amino acid sequence of the RuvCI domain, the RuvCII domain, or the RuvCIII domain of SEQ ID NO:
57.
4. The engineered nuclease system of claim 2, wherein the endonuclease comprises the RuvCI domain, and the RuvCI domain comprises amino acid residues 267 to 425 of SEQ ID NO:
57.
5. The engineered nuclease system described in claim 2, wherein the endonuclease comprises the RuvCII domain, and the RuvCII domain comprises amino acid residues 467 to 488 of SEQ ID NO:
57.
6. The engineered nuclease system of claim 1, wherein the endonuclease comprises a WED II domain comprising an amino acid sequence having at least 75% sequence identity to the WED II domain of SEQ ID NO:
57.
7. The engineered nuclease system of claim 1, wherein the endonuclease comprises a WED II domain, and the WED II domain comprises the amino acid sequence of a WED II domain of SEQ ID NO:
57.
8. The engineered nuclease system of claim 1, wherein the endonuclease comprises a WED II domain, and wherein the WED II domain comprises an amino acid sequence having at least 70% sequence identity to residues 183-257 of SEQ ID NO:
57.
9. The engineered nuclease system of claim 1, wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least 80% sequence identity to the non-degenerate nucleotides of SEQ ID NO:
446.
10. The engineered nuclease system described in claim 9, wherein the engineered guide ribonucleic acid comprises a nucleotide sequence having at least 90% sequence identity to SEQ ID NO:
446.
11. The engineered nuclease system described in claim 10, wherein the engineered guide ribonucleic acid comprises the nucleotide sequence of SEQ ID NO:
446.
12. The engineered nuclease system of claim 1, wherein the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease.
13. The engineered nuclease system described in claim 12, wherein the one or more NLSs comprise a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 630-645.
14. The engineered nuclease system described in claim 13, wherein the one or more NLSs include sequence number 631 or sequence number 630.
15. The engineered nuclease system of claim 1, wherein the endonuclease comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:
57.
16. The engineered nuclease system described in claim 15, wherein the endonuclease comprises the amino acid sequence of SEQ ID NO:
57.
17. The engineered nuclease system of claim 1, wherein the endonuclease comprises the amino acid sequence of SEQ ID NO:57 and the engineered guide ribonucleic acid comprises the nucleotide sequence of SEQ ID NO:
446.
18. A method for modifying a target nucleic acid locus, the method comprising delivering an engineered nuclease system described in any one of claims 1 to 17 to the target nucleic acid locus.
19. A method for modifying a cell, the method comprising contacting the cell with a composition comprising an endonuclease or a nucleic acid encoding the endonuclease, wherein the endonuclease comprises an amino acid sequence having at least 75% sequence identity to SEQ ID NO:
57.
20. The method of claim 19, wherein the contacting is in vitro or ex vivo.