Enzymes with ruvc domains
The engineered nuclease system, featuring a class 2 type IICas endonuclease with RuvC_III and HNH domains, and a customized guide RNA structure, addresses the limitations of current CRISPR/Cas systems by providing high specificity and efficiency in DNA targeting and cleavage.
Patent Information
- Application Number
- JP2025027249
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-04-27
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current CRISPR/Cas systems for DNA manipulation and gene editing lack specificity and efficiency, particularly in targeting and cleaving specific nucleic acid sequences with high precision.
An engineered nuclease system comprising a class 2 type IICas endonuclease with an RuvC_III domain and an HNH domain, derived from uncultured microorganisms, paired with an engineered guide ribonucleic acid structure that includes a guide RNA sequence complementary to a target DNA sequence and a tracr RNA sequence for binding to the endonuclease.
The engineered nuclease system achieves high specificity and efficiency in targeting and cleaving specific DNA sequences, enhancing the precision and effectiveness of DNA manipulation and gene editing applications.
Smart Images

Figure 2025090617000001_ABST
Abstract
Description
Technical Field
[0001] Related Applications This application claims priority to U.S. Provisional Application No. 63 / 022,320, filed May 8, 2020, entitled "ENZYMES WITH RUVC DOMAINS"; U.S. Provisional Application No. 63 / 032,464, filed May 29, 2020, entitled "ENZYMES WITH RUVC DOMAINS"; U.S. Provisional Application No. 63 / 116,155, filed November 19, 2020, entitled "ENZYMES WITH RUVC DOMAINS"; and U.S. Provisional Application No. 63 / 180,570, filed April 27, 2021, entitled "ENZYMES WITH RUVC DOMAINS", all of which are hereby incorporated by reference in their entirety.
Background Art
[0002] Cas enzymes, along with their associated Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) guide ribonucleic acid (RNA), appear to be a widespread (about 45% of bacteria, about 84% of archaea) component of prokaryotic immune systems that help protect such microorganisms from non-self nucleic acids such as infectious viruses and plasmids by CRISPR-RNA-guided nucleic acid cleavage. Deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements can be relatively conserved in structure and length, but their CRISPR-associated (Cas) proteins are extremely diverse and contain a wide variety of nucleic acid interaction domains. Although CRISPR DNA elements were observed as early as 1987, the programmable endonuclease cleavage ability of CRISPR / Cas complexes was only recently recognized, leading to the use of recombinant CRISPR / Cas systems in a variety of DNA manipulation and gene editing applications.
[0003] Sequence Listing This application includes a Sequence Listing submitted electronically in ASCII format, which is hereby incorporated by reference in its entirety. The name of the ASCII copy created on May 29, 2020 is 55921_712_601_SL.txt, and the size is 24,659,439 bytes.
Summary of the Invention
[0004] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising an RuvC_III domain and an HNH domain, the endonuclease being derived from an uncultured microorganism and being a class 2 type IICas endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the RuvC_III domain comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% sequence identity to any one of SEQ ID NOs: 1827-3637.
[0005] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising an RuvC_III domain having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% sequence identity to any one of SEQ ID NOs: 1827 - 3637; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease.
[0006] In some aspects, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising SEQ ID NOs: 5512 - 5537, wherein the endonuclease is a class 2 type IICas endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a tracr ribonucleic acid sequence configured to bind to the endonuclease.
[0007] In some embodiments, the endonuclease is derived from uncultured microorganisms. In some embodiments, the endonuclease has not been engineered to bind to different PAM sequences. In some embodiments, the endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the endonuclease has less than 80% identity to the Cas9 endonuclease. In some embodiments, the endonuclease further comprises an HNH domain. In some embodiments, the tracr ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to about 60-90 consecutive nucleotides selected from any one of SEQ ID NOs: 5476-5511 and SEQ ID NO: 5538.
[0008] In some aspects, the disclosure provides an engineered nuclease system comprising: (a) an engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a tracr ribonucleic acid sequence configured to bind to an endonuclease, wherein the tracr ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to about 60-90 consecutive nucleotides selected from any one of SEQ ID NOs: 5476-5511 and SEQ ID NO: 5538, and (b) a class 2 type II Cas endonuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs: 5512-5537.
[0009] In some embodiments, the engineered guide ribonucleic acid structure comprises at least two ribonucleic acid polynucleotides. In some embodiments, the engineered guide ribonucleic acid structure comprises one ribonucleic acid polynucleotide comprising a guide ribonucleic acid sequence and a tracr ribonucleic acid sequence.
[0010] In some embodiments, the guide ribonucleic acid sequence is complementary to a prokaryotic, bacterial, archaeal, eukaryotic, fungal, plant, mammalian, or human genomic sequence. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLSs) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612.
[0011] In some embodiments, the engineered nuclease system further comprises a single-stranded or double-stranded DNA repair template comprising, in order from 5' to 3', a first homology arm comprising at least 20 nucleotides that are 5' to the target deoxyribonucleic acid sequence, at least 10 nucleotides of a synthetic DNA sequence, and a second homology arm comprising at least 20 nucleotides that are 3' to the target sequence. In some embodiments, the first homology arm or the second homology arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides.
[0012] In some embodiments, the system further comprises a source of Mg2+.
[0013] In some embodiments, the endonuclease and the tracr ribonucleic acid sequence are derived from different bacterial species within the same phylum. In some embodiments, the endonuclease is derived from a bacterium belonging to the genus Dermabacter. In some embodiments, the endonuclease is derived from a bacterium belonging to the phylum Verrucomicrobia, the phylum Candidatus Peregrinibacteria, or the phylum Candidatus Melainabacteria. In some embodiments, the endonuclease is derived from a bacterium comprising a 16S rRNA gene having at least 90% identity to any one of SEQ ID NOs: 5592-5595.
[0014] In some embodiments, the HNH domain comprises a sequence having at least 70% or at least 80% identity to any one of SEQ ID NOs: 5638-5460. In some embodiments, the endonuclease comprises SEQ ID NOs: 1-1826, or a variant thereof having at least 55% identity thereto. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 1827-1830 or SEQ ID NOs: 1827-2140.
[0015] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 3638-3641 or SEQ ID NOs: 3638-3954. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NOs: 5615-5632. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 1-4 or SEQ ID NOs: 1-319.
[0016] In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 5461-5464, SEQ ID NOs: 5476-5479, or SEQ ID NOs: 5476-5489. In some embodiments, the guide RNA structure comprises an RNA sequence predicted to comprise a hairpin consisting of a stem and a loop, wherein the stem comprises at least 10, at least 12, or at least 14 base pairs of ribonucleotides and an asymmetric bulge within 4 base pairs of the loop.
[0017] In some embodiments, the endonuclease is configured to bind to a PAM and comprises a sequence selected from the group consisting of SEQ ID NOs: 5512-5515 or SEQ ID NOs: 5527-5530.
[0018] In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1827, (b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5461 or SEQ ID NO: 5476, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5512 or SEQ ID NO: 5527. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1828, (b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5462 or SEQ ID NO: 5477, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5513 or SEQ ID NO: 5528. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1829, (b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5463 or SEQ ID NO: 5478, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5514 or SEQ ID NO: 5529. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to SEQ ID NO: 1830, (b) the guide RNA structure comprises a sequence that is at least 70%, at least 80%, or at least 90% identical to at least one of SEQ ID NO: 5464 or SEQ ID NO: 5479, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5515 or SEQ ID NO: 5530.
[0019] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 2141-2142 or SEQ ID NO: 2141-2241. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3955-3956 or SEQ ID NO: 3955-4055. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NO: 5632-5638. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 320-321 or SEQ ID NO: 320-420. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5465, SEQ ID NO: 5490-5491, or SEQ ID NO: 5490-5494. In some embodiments, the guide RNA structure comprises a tracr ribonucleic acid sequence comprising a hairpin comprising at least 8, at least 10, or at least 12 base pairs of ribonucleotides. In some embodiments, the endonuclease is configured to bind to a PAM and comprises a sequence selected from the group consisting of SEQ ID NO: 5516 or SEQ ID NO: 5531. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2141, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5490, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5531.In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2142, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5465 or SEQ ID NO: 5491, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5516.
[0020] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2245 - 2246. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4059 - 4060. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NOs: 5639 - 5648. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 424 - 425. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 5498 - 5499 and SEQ ID NO: 5539. In some embodiments, the guide RNA structure comprises a guide ribonucleic acid sequence predicted to comprise a hairpin having an uninterrupted base pair region comprising at least 8 nucleotides of the guide ribonucleic acid sequence and at least 8 nucleotides of the tracr ribonucleic acid sequence, wherein the tracr ribonucleic acid sequence comprises a first hairpin and a second hairpin in the 5' to 3' direction, and the first hairpin has a longer stem than the second hairpin.
[0021] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2242-2244 or SEQ ID NOs: 2247-2249. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4056-4058 or SEQ ID NOs: 4061-4063. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NOs: 5639-5648. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 421-423 or SEQ ID NOs: 426-428. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 5466-5467, SEQ ID NOs: 5495-5497, SEQ ID NOs: 5500-5502, and SEQ ID NO: 5539. In some embodiments, the guide RNA structure comprises a guide ribonucleic acid sequence predicted to comprise a hairpin having an uninterrupted base pair region comprising at least 8 nucleotides of the guide ribonucleic acid sequence and at least 8 nucleotides of the tracr ribonucleic acid sequence, wherein the tracr ribonucleic acid sequence comprises a first hairpin and a second hairpin in the 5' to 3' direction, and the first hairpin has a longer stem than the second hairpin. In some embodiments, the endonuclease is configured to bind to a PAM and comprises a sequence selected from the group consisting of SEQ ID NOs: 5517-5518 or SEQ ID NOs: 5532-5534. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2247, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5500, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5517 or SEQ ID NO: 5532.In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2248, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5501, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5518 or SEQ ID NO: 5533. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2249, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5502, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5534.
[0022] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 2253 or SEQ ID NOs: 2253-2481. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 4067 or SEQ ID NOs: 4067-4295. In some embodiments, the endonuclease comprises a peptide motif according to SEQ ID NO: 5649. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 432 or SEQ ID NOs: 432-660. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5468 or SEQ ID NO: 5503. In some embodiments, the endonuclease is configured to bind to a PAM and comprises a sequence selected from the group consisting of SEQ ID NO: 5519. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2253, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5468 or SEQ ID NO: 5503, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5519.
[0023] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2482-2489. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4296-4303. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 661-668. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 2490-2498. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 4304-4312. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 669-677. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5504.
[0024] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 2499 or SEQ ID NO: 2499-2750. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 4313 or SEQ ID NO: 4313-4564. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NO: 5650-5667. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 678 or SEQ ID NO: 678-929. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5469 or SEQ ID NO: 5505. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5520 or SEQ ID NO: 5535. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2499, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5469 or SEQ ID NO: 5505, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5520 or SEQ ID NO: 5535.
[0025] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 2751 or SEQ ID NO: 2751-2913. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 4565 or SEQ ID NO: 4565-4727. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NO: 5668-5678. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 930 or SEQ ID NO: 930-1092. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5470 or SEQ ID NO: 5506. In some embodiments, the endonuclease is configured to bind to a PAM and comprises a sequence selected from the group consisting of SEQ ID NO: 5521 or SEQ ID NO: 5536. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2751, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5470 or SEQ ID NO: 5506, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5521 or SEQ ID NO: 5536.
[0026] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 2914 or SEQ ID NOs: 2914 - 3174. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 4728 or SEQ ID NOs: 4728 - 4988. In some embodiments, the endonuclease comprises at least 1, at least 2, or at least 3 peptide motifs selected from the group consisting of SEQ ID NOs: 5676 - 5678. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1093 or SEQ ID NOs: 1093 - 1353. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5471, SEQ ID NO: 5507, and SEQ ID NOs: 5540 - 5542. In some embodiments, the guide RNA structure comprises a tracr ribonucleic acid sequence predicted to comprise at least two hairpins containing ribonucleotides less than 5 base pairs. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5522. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 2914, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5471 or SEQ ID NO: 5507, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5522.
[0027] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3175 or SEQ ID NO: 3175-3330. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 4989 or SEQ ID NO: 4989-5146. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NO: 5679-5686. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1354 or SEQ ID NO: 1354-1511. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5472 or SEQ ID NO: 5508. In some embodiments, the endonuclease is configured to bind to a PAM and comprises a sequence selected from the group consisting of SEQ ID NO: 5523 or SEQ ID NO: 5537. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 3175, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5472 or SEQ ID NO: 5508, and (c) the endonuclease is configured to bind to a PAM that comprises SEQ ID NO: 5523 or SEQ ID NO: 5537.
[0028] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3331 or SEQ ID NOs: 3331-3474. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5147 or SEQ ID NOs: 5147-5290. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NOs: 5674-5675 and SEQ ID NOs: 5687-5693. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1512 or SEQ ID NOs: 1512-1655. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5473 or SEQ ID NO: 5509. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5524. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 3331, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5473 or SEQ ID NO: 5509, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5524.
[0029] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3475 or SEQ ID NO: 3475-3568. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5291 or SEQ ID NO: 5291-5389. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NO: 5694-5699. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1656 or SEQ ID NO: 1656-1755. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5474 or SEQ ID NO: 5510. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5525. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 3475, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5474 or SEQ ID NO: 5510, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5525.
[0030] In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 3569 or SEQ ID NOs: 3569-3637. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 5390 or SEQ ID NOs: 5390-5460. In some embodiments, the endonuclease comprises at least 1, at least 2, at least 3, at least 4, or at least 5 peptide motifs selected from the group consisting of SEQ ID NOs: 5700-5717. In some embodiments, the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to a sequence selected from the group consisting of SEQ ID NO: 1756 or SEQ ID NOs: 1756-1826. In some embodiments, the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5475 or SEQ ID NO: 5511. In some embodiments, the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5526. In some embodiments, (a) the endonuclease comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 3569, (b) the guide RNA structure comprises a sequence that is at least 70%, 80%, or 90% identical to SEQ ID NO: 5475 or SEQ ID NO: 5511, and (c) the endonuclease is configured to bind to a PAM comprising SEQ ID NO: 5526. In some embodiments, sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) parameter of 10, and a BLOSUM62 scoring matrix gap cost of 11 for existence and 1 for extension, and using conditional composition score matrix adjustment.
[0031] In some embodiments, the present disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and (b) a protein binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other using intervening nucleotides, and the engineered guide ribonucleic acid polynucleotide forms a complex with an endonuclease comprising an RuvC_III domain having a sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% sequence identity to any one of SEQ ID NOs: 1827 to 3637, or is configured to target the complex to the target sequence of the target DNA molecule. In some embodiments, the DNA targeting segment is located 5' to both of the two complementary stretches of nucleotides.
[0032] In some embodiments, (a) the protein-binding segment comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 88%, at least 90%, at least 92%, at least 95%, or at least 98% identity to a sequence selected from the group consisting of SEQ ID NOs: 5476-5479 or SEQ ID NOs: 5476-5489, (b) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to a sequence selected from the group consisting of (SEQ ID NOs: 5490-5491 or SEQ ID NOs: 5490-5494) and SEQ ID NO: 5538, (c) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5498-5499, (d) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5495-5497 and SEQ ID NOs: 5500-5502, (e) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5503, (f) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5504, (g) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5505, (h) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5506, (i) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5507, (j) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5508, (k) the protein-binding segment comprises a sequence having at least 70%, at least 80%,or comprises a sequence having at least 90% identity, and (l) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5510, or (m) the protein-binding segment comprises a sequence having at least 70%, at least 80%, or at least 90% identity to SEQ ID NO: 5511.,
[0033] In some embodiments, (a) the guide ribonucleic acid polynucleotide comprises an RNA sequence comprising a hairpin comprising a stem and a loop, wherein the stem comprises at least 10, at least 12, or at least 14 ribonucleotide base pairs and an asymmetric bulge within 4 base pairs of the loop, (b) the guide ribonucleic acid polynucleotide is predicted to comprise a hairpin comprising at least 8, at least 10, or at least 12 ribonucleotide base pairs and to comprise a tracr ribonucleic acid sequence, (c) the guide ribonucleic acid polynucleotide comprises a guide ribonucleic acid sequence predicted to comprise a hairpin having an uninterrupted base pair region comprising at least 8 nucleotides of the guide ribonucleic acid sequence and at least 8 nucleotides of the tracr ribonucleic acid sequence, wherein the tracr ribonucleic acid sequence comprises a first hairpin and a second hairpin in the 5' to 3' direction, and the first hairpin has a longer stem than the second hairpin, or (d) the guide ribonucleic acid polynucleotide is predicted to comprise at least two hairpins comprising less than 5 ribonucleotide base pairs and to comprise a tracr ribonucleic acid sequence.,
[0034] In some aspects, the disclosure provides a deoxyribonucleic acid polynucleotide encoding any of the engineered guide ribonucleic acid polynucleotides described herein.,
[0035] In some embodiments, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a Class 2 type IIC Cas endonuclease comprising an RuvC_III domain and an HNH domain, and the endonuclease is derived from an uncultured microorganism.
[0036] In some embodiments, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes an endonuclease comprising an RuvC_III domain having at least 70% sequence identity to any one of SEQ ID NOs: 1827-3637. In some embodiments, the endonuclease comprises an HNH domain having at least 70% or at least 80% sequence identity to any one of SEQ ID NOs: 3638-5460. In some embodiments, the endonuclease comprises SEQ ID NOs: 5572-5591, or a variant thereof having at least 70% sequence identity thereto. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLSs) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612.
[0037] In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human. In some embodiments, the organism is Escherichia coli (E. coli), and (a) the nucleic acid sequence has at least 70%, 80%, or 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5572-5575, (b) the nucleic acid sequence has at least 70%, 80%, or 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5576-5577, (c) the nucleic acid sequence has at least 70%, 80%, or 90% identity to a sequence selected from the group consisting of SEQ ID NOs: 5578-5580, (d) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5581, (e) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5582, (f) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5583, (g) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5584, (h) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5585, (i) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5586, or (j) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5587. In some embodiments, the organism is a human, and (a) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5588 or SEQ ID NO: 5589, or (b) the nucleic acid sequence has at least 70%, 80%, or 90% identity to SEQ ID NO: 5590 or SEQ ID NO: 5591.
[0038] In some aspects, the disclosure provides a vector comprising a nucleic acid sequence encoding a Class 2 type IIC Cas endonuclease comprising an RuvC_III domain and an HNH domain, wherein the endonuclease is derived from an uncultured microorganism.
[0039] In some aspects, the present disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with an endonuclease, the endonuclease comprising a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence and a tracr ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
[0040] In some aspects, the present disclosure provides a cell comprising any of the vectors described herein.
[0041] In some aspects, the present disclosure provides a method for producing an endonuclease, comprising culturing any of the cells described herein.
[0042] In some aspects, the present disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide in a complex with a class 2 type IIC Cas endonuclease and an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide; (b) the double-stranded deoxyribonucleic acid polynucleotide comprising a protospacer adjacent motif (PAM); and (c) the PAM comprising a sequence selected from the group consisting of SEQ ID NOs: 5512-5526 or SEQ ID NOs: 5527-5537. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to the sequence of the engineered guide ribonucleic acid structure and a second strand comprising the PAM. In some embodiments, the PAM is directly adjacent to the 3' end of the sequence complementary to the sequence of the engineered guide ribonucleic acid structure.
[0043] In some embodiments, the Class 2 type II Cas endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease. In some embodiments, the Class 2 type II Cas endonuclease is derived from uncultured microorganisms. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0044] In some embodiments, (a) the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 5512-5515 and SEQ ID NOs: 5527-5530, (b) the PAM comprises SEQ ID NO: 5516 or SEQ ID NO: 5531, (c) the PAM comprises SEQ ID NO: 5539, (d) the PAM comprises SEQ ID NO: 5517 or SEQ ID NO: 5518, (e) the PAM comprises SEQ ID NO: 5519, (f) the PAM comprises SEQ ID NO: 5520 or SEQ ID NO: 5535, (g) the PAM comprises SEQ ID NO: 5521 or SEQ ID NO: 5536, (h) the PAM comprises SEQ ID NO: 5522, (i) the PAM comprises SEQ ID NO: 5523 or SEQ ID NO: 5537, (j) the PAM comprises SEQ ID NO: 5524, (k) the PAM comprises SEQ ID NO: 5525, or (l) the PAM comprises SEQ ID NO: 5526.
[0045] In some aspects, the present disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus any of the engineered nuclease systems described herein, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, and the complex is configured to modify the target nucleic acid locus when the complex binds to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is intracellular. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell.
[0046] In some embodiments, delivering an engineered nuclease system to a target nucleic acid locus comprises delivering any of the nucleic acids described herein or any of the vectors described herein. In some embodiments, delivering an engineered nuclease system to a target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding an endonuclease. In some embodiments, the nucleic acid comprises a promoter operably linked to an open reading frame encoding an endonuclease. In some embodiments, the engineered nuclease system for a target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding an endonuclease. In some embodiments, the engineered nuclease system for a target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, the engineered nuclease system for a target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide ribonucleic acid (RNA) structure operably linked to an RNA pol III promoter. In some embodiments, the endonuclease induces a single-strand break or a double-strand break at or near the target locus.
[0047] In some embodiments, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846 or SEQ ID NO: 6257; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the present disclosure provides an engineered nuclease system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising SEQ ID NOs: 5847-5861 or 6258-6278, wherein the endonuclease is a class 2 type IICas endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (i) a guide ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (ii) a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the endonuclease is derived from an uncultured microorganism. In some embodiments, the endonuclease has not been engineered to bind to different PAM sequences. In some embodiments, the endonuclease has less than 80% identity to Cas9 endonuclease. In some embodiments, the endonuclease is not a Cas9 endonuclease, Cas14 endonuclease, Cas12a endonuclease, Cas12b endonuclease, Cas12c endonuclease, Cas12d endonuclease, Cas12e endonuclease, Cas13a endonuclease, Cas13b endonuclease, Cas13c endonuclease, or Cas13d endonuclease.In some embodiments, the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to any one of (a) any one of SEQ ID NOs: 5886-5887, 5891, 5893, or 5894, or (b) any one of SEQ ID NOs: 5862-5885, 5888-5890, 5892, 5895-5896, or 6279-6301 non-degenerate nucleotides. In some aspects, the present disclosure provides an engineered nuclease system comprising (a) an engineered guide ribonucleic acid structure comprising (i) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) a ribonucleic acid sequence configured to bind to an endonuclease, wherein the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to any one of (a) any one of SEQ ID NOs: 5886-5887, 5891, 5893, or 5894, or (b) any one of SEQ ID NOs: 5862-5885, 5888-5890, 5892, 5895-5896, or 6279-6301 non-degenerate nucleotides, and (b) a class 2 type IICas endonuclease configured to bind to the engineered guide ribonucleic acid. In some embodiments, the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs: 5847-5861 or SEQ ID NOs: 6258-6278. In some embodiments, the guide ribonucleic acid sequence is 15-24 nucleotides in length or 19-24 nucleotides in length. In some embodiments, the endonuclease comprises one or more nuclear localization sequences (NLSs) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612. In some embodiments, the system further comprises a single-stranded or double-stranded DNA repair template comprising, in order from 5' to 3', a first homology arm comprising at least 20 nucleotides that are 5' to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homology arm comprising at least 20 nucleotides that are 3' to the target sequence.In some embodiments, the first homology arm or the second homology arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. In some embodiments, the sequence identity is determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) parameter of 10, and a BLOSUM62 scoring matrix gap cost of 11 for existence and 1 for extension, and using a composition-based score matrix adjustment.
[0048] In some aspects, the disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and (b) a protein binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other using intervening nucleotides, and the engineered guide ribonucleic acid polynucleotide forms a complex with an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846 or SEQ ID NO: 6257 and is configured to target the complex to the target sequence of the target DNA molecule. In some embodiments, the DNA targeting segment is located 5' to both of the two complementary stretches of nucleotides.
[0049] In some aspects, the disclosure provides a deoxyribonucleic acid polynucleotide encoding any of the engineered guide ribonucleic acid polynucleotides described herein.
[0050] In some embodiments, the present disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846 or SEQ ID NO: 6257. In some embodiments, the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLSs) proximal to the N-terminus or C-terminus of the endonuclease. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612. In some embodiments, the organism is a prokaryote, bacterium, eukaryote, fungus, plant, mammal, rodent, or human.
[0051] In some embodiments, the present disclosure provides a vector comprising any of the nucleic acids described herein. In some embodiments, the vector further comprises a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, the engineered guide ribonucleic acid structure comprising: (a) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence; and (b) a ribonucleic acid sequence configured to bind to the endonuclease. In some embodiments, the vector is a plasmid, minicircle, CELiD, adeno-associated virus (AAV)-derived virion, or lentivirus.
[0052] In some embodiments, the present disclosure provides a cell comprising any of the vectors described herein.
[0053] In some embodiments, the present disclosure provides a method for producing an endonuclease, the method comprising culturing any of the cells described herein.
[0054] In some aspects, the present disclosure provides a method for binding, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide, the method comprising contacting the double-stranded deoxyribonucleic acid polynucleotide in a complex with a Class 2 type IIC Cas endonuclease and an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and the PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 5847-5861 or SEQ ID NOs: 6258-6278. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to the sequence of the engineered guide ribonucleic acid structure and a second strand comprising the PAM. In some embodiments, the PAM is directly adjacent to the 3' end of the sequence complementary to the sequence of the engineered guide ribonucleic acid structure. In some embodiments, the Class 2 type IIC Cas endonuclease is not Cas9 endonuclease, Cas14 endonuclease, Cas12a endonuclease, Cas12b endonuclease, Cas12c endonuclease, Cas12d endonuclease, Cas12e endonuclease, Cas13a endonuclease, Cas13b endonuclease, Cas13c endonuclease, or Cas13d endonuclease. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0055] In some embodiments, the present disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus any of the engineered nuclease systems described herein, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, and the complex is configured to modify the target nucleic acid locus when the complex binds to the target nucleic acid locus. In some embodiments, the target nucleic acid locus is modified by binding, nicking, cleaving, or marking the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is intracellular. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering any of the nucleic acids described herein or any of the vectors described herein. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease. In some embodiments, the nucleic acid comprises a promoter operably linked to the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease. In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide.In some embodiments, delivering the engineered nuclease system to the target nucleic acid locus comprises delivering deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid (RNA) structure operably linked to an RNA pol III promoter. In some embodiments, the endonuclease induces a single-strand break or a double-strand break at or near the target locus.
[0056] In some aspects, the present disclosure provides a method of editing the TRAC locus within a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the TRAC locus, and the engineered guide RNA comprises a target sequence having at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to any one of SEQ ID NOs: 5950-5958 or SEQ ID NOs: 5959-5965 for at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 consecutive nucleotides. In some embodiments, the RNA-guided endonuclease is a Class II type IICas endonuclease. In some embodiments, the RNA-guided endonuclease comprises an RuvCIII domain having at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5950-5958, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5959-5965, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5953-5957. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5960-5961 or SEQ ID NOs: 5963-5964.
[0057] In some embodiments, the present disclosure provides a method for editing the TRBC locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the TRBC locus, and the engineered guide RNA comprises a target sequence having at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to any one of SEQ ID NOs: 5966-6004 or SEQ ID NOs: 6005-6025 for at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 consecutive nucleotides. In some embodiments, the RNA-guided endonuclease is a Class II type IICas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain having at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5966 - 6004, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6005 - 6025, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5970, 5971, 5983, or 5984. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6006, 6010, 6011, or 6012.
[0058] In some aspects, the present disclosure provides a method of editing the GR (NR3C1) locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the GR (NR3C1) locus, and the engineered guide RNA comprises a target sequence having at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to any one of SEQ ID NOs: 6026-6090 or SEQ ID NOs: 6091-6121 for at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 consecutive nucleotides. In some embodiments, the RNA-guided endonuclease is a Class II type IICas endonuclease. In some embodiments, the RNA-guided endonuclease comprises an RuvCIII domain having at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6026 to 6090, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6091 to 6121, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6027 to 6028, 6029, 6038, 6043, 6049, 6076, 6080, 6081, or 6086. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6092, 6115, or 6119.
[0059] In some embodiments, the present disclosure provides a method of editing the AAVS1 locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the AAVS1 locus, and the engineered guide RNA comprises a target sequence having at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to any one of SEQ ID NOs: 6122-6152 for at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 consecutive nucleotides. In some embodiments, the RNA-guided endonuclease is a Class II type IICas endonuclease. In some embodiments, the RNA-guided endonuclease comprises an RuvCIII domain having at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NO: 6122, 6125-6126, 6128, 6131, 6133, 6136, 6141, 6143, or 6148.
[0060] In some aspects, the present disclosure provides a method of editing a TIGIT locus in a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the TIGIT locus, and the engineered guide RNA comprises a target sequence having at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to any one of SEQ ID NOs: 6153-6181 for at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 consecutive nucleotides. In some embodiments, the RNA-guided endonuclease is a Class II type IIC Cas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423.In some embodiments, the RNA-guided endonuclease comprises an RuvCIII domain comprising a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NO: 66155, 6159, 616, or 6172.
[0061] In some aspects, the present disclosure provides a method of editing the CD38 locus within a cell, the method comprising contacting the cell with (a) an RNA-guided endonuclease and (b) an engineered guide RNA, wherein the engineered guide RNA is configured to form a complex with the endonuclease, the engineered guide RNA comprises a spacer sequence configured to hybridize to a region of the CD38 locus, and the engineered guide RNA comprises a target sequence having at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to any one of SEQ ID NOs: 6182-6248 or SEQ ID NOs: 6249-6256 for at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 consecutive nucleotides. In some embodiments, the RNA-guided endonuclease is a class II type IICas endonuclease. In some embodiments, the RNA-guided endonuclease comprises a RuvCIII domain having at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain.In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75% identity, at least 80% identity, at least 82% identity, at least 84% identity, at least 86% identity, at least 88% identity, at least 90% identity, at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6182-6248, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6249-6256, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 423. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6182-6183, 6189, 6191, 6208, 6210, 6211, or 6215. In some embodiments, the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of SEQ ID NO: 6251.
[0062] In some embodiments of any of the methods of editing a specific locus in the above cells, the cells are peripheral blood mononuclear cells, T cells, NK cells, hematopoietic stem cells (HSCT), or B cells, or any combination thereof.
[0063] In some embodiments, the present disclosure provides an engineered guide ribonucleic acid polynucleotide comprising: (a) a DNA targeting segment comprising a nucleotide sequence complementary to a target sequence in a target DNA molecule; and (b) a protein binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex, wherein the two complementary stretches of nucleotides are covalently linked to each other using intervening nucleotides, and the engineered guide ribonucleic acid polynucleotide is configured to form a complex with a Class 2 type IIC Cas endonuclease and target the complex to the target sequence of the target DNA molecule, and the DNA targeting segment comprises a sequence having at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 consecutive nucleotides of any one of SEQ ID NOs: 5950-5965, 5966-6025, 6026-6121, 6122-6152, 6153-6181, or 6182-6256. In some embodiments, the protein binding segment comprises a sequence having at least 85% identity to any one of SEQ ID NO: 5466 or SEQ ID NO: 6304.
[0064] In some embodiments, the present disclosure provides a system for generating engineered immune cells, comprising: (a) an RNA-guided endonuclease; (b) an engineered guide ribonucleic acid polynucleotide as recited in claim 97, configured to bind to the RNA-guided endonuclease; and (c) a single-stranded or double-stranded DNA repair template comprising a first homology arm and a second homology arm adjacent to a sequence encoding a chimeric antigen receptor (CAR). In some embodiments, the cells are peripheral blood mononuclear cells, T cells, NK cells, hematopoietic stem cells (HSCT), or B cells, or any combination thereof. In some embodiments, the RNA-guided endonuclease is a class II type IICas endonuclease. In some embodiments, the RNA-guided endonuclease comprises an RuvCIII domain having a sequence with at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244. In some embodiments, the RNA-guided endonuclease further comprises an HNH domain. In some embodiments, the RNA-guided endonuclease comprises a sequence having at least 75%, at least 80%, at least 82%, at least 84%, at least 86%, at least 88%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identity to SEQ ID NO: 421 or SEQ ID NO: 423.
[0065] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which illustrates only exemplary embodiments of the present disclosure. As will be understood, the present disclosure is capable of other and different embodiments, and some of the details thereof are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive. Incorporation by reference
[0066] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference into this specification to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the invention will be obtained from the following detailed description that illustrates exemplary embodiments in which the principles of the invention are utilized, and from the appended drawings (also referred to herein as "Figures" and "FIGs").
[0068]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 9C
Figure 9D
Figure 9E
Figure 9F
Figure 9G
Figure 9H
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Figure 63
Figure 64
Figure 65
Figure 66
Figure 67
Figure 68
Figure 69
Figure 70
Figure 71
Figure 72
Figure 73
[0069] Brief Description of the Sequence Listing The sequence listing submitted with this specification provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. The following is an exemplary description of the sequences therein.
[0070] MG1
[0071] SEQ ID NOs: 1 to 319 represent the full-length peptide sequences of the MG1 nuclease.
[0072] SEQ ID NOs: 1827 to 2140 represent the peptide sequences of the RuvC_III domain of the above-mentioned MG1 nuclease.
[0073] SEQ ID NOs: 3638 to 3955 represent the peptides of the HNH domain of the above-mentioned MG1 nuclease.
[0074] SEQ ID NOs: 5476 to 5479 represent the nucleotide sequences of MG1 tracrRNA derived from the same locus as the above-mentioned MG1 nuclease (e.g., the same locus as SEQ ID NOs: 1 to 4 respectively).
[0075] SEQ ID NOs: 5461 to 5464 show the nucleotide sequences of sgRNAs engineered to function with an MG1 nuclease (e.g., SEQ ID NOs: 1 to 4, respectively), where Ns denote the nucleotides of the targeting sequence.
[0076] SEQ ID NOs: 5572 to 5575 show the nucleotide sequences of the E. coli codon-optimized coding sequences of the MG1 family enzymes (SEQ ID NOs: 1 to 4).
[0077] SEQ ID NOs: 5588 to 5589 show the nucleotide sequences of the human codon-optimized coding sequences for the MG1 family enzymes (SEQ ID NOs: 1 and 3).
[0078] SEQ ID NOs: 5616 to 5632 show peptide motifs characteristic of the MG1 family enzymes.
[0079] MG2
[0080] SEQ ID NOs: 320 to 420 show the full-length peptide sequence of the MG2 nuclease.
[0081] SEQ ID NOs: 2141 to 2241 show the peptide sequence of the RuvC_III domain of the above-mentioned MG2 nuclease.
[0082] SEQ ID NOs: 3955 to 4055 show the peptide of the HNH domain of the above-mentioned MG2 nuclease.
[0083] SEQ ID NOs: 5490 to 5494 show the nucleotide sequences of MG2 tracrRNAs derived from the same loci as the above-mentioned MG2 nuclease (e.g., the same loci as SEQ ID NOs: 320, 321, 323, 325, and 326, respectively).
[0084] SEQ ID NO: 5465 shows the nucleotide sequence of an sgRNA engineered to function with an MG2 nuclease (e.g., SEQ ID NO: 321 above).
[0085] SEQ ID NOs: 5572 to 5575 show the nucleotide sequences of the E. coli codon-optimized coding sequences of the MG2 family enzymes.
[0086] SEQ ID NOs: 5631 to 5638 show the peptide sequences characteristic of the MG2 family enzymes.
[0087] MG3
[0088] SEQ ID NOs: 421 to 431 show the full-length peptide sequence of the MG3 nuclease.
[0089] SEQ ID NOs: 2242 to 2252 show the peptide sequence of the RuvC_III domain of the above-mentioned MG3 nuclease.
[0090] SEQ ID NOs: 4056 to 4066 show the peptide of the HNH domain of the above-mentioned MG3 nuclease.
[0091] SEQ ID NOs: 5495 to 5502 show the nucleotide sequences of the MG3 tracrRNA derived from the same locus as the above-mentioned MG3 nuclease (for example, the same locus as SEQ ID NOs: 421 to 428, respectively).
[0092] SEQ ID NOs: 5466 to 5467 show the nucleotide sequences of the sgRNA engineered to function with the MG3 nuclease (for example, SEQ ID NOs: 421 to 423).
[0093] SEQ ID NOs: 5578 to 5580 show the nucleotide sequences of the E. coli codon-optimized coding sequences of the MG3 family enzymes.
[0094] SEQ ID NOs: 5639 to 5648 show the peptide sequences characteristic of the MG3 family enzymes.
[0095] MG4
[0096] SEQ ID NOs: 432 to 660 show the full-length peptide sequence of the MG4 nuclease.
[0097] Sequence numbers 2253 to 2481 show the peptide sequences of the RuvC_III domain of the above-mentioned MG4 nuclease.
[0098] Sequence numbers 4067 to 4295 show the peptides of the HNH domain of the above-mentioned MG4 nuclease.
[0099] Sequence number 5503 shows the nucleotide sequence of the MG4 tracrRNA derived from the same locus as the above-mentioned MG4 nuclease.
[0100] Sequence number 5468 shows the nucleotide sequence of the sgRNA engineered to function with the MG4 nuclease.
[0101] Sequence number 5649 shows the peptide sequence characteristic of the MG4 family of enzymes.
[0102] MG6
[0103] Sequence numbers 661 to 668 show the full-length peptide sequence of the MG6 nuclease.
[0104] Sequence numbers 2482 to 2489 show the peptide sequences of the RuvC_III domain of the above-mentioned MG6 nuclease.
[0105] Sequence numbers 4296 to 4303 show the peptides of the HNH domain of the above-mentioned MG3 nuclease.
[0106] MG7
[0107] Sequence numbers 669 to 677 show the full-length peptide sequence of the MG7 nuclease.
[0108] Sequence numbers 2490 to 2498 show the peptide sequences of the RuvC_III domain of the above-mentioned MG7 nuclease.
[0109] Sequence numbers 4304 to 4312 show the peptides of the HNH domain of the above-mentioned MG3 nuclease.
[0110] SEQ ID NO: 5504 shows the nucleotide sequence of the MG7 tracrRNA derived from the same locus as the above-mentioned MG7 nuclease.
[0111] MG14
[0112] SEQ ID NOs: 678 to 929 show the full-length peptide sequence of the MG14 nuclease.
[0113] SEQ ID NOs: 2499 to 2750 show the peptide sequence of the RuvC_III domain of the above-mentioned MG14 nuclease.
[0114] SEQ ID NOs: 4313 to 4564 show the peptide of the HNH domain of the above-mentioned MG14 nuclease.
[0115] SEQ ID NO: 5505 shows the nucleotide sequence of the MG14 tracrRNA derived from the same locus as the above-mentioned MG14 nuclease.
[0116] SEQ ID NO: 5581 shows the nucleotide sequence of the E. coli codon-optimized coding sequence of the MG14 family enzyme.
[0117] SEQ ID NOs: 5650 to 5667 show the peptide sequence characteristic of the MG14 family enzyme.
[0118] MG15
[0119] SEQ ID NOs: 930 to 1092 show the full-length peptide sequence of the MG15 nuclease.
[0120] SEQ ID NOs: 2751 to 2913 show the peptide sequence of the RuvC_III domain of the above-mentioned MG15 nuclease.
[0121] SEQ ID NOs: 4565 to 4727 show the peptide of the HNH domain of the above-mentioned MG15 nuclease.
[0122] SEQ ID NO: 5506 shows the nucleotide sequence of the MG15 tracrRNA derived from the same locus as the above-mentioned MG15 nuclease.
[0123] SEQ ID NO: 5470 shows the nucleotide sequence of the sgRNA engineered to function with the MG15 nuclease.
[0124] SEQ ID NO: 5582 shows the nucleotide sequence of the E. coli codon-optimized coding sequence of the MG15 family enzyme.
[0125] SEQ ID NOs: 5668 to 5675 show the peptide sequences characteristic of the MG15 family enzyme.
[0126] MG16
[0127] SEQ ID NOs: 1093 to 1353 show the full-length peptide sequence of the MG16 nuclease.
[0128] SEQ ID NOs: 2914 to 3174 show the peptide sequence of the RuvC_III domain of the above-mentioned MG16 nuclease.
[0129] SEQ ID NOs: 4728 to 4988 show the peptide of the HNH domain of the above-mentioned MG16 nuclease.
[0130] SEQ ID NO: 5507 shows the nucleotide sequence of the MG16 tracrRNA derived from the same locus as the above-mentioned MG3 nuclease.
[0131] SEQ ID NO: 5471 shows the nucleotide sequence of the sgRNA engineered to function with the MG16 nuclease.
[0132] SEQ ID NO: 5583 shows the nucleotide sequence of the E. coli codon-optimized coding sequence of the MG16 family enzyme.
[0133] SEQ ID NOs: 5676 to 5678 show the peptide sequences characteristic of the MG16 family enzyme.
[0134] MG18
[0135] SEQ ID NOs: 1354 to 1511 show the full-length peptide sequence of the MG18 nuclease.
[0136] SEQ ID NOs: 3175 to 3330 show the peptide sequence of the RuvC_III domain of the above-mentioned MG18 nuclease.
[0137] SEQ ID NOs: 4989 to 5146 show the peptide of the HNH domain of the above-mentioned MG18 nuclease.
[0138] SEQ ID NO: 5508 shows the nucleotide sequence of the MG18 tracrRNA derived from the same locus as the above-mentioned MG18 nuclease.
[0139] SEQ ID NO: 5472 shows the nucleotide sequence of the sgRNA engineered to function with the MG18 nuclease.
[0140] SEQ ID NO: 5584 shows the nucleotide sequence of the E. coli codon-optimized coding sequence of the MG18 family enzyme.
[0141] SEQ ID NOs: 5679 to 5686 show the peptide sequence characteristic of the MG18 family enzyme.
[0142] MG21
[0143] SEQ ID NOs: 1512 to 1655 show the full-length peptide sequence of the MG21 nuclease.
[0144] SEQ ID NOs: 3331 to 3474 show the peptide sequence of the RuvC_III domain of the above-mentioned MG21 nuclease.
[0145] SEQ ID NOs: 5147 to 5290 show the peptide of the HNH domain of the above-mentioned MG21 nuclease.
[0146] SEQ ID NO: 5509 shows the nucleotide sequence of the MG21 tracrRNA derived from the same locus as the above-mentioned MG21 nuclease.
[0147] SEQ ID NO: 5473 shows the nucleotide sequence of the sgRNA engineered to function with the MG21 nuclease.
[0148] SEQ ID NO: 5585 shows the nucleotide sequence of the E. coli codon-optimized coding sequence of the MG21 family enzyme.
[0149] SEQ ID NOs: 5687-5692 and SEQ ID NOs: 5674-5675 show the peptide sequences characteristic of the MG21 family enzyme.
[0150] MG22
[0151] SEQ ID NOs: 1656-1755 show the full-length peptide sequence of the MG22 nuclease.
[0152] SEQ ID NOs: 3475-3568 show the peptide sequence of the RuvC_III domain of the above-mentioned MG22 nuclease.
[0153] SEQ ID NOs: 5291-5389 show the peptide of the HNH domain of the above-mentioned MG22 nuclease.
[0154] SEQ ID NO: 5510 shows the nucleotide sequence of the MG22 tracrRNA derived from the same locus as the above-mentioned MG22 nuclease.
[0155] SEQ ID NO: 5474 shows the nucleotide sequence of the sgRNA engineered to function with the MG22 nuclease.
[0156] SEQ ID NO: 5586 shows the nucleotide sequence of the E. coli codon-optimized coding sequence of the MG22 family enzyme.
[0157] Accession numbers 5694 to 5699 represent peptide sequences characteristic of the MG22 family of enzymes.
[0158] MG23
[0159] Accession numbers 1756 to 1826 represent the full-length peptide sequence of the MG23 nuclease.
[0160] Accession numbers 3569 to 3637 represent the peptide sequence of the RuvC_III domain of the above-mentioned MG23 nuclease.
[0161] Accession numbers 5390 to 5460 represent the peptide of the HNH domain of the above-mentioned MG23 nuclease.
[0162] Accession number 5511 represents the nucleotide sequence of the MG23 tracrRNA derived from the same locus as the above-mentioned MG23 nuclease.
[0163] Accession number 5475 represents the nucleotide sequence of the sgRNA engineered to function with the MG23 nuclease.
[0164] Accession number 5587 represents the nucleotide sequence of the E. coli codon-optimized coding sequence of the MG23 family of enzymes.
[0165] Accession numbers 5700 to 5717 represent peptide sequences characteristic of the MG23 family of enzymes.
[0166] MG40
[0167] Accession numbers 5718 to 5750 represent the full-length peptide sequence of the MG40 nuclease.
[0168] Accession numbers 5847 to 5852 represent the protospacer adjacent motif related to the MG40 nuclease.
[0169] SEQ ID NOs: 5862 to 5873 show the nucleotide sequences of sgRNAs engineered to function with the MG40 nuclease.
[0170] MG47
[0171] SEQ ID NOs: 5751 to 5768 show the full-length peptide sequences of the MG47 nuclease.
[0172] SEQ ID NOs: 5853 to 5854 show the protospacer adjacent motifs related to the MG47 nuclease.
[0173] SEQ ID NOs: 5878 to 5881 show the nucleotide sequences of sgRNAs engineered to function with the MG47 nuclease.
[0174] MG48
[0175] SEQ ID NOs: 5769 to 5804 show the full-length peptide sequences of the MG48 nuclease.
[0176] SEQ ID NOs: 5855 to 5856 show the protospacer adjacent motifs related to the MG48 nuclease.
[0177] SEQ ID NOs: 5886, 5890, and 5893 show the nucleotide sequences of the MG48 tracrRNA derived from the same locus as the above-mentioned MG48 nuclease.
[0178] SEQ ID NOs: 5887, 5891, and 5894 show the CRISPR repeats related to the MG48 nuclease described herein.
[0179] SEQ ID NOs: 5888 to 5889, 5892, and 5895 to 5896 show the putative sgRNAs designed to function with the MG48 nuclease.
[0180] MG49
[0181] Accession numbers 5805 to 5823 represent the full-length peptide sequences of MG49 nuclease.
[0182] Accession numbers 5857 to 5858 represent the protospacer adjacent motifs related to MG49 nuclease.
[0183] Accession numbers 5862 to 5873 represent the nucleotide sequences of sgRNAs engineered to function with MG40 nuclease.
[0184] Accession numbers 5876 to 5877 represent the nucleotide sequences of sgRNAs engineered to function with MG49 nuclease.
[0185] MG50
[0186] Accession numbers 5824 to 5826 represent the full-length peptide sequences of MG50 nuclease.
[0187] Accession number 5859 represents the protospacer adjacent motif related to MG50 nuclease.
[0188] Accession numbers 5884 to 5885 represent the nucleotide sequences of sgRNAs engineered to function with MG50 nuclease.
[0189] MG51
[0190] Accession numbers 5827 to 5830 represent the full-length peptide sequences of MG51 nuclease.
[0191] Accession number 5860 represents the protospacer adjacent motif related to MG51 nuclease.
[0192] Accession numbers 5882 to 5883 represent the nucleotide sequences of sgRNAs engineered to function with MG51 nuclease.
[0193] MG52
[0194] Accession numbers 5831 to 5846 represent the full-length peptide sequence of MG52 nuclease.
[0195] Accession number 5861 represents the protospacer adjacent motif related to MG52 nuclease.
[0196] Accession numbers 5874 to 5875 represent the nucleotide sequences of sgRNAs engineered to function with MG42 nuclease.
Modes for Carrying Out the Invention
[0197] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, modifications, and substitutions may occur to those skilled in the art without departing from the present invention. It should be understood that various alternative forms to the embodiments of the present invention described herein may be employed.
[0198] The practice of several methods disclosed herein employs, unless otherwise indicated, the techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. For example, Sambrook and Green, "Molecular Cloning: A Laboratory Manual," 4th Edition (2012); "the series Current Protocols in Molecular Biology" (edited by F.M. Ausubel et al.); "the series Methods In Enzymology" (Academic Press, Inc.), "PCR 2: A Practical Approach" (edited by M.J. MacPherson, B.D. Hames and G.R. Taylor (1995)), edited by Harlow and Lane (1988), "Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications," 6th Edition (edited by R.I. Freshney (2010) (which is hereby incorporated by reference in its entirety herein).
[0199] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, the terms "including," "includes," "having," "has," "with," or variations thereof, as used in any of the detailed description and / or claims, are intended to be inclusive in a manner similar to the term "comprising."
[0200] The terms "about" or "approximately" mean within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or two standard deviations, according to the conventions of the relevant art. Alternatively, "about" can mean a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0201] As used herein, "cell" generally refers to a biological cell. A cell can be the basic structural, functional, and / or biological unit of an organism. A cell can be derived from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of single-celled eukaryotes, protozoal cells, cells derived from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, clubmosses, spike mosses, quillworts, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, and Sargassum patens C. Agardh, etc.), seaweeds (e.g., kelp), fungal cells (e.g., yeast cells, cells derived from mushrooms), animal cells, cells derived from invertebrates (e.g., Drosophila, cnidarians, echinoderms, nematodes, etc.), cells derived from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), and cells derived from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). A cell may not be derived from a natural organism (e.g., a cell may be synthetically produced and may be referred to as an artificial cell).
[0202] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. Nucleotides can include synthetic nucleotides. Nucleotides can include synthetic nucleotide analogs. Nucleotides can be the monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term "nucleotide" can include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives can include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. The term "nucleotide" as used herein can refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Exemplary examples of dideoxyribonucleoside triphosphates can include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled, such as by using a moiety that includes an optically detectable moiety (e.g., a fluorophore). Labeling can also be performed using quantum dots. Detectable labels can include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels.Examples of fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2’7’-dimethoxy-4’5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N’,N’-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4’-dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2’-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer, Foster City, CA; FluoroLink deoxyribonucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, available from Amersham, Arlington Heights, IL; fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, available from Boehringer Mannheim, Indianapolis, IN; and BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, available from Molecular Probes, Eugene, OR. Nucleotides can also be labeled or marked by chemical modification. A chemically modified single nucleotide can be biotin-dNTP.Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., biotin-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0203] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length, either in single-stranded, double-stranded, or multi-stranded form, of either deoxyribonucleotides or ribonucleotides, or analogs thereof. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can be present in a cell-free environment. A polynucleotide can be a gene or a fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three-dimensional structure and can perform any function. A polynucleotide can contain one or more analogs (e.g., modified backbone, sugar, or nucleobase). When present, modifications to the nucleotide structure can be imparted before or after polymerization of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophore (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotide, biotin-binding nucleotide, fluorescent base analog, CpG island, methyl-7-guanosine, methylated nucleotide, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. Non-limiting examples of polynucleotides include the coding or non-coding regions of a gene or gene fragment, locus (loci) defined from linkage analysis, exon, intron, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozyme, cDNA, recombinant polynucleotide, branched polynucleotide, plasmid, vector, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotide including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probe, and primer. The nucleotide sequence can be interrupted by non-nucleotide constituents.
[0204] The terms "transfection" or "transfected" generally refer to the introduction of nucleic acids into cells by non-viral or virus-based methods. The nucleic acid molecule can be a gene sequence encoding a complete protein or a functional portion thereof. See, for example, Sambrook et al., "Molecular Cloning: A Laboratory Manual", pages 18.1 - 18.88, 1989.
[0205] The terms "peptide", "polypeptide", and "protein" are used interchangeably herein generally to refer to a polymer of at least two amino acid residues linked by peptide bond(s). The term does not imply a polymer of a specific length, nor is it intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical synthesis, or enzymatic synthesis, or is naturally occurring. The term applies to naturally occurring amino acid polymers and amino acid polymers containing at least one modified amino acid. In some cases, the polymer can be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary and / or tertiary structures (e.g., domains). The term also encompasses amino acid polymers modified by any other operation, such as disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. As used herein, the terms "amino acid(s)" generally refer to natural and non-natural amino acids, including modified amino acids and amino acid analogs, but are not limited thereto. Modified amino acids may include natural and non-natural amino acids that are chemically modified to contain a group or chemical moiety not naturally present on the amino acid. Amino acid analogs may refer to amino acid derivatives. The term "amino acid" includes both D-amino acids and L-amino acids.
[0206] As used herein, "non-natural" can generally refer to a nucleic acid or polypeptide sequence not found in natural nucleic acids or proteins. Non-natural can refer to an affinity tag. Non-natural can refer to a fusion. Non-natural can refer to a naturally occurring nucleic acid or polypeptide sequence that includes mutations, insertions, and / or deletions. A non-natural sequence can also present and / or encode an activity (e.g., enzymatic activity, methyltransferase activity, acetyltransferase activity, kinase activity, ubiquitination activity, etc.) conferred by the nucleic acid and / or polypeptide sequence to which the non-natural sequence is fused. A non-natural nucleic acid or polypeptide sequence can be linked by genetic engineering to a naturally occurring nucleic acid or polypeptide sequence (or variant thereof) to generate a chimeric nucleic acid and / or polypeptide sequence encoding a chimeric nucleic acid and / or polypeptide.
[0207] As used herein, the term "promoter" generally refers to a regulatory DNA region that can be adjacent to or overlapping with a nucleotide or region of nucleotides that controls the transcription or expression of a gene and at which RNA transcription is initiated. A promoter can contain a specific DNA sequence that binds to a protein factor, often called a transcription factor, that facilitates the binding of RNA polymerase to DNA to effect gene transcription. A "basal promoter," also referred to as a "core promoter," can generally refer to a promoter that contains all of the essential elements necessary to facilitate the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters typically contain, but not necessarily, a TATA box and / or a CAAT box.
[0208] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (such as into mRNA or other RNA transcripts), and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may collectively be referred to as a "gene product". When the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in eukaryotic cells.
[0209] As used herein, "operably linked", "operable linkage", "operatively linked", or their grammatical equivalents generally refer to the juxtaposition of genetic elements, such as a promoter, enhancer, polyadenylation sequence, etc., in a relationship that enables these elements to function in their expected manner. As an example, a regulatory element that may include a promoter sequence and / or enhancer sequence is operably linked to a coding region if the regulatory element aids in the initiation of transcription of the coding sequence. Intervening residues may be present between the regulatory element and the coding region as long as this functional relationship is maintained.
[0210] As used herein, a "vector" generally refers to a macromolecule or aggregate of macromolecules that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide into a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. Vectors generally contain genetic elements, such as regulatory elements, that are operably linked to a gene to facilitate expression of the gene in a target.
[0211] As used herein, the terms "expression cassette" and "nucleic acid cassette" are used interchangeably to generally refer to a combination of nucleic acid sequences or elements that are co-expressed or operably linked for expression. In some cases, an expression cassette refers to a combination of regulatory elements and one or more genes to which they are operably linked for expression.
[0212] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (either functional or structural) that is substantially similar to the biological activity of the full-length DNA or protein sequence. The biological activity of a DNA sequence can be its ability to affect expression in a manner known to be attributable to the full-length sequence.
[0213] As used herein, an "engineered" object generally indicates that the object has been altered by human intervention. By way of non-limiting example, a nucleic acid can be modified by changing its sequence to a sequence that does not naturally exist. A nucleic acid can be modified by ligating it to a nucleic acid to which it does not naturally associate such that the ligation product has a function that does not exist in the original nucleic acid. Engineered nucleic acids can be synthesized in vitro using sequences that do not naturally exist. A protein can be modified by changing its amino acid sequence to a sequence that does not naturally exist. Engineered proteins can acquire new functions or properties. An "engineered" system includes at least one engineered component.
[0214] As used herein, the terms "synthetic" and "artificial" are used interchangeably to refer to a protein or domain thereof that has a low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to a naturally occurring human protein. For example, the VPR domain and the VP64 domain are synthetic transactivation domains.
[0215] As used herein, the term "tracrRNA" or "tracr sequence" generally refers to a nucleic acid having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or sequence similarity to an exemplary wild-type tracrRNA sequence (e.g., tracrRNA derived from Streptococcus pyogenes, Staphylococcus aureus, etc., or SEQ ID NOs: 5476-5511). A tracrRNA can refer to a nucleic acid having up to about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to an exemplary wild-type tracrRNA sequence (e.g., tracrRNA derived from Streptococcus pyogenes, Staphylococcus aureus, etc.). A tracrRNA can refer to a modified form of a tracrRNA that can include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeras. A tracrRNA can refer to a nucleic acid that is at least about 60% identical over a stretch of at least 6 contiguous nucleotides to an exemplary wild-type tracrRNA sequence (e.g., tracrRNA derived from Streptococcus pyogenes, Staphylococcus aureus, etc.). For example, a tracrRNA sequence can be at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical over a stretch of at least 6 contiguous nucleotides to an exemplary wild-type tracrRNA sequence (e.g., tracrRNA derived from Streptococcus pyogenes, Staphylococcus aureus, etc.). Type II tracrRNA sequences can be predicted on a genomic sequence by identifying regions having complementarity to a portion of the repeat sequence in an adjacent CRISPR array.
[0216] As used herein, "guide nucleic acid" can generally refer to a nucleic acid that can hybridize to another nucleic acid. The guide nucleic acid can be RNA. The guide nucleic acid can be DNA. The guide nucleic acid can be programmed to specifically bind to a site in the nucleic acid sequence. The nucleic acid to be targeted, or target nucleic acid, can contain nucleotides. The guide nucleic acid can contain nucleotides. A portion of the target nucleic acid can be complementary to a portion of the guide nucleic acid. The strand of the double-stranded target polynucleotide that is complementary to the guide nucleic acid and hybridizes to the guide nucleic acid can be referred to as the complementary strand. The strand of the double-stranded target polynucleotide that is complementary to the complementary strand and thus may not be complementary to the guide nucleic acid can be called the non-complementary strand. The guide nucleic acid can include a polynucleotide strand and can be referred to as a "single guide nucleic acid". The guide nucleic acid can include two polynucleotide strands and can be referred to as a "double guide nucleic acid". Unless otherwise specified, the term "guide nucleic acid" can be an inclusive one that refers to both single guide nucleic acids and double guide nucleic acids. The guide nucleic acid can include a segment that can be referred to as a "nucleic acid targeting segment" or "nucleic acid targeting sequence". The nucleic acid targeting segment can include a sub-segment that can be referred to as a "protein binding segment" or "protein binding sequence" or "Cas protein binding segment".
[0217] The terms "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refer to sequences of two or more (e.g., in a pair alignment, e.g., in a multiple sequence alignment) that are the same or have a specified percentage of amino acid residues or nucleotides that are the same when compared over a local or global comparison window and aligned with maximum correspondence, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using parameters of a word length of 3, an expectation value I of 10, and a BLOSUM62 scoring matrix setting gap cost of 11 for existence and 1 for extension, and a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues. BLASTP using parameters of a word length (W) of 2, an expectation value (E) of 1,000,000, and a PAM30 scoring matrix setting gap has a cost of 9 to open a gap and 1 to extend a gap for sequences less than 30 residues (these are the default parameters of BLAST in a set of BLASTs available at https: / / blast.ncbi.nlm.nih.gov); the Smith-Waterman homology search algorithm with parameters of 2 for match, -1 for mismatch, and -1 for gap; MUSCLE with default parameters; MAFFT with parameters of 2 for retree and 1000; Novafold with default parameters; CLUSTALW with parameters of HMMER hmmalign with default parameters.
[0218] The present disclosure includes any variant of the enzymes described herein having one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of the polypeptide without disrupting the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids having similar hydrophobicity, polarity, and R-chain length with each other. Additionally or alternatively, by comparing the aligned sequences of homologous proteins from different species, conservative substitutions can be identified by placing amino acid residues that are mutated between species (e.g., non-conserved residues) without altering the basic function of the encoded protein. Such conservatively substituted variants can have at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of the endonuclease protein sequences described herein (e.g., the MG1, MG2, MG3, MG4, MG6, MG7, MG14, MG15, MG16, MG18, MG21, MG22, or MG23 family endonucleases described herein). In some embodiments, such conservative substitution variants are functional variants. Such functional variants can include sequences having substitutions such that the activity of important active site residues of the endonuclease is not disrupted. In some embodiments, any functional variant of the proteins described herein lacks substitution of at least one of the conserved or functional residues referred to in FIGS. 6, 7, 8, 9A, 9B, 9C, 9D, 9E, 9F, 9G, or 9H.In some embodiments, any functional variant of the proteins described herein lacks all substitutions of conserved or functional residues referred to in FIGS. 6, 7, 8, 9A, 9B, 9C, 9D, 9E, 9F, 9G, or 9H.
[0219] Tables of conservative substitutions that provide functionally similar amino acids are available from a variety of references (e.g., Creighton, "Proteins: Structures and Molecular Properties" (W H Freeman & Co.; 2nd Edition (December 1993)). The following eight groups each contain amino acids that are conservative substitutions for one another. (1) Alanine (A), Glycine (G); (2) Aspartic acid (D), Glutamic acid (E); (3) Asparagine (N), Glutamine (Q); (4) Arginine (R), Lysine (K); (5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); (6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); (7) Serine (S), Threonine (T); and (8) Cysteine (C), Methionine (M)
[0220] As used herein, the term "RuvC_III domain" generally refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three discontinuous segments, RuvC_I, RuvC_II, and RuvC_III). The RuvC domain or segments thereof can generally be identified by alignment to known domain sequences, by structural alignment to proteins with annotated domains, or by comparison to a Hidden Markov Model (HMM) constructed based on known domain sequences (e.g., for RuvC_III, Pfam HMM PF18541).
[0221] As used herein, the term "HNH domain" generally refers to an endonuclease domain having characteristic histidine and asparagine residues. The HNH domain can generally be identified by alignment to known domain sequences, by structural alignment to proteins having annotated domains, or by comparison to a Hidden Markov Model (HMM) constructed based on known domain sequences (e.g., Pfam HMM PF01844 for domain HNH).
[0222] Overview
[0223] The discovery of novel Cas enzymes with unique functionality and structure may further disrupt deoxyribonucleic acid (DNA) editing technologies and offer the potential to improve speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeat (CRISPR) systems in microorganisms and the sheer diversity of microbial species, relatively few CRISPR / Cas enzymes have been functionally characterized in the literature. This is because a vast number of microbial species may not be easily cultured under laboratory conditions. Metagenomic sequencing from natural environmental niches representing numerous microbial species may dramatically increase the number of known new CRISPR / Cas systems and offer the potential to accelerate the discovery of new oligonucleotide editing functions. A recent example of the utility of such an approach is demonstrated by the 2016 discovery of the CasX / CasY CRISPR systems from metagenomic analysis of natural microbial communities.
[0224] The CRISPR / Cas system is an RNA-guided nuclease complex that has been described as functioning as an adaptive immune system in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) operons or loci, which generally consist of two parts: (i) an array of short repeat sequences (30 - 40 bp) separated by equally short spacer sequences that encode RNA-based targeting elements, and (ii) an ORF encoding Cas that encodes a nuclease polypeptide induced by the RNA-based targeting element together with accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6 - 8 nucleic acids of the target (the target seed) and the crRNA guide, and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (the PAM is usually a sequence not commonly represented within the host genome). Depending on the exact function and composition of the system, CRISPR-Cas systems are generally organized into two classes, five types, and 16 subtypes based on shared functional features and evolutionary similarities.
[0225] Class I CRISPR-Cas systems have large multi-subunit effector complexes and include type I, type III, and type IV.
[0226] Type I CRISPR-Cas systems are thought to have a moderate complexity with respect to their components. In Type I CRISPR-Cas systems, arrays of RNA targeting elements are transcribed as long precursor crRNAs (pre-crRNAs) that are processed by repeat elements to release short mature crRNAs that direct the nuclease complex to nucleic acid targets when the nucleic acid target is followed by a suitable short consensus sequence called the protospacer adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called Cascade, which also includes the nuclease (Cas3) protein component of the crRNA-guided nuclease complex. Cas I nucleases function primarily as DNA nucleases.
[0227] Type III CRISPR systems may be characterized by the presence of a central nuclease known as Cas10, along with repeat-associated mysterious proteins (RAMPs) that include Csm or Cmr protein subunits. Similar to Type I systems, mature crRNAs are processed from pre-crRNAs using a Cas6-like enzyme. Unlike Type I and II systems, Type III systems appear to target and cleave DNA-RNA duplexes (such as the DNA strand being used as a template by RNA polymerase).
[0228] Type IV CRISPR-Cas systems carry an effector complex consisting of a highly reduced large subunit nuclease (csf1), two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) groups, and, optionally, a gene for a predicted small subunit. Such systems are commonly found on endogenous plasmids.
[0229] Class II CRISPR-Cas systems generally have a single polypeptide multi-domain nuclease effector and include Type II, Type V, and Type VI.
[0230] Type II CRISPR-Cas systems are thought to be the simplest in terms of their components. In Type II CRISPR-Cas systems, processing of the CRISPR array into mature crRNAs does not require the presence of a special endonuclease subunit, but rather requires the presence of a small trans-encoded crRNA (tracrRNA) with regions complementary to the array repeat sequences. This tracrRNA interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequences to form a precursor dsRNA structure, which is cleaved by endogenous RNAse III to generate a mature effector enzyme loaded with both tracrRNA and crRNA. Cas II nuclease is known as a DNA nuclease. Type II effectors generally adopt a structure consisting of an RuvC-like endonuclease domain that incorporates an unrelated HNH nuclease domain inserted within the fold of the RuvC-like nuclease domain. The RuvC-like domain is responsible for cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is responsible for cleavage of the displaced DNA strand.
[0231] Type V CRISPR-Cas systems feature nuclease effector (e.g., Cas12) structures similar to those of type II effectors that contain an RuvC-like domain. Similar to type II, most (but not all) type V CRISPR systems use tracrRNA to process pre-crRNA into mature crRNA. However, unlike type II systems that require RNase III to cleave pre-crRNA into multiple crRNAs, type V systems can cleave pre-crRNA using the effector nuclease itself. Similar to type II CRISPR-Cas systems, type V CRISPR-Cas systems are also known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to have robust single-stranded non-specific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of a double-stranded target sequence.
[0232] Type VI CRIPSR-Cas systems have an RNA-guided RNA endonuclease. Instead of an RuvC-like domain, the single polypeptide effector (e.g., Cas13) of type VI systems contains two HEPN ribonuclease domains. Unlike either type II or type V systems, type VI systems also do not appear to require tracrRNA to process pre-crRNA into crRNA. However, similar to type V systems, some type VI systems (e.g., C2C2) appear to have robust single-stranded non-specific nuclease (ribonuclease) activity that is activated by the first crRNA-directed cleavage of a target RNA.
[0233] Due to their simpler structures, class II CRISPR-Cas have been the most widely adopted for engineering and development as designer nuclease / genome editing applications.
[0234] One of the initial uses of such systems for in vitro use can be found in Jinek et al. ("Science.", August 17, 2012; Vol. 337 (No. 6096): pp. 816 - 21, which is hereby incorporated by reference in its entirety). Jinek's study first described (i) recombinantly expressed and purified full - length Cas9 isolated from Streptococcus pyogenes SF370 (e.g., a class II, type II Cas enzyme), (ii) a purified mature ~42nt carrying approximately 20nt of 5' complementary to the target DNA sequence where cleavage is desired, followed by a 3' tracr - binding sequence (the entire crRNA is in vitro transcribed from a synthetic DNA template carrying a T7 promoter sequence), (iii) a purified tracrRNA in vitro transcribed from a synthetic DNA template carrying a T7 promoter sequence, and (iv) a system containing Mg2+. Jinek later described an improved and engineered system where the crRNA of (ii) is linked to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) that can direct Cas9 to the target alone (compare the upper and lower panels of Figure 2).
[0235] Mali et al. (''Science.'', February 15, 2013; Vol. 339 (No. 6121): pp. 823 - 826, which is hereby incorporated by reference in its entirety) later adapted this system for use in mammalian cells by providing a DNA vector encoding (i) an ORF encoding codon - optimized Cas9 (e.g., a class II type II Cas enzyme) under a suitable mammalian promoter having a C - terminal nuclear localization sequence (e.g., SV40NLS) and a suitable polyadenylation signal (e.g., TKpA signal), and (ii) an ORF encoding an sgRNA (a 20nt complementary targeting nucleic acid sequence, a linker, and a tracrRNA sequence linked to a 5' sequence starting with G and followed by a 3' tracr - binding sequence) under a suitable polymerase III promoter (e.g., the U6 promoter).
[0236] MG enzyme
[0237] In one aspect, the present disclosure provides engineered nuclease systems discovered by metagenomic sequencing. Optionally, metagenomic sequencing is performed on a sample. Optionally, the sample can be collected by various environments. Such environments can be the human microbiota, animal microbiota, high-temperature environments, low-temperature environments. Such environments can include precipitates. Examples of such types of environments of the engineered nuclease systems described herein can be found in FIG. 45.
[0238] MG1 enzyme
[0239] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 1827 - 2140. Optionally, the endonuclease may comprise an RuvC_III domain having, to any one of SEQ ID NOs: 1827 - 2140, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity. Optionally, the endonuclease may comprise an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 1827 - 2140. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 1827 - 1831. Optionally, the endonuclease may comprise an RuvC_III domain having, to any one of SEQ ID NOs: 1827 - 1831, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 1827 to 1831. In some cases, the endonuclease may contain an RuvC_III domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NO: 1827. In some cases, the endonuclease may contain an RuvC_III domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NO: 1828. In some cases, the endonuclease may contain an RuvC_III domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NO: 1829. In some cases, the endonuclease may contain an RuvC_III domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NO: 1830.In some cases, the endonuclease may comprise an RuvC_III domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to SEQ ID NO: 1831.
[0240] The endonuclease may contain an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3638 to 3955. In some cases, the endonuclease may contain an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3638 to 3955. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 3638 to 3955. The endonuclease may contain an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3638 to 3955. In some cases, the endonuclease may contain an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3638 to 3955. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 3638 to 3955. The endonuclease may contain an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 3638 to 3641. In some cases, the endonuclease may contain an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3638 to 3641. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 3638 to 3641.The endonuclease may contain an HNH domain having at least about 70% identity to any one of SEQ ID NO: 3638. In some cases, the endonuclease may contain an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NO: 3638. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NO: 3638. The endonuclease may contain an HNH domain having at least about 70% identity to any one of SEQ ID NO: 3639. In some cases, the endonuclease may contain an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NO: 3639. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NO: 3639. The endonuclease may contain an HNH domain having at least about 70% identity to any one of SEQ ID NO: 3640. In some cases, the endonuclease may contain an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NO: 3640. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NO: 3640. The endonuclease may contain an HNH domain having at least about 70% identity to any one of SEQ ID NO: 3641.In some cases, the endonuclease may contain an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NO: 3641. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NO: 3641.
[0241] In some cases, the endonuclease may contain a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-6 or SEQ ID NOs: 9-319. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1-6 or SEQ ID NOs: 9-319. In some cases, the endonuclease may contain a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1-4. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1-4. In some cases, the endonuclease may contain a peptide motif that is substantially identical to any one of SEQ ID NOs: 5615, 5616, or 5617.
[0242] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 1 to 6 or SEQ ID NOs: 9 to 319, or with respect to any one of SEQ ID NOs: 1 to 319 with at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity. The NLS can be the SV40 large T antigen NLS. The NLS can be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity with respect to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1 below, or combinations thereof.
[0243]
Table 1
[0244] In some cases, the endonuclease can be a recombinant (e.g., cloned, expressed, and purified by suitable methods such as expression in Escherichia coli (E. coli) followed by epitope tag purification). In some cases, the endonuclease can be derived from a bacterium having a 16S rRNA gene with at least about 90% identity to any one of SEQ ID NOs: 5592 - 5595. The endonuclease can be derived from a species having a 16S rRNA gene with at least about 80%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5592 - 5595. The endonuclease can be derived from a species having a 16S rRNA gene that is substantially identical to any one of SEQ ID NOs: 5592 - 5595. The endonuclease can be derived from a bacterium belonging to the phylum Verrucomicrobia or the phylum Candidatus Peregrinibacteria.
[0245] In some cases, the sequence identity can be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith - Waterman homology search algorithm. The sequence identity can be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix gap cost of 11 for existence and 1 for extension, and using conditional composition score matrix adjustment.
[0246] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' of the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' of the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0247] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5476 - 5489. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60 to 90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5476 - 5489. In some cases, the tracrRNA may be substantially identical to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5476 - 5489. The tracrRNA may include any one of SEQ ID NOs: 5476 - 5489.
[0248] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to any one of SEQ ID NOs: 5461-5464. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5461-5464. The sgRNA may comprise a sequence substantially identical to any one of SEQ ID NOs: 5461-5464.
[0249] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, wherein the second region is 3' to the first region. In some cases, the above system, from 5' to 3', may comprise a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, a synthetic DNA sequence of at least about 10 nucleotides, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and may comprise a single-stranded or double-stranded DNA repair template.
[0250] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus. The method can include delivering to the target nucleic acid locus any of the non-natural systems disclosed herein that include an enzyme and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus, it can modify the target nucleic acid locus. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the target locus of interest. Optionally, the target nucleic acid locus includes deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0251] When the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 1827-2140. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a substantially identical sequence in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5572-5575 or to any one of SEQ ID NOs: 5572-5575. Optionally, the nucleic acid contains a promoter operably linked to an open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0252] In some cases, the present disclosure may provide a system disclosed herein or an expression cassette comprising a nucleic acid described herein. In some cases, the expression cassette or nucleic acid may be provided as a vector. In some cases, the expression cassette, nucleic acid, or vector may be provided intracellularly. In some cases, the cell is a bacterial cell having a 16S rRNA gene having at least about 90% (e.g., at least about 99%) identity to any one of SEQ ID NOs: 5592-5595.
[0253] MG2 enzyme
[0254] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2141 - 2241. Optionally, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2141 - 2241. Optionally, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2141 - 2142. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2141 - 2142. Optionally, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2141 - 2142.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2141 to 2142.
[0255] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 3955 to 4055. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3955 to 4055. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 2141 to 2142. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 3955 to 3956. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 3955 to 3956. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 3955 to 3956.
[0256] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 320 - 420. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 320 - 420. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 320 - 321. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 320 - 321.
[0257] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximate to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus relative to any one of SEQ ID NOs: 320-420, or to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 320-420. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593-5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593-5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0258] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix setting gap cost of 11 for existence and 1 for extension, and using a conditional composition score matrix adjustment.
[0259] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5'-targeting region complementary to the desired cleavage sequence. In some cases, the 5'-targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5'-targeting region may be 15 to 23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' to the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0260] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5490 - 5494. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60 to 90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5490 - 5494. In some cases, the tracrRNA may be substantially identical to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5490 - 5494. The tracrRNA may include any one of SEQ ID NOs: 5490 - 5494.
[0261] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5465. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5465. The sgRNA may comprise a sequence that is substantially identical to SEQ ID NO: 5465.
[0262] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, wherein the second region is 3' to the first region. In some cases, the above system comprises, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides relative to a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and may comprise a single-stranded or double-stranded DNA repair template.
[0263] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein that include an enzyme and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, the target nucleic acid locus of interest can be modified. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0264] When the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2141-2241. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a substantially identical sequence in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5576-5577 or any one of SEQ ID NOs: 5576-5577. Optionally, the nucleic acid contains a promoter operably linked to an open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0265] MG3 enzyme
[0266] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise a RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2242 - 2251. Optionally, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2242 - 2251. Optionally, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2242 - 2251. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2242 - 2244. Optionally, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2242 - 2244.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2242 to 2244.
[0267] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4056 to 4066. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4056 to 4066. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4056 to 4066. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4056 to 4058. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4056 to 4058. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4056 to 4058.
[0268] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 421 - 431. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 421 - 431. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 421 - 423. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 421 - 423.
[0269] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 421 to 431, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 421 to 431. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0270] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix setting gap cost of 11 for existence and 1 for extension, and using a conditional composition score matrix adjustment.
[0271] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15 to 23 nucleotides in length. The guide sequence and the tracr sequence may be supplied as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA-tracrRNA binding sequence at the 3' of the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' of the crRNA-tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0272] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5495 - 5502. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60 to 90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5495 - 5502. In some cases, the tracrRNA may be substantially identical to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of any one of SEQ ID NOs: 5495 - 5502. The tracrRNA may include any one of SEQ ID NOs: 5495 - 5502.
[0273] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to any one of SEQ ID NOs: 5466 - 5467. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5466 - 5467. The sgRNA may comprise a sequence substantially identical to any one of SEQ ID NOs: 5466 - 5467.
[0274] In some cases, the above system may comprise two different sgRNAs targeting a first region and a second region for cleavage at a target DNA locus, wherein the second region is 3' to the first region. In some cases, the above system, from 5' to 3', may comprise a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, a synthetic DNA sequence of at least about 10 nucleotides, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and may comprise a single-stranded or double-stranded DNA repair template.
[0275] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein that include an enzyme and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, the target nucleic acid locus of interest can be modified. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0276] When the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2242-2251. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a substantially identical sequence in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5578-5580 or to any one of SEQ ID NOs: 5578-5580. Optionally, the nucleic acid contains a promoter operably linked to an open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0277] MG4 enzyme
[0278] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise a RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2253-2481. Optionally, the endonuclease may comprise a RuvC_III domain, where the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2253-2481. Optionally, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2253-2481. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2253-2481. Optionally, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2253-2481.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2253 to 2481.
[0279] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4067 to 4295. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4067 to 4295. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4067 to 4295. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4067 to 4295. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4067 to 4295. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4067 to 4295.
[0280] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 432 to 660. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 432 to 660. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 432 to 660. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 432 to 660.
[0281] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 432 to 660, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 432 to 660. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0282] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using parameters of 3 word length (W), 10 expectation value (E), and BLOSUM62 scoring matrix setting gap costs of 11 existence, 1 extension, and conditional composition score matrix adjustment.
[0283] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15 to 23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' to the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0284] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5503. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5503. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5503. The tracrRNA may include SEQ ID NO: 5503.
[0285] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease can include a sequence having at least about 80% identity to SEQ ID NO: 5468. The sgRNA can include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5468. The sgRNA can include a sequence that is substantially identical to SEQ ID NO: 5468.
[0286] In some cases, the above system may include two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, where the second region is 3' to the first region. In some cases, the above system, from 5' to 3', relative to a first homology arm that includes a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, a synthetic DNA sequence of at least about 10 nucleotides, and relative to the second region, a second homology arm that includes a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and can include a single-stranded or double-stranded DNA repair template.
[0287] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein that include an enzyme and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, the target nucleic acid locus of interest can be modified. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0288] When the target nucleic acid locus can be intracellular, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2253-2481. Optionally, the nucleic acid comprises a promoter operably linked to an open reading frame encoding an endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be supplied as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be supplied as a translated polypeptide. At least one engineered sgRNA can be supplied as deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA polymerase III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0289] MG6 enzyme
[0290] In one aspect, the present disclosure provides an engineered nuclease system that includes (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may include an RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2482-2489. Optionally, the endonuclease may include an RuvC_III domain, where the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 2482-2489. Optionally, the endonuclease may include an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2482-2489.
[0291] The endonuclease may include an HNH domain having at least about 70% identity to any one of SEQ ID NOs: 4296-4303. Optionally, the endonuclease may include an HNH domain having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4296-4303. The endonuclease may include an HNH domain that is substantially identical to any one of SEQ ID NOs: 4056-4066.
[0292] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 661 - 668. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 661 - 668.
[0293] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N - terminus or C - terminus of the endonuclease. The NLS may be added to the N - terminus or C - terminus with respect to any one of SEQ ID NOs: 661 - 668, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 661 - 668. The NLS may be the SV40 large T antigen NLS. The NLS may be the c - myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 - 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 - 5608. The NLS may include any one of the sequences in Table 1, or combinations thereof.
[0294] In some cases, the sequence identity can be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity can be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix with a gap cost of 11 for existence and 1 for extension, and using a conditional composition score matrix adjustment.
[0295] In some cases, the system can include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to a desired cleavage sequence. In some cases, the 5' targeting region can include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region can be G. In some cases, the 5' targeting region can be 15-23 nucleotides in length. The guide sequence and the tracr sequence can be supplied as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA can include a crRNA tracrRNA binding sequence 3' to the targeting region. The guide RNA can include a tracrRNA sequence preceded by a 4-nucleotide linker 3' to the crRNA tracrRNA binding region. The sgRNA can include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0296] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA array.
[0297] In some cases, the above system may include two different guide RNAs that target a first region and a second region for cleavage at a target DNA locus, where the second region is 3' to the first region. In some cases, the system may include a single-stranded or double-stranded DNA repair template that includes a first homology arm that includes a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, a synthetic DNA sequence of at least about 10 nucleotides, and a second homology arm that includes a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, from 5' to 3'.
[0298] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein, including the enzymes and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it can modify the target nucleic acid locus of interest. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0299] When the target nucleic acid locus can be intracellular, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2482-2489. Optionally, the nucleic acid includes a promoter operably linked to an open reading frame encoding an endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be supplied as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be supplied as a translated polypeptide. At least one engineered sgRNA can be supplied as deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA polymerase III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0300] MG7 enzyme
[0301] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise an RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2490 - 2498. Optionally, the endonuclease may comprise an RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2490 - 2498. Optionally, the endonuclease may comprise an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2490 - 2498. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2490 - 2498. Optionally, the endonuclease may comprise an RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2490 - 2498.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2490 to 2498.
[0302] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4304 to 4312. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4304 to 4312. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4304 to 4312. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4304 to 4312. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4304 to 4312. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4304 to 4312.
[0303] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 669-677. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 669-677. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 669-677. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 669-677.
[0304] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 669 to 677, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 669 to 677. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0305] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix setting gap cost of 11 for existence and 1 for extension, and using a conditional composition score matrix adjustment.
[0306] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' of the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' of the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence that can hybridize to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0307] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5504. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60 to 90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5504. In some cases, the tracrRNA may be substantially identical to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5504. The tracrRNA may include SEQ ID NO: 5504.
[0308] In some cases, the above system may include two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, where the second region is 3' to the first region. In some cases, the system may include, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides relative to a first homology arm that includes a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, and a second homology arm that includes a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and may include a single-stranded or double-stranded DNA repair template.
[0309] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein, including the enzymes and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, the target nucleic acid locus of interest can be modified. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0310] When the target nucleic acid locus can be intracellular, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2490-2498. Optionally, the nucleic acid includes a promoter operably linked to an open reading frame encoding an endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be supplied as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be supplied as a translated polypeptide. At least one engineered sgRNA can be supplied as deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA polymerase III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0311] MG14 enzyme
[0312] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise an RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2499 - 2750. Optionally, the endonuclease may comprise an RuvC_III domain, where the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2499 - 2750. Optionally, the endonuclease may comprise an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2499 - 2750. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2499 - 2750. Optionally, the endonuclease may comprise an RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2499 - 2750.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2499 to 2750.
[0313] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4313 to 4564. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4313 to 4564. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4313 to 4564. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4313 to 4564. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4067 to 4295. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4313 to 4564.
[0314] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 678 - 929. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 678 - 929. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 678 - 929. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 678 - 929.
[0315] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 678 to 929, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 678 to 929. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of the sequences in Table 1, or a combination thereof.
[0316] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using parameters of 3 word length (W), 10 expectation value (E), and BLOSUM62 scoring matrix setting gap costs of 11 existence, 1 extension, and conditional composition score matrix adjustment.
[0317] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' to the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0318] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5505. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5505. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5505. The tracrRNA may include SEQ ID NO: 5505.
[0319] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5469. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5469. The sgRNA may comprise a sequence that is substantially identical to SEQ ID NO: 5469.
[0320] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, where the second region is 3' to the first region. In some cases, the above system comprises a single-stranded or double-stranded DNA repair template comprising, 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides relative to a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region.
[0321] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein, including the enzymes and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it can modify the target nucleic acid locus of interest. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0322] When the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2499-2750. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a sequence substantially identical to SEQ ID NO: 5581, or in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5581. Optionally, the nucleic acid contains a promoter operably linked to the open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0323] MG15 enzyme
[0324] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 2751-2913. Optionally, the endonuclease may comprise a RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2751-2913. Optionally, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2751-2913. The endonuclease may comprise a RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2751-2913. Optionally, the endonuclease may comprise a RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2751-2913.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2751 to 2913.
[0325] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4565 to 4727. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4565 to 4727. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4565 to 4727. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4565 to 4727. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4565 to 4727. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4565 to 4727.
[0326] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 930 - 1092. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 930 - 1092. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 930 - 1092. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 930 - 1092.
[0327] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 930 to 1092, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 930 to 1092. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0328] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix setting gap cost of 11 for existence and 1 for extension, and using a conditional composition score matrix adjustment.
[0329] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15 to 23 nucleotides in length. The guide sequence and the tracr sequence may be supplied as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA-tracrRNA binding sequence at the 3' of the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' of the crRNA-tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0330] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5506. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5506. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5506. The tracrRNA may include SEQ ID NO: 5506.
[0331] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5470. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5470. The sgRNA may comprise a sequence substantially identical to SEQ ID NO: 5470.
[0332] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, wherein the second region is 3' to the first region. In some cases, the system may comprise a single-stranded or double-stranded DNA repair template comprising, 5' to 3', a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, a synthetic DNA sequence of at least about 10 nucleotides, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region.
[0333] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein that include an enzyme disclosed herein and at least one synthetic guide RNA (sgRNA). The enzyme can form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, the target nucleic acid locus of interest can be modified. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus includes deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or proximal to the target locus of interest.
[0334] When the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2751-2913. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a sequence substantially identical to SEQ ID NO: 5582 or in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5582. Optionally, the nucleic acid contains a promoter operably linked to an open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0335] MG16 enzyme
[0336] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2914-3174. Optionally, the endonuclease may comprise an RuvC_III domain, where the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2914-3174. Optionally, the endonuclease may comprise an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2914-3174. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 2914-3174. Optionally, the endonuclease may comprise an RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 2914-3174.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 2914 to 3174.
[0337] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4728 to 4988. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4728 to 4988. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4728 to 4988. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4728 to 4988. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4728 to 4988. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4728 to 4988.
[0338] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1093 to 1353. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1093 to 1353. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1093 to 1353. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1093 to 1353.
[0339] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 1093 to 1353, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1093 to 1353. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0340] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using parameters of 3 word length (W), 10 expectation value (E), and a BLOSUM62 scoring matrix setting gap cost of 11 existence, 1 extension, and using conditional composition score matrix adjustment.
[0341] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be provided as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' of the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' of the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0342] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5507. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60 to 90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5507. In some cases, the tracrRNA may be substantially identical to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5507. The tracrRNA may include SEQ ID NO: 5507.
[0343] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5471. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5471. The sgRNA may comprise a sequence that is substantially identical to SEQ ID NO: 5471.
[0344] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, wherein the second region is 3' to the first region. In some cases, the above system comprises, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides relative to a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and may comprise a single-stranded or double-stranded DNA repair template.
[0345] In another aspect, the present disclosure provides a method for modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein that include an enzyme and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with the at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it can modify the target nucleic acid locus of interest. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0346] When the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 2914-3174. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a sequence substantially identical in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5583. Optionally, the nucleic acid contains a promoter operably linked to an open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0347] MG18 enzyme
[0348] In one aspect, the present disclosure provides an engineered nuclease system that includes (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may include an RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 3175 - 3300. Optionally, the endonuclease may include an RuvC_III domain, where the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 3175 - 3300. Optionally, the endonuclease may include an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3175 - 3300. The endonuclease may include an RuvC_III domain that has at least about 70% sequence identity to any one of SEQ ID NOs: 3175 - 3300. Optionally, the endonuclease may include an RuvC_III domain that has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 3175 - 3300.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3175 to 3300.
[0349] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4989 to 5146. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4989 to 5146. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4989 to 5146. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 4989 to 5146. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 4989 to 5146. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 4989 to 5146.
[0350] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1354 to 1511. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1354 to 1511. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1354 to 1511. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1354 to 1511.
[0351] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 1354 to 1511, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1354 to 1511. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0352] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix setting gap cost of 11 for existence and 1 for extension, and using a conditional composition score matrix adjustment.
[0353] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15 to 23 nucleotides in length. The guide sequence and the tracr sequence may be supplied as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' of the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' of the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0354] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5508. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) of the contiguous nucleotides of SEQ ID NO: 5508. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) contiguous nucleotides of SEQ ID NO: 5508. The tracrRNA may include SEQ ID NO: 5508.
[0355] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5472. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5472. The sgRNA may comprise a sequence that is substantially identical to SEQ ID NO: 5472.
[0356] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, where the second region is 3' to the first region. In some cases, the system may comprise a single-stranded or double-stranded DNA repair template comprising, 5' to 3', a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, a synthetic DNA sequence of at least about 10 nucleotides, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region.
[0357] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein, including the enzymes and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, it can modify the target nucleic acid locus of interest. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0358] When the target nucleic acid locus can be intracellular, the enzyme can be provided as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain having at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 3175-3300. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a sequence substantially identical in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5584. Optionally, the nucleic acid contains a promoter operably linked to the open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be provided as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be provided as a translated polypeptide. At least one engineered sgRNA can be provided as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0359] MG21 enzyme
[0360] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise an RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 3331-3474. Optionally, the endonuclease may comprise an RuvC_III domain, wherein the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 3331-3474. Optionally, the endonuclease may comprise an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3331-3474. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 3331-3474. Optionally, the endonuclease may comprise an RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 3331-3474.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3331 to 3474.
[0361] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 5147 to 5290. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5147 to 5290. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 5147 to 5290. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 5147 to 5290. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5147 to 5290. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 5147 to 5290.
[0362] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1512 to 1655. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1512 to 1655. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1512 to 1655. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1512 to 1655.
[0363] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 1512 to 1655, or with respect to variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1512 to 1655. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% identity to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0364] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using a word length (W) of 3, an expectation value (E) of 10, and a BLOSUM62 scoring matrix setting gap cost of 11 for existence and 1 for extension, and using a conditional composition score matrix adjustment.
[0365] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15-23 nucleotides in length. The guide sequence and the tracr sequence may be supplied as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' to the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0366] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5509. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60-90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5509. In some cases, the tracrRNA may be substantially identical to at least about 60-100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5509. The tracrRNA may include SEQ ID NO: 5509.
[0367] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease may comprise a sequence having at least about 80% identity to SEQ ID NO: 5473. The sgRNA may comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5473. The sgRNA may comprise a sequence that is substantially identical to SEQ ID NO: 5473.
[0368] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, where the second region is 3' to the first region. In some cases, the above system comprises, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides relative to a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and may comprise a single-stranded or double-stranded DNA repair template.
[0369] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein, including the enzymes and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, the target nucleic acid locus of interest can be modified. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0370] When the target nucleic acid locus can be intracellular, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 3331-3474. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a sequence substantially identical in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5585. Optionally, the nucleic acid contains a promoter operably linked to an open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta-actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be supplied as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be supplied as a translated polypeptide. At least one engineered sgRNA can be supplied as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0371] MG22 enzyme
[0372] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 3475 - 3568. Optionally, the endonuclease may comprise an RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 3475 - 3568. Optionally, the endonuclease may comprise an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3475 - 3568. The endonuclease may comprise an RuvC_III domain having at least about 70% sequence identity to any one of SEQ ID NOs: 3475 - 3568. Optionally, the endonuclease may comprise an RuvC_III domain having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 3475 - 3568.In some cases, the endonuclease may contain an RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3475 to 3568.
[0373] The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 5291 to 5389. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5291 to 5389. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 5291 to 5389. The endonuclease may contain an HNH domain that has at least about 70% identity to any one of SEQ ID NOs: 5291 to 5389. In some cases, the endonuclease may contain an HNH domain that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 5291 to 5389. The endonuclease may contain an HNH domain that is substantially identical to any one of SEQ ID NOs: 5291 to 5389.
[0374] In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1656 to 1755. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1656 to 1755. In some cases, the endonuclease may include variants having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 1656 to 1755. In some cases, the endonuclease may be substantially identical to any one of SEQ ID NOs: 1656 to 1755.
[0375] In some cases, the endonuclease may include variants having one or more nuclear localization sequences (NLSs). The NLS may be proximal to the N-terminus or C-terminus of the endonuclease. The NLS may be added to the N-terminus or C-terminus with respect to any one of SEQ ID NOs: 432 to 660, or with respect to any one of SEQ ID NOs: 1656 to 1755, with at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity. The NLS may be the SV40 large T antigen NLS. The NLS may be the c-myc NLS. The NLS may include a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99% identity with respect to any one of SEQ ID NOs: 5593 to 5608. The NLS may include a sequence substantially identical to any one of SEQ ID NOs: 5593 to 5608. The NLS may include any of the sequences in Table 1, or combinations thereof.
[0376] In some cases, the sequence identity may be determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, Novafold, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. The sequence identity may be determined by the BLASTP homology search algorithm using the parameters of 3 word length (W), 10 expectation value (E), and the BLOSUM62 scoring matrix setting gap cost of 11 existence, 1 extension, and using conditional composition score matrix adjustment.
[0377] In some cases, the system may include at least one engineered synthetic guide ribonucleic acid (sgRNA) that can form a complex with an endonuclease carrying a 5' targeting region complementary to the desired cleavage sequence. In some cases, the 5' targeting region may include a PAM sequence compatible with the endonuclease. In some cases, most of the nucleotides at the 5' of the targeting region may be G. In some cases, the 5' targeting region may be 15 to 23 nucleotides in length. The guide sequence and the tracr sequence may be supplied as separate ribonucleic acids (RNAs) or as a single ribonucleic acid (RNA). The guide RNA may include a crRNA tracrRNA binding sequence at the 3' to the targeting region. The guide RNA may include a tracrRNA sequence preceded by a 4-nucleotide linker at the 3' to the crRNA tracrRNA binding region. The sgRNA may include, from 5' to 3', a target sequence in the cell and a non-natural guide nucleic acid sequence capable of hybridizing to the tracr sequence. In some cases, the non-natural guide nucleic acid sequence and the tracr sequence are covalently linked.
[0378] In some cases, the tracr array may have a specific array. The tracr array may have at least about 80% identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of the native tracrRNA sequence. The tracr array may have at least about 80% sequence identity to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5510. In some cases, the tracrRNA may have at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to at least about 60 to 90 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5510. In some cases, the tracrRNA may be substantially identical to at least about 60 to 100 (e.g., at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, or at least about 90) consecutive nucleotides of SEQ ID NO: 5510. The tracrRNA may include SEQ ID NO: 5510.
[0379] In some cases, at least one engineered synthetic guide ribonucleic acid (sgRNA) capable of forming a complex with an endonuclease can comprise a sequence having at least about 80% identity to SEQ ID NO: 5474. The sgRNA can comprise a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5474. The sgRNA can comprise a sequence that is substantially identical to SEQ ID NO: 5474.
[0380] In some cases, the above system may comprise two different sgRNAs that target a first region and a second region for cleavage at a target DNA locus, where the second region is 3' to the first region. In some cases, the above system comprises, from 5' to 3', a synthetic DNA sequence of at least about 10 nucleotides relative to a first homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 5' to the first region, and a second homology arm comprising a sequence of at least about 20 (e.g., at least about 40, 80, 120, 150, 200, 300, 500, or 1 kb) nucleotides 3' to the second region, and may comprise a single-stranded or double-stranded DNA repair template.
[0381] In another aspect, the present disclosure provides a method of modifying a target nucleic acid locus of interest. The method can include delivering to the target nucleic acid locus of interest any of the non-natural systems disclosed herein that include an enzyme and at least one synthetic guide RNA (sgRNA) disclosed herein. The enzyme can form a complex with at least one sgRNA, and when the complex binds to the target nucleic acid locus of interest, the target nucleic acid locus of interest can be modified. Delivering the enzyme to the locus can include transfecting a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include electroporating a cell with the system or a nucleic acid encoding the system. Delivering the nuclease to the locus can include incubating the system in a buffer with a nucleic acid containing the locus of interest. Optionally, the target nucleic acid locus can include deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The target nucleic acid locus can include genomic DNA, viral DNA, viral RNA, or bacterial DNA. The target nucleic acid locus can be intracellular. The target nucleic acid locus can be in vitro. The target nucleic acid locus can be within a eukaryotic or prokaryotic cell. The cell can be an animal cell, a human cell, a bacterial cell, an archaeal cell, or a plant cell. The enzyme can induce a single-stranded or double-stranded break at or near the target locus of interest.
[0382] When the target nucleic acid locus can be intracellular, the enzyme can be supplied as a nucleic acid containing an open reading frame encoding an enzyme having an RuvC_III domain with at least about 75% (e.g., at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%) identity to any one of SEQ ID NOs: 3475 - 3568. The deoxyribonucleic acid (DNA) containing the open reading frame encoding the endonuclease can contain a sequence substantially identical to SEQ ID NO: 5586, or in a variant having at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 5586. Optionally, the nucleic acid contains a promoter operably linked to the open reading frame encoding the endonuclease. The promoter can be a CMV, EF1a, SV40, PGK1, Ubc, human beta - actin, CAG, TRE, or CaMKIIa promoter. The endonuclease can be supplied as a capped mRNA containing the open reading frame encoding the endonuclease. The endonuclease can be supplied as a translated polypeptide. At least one engineered sgRNA can be supplied as a deoxyribonucleic acid (DNA) containing a gene sequence encoding the at least one engineered sgRNA operably linked to an RNA pol III promoter. Optionally, the organism can be a eukaryote. Optionally, the organism can be a fungus. Optionally, the organism can be a human.
[0383] MG23 enzyme
[0384] In one aspect, the present disclosure provides an engineered nuclease system comprising (a) an endonuclease. Optionally, the endonuclease is a Cas endonuclease. Optionally, the endonuclease is a type II, class II Cas endonuclease. The endonuclease may comprise a RuvC_III domain, and the RuvC_III domain has at least about 70% sequence identity to any one of SEQ ID NOs: 3569 - 3637. Optionally, the endonuclease may comprise a RuvC_III domain, where the RuvC_III domain has at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identity to any one of SEQ ID NOs: 3569 - 3637. Optionally, the endonuclease may comprise a RuvC_III domain that is substantially identical to any one of SEQ ID NOs: 3569 - 363...
Claims
1. 1. An engineered nuclease system comprising: (a) an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846 or SEQ ID NO: 6257; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, (i) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) an engineered guide ribonucleic acid structure comprising a ribonucleic acid sequence configured to bind to said endonuclease; An engineered nuclease system comprising:
2. 1. An engineered nuclease system comprising: (a) an endonuclease configured to bind to a protospacer adjacent motif (PAM) sequence comprising SEQ ID NOs: 5847-5861 or 6258-6278, wherein the endonuclease is a Class 2 Type II Cas endonuclease; and (b) an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease, (i) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) an engineered guide ribonucleic acid structure comprising a ribonucleic acid sequence configured to bind to said endonuclease; An engineered nuclease system comprising:
3. 3. The engineered nuclease system of claim 1 or 2, wherein the endonuclease is derived from an uncultured microorganism.
4. 4. The engineered nuclease system of any one of claims 1 to 3, wherein the endonuclease is not engineered to bind to a different PAM sequence.
5. 5. The engineered nuclease system of any one of claims 1 to 4, wherein the endonuclease is not Cas9 endonuclease, Cas14 endonuclease, Cas12a endonuclease, Cas12b endonuclease, Cas12c endonuclease, Cas12d endonuclease, Cas12e endonuclease, Cas13a endonuclease, Cas13b endonuclease, Cas13c endonuclease, or Cas13d endonuclease.
6. 6. The engineered nuclease system of any one of claims 1 to 5, wherein the endonuclease has less than 80% identity to a Cas9 endonuclease.
7. 7. The engineered nuclease system of any one of claims 1-6, wherein the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of (a) any one of SEQ ID NOs: 5886-5887, 5891, 5893, or 5894; or (b) any one of SEQ ID NOs: 5862-5885, 5888-5890, 5892, 5895-5896, or 6279-6301.
8. 1. An engineered nuclease system comprising: (a) an engineered guide ribonucleic acid structure, (i) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (ii) comprises a ribonucleic acid sequence configured to bind to an endonuclease; wherein the ribonucleic acid sequence comprises a sequence having at least 80% sequence identity to the non-degenerate nucleotides of any one of (a) SEQ ID NOs: 5886-5887, 5891, 5893, or 5894; or (b) SEQ ID NOs: 5862-5885, 5888-5890, 5892, 5895-5896, or 6279-6301; and (b) a class 2 type II Cas endonuclease configured to bind to the engineered guide ribonucleic acid. An engineered nuclease system comprising:
9. 9. The engineered nuclease system of claim 8, wherein the endonuclease is configured to bind to a protospacer adjacent motif (PAM) sequence selected from the group consisting of SEQ ID NOs:5847-5861 or SEQ ID NOs:6258-6278.
10. 10. The engineered nuclease system of any one of claims 8 to 9, wherein the guide ribonucleic acid sequence is 15-24 nucleotides in length or 19-24 nucleotides in length.
11. 11. The engineered nuclease system of any one of claims 1 to 10, wherein the endonuclease comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease.
12. 12. The engineered nuclease system of any one of claims 1 to 11, wherein the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612.
13. 13. The engineered nuclease system of any one of claims 1 to 12, further comprising a single-stranded or double-stranded DNA repair template comprising, in 5' to 3' order, a first homology arm comprising a sequence of at least 20 nucleotides that is 5' to the target deoxyribonucleic acid sequence, a synthetic DNA sequence of at least 10 nucleotides, and a second homology arm comprising a sequence of at least 20 nucleotides that is 3' to the target sequence.
14. 14. The engineered nuclease system of claim 13, wherein the first homology arm or the second homology arm comprises a sequence of at least 40, 80, 120, 150, 200, 300, 500, or 1,000 nucleotides.
15. 15. The engineered nuclease system of any one of claims 1 to 14, wherein the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using parameters of the Smith-Waterman homology search algorithm.
16. The engineered nuclease system of claim 15, wherein the sequence identity is determined by the BLASTP homology search algorithm with parameters of word length (W) of 3, expectation (E) of 10, and a BLOSUM62 scoring matrix setting gap costs at presence of 11, extension of 1, and with a conditional composition score matrix adjustment.
17. An engineered guide ribonucleic acid polynucleotide, (a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment comprising two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex; wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide, The engineered guide ribonucleic acid polynucleotide is configured to form a complex with an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 5718-5846 or SEQ ID NO: 6257, and to target the complex to the target sequence of the target DNA molecule. An engineered guide ribonucleic acid polynucleotide.
18. 18. The engineered guide ribonucleic acid polynucleotide of Claim 17, wherein the DNA targeting segment is located 5' to both of the two complementary stretches of nucleotides.
19. A deoxyribonucleic acid polynucleotide encoding an engineered guide ribonucleic acid polynucleotide or structure according to any one of claims 17 to 18.
20. 1. A nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, said nucleic acid encoding an endonuclease comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs:5718-5846 or SEQ ID NO:6257.
21. 21. The nucleic acid of claim 20, wherein the endonuclease comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the endonuclease.
22. 22. The nucleic acid of claim 21, wherein the NLS comprises a sequence selected from SEQ ID NOs: 5597-5612.
23. 23. The nucleic acid of any one of claims 20 to 22, wherein the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.
24. A vector comprising the nucleic acid according to any one of claims 20 to 23.
25. (a) a ribonucleic acid sequence configured to hybridize to a target deoxyribonucleic acid sequence, and (b) a ribonucleic acid sequence configured to bind to the endonuclease.
25. The vector of claim 24, further comprising a nucleic acid encoding an engineered guide ribonucleic acid structure configured to form a complex with the endonuclease comprising:
26. The vector of any one of claims 24 to 25, which is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus.
27. A cell comprising the vector according to any one of claims 24 to 26.
28. 28. A method for producing an endonuclease comprising culturing the cell of claim 27.
29. 1. A method for binding, cleaving, marking or modifying a double-stranded deoxyribonucleic acid polynucleotide, comprising: contacting the double-stranded deoxyribonucleic acid polynucleotide in a complex with a Class 2 Type II Cas endonuclease and an engineered guide ribonucleic acid structure configured to bind to the endonuclease and the double-stranded deoxyribonucleic acid polynucleotide; wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a protospacer adjacent motif (PAM), and The PAM comprises a sequence selected from the group consisting of SEQ ID NOs: 5847-5861 or SEQ ID NOs: 6258-6278; method.
30. 30. The method of claim 29, wherein the double-stranded deoxyribonucleic acid polynucleotide comprises a first strand comprising a sequence complementary to a sequence of the engineered guide ribonucleic acid structure, and a second strand comprising the PAM.
31. The method of claim 30, wherein the PAM is immediately adjacent to the 3' end of the sequence complementary to the sequence of the engineered guide ribonucleic acid structure.
32. 32. The method of any one of claims 29 to 31, wherein the class 2 Type II Cas endonuclease is not a Cas9 endonuclease, a Cas14 endonuclease, a Cas12a endonuclease, a Cas12b endonuclease, a Cas12c endonuclease, a Cas12d endonuclease, a Cas12e endonuclease, a Cas13a endonuclease, a Cas13b endonuclease, a Cas13c endonuclease, or a Cas13d endonuclease.
33. 33. The method of any one of claims 29 to 32, wherein the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
34. 17. A method of modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus the engineered nuclease system of any one of claims 1 to 16, wherein the endonuclease is configured to form a complex with the engineered guide ribonucleic acid structure, the complex being configured such that when the complex binds to the target nucleic acid locus, the complex modifies the target nucleic acid locus.
35. 35. The method of claim 34, wherein modifying the target nucleic acid locus comprises binding, nicking, cleaving, or marking the target nucleic acid locus.
36. 36. The method of any one of claims 34 to 35, wherein the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
37. 37. The method of claim 36, wherein the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA.
38. 38. The method of any one of claims 34 to 37, wherein the target nucleic acid locus is in vitro.
39. The method of any one of claims 34 to 37, wherein the target nucleic acid locus is intracellular.
40. 40. The method of claim 39, wherein the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell.
41. 41. The method of any one of claims 34 to 40, wherein delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid according to any one of claims 20 to 23, or a vector according to any one of claims 24 to 26.
42. 41. The method of any one of claims 34 to 40, wherein delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the endonuclease.
43. 42. The method of claim 41 , wherein the nucleic acid comprises a promoter to which the open reading frame encoding the endonuclease is operably linked.
44. 42. The method of any one of claims 34-41, wherein delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the endonuclease.
45. 42. The method of any one of claims 34-41, wherein delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a translated polypeptide.
46. 42. The method of any one of claims 34-41, wherein delivering the engineered nuclease system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding the engineered guide ribonucleic acid structure operably linked to a ribonucleic acid (RNA) pol III promoter.
47. 47. The method of any one of claims 34 to 46, wherein the endonuclease induces a single-stranded or double-stranded break at or adjacent to the target locus.
48. 1. A method for editing a TRAC locus in a cell, comprising administering to the cell: (a) an RNA-guided endonuclease; and (b) an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the TRAC locus. contacting the wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5950-5958 or 5959-5965; method.
49. 49. The method of claim 48, wherein the RNA-guided endonuclease is a Class II Type II Cas endonuclease.
50. 50. The method of claim 48 or claim 49, wherein the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity to SEQ ID NO:2242 or SEQ ID NO:2244.
51. 51. The method of claim 50, wherein the RNA-guided endonuclease further comprises an HNH domain.
52. 52. The method of any one of claims 48 to 51, wherein the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO:
423.
53. 53. The method of any one of claims 48-52, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5950-5958, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
421.
54. 53. The method of any one of claims 48-52, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5959-5965, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
423.
55. 53. The method of any one of claims 48-52, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5953-5957.
56. 53. The method of any one of claims 48-52, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5960-5961 or SEQ ID NOs: 5963-5964.
57. 1. A method for editing a TRBC locus in a cell, comprising administering to said cell: (a) an RNA-guided endonuclease; and (b) an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the TRBC locus. contacting the wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5966-6004 or 6005-6025; method.
58. 58. The method of claim 57, wherein the RNA-guided endonuclease is a Class II Type II Cas endonuclease.
59. 59. The method of claim 57 or claim 58, wherein the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity to SEQ ID NO:2242 or SEQ ID NO:2244.
60. 60. The method of claim 59, wherein the RNA-guided endonuclease further comprises an HNH domain.
61. 61. The method of any one of claims 57 to 60, wherein the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO:
423.
62. 62. The method of any one of claims 57-61, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5966-6004, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
421.
63. 62. The method of any one of claims 57-61, wherein said engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6005-6025, and said endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
423.
64. 62. The method of any one of claims 57-61, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 5970, 5971, 5983, or 5984.
65. 62. The method of any one of claims 57-61, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6006, 6010, 6011, or 6012.
66. 1. A method for editing the GR (NR3C1) locus in a cell, comprising administering to the cell (a) an RNA-guided endonuclease; and (b) an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the GR(NR3C1) locus. contacting the wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6026-6090 or 6091-6121; method.
67. 67. The method of claim 66, wherein the RNA-guided endonuclease is a Class II Type II Cas endonuclease.
68. 68. The method of claim 66 or claim 67, wherein the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244.
69. 69. The method of claim 68, wherein the RNA-guided endonuclease further comprises an HNH domain.
70. 70. The method of any one of claims 66 to 69, wherein the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO:
423.
71. 71. The method of any one of claims 66-70, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6026-6090, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
421.
72. 71. The method of any one of claims 66-70, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6091-6121, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
423.
73. 71. The method of any one of claims 66-70, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6027-6028, 6029, 6038, 6043, 6049, 6076, 6080, 6081, or 6086.
74. 71. The method of any one of claims 66-70, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6092, 6115, or 6119.
75. 1. A method for editing the AAVS1 locus in a cell, comprising administering to the cell: (a) an RNA-guided endonuclease; and (b) an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the AAVS1 locus. contacting the wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6122-6152; method.
76. 76. The method of claim 75, wherein the RNA-guided endonuclease is a Class II Type II Cas endonuclease.
77. 77. The method of claim 75 or claim 76, wherein the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244.
78. 69. The method of claim 68, wherein the RNA-guided endonuclease further comprises an HNH domain.
79. 79. The method of any one of claims 75 to 78, wherein the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO:
423.
80. 80. The method of any one of Claims 75-79, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6122, 6125-6126, 6128, 6131, 6133, 6136, 6141, 6143, or 6148.
81. 1. A method for editing a TIGIT locus in a cell, comprising administering to the cell: (a) an RNA-guided endonuclease; and (b) an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the TIGIT locus. contacting the wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6153-6181; method.
82. 82. The method of claim 81 , wherein the RNA-guided endonuclease is a Class II Type II Cas endonuclease.
83. 83. The method of claim 81 or claim 82, wherein the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO:
423.
84. 84. The method of any one of claims 81 to 83, wherein the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity to SEQ ID NO:2242 or SEQ ID NO:2244.
85. 85. The method of claim 84, wherein the RNA-guided endonuclease further comprises an HNH domain.
86. 86. The method of any one of claims 81-85, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 66155, 6159, 616, or 6172.
87. 1. A method for editing the CD38 locus in a cell, comprising administering to the cell: (a) an RNA-guided endonuclease; and (b) an engineered guide RNA, the engineered guide RNA configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize to a region of the CD38 locus; contacting the wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6182-6248 or 6249-6256; method.
88. 88. The method of claim 87, wherein the RNA-guided endonuclease is a Class II Type II Cas endonuclease.
89. 89. The method of claim 87 or claim 88, wherein the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity to SEQ ID NO: 2242 or SEQ ID NO: 2244.
90. 90. The method of claim 89, wherein the RNA-guided endonuclease further comprises an HNH domain.
91. 91. The method of any one of claims 87 to 90, wherein the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO: 421 or SEQ ID NO:
423.
92. 92. The method of any one of claims 87-91, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6182-6248, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
421.
93. 92. The method of any one of claims 87-91, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6249-6256, and the endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:
423.
94. 92. The method of any one of Claims 87-91, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of any one of SEQ ID NOs: 6182-6183, 6189, 6191, 6208, 6210, 6211, or 6215.
95. 92. The method of any one of claims 87-91, wherein the engineered guide RNA comprises a target sequence having at least 85% identity to at least 18 consecutive nucleotides of SEQ ID NO: 6251.
96. 96. The method of any one of claims 48 to 95, wherein the cell is a peripheral blood mononuclear cell, a T cell, a NK cell, a hematopoietic stem cell (HSCT), or a B cell.
97. An engineered guide ribonucleic acid polynucleotide, (a) a DNA-targeting segment comprising a nucleotide sequence that is complementary to a target sequence in a target DNA molecule; and (b) a protein-binding segment that comprises two complementary stretches of nucleotides that hybridize to form a double-stranded RNA (dsRNA) duplex; Including, wherein the two complementary stretches of nucleotides are covalently linked to each other with an intervening nucleotide, The engineered guide ribonucleic acid polynucleotide is configured to form a complex with a Class 2 Type II Cas endonuclease and target the complex to the target sequence of the target DNA molecule, wherein the DNA-targeting segment comprises a sequence having at least 85% identity to any one of SEQ ID NOs: 5950-5965, 5966-6025, 6026-6121, 6122-6152, 6153-6181, or 6182-6256; An engineered guide ribonucleic acid polynucleotide.
98. 98. The engineered guide ribonucleic acid polynucleotide of Claim 97, wherein the protein-binding segment comprises a sequence having at least 85% identity to any one of SEQ ID NO:5466 or SEQ ID NO:6304.
99. 1. A system for generating edited immune cells, comprising: (a) an RNA-guided endonuclease; (b) an engineered guide ribonucleic acid polynucleotide of Claim 97 configured to bind to said RNA-guided endonuclease; and (c) a single-stranded or double-stranded DNA repair template comprising a first homology arm and a second homology arm flanking a sequence encoding a chimeric antigen receptor (CAR); Including, the system.
100. 100. The system of claim 99, wherein the cell is a peripheral blood mononuclear cell, a T cell, a NK cell, a hematopoietic stem cell (HSCT), or a B cell.
101. 101. The system of claim 99 or 100, wherein the RNA-guided endonuclease is a Class II Type II Cas endonuclease.
102. 102. The method of any one of claims 99 to 101, wherein the RNA-guided endonuclease comprises a RuvCIII domain comprising a sequence having at least 75% identity to SEQ ID NO:2242 or SEQ ID NO:2244.
103. 103. The method of claim 102, wherein the RNA-guided endonuclease further comprises an HNH domain.
104. 104. The method of any one of claims 99 to 103, wherein the RNA-guided endonuclease comprises a sequence having at least 75% identity to SEQ ID NO:421 or SEQ ID NO:423.
Citation Information
Patent Citations
CRISPR-CPF1-Related Methods, Compositions, and Components for Cancer Immunotherapy
JP2019507599A
Novel CAS9 orthologs
US20190264232A1
Materials and methods for engineering cells and uses thereof in immuno-oncology
WO2019097305A2
Modification of immune-related genomic LOCI using paired crispr nickase ribonucleoproteins
WO2019200306A1
Compositions and methods for immunotherapy
WO2020081613A1