Compositions and methods for nucleic acid modification

By using nucleases with specific amino acid sequence identity and specific gRNA compositions, the problems of low editing efficiency and inaccurate targeting in eukaryotes are solved, and more efficient gene editing effects are achieved.

CN119948157APending Publication Date: 2025-05-06ACRIGEN BIOSCIENCES
View PDF 58 Cites 0 Cited by

Patent Information

Application Number
CN202380054428.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-02
Filing Date
2023-06-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing CRISPR/Cas systems have problems in eukaryotes with low editing efficiency, off-target events, target sequence preferences and nuclease delivery difficulties.

Method used

Compositions containing specific nucleases that have a sequence that is highly identical to a specific amino acid sequence and may include a nuclear localization sequence and purification tag in combination with a specific guide RNA (gRNA) to improve editing efficiency and targeting accuracy.

Benefits of technology

By improving the editing efficiency and targeting accuracy of nucleases, reducing off-target events and target sequence preferences, enhancing the delivery and expression of nucleases, and achieving more effective gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948157A_ABST
    Figure CN119948157A_ABST
Patent Text Reader

Abstract

The present disclosure provides nuclease enzymes for nucleic acid modification, and compositions, methods, and systems thereof. More specifically, the present disclosure provides compositions and systems comprising a nuclease comprising an amino acid sequence having at least 70% identity to any of SEQ ID NO: 1-250, and at least one gRNA for target nucleic acid modification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to nucleases and compositions, methods and systems thereof for nucleic acid modification.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Application No. 63 / 351,140 filed on June 10, 2022, U.S. Provisional Application No. 63 / 383,107 filed on November 10, 2022, and U.S. Provisional Application No. 63 / 482,936 filed on February 2, 2023, the contents of which are incorporated herein by reference in their entirety.

[0004] Sequence Listing Statement

[0005] The contents of the electronic sequence listing entitled ACRIG_404894_601.xml (size: 579,833 bytes; creation date: June 8, 2023) are incorporated herein by reference in their entirety. Background Art

[0006] Clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) nucleases dominate the nucleic acid editing landscape because they are versatile, rapid, and easy-to-use editing tools. The most fully characterized CRISPR-Cas nuclease, Cas9, utilizes one or more RNAs as sequence-specific targeting elements that attach the nuclease to the target nucleic acid. However, current CRISPR / Cas systems have some limitations in use, especially in eukaryotes, including low editing efficiency, off-target events, target sequence preferences, and efficient delivery and expression of nucleases. Summary of the invention

[0007] Provided herein are compositions comprising a nuclease, wherein the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or greater than 99% identity to any one of SEQ ID NOs: 1-250. In some embodiments, the amino acid sequence of the nuclease comprises any one of SEQ ID NOs: 1-250.

[0008] In some embodiments, the nuclease further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is located at the N-terminus, the C-terminus, or both the N-terminus and the C-terminus of the nuclease. In some embodiments, the NLS at the N-terminus of the nuclease and the NLS at the C-terminus are different sequences.

[0009] Also provided are nucleic acid molecules comprising a first polynucleotide sequence encoding a nuclease and vectors comprising the nucleic acid molecules. In some embodiments, the vector further comprises a promoter operably connected to the first polynucleotide sequence. In some embodiments, the vector further comprises a second polynucleotide sequence encoding a guide RNA (gRNA). In some embodiments, the vector further comprises a promoter operably connected to the second polynucleotide.

[0010] In some embodiments, the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 251-422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 251-343. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 344-422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 472-482. In some embodiments, the gRNA comprises SEQ ID NOs: 346, 420, 481, or 479.

[0011] In some embodiments, the gRNA comprises a tracr sequence, and the gRNA comprises one or more sequence deletions in or near the region comprising the tracr sequence. In some embodiments, one or more sequence deletions include sequences predicted to form a stem-loop structure. In some embodiments, one or more sequence deletions include sequences predicted to form a stem-loop structure at or near the 5' end of the gRNA. In some embodiments, the gRNA comprises SEQ ID NO: 346, 420, 481 or 479.

[0012] In some embodiments, the gRNA comprises a spacer sequence of at least 18 nucleotides in length. In some embodiments, the gRNA comprises a spacer sequence of between 18 and 20 nucleotides in length.

[0013] In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 410-419.

[0014] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 20, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 410-419.

[0015] In some embodiments, the nuclease comprises SEQ ID NO: 21, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and wherein at least one gRNA comprises any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482.

[0016] In some embodiments, the nuclease comprises SEQ ID NO: 22, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 311, 346, 381, and 398-399. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381, and 398-399.

[0017] In some embodiments, the nuclease comprises SEQ ID NO: 23, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 23, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 312, 346, and 382.

[0018] In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 350-355, 358, 361-363, 367-372, and 389-392. In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.

[0019] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and wherein at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 350-355, 358, 361-363, 367-372, and 389-392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and wherein at least one gRNA comprises any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.

[0020] In some embodiments, the nuclease comprises SEQ ID NO: 25, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 25, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 314, 346, 383, and 400.

[0021] In some embodiments, the nuclease comprises SEQ ID NO: 26, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 26, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 346, 384, and 392.

[0022] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and wherein at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and wherein at least one gRNA comprises any one of SEQ ID NOs: 346, 384, and 392.

[0023] In some embodiments, the nuclease comprises SEQ ID NO: 27, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385, and 401.

[0024] In some embodiments, the nuclease comprises SEQ ID NO: 28, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 28, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 317, 346, 386, and 402.

[0025] In some embodiments, the nuclease comprises SEQ ID NO: 29, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 318, 346, 387, and 403.

[0026] In some embodiments, the nuclease comprises SEQ ID NO: 36, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 356-360, and 373-378. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 36, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356-360, and 373-378.

[0027] Also provided is a system for modifying a first target nucleic acid, the system comprising: a) a nuclease or a first nucleic acid sequence encoding the nuclease, the nuclease comprising an amino acid sequence having 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, greater than 99%, or 100% identity to any one of SEQ ID NOs: 1-250; and b) at least one guide RNA (gRNA) or a nucleic acid encoding the at least one gRNA, the guide RNA comprising a sequence complementary to at least a portion of the first target nucleic acid and a region that associates with the nuclease.

[0028] In some embodiments, the nuclease is capable of recognizing a protospacer adjacent motif (PAM) sequence selected from the group comprising ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. In some embodiments, the gRNA comprises a spacer sequence complementary to the first strand sequence of the target nucleic acid, and wherein the first strand sequence is directly adjacent to a protospacer adjacent motif (PAM) sequence selected from the group comprising ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. In some embodiments, the PAM sequence comprises DTTR, wherein D is A, G, or T, and R is A or G.

[0029] In some embodiments, the nuclease is capable of preferentially modifying a first target nucleic acid comprising the PAM sequence ATTA, wherein R is A or G, compared to a first target nucleic acid comprising the PAM sequence TTTR.

[0030] In some embodiments, the nuclease is capable of modifying the target nucleic acid more efficiently than the modification efficiency of the nuclease SEQ ID NO: 471 of the target nucleic acid, wherein the PAM sequence contained in the target nucleic acid is ATTA.

[0031] In some embodiments, in the presence of gRNA, the nuclease is capable of modifying the first target nucleic acid. In some embodiments, the modification includes nucleic acid cleavage. In some embodiments, the modification includes one or more of modification of the target nucleic acid, regulation of target nucleic acid transcription, and modification of a polypeptide associated with the target nucleic acid.

[0032] In some embodiments, the nuclease further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is located at the N-terminus, the C-terminus, or both the N-terminus and the C-terminus of the nuclease. In some embodiments, the NLS at the N-terminus of the nuclease and the NLS at the C-terminus are different sequences. In some embodiments, the nuclease further comprises a purification tag.

[0033] In some embodiments, the gRNA further comprises a sequence that is complementary to at least a portion of a second target nucleic acid.

[0034] In some embodiments, the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 251-422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 251-343. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 344-422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 472-482. In some embodiments, the gRNA comprises SEQ ID NOs: 346, 420, 481, or 479.

[0035] In some embodiments, the gRNA comprises a tracr sequence, and the gRNA comprises one or more sequence deletions in or near the region comprising the tracr sequence. In some embodiments, one or more sequence deletions include sequences predicted to form a stem-loop structure. In some embodiments, one or more sequence deletions include sequences predicted to form a stem-loop structure at or near the 5' end of the gRNA. In some embodiments, the gRNA comprises SEQ ID NO: 346, 420, 481, or 479.

[0036] In some embodiments, the gRNA comprises a spacer sequence of at least 18 nucleotides in length. In some embodiments, the gRNA comprises a spacer sequence of between 18 and 20 nucleotides in length.

[0037] In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 410-419.

[0038] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 20, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 410-419.

[0039] In some embodiments, the nuclease comprises SEQ ID NO: 21, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and wherein at least one gRNA comprises any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482.

[0040] In some embodiments, the nuclease comprises SEQ ID NO: 22, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 311, 346, 381, and 398-399. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381, and 398-399.

[0041] In some embodiments, the nuclease comprises SEQ ID NO: 23, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 23, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 312, 346, and 382.

[0042] In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 350-355, 358, 361-363, 367-372, and 389-392. In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.

[0043] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and wherein at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 350-355, 358, 361-363, 367-372, and 389-392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and wherein at least one gRNA comprises any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.

[0044] In some embodiments, the nuclease comprises SEQ ID NO: 25, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 25, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 314, 346, 383, and 400.

[0045] In some embodiments, the nuclease comprises SEQ ID NO: 26, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 26, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 346, 384, and 392.

[0046] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and wherein at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and wherein at least one gRNA comprises any one of SEQ ID NOs: 346, 384, and 392.

[0047] In some embodiments, the nuclease comprises SEQ ID NO: 27, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385, and 401.

[0048] In some embodiments, the nuclease comprises SEQ ID NO: 28, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 28, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 317, 346, 386, and 402.

[0049] In some embodiments, the nuclease comprises SEQ ID NO: 29, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 318, 346, 387, and 403.

[0050] In some embodiments, the nuclease comprises SEQ ID NO: 36, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 356-360, and 373-378. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 36, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356-360, and 373-378.

[0051] In some embodiments, the nucleic acid molecule encoding each or both of the nuclease and gRNA is a DNA molecule, such as a vector, a plasmid, or a linear nucleic acid. In some embodiments, the nuclease is encoded in a messenger RNA. In some embodiments, the gRNA is contained in a small RNA.

[0052] In some embodiments, the nuclease and gRNA are encoded on the same nucleic acid. In some embodiments, the nuclease and gRNA are encoded on different nucleic acids.

[0053] Also provided are vectors comprising the disclosed systems. In some embodiments, the vector further comprises a first promoter operably linked to a nucleic acid encoding a nuclease and a second promoter operably linked to a nucleic acid encoding at least one gRNA. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is an AAV vector. In some embodiments, the first promoter and the second promoter are active in mammalian cells.

[0054] In some embodiments, the system further comprises a target nucleic acid.

[0055] In some embodiments, the system is a cell-free system.

[0056] Also provided are cells comprising the disclosed compositions and systems. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell or a human cell).

[0057] Also provided are methods for modifying a target nucleic acid, the methods comprising contacting the target nucleic acid with a nuclease, composition, vector or system described herein.

[0058] In some embodiments, the target nucleic acid sequence is located in a cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell or a human cell).

[0059] In some embodiments, introducing the system or composition into a cell comprises administering the system or composition to a subject. In some embodiments, administering comprises administering in vivo.

[0060] Also provided are kits comprising any or all of the components of the compositions or systems described herein. In some embodiments, the kits further comprise one or more reagents, shipping and / or packaging containers, one or more buffers, delivery devices, instructions, software, computing devices, or a combination thereof.

[0061] Other aspects and embodiments of the present disclosure will be apparent from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a graph of the editing activity of nucleases having SEQ ID NOs: 21, 24, and 36 (sgRNAs having SEQ ID NOs: 310, 131, and 325, respectively) in human cells.

[0063] Figure 2 is a graph of the editing activity of nucleases having SEQ ID NO: 21 (1-8), SEQ ID NO: 24 (9-16), and SEQ ID NO: 36 (17-24) in human cells using single guide RNAs (sgRNAs) having different lengths.

[0064] Figure 3 is a graph of the editing activity of the Kim-T1 target with a single guide RNA (sgRNA) of SEQ ID NO:346.

[0065] Figure 4 is a graph of the editing activity of the sgRNA off-target panel, where each sgRNA contained a mismatch at the indicated position.

[0066] FIG. 5A to FIG. 5D is SEQ ID NO:20( Figure 5A and Figure 5D )、SEQ ID NO:24( Figure 5B ) and SEQ ID NO:26( Figure 5C )'s nuclease editing activity against the Kim-T1 target with sgRNA. Figure 5E Schematic diagram of the predicted structure of tracrRNA (SEQ ID NO: 508) with a truncated middle region of the third and major RNA stems.

[0067] Figure 6 is a graph of the editing activities of the nucleases and Un1Cas12f1 of SEQ ID NOs: 20, 24, and 26 across different genomic target sequences.

[0068] Fig. 7A Schematic diagram of the predicted structure of tracrRNA with a complete repeat sequence (upper figure; SEQ ID NO: 509) and a truncated repeat sequence modified from SEQ ID NO: 346 (lower figure; SEQ ID NO: 510). Figure 7B is SEQ ID NO:20 and Fig. 7A Plot of the editing efficiency of the indicated tracrRNAs on the Kim-T1 target. Figure 7C Schematic diagram of the predicted structure of tracrRNA (SEQ ID NO: 508) with stem stability and A-kink modification modified from SEQ ID NO: 346. Fig.7D and Fig. 7E Graph of editing efficiency of nucleases of SEQ ID NOs: 24 and 20, respectively, with modified tracrRNA as indicated for Kim-T1 target.

[0069] Figure 8 is a graph of the editing efficiency of different length spacers (as indicated) of the nuclease of SEQ ID NO: 20. Un1Cas12f1 was used as a positive control, and NT stands for non-targeted cells, for determining the level of detection (LOD).

[0070] Fig. 9A and Fig. 9B is a graph of the editing efficiency of the nucleases of SEQ ID NOs: 20 and 26 and the indicated spacer sequences.

[0071] Fig.10 is a schematic diagram of representative AAV vector designs.

[0072] Fig.11Graph of editing efficiency of AAV constructs encoding the nuclease of SEQ ID NO: 20 with different guides. The guides shown here are: PCSK9_1 = GSp380, PCSK9_2 = GSp376, PCSK9_3 = GSp377, TTR_1 = GSp368, TTR_2 = GSp356, PRSS1 = GSp342, SMN2 = GSp251.

[0073] Fig.12 is a comparison of editing with AAV and nucleases of SEQ ID NO: 20 with different targets in the presence and absence of etoposide treatment. NT is a sample to which no AAV was added but which was processed, amplified, and sequenced using the same method as the AAV-treated sample. DETAILED DESCRIPTION

[0074] The disclosed compositions, systems, kits and methods include nucleases useful for nucleic acid modification. The disclosed nucleases allow for the use of gene editing with improved efficacy and safety in eukaryotic (e.g., mammalian (e.g., human)) therapeutics, diagnostics and research in vivo and ex vivo applications.

[0075] The section headings used in this section and throughout the disclosure herein are for organizational purposes only and are not intended to be limiting.

[0076] definition

[0077] As used herein, the terms "comprising," "including," "having," "has," "may," "containing," and variations thereof are intended to be open transitional phrases, terms, or words that do not exclude the possibility of additional behavior or structure. As used herein, comprising a certain sequence or a certain SEQ ID NO generally means that at least one copy of the sequence is present in the peptide or polynucleotide. However, two or more copies are also contemplated. Unless the context clearly specifies otherwise, the singular forms "a," "an," and "the" include plural referents. The present disclosure also contemplates other embodiments "comprising the embodiments or elements set forth herein," "consisting of the embodiments or elements set forth herein," and "consisting essentially of the embodiments or elements set forth herein," whether or not explicitly set forth.

[0078] For the recitation of numerical ranges herein, each intermediate value with the same degree of precision is expressly contemplated. For example, for the range of 6-9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range of 6.0-7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9 and 7.0 are expressly contemplated.

[0079] Unless otherwise defined herein, scientific and technical terms used in conjunction with the present disclosure should have the meanings commonly understood by those of ordinary skill in the art. For example, any terminology used in conjunction with cell and tissue culture, molecular biology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization as described herein and the techniques of the described fields are those well known and commonly used in the art. The meaning and scope of the terms should be clear; however, if there is any potential ambiguity, the definitions provided herein take precedence over any dictionary or external definition. In addition, unless the context otherwise requires, singular terms should include plural terms and plural terms should include singular terms.

[0080] As used herein, "nucleic acid", "nucleic acid sequence" refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, at pages 793-800 (Worth Pub. 1982)). The present technology contemplates any deoxyribonucleotide, ribonucleotide or peptide nucleic acid component and any chemical variant thereof, such as methylated, hydroxymethylated or glycosylated forms of these bases, etc. The polymer or oligomer may be heterogeneous or homogeneous in composition and may be isolated from a naturally occurring source or may be artificially or synthetically produced. In addition, the nucleic acid may be DNA or RNA or a mixture thereof, and may exist permanently or transiently in single-stranded or double-stranded form, including homoduplex, heteroduplex and hybrid states. In some embodiments, the nucleic acid or nucleic acid sequence comprises other types of nucleic acid structures, such as, for example, DNA / RNA helices, peptide nucleic acids (PNAs), morpholino nucleic acids (see, for example, Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002)) and U.S. Pat. No. 5,034,506), locked nucleic acids (LNAs; see Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97:5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122:8595-8602 (2000)) and / or ribozymes. Therefore, the term "nucleic acid" or "nucleic acid sequence" may also encompass chains comprising non-natural nucleotides, modified nucleotides, and / or non-nucleotide building blocks (e.g., "nucleotide analogs") that may exhibit the same functions as natural nucleotides; in addition, as used herein, the term "nucleic acid sequence" refers to oligonucleotides, nucleotides, or polynucleotides, and fragments or portions thereof, as well as DNA or RNA of genomic or synthetic origin, which may be single-stranded or double-stranded, and represent sense or antisense strands. The terms "nucleic acid," "polynucleotide," "nucleotide sequence," and "oligonucleotide" are used interchangeably. It refers to a polymeric form of nucleotides (deoxyribonucleotides or ribonucleotides, or analogs thereof) of any length.

[0081] Nucleic acid or amino acid sequence "identity" as described herein can be determined by comparing the nucleic acid or amino acid sequence of interest to a reference nucleic acid or amino acid sequence. The percent identity is the number of nucleotides or amino acid residues that are identical (e.g., equal) between the sequence of interest and the reference sequence divided by the length of the longest sequence (e.g., the length of the sequence of interest or the reference sequence, whichever is longer). Many mathematical algorithms for obtaining optimal alignments and calculating the identity between two or more sequences are known and are incorporated into many available software programs. Examples of such programs include CLUSTAL-W, T-Coffee, and ALIGN (for alignment of nucleic acid and amino acid sequences), BLAST programs (e.g., BLAST2.1, BL2SEQ and subsequent versions thereof), and FASTA programs (e.g., FASTA3x, FASTSEQ, and subsequent versions thereof). TM and SSEARCH) (for sequence alignment and sequence similarity searches). Sequence alignment algorithms are also disclosed in, for example, Altschul et al., J. Molecular Biol., 215(3):403-410 (1990); Beigert et al., Proc. Natl. Acad. Sci. USA, 106(10):3770-3775 (2009); Durbin et al., eds., Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press, Cambridge, UK (2009); Soding, Bioinformatics, 21(7):951-960 (2005); Altschul et al., Nucleic Acids Res., 25(17):3389-3402 (1997); and Gusfield, Algorithms on Strings, Trees and Sequences, Cambridge University Press, Cambridge UK (1997)).

[0082] The terms "non-naturally occurring," "engineered," and "synthetic" are used interchangeably and indicate that the human hand is involved. When referring to a nucleic acid molecule or polypeptide, the term means that the nucleic acid molecule or polypeptide is at least substantially free of at least one other component with which it is naturally associated in nature and found in nature and / or the nucleic acid molecule or polypeptide is associated with at least one other component with which it is not naturally associated in nature and / or there are one or more changes in the nucleic acid or amino acid sequence compared to the sequence found in nature.

[0083] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus or cosmid, to which another DNA segment, such as an "insert sequence," can be attached or integrated so as to cause the attached segment to replicate in the cell.

[0084] When exogenous DNA, such as a recombinant expression vector, has been introduced into a cell, the cell has been "genetically modified," "transformed," or "transfected" by such DNA. The presence of exogenous DNA results in permanent or transient genetic changes. The transforming DNA may or may not be integrated (covalently linked) into the cell genome. For example, the transforming DNA may be maintained on an episomal element such as a plasmid. Relative to eukaryotic cells, stably transformed cells are cells in which the transforming DNA is gradually integrated into the chromosome so that it is inherited to daughter cells by chromosome replication. This stability is demonstrated by the ability of eukaryotic cells to establish cell lines or clones containing a group of daughter cells containing the transforming DNA. A "clone" is a group of cells derived from a single cell or a common ancestor by mitosis. A "cell line" is a clone of primary cells that can stably grow for many generations in vitro.

[0085] The term "contacting" as used herein means to contact, to be in contact, or to be in contact. The term "contacting" as used herein means the state or condition of touching or being in direct or local proximity. The composition may be contacted with a target destination such as, but not limited to, an organ, tissue, cell, or tumor by any means of administration known to those skilled in the art.

[0086] As used herein, the terms "providing," "administering," and "introducing" are used interchangeably herein and refer to placing a composition or system of the present disclosure into a cell, an organism, or a subject by a method or route that results in at least partial localization to a desired site. The composition or system can be administered by any appropriate route that results in delivery to a desired site in a cell, an organism, or a subject.

[0087] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, preferred methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and are not intended to be limiting.

[0088] Nuclease

[0089] The progress and development of CRISPR-Cas genome editing tools (including nucleases and other Cas proteins) have driven major advances in nucleic acid editing. Nucleic acid editing has many uses, including in the field of diagnosis and treatment. Such breadth is accompanied by the diversity of nucleic acid targets and environments that are engineered for editing activity. Therefore, it is necessary to provide different and additional nucleases and related methods for the toolbox for nucleic acid editing.

[0090] Disclosed herein are compositions comprising nucleases having Cas-like activity. The disclosed nucleases comprise sequences having at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 93%, at least 95%, at least 98%, at least 99% or 100% identity) to the amino acid sequences of SEQ ID NO: 1-250. In some embodiments, the nuclease comprises a sequence having at least 90% identity to the amino acid sequences of SEQ ID NO: 1-250. In certain embodiments, the nuclease comprises the amino acid sequences of SEQ ID NO: 1-250.

[0091] Any of the nucleases described herein may comprise one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 150, etc.) amino acid substitutions. An amino acid "substitution" or "replacement" refers to the replacement of an amino acid at a given position or residue in a polypeptide sequence by another amino acid at the same position or residue. Amino acids are broadly classified as "aromatic" or "aliphatic". Aromatic amino acids include aromatic rings. Examples of "aromatic" amino acids include histidine (H or His), phenylalanine (F or Phe), tyrosine (Y or Tyr) and tryptophan (W or Trp). Non-aromatic amino acids are broadly classified as "aliphatic". Examples of "aliphatic" amino acids include glycine (G or Gly), alanine (A or Ala), valine (V or Val), leucine (L or Leu), isoleucine (I or He), methionine (M or Met), serine (S or Ser), threonine (T or Thr), cysteine ​​(C or Cys), proline (P or Pro), glutamate (E or Glu), aspartic acid (A or Asp), asparagine (N or Asn), glutamine (Q or GIn), lysine (K or Lys), and arginine (R or Arg).

[0092] Amino acid replacement or substitution can be conservative, semi-conservative or non-conservative. The phrase "conservative amino acid replacement" or "conservative mutation" refers to an amino acid replacement by another amino acid with common properties. A functional method for defining common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz and Schirmer, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such analysis, amino acid groups can be defined, wherein the amino acids within the group are preferentially exchanged with each other, and therefore are most similar to each other in terms of their impact on the entire protein structure (Schulz and Schirmer, supra). Examples of conservative amino acid replacements include amino acid replacements within the above-mentioned subgroups, for example, lysine replaces arginine and vice versa, so that a positive charge can be maintained, glutamic acid replaces aspartic acid and vice versa, so that a negative charge can be maintained, serine replaces threonine, so that free-OH can be maintained, and glutamine replaces asparagine, so that free-NH 2 "Semi-conservative mutations" include amino acid substitutions of amino acids within the same group listed above but not within the same subgroup. For example, substitutions of aspartic acid for asparagine or asparagine for lysine involve amino acids in the same group but in different subgroups. "Non-conservative mutations" involve substitutions of amino acids between different groups (e.g., lysine for tryptophan or phenylalanine for serine, etc.).

[0093] In some embodiments, the nuclease comprises one or more amino acid substitutions and has an amino acid sequence that is at least 70% identical (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 93%, at least 95%, at least 98%, at least 99% identical, or 100% identical) to the amino acid sequence of SEQ ID NO: 1-250. In some embodiments, the nuclease comprises one or more amino acid substitutions compared to SEQ ID NO: 1-250, and the one or more substitutions improve the editing efficiency of the nuclease.

[0094] The nucleases disclosed herein are capable of recognizing a wide range of protospacer adjacent motifs (PAMs) flanking a target nucleic acid. In certain embodiments, the nuclease can only cleave a target nucleic acid if an appropriate PAM is present. In certain embodiments, the nuclease has a broad ability to recognize target nucleic acids (e.g., those lacking a PAM or those recognized by a broad PAM).

[0095] The PAM is usually close to the target sequence. For example, the PAM can be a sequence that is closely or directly adjacent to the target nucleic acid. The PAM can be located 5' or 3' to the target sequence. The PAM can be located upstream or downstream of the target sequence. In one embodiment, the target nucleic acid is closely flanked by the PAM on the 3' end. In one embodiment, the target nucleic acid is closely flanked by the PAM on the 5' end.

[0096] The length of the PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. In certain embodiments, the length of the PAM is between 2-6 nucleotides.

[0097] Non-limiting examples of PAM sequences include: CC, CA, AG, GT, TA, AC, CA, GC, CG, GG, CT, TG, GA, AGG, TGG, T-rich PAMs (such as TTT, TTG, TTC, etc.), NGG, NGA, NAG, NGGNG and NNAGAAW, NNNNGATT, NAAR (R=A or G), NNGRR (R=A or G), NNAGAA and NAAAC, where "N" is any nucleotide.

[0098] In some embodiments, the nuclease disclosed herein is capable of recognizing a protospacer adjacent motif (PAM) sequence selected from the group comprising ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. In some embodiments, the PAM sequence comprises DTTR, wherein D is A, G, or T, and R is A or G.

[0099] Different PAM sequences can confer different preferences and efficiencies for nuclease cleavage or desired nuclease modification. In some embodiments, the nuclease preferentially modifies a first target nucleic acid comprising the PAM sequence ATTA compared to a target nucleic acid comprising the PAM sequence TTTR (wherein R is A or G). In some embodiments, the modification efficiency of the target nucleic acid by the nuclease disclosed herein is observed to be higher compared to the modification efficiency of the nuclease SEQ ID NO: 471. In some embodiments, when the PAM sequence contained in the target nucleic acid is ATTA, the modification efficiency of the target nucleic acid by the nuclease disclosed herein is observed to be higher compared to the modification efficiency of the nuclease SEQ ID NO: 471.

[0100] In some embodiments, the nuclease further comprises a nuclear localization sequence (NLS). The nuclear localization sequence can be attached to, for example, one or both of the N-terminus and the C-terminus. In some embodiments, the nuclease comprises two or more NLS. The two or more NLS can be in series, separated by a linker at the N-terminus or the C-terminus of the protein, or one or more can be inside the open reading frame of the nuclease.

[0101] The nuclear localization sequence may comprise any amino acid sequence known in the art to functionally mark or direct the protein to be introduced into the nucleus (e.g., for nuclear transport). Typically, the nuclear localization sequence comprises one or more positively charged amino acids, such as lysine and arginine.

[0102] In some embodiments, the NLS is a single component sequence. The single component NLS comprises a single cluster of positively charged or basic amino acids. In some embodiments, the single component NLS comprises a sequence of KK / RXK / R, wherein X can be any amino acid. Exemplary single component NLS sequences include sequences from SV40 large T antigen, c-myc, and TUS protein. In selected embodiments, the NLS includes an NLS of SV40 large T antigen, which comprises the amino acid sequence of PKKKRKV (SEQ ID NO: 504).

[0103] In some embodiments, the NLS is a two-component sequence. The two-component NLS comprises two clusters of basic amino acids separated by a spacer of about 9-12 amino acids. Exemplary two-component NLSs include the nuclear localization sequence of nucleoplasmin, EGL-12, or two-component SV40. In the embodiment selected, the NLS includes the NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK (SEQ ID NO: 505).

[0104] In some embodiments, two or more NLSs may have the same or different sequences. For example, in some embodiments, the nuclease comprises two NLSs, one sequence from the SV40 large T antigen and one sequence from a nucleoplasmic protein.

[0105] The NLS can be attached to the nuclease via a linker. The linker can be a polypeptide of any amino acid sequence and length. The linker can serve as a spacer peptide. In some embodiments, the linker is flexible. In some embodiments, the linker comprises at least one glycine and at least one serine. In some embodiments, the linker comprises (Gly 2 Ser) n An amino acid sequence composed of the following: wherein n is the number of repeated sequences which is an integer ranging from 2 to 20.

[0106] In some embodiments, the nuclease may include a tag (e.g., 3xFLAG tag, HA tag, Myc tag, etc.). The tag may facilitate tracking, separation, or purification of the nuclease. In some embodiments, the tag may be adjacent to the nuclear localization sequence, upstream thereof, or downstream thereof. The tag may be located at the N-terminus, the C-terminus, or a combination thereof of the nuclease.

[0107] In some embodiments, the nuclease is covalently linked to a peptide or protein in a fusion protein. The nuclease can be a part of a fusion protein comprising another protein or protein domain. For example, the nuclease can be fused to another protein or protein domain (e.g., GFP) that provides a label or visualization. The nuclease can be fused to a protein or protein domain that has other functionalities or activities (e.g., such as nuclease activity provided by FokI nuclease, protein modification activity such as histone modification activity (including acetylation or deacetylation or demethylation or methyltransferase activity), transcriptional regulatory activity such as the activity of a transcription activator or repressor, base editing activity such as deaminase activity, DNA modification activity such as DNA methylation activity, etc.) that can be used to target certain DNA sequences.

[0108] In some embodiments, the nuclease can be fused with one or more (e.g., two, three, four or more) protein transduction structures or PTDs (also referred to as CPP-cell penetrating peptides). The protein transduction domain is a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates passage through a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. The PTD attached to another molecule facilitates the passage of the molecule through the membrane, for example, from the extracellular space to the intracellular space, or from the cytosol to the organelle. In some embodiments, the PTD is covalently attached to the end (e.g., N-terminal, C-terminal, or both) of the nuclease. In some embodiments, the PTD is internally inserted at a suitable insertion site. Examples of PTDs include, but are not limited to, a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT); a polyarginine sequence comprising a sufficient number of arginines (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines) to direct entry into cells; a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); a Drosophila antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); a truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21:1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); a cell-penetrating peptide (Transportan), etc.

[0109] Nuclease can be fused via a joint polypeptide. The joint polypeptide can have any of a variety of amino acid sequences. Proteins can be connected by a spacer peptide, which is generally flexible, but other chemical bonds are not excluded. Suitable joints include polypeptides between 4 amino acids and 40 amino acids in length or between 4 amino acids and 25 amino acids in length. These joints can be produced by coupling proteins using synthetic oligonucleotides encoding joints, or can be encoded by a nucleic acid sequence encoding a fusion protein. Peptide joints with a certain degree of flexibility can be used. The connecting peptide can actually have any amino acid sequence, keeping in mind that a preferred joint will have a sequence that produces a generally flexible peptide. The use of small amino acids such as glycine and alanine is useful in forming flexible peptides. The creation of such sequences is conventional for those skilled in the art. A variety of different joints are commercially available and are considered to be suitable for use, including but not limited to glycine-serine polymers, glycine-alanine polymers and alanine-serine polymers.

[0110] Composition and system

[0111] Also disclosed herein are compositions comprising a nuclease as described herein or a nucleic acid molecule comprising a sequence encoding the nuclease.

[0112] Also disclosed herein is a system for modifying a target nucleic acid, the system comprising a nuclease as described herein (e.g., a nuclease comprising an amino acid sequence having at least 70% identity to an amino acid sequence of SEQ ID NOs: 1-250 (e.g., an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 93%, at least 95%, at least 98%, at least 99% identity, or 100% identity to an amino acid sequence of SEQ ID NOs: 1-250)) or a nucleic acid molecule comprising a sequence encoding the nuclease.

[0113] In some embodiments, the components of the system can be in the form of a composition. In some embodiments, the components of the composition or system of the present invention can be mixed with a carrier individually or in any combination, which is also within the scope of the present disclosure. Exemplary carriers include buffers, antioxidants, preservatives, carbohydrates, surfactants, etc.

[0114] Also disclosed are cells comprising the compositions or systems described herein. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells.

[0115] The compositions or systems disclosed herein may also include at least one gRNA or a nucleic acid encoding the at least one gRNA, the gRNA comprising a sequence complementary to at least a portion of the first target nucleic acid and a region associated with a nuclease. In some embodiments, at least one gRNA also comprises a sequence complementary to at least a portion of the second target nucleic acid. In the case where the composition or system comprises more than one gRNA, each gRNA may be encoded on a nucleic acid that is the same or different from another gRNA.

[0116] The gRNA can be crRNA, crRNA / tracrRNA (or single guide RNA, sgRNA). The terms "gRNA", "guide RNA" and "CRISPR guide sequence" are used interchangeably throughout the text and refer to a nucleic acid comprising a sequence that associates with a nuclease and determines the sequence specificity of the nuclease. The gRNA can be engineered to hybridize with (e.g., partially or completely complementarity with) a target nucleic acid sequence (e.g., a genome in a host cell).

[0117] In some embodiments, at least one gRNA is encoded in a CRISPR RNA (crRNA) array. The CRISPR array contains a series of direct repeat sequences separated by short sequences called spacers. The nucleases described herein may have a preference for direct repeat sequences. For example, a CRISPR RNA (crRNA) may contain multiple gRNAs or may contain more than one different sequence, each sequence being constructed to hybridize with a different target nucleic acid sequence.

[0118] The length of the gRNA or its part hybridized with the target nucleic acid (target site) can be between 15-40 nucleotides. In some embodiments, the length of the gRNA sequence hybridized with the target nucleic acid is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40 nucleotides. 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145

[0119] In addition to the sequence that binds to the target nucleic acid, in some embodiments, the gRNA may also include a scaffold sequence (e.g., tracrRNA). In some embodiments, such a chimeric gRNA may be referred to as a single guide RNA (sgRNA). Exemplary scaffold sequences are obvious to those skilled in the art and can be found in, for example, Jinek et al., Science (2012) 337(6096): 816-821 and Ran et al., Nature Protocols (2013) 8: 2281-2308, which are incorporated herein by reference in their entirety.

[0120] In some embodiments, the gRNA sequence does not include a scaffold sequence, and the scaffold sequence is expressed as a separate transcript. In such embodiments, the gRNA sequence also includes an additional sequence complementary to a portion of the scaffold sequence and is used to bind (hybridize) the scaffold sequence.

[0121] In some embodiments, the gRNA comprises a sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or at least 100% complementary to the target nucleic acid. In some embodiments, the sequence is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or at least 100% complementary to the 3' end of the target nucleic acid (e.g., the last 5, 6, 7, 8, 9 or 10 nucleotides of the 3' end of the target nucleic acid).

[0122] In some embodiments, the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 251-422 and 472-482. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 251-343. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 344-422. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 472-482.

[0123] The gRNA of the present disclosure may comprise a sequence having one or more nucleotide substitutions or mutations, truncations or insertions relative to any one of SEQ ID NOs: 251-343. Nucleotide substitutions or mutations, truncations or insertions may increase stability, modify secondary structural elements, increase binding efficiency with homologous nucleases or target chains, and increase. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 344-422. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 472-482. In some embodiments, the gRNA comprises SEQ ID NO: 346. In some embodiments, the gRNA comprises SEQ ID NO: 420. In some embodiments, the gRNA comprises SEQ ID NO: 481. In some embodiments, the gRNA comprises SEQ ID NO: 479.

[0124] In some embodiments, the gRNA comprises a spacer sequence. The spacer sequence can be any length or sequence. In some embodiments, the length of the spacer sequence is at least 18 (e.g., 18, 19, 20, 21, 22, 23, 24, etc.) nucleotides. In some embodiments, the length of the spacer sequence is between 18 and 20 nucleotides. Therefore, in certain embodiments, the length of the spacer sequence is 18 nucleotides. In certain embodiments, the length of the spacer sequence is 19 nucleotides. In certain embodiments, the length of the spacer sequence is 20 nucleotides.

[0125] In some embodiments, the gRNA comprises a spacer sequence complementary to the first strand sequence of the target nucleic acid. In some embodiments, the first strand sequence is directly adjacent to a protospacer adjacent motif (PAM) sequence selected from the group comprising ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG.

[0126] In some embodiments, the nuclease comprises SEQ ID NO: 21, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and wherein at least one gRNA comprises any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482. In some embodiments, the nuclease comprises SEQ ID NO:21 or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:21, and the gRNA comprises SEQ ID NO:346 or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical to SEQ ID NO:346.

[0127] In some embodiments, the nuclease comprises SEQ ID NO:24, and at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs:310, 313, 325, 346, 350-355, 358, 361-363, 367-372, and 389-392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and wherein at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 350-355, 358, 361-363, 367-372, and 389-392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and wherein at least one gRNA comprises any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392. In some embodiments, the nuclease comprises SEQ ID NO:24, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:24, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346. In some embodiments, the nuclease comprises SEQ ID NO:24, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:24, and the gRNA comprises SEQ ID NO:352, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:352.

[0128] In some embodiments, the nuclease comprises SEQ ID NO: 36, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 356-360, and 373-378. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 36, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356-360, and 373-378. In some embodiments, the nuclease comprises SEQ ID NO:36, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:36, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346. In some embodiments, the nuclease comprises SEQ ID NO:36, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:36, and the gRNA comprises SEQ ID NO:358, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:358.

[0129] In some embodiments, the nuclease comprises SEQ ID NO: 1, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 251-256. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 1, and wherein at least one gRNA comprises any one of SEQ ID NOs: 251-256.

[0130] In some embodiments, the nuclease comprises SEQ ID NO: 2, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 257-259. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 2, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 257-259.

[0131] In some embodiments, the nuclease comprises SEQ ID NO: 3, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 260-262. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 3, and wherein at least one gRNA comprises any one of SEQ ID NOs: 260-262.

[0132] In some embodiments, the nuclease comprises SEQ ID NO: 4, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 263-265. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 4, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 263-265.

[0133] In some embodiments, the nuclease comprises SEQ ID NO: 5, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 266-268. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 5, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 266-268.

[0134] In some embodiments, the nuclease comprises SEQ ID NO: 6, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 269-271. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 6, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 269-271.

[0135] In some embodiments, the nuclease comprises SEQ ID NO: 7, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 272-274. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 7, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 272-274.

[0136] In some embodiments, the nuclease comprises SEQ ID NO: 8, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 275-277. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 8, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 275-277.

[0137] In some embodiments, the nuclease comprises SEQ ID NO: 9, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 278-280. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 9, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 278-280.

[0138] In some embodiments, the nuclease comprises SEQ ID NO: 10, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 281-283. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 10, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 281-283.

[0139] In some embodiments, the nuclease comprises SEQ ID NO: 11, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 284-286. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 11, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 284-286.

[0140] In some embodiments, the nuclease comprises SEQ ID NO: 12, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 287-289. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 12, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 287-289.

[0141] In some embodiments, the nuclease comprises SEQ ID NO: 13, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 290-292. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 13, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 290-292.

[0142] In some embodiments, the nuclease comprises SEQ ID NO: 14, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 293-295. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 14, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 293-295.

[0143] In some embodiments, the nuclease comprises SEQ ID NO: 15, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 296-298. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 15, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 296-298.

[0144] In some embodiments, the nuclease comprises SEQ ID NO: 16, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 299-301. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 16, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 299-301.

[0145] In some embodiments, the nuclease comprises SEQ ID NO: 17, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 302-304. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 17, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 302-304.

[0146] In some embodiments, the nuclease comprises SEQ ID NO: 18, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 305-307. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 18, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 305-307.

[0147] In some embodiments, the nuclease comprises SEQ ID NO: 19, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NO: 308 or 379. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 19, and wherein the at least one gRNA comprises any one of SEQ ID NO: 308 or 379.

[0148] In some embodiments, the nuclease comprises SEQ ID NO:20, and at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs:309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479, and 481. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 20, and wherein at least one gRNA comprises any one of SEQ ID NO: 309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479, and 481. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 20, and wherein at least one gRNA comprises any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392, and 417, or any one of SEQ ID NOs: 346 and 362, or any one of SEQ ID NOs: 410-419. In some embodiments, the nuclease comprises SEQ ID NO:20, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:20, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346.

[0149] In some embodiments, the nuclease comprises SEQ ID NO: 22, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 311, 346, 381, and 398-399. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381, and 398-399. In some embodiments, the nuclease comprises SEQ ID NO:22, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:22, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346.

[0150] In some embodiments, the nuclease comprises SEQ ID NO: 23, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 23, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises SEQ ID NO:23, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:23, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical to SEQ ID NO:346.

[0151] In some embodiments, the nuclease comprises SEQ ID NO: 25, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 25, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises SEQ ID NO:25, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:25, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346.

[0152] In some embodiments, the nuclease comprises SEQ ID NO: 26, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:26, and wherein at least one gRNA comprises any one of SEQ ID NO:346, 384, and 392.

[0153] In some embodiments, the nuclease comprises SEQ ID NO:26, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:26, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346.

[0154] In some embodiments, the nuclease comprises SEQ ID NO: 27, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises SEQ ID NO:27, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:27, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346.

[0155] In some embodiments, the nuclease comprises SEQ ID NO: 28, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 28, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises SEQ ID NO:28, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:28, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346.

[0156] In some embodiments, the nuclease comprises SEQ ID NO: 29, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity, or 100% identity to any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises SEQ ID NO:29, or a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:29, and the gRNA comprises SEQ ID NO:346, or a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO:346.

[0157] In some embodiments, the nuclease comprises SEQ ID NO: 30, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 319. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 30, and wherein the at least one gRNA comprises SEQ ID NO: 319.

[0158] In some embodiments, the nuclease comprises SEQ ID NO: 31 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 320. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 31, and wherein the at least one gRNA comprises SEQ ID NO: 320.

[0159] In some embodiments, the nuclease comprises SEQ ID NO: 32, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 321. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 32, and wherein the at least one gRNA comprises SEQ ID NO: 321.

[0160] In some embodiments, the nuclease comprises SEQ ID NO: 33, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 322. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 33, and wherein the at least one gRNA comprises SEQ ID NO: 322.

[0161] In some embodiments, the nuclease comprises SEQ ID NO: 34, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NO: 323 or 388. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 34, and wherein the at least one gRNA comprises any one of SEQ ID NO: 323 or 388.

[0162] In some embodiments, the nuclease comprises SEQ ID NO: 35, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 324. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 35, and wherein the at least one gRNA comprises SEQ ID NO: 324.

[0163] In some embodiments, the nuclease comprises SEQ ID NO: 37, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 326. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 37, and wherein the at least one gRNA comprises SEQ ID NO: 326.

[0164] In some embodiments, the nuclease comprises SEQ ID NO: 38, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 327. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 38, and wherein the at least one gRNA comprises SEQ ID NO: 327.

[0165] In some embodiments, the nuclease comprises SEQ ID NO: 39, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 328. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 39, and wherein the at least one gRNA comprises SEQ ID NO: 328.

[0166] In some embodiments, the nuclease comprises SEQ ID NO: 40, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 329. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 40, and wherein the at least one gRNA comprises SEQ ID NO: 329.

[0167] In some embodiments, the nuclease comprises SEQ ID NO: 41, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 330. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 41, and wherein the at least one gRNA comprises SEQ ID NO: 330.

[0168] In some embodiments, the nuclease comprises SEQ ID NO: 42, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 331. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 42, and wherein the at least one gRNA comprises SEQ ID NO: 331.

[0169] In some embodiments, the nuclease comprises SEQ ID NO: 43, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 332. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 43, and wherein the at least one gRNA comprises SEQ ID NO: 332.

[0170] In some embodiments, the nuclease comprises SEQ ID NO: 44, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 333. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 44, and wherein the at least one gRNA comprises SEQ ID NO: 333.

[0171] In some embodiments, the nuclease comprises SEQ ID NO: 45, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 334. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 45, and wherein the at least one gRNA comprises SEQ ID NO: 334.

[0172] In some embodiments, the nuclease comprises SEQ ID NO: 46, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 335. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 46, and wherein the at least one gRNA comprises SEQ ID NO: 335.

[0173] In some embodiments, the nuclease comprises SEQ ID NO: 47, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 336. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 47, and wherein the at least one gRNA comprises SEQ ID NO: 336.

[0174] In some embodiments, the nuclease comprises SEQ ID NO: 48, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 337. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 48, and wherein the at least one gRNA comprises SEQ ID NO: 337.

[0175] In some embodiments, the nuclease comprises SEQ ID NO: 49, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 338. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49, and wherein the at least one gRNA comprises SEQ ID NO: 338.

[0176] In some embodiments, the nuclease comprises SEQ ID NO: 50, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 339. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 50, and wherein the at least one gRNA comprises SEQ ID NO: 339.

[0177] In some embodiments, the nuclease comprises SEQ ID NO: 51, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 340. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 51, and wherein the at least one gRNA comprises SEQ ID NO: 340.

[0178] In some embodiments, the nuclease comprises SEQ ID NO: 52, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 341. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 52, and wherein the at least one gRNA comprises SEQ ID NO: 341.

[0179] In some embodiments, the nuclease comprises SEQ ID NO: 53, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 342. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 53, and wherein the at least one gRNA comprises SEQ ID NO: 342.

[0180] In some embodiments, the nuclease comprises SEQ ID NO: 54, and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 343. In some embodiments, the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 54, and wherein the at least one gRNA comprises SEQ ID NO: 343.

[0181] In some embodiments, the nuclease comprises any one of SEQ ID NOs: 1-19 and 30-54, or a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to any one of SEQ ID NOs: 1-19 and 30-54, and the gRNA comprises SEQ ID NO: 346, or a sequence at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to SEQ ID NO: 346.

[0182] In some embodiments, the gRNA described herein may comprise one or more nucleotide substitutions or mutations (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, etc.) relative to any one of SEQ ID NOs: 251-343.

[0183] In some embodiments, relative to any one of SEQ ID NO: 251-343, gRNA comprises one or more truncations or deletions of one or more nucleotides. The truncation or deletion may be located at one or both of the 3' end and the 5' end of the sequence, or located within or inside a sequence associated with any one of SEQ ID NO: 251-343. The truncation or deletion may comprise a single nucleotide or may comprise a series of two or more consecutive nucleotides (e.g., 2, 3, 4, 5, 10, 15, 20, etc.) deletions or truncations. In some embodiments, the gRNA of the present invention may comprise a truncated sequence corresponding to or estimated as a crRNA:tracrRNA stem.

[0184] In some embodiments, the gRNA comprises a tracr sequence. The gRNA may comprise one or more sequence deletions in or near the region comprising the tracr sequence. For example, one or more sequence deletions may comprise a sequence predicted to form a stem-loop structure. In some embodiments, one or more sequence deletions include a sequence predicted to form a stem-loop structure at or near the 5' end of the gRNA. In some embodiments, the gRNA comprises SEQ ID NO: 346. In some embodiments, the gRNA comprises SEQ ID NO: 420. In some embodiments, the gRNA comprises SEQ ID NO: 481. In some embodiments, the gRNA comprises SEQ ID NO: 479.

[0185] In some embodiments, relative to any one of SEQ ID NO: 251-343, gRNA comprises one or more insertions or additions of one or more nucleotides. The insertion or addition may be located at one or both of the 3' end and the 5' end of the sequence, or located within a sequence associated with any one of SEQ ID NO: 251-343. The insertion or addition may comprise a single nucleotide or may comprise a series of two or more consecutive nucleotides (e.g., 2, 3, 4, 5, 10, 15, 20, etc.) deletions or truncations. In some embodiments, the gRNA of the present invention may comprise an artificial stem-loop between crRNA and tracrRNA.

[0186] The gRNA may be a non-naturally occurring gRNA.

[0187] In certain embodiments, engineering of nucleases for eukaryotic cells may include codon optimization. It should be appreciated that changing natural codons to those codons most commonly used in mammals allows for maximum expression of system proteins in mammalian cells (e.g., human cells). Such modified nucleic acid sequences are generally described in the art as "codon optimized," or as utilizing "mammal preferred" or "human preferred" codons. In some embodiments, if at least about 60% (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 98%) of the codons encoded in the nucleic acid sequence are mammal preferred codons, the nucleic acid sequence is considered to be codon optimized.

[0188] In some cases, the compositions or systems disclosed herein may also include donor polynucleotides. For example, in applications where it is desired to insert a polynucleotide sequence into a genome where a target sequence is cleaved, a donor polynucleotide (a nucleic acid comprising a donor sequence) may also be provided to a cell. The so-called "donor sequence" or "donor polynucleotide" or "donor template" means a nucleic acid sequence at a site targeted by a nuclease to be inserted (e.g., after dsDNA cleavage, after cutting the target DNA, after double cutting the target DNA, etc.). In some cases, the donor sequence is provided to the cell as a single-stranded DNA. In some cases, the donor template is provided to the cell as a double-stranded DNA. It can be introduced into the cell in a linear or circular form. If introduced in a linear form, the ends of the donor sequence can be protected by any convenient method (e.g., from exonucleolytic degradation), and such methods are known to those skilled in the art. For example, one or more dideoxynucleotide residues may be added to the 3' end of a linear molecule and / or a self-complementary oligonucleotide may be connected to one or both ends. The donor template may be introduced into the cell as part of a vector molecule having additional sequences such as an origin of replication, a promoter, and genes encoding antibiotic resistance. Furthermore, the donor template can be introduced as naked nucleic acid, as nucleic acid complexed with an agent such as a liposome or poloxamer, or can be delivered by virus (eg, adenovirus, AAV).

[0189] The present disclosure also provides one or more nucleic acids encoding the nucleases and gRNA disclosed herein, vectors containing these nucleic acids, and cells containing these vectors. The vector can be used to propagate the fragments in appropriate cells and / or allow expression from the fragments (e.g., expression vectors). Those of ordinary skill in the art will know various vectors that can be used for the propagation and expression of nucleic acid sequences.

[0190] In some embodiments, one or more nucleic acids include one or more messenger RNAs, one or more vectors, or any combination thereof. In some embodiments, one or more nucleic acids include messenger RNAs for expressing nucleases, and at least one nucleic acid provides gRNA. A single nucleic acid can encode a nuclease and at least one gRNA, or a nuclease can be encoded on a nucleic acid different from at least one gRNA.

[0191] In some embodiments, nuclease is provided in the form of cleavage nuclease (for example, in some cases, nuclease can be delivered as cleavage nuclease or nucleic acid encoding cleavage nuclease), so that two separate proteins form functional nuclease together. In some such cases, the sequences encoding the two parts of the split nuclease protein are present on the same carrier. In some cases, they are present on a separate carrier, for example, as a part of a carrier system encoding nuclease, gRNA and its system.

[0192] The present disclosure also provides engineered, non-naturally occurring vectors and vector systems that can encode one or more or all of the components of the system of the present invention.The vector can be introduced into a cell capable of expressing the polypeptide encoded therein, including any suitable prokaryotic or eukaryotic cell.

[0193] The vectors of the present disclosure can be delivered to eukaryotic cells of a subject, such as a mammalian subject, such as a human subject. Modification of eukaryotic cells via the system of the present invention can be performed in cell culture.

[0194] Viral and non-viral gene transfer methods can be used to introduce nucleic acids encoding components of the system of the present invention into cells, tissues or subjects. Such methods can be used to apply nucleic acids encoding components of the system of the present invention to cells in culture or in host organisms. Non-viral vector delivery systems include DNA plasmids, cosmids, RNA (e.g., transcripts of vectors described herein), nucleic acids and nucleic acids compounded with delivery vehicles. Viral vector delivery systems include DNA and RNA viruses, which have additional or integrated genomes after being delivered to cells. Viral vectors include, for example, retroviruses, slow viruses, adenoviruses, adeno-associated viruses and herpes simplex virus vectors.

[0195] In certain embodiments, non-replicating plasmids or plasmids that can be cured by high temperature can be used so that any or all essential components of the composition or system can be removed from the cell under certain conditions. For example, this can allow DNA integration by transforming the bacteria of interest, leaving an engineered strain that does not have a memory of the plasmid or vector used for integration.

[0196] A variety of viral constructs can be used to deliver the compositions or systems of the present invention (such as nucleases and one or more gRNAs) to target cells and / or subjects. Non-limiting examples of such recombinant viruses include recombinant adeno-associated virus (AAV), recombinant adenovirus, recombinant lentivirus, recombinant retrovirus, recombinant herpes simplex virus, recombinant poxvirus, bacteriophage, etc. The present disclosure provides vectors that can be integrated into the host genome, such as retroviruses or lentiviruses. See, for example, Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1989; Kay, MA et al., 2001 Nat. Medic. 7 (1): 33-40; and Walther W. and Stein U., 2000 Drugs, 60 (2): 249-71, which are incorporated herein by reference.

[0197] In one embodiment, the DNA fragment encoding the nuclease is contained in a plasmid vector that allows the expression of the protein and subsequent isolation and purification of the protein produced by the recombinant vector. Thus, the nuclease disclosed herein can be purified after expression, obtained by chemical synthesis, or obtained by recombinant methods.

[0198] In order to construct cells expressing the system of the present invention, an expression vector for stable or transient expression of the system or any component thereof can be constructed as described herein or by methods known in the art, and introduced into cells. For example, the nucleic acid encoding the components of the system of the present invention can be cloned into a suitable expression vector, such as a plasmid or viral vector operably connected to a suitable promoter. The selection of expression vector / plasmid / viral vector should be suitable for integration and replication in eukaryotic cells. In some embodiments, a single nucleic acid comprises a first promoter operably connected to a nuclease and a second promoter operably connected to a gRNA. In some cases, a single nucleic acid is a vector.

[0199] In certain embodiments, one or more promoters can drive expression of one or more sequences (e.g., nucleases and / or gRNAs) in prokaryotes. Useful promoters include T7 RNA polymerase promoters, constitutive E. coli promoters, and promoters that are widely recognized by the transcriptional machinery in a wide range of bacterial organisms. The composition or system can be used with a variety of bacterial hosts.

[0200] In certain embodiments, one or more promoters can drive expression of one or more sequences (e.g., nucleases and / or gRNAs) in mammalian cells, such as when contained in a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, Nature (1987) 329: 840, which is incorporated herein by reference) and pMT2PC (Kaufman, et al., EMBO J. (1987) 6: 187, which is incorporated herein by reference). When used in mammalian cells, the control functions of the expression vector are generally provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma virus, adenovirus 2, cytomegalovirus, simian virus 40, and other promoters disclosed herein and known in the art. For other suitable expression systems for prokaryotic and eukaryotic cells, see, e.g., Sambrook et al., MOLECULAR CLONING: A LABORATORY MANUAL., 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989, Chapters 16 and 17, which are incorporated herein by reference.

[0201] Promoters for expressing the nucleases and gRNAs herein may include any of a variety of promoters known in the art, wherein the promoter is constitutive, regulated or inducible, cell type specific, tissue specific or species specific. In addition to sequences sufficient to direct transcription, promoter sequences of the present invention may also include sequences of other regulatory elements involved in regulating transcription (e.g., enhancers, Kozak sequences and introns). Many promoter / regulatory sequences that can be used to drive constitutive expression of genes are available in the art, including, but not limited to, for example, CMV (cytomegalovirus promoter), EF1a (human elongation factor 1α promoter), SV40 (simian vacuolating virus 40 promoter), PGK (mammalian phosphoglycerate kinase promoter), Ubc (human ubiquitin C promoter), human β-actin promoter, rodent β-actin promoter, CBh (chicken β-actin promoter), CAG (hybrid promoter containing CMV enhancer, chicken β-actin promoter and rabbit β-globin splicing acceptor), TRE (tetracycline response element promoter), H1 (human polymerase III RNA promoter), U6 (human U6 small nuclear promoter), etc. Other promoters that can be used to express the components of the system of the present invention include, but are not limited to, cytomegalovirus (CMV) intermediate early promoter, viral LTRs such as Rous sarcoma virus LTRs, HIV-LTRs, HTLV-1LTRs, Moloney murine leukemia virus (MMLV) LTRs, myeloproliferative sarcoma virus (MPSV) LTRs, spleen focus forming virus (SFFV) LTRs, simian virus 40 (SV40) early promoters, herpes simplex virus tk virus promoters, elongation factor 1-α (EF1-α) promoters with or without EF1-α introns. Additional promoters include any constitutively active promoters. Alternatively, any regulatable promoter can be used so that its expression can be regulated in the cell. In an embodiment, a polymerase II promoter is used to drive the expression of a nuclease (e.g., a CMV promoter), and a polymerase III promoter (e.g., a U6 promoter) is used to drive the expression of a gRNA.

[0202] Different promoters and regulatory elements can be used to achieve the appropriate balance (expression level ratio) between the components of the system (e.g., nuclease, at least one gRNA). For example, in some cases, the nucleic acid comprises a promoter and a regulatory element operably connected to the sequence encoding the nuclease (and thus regulates / adjusts its translation). In some cases, the subject nucleic acid comprises a promoter and a regulatory element operably connected to the sequence encoding the gRNA. In some cases, the sequence encoding the nuclease and the sequence encoding the gRNA are both operably connected to the same promoter and regulatory element.

[0203] A variety of promoter types are suitable for use. A promoter can be a constitutively active promoter (e.g., a promoter having a constitutively active / "ON" state), it can be an inducible promoter (e.g., a promoter whose activity / "ON" or inactive / "OFF" state is controlled by an external stimulus (e.g., the presence of a specific temperature, compound, or protein), it can be a spatially restricted promoter (e.g., a tissue-specific promoter, a cell type-specific promoter, etc.), and it can be a temporally restricted promoter (e.g., a promoter that is in an "ON" state or an "OFF" state during a specific stage of embryonic development or during a specific stage of a biological process (e.g., the hair follicle cycle of a mouse).

[0204] In addition, the inducible and tissue-specific expression of RNA or protein can be achieved by placing the nucleic acid encoding such a molecule under the control of an inducible or tissue-specific promoter / regulatory sequence. The promoter can guide the expression of nucleic acids in specific cell types (for example, tissue-specific regulatory elements are used to express nucleic acids). Such regulatory elements include promoters that can be tissue-specific or cell-specific. When applied to promoters, the term "tissue-specific" refers to a promoter that can guide the selective expression of a nucleotide sequence of interest to a specific type of tissue (for example, seed) when the same nucleotide sequence of interest is not expressed relatively in different types of tissues. The term "cell type specificity" as applied to promoters refers to a promoter that can guide the selective expression of a nucleotide sequence of interest in a specific type of cell when the same nucleotide sequence of interest is not expressed relatively in different types of cells in the same tissue. When applied to promoters, the term "cell type specificity" also means a promoter that can promote the selective expression of a nucleotide sequence of interest in a region within a single tissue. The cell type specificity of a promoter can be assessed using methods well known in the art such as immunohistochemical staining.

[0205] Examples of tissue-specific or inducible promoters / regulatory sequences that can be used for this purpose include, but are not limited to, the rhodopsin promoter, the MMTV LTR inducible promoter, the SV40 late enhancer / promoter, the synapsin 1 promoter, the ET hepatocyte promoter, the GS glutamine synthase promoter, and many other promoters. Various commercially available ubiquitous promoters as well as tissue-specific promoters and tumor-specific promoters can be purchased from, for example, InvivoGen. In addition, promoters well known in the art can be induced in response to inducing agents such as metals, glucocorticoids, tetracyclines, hormones, etc., and are also contemplated for use in the present invention. Therefore, it should be understood that the present disclosure includes the use of any promoter / regulatory sequence known in the art that is capable of driving the expression of a desired nuclease or gRNA operably linked thereto.

[0206] Examples of spatially restricted promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, photoreceptor-specific promoters, etc. Neuron-specific spatially restricted promoters include, but are not limited to, neuron-specific enolase (NSE) promoter (see, e.g., EMBLHSENO2, X51956); aromatic amino acid decarboxylase (AADC) promoter; neurofilament promoter (see, e.g., GenBank HUMNFL, L04147); synapsin promoter (see, e.g., GenBank HUMSYNIB, M55301); thy-1 promoter; serotonin receptor promoter (see, e.g., GenBank S62283); tyrosine hydroxylase promoter (TH); GnRH promoter; L7 promoter; DNMT promoter; enkephalin; myelin basic protein (MBP) promoter; Ca2+ calmodulin-dependent protein kinase II-α (CamKIIα) promoter; CMV enhancer / platelet-derived growth factor-β promoter, etc. In some cases, suitable liver-specific promoters may include, but are not limited to, TTR, albumin, and AAT promoters. In some cases, suitable CNS-specific promoters may include, but are not limited to, synaptophysin 1, BM88, CHNRB2, GFAP, and CAMK2a promoters. In some cases, suitable muscle-specific promoters may include, but are not limited to, MYOD1, MYLK2, SPc5-12 (synthetic), α-MHC, MLC-2, MCK, MHCK7, human cardiac troponin C (cTnC), and desmin promoters. Adipocyte-specific spatially restricted promoters include, but are not limited to, aP2 gene promoter / enhancers, such as -5.4kb to +21bp regions of human aP2; glucose transporter 4 (GLUT4); fatty acid translocase (FAT / CD36) promoter; stearoyl-CoA desaturase 1 (SCD1) promoter; leptin promoter; adiponectin promoter; lipopolysaccharide promoter; resistin promoter, etc. Cardiomyocyte-specific spatially restricted promoters include, but are not limited to, control sequences derived from the following genes: myosin light chain-2, α-myosin heavy chain, AE3, cardiac troponin C, cardiac actin, etc. Smooth muscle-specific spatially restricted promoters include, but are not limited to, SM22α promoter; smoothelin promoter; α-smooth muscle actin promoter, etc. For example, the 0.4kb region of the SM22α promoter (in which there are two CArG elements) has been shown to mediate vascular smooth muscle cell specificity. Photoreceptor-specific spatially restricted promoters include, but are not limited to, rhodopsin promoter; rhodopsin kinase promoter; β-phosphodiesterase gene; retinitis pigmentosa gene promoter; inter-photoreceptor retinoid binding protein (IRBP) gene enhancer; IRBP gene promoter, etc.

[0207] Examples of inducible promoters include, but are not limited to, heat shock promoters, tetracycline-regulated promoters, steroid-regulated promoters, metal-regulated promoters, estrogen receptor-regulated promoters, etc. Thus, inducible promoters can be regulated by molecules including, but not limited to, doxycycline, estrogen receptors, estrogen receptor fusions, estrogen analogs, IPTG, etc. Inducible promoters suitable for use include any inducible promoter described herein or known to those of ordinary skill in the art. Examples of inducible promoters include, but are not limited to, chemically / biochemically regulated and physically regulated promoters, such as alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc) responsive promoters and other tetracycline-responsive promoter systems, including tetracycline repressor protein (tetR), tetracycline operator sequence (tetO), and tetracycline transactivator fusion protein (tTA)), steroid-regulated promoters (e.g., promoters based on rat glucocorticoid receptor, human estrogen receptor, moth ecdysone receptor, and promoters from the steroid / retinoid / thyroid receptor superfamily), metal-regulated promoters (e.g., promoters from metallothionein (proteins that bind and chelate metal ions) genes from yeast, mouse, and human), pathogenesis-regulated promoters (e.g., induced by salicylic acid, ethylene, or benzothiadiazole (BTH)), temperature / heat-inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light-responsive promoters from plant cells).

[0208] Inducible promoters include sugar inducible promoters (e.g., lactose inducible promoter; arabinose inducible promoter); amino acid inducible promoter; alcohol inducible promoter, etc. Suitable promoters include, for example, lactose regulatory system (e.g., lactose operator system, sugar regulatory system, isopropyl-β-D-thiogalactoside (IPTG) inducible system, arabinose regulatory system (e.g., arabinose operator system, such as ARA operator promoter, pBAD, pARA, parts thereof, combinations thereof, etc.)), synthetic amino acid regulatory system, fructose repressor, tac promoter / operator (pTac), tryptophan promoter, PhoA promoter, recA promoter, proU promoter, cst-1 promoter, tetA promoter, cadA promoter, nar promoter, P LIn some cases, the promoter comprises Lac-Z or a portion thereof. In some cases, the promoter comprises a Lac operon or a portion thereof. In some cases, the inducible promoter comprises an ARA operon promoter or a portion thereof. In certain embodiments, the inducible promoter comprises an arabinose promoter or a portion thereof. The arabinose promoter can be obtained from any suitable bacteria. In some cases, the inducible promoter comprises an arabinose operon of Escherichia coli or Bacillus subtilis. In some cases, the inducible promoter is activated by the presence of sugar or its analogs. Non-limiting examples of sugars and sugar analogs include lactose, arabinose (e.g., L-arabinose), glucose, sucrose, fructose, IPTG, etc. Suitable promoters include T7 promoters; pBAD promoters; lacIQ promoters, etc. In some cases, the promoter is a J23119 promoter. Many bacterial promoters are known in the art; bacterial promoters can be found on the Internet via parts(dot)igem(dot)org / promoters.

[0209] In some cases, the promoter is a reversible promoter. Suitable reversible promoters are known in the art, including reversible inducible promoters. Such reversible promoters can be separated and derived from many organisms. Such reversible promoters can be separated and derived from many organisms (e.g., eukaryotes and prokaryotes). Modification of a reversible promoter derived from a first organism for a second organism is well known in the art. Modification of a reversible promoter derived from a first organism (e.g., a first prokaryote and a second eukaryote, a first eukaryote and a second prokaryote, etc.) for a second organism is well known in the art. Such reversible promoters and systems based on such reversible promoters but further comprising additional control proteins include, but are not limited to, alcohol-regulated promoters (e.g., alcohol dehydrogenase I (alcA) gene promoter, promoters responsive to alcohol transactivator (AlcR)), tetracycline-regulated promoters (e.g., promoter systems including TetActivators, TetON, TetOFF), steroid-regulated promoters (e.g., rat glucocorticoid receptor promoter system, human estrogen receptor promoter system, retinoid promoter system, thyroid promoter system, ecdysone promoter system, mifepristone promoter system), metal-regulated promoters (e.g., metallothionein promoter system), pathogenesis-related regulated promoters (e.g., salicylic acid-regulated promoter, ethylene-regulated promoter, benzothiadiazole-regulated promoter), temperature-regulated promoters (e.g., heat shock-inducible promoters (e.g., HSP-70, HSP-90, soybean heat shock promoter)), light-regulated promoters, synthetic inducible promoters, and the like.

[0210] Thus, it should be understood that the present disclosure encompasses any promoter / regulatory sequence capable of driving expression of a desired nuclease or RNA to which it is operably linked.

[0211] In addition, the vectors described herein for expressing nucleases and / or gRNAs may include, for example, some or all of the following: a selectable marker gene, such as the neomycin gene for selecting stable or transient transfectants in host cells; an enhancer / promoter sequence from the immediate early gene of human CMV for high-level transcription; transcription termination and RNA processing signals from SV40 for mRNA stability; 5'- and 3'-untranslated regions for mRNA stability and translation efficiency from highly expressed genes such as α-globin or β-globin; the SV40 polyoma virus replication origin and ColE1 for appropriate episomal replication; internal ribosome binding sites (IRESes), a universal multiple cloning site; T7 and SP6 RNA promoters for in vitro transcription of sense and antisense RNA; a "suicide switch" or "suicide gene" that, when triggered, causes the death of cells carrying the vector (e.g., HSV thymidine kinase, inducible caspases, such as iCasp9), and a reporter gene for evaluating chimeric receptor expression. Suitable vectors and methods for producing vectors containing transgenes are well known and available in the art. Selectable markers also include chloramphenicol resistance, tetracycline resistance, spectinomycin resistance, streptomycin resistance, erythromycin resistance, rifampicin resistance, bleomycin resistance, heat-adapted kanamycin resistance, gentamicin resistance, hygromycin resistance, trimethoprim resistance, dihydrofolate reductase (DHFR), GPT; URA3, HIS4, LEU2 and TRP1 genes of Saccharomyces cerevisiae.

[0212] When introduced into a cell, the vector may be maintained as an autonomously replicating sequence or extrachromosomal element, or may be integrated into the host DNA.

[0213] The compositions and systems of the invention (e.g., proteins, polynucleotides encoding these proteins, or compositions comprising proteins and / or polynucleotides described herein) can be delivered by any suitable means. In certain embodiments, the compositions or systems are delivered in vivo. In other embodiments, the compositions or systems are delivered in vitro to isolated / cultured cells (e.g., autologous iPS cells).

[0214] Vectors and nucleic acids according to the present disclosure can be transformed, transfected or otherwise introduced into a variety of host cells. Transfection refers to the host cell uptake of nucleic acid, whether or not any coding sequence is expressed in fact. Many transfection methods are known to those of ordinary skill in the art, such as lipofectamine, calcium phosphate coprecipitation, electroporation, DEAE-dextran treatment, microinjection, viral infection and other methods known in the art. Transduction refers to the virus entering the cell and expressing (e.g., transcribing and / or translating) the sequence delivered by the viral vector genome. With respect to recombinant vectors, "transduction" generally refers to the recombinant viral vector entering the cell and expressing the nucleic acid of interest delivered by the vector genome.

[0215] Any vector comprising the nucleic acid sequence of the component of the coding composition and system of the present invention is also within the scope of the present disclosure.Such a vector can be delivered to the host cell by a suitable method.The method of delivering the vector to the cell is well known in the art, and may include DNA or RNA electroporation, transfection reagents such as liposomes or nanoparticles to deliver DNA or RNA, deliver DNA, RNA or protein or viral transduction by mechanical deformation.In some embodiments, the vector is delivered to the host cell by viral transduction.Nucleic acid can be delivered as a part of a larger construct (such as a plasmid or viral vector), or directly delivered, for example, delivered by electroporation, lipid vesicles, viral transporters, microinjection and biolistics (biolistics) (high-speed particle bombardment).Similarly, the construct containing one or more transgenics can be delivered by any method suitable for introducing nucleic acid into the cell.

[0216] In addition, delivery vehicles such as mRNA or protein delivery systems based on nanoparticles and lipids can be used. Other examples of delivery vehicles include lentiviral vectors, ribonucleoprotein (RNP) complexes, lipid-based delivery systems, gene guns, fluid mechanics, electroporation or nuclear transfection microinjection, bio-ballistics, etc.

[0217] In some embodiments, the vector is a viral construct, such as a recombinant adeno-associated virus construct, a recombinant adenovirus construct, a recombinant lentivirus construct, a recombinant retrovirus construct, etc. Suitable viral vectors include, but are not limited to, vaccinia virus-based viral vectors; poliovirus; adenovirus; adeno-associated virus; SV40; herpes simplex virus; human immunodeficiency virus; retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses, such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.

[0218] In some embodiments, the vector is an AAV vector. By "adeno-associated virus" or "AAV" is meant the virus itself or a derivative thereof. Unless otherwise required, the term encompasses all subtypes and naturally occurring and recombinant forms, such as AAV type 1 (AAV-1), AAV type 2 (AAV-2), AAV type 3 (AAV-3), AAV type 4 (AAV-4), AAV type 5 (AAV-5), AAV type 6 (AAV-6), AAV type 7 (AAV-7), AAV type 8 (AAV-8), AAV type 9 (AAV-9), AAV AAV-10, AAV-11, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, ovine AAV, hybrid AAV (i.e., AAV comprising a capsid protein of one AAV subtype and genomic material of another subtype), AAV comprising a mutant AAV capsid protein or a chimeric AAV capsid (i.e., a capsid protein having regions or domains or single amino acids derived from two or more different serotypes of AAV, such as AAV-DJ, AAV-LK3, AAV-LK19). "Primate AAV" refers to AAV that infects primates, "non-primate AAV" refers to AAV that infects non-primate mammals, "bovine AAV" refers to AAV that infects bovine mammals, and the like.

[0219] The so-called "recombinant AAV vector" or "rAAV vector" means an AAV virus or AAV viral chromosomal material, which contains a polynucleotide sequence that is not derived from AAV (e.g., a polynucleotide heterologous to AAV), usually a nucleic acid sequence of interest to be integrated into a cell according to the method of the present invention. Generally speaking, the heterologous polynucleotide is flanked by at least one, usually two, AAV inverted terminal repeats (ITRs). In some cases, the recombinant viral vector also contains viral genes that are important for the packaging of the recombinant viral vector material. Packaging refers to a series of intracellular events that lead to the assembly and encapsulation of viral particles (e.g., AAV viral particles). Examples of nucleic acid sequences that are important for AAV packaging include the AAV "rep" and "cap" genes, which encode replication and encapsulation proteins of adeno-associated viruses, respectively. The term rAAV vector includes rAAV vector particles and rAAV vector plasmids.

[0220] "Viral particle" refers to a single unit of a virus containing a capsid encapsulating a virus-based polynucleotide, such as a viral genome (as in a wild-type virus), or, for example, a subject targeting vector (as in a recombinant virus). An AAV viral particle refers to a viral particle consisting of at least one AAV capsid protein (usually all capsid proteins of a wild-type AAV) and an encapsulated polynucleotide AAV vector. If the particle contains a heterologous polynucleotide (e.g., a polynucleotide different from the wild-type AAV genome, such as a transgene to be delivered to a mammalian cell), it is generally referred to as a "rAAV vector particle" or simply "rAAV vector". Therefore, the production of rAAV particles necessarily includes the production of rAAV vectors, because such vectors are contained within rAAV particles.

[0221] rAAV virions can be constructed by a variety of methods. For example, heterologous sequences can be directly inserted into the AAV genome, which has a major AAV open reading frame ("ORF") excised therefrom. Other parts of the AAV genome can also be deleted, as long as enough ITR parts are retained to allow replication and packaging functions. In order to produce rAAV virions, known techniques such as AAV expression vectors can be introduced into suitable host cells by transfection. Particularly suitable transfection methods include calcium phosphate co-injection, direct microinjection into cultured cells, electroporation, liposome-mediated gene transfer, lipid-mediated transduction, and nucleic acid delivery using high-speed microparticles. Suitable cells for producing rAAV virions include microorganisms, yeast cells, insect cells, and mammalian cells, which can be used as or have been used as receptors for heterologous DNA molecules.

[0222] The AAV virus produced can be replication-competent or non-replication-competent. "Replication-competent" virus (e.g., replication-competent AAV) refers to a phenotypic wild-type virus that is infective and can also replicate in infected cells (e.g., in the presence of a helper virus or helper virus function). With respect to AAV, replication ability generally requires the presence of functional AAV packaging genes. Generally speaking, due to the lack of one or more AAV packaging genes, rAAV vectors as described herein are non-replication-competent in mammalian cells (especially in human cells). Typically, such rAAV vectors lack any AAV packaging gene sequences to minimize the possibility of producing replication-competent AAVs by recombination between AAV packaging genes and introduced rAAV vectors.

[0223] Retroviruses such as lentiviruses are suitable for use in the methods disclosed herein. Commonly used retroviral vectors cannot produce the viral proteins required for proliferative infection. Moreover, vector replication requires growth in a packaging cell line. In order to produce viral particles containing a nucleic acid of interest, a retroviral nucleic acid containing the nucleic acid is packaged into a viral capsid by a packaging cell line. Different packaging cell lines provide different envelope proteins (ecotropic, amphotropic or heterotropic) to be integrated into the capsid, which determines the specificity of the viral particles to the cell (ecotropic for mice and rats; amphotropic for most mammalian cell types including humans, dogs and mice; and heterotropic for most mammalian cell types except mouse cells). Appropriate packaging cell lines can be used to ensure that cells are targeted by packaged viral particles. Methods for introducing a subject vector expression vector into a packaging cell line and collecting viral particles produced by a packaging cell line are well known in the art. Nucleic acids can also be introduced by direct microinjection (e.g., injection of RNA).

[0224] As described elsewhere herein, proteins can be provided to cells as can RNA (e.g., RNA comprising translational control elements as discussed elsewhere herein). Methods for introducing RNA into cells may include, for example, direct injection, transfection, or any other method for introducing DNA. Nucleases can also be introduced directly into host cells as proteins. In this case, the nuclease can be delivered as an RNP (ribonucleoprotein complex) in which it has been complexed with an appropriate guide RNA.

[0225] Any convenient method can be used to deliver the disclosed nucleic acids (e.g., vectors) and proteins to cells. Suitable methods include, for example, viral infection (e.g., AAV, adenovirus, slow virus), transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc.

[0226] In some cases, the nuclease is delivered to the cell in or associated with the particle. In some cases, the nuclease is delivered with a cationic lipid and a hydrophilic polymer, for example, wherein the cationic lipid comprises 1,2-dioleoyl-3-trimethylammonium-propane (DOTAP) or 1,2-dimyristoyl-sn-glycero-3-phosphocholine (DMPC), and / or wherein the hydrophilic polymer comprises ethylene glycol or polyethylene glycol (PEG); and / or wherein the particle also comprises cholesterol.

[0227] Nuclease can be delivered using particles or lipid envelopes. For example, biodegradable core-shell structured nanoparticles with poly (β-amino ester) (PBAE) cores encapsulated by a phospholipid bilayer shell can be used. In some cases, particles / nanoparticles based on self-assembling bioadhesive polymers are used; such particles / nanoparticles can be applied to oral delivery of peptides, intravenous delivery of peptides, and nasal delivery of peptides, for example, to the brain. Other embodiments are also envisioned, such as oral absorption and ocular delivery of hydrophobic drugs. Molecular envelope technology can be used, which involves an engineered polymer envelope that is protected and delivered to the desired cells.

[0228] Lipidoid compounds (e.g., as described in U.S. Patent Application Publication No. 2011 / 0293703) can also be used to deliver polynucleotides, and can be used to deliver the disclosed nucleases (or RNA or DNA encoding them). In one aspect, the amino alcohol lipidoid compound is combined with the agent to be delivered to the cell to form microparticles, nanoparticles, liposomes or micelles. The amino alcohol lipidoid compound can be combined with other amino alcohol lipidoid compounds, polymers (synthetic or natural), surfactants, cholesterol, carbohydrates, proteins, lipids, etc. to form particles. These particles can then be optionally combined with a pharmaceutical excipient to form a pharmaceutical composition.

[0229] Poly(β-amino alcohols) (PBAAs) can be used to deliver nucleases or nucleic acids encoding the same and gRNAs or nucleic acids encoding the same to target cells. U.S. Patent Application Publication No. 2013 / 0302401 relates to a class of poly(β-amino alcohols) (PBAAs) that have been prepared using combinatorial polymerization.

[0230] Sugar-based particles such as GalNAc as described in International Patent Publication No. WO2014118272 (incorporated herein by reference in its entirety, and Nair, JK et al., 2014, Journal of the American Chemical Society 136(49), 16958-16961) can be used to deliver nucleases or nucleic acids encoding the same and gRNA or nucleic acids encoding the same to target cells.

[0231] In some cases, lipid nanoparticles (LNPs) are used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells. Negatively charged polymers (such as RNA) can be loaded into LNPs at low pH values ​​(e.g., pH 4), where ionizable lipids show positive charges. However, at physiological pH values, LNPs show low surface charges compatible with longer circulation times. Four ionizable cationic lipids have been concerned, i.e., 1,2-dilinoleyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N, N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxy-keto-N, N-dimethyl-3-aminopropane (DLinKDMA) and 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA). Preparation of LNPs is described, for example, in Rosin et al. (2011) Molecular Therapy 19: 1286-2200. Cationic lipids 1,2-dilinoleoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxyketo-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA), (3-o-[2"-(methoxypolyethylene glycol 2000) succinyl]-1,2-dimyristoyl-sn -diol (PEG-S-DMG) and R-3-[(ω-methoxy-poly(ethylene glycol) 2000)carbamoyl]-1,2-dimyristyloxypropyl-3-amine (PEG-C-DOMG). Nucleic acids can be encapsulated in LNPs containing DLinDAP, DLinDMA, DLinK-DMA and DLinKC2-DMA (40:10:40:10 molar ratio of cationic lipid: DSPC: CHOL: PEGS-DMG or PEG-C-DOMG). In some cases, 0.2% SP-DiOC18 was incorporated.

[0232] Spherical nucleic acid (SNA) TM ) constructs and other nanoparticles (particularly gold nanoparticles) can be used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells.

[0233] Self-assembled nanoparticles with RNA can be constructed with polyethyleneimine (PEI) PEGylated with an Arg-Gly-Asp (RGD) peptide ligand attached at the distal end of polyethylene glycol (PEG).

[0234] Nanoparticles suitable for delivering nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells can be provided in different forms, for example, as solid nanoparticles (e.g., metals (such as silver, gold, iron, titanium), non-metals, lipid-based solids, polymers), suspensions of nanoparticles, or combinations thereof. Metallic, dielectric and semiconductor nanoparticles, as well as hybrid structures (e.g., core-shell nanoparticles) can be prepared. If the nanoparticles made of semiconductor materials are small enough (usually less than 10 nm) so that quantization of electronic energy levels occurs, they can also be labeled as quantum dots. Such nanoscale particles are used as drug carriers or imaging agents in biomedical applications, and may be suitable for similar purposes in the present disclosure. Generally speaking, "nanoparticle" refers to any particle with a diameter less than 1000 nm. In some cases, the nanoparticles suitable for delivering nucleases or nucleic acids to target cells have a diameter of 500nm or less, such as 25nm to 35nm, 35nm to 50nm, 50nm to 75nm, 75nm to 100nm, 100nm to 150nm, 150nm to 200nm, 200nm to 300nm, 300nm to 400nm or 400nm to 500nm. In some cases, the nanoparticles suitable for delivering nucleases or nucleic acids to target cells have a diameter of 25nm to 200nm.

[0235] In some cases, exosomes are used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells. Exosomes are endogenous nanovesicles that transport RNA and proteins, which can deliver RNA to the brain and other target organs.

[0236] In some cases, liposomes are used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells. Liposomes are spherical vesicle structures consisting of a monolayer or multilayer lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable external lipophilic phospholipid bilayer. Liposomes can be made of several different types of lipids; however, phospholipids are most commonly used to generate liposomes. Although the formation of liposomes is spontaneous when the lipid film is mixed with an aqueous solution, the formation of liposomes can also be accelerated by applying force in the form of shaking using a homogenizer, a sonicator, or an extrusion device. Several other additives can be added to liposomes to change their structure and properties. For example, cholesterol or sphingomyelin can be added to the liposome mixture to help stabilize the liposome structure and prevent leakage of cargo inside the liposome. Liposome preparations can be mainly composed of natural phospholipids and lipids such as 1,2-distearoyl-sn-glyceryl-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine, and monosialic acid ganglioside.

[0237] Stable nucleic acid-lipid particles (SNALP) can be used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells. SNALP formulations can contain lipids 3-N-[(methoxypoly(ethylene glycol) 2000)carbamoyl]-1,2-dimyristoyloxy-propylamine (PEG-C-DMA), 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC) and cholesterol in a molar percentage of 2:40:10:48. SNALP liposomes can be prepared by formulating D-lin-DMA and PEG-C-DMA with distearoylphosphatidylcholine (DSPC), cholesterol and siRNA using a 25:1 lipid / siRNA ratio and a 48 / 40 / 10 / 2 molar ratio of cholesterol / D-Lin-DMA / DSPC / PEG-C-DMA. The size of the resulting SNALP liposomes can be about 80-100 nm. SNALP can include synthetic cholesterol (Sigma-Aldrich, St Louis, Mo., USA), dipalmitoylphosphatidylcholine (Avanti Polar Lipids, Alabaster, Ala., USA), 3-N-[(w-methoxypoly(ethylene glycol) 2000)carbamoyl]-1,2-dimyristyloxypropylamine and cationic 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane. SNALP can include synthetic cholesterol (Sigma-Aldrich), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC; Avanti Polar Lipids Inc.), PEG-cDMA and 1,2-dilinoleyloxy-3-(N;N-dimethyl)aminopropane (DLinDMA).

[0238] Other cationic lipids such as amino lipid 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane (DLin-KC2-DMA) can be used to deliver nucleases or nucleic acids to target cells. Preformed vesicles with the following lipid composition can be envisioned: amino lipids, distearoylphosphatidylcholine (DSPC), cholesterol, and (R)-2,3-bis(octadecyloxy)propyl-1-(methoxypoly(ethylene glycol) 2000)propylcarbamate (PEG-lipid) in a molar ratio of 40 / 10 / 40 / 10, respectively, and a FVII siRNA / total lipid ratio of about 0.05 (w / w). To ensure a narrow particle size distribution in the range of 70-90 nm and a low polydispersity index of 0.11.+-.0.04 (n=56), the particles can be extruded through an 80 nm membrane up to three times before adding the guide RNA. Particles containing the highly potent amino lipid 16 can be used, wherein the molar ratios of the four lipid components 16, DSPC, cholesterol, and PEG-lipid (50 / 10 / 38.5 / 1.5) can be further optimized to enhance in vivo activity.

[0239] Lipids can be formulated with nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to form lipid nanoparticles (LNPs). Suitable lipids include, but are not limited to, DLin-KC2-DMA4, C12-200 and co-lipids distearoylphosphatidylcholine, cholesterol and PEG-DMG, which can be formulated with nucleases or nucleic acids using a spontaneous vesicle formation method.

[0240] Nucleases or nucleic acids encoding the same and gRNA or nucleic acids encoding the same can be delivered encapsulated in PLGA microspheres, such as those further described in U.S. Published Applications 20130252281, 20130245107, and 20130244279.

[0241] Supercharged proteins can be used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells. Supercharged proteins are a class of engineered or naturally occurring proteins with abnormally high positive or negative net theoretical charges. Both supernegatively and superpositively charged proteins show the ability to withstand thermally or chemically induced aggregation. Superpositively charged proteins are also able to penetrate mammalian cells. The association of cargo with these proteins such as plasmid DNA, RNA or other proteins can facilitate the functional delivery of these macromolecules into mammalian cells in vitro and in vivo.

[0242] Cell penetrating peptides (CPPs) can be used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to target cells. CPPs typically have an amino acid composition that contains high relative abundance of positively charged amino acids, such as lysine or arginine, or a sequence that contains an alternating pattern of polar / charged amino acids and nonpolar hydrophobic amino acids.

[0243] method

[0244] The present disclosure also provides a method for modifying a target nucleic acid sequence (e.g., DNA or RNA). As used herein, the phrase "modified nucleic acid sequence" refers to at least one physical feature of a nucleic acid sequence of interest that is modified. Nucleic acid modifications include, for example, single-strand or double-strand breaks, deletions, or insertions of one or more nucleotides, and other modifications that affect the structural integrity of the nucleic acid sequence or the nucleotide sequence. Modifications may include one or more of the modification of the target nucleic acid, the regulation of target nucleic acid transcription, and the modification of a polypeptide associated with the target nucleic acid. The method includes contacting the target nucleic acid sequence with a composition as disclosed herein, a system disclosed herein, or a composition comprising the system.

[0245] In one embodiment, the method introduces single-strand or double-strand breaks in the target nucleic acid sequence. In this regard, the disclosed system can direct the cleavage of one or both strands of the target DNA sequence, such as within the target genomic DNA sequence and / or within the complement of the target sequence.

[0246] In some embodiments, contacting the target nucleic acid sequence comprises introducing a composition or system described herein into a cell. As described above, the composition or system can be introduced into a eukaryotic or prokaryotic cell by methods known in the art.

[0247] The cell can be a prokaryotic cell, a plant cell, an insect cell, a vertebrate cell, an invertebrate cell, an animal cell, a mammalian cell or a human cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a vertebrate cell. In some embodiments, the cell is an invertebrate cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some cases, the cell is in vitro (e.g., fresh isolate-early passage). In some cases, the cell is in vivo. In some cases, the cell is cultured in vitro (e.g., immortalized cell line).

[0248] The cells can be from an established cell line or they can be primary cells, where "primary cells," "primary cell lines," and "primary cultures" are used interchangeably herein to refer to cells and cell cultures that are derived from a subject and allowed to grow in vitro for a limited number of passages. For example, a primary culture is a culture that can be passaged 0, 1, 2, 4, 5, 10, or 15 times but has not been passaged a sufficient number of times to pass the breakout phase. Typically, primary cell lines are maintained in culture for less than 10 passages.

[0249] Suitable cells include, but are not limited to, bacterial cells; archaeal cells; eukaryotic cells; cells of unicellular eukaryotic organisms; plant cells; protozoan cells; algal cells, such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum spratense, patens) C. agardh, etc.; fungal cells (e.g., yeast cells); animal cells; cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.); cells of insects (e.g., mosquitoes; bees; agricultural pests, etc.); cells of arachnids (e.g., spiders; ticks, etc.); cells of vertebrates (e.g., fish, amphibians, reptiles, birds, mammals); mammalian cells (e.g., rodent cells; human cells; cells of non-human mammals; cells of rodents (e.g., mice, rats); cells of lagomorphs (e.g., rabbits); cells of ungulates (e.g., cattle, horses, camels, llamas, llamas, sheep, goats, etc.); cells of marine mammals (e.g., whales, seals, elephant seals, dolphins, sea lions, etc.)), etc. Any type of cell can be of interest (e.g., stem cells, such as embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells (e.g., oocytes, sperm, oogonia, spermatogonia, etc.), adult stem cells, somatic cells, such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells; in vitro or in vivo embryonic cells of embryos at any stage, such as zebrafish embryos at the 1-cell, 2-cell, 4-cell, 8-cell, etc. stages, etc.). In some cases, the cell is a cell that is not derived from a natural organism (e.g., the cell can be a synthetically prepared cell; also referred to as an artificial cell).

[0250] Non-limiting examples of plant cells include cells from plant crops, fruits, vegetables, cereals, soy, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, squash, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, angiosperms, ferns, lycophytes, hornworts, liverworts, mosses, dicots, monocots, algae (e.g., kelp), and the like.

[0251] Suitable cells include stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells; germ cells (e.g., oocytes, sperm, oogonia, spermatogonia, etc.); somatic cells, such as fibroblasts, oligodendrocytes, glial cells, hematopoietic cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells, etc.

[0252] Suitable cells include human embryonic stem cells, fetal cardiomyocytes, myofibroblasts, mesenchymal stem cells, autologous transplant expanded cardiomyocytes, adipocytes, totipotent cells, pluripotent cells, blood stem cells, myoblasts, adult stem cells, bone marrow cells, mesenchymal cells, embryonic stem cells, parenchymal cells, epithelial cells, endothelial cells, mesothelial cells, fibroblasts, osteoblasts, chondrocytes, exogenous cells, endogenous cells, stem cells, hematopoietic stem cells, bone marrow derived progenitor cells, cardiomyocytes, skeletal cells, fetal cells, undifferentiated cells, multipotent progenitor cells, unipotent progenitor cells, monocytes, cardiac myoblasts, skeletal myoblasts, macrophages, capillary endothelial cells, xenogeneic cells, allogeneic cells and postnatal stem cells.

[0253] In some cases, the cell is an immune cell, a neuron, an epithelial cell and an endothelial cell, or a stem cell. In some cases, the immune cell is a T cell, a B cell, a monocyte, a natural killer cell, a dendritic cell, or a macrophage. In some cases, the immune cell is a cytotoxic T cell. In some cases, the immune cell is a helper T cell. In some cases, the immune cell is a regulatory T cell (Treg).

[0254] In some cases, the cell is a stem cell. Stem cells include adult stem cells. Adult stem cells are also called somatic stem cells.

[0255] Adult stem cells reside in differentiated tissues but retain the properties of self-renewal and the ability to generate a variety of cell types, often typical of the tissue in which the stem cells are present. Many examples of somatic stem cells are known to those skilled in the art, including muscle stem cells; hematopoietic stem cells; epithelial stem cells; neural stem cells; mesenchymal stem cells; mammary stem cells; intestinal stem cells; mesoderm stem cells; endothelial stem cells; olfactory stem cells; neural crest stem cells, etc.

[0256] The stem cells of interest include mammalian stem cells, wherein the term "mammal" refers to any animal classified as a mammal, including humans; non-human primates; livestock and farm animals; and zoo animals, laboratory animals, sports animals, or pet animals, such as dogs, horses, cats, cows, mice, rats, rabbits, etc. In some cases, the stem cells are human stem cells. In some cases, the stem cells are rodent (e.g., mouse; rat) stem cells. In some cases, the stem cells are non-human primate stem cells.

[0257] In some embodiments, the stem cell is a hematopoietic stem cell (HSC). HSCs are mesodermal cells that can be isolated from bone marrow, blood, umbilical cord blood, fetal liver, and yolk sac. HSCs are characterized by CD34 + and CD3 - HSCs can repopulate the erythroid, neutrophil-macrophage, megakaryocyte, and lymphohematopoietic lineages in vivo. In vitro, HSCs can be induced to undergo at least some self-renewing cell divisions, and can be induced to differentiate into the same lineages as seen in vivo. Thus, HSCs can be induced to differentiate into one or more of erythroid cells, megakaryocytes, neutrophils, macrophages, and lymphoid cells.

[0258] In other embodiments, the stem cell is a neural stem cell (NSC). Neural stem cells (NSC) can differentiate into neurons and glial cells (including oligodendrocytes and astrocytes). Neural stem cells are multipotent stem cells that can divide multiple times, and under certain conditions, daughter cells can be produced, which are neural stem cells or neural progenitor cells, which can be neuroblasts or glial cells, such as cells that are stereotyped as one or more types of neurons and glial cells, respectively. Methods for obtaining NSC are known in the art.

[0259] In other embodiments, the stem cell is a mesenchymal stem cell (MSC). MSCs, originally derived from embryonic mesoderm and isolated from adult bone marrow, can differentiate into muscle, bone, cartilage, fat, bone marrow stroma, and tendon. Methods for isolating MSCs are known in the art; and any known method can be used to obtain MSCs. See, for example, U.S. Patent No. 5,736,396, which describes the isolation of human MSCs.

[0260] In some embodiments, the cell is a T cell. The present invention is not limited by the type of T cells. T cells can be selected from, for example, CD3+T cells, CD8+T cells, CD4+T cells, natural killer (NK) T cells, αβT cells, γδT cells, or any combination thereof (e.g., a combination of CD4+ and CD8+T cells).

[0261] In some embodiments, T cells are naturally occurring T cells. For example, T cells can be separated from a subject sample. In some embodiments, T cells are anti-tumor T cells (e.g., T cells with activity against tumors (e.g., autologous tumors), which become activated and amplified in response to antigens). Anti-tumor T cells include, but are not limited to, T cells (e.g., tumor infiltrating lymphocytes (TILs)) obtained from resected tumors or tumor biopsies and polyclonal or monoclonal tumor reactive T cells (e.g., obtained by apheresis, amplified in vitro for tumor antigens presented by autologous or artificial antigen presenting cells). In some embodiments, T cells are amplified in vitro.

[0262] In some cases, the cell is a plant cell. The plant cell can be a cell of a monocot. The plant cell can be a cell of a dicot. The cell can be a root cell, a leaf cell, a xylem cell, a phloem cell, a cambium cell, a apical meristem cell, a parenchyma cell, a collenchyma cell, a sclerenchyma cell, etc. The plant cell includes cells of agricultural crops such as wheat, corn, rice, sorghum, millet, soybean, etc. The plant cell includes cells of agricultural fruit and nut plants (e.g., plants producing apricots, oranges, lemons, apples, plums, pears, almonds, etc.).

[0263] The plant cell can be a cell of a major agricultural plant (e.g., barley, beans (dry food), canola, corn, cotton (Pima), cotton (Upland), linseed, hay (alfalfa), hay (non-alfalfa), oats, peanuts, rice, sorghum, soybeans, sugar beets, sugar cane, sunflower (oil), sunflower (non-oil), sweet potato, tobacco (Burley), tobacco (flue-cured), tomato, wheat (Durum), wheat (Spring), wheat (Winter), etc.). In another example, the cell is a cell of a vegetable crop, which includes, but is not limited to, alfalfa sprouts, aloe vera leaves, cassava, arrowhead, artichoke, asparagus, bamboo shoots, banana flowers, bean sprouts, beans, beet tops, beets, bitter melon, cabbage, broccoli, kale (watercress heart), Brussels sprouts, cabbage, cabbage sprouts, cactus leaves (nopal cactus), pumpkin, artichoke, carrot, cauliflower, celery, chayote, Chinese artichoke (manna seed), Chinese cabbage, Chinese celery, leek, vegetable heart, chrysanthemum leaf (chrysanthemum), kale, corn stalks, sweet corn, cucumber, radish, dandelion greens, taro, dau mue (pea sprouts), donqua (winter melon), eggplant, lettuce, broadleaf endive, young fern, field cress, endive, mustard greens (Chinese mustard), kale, galangal (siam, Thai ginger), garlic, ginger, burdock, greens, Hanover salad, huauzontle, Jerusalem artichoke, jicama, kale, kohlrabi, lamb's leg quinoa (Quilete), lettuce (Bibb lettuce), lettuce (Boston lettuce), lettuce (Boston red lettuce), lettuce (green leaf lettuce), lettuce (iceberg lettuce), lettuce (Rosa lettuce), lettuce (green oak leaf lettuce), lettuce (red oak leaf lettuce), lettuce (processed lettuce), lettuce (red leaf lettuce), lettuce (romaine lettuce), lettuce (ruby romaine lettuce), lettuce (Russian red mustard lettuce), linkok, white radish, long bean, lotus root, mache, agave leaf, yellow meat taro, mesculin mix, mizuna, moap (smooth loofah), moo, hairy melon, mushroom, mustard, yam, okra, water spinach, onion, gourd (long melon), ornamental corn, ornamental gourd, parsley, parsnip, peas, peppers (sweet peppers), pepper, pumpkin, endive, radish sprouts, radish, rapeseed, rapeseed, rhubarb, romaine lettuce (small red lettuce), rutabaga, sea beans, loofah (oblique strips / horizontal strips), spinach, pumpkin, straw bale, sugarcane, sweet potato, Swiss chard, tamarind, taro, taro leaves, taro shoots, tamarind, tepeguaje (river tamarind), vine melon, tomatillo, tomato, tomato (cherry tomato), tomato (grape tomato), tomato (plum tomato), turmeric, radish leaves, turnip, water chestnut, potato, yam, rapeseed, cassava, etc.

[0264] In some cases, the cell is an arthropod cell. For example, the cell can be a cell of a suborder, family, subfamily, group, subgroup, or species, such as Chelicerata, Myriapodia, Hexipodia, Arachnida, Insecta, Archaeognatha, Thysanura, Palaeoptera, Ephemeroptera, Odonata, or a suborder, family, subfamily, group, subgroup, or species. nata), Anisoptera, Zygoptera, Neoptera, Exopterygota, Plecoptera, Embioptera ), Orthoptera, Zoraptera, Dermaptera, Dictyoptera, Notoptera, Grylloblattidae, Mantis The order Mantophasmatidae, Phasmatodea, Blattaria, Isoptera, Mantodea, Parapneuroptera, Psocoptera, Thysanoptera, Phthiraptera, Hemiptera, Endopterygota or Holometabola, Hymenoptera, Coleoptera, Strepsiptera, Raphidioptera, Megaloptera, Neuroptera, Mecoptera, Siphonaptera, Diptera, Trichoptera or Lepidoptera.

[0265] In some cases, the cell is an insect cell. For example, in some cases, the cell is a cell of a mosquito, grasshopper, true bug, fly, flea, bee, wasp, ant, louse, moth, or beetle.

[0266] In some embodiments, introducing the system into a cell comprises administering the system to a subject. In some embodiments, the subject is a human. Administration may include in vivo administration. In alternative embodiments, the vector is contacted with the cell in vitro or ex vivo, and the treated cells containing the system are transplanted into the subject.

[0267] In some embodiments, the target nucleic acid is a nucleic acid endogenous to the target cell. In some embodiments, the target nucleic acid is a genomic DNA sequence. As used herein, the term "genome" refers to a nucleic acid sequence (e.g., a gene or locus) located on the chromosome of a cell.

[0268] In some embodiments, the target nucleic acid encodes a gene or gene product. As used herein, the term "gene product" refers to any biochemical product produced by gene expression. The gene product can be RNA or protein. RNA gene products include non-coding RNA, such as tRNA, rRNA, microRNA (miRNA) and small interfering RNA (siRNA), and coding RNA, such as messenger RNA (mRNA). In some embodiments, the target nucleic acid sequence encodes a protein or polypeptide.

[0269] The disclosed methods can modify a target DNA sequence in a host cell to modulate the expression of the target DNA sequence, for example, the expression of the target DNA sequence is increased, decreased, or completely eliminated (eg, via deletion of the gene).

[0270] In another embodiment, the method of modifying a target sequence can be used to delete a nucleic acid sequence or a portion thereof from a target sequence in a host cell by cleaving the target sequence and allowing the host cell to repair the cleaved sequence in the absence of an exogenously provided donor nucleic acid molecule. Deleting a nucleic acid sequence in this manner can be used for a variety of applications, such as removing disease-causing trinucleotide repeat sequences in neurons, generating gene knockouts or knockdowns, and generating mutations for disease models in research.

[0271] In some embodiments, the systems and methods described herein can be used to insert a gene or a fragment thereof into a cell. In certain embodiments, the disclosed system can be used to generate cells expressing recombinant receptors. In some embodiments, the recombinant receptor is a T cell receptor (TCR) or a chimeric antigen receptor (CAR). Also provided herein are cells, such as T cells, comprising recombinant receptors as described herein and / or nucleic acids encoding them and systems (e.g., nucleases and at least one gRNA).

[0272] In some embodiments, the systems and methods described herein can be used for genetically modified plants or plant cells. As used herein, genetically modified plants include plants into which exogenous polynucleotides have been introduced. Genetically modified plants also include plants that have been genetically manipulated so that endogenous nucleotides have been changed to include mutations, such as deletions, insertions, conversions, transversions or combinations thereof. For example, endogenous coding regions may be deleted. Such mutations may result in polypeptides having amino acid sequences different from the amino acid sequences encoded by endogenous polynucleotides. Another example of a genetically modified plant is a plant having a changed regulatory sequence (such as a promoter) to cause an increase or decrease in the expression of an operably connected endogenous coding region. Genetically modified plants can promote desired phenotypes or genotype plant traits.

[0273] Genetically modified plants can potentially have improved crop yields, enhanced nutritional value and extended storage life. They can also resist adverse environmental conditions, insects and pesticides. The system and method of the present invention has a wide range of applications in gene discovery and verification, mutation and cisgene breeding and cross breeding. The system and method of the present invention can promote the production of a new generation of genetically modified crops with various improved agronomic traits, such as herbicide resistance, herbicide tolerance, drought tolerance, male sterility, insect resistance, abiotic stress tolerance, improved fatty acid metabolism, improved carbohydrate metabolism, improved seed yield, improved oil percentage, improved protein percentage, resistance to bacterial diseases, disease (for example, bacteria, fungi and viruses) resistance, high yield and good quality. The system and method of the present invention can also promote the production of a new generation of genetically modified crops with optimized fragrance, nutritional value, storage life, pigmentation (for example, lycopene content), starch content (for example, low gluten wheat), toxin level, breeding and / or breeding and growth time. See, e.g., CRISPR / Cas Genome Editing and Precision Plant Breeding in Agriculture (Chen et al., Annu Rev Plant Biol. 2019 Apr 29;70:667-69), which is incorporated herein by reference.

[0274] The systems and methods of the invention can confer one or more of the following traits to a plant cell: herbicide tolerance, drought tolerance, male sterility, insect resistance, abiotic stress tolerance, improved fatty acid metabolism, improved carbohydrate metabolism, improved seed yield, improved oil percentage, improved protein percentage, resistance to bacterial diseases, resistance to fungal diseases, and resistance to viral diseases.

[0275] The present disclosure provides modified plant cells produced by the systems and methods of the present invention, plants comprising the plant cells, and seeds, fruits, plant parts or propagation materials of the plants. The transformed or genetically modified plant cells of the present disclosure can be used as cell groups, or as tissues, seeds, whole plants, stems, fruits, leaves, roots, flowers, stems, tubers, grains, animal feed, plant fields, etc. The present disclosure provides a transgenic plant. The transgenic plant can be homozygous or heterozygous for genetic modification. The present disclosure also provides transformed or genetically modified plant cells, tissues, plants, and products containing the transformed or genetically modified plant cells. The present disclosure also includes offspring, clones, cell lines or cells of transgenic plants.

[0276] The systems and methods of the present invention can be used to modify plant stem cells. The present disclosure also provides progeny of genetically modified cells, wherein the progeny may contain the same genetic modification as the genetically modified cell from which it is derived. The present disclosure also provides a composition comprising genetically modified cells.

[0277] In one embodiment, transformed or genetically modified cells and tissues and products comprise nucleic acids integrated into the genome, as well as gene products produced by the plant cell as a result of the transformation or genetic modification.

[0278] Methods for introducing exogenous nucleic acids into plant cells are well known in the art. Such plant cells are considered to be "transformed". DNA constructs can be introduced into plant cells by various methods, including but not limited to protoplast transformation mediated by PEG or electroporation, tissue culture or plant tissue transformation by bioballistic bombardment, or instantaneous and stable transformation mediated by Agrobacterium. Transformation can be instantaneous or stable transformation. Suitable methods also include viral infection (such as double-stranded DNA viruses), transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whisker technology, Agrobacterium-mediated transformation, etc. The selection of methods generally depends on the type of cells transformed and the environment in which the transformation occurs (i.e., in vitro, in vitro or in vivo). Transformation methods based on soil bacteria Agrobacterium tumefaciens can be used to introduce exogenous nucleic acid molecules into vascular plants. The wild-type form of Agrobacterium contains a Ti (tumor induction) plasmid, which instructs the growth of tumor-causing crown galls on host plants. Transfer of the tumor-inducing T-DNA region of the Ti plasmid to the plant genome requires the Ti plasmid to encode virulence genes as well as a T-DNA border sequence, which is a series of direct DNA repeats that delineate the region to be transferred. Agrobacterium-based vectors are modified forms of the Ti plasmid in which the tumor-inducing function is replaced by a nucleic acid sequence of interest to be introduced into the plant host.

[0279] Agrobacterium-mediated transformation generally employs a cointegrate vector or a binary vector system in which the components of the Ti plasmid are distributed between a helper vector that is permanently present in the Agrobacterium host and carries the virulence genes and a shuttle vector that contains the gene of interest defined by the T-DNA sequence. Various binary vectors are well known in the art and are commercially available, for example, from Clontech (Palo Alto, Calif.). Methods for co-culturing Agrobacterium with cultured plant cells or wounded tissues, such as leaf tissue, root explants, underground cotyledons, stem segments, or tubers are also well known in the art. See, for example, Glick and Thompson (eds.), Methods in Plant Molecular Biology and Biotechnology, Boca Raton, Fla.: CRC Press (1993), which is incorporated herein by reference.

[0280] Microparticle-mediated transformation can also be used to generate transgenic plants. This method, first described by Klein et al. (Nature 327:70-73 (1987), which is incorporated herein by reference), relies on microparticles such as gold or tungsten that are coated with the desired nucleic acid molecules by precipitation with calcium chloride, spermidine or polyethylene glycol. Microparticle particles are accelerated into angiosperm tissue at high speed using a device such as the BIOLISTIC PD-1000 (Biorad; Hercules Calif.).

[0281] In one embodiment, the system and method of the present invention are applicable to plants. In one embodiment, a series of plant-specific RNA-guided genome editing vectors (pRGE plasmids) for expressing the system of the present invention in plants are provided. The vector can be optimized to transiently express the system of the present invention in plant protoplasts, or stably integrate and express in complete plants via Agrobacterium-mediated transformation. In one aspect, the vector construct comprises: a nucleotide sequence containing a DNA-dependent RNA polymerase III promoter, wherein the promoter is operably linked to a gRNA molecule and a Pol III termination sequence; and a nucleotide sequence containing a DNA-dependent RNA polymerase II promoter, which is operably linked to a nucleic acid sequence encoding a nuclease.

[0282] In certain embodiments, the system and method of the present invention use a monocot promoter to drive the expression of one or more components (e.g., gRNA) of the system of the present invention in monocots. In certain embodiments, the system and method of the present invention use a dicot promoter to drive the expression of one or more components (e.g., gRNA) of the system of the present invention in dicots. In some embodiments, the system of the present invention is transiently expressed in plant protoplasts. The vectors used for plant transient transformation include but are not limited to pRGE3, pRGE6, pRGE31 and pRGE32. In some embodiments, the vector can be optimized for specific plant types or species, such as pStGE3.

[0283] In one embodiment, the system of the present invention can be stably integrated into the plant genome, for example, via Agrobacterium-mediated transformation. Thereafter, one or more components of the system of the present invention (e.g., transgenes) can be removed by genetic hybridization and separation, which can result in the production of non-transgenic but genetically modified plants or crops. In one embodiment, the vector is optimized for Agrobacterium-mediated transformation. In one embodiment, the vector used for stable integration is pRGEB3, pRGEB6, pRGEB31, pRGEB32, or pStGEB3.

[0284] The system of the invention can be used in a variety of bacterial hosts, including medically important human pathogens, bacterial pests that are key targets in the agricultural industry, and antibiotic-resistant forms thereof.

[0285] The systems and methods can be designed to target any gene or any group of genes, such as virulence or metabolic genes, for clinical and industrial applications in other embodiments. For example, the systems and methods of the present invention can be used to target and eliminate virulence genes from a population for in situ gene knockout, or to stably introduce new genetic elements into a metagenomic library of a microbiome. The systems and methods of the present invention can be used to treat multidrug-resistant bacterial infections in a subject. The systems and methods of the present invention can be used for genome engineering in complex bacterial consortia.

[0286] The systems and methods of the present invention can be used to inactivate microbial genes. In some embodiments, the gene is an antibiotic resistance gene. For example, the coding sequence of a bacterial resistance gene can be destroyed in vivo by inserting a DNA sequence, resulting in non-selective resensitization to drug treatment.

[0287] The components of the composition or system can be administered as a pharmaceutical composition together with a pharmaceutically acceptable carrier or excipient. In some embodiments, the components of the system of the present invention can be mixed with a pharmaceutically acceptable carrier alone or in any combination to form a pharmaceutical composition, which is also within the scope of the present disclosure.

[0288] In some embodiments, an effective amount of a component of a system or composition of the invention as described herein may be administered.In the context of the present disclosure, the term "effective amount" refers to the amount of a component of a system, which amount enables modification of a target nucleic acid to be achieved.

[0289] The methods described herein are also used to treat a disease or condition in a subject. In some embodiments, the system and method are used to treat a pathogen or parasite on or in a subject by changing the pathogen or parasite. In some embodiments, the system and method target "disease-related" genes. The term "disease-related gene" refers to any gene or polynucleotide whose gene product is expressed at an abnormal level or in an abnormal form in a cell obtained from an individual affected by the disease compared to a tissue or cell obtained from an individual not affected by the disease. Disease-related genes can be expressed at abnormally high levels or abnormally low levels, wherein the expression of the changes is associated with the occurrence and / or progression of the disease. Disease-related genes also refer to genes whose mutations or genetic variations are directly responsible for the cause of the disease or are not in equilibrium with the gene linkage responsible for the cause of the disease. Examples of genes responsible for such "single gene" or "monogenic" diseases include, but are not limited to, adenosine deaminase, alpha-1 antitrypsin, cystic fibrosis transmembrane conductance regulator (CFTR), beta-hemoglobin (HBB), oculocutaneous albinism II (OCA2), huntingtin (HTT), myotonic dystrophy protein kinase (DMPK), low-density lipoprotein receptor (LDLR), apolipoprotein B (APOB), neurofibromin 1 (NF1), polycystic kidney disease 1 (PKD1), polycystic kidney disease 2 (PKD2), coagulation factor VIII (F8), dystrophin (DMD), phosphate-regulated endopeptidase homolog, X-linked (PHEX), methyl-CpG-binding protein 2 (MECP2), and ubiquitin-specific peptidase 9Y, Y-linked (USP9Y). Other single genes or monogenic diseases are known in the art and are described in, for example, Chial, H. Rare Genetic Disorders: Learning About Genetic Disease Through Gene Mapping, SNPs, and Microarray Data, Nature Education 1(1): 192 (2008); Online Mendelian Inheritance in Man (OMIM); Human Gene Mutation Database (HGMD). In another embodiment, the target genomic DNA sequence may comprise a gene whose mutations in combination with mutations in other genes result in a specific disease. Diseases caused by the contribution of multiple genes that lack a simple (i.e., Mendelian) inheritance pattern are referred to in the art as "multifactorial" or "polygenic" diseases. Examples of multifactorial or polygenic diseases include, but are not limited to, asthma, diabetes, epilepsy, hypertension, bipolar disorder, and schizophrenia. Certain developmental abnormalities may also be inherited in a multifactorial or polygenic pattern, including, for example, cleft lip / palate, congenital heart defects, and neural tube defects. In another embodiment, the target DNA sequence may comprise an oncogene.

[0290] The present disclosure provides gene editing methods that can ablate disease-related genes (e.g., oncogenes), which can then be used for in vivo gene therapy in patients. In some embodiments, the gene editing method includes a donor nucleic acid containing a therapeutic gene.

[0291] When used as a therapeutic method, the effective amount may depend on the specific condition being treated, the severity of the condition, individual patient parameters including age, physical condition, size, sex, and weight, the duration of treatment, the nature of concurrent therapy (if any), the specific route of administration, and similar factors within the knowledge and expertise of a health practitioner. In some embodiments, the effective amount alleviates, relieves, ameliorates, improves, reduces symptoms, or delays the progression of any disease or condition in the subject. In some embodiments, the subject is a human.

[0292] A variety of additional therapies can be used in conjunction with the methods of the present disclosure. Additional therapies can be the administration of additional therapeutic agents, or can be additional therapies that are unrelated to the administration of another agent. Such additional therapies include, but are not limited to, surgery, immunotherapy, and radiotherapy. Additional therapies can be administered simultaneously with the above methods. In some embodiments, additional therapies can be administered at intervals of several hours to several months before or after the treatment of the disclosed methods.

[0293] In some embodiments, a therapeutically effective amount of a system as described herein (e.g., a nuclease and / or gRNA) or composition is administered alone or in combination with a therapeutically effective amount of at least one additional therapeutic agent. In some embodiments, effective combined therapy is achieved with a single composition or pharmacological preparation or with two different compositions or preparations, which are administered simultaneously or at intervals of time. The at least one additional therapeutic agent may include any treatment modality, including proteins, small molecules, nucleic acids, etc. For example, exemplary additional therapeutic agents include, but are not limited to, immunomodulators, chemotherapeutics, nucleic acids (e.g., mRNA, aptamers, antisense oligonucleotides, ribozyme nucleic acids, interfering RNA, antigenomic nucleic acids), decongestants, steroids, analgesics, antimicrobials, immunotherapy, or any combination thereof.

[0294] In the context of the present disclosure, in the context of any disease condition described herein, the terms "treat", "treatment", etc., mean to alleviate or relieve at least one symptom associated with such condition, or to slow or reverse the progression of such condition. Within the meaning of the present disclosure, the term "treatment" also means to prevent, delay the onset (e.g., the period before the clinical manifestation of the disease) and / or reduce the risk of disease development or worsening. For example, in the case of cancer, the term "treatment" may mean to eliminate or reduce the patient's tumor burden, or to prevent, delay or inhibit metastasis, etc.

[0295] The phrase "pharmaceutically acceptable" as used in conjunction with the compositions and / or cells of the present disclosure refers to the molecular entities and other ingredients of such compositions that are physiologically tolerable and generally do not produce adverse reactions when administered to a subject (e.g., a mammal, a human). Preferably, as used herein, the term "pharmaceutically acceptable" means approved by a regulatory agency of a federal or state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopoeia for use in mammals, and more specifically in humans. "Acceptable" means that the carrier is compatible with the active ingredients of the composition (e.g., nucleic acids, vectors, cells, or therapeutic antibodies) and does not adversely affect the subject to whom the composition is administered. Any pharmaceutical composition and / or cell used in the methods of the present invention may include a pharmaceutically acceptable carrier, excipient, or stabilizer in the form of a lyophilized formulation or an aqueous solution.

[0296] Pharmaceutically acceptable carriers, including buffers, are well known in the art and may include phosphates, citrates and other organic acids; antioxidants, including ascorbic acid and methionine; preservatives; low molecular weight polypeptides; proteins, such as serum albumin, gelatin or immunoglobulins; amino acids; hydrophobic polymers; monosaccharides; disaccharides; and other carbohydrates; metal complexes; and / or nonionic surfactants. See, for example, Remington: The Science and Practice of Pharmacy 20th Edition (2000) Lippincott Williams and Wilkins, Ed. KE Hoover.

[0297] In some cases, the desired delivery system provides a generally uniform distribution and has a controlled release rate of its components (e.g., carriers, proteins, nucleic acids) in vivo. Various different media that can be used to construct a composition delivery system are described below. This does not mean that any one medium will limit the present invention. It should be noted that any medium can be combined with another medium or carrier; for example, in one embodiment, polymer microparticles attached to the compound can be combined with a gel medium. Implantable devices can be used to deliver nucleases or nucleic acids encoding them and gRNA or nucleic acids encoding them to, for example, target cells in vivo.

[0298] Contemplated carriers or vehicles include materials such as gelatin, collagen, cellulose esters, dextran sulfate, pentosan polysulfate, chitin, carbohydrates, albumin, fibrin adhesives, synthetic polyvinyl pyrrolidone, polyethylene oxide, polypropylene oxide, block polymers of polyethylene oxide and polypropylene oxide, polyethylene glycol, acrylates, acrylamides, methacrylates (including but not limited to 2-hydroxyethyl methacrylate), poly(orthoesters), cyanoacrylates, gelatin-resorcinol-aldehyde type bioadhesives, polyacrylic acid, and copolymers and block copolymers thereof.

[0299] In some cases, the carrier / medium may comprise microparticles. Microparticles may include, but are not limited to, liposomes, nanoparticles, microspheres, nanospheres, microcapsules and nanocapsules. In some cases, microparticles may include one or more of: poly(lactide-co-glycolide), aliphatic polyesters (including but not limited to polyethylene glycol acid and polylactic acid), hyaluronic acid, modified polysaccharides, chitosan, cellulose, dextran, polyurethane, polyacrylic acid, pseudo-poly(amino acid), copolymers related to polyhydroxybutyrate, polyanhydrides, polymethyl methacrylate, poly(ethylene oxide), lecithin and phospholipids, with any combination thereof.

[0300] In some cases, the carrier / medium may include liposomes that can attach and release therapeutic agents (e.g., subject nucleic acids and / or proteins). Liposomes are microscopic spherical lipid bilayers that surround an aqueous core made of amphiphilic molecules such as phospholipids. For example, liposomes can trap therapeutic agents between the hydrophobic tails of phospholipid micelles. Water-soluble agents can be embedded in the core, while fat-soluble agents can be dissolved in the shell-like bilayer. Liposomes have special characteristics because they can use water-soluble and water-insoluble chemicals together in the medium without using surfactants or other emulsifiers. Liposomes can be spontaneously formed by mixing phospholipids in an aqueous medium. Water-soluble compounds are dissolved in aqueous solutions that can hydrate phospholipids. Therefore, when liposomes are formed, these compounds are trapped in the center of aqueous liposomes. The liposome wall as a phospholipid membrane accommodates fat-soluble substances, such as oil. Liposomes provide controlled release of incorporated compounds. In addition, liposomes can be coated with water-soluble polymers such as polyethylene glycol to increase the pharmacokinetic half-life.

[0301] In some embodiments, cationic or anionic liposomes are used as part of the subject composition or method, or liposomes with neutral lipids can also be used. Cationic liposomes can contain negatively charged substances by mixing negatively charged substances with fatty acid liposome components and allowing them to charge associate. The selection of cationic or anionic liposomes depends on the desired pH of the final liposome mixture.

[0302] Where appropriate, any elements of any suitable CRISPR / Cas gene editing system known in the art can be used in the systems and methods described herein. CRISPR / Cas gene editing technology is described in detail, for example, in U.S. Patent Nos. 8,546,553, 8,697,359, 8,771,945, 8,795,965, 8,865,406, 8,871,445, 8,889,356, 8,889,418, 8,895,308, 8,9066,616, 8,932,814, 8,945,839, 8,993,233, 8,999,641, 9 , 115,348, 9,149,049, 9,493,844, 9,567,603, 9,637,739, 9,663,782, 9,404,098, 9,885,026, 9,951,342, 10,087,431, 10,227,610, 10,266,850, 10,601,748, 10,604,771, and 10,760,064; and U.S. Patent Application Publication No. US201 0 / 0076057、US2014 / 0113376、US2015 / 0050699、US2015 / 0031134、US2014 / 0357530、US2014 / 0349400、 US2014 / 0315985, US2014 / 0310830, US2014 / 0310828, US2014 / 0309487, US2014 / 0294773, US2014 / 0287 938, US2014 / 0273230, US2014 / 0242699, US2014 / 0242664, US2014 / 0212869, US2014 / 0201857, US2014 / 0199767, US2014 / 0189896, US2014 / 0186919, US2014 / 0186843 and US2014 / 0179770, each of which is incorporated herein by reference.

[0303] Reagent test kit

[0304] Kits comprising the compositions, systems, or components thereof as disclosed herein are also within the scope of the present disclosure.

[0305] For example, the kit may include one or more reagents or other components useful, necessary, or sufficient to perform any of the methods described herein, such as editing reagents (nucleases, guide RNAs, vectors, compositions, etc.), transfection or administration reagents, negative and positive control samples (e.g., cells, template DNA), cells, containers for holding one or more components (e.g., microcentrifuge tubes, boxes), detectable labels, detection and analysis instruments, software, instructions, etc.

[0306] The kit may include instructions for use of any of the methods described herein. The instructions may include a description of applying the system or composition of the present invention to a subject to achieve the desired effect. The instructions typically include information about the dosage, dosing regimen, and route of administration for the intended treatment. The kit may also include instructions for selecting a subject suitable for treatment based on identifying whether the subject needs treatment.

[0307] The kits provided herein are packaged in suitable packaging. Suitable packaging includes, but is not limited to, vials, bottles, jars, soft packaging, etc. The kit may have a sterile inlet (e.g., the container may be an intravenous bag or vial with a stopper that can be pierced by a hypodermic needle). The container may also have a sterile inlet.

[0308] The packaging can be a unit dose, a bulk package (e.g., a multi-dose package), or a subunit dose. The instructions provided in the kit of the present disclosure are typically written instructions on a label or package insert. The label or package insert indicates that the pharmaceutical composition is used to treat a disease or condition in a subject, delay the onset of the disease or condition, and / or alleviate the disease or condition.

[0309] The kit may optionally provide additional components, such as buffers and interpretative information. Typically, the kit includes a container and a label or package insert on or associated with the container. In some embodiments, the disclosure provides an article of manufacture comprising the contents of the above-described kit.

[0310] The kit may also include a device for holding or administering the system or composition of the invention. The device may include an infusion set, an intravenous solution bag, a hypodermic needle, a vial, and / or a syringe.

[0311] Example

[0312] The following are examples of the present invention and should not be construed as limiting.

[0313] Example 1

[0314] Nuclease and guide RNA vectors

[0315] The identification of the single guide RNA vector group Nuclease sequences (SEQ ID NO: 1-250) were identified as candidate CRISPR V-type nucleases with Cas12f-like features. Based on its predicted crRNA and tracrRNA binding and folding patterns, single guide RNA (sgRNA) vectors were designed for nuclease SEQ ID NO: 1-54 (Table 5). The designed sgRNA was placed downstream of the U6 promoter with an initiator G and then placed upstream of the spacer sequence (Table 6).

[0316] Nuclease expression vectors synthesize codon-optimized genes encoding candidate nucleases (nuclease amino acid sequences SEQ ID NO: 20-29 and 36) and clone them into mammalian expression vectors under CMV promoter pTwist_CMV (Twist Biosciences). The cloned nuclease is placed in an expression vector having an SV40 nuclear localization sequence (NLS) fused to its N-terminus and a nucleoplasmin NLS located on its C-terminus, followed by a 3xHA tag. A similar vector was constructed with Un1Cas12f1 (SEQ ID NO: 471).

[0317] Example 2

[0318] Editing activity in human cells

[0319] Nucleases SEQ ID NO: 21, 24 and 36 were tested in HEK293T cells by plasmid transfection using Mirus Transit X2 reagent. 50,000 cells were plated in each well of a 96-well plate and immediately transfected with 100 ng of the nuclease expression vector and 100 ng of the corresponding sgRNA vector shown in Table 1.

[0320] Table 1

[0321]

[0322] The samples were incubated for 72 hours and harvested with QuickExtract (Lucigen). About 200ng of genomic DNA was amplified using KAPA HiFi polymerase and primers with Illumina adapters ACACTCTTTCCCTACACGACGCTCTTCCGATCTgtaatgagcaaccttgagggatcagg (SEQ ID NO: 506) and GACTGGAGTTCAGACGTGTGCTCTTCCGATCTctcatggcaaaagcagtaatcagaac (SEQ ID NO: 507) specific for the target region on chromosome 3. 2uL of the first 25uL PCR was input into the second PCR using the Illumina P7 barcode primer from New England BioLabs test kit #E6609S. The purity of the PCR product was checked on a 2% agarose gel and cleaned by ZYMO test kit #D4034. The samples were then sequenced on an Illumina MiSeq system, returning 100,000-400,000 150bp paired-end reads per sample. Editing analysis was performed by CRISPResso2 with option "--cleavage_offset 1" (Clement, Kendell et al., "CRISPResso2 provides accurate and rapid genomeediting sequence analysis." Nature biotechnology 37.3 (2019): 224-226.). The percentage of nucleotide insertion or deletion mutations (indels) around the cleavage site was calculated for transfected and non-transfected (NT) cells, excluding mutations with only substitutions. The percentage of indels of transfected cells was divided by the percentage of indels of non-transfected cells to calculate the fold change of editing. The results are shown in Figure 1 middle.

[0323] Example 3

[0324] Engineering single guide RNA

[0325] Nuclease SEQ ID NO:21,24 and 36 engineered single guide RNA (sgRNA) vectors are designed to have different lengths, as shown in Table 2. The designed sgRNA is placed downstream of the U6 promoter with a start G, and then placed upstream of the spacer sequence CACACACACAGTGGGCTACC (SEQ ID NO:423), which targets the intergenic region of chromosome 3 of the human genome and has a 5'TTTG PAM sequence. Nuclease SEQ ID NO:21,24 and 36 were tested in HEK293T cells by plasmid transfection using Mirus Transit X2 reagent. 50,000 cells were plated in each well of a 96-well plate and immediately transfected with 100ng of the nuclease expression vector and 100ng of the corresponding sgRNA vector. The samples were incubated for 72 hours and harvested with QuickExtract (Lucigen). Genomic DNA was amplified around the target region on chromosome 3 and sequenced by Sanger sequencing. TIDE (Tracking Indels by Decomposition) analysis was performed according to the method of Brinkman et al. (Brinkman EK, Chen T, Amendola M, van Steensel B. Nucleic Acids Res. 2014; 42(22): e168, which is incorporated herein by reference in its entirety) and the recommendations on tide.nki.nl. The results are shown in Figure 2 Table 3 shows the corresponding nuclease and guide RNA sequences for each numbered sample. Certain truncations using sgRNAs improved editing.

[0326] Example 4

[0327] Editing activity in human cells

[0328] Following the method described in Example 2, the editing activity of nucleases SEQ ID NOs: 20-29 and 36 was tested in HEK293T cells targeting Kim-T1 (SEQ ID NO: 423) with sgRNA of SEQ ID NO: 346. Figure 3 The results shown demonstrate that the selected nucleases have editing activity in human cells.

[0329] Example 5

[0330] Off-target editing activity

[0331] As described in Example 3, the nuclease SEQ ID NO: 20 was tested with a guide (SEQ-ID NO: 430) matching the TCRA gene or with a single mismatch of TCRA at different positions (SEQ-ID NO: 433-452). The mismatched guide was used as an artificial off-target to determine the tendency of the nuclease to edit when there was a mismatch at each position of the guide. As described in Example 3, the editing efficiency of the matched guide and the mismatched guide was determined by Sanger sequencing. The resulting amplicon was subjected to Sanger sequencing, and TIDE analysis was performed according to the recommendations of Brinkman et al., 2014 and the TIDE website (tide.nki.nl). Untransfected cells were also harvested, amplified and sequenced via the same method to set a limit of detection (LOD), at which the editing level could not be determined. The results of the editing efficiency using a single mismatched guide RNA are shown in Figure 4 middle.

[0332] Example 6

[0333] Guide RNA modification

[0334] Based on the predicted crRNA and tracrRNA binding and folding patterns, single guide RNA (sgRNA) constructs for targeting Kim-T1 were designed and cloned into vectors as described in Example 1. The sgRNAs were tested with nucleases having SEQ ID NOs: 20, 24, and 26 as described in Example 3 (Table 8). The results for each of SEQ ID NOs: 20, 24, and 26 are shown in Table 8, respectively. FIG. 5A to FIG. 5C The results with the additional sequence of SEQ ID NO: 20 are shown in Figure 5D The putative structures and modifications of sgRNA are shown in Figure 5E Surprisingly, some modifications such as those in SEQ ID NO: 346 removed the predicted stem-loop, allowing the sgRNA construct to function well with a variety of nucleases. Also surprisingly, multiple truncations located in the stem and upper loop retained functionality when paired with the nuclease SEQ ID NO: 20.

[0335] Example 7

[0336] Guide RNA modification

[0337] According to the method described in Example 3, the editing activities of nucleases with SEQ ID NOs: 20, 24, 26 and Un1Cas12f1 (SEQ ID NO: 471) were compared at different target sites using sgRNA with SEQ ID NO: 346. The results are shown in Figure 6The results showed that each nuclease was able to edit to different levels at different genomic target sites. Surprisingly, when paired with the sgRNA with SEQ ID NO: 346, Un1Cas12f1 did not show editing above background levels at the Kim-T1 site, while the other three nucleases showed editing activity against this sgRNA.

[0338] Example 8

[0339] TracrRNA modification

[0340] The editing activities of nucleases SEQ ID NOs: 20 and 21 were compared with sgRNAs with small deletions in the tracrRNA sequence as described in Example 3. The tracrRNA deletion and editing results are shown in Table 9.

[0341] Nuclease SEQ ID NO: 20 was then tested for various sgRNA modifications that altered the predicted structure of the tracrRNA sequence. Two configurations with either longer repeats or truncated repeats were tested (see Fig. 7A ) and compared with a modification with a truncated 5' stem (SEQ ID NO: 346). Notably, having a full repeat sequence was detrimental to editing activity compared to the other truncated forms ( Figure 7B ).

[0342] To further investigate the relationship of the tracrRNA sequence to these nucleases, further modifications were made. Starting from SEQ ID NO: 346, a portion of the 5' stem and 3' tail of the tracrRNA were removed to evaluate their importance in editing efficiency ( Figure 7C Further removal of the 5' stem did not affect editing, whereas removal of the 3' tail of the tracrRNA was highly deleterious for editing and had efficiencies similar to those observed for non-targeted cells ( Fig.7D ).

[0343] To further evaluate the role of the bases of the stem, the sequence was modified by changing AT to GC to display “stem stability” and separately by removing the kink inserted by the single nucleotide at the unpaired A to enhance base pairing ( Figure 7C ). Increasing the stability of the stem changed the predicted ΔG of the structure, but it did not improve the editing efficiency of the nuclease SEQ ID NO:20. Removal of the A-kink completely abolished the editing ability of the nuclease ( Fig. 7E ).

[0344] Example 9

[0345] Spacer modification

[0346] The editing activity of nuclease SEQ ID NO: 20 was evaluated for sgRNAs with varying lengths of intervening sequences according to the method described in Example 3. The editing results are shown in Figure 8 Medium. A spacer length of 18–20 nucleotides is optimal for editing activity.

[0347] Example 10

[0348] PAM Preferences

[0349] According to the method of using spacer 3 of Walton et al., the effect of PAM sequence on the efficiency of nuclease editing was tested. (Walton RT et al., Science. April 17, 2020; 368(6488): 290-296, the entire document is incorporated herein by reference). In short, the spacer capable of targeting the randomized PAM plasmid library is incorporated downstream of the repeat region of TracrRNA and gRNA, and the randomized PAM plasmid library is prepared with a 10-bp randomized PAM. In this process, the effective PAM of the nuclease is exhausted, and the remaining PAM is displayed by next generation sequencing (NGS). The preferred PAM sequences of nuclease SEQ ID NO: 20 and 26 are listed in Table 10. Based on the calculated values ​​described by Walton et al., and the PAM preference is listed in the order of preference (the top of each list indicates a more preferred sequence).

[0350] In the case of multiple spacers in the sgRNA, the editing activity of the identified PAM sequences was tested with nucleases SEQ ID NOs: 20 and 26. Fig. 9A and Fig. 9B Target sequences with higher editing levels are shown (X axis) ( Fig. 9A ) and target sequences with lower editing levels ( Fig. 9B ) and various PAM sequences (PAM sequences are shown in brackets on the bars). Surprisingly, the nuclease has a PAM preference different from that of known Cas12f nucleases such as Un1Cas12f1, AsCas12f, and SpaCas12f1. For the tested nucleases (SEQ ID NOs: 20, 21, and 26), the preferred PAM sequence is DTTR, where D is A, G, or T, and R is A or G; there is a strong preference for ATTA PAM. In contrast, for Un1Cas12f1 and AsCas12f, the PAM preference is TTTR, while for SpaCas12f1, the PAM preference is NTTY, where N can be any base.

[0351] Embodiment 11

[0352] AAV vector design and editing in mammalian cells

[0353] A single AAV vector was designed to deliver the nuclease of SEQ ID NO: 20 and sgRNA to mammalian cells using a CMV promoter and an SV40 nuclear localization sequence at the 5' end of the nuclease and an HA tag and a nucleoplasmic protein localization sequence at the 3' end, followed by a U6 promoter to drive expression of the sgRNA (at Fig.10 The vector is shown as Tracr in Fig.10 middle.

[0354] Using this vector design, a set of constructs with the same nuclease but with different sgRNAs designed for different targets were constructed, as shown in Table 11.

[0355] Constructs for human targets were tested in HEK293T cells and constructs for mouse targets were tested in NIH3T3 cells. On day 0, cells were plated at 3x10 5 The cells were plated at a confluence of 10 cells / ml. On day 1, the cells were transduced at an MOI of 100K. On day 2, etoposide (to enhance AAV delivery) was added to the cells to a final concentration of 60 mM, and the cells were imaged on day 3. The cells were cultured for 72 hours and then harvested according to the method of Example 2. After DNA extraction, samples for NGS were prepared by amplifying each region with NGS-specific primers listed in Table 12. NGS reads were processed using the CRISPRESSO2 tool (Clement, Kendell et al., Nature biotechnology 37.3 (2019): 224-226, which is incorporated herein by reference in its entirety). The edited data for each construct are shown in Fig.11 middle.

[0356] SMN2 and TTR constructs were further tested for editing in HEK293T cells and NIH3T3 cells with and without etoposide treatment. Etoposide was added to cells on day 1, AAV vectors were added on day 2, and cells were harvested on day 7 as described above, but at an MOI of 10K. Samples were prepared for NGS using the primers in Table 9. NGS paired reads were processed using CRISPRESSO2 (Clement et al., 2019). Editing efficiencies are shown in Fig.12 NIH3T3 cells tolerated etoposide treatment, and in general, editing was improved in treated cells. In contrast, HEK293T cells showed signs of toxicity, and editing was reduced in treated cells compared to cells not treated with etoposide.

[0357] Table 2

[0358]

[0359]

[0360]

[0361]

[0362]

[0363]

[0364]

[0365]

[0366]

[0367]

[0368] Table 4

[0369]

[0370]

[0371]

[0372]

[0373]

[0374]

[0375]

[0376]

[0377]

[0378]

[0379]

[0380]

[0381]

[0382]

[0383]

[0384]

[0385]

[0386]

[0387]

[0388]

[0389]

[0390]

[0391]

[0392]

[0393]

[0394]

[0395]

[0396]

[0397]

[0398]

[0399]

[0400]

[0401]

[0402]

[0403]

[0404]

[0405]

[0406]

[0407]

[0408]

[0409]

[0410]

[0411]

[0412]

[0413]

[0414]

[0415]

[0416]

[0417]

[0418]

[0419]

[0420]

[0421]

[0422]

[0423] Table 5

[0424]

[0425]

[0426]

[0427]

[0428]

[0429]

[0430]

[0431]

[0432]

[0433]

[0434] Table 6

[0435]

[0436]

[0437] Table 7

[0438]

[0439]

[0440] Table 9

[0441]

[0442] Table 10: PAM sequence preferences

[0443]

[0444] Table 11: Constructs with the nuclease of SEQ ID NO: 20 for AAV studies with sgRNA targeting PRSS1, SMN2, PCSK9 and TTR.

[0445]

[0446] Table 12: Amplification primer sequences

[0447]

[0448]

[0449] The scope of the present invention is not limited by the content specifically shown and described above. Those skilled in the art will recognize that there are suitable alternatives for the examples of described materials, configurations, constructions and sizes. Without departing from the spirit and scope of the present invention, those of ordinary skill in the art can think of changes, modifications and other embodiments of the content described herein.

[0450] Many references, including patents and various publications, are cited and discussed in the specification of the present invention. The citation and discussion of these references are provided only to illustrate the description of the present invention, and no admission is made that any reference is prior art to the present invention described herein. All references cited and discussed in this specification are incorporated herein by reference in their entirety.

Claims

1. A composition comprising a nuclease, wherein the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 99% identical to any one of SEQ ID NOs: 1-250.

2. The composition of claim 1, wherein the amino acid sequence of the nuclease comprises any one of SEQ ID NOs: 1-250.

3. The composition of claim 1 or 2, wherein the nuclease further comprises a nuclear localization sequence (NLS) at the N-terminus, the C-terminus, or both the N-terminus and the C-terminus of the nuclease.

4. The composition of claim 3, wherein the NLS at the N-terminus and the NLS at the C-terminus of the nuclease are different sequences. 5 . A nucleic acid comprising a first polynucleotide sequence encoding the nuclease according to claim 1 . A vector comprising the nucleic acid according to claim 5 .

7. The vector of claim 6, further comprising a promoter operably linked to the first polynucleotide.

8. The vector of claim 6 or 7, further comprising a second polynucleotide sequence encoding a guide RNA (gRNA).

9. The vector of claim 8, further comprising a promoter operably linked to the second polynucleotide sequence.

10. The vector of claim 8 or 9, wherein the gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical to any one of SEQ ID NOs: 251-422 and 472-482.

11. The vector of any one of claims 8 to 10, wherein the gRNA comprises any one of SEQ ID NOs: 251-343.

12. The vector of any one of claims 8 to 10, wherein the gRNA comprises any one of SEQ ID NOs: 344-422.

13. The vector of any one of claims 8 to 10, wherein the gRNA comprises any one of SEQ ID NOs: 472-482.

14. The vector of any one of claims 8 to 13, wherein the gRNA comprises a tracr sequence, and the gRNA comprises one or more sequence deletions in or near a region comprising the tracr sequence.

15. The vector of claim 14, wherein the one or more sequence deletions include sequences predicted to form a stem-loop structure.

16. The vector of claim 14 or 15, wherein the one or more sequence deletions include a sequence predicted to form a stem-loop structure at or near the 5' end of the gRNA.

17. The vector of any one of claims 14 to 16, wherein the gRNA comprises SEQ ID NO:

346.

18. The vector of any one of claims 14 to 16, wherein the gRNA comprises SEQ ID NO:

420.

19. The vector of any one of claims 14 to 16, wherein the gRNA comprises SEQ ID NO:

481.

20. The vector of any one of claims 14 to 16, wherein the gRNA comprises SEQ ID NO:

479.

21. The vector of any one of claims 8 to 20, wherein the gRNA comprises a spacer sequence of at least 18 nucleotides in length or between 18 and 20 nucleotides in length.

22. A system for modifying a target nucleic acid, the system comprising: a) a nuclease or a nucleic acid encoding the same, the nuclease comprising an amino acid sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 1-250; and b) at least one guide RNA (gRNA) or a nucleic acid encoding the at least one gRNA, the gRNA comprising a sequence complementary to at least a portion of the target nucleic acid and a region that associates with the nuclease.

23. The system of claim 22, wherein the nuclease is capable of recognizing a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA and CTTG.

24. The system of claim 22 or 23, wherein the gRNA comprises a spacer sequence complementary to the first strand sequence of the target nucleic acid, and wherein the first strand sequence is directly adjacent to a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG.

25. The system of claim 23 or 24, wherein the PAM sequence comprises DTTR, wherein D is A, G, or T, and R is A or G.

26. The system of any one of claims 22 to 25, wherein the nuclease is capable of preferentially modifying a target nucleic acid comprising the PAM sequence ATTA, wherein R is A or G, compared to a target nucleic acid comprising the PAM sequence TTTR.

27. The system of any one of claims 22 to 25, wherein the nuclease is capable of modifying the target nucleic acid more efficiently than the efficiency of modifying the target nucleic acid with the nuclease SEQ ID NO: 471, wherein the PAM sequence contained in the target nucleic acid is ATTA.

28. The system of any one of claims 22 to 27, wherein modification comprises nucleic acid cleavage.

29. The system of any one of claims 22 to 28, wherein modification comprises one or more of modification of the target nucleic acid, regulation of transcription of the target nucleic acid, and modification of a polypeptide associated with the target nucleic acid.

30. The system of any one of claims 22 to 29, wherein the nuclease further comprises a nuclear localization sequence (NLS) at the N-terminus, the C-terminus, or both the N-terminus and the C-terminus of the nuclease.

31. The system of claim 30, wherein the NLS at the N-terminus and the NLS at the C-terminus of the nuclease are different sequences.

32. The system of any one of claims 22 to 31, wherein the nuclease further comprises a purification tag.

33. The system of any one of claims 22 to 32, wherein the at least one gRNA further comprises a sequence complementary to at least a portion of a second target nucleic acid.

34. The system of any one of claims 22 to 33, wherein the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 251-422.

35. The system of claim 34, wherein the at least one gRNA comprises any one of SEQ ID NOs: 251-343.

36. The system of claim 34, wherein the at least one gRNA comprises any one of SEQ ID NOs: 344-422.

37. The system of claim 34, wherein the at least one gRNA comprises any one of SEQ ID NOs: 472-482.

38. The system of claim 34, wherein the at least one gRNA comprises SEQ ID NO:

346.

39. The system of claim 34, wherein the at least one gRNA comprises SEQ ID NO:

420.

40. The system of claim 34, wherein the at least one gRNA comprises SEQ ID NO:

481.

41. The system of claim 34, wherein the at least one gRNA comprises SEQ ID NO:

479.

42. The system of any one of claims 22 to 41, wherein the at least one gRNA comprises a spacer sequence of at least 18 nucleotides in length or between 18 and 20 nucleotides in length.

43. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 20 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479 and 481, or any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392 and 417, or any one of SEQ ID NOs: 346 and 362, or any one of SEQ ID NOs: 410-419.

44. The system of any one of claims 22 to 43, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 20, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 309, 346, 352, 358, 362-364, 380, 392-395, 410-420, 472-479 and 481 or any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392 and 417 or any one of SEQ ID NOs: 346 and 362 or any one of SEQ ID NOs: 410-419.

45. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 21 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical, or 100% identical to any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482.

46. ​​The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 21, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422 and 479-482.

47. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 22 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 311, 346, 381, and 398-399.

48. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 22, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381 and 398-399.

49. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 23 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 312, 346 and 382.

50. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:23, and wherein the at least one gRNA comprises any one of SEQ ID NOs:312, 346 and 382.

51. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 24 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 310, 313, 325, 346, 350-355, 358, 361-363, 367-372, and 389-392 or any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.

52. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 24, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 350-355, 358, 361-363, 367-372 and 389-392 or any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369 and 392.

53. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 25 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 314, 346, 383 and 400.

54. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 25, and wherein the at least one gRNA comprises any one of SEQ ID NO: 314, 346, 383 and 400.

55. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 26 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479 and 481 or any one of SEQ ID NOs: 346, 384 and 392.

56. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 26, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479 and 481 or any one of SEQ ID NOs: 346, 384 and 392.

57. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 27 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 316, 346, 385 and 401.

58. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 27, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385 and 401.

59. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 28 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 317, 346, 386 and 402.

60. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 28, and wherein the at least one gRNA comprises any one of SEQ ID NO: 317, 346, 386 and 402.

61. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 29 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 318, 346, 387 and 403.

62. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 29, and wherein the at least one gRNA comprises any one of SEQ ID NO: 318, 346, 387 and 403.

63. The system of any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 36 and the at least one gRNA comprises a sequence that is at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identical or 100% identical to any one of SEQ ID NOs: 310, 313, 325, 346, 356-360, and 373-378.

64. The system of any one of claims 22 to 42, wherein the nuclease comprises a sequence at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 36, and wherein the at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356-360 and 373-378.

65. The system of any one of claims 22 to 64, wherein the nucleic acid molecule encoding each or both of the nuclease and the at least one gRNA comprises a messenger RNA, a vector, or a combination thereof.

66. The system of any one of claims 22 to 65, wherein the nuclease and the at least one gRNA are encoded on one nucleic acid.

67. The system of claim 66, wherein the nuclease and the at least one gRNA are operably linked to different promoters.

68. The system of claim 66 or 67, wherein the one nucleic acid is a vector.

69. The system of claim 68, wherein the vector is a viral vector.

70. The system of claim 69, wherein the viral vector is an AAV vector.

71. A kit comprising the system of any one of claims 22 to 70.

72. A cell comprising the system of any one of claims 22 to 70.

73. The cell of claim 72, wherein the cell is a prokaryotic cell or a eukaryotic cell.

74. The cell of claim 72 or 73, wherein the cell is a mammalian cell.

75. The cell of any one of claims 72 to 74, wherein the cell is a human cell.

76. A method of modifying a selected target nucleic acid sequence, the method comprising contacting the selected target nucleic acid with a composition as described in any one of claims 1 to 4, a nucleic acid as described in claim 5, a vector as described in any one of claims 6 to 21, or a system as described in any one of claims 22 to 70.

77. The method of claim 76, wherein the target nucleic acid sequence is located in a cell.

78. The method of claim 77, wherein the cell is a prokaryotic cell or a eukaryotic cell.

79. The method of claim 77 or 78, wherein the cell is a mammalian cell.

80. The method of any one of claims 76 to 78, wherein the cell is a human cell.

81. The method of any one of claims 76 to 80, wherein the contacting comprises introducing into the cell a composition as described in any one of claims 1 to 4, a nucleic acid as described in claim 5, a vector as described in any one of claims 6 to 21, or a system as described in any one of claims 22 to 69.

82. The method of any one of claims 75 to 80, wherein the contacting comprises administering to the subject a composition of any one of claims 1 to 4, a nucleic acid of claim 5, a vector of any one of claims 6 to 21, or a system of any one of claims 22 to 70.

83. The method of any one of claims 76 to 82, wherein the selected target nucleic acid sequence encodes a gene product.

84. The composition of any one of claims 1 to 4, the nucleic acid of claim 5, the vector of any one of claims 6 to 21, or the system of any one of claims 22 to 70 for use in modifying a selected target nucleic acid sequence.

85. A kit comprising the composition of any one of claims 1 to 4, the nucleic acid of claim 5, the vector of any one of claims 6 to 21, or the system of any one of claims 22 to 70 for modifying a selected target nucleic acid sequence in an in vitro assay.

Citation Information

Patent Citations

  • Methods of generating nucleic acid fragments

    US10087431B2

  • Methods and compositions for enhancing nuclease-mediated gene disruption

    US10227610B2

  • Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription

    US10266850B2

  • Information processing method and device

    US10601748B2

  • Delivery methods and compositions for nuclease-mediated genome engineering

    US10604771B2