Compositions and Methods for Nucleic Acid Modification
The use of nucleases with high sequence identity and guide RNAs with complementary sequences, optionally with nuclear localization sequences, addresses the inefficiencies and specificity issues of current CRISPR/Cas systems, enhancing gene editing in eukaryotic cells.
Patent Information
- Application Number
- JP2024573092
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-02
- Filing Date
- 2023-06-09
- Publication Date
- 2025-06-26
AI Technical Summary
Current CRISPR/Cas systems face limitations such as low editing efficiency, off-target events, target sequence preference, and challenges in efficient delivery and expression of nucleases, especially in eukaryotes.
A composition comprising a nuclease with a sequence having at least 70% to 99% identity to specific SEQ ID NOs, potentially combined with a guide RNA (gRNA) having complementary sequences, and optionally including a nuclear localization sequence (NLS) for enhanced targeting and delivery.
The described nuclease and gRNA system achieves improved gene editing efficiency and specificity, potentially overcoming the limitations of existing CRISPR/Cas systems, particularly in eukaryotic cells.
Smart Images

Figure 2025519628000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to nucleases for nucleic acid modification and compositions, methods, and systems thereof.
[0002] Cross - reference to related applications This application claims the benefit of U.S. Provisional Application No. 63 / 351,140, filed on June 10, 2022, U.S. Provisional Application No. 63 / 383,107, filed on November 10, 2022, and U.S. Provisional Application No. 63 / 482,936, filed on February 2, 2023, the contents of which are hereby incorporated by reference in their entirety.
[0003] Description regarding the sequence listing The contents of the electronic sequence listing entitled ACRIG_404894_601.xml (size: 579,833 bytes; and creation date: June 8, 2023) are hereby incorporated by reference in their entirety.
Background Art
[0004] Clustered regularly interspaced short palindromic repeat (CRISPR) - associated (Cas) nucleases occupy a dominant position in the context of nucleic acid editing because they are versatile, rapid, and easy - to - use editing tools. Cas9, the most well - characterized CRISPR - Cas nuclease, uses one or more RNAs to act as a sequence - specific targeting element that links the nuclease to the target nucleic acid. However, current CRISPR / Cas systems have several limitations in use, including low editing efficiency, off - target events, target sequence preference, and efficient delivery and expression of nucleases, especially in eukaryotes.
Summary of the Invention
[0005] Provided herein is a composition comprising a nuclease, wherein the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or more than 99% identity to any one of SEQ ID NOs: 1-250. In some embodiments, the amino acid sequence of the nuclease comprises any one of SEQ ID NOs: 1-250.
[0006] In some embodiments, the nuclease further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is at the N-terminus, C-terminus or both the N-terminus and C-terminus of the nuclease. In some embodiments, the NLS at the N-terminus and the NLS at the C-terminus of the nuclease are different sequences.
[0007] Also provided are a nucleic acid molecule comprising a first polynucleotide sequence encoding a nuclease and a vector comprising the nucleic acid molecule. In some embodiments, the vector further comprises a promoter operably linked to the first polynucleotide sequence. In some embodiments, the vector further comprises a second polynucleotide sequence encoding a guide RNA (gRNA). In some embodiments, the vector further comprises a promoter operably linked to the second polynucleotide.
[0008] In some embodiments, the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 251-422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 251-343. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 344-422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 472-482. In some embodiments, the gRNA comprises SEQ ID NO: 346, 420, 481, or 479.
[0009] In some embodiments, the gRNA comprises a tracr sequence, and the gRNA comprises one or more sequence deletions within or near the region encompassing the tracr sequence. In some embodiments, the one or more sequence deletions comprise sequences predicted to form a stem-loop structure. In some embodiments, the one or more sequence deletions comprise sequences predicted to form a stem-loop structure at or near the 5' end of the gRNA. In some embodiments, the gRNA comprises SEQ ID NO: 346, 420, 481, or 479.
[0010] In some embodiments, the gRNA comprises a spacer sequence that is at least 18 nucleotides in length. In some embodiments, the gRNA comprises a spacer sequence that is 18 to 20 nucleotides in length.
[0011] In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 410 - 419.
[0012] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 20, and at least one gRNA comprises any one of SEQ ID NO: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NO: 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NO: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NO: 410 - 419.
[0013] In some embodiments, the nuclease comprises SEQ ID NO: 21, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NO: 310, 344 - 349, 361 - 366, 404 - 422, and 479 - 482. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and at least one gRNA comprises any one of SEQ ID NO: 310, 344 - 349, 361 - 366, 404 - 422, and 479 - 482.
[0014] In some embodiments, the nuclease comprises SEQ ID NO: 22, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 311, 346, 381, and 398 - 399. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381, and 398 - 399.
[0015] In some embodiments, the nuclease comprises SEQ ID NO: 23, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 23, and at least one gRNA comprises any one of SEQ ID NOs: 312, 346, and 382.
[0016] In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 350 - 355, 358, 361 - 363, 367 - 372, and 389 - 392. In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.
[0017] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 24, and at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 350 - 355, 358, 361 - 363, 367 - 372, and 389 - 392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 24, and at least one gRNA comprises any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.
[0018] In some embodiments, the nuclease comprises SEQ ID NO: 25, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 25, and at least one gRNA comprises any one of SEQ ID NOs: 314, 346, 383, and 400.
[0019] In some embodiments, the nuclease comprises SEQ ID NO: 26, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 26, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 346, 384, and 392.
[0020] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and at least one gRNA comprises any one of SEQ ID NOs: 346, 384, and 392.
[0021] In some embodiments, the nuclease comprises SEQ ID NO: 27, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385, and 401.
[0022] In some embodiments, the nuclease comprises SEQ ID NO: 28, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 28, and at least one gRNA comprises any one of SEQ ID NOs: 317, 346, 386, and 402.
[0023] In some embodiments, the nuclease comprises SEQ ID NO: 29, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and at least one gRNA comprises any one of SEQ ID NOs: 318, 346, 387, and 403.
[0024] In some embodiments, the nuclease comprises SEQ ID NO: 36, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 36, and at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378.
[0025] Furthermore, provided is a system for modifying a first target nucleic acid, comprising: a) a nuclease comprising an amino acid sequence having 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, more than 99% or 100% identity to any one of SEQ ID NOs: 1 - 250, or a first nucleic acid sequence encoding the nuclease; and b) at least one guide RNA (gRNA) comprising a sequence complementary to at least a portion of the first target nucleic acid and a region that associates with the nuclease, or a nucleic acid encoding at least one gRNA.
[0026] In some embodiments, the nuclease can recognize a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. In some embodiments, the gRNA comprises a spacer sequence complementary to the first strand sequence of the target nucleic acid, wherein the first strand sequence is directly adjacent to a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. In some embodiments, the PAM sequence comprises DTTR, wherein D is A, G, or T, and R is A or G.
[0027] In some embodiments, the nuclease can preferentially modify a first target nucleic acid containing the PAM sequence ATTA as compared to a first target nucleic acid containing the PAM sequence TTTR (where R is A or G).
[0028] In some embodiments, the nuclease can modify the target nucleic acid with higher efficiency as compared to the efficiency of the modification of the target nucleic acid by the nuclease of SEQ ID NO: 471, where the target nucleic acid contains a PAM sequence that is ATTA.
[0029] In some embodiments, the nuclease can modify a first target nucleic acid in the presence of the gRNA. In some embodiments, the modification includes nucleic acid cleavage. In some embodiments, the modification includes one or more of modification of the target nucleic acid, regulation of transcription from the target nucleic acid, and modification of a polypeptide associated with the target nucleic acid.
[0030] In some embodiments, the nuclease further includes a nuclear localization sequence (NLS). In some embodiments, the NLS is at the N-terminus, C-terminus, or both the N-terminus and C-terminus of the nuclease. In some embodiments, the NLS at the N-terminus and the NLS at the C-terminus of the nuclease are different sequences. In some embodiments, the nuclease further includes a purification tag.
[0031] In some embodiments, the gRNA further includes a sequence complementary to at least a portion of a second target nucleic acid.
[0032] In some embodiments, the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 251 - 422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 251 - 343. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 344 - 422. In some embodiments, the gRNA comprises any one of SEQ ID NOs: 472 - 482. In some embodiments, the gRNA comprises SEQ ID NO: 346, 420, 481, or 479.
[0033] In some embodiments, the gRNA comprises a tracr sequence and the gRNA comprises one or more sequence deletions within or near the region encompassing the tracr sequence. In some embodiments, the one or more sequence deletions comprise sequences predicted to form a stem - loop structure. In some embodiments, the one or more sequence deletions comprise sequences predicted to form a stem - loop structure at or near the 5' end of the gRNA. In some embodiments, the gRNA comprises SEQ ID NO: 346, 420, 481, or 479.
[0034] In some embodiments, the gRNA comprises a spacer sequence that is at least 18 nucleotides in length. In some embodiments, the gRNA comprises a spacer sequence that is 18 - 20 nucleotides in length.
[0035] In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 410 - 419.
[0036] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 20, and at least one gRNA comprises any one of SEQ ID NOs: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392, and 417. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 346 and 362. In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 410 - 419.
[0037] In some embodiments, the nuclease comprises SEQ ID NO: 21, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 310, 344 - 349, 361 - 366, 404 - 422, and 479 - 482. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and at least one gRNA comprises any one of SEQ ID NOs: 310, 344 - 349, 361 - 366, 404 - 422, and 479 - 482.
[0038] In some embodiments, the nuclease comprises SEQ ID NO: 22, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 311, 346, 381, and 398-399. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381, and 398-399.
[0039] In some embodiments, the nuclease comprises SEQ ID NO: 23, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 23, and at least one gRNA comprises any one of SEQ ID NOs: 312, 346, and 382.
[0040] In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 350 - 355, 358, 361 - 363, 367 - 372, and 389 - 392. In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.
[0041] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 350 - 355, 358, 361 - 363, 367 - 372, and 389 - 392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and at least one gRNA comprises any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392.
[0042] In some embodiments, the nuclease comprises SEQ ID NO: 25, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 25, and at least one gRNA comprises any one of SEQ ID NOs: 314, 346, 383, and 400.
[0043] In some embodiments, the nuclease comprises SEQ ID NO: 26, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 315, 346, 384, 392, 396 - 397, 420, 479, and 481. In some embodiments, the nuclease comprises SEQ ID NO: 26, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 346, 384, and 392.
[0044] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396 - 397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and at least one gRNA comprises any one of SEQ ID NOs: 346, 384, and 392.
[0045] In some embodiments, the nuclease comprises SEQ ID NO: 27, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385, and 401.
[0046] In some embodiments, the nuclease comprises SEQ ID NO: 28, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 28, and at least one gRNA comprises any one of SEQ ID NOs: 317, 346, 386, and 402.
[0047] In some embodiments, the nuclease comprises SEQ ID NO: 29, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and at least one gRNA comprises any one of SEQ ID NOs: 318, 346, 387, and 403.
[0048] In some embodiments, the nuclease comprises SEQ ID NO: 36, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 36, and at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378.
[0049] In some embodiments, the nucleic acid molecule encoding one or both of the nuclease and the gRNA is a DNA molecule such as a vector, plasmid, or linear nucleic acid. In some embodiments, the nuclease is encoded by a messenger RNA. In some embodiments, the gRNA is included in a small molecule RNA.
[0050] In some embodiments, the nuclease and the gRNA are encoded on the same nucleic acid. In some embodiments, the nuclease and the gRNA are encoded on different nucleic acids.
[0051] Also provided are vectors comprising the disclosed system. In some embodiments, the vector further comprises a first promoter operably linked to a nucleic acid encoding a nuclease and a second promoter operably linked to a nucleic acid encoding at least one gRNA. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is an AAV vector. In some embodiments, the first promoter and the second promoter are active in mammalian cells.
[0052] In some embodiments, the system further comprises a target nucleic acid.
[0053] In some embodiments, the system is a cell-free system.
[0054] Also provided are cells comprising the disclosed compositions and systems. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are eukaryotic cells (e.g., mammalian cells or human cells).
[0055] Further provided is a method for modifying a target nucleic acid, the method comprising contacting the target nucleic acid with a nuclease, composition, vector, or system described herein.
[0056] In some embodiments, the target nucleic acid sequence is intracellular. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell (e.g., mammalian cells or human cells).
[0057] In some embodiments, introducing the system or composition into a cell comprises administering the system or composition to a subject. In some embodiments, the administration comprises in vivo administration.
[0058] Also provided is a kit comprising any or all of the components of the compositions or systems described herein. In some embodiments, the kit further comprises one or more reagents, transport containers and / or packaging containers, one or more buffers, delivery devices, instructions, software, computing devices, or combinations thereof.
[0059] Other aspects and embodiments of the present disclosure will become apparent in light of the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0060]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 5C
Figure 5D
Figure 5E
Figure 6
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Figure 7E
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
BRIEF DESCRIPTION OF THE INVENTION
[0061] The disclosed compositions, systems, kits, and methods include nucleases useful for nucleic acid modification. The disclosed nucleases enable gene editing with improved efficiency and safety for in vivo and ex vivo uses in the treatment, diagnosis, and research of eukaryotes (e.g., mammals (e.g., humans)).
[0062] The section headings used in this section and throughout the disclosure of this specification are for organizational purposes only and are not intended to be limiting.
[0063] Definitions The terms “comprise(s),” “include(s),” “having,” “has,” “can,” “contain(s),” and variations thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. As used herein, including a particular sequence or a particular SEQ ID NO. generally means that at least one copy of the sequence is present in the recited peptide or polynucleotide. However, two or more copies are also contemplated. The singular forms “a,” “and,” and “the” include plural referents unless the context clearly dictates otherwise. The present disclosure contemplates other embodiments “comprising,” “consisting of,” and “consisting essentially of” the embodiments or elements presented herein, whether or not explicitly recited.
[0064] Regarding the references to numerical ranges in this specification, each numerical value intervening between them is explicitly assumed with the same degree of precision. For example, in the case of the range of 6 to 9, in addition to 6 and 9, the numbers 7 and 8 are also contemplated, and in the case of the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are also explicitly contemplated.
[0065] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings commonly understood by those of ordinary skill in the art. For example, any academic and technical terms and techniques used in connection with the culturing of cells and tissues, molecular biology, microbiology, genetics, and the chemistry and hybridization of proteins and nucleic acids described herein are well known and commonly used in the relevant technical fields. The meanings and scopes of the terms need to be clear, but if any potential ambiguity arises, the definitions provided herein shall take precedence over any dictionary or external definition. Further, unless otherwise required by the context, singular terms shall include the plural, and plural terms shall include the singular.
[0066] As used herein, "nucleic acid" or "nucleic acid sequence" refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, at 793-800 (Worth Pub. 1982)). The technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, and any chemical variants thereof (e.g., methylated, hydroxymethylated, or glycosylated forms of these bases). The polymer or oligomer can have a heterogeneous or homogeneous composition and can be isolated from a naturally occurring source or generated artificially or synthetically. Further, the nucleic acid can be DNA or RNA, or a mixture thereof, and can exist permanently or transiently in single-stranded or double-stranded forms, including homoduplex, heteroduplex, and hybrid states. In some embodiments, the nucleic acid or nucleic acid sequence includes other types of nucleic acid structures such as, for example, DNA / RNA helices, peptide nucleic acids (PNA), morpholino nucleic acids (see, e.g., Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002), and U.S. Patent No. 5,034,506), locked nucleic acids (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. U.S.A., 97:5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122:8595-8602 (2000)) and / or ribozymes.Accordingly, the term "nucleic acid" or "nucleic acid sequence" may also include a strand containing unnatural nucleotides, modified nucleotides, and / or non-nucleotide building blocks (e.g., "nucleotide analogs") that can exhibit the same function as natural nucleotides; further, the term "nucleic acid sequence" as used herein refers to oligonucleotides, nucleotides, or polynucleotides, and fragments or portions thereof, as well as DNA or RNA of genomic or synthetic origin, which may be single-stranded or double-stranded and may represent sense or antisense strands. The terms "nucleic acid", "polynucleotide", "nucleotide sequence", and "oligonucleotide" are used interchangeably. These refer to polymeric forms of nucleotides of any length, which may be either deoxyribonucleotides or ribonucleotides, or analogs thereof.
[0067] The "identity" of a nucleic acid or amino acid sequence described herein can be determined by comparing the nucleic acid or amino acid sequence of interest to a reference nucleic acid or reference amino acid sequence. The percent identity is obtained by dividing the number of nucleotide or amino acid residues that are the same (e.g., identical) between the sequence of interest and the reference sequence by the length of the longest sequence (e.g., the length of the longer of the sequence of interest or the reference sequence) to obtain an optimal alignment. Numerous mathematical algorithms are known for calculating the identity between two or more sequences and are incorporated into many available software programs. Examples of such programs include CLUSTAL-W, T-Coffee, and ALIGN (for alignment of nucleic acid and amino acid sequences), BLAST programs (e.g., BLAST 2.1, BL2SEQ, and their later versions), and FASTA programs (e.g., FASTA 3x, FAS™, and SSEARCH) (for sequence alignment and sequence similarity searching). Sequence alignment algorithms are also disclosed, for example, in Altschul et al., J. Molecular Biol., 215(3):403-410 (1990), Beigert et al., Proc. Natl. Acad. Sci. USA, 106(10):3770-3775 (2009), Durbin et al., eds., Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press, Cambridge, UK (2009), Soding, Bioinformatics, 21(7):951-960 (2005), Altschul et al., Nucleic Acids Res., 25(17):3389-3402 (1997), and Gusfield, Algorithms on Strings, Trees and Sequences, Cambridge University Press, Cambridge UK (1997)).
[0068] The terms "non-naturally occurring", "engineered", and "synthetic" are used interchangeably and indicate the involvement of human hands. When referring to a nucleic acid molecule or polypeptide, these terms mean that the nucleic acid molecule or polypeptide does not contain, at least substantially, at least one other component that is naturally associated and found in nature, and / or that the nucleic acid molecule or polypeptide is associated with at least one other component that is not naturally associated and / or has one or more changes in the sequence compared to the nucleic acid or amino acid sequence found in nature.
[0069] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, that can bind or incorporate another DNA segment, e.g., an "insert", and effect the replication of the ligated segment intracellularly.
[0070] A cell is "genetically modified", "transformed", or "transfected" by exogenous DNA, e.g., a recombinant expression vector, when introduced into the cell. The presence of exogenous DNA results in a permanent or transient genetic change. The transforming DNA may or may not be integrated (covalently bound) into the genome of the cell. For example, the transforming DNA may be maintained on an episomal element such as a plasmid. For eukaryotic cells, a stably transformed cell is a cell in which the transforming DNA has been integrated into the chromosome and is passed on to daughter cells through chromosomal replication. This stability is demonstrated by the ability of eukaryotic cells to establish cell lines or clones containing a population of daughter cells containing the DNA to be transformed. A "clone" is a population of cells derived from a single cell or common ancestor by mitosis. A "cell line" is a clone of primary cells that can grow stably in vitro for many generations.
[0071] As used herein, the term "contacting" refers to bringing about contact, resulting in contact, being in a state of contact, or coming into a state of contact. The term "contacting" as used herein refers to a situation or state of being in contact or immediately adjacent or locally proximate. Contact of a composition to a target destination such as, but not limited to, an organ, tissue, cell, or tumor can be effected by any means of administration known to those of skill in the art.
[0072] As used herein, the terms "providing", "administering", and "introducing" are used interchangeably herein and refer to disposing a composition or system of the present disclosure to a cell, organism, or subject by a method or route that results in at least partial localization to a desired site. The composition or system can be administered by any suitable route that results in delivery to the desired location of the cell, organism, or subject.
[0073] Preferred methods and materials are described below, but methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and not intended to be limiting.
[0074] Nuclease Advances and developments in CRISPR-Cas genome editing tools, including nucleases and other Cas proteins, have led to significant progress in nucleic acid editing. Nucleic acid editing has many applications, including in the fields of diagnosis and therapy. Such breadth is accompanied by a diversity of nucleic acid targets and environments for manipulating editing activity. Accordingly, there is a need for diverse and additional nucleases and related methods that provide a toolbox for nucleic acid editing.
[0075] Disclosed herein are compositions comprising a nuclease having Cas-like activity. The disclosed nuclease comprises a sequence having at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 93%, at least 95%, at least 98%, at least 99%, or 100% identity) to the amino acid sequence of SEQ ID NOs: 1-250. In some embodiments, the nuclease comprises a sequence having at least 90% identity to the amino acid sequence of SEQ ID NOs: 1-250. In certain embodiments, the nuclease comprises the amino acid sequence of SEQ ID NOs: 1-250.
[0076] Any of the nucleases described herein may include one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 150, etc.) amino acid substitutions. An "exchange" or "substitution" of an amino acid refers to replacing one amino acid at a given position or residue within a polypeptide sequence with another amino acid at the same position or residue. Amino acids are broadly classified as "aromatic" or "aliphatic". Aromatic amino acids contain an aromatic ring. Examples of "aromatic" amino acids include histidine (H or His), phenylalanine (F or Phe), tyrosine (Y or Tyr), and tryptophan (W or Trp). Non-aromatic amino acids are broadly classified as "aliphatic". Examples of "aliphatic" amino acids include glycine (G or Gly), alanine (A or Ala), valine (V or Val), leucine (L or Leu), isoleucine (I or Ile), methionine (M or Met), serine (S or Ser), threonine (T or Thr), cysteine (C or Cys), proline (P or Pro), glutamic acid (E or Glu), aspartic acid (A or Asp), asparagine (N or Asn), glutamine (Q or Gln), lysine (K or Lys), and arginine (R or Arg).
[0077] Amino acid replacements or substitutions can be conservative, semi-conservative, or non-conservative. The phrase "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid by another amino acid with common properties. A functional way to define the common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of the same species (Schulz and Schirmer, Principles of Protein Structure, Springer-Verlag, New York (1979)). According to such an analysis, amino acids within a group are preferentially exchanged with each other, and thus groups of amino acids with the most similar effects on the overall protein structure can be defined (Schulz and Schirmer (supra)). Examples of conservative amino acid substitutions include substitutions of amino acids within the above-mentioned subgroups, for example, the substitution of arginine with lysine and vice versa so that the positive charge can be maintained, the substitution of aspartic acid with glutamic acid and vice versa so that the negative charge can be maintained, the substitution of threonine with serine so that the free -OH can be maintained, and the substitution of asparagine with glutamine so that the free -NH2 can be maintained. "Semi-conservative mutations" include amino acid substitutions within the same group listed above, but not amino acid substitutions within the same subgroup. For example, the substitution of asparagine with aspartic acid, or the substitution of lysine with asparagine, involves amino acids within the same group but different subgroups. "Non-conservative mutations" involve amino acid substitutions between different groups, for example, those involving tryptophan to lysine, or serine to phenylalanine, etc.
[0078] In some embodiments, the nuclease comprises one or more amino acid substitutions and has an amino acid sequence that has at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 93%, at least 95%, at least 98%, at least 99% identity, or 100% identity) to the amino acid sequence of SEQ ID NOs: 1-250. In some embodiments, the nuclease comprises one or more amino acid substitutions compared to SEQ ID NOs: 1-250, and one or more of the substitutions improve the editing efficiency of the nuclease.
[0079] The nucleases disclosed herein may be capable of recognizing a wide range of protospacer adjacent motifs (PAMs) adjacent to the target nucleic acid. In certain embodiments, the nuclease can cleave the target nucleic acid only when an appropriate PAM is present. In certain embodiments, the nuclease has a broad ability for recognition of target nucleic acids, such as target nucleic acids lacking a PAM, or broad PAM recognition.
[0080] The PAM is generally adjacent to the target sequence. For example, the PAM can be the sequence immediately adjacent to or directly adjacent to the target nucleic acid. The PAM can be 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. In one embodiment, the target nucleic acid is located immediately adjacent to the PAM at the 3' end. In one embodiment, the target nucleic acid is located immediately adjacent to the PAM at the 5' end.
[0081] The PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. In certain embodiments, the PAM is 2-6 nucleotides in length.
[0082] Non-limiting examples of PAM sequences include CC, CA, AG, GT, TA, AC, CA, GC, CG, GG, CT, TG, GA, AGG, TGG, T-rich PAMs (e.g., TTT, TTG, TTC, etc.), NGG, NGA, NAG, NGGNG and NNAGAAW, NNNNGATT, NAAR (R = A or G), NNGRR (R = A or G), NNAGAA and NAAAAC, where "N" is any nucleotide.
[0083] In some embodiments, the nucleases disclosed herein are capable of recognizing a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. In some embodiments, the PAM sequence includes DTTR, where D is A, G, or T, and R is A or G.
[0084] Different PAM sequences can confer different preferences and efficiencies for nuclease cleavage or modification by a desired nuclease. In some embodiments, the nuclease preferentially modifies a first target nucleic acid comprising the PAM sequence ATTA as compared to a target nucleic acid comprising the PAM sequence TTTR (R is A or G). In some embodiments, a more efficient modification of the target nucleic acid by the nucleases disclosed herein is observed as compared to the efficiency of modification by the nuclease of SEQ ID NO: 471. In some embodiments, when the target nucleic acid is a PAM sequence that is ATTA, a more efficient modification of the target nucleic acid by the nucleases disclosed herein is observed as compared to the modification efficiency by the nuclease of SEQ ID NO: 471.
[0085] In some embodiments, the nuclease further comprises a nuclear localization sequence (NLS). The nuclear localization sequence can be added, for example, to one or both of the N-terminus and the C-terminus. In some embodiments, the nuclease comprises two or more NLSs. The two or more NLSs may be present tandemly, separated by a linker, at either the N-terminus or the C-terminus of the protein, or one or more may be within the open reading frame of the nuclease.
[0086] The nuclear localization sequence can include any amino acid sequence known in the art that functions to tag a protein or direct a protein for translocation to the cell nucleus (e.g., nuclear transport). Typically, the nuclear localization sequence includes one or more positively charged amino acids such as lysine and arginine.
[0087] In some embodiments, the NLS is a monopartite sequence. A monopartite NLS includes a single cluster of positively charged or basic amino acids. In some embodiments, the monopartite NLS includes the sequence K-K / R-X-K / R, where X can be any amino acid. Exemplary monopartite NLS sequences include those derived from the SV40 large T antigen, c-Myc, and the TUS protein. In a selected embodiment, the NLS includes the NLS of the SV40 large T antigen, having the amino acid sequence PKKKRKV (SEQ ID NO: 504).
[0088] In some embodiments, the NLS is a bipartite sequence. A bipartite NLS includes two clusters of basic amino acids separated by a spacer of about 9-12 amino acids. Exemplary bipartite NLSs include the nuclear localization sequences of nucleoplasmin, EGL-12, or bipartite SV40. In a selected embodiment, the NLS includes the NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK (SEQ ID NO: 505).
[0089] In some embodiments, two or more NLSs can have the same or different sequences. For example, in some embodiments, the nuclease comprises two NLSs, one being a sequence derived from the SV40 large T antigen and one being a sequence derived from nucleoplasmin.
[0090] The NLS can be added to the nuclease by a linker. The linker can be a polypeptide of any amino acid sequence and length. The linker can act as a spacer peptide. In some embodiments, the linker is flexible. In some embodiments, the linker comprises at least one glycine and at least one serine. In some embodiments, the linker comprises an amino acid sequence consisting of (Gly2Ser) n where n is the number of repeats including integers from 2 to 20.
[0091] In some embodiments, the nuclease can comprise a tag (e.g., 3xFLAG tag, HA tag, Myc tag, etc.). The tag can facilitate tracking, isolation, or purification of the nuclease. In some embodiments, the tag may be adjacent to either upstream or downstream of the nuclear localization sequence. The tag can be at the N-terminus, C-terminus, or a combination thereof of the nuclease.
[0092] In some embodiments, the nuclease is covalently attached to a peptide or protein of a fusion protein. The nuclease may be part of a fusion protein that includes another protein or protein domain. For example, the nuclease can be fused to another protein or protein domain that provides tagging or visualization (e.g., GFP). The nuclease can be fused to a protein or protein domain having another functionality or activity useful for targeting to a particular DNA sequence (e.g., nuclease activity as provided by a FokI nuclease, protein modification activities such as histone modification activities including acetylation or deacetylation or demethylation or methyltransferase activity, transcriptional regulatory activities such as the activity of a transcriptional activator or transcriptional repressor, base editing activities such as deaminase activity, DNA modification activities such as DNA methylation activity, etc.).
[0093] In some embodiments, the nuclease can be fused to a protein transduction domain or PTD, also known as one or more (e.g., 2, 3, 4, or more) CPPs (cell-penetrating peptides). A protein transduction domain is a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates passage across a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule facilitates the passage of the molecule across a membrane, e.g., from the extracellular space into the intracellular space or from the cytosol into an organelle. In some embodiments, the PTD is covalently linked to the terminus (e.g., N-terminus, C-terminus, or both) of the nuclease. In some embodiments, the PTD is inserted internally at a suitable insertion site. Examples of PTDs include the minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT); polyarginine sequences containing a sufficient number of arginines (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines) to induce entry into cells; the VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); the Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); the truncated human calcitonin peptide (Trehin et al. (2004) Pharm. Research 21:1248-1256); polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008); transportan, etc., but are not limited thereto.
[0094] The nuclease can be fused via a linker polypeptide. The linker polypeptide can have any of various amino acid sequences. Proteins can generally be linked by a flexible spacer peptide, although other chemical linkages are not excluded. Suitable linkers include polypeptides that are 4 to 40 amino acids in length, or 4 to 25 amino acids in length. These linkers can also be generated by using oligonucleotides encoding synthetic linkers that link proteins, or may be encoded by nucleic acid sequences encoding fusion proteins. Peptide linkers having a certain degree of flexibility can be used. The linking peptide can have substantially any amino acid sequence, but it should be noted that preferred linkers generally have sequences that result in flexible peptides. The use of small amino acids such as glycine and alanine is useful for making flexible peptides. The production of such sequences is routine for those skilled in the art. Various different linkers are commercially available and considered suitable for use, including but not limited to glycine-serine polymers, glycine-alanine polymers, and alanine-serine polymers.
[0095] Compositions and Systems Also disclosed herein are compositions comprising a nuclease described herein or a nucleic acid molecule comprising a sequence encoding such nuclease.
[0096] Furthermore, disclosed herein is a system for modifying a target nucleic acid, comprising a nuclease described herein (e.g., a nuclease comprising an amino acid sequence having at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 93%, at least 95%, at least 98%, at least 99% identity or 100% identity) to the amino acid sequences of SEQ ID NOs: 1-250) or a nucleic acid molecule comprising a sequence encoding such nuclease.
[0097] In some embodiments, the components of the system can be in the form of a composition. In some embodiments, the components of the present composition or system can be mixed with a carrier, either individually or in any combination, and these are also within the scope of the present disclosure. Exemplary carriers include buffers, antioxidants, preservatives, carbohydrates, surfactants, and the like.
[0098] Also disclosed are cells comprising the compositions or systems described herein. In some embodiments, the cells are prokaryotic cells. In some embodiments, the cells are eukaryotic cells. In some embodiments, the cells are mammalian cells. In some embodiments, the cells are human cells.
[0099] The compositions or systems disclosed herein may further comprise at least one gRNA comprising a sequence complementary to at least a portion of a first target nucleic acid and a region that associates with a nuclease, or a nucleic acid encoding at least one gRNA. In some embodiments, the at least one gRNA further comprises a sequence complementary to at least a portion of a second target nucleic acid. If the composition or system comprises multiple gRNAs, each of them can be encoded by the same or a different nucleic acid from the other gRNAs.
[0100] The gRNA can be a crRNA, a crRNA / tracrRNA (or single-guide RNA, sgRNA). The terms "gRNA", "guide RNA", and "CRISPR guide sequence" can be used interchangeably throughout and refer to a nucleic acid that associates with a nuclease and contains a sequence that determines the sequence specificity of the nuclease. The gRNA can be engineered to hybridize to a target nucleic acid sequence (e.g., the genome within a host cell) (e.g., be partially or fully complementary).
[0101] In some embodiments, at least one gRNA is encoded in a CRISPR RNA (crRNA) array. A CRISPR array contains a series of unidirectional repeats separated by short sequences called spacers. The nucleases described herein may prefer unidirectional repeat sequences. For example, a CRISPR RNA (crRNA) may contain multiple gRNAs or multiple different sequences, each configured to hybridize to a different target nucleic acid sequence.
[0102] The gRNA or a portion thereof that hybridizes to the target nucleic acid (target site) can be 15 to 40 nucleotides in length. In some embodiments, the gRNA sequence that hybridizes to the target nucleic acid is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides in length. The gRNA or sgRNA(s) used in this disclosure can be about 5-100 nucleotides in length, or longer (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 11 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in length or more).
[0103] In addition to the sequence that binds to the target nucleic acid, in some embodiments, the gRNA may include a scaffold sequence (e.g., tracrRNA). In some embodiments, such chimeric gRNAs may be referred to as single-guide RNAs (sgRNAs). Exemplary scaffold sequences will be apparent to those skilled in the art and can be found, for example, in Jinek, et al. Science (2012) 337(6096):816-821, and Ran, et al. Nature Protocols (2013) 8:2281-2308, which are hereby incorporated by reference in their entirety.
[0104] In some embodiments, the gRNA sequence does not include a scaffold sequence, and the scaffold sequence is expressed as a separate transcript. In such embodiments, the gRNA sequence further includes an additional sequence that is complementary to a portion of the scaffold sequence and functions to bind (hybridize) to the scaffold sequence.
[0105] In some embodiments, the gRNA includes a sequence that is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to the target nucleic acid. In some embodiments, the sequence is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to the 3' end of the target nucleic acid (e.g., the last 5, 6, 7, 8, 9, or 10 nucleotides of the 3' end of the target nucleic acid).
[0106] In some embodiments, the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 251-422 and 472-482. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 251-343. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 344-422. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 472-482.
[0107] The gRNA of the present disclosure may comprise a sequence having one or more nucleotide substitutions or mutations, deletions, or insertions relative to any of SEQ ID NOs: 251-343. The nucleotide substitutions or mutations, deletions, or insertions may increase stability, modify secondary structure elements, increase binding efficiency to cognate nucleases or target strands, and may increase. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 344-422. In some embodiments, at least one gRNA comprises any one or more of SEQ ID NOs: 472-482. In some embodiments, the gRNA comprises SEQ ID NO: 346. In some embodiments, the gRNA comprises SEQ ID NO: 420. In some embodiments, the gRNA comprises SEQ ID NO: 481. In some embodiments, the gRNA comprises SEQ ID NO: 479.
[0108] In some embodiments, the gRNA comprises a spacer sequence. The spacer sequence can be of any length or sequence. In some embodiments, the spacer sequence is at least 18 (e.g., 18, 19, 20, 21, 22, 23, 24, etc.) nucleotides in length. In some embodiments, the spacer sequence is 18 to 20 nucleotides in length. Thus, in certain embodiments, the spacer sequence is 18 nucleotides in length. In certain embodiments, the spacer sequence is 19 nucleotides in length. In certain embodiments, the spacer sequence is 20 nucleotides in length.
[0109] In some embodiments, the gRNA comprises a spacer sequence complementary to the first strand sequence of the target nucleic acid. In some embodiments, the first strand sequence is directly adjacent to a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG.
[0110] In some embodiments, the nuclease comprises SEQ ID NO: 21, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 310, 344 - 349, 361 - 366, 404 - 422, and 479 - 482. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and at least one gRNA comprises any one of SEQ ID NOs: 310, 344 - 349, 361 - 366, 404 - 422, and 479 - 482. In some embodiments, the nuclease comprises SEQ ID NO: 21 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or SEQ ID NO: 346.
[0111] In some embodiments, the nuclease comprises SEQ ID NO: 24, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 350 - 355, 358, 361 - 363, 367 - 372, and 389 - 392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 350 - 355, 358, 361 - 363, 367 - 372, and 389 - 392. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and at least one gRNA comprises any one of SEQ ID NOs: 346, 352, 358, 361, 362, 368, 369, and 392. In some embodiments, the nuclease comprises SEQ ID NO: 24 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and the gRNA comprises SEQ ID NO: 346 or a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346. In some embodiments, the nuclease comprises SEQ ID NO: 24 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and the gRNA comprises SEQ ID NO: 352 or a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 352.
[0112] In some embodiments, the nuclease comprises SEQ ID NO: 36, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 36, and at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378. In some embodiments, the nuclease comprises SEQ ID NO: 36 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 36, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or a sequence having the same identity to SEQ ID NO: 346. In some embodiments, the nuclease comprises SEQ ID NO: 36 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 36, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 358 or a sequence having the same identity to SEQ ID NO: 358.
[0113] In some embodiments, the nuclease comprises SEQ ID NO: 1, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 251-256. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 1, and at least one gRNA comprises any one of SEQ ID NOs: 251-256.
[0114] In some embodiments, the nuclease comprises SEQ ID NO: 2, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 257-259. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 2, and at least one gRNA comprises any one of SEQ ID NOs: 257-259.
[0115] In some embodiments, the nuclease comprises SEQ ID NO: 3, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 260-262. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 3, and at least one gRNA comprises any one of SEQ ID NOs: 260-262.
[0116] In some embodiments, the nuclease comprises SEQ ID NO: 4, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 263-265. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 4, and at least one gRNA comprises any one of SEQ ID NOs: 263-265.
[0117] In some embodiments, the nuclease comprises SEQ ID NO: 5, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 266-268. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 5, and at least one gRNA comprises any one of SEQ ID NOs: 266-268.
[0118] In some embodiments, the nuclease comprises SEQ ID NO: 6, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 269-271. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 6, and at least one gRNA comprises any one of SEQ ID NOs: 269-271.
[0119] In some embodiments, the nuclease comprises SEQ ID NO: 7, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 272-274. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 7, and at least one gRNA comprises any one of SEQ ID NOs: 272-274.
[0120] In some embodiments, the nuclease comprises SEQ ID NO: 8, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 275-277. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 8, and at least one gRNA comprises any one of SEQ ID NOs: 275-277.
[0121] In some embodiments, the nuclease comprises SEQ ID NO: 9, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 278-280. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 9, and at least one gRNA comprises any one of SEQ ID NOs: 278-280.
[0122] In some embodiments, the nuclease comprises SEQ ID NO: 10, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 281 - 283. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 10, and at least one gRNA comprises any one of SEQ ID NOs: 281 - 283.
[0123] In some embodiments, the nuclease comprises SEQ ID NO: 11, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 284 - 286. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 11, and at least one gRNA comprises any one of SEQ ID NOs: 284 - 286.
[0124] In some embodiments, the nuclease comprises SEQ ID NO: 12, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 287 - 289. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 12, and at least one gRNA comprises any one of SEQ ID NOs: 287 - 289.
[0125] In some embodiments, the nuclease comprises SEQ ID NO: 13, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 290-292. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 13, and at least one gRNA comprises any one of SEQ ID NOs: 290-292.
[0126] In some embodiments, the nuclease comprises SEQ ID NO: 14, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 293-295. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 14, and at least one gRNA comprises any one of SEQ ID NOs: 293-295.
[0127] In some embodiments, the nuclease comprises SEQ ID NO: 15, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 296-298. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 15, and at least one gRNA comprises any one of SEQ ID NOs: 296-298.
[0128] In some embodiments, the nuclease comprises SEQ ID NO: 16, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 299-301. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 16, and at least one gRNA comprises any one of SEQ ID NOs: 299-301.
[0129] In some embodiments, the nuclease comprises SEQ ID NO: 17, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 302-304. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 17, and at least one gRNA comprises any one of SEQ ID NOs: 302-304.
[0130] In some embodiments, the nuclease comprises SEQ ID NO: 18, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 305-307. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 18, and at least one gRNA comprises any one of SEQ ID NOs: 305-307.
[0131] In some embodiments, the nuclease comprises SEQ ID NO: 19, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to either SEQ ID NO: 308 or 379. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity to SEQ ID NO: 19, and at least one gRNA comprises either SEQ ID NO: 308 or 379.
[0132] In some embodiments, the nuclease comprises SEQ ID NO: 20, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 20, and at least one gRNA comprises any one of SEQ ID NOs: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479 and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 20, and at least one gRNA comprises any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392, and 417, or any one of SEQ ID NOs: 346 and 362, or any one of SEQ ID NOs: 410 - 419. In some embodiments, the nuclease comprises SEQ ID NO: 20 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 20, and the gRNA comprises SEQ ID NO: 346 or a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346.
[0133] In some embodiments, the nuclease comprises SEQ ID NO: 22, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 311, 346, 381, and 398-399. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381, and 398-399. In some embodiments, the nuclease comprises SEQ ID NO: 22 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or SEQ ID NO: 346.
[0134] In some embodiments, the nuclease comprises SEQ ID NO: 23, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 23, and at least one gRNA comprises any one of SEQ ID NOs: 312, 346, and 382. In some embodiments, the nuclease comprises SEQ ID NO: 23 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 23, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or SEQ ID NO: 346.
[0135] In some embodiments, the nuclease comprises SEQ ID NO: 25, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 25, and at least one gRNA comprises any one of SEQ ID NOs: 314, 346, 383, and 400. In some embodiments, the nuclease comprises SEQ ID NO: 25 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 25, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or SEQ ID NO: 346.
[0136] In some embodiments, the nuclease comprises SEQ ID NO: 26, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 315, 346, 384, 392, 396 - 397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396 - 397, 420, 479, and 481. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and at least one gRNA comprises any one of SEQ ID NOs: 346, 384, and 392.
[0137] In some embodiments, the nuclease comprises SEQ ID NO: 26 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and the gRNA comprises SEQ ID NO: 346 or a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346.
[0138] In some embodiments, the nuclease comprises SEQ ID NO: 27, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385, and 401. In some embodiments, the nuclease comprises SEQ ID NO: 27 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or SEQ ID NO: 346.
[0139] In some embodiments, the nuclease comprises SEQ ID NO: 28, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 28, and at least one gRNA comprises any one of SEQ ID NOs: 317, 346, 386, and 402. In some embodiments, the nuclease comprises SEQ ID NO: 28 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 28, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or SEQ ID NO: 346.
[0140] In some embodiments, the nuclease comprises SEQ ID NO: 29, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and at least one gRNA comprises any one of SEQ ID NOs: 318, 346, 387, and 403. In some embodiments, the nuclease comprises SEQ ID NO: 29 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 346 or SEQ ID NO: 346.
[0141] In some embodiments, the nuclease comprises SEQ ID NO: 30, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 319. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 30, and at least one gRNA comprises SEQ ID NO: 319.
[0142] In some embodiments, the nuclease comprises SEQ ID NO: 31, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 320. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 31, and at least one gRNA comprises SEQ ID NO: 320.
[0143] In some embodiments, the nuclease comprises SEQ ID NO: 32, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 321. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 32, and at least one gRNA comprises SEQ ID NO: 321.
[0144] In some embodiments, the nuclease comprises SEQ ID NO: 33, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 322. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 33, and at least one gRNA comprises SEQ ID NO: 322.
[0145] In some embodiments, the nuclease comprises SEQ ID NO: 34, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to either SEQ ID NO: 323 or 388. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 34, and at least one gRNA comprises either SEQ ID NO: 323 or 388.
[0146] In some embodiments, the nuclease comprises SEQ ID NO: 35, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 324. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 35, and at least one gRNA comprises SEQ ID NO: 324.
[0147] In some embodiments, the nuclease comprises SEQ ID NO: 37, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 326. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 37, and at least one gRNA comprises SEQ ID NO: 326.
[0148] In some embodiments, the nuclease comprises SEQ ID NO: 38, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 327. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 38, and at least one gRNA comprises SEQ ID NO: 327.
[0149] In some embodiments, the nuclease comprises SEQ ID NO: 39, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 328. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 39, and at least one gRNA comprises SEQ ID NO: 328.
[0150] In some embodiments, the nuclease comprises SEQ ID NO: 40, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 329. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 40, and at least one gRNA comprises SEQ ID NO: 329.
[0151] In some embodiments, the nuclease comprises SEQ ID NO: 41, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 330. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 41, and at least one gRNA comprises SEQ ID NO: 330.
[0152] In some embodiments, the nuclease comprises SEQ ID NO: 42, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 331. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 42, and at least one gRNA comprises SEQ ID NO: 331.
[0153] In some embodiments, the nuclease comprises SEQ ID NO: 43, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 332. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 43, and at least one gRNA comprises SEQ ID NO: 332.
[0154] In some embodiments, the nuclease comprises SEQ ID NO: 44, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 333. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 44, and at least one gRNA comprises SEQ ID NO: 333.
[0155] In some embodiments, the nuclease comprises SEQ ID NO: 45, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 334. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 45, and at least one gRNA comprises SEQ ID NO: 334.
[0156] In some embodiments, the nuclease comprises SEQ ID NO: 46, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 335. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 46, and at least one gRNA comprises SEQ ID NO: 335.
[0157] In some embodiments, the nuclease comprises SEQ ID NO: 47, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 336. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 47, and at least one gRNA comprises SEQ ID NO: 336.
[0158] In some embodiments, the nuclease comprises SEQ ID NO: 48, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 337. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 48, and at least one gRNA comprises SEQ ID NO: 337.
[0159] In some embodiments, the nuclease comprises SEQ ID NO: 49, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 338. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 49, and at least one gRNA comprises SEQ ID NO: 338.
[0160] In some embodiments, the nuclease comprises SEQ ID NO: 50, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 339. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 50, and at least one gRNA comprises SEQ ID NO: 339.
[0161] In some embodiments, the nuclease comprises SEQ ID NO: 51, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 340. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 51, and at least one gRNA comprises SEQ ID NO: 340.
[0162] In some embodiments, the nuclease comprises SEQ ID NO: 52, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 341. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 52, and at least one gRNA comprises SEQ ID NO: 341.
[0163] In some embodiments, the nuclease comprises SEQ ID NO: 53, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 342. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 53, and at least one gRNA comprises SEQ ID NO: 342.
[0164] In some embodiments, the nuclease comprises SEQ ID NO: 54, and at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 343. In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 54, and at least one gRNA comprises SEQ ID NO: 343.
[0165] In some embodiments, the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to any of SEQ ID NOs: 1-19 and 30-54 or any of SEQ ID NOs: 1-19 and 30-54, and the gRNA comprises SEQ ID NO: 346 or a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to SEQ ID NO: 346.
[0166] In some embodiments, the gRNA described herein may comprise one or more nucleotide substitutions or mutations (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, etc.) relative to any of SEQ ID NOs: 251-343.
[0167] In some embodiments, the gRNA comprises one or more cleavages or deletions of one or more nucleotides relative to any of SEQ ID NOs: 251-343. The cleavage or deletion can be at one or both of the 3' and 5' ends of the sequence, or within or inside a sequence associated with any of SEQ ID NOs: 251-343. The cleavage or deletion may include a single nucleotide, or may include a deletion or cleavage of a series of two or more consecutive nucleotides (e.g., 2, 3, 4, 5, 10, 15, 20, etc.). In some embodiments, the gRNA of the present invention may comprise a cleavage sequence corresponding to or presumed to be a crRNA:tracrRNA stem.
[0168] In some embodiments, the gRNA comprises a tracr sequence. The gRNA may include one or more sequence deletions within or near the region encompassing the tracr sequence. For example, the one or more sequence deletions may include sequences predicted to form a stem-loop structure. In some embodiments, the one or more sequence deletions include sequences predicted to form a stem-loop structure at or near the 5' end of the gRNA. In some embodiments, the gRNA comprises SEQ ID NO: 346. In some embodiments, the gRNA comprises SEQ ID NO: 420. In some embodiments, the gRNA comprises SEQ ID NO: 481. In some embodiments, the gRNA comprises SEQ ID NO: 479.
[0169] In some embodiments, the gRNA comprises one or more insertions or additions of one or more nucleotides relative to any of SEQ ID NOs: 251-343. The insertion or addition can be at one or both of the 3' and 5' ends of the sequence, or within a sequence associated with any of SEQ ID NOs: 251-343. The insertion or addition may include a single nucleotide, or may include a deletion or cleavage of a series of two or more consecutive nucleotides (e.g., 2, 3, 4, 5, 10, 15, 20, etc.). In some embodiments, the gRNA of the present invention may include an artificial stem-loop between the crRNA and the tracrRNA.
[0170] The gRNA can be a non-naturally occurring gRNA.
[0171] In certain embodiments, engineering a nuclease for use in eukaryotic cells can involve codon optimization. It will be appreciated that by changing native codons to those most frequently used in mammals, maximal expression of the system protein can be achieved in mammalian cells (e.g., human cells). Nucleic acid sequences so modified are generally described in the art as utilizing “codon-optimized” codons, or “mammal-preferred” or “human-preferred” codons. In some embodiments, a nucleic acid sequence is considered to be codon-optimized if at least about 60% (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 98%) of the codons encoded by the sequence are mammal-preferred codons.
[0172] In some cases, the compositions or systems disclosed herein may further comprise a donor polynucleotide. For example, in applications where it is desirable to insert a polynucleotide sequence into a genome where a target sequence is cleaved, a donor polynucleotide (a nucleic acid comprising a donor sequence) may also be provided to the cell. A "donor sequence" or "donor polynucleotide" or "donor template" means a nucleic acid sequence that is inserted at the site targeted by the nuclease (e.g., after dsDNA cleavage, after nicking of the target DNA, after double nicking of the target DNA, etc.). In some cases, the donor sequence is provided to the cell as single-stranded DNA. In some cases, the donor template is provided to the cell as double-stranded DNA. It can be introduced into the cell in linear or circular form. When introduced in linear form, the ends of the donor sequence can be protected by any convenient method (e.g., from exonuclease digestion), such methods being known to those of skill in the art. For example, one or more dideoxynucleotide residues can be added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides can be ligated to one or both ends. The donor template can be introduced into the cell, for example, as part of a vector molecule having additional sequences such as an origin of replication, a promoter, and a gene encoding antibiotic resistance. Further, the donor template can be introduced as a naked nucleic acid, as a nucleic acid complexed with an agent such as a liposome or a poloxamer, or can be delivered by a virus (e.g., an adenovirus, AAV).
[0173] The present disclosure also provides one or more nucleic acids encoding the nucleases and gRNAs disclosed herein, vectors containing these nucleic acids, and cells containing the vectors. The vectors can be used to propagate the segments in a suitable cell and / or to enable expression from the segments (e.g., expression vectors). Those of skill in the art will be aware of the various vectors available for the propagation and expression of nucleic acid sequences.
[0174] In some embodiments, the one or more nucleic acids comprise one or more messenger RNAs, one or more vectors, or any combination thereof. In some embodiments, the one or more nucleic acids comprise messenger RNA for the expression of a nuclease, and at least one nucleic acid provides a gRNA. A single nucleic acid may encode a nuclease and at least one gRNA, or the nuclease may be encoded on a nucleic acid different from the at least one gRNA.
[0175] In some embodiments, the nuclease is provided as a split nuclease such that two separate proteins come together to form a functional nuclease (e.g., the nuclease can be delivered in some cases as a split nuclease, or a nucleic acid(s) encoding a split nuclease). In some such cases, the sequences encoding the two parts of the split nuclease protein are present on the same vector. In some cases, the sequences are present on separate vectors, e.g., as part of a vector system encoding a nuclease, gRNA(s), and its system.
[0176] The present disclosure further provides engineered non-naturally occurring vectors and vector systems that can encode one or more or all of the components of the system. The vector(s) can be introduced into a cell comprising any suitable prokaryotic or eukaryotic cell capable of expressing the polypeptide encoded by the vector.
[0177] The vectors of the present disclosure can be delivered to eukaryotic cells of a subject, e.g., a mammalian subject, e.g., a human subject. Modification of eukaryotic cells via the system can be performed in cell culture.
[0178] Using viral and non-viral based gene transfer methods, nucleic acids encoding the components of the system can be introduced into cells, tissues, or subjects. Such methods can be used to administer nucleic acids encoding the components of the system to cells in culture or in a host organism. Non-viral vector delivery systems include DNA plasmids, cosmids, RNA (e.g., transcripts of the vectors described herein), nucleic acids, and nucleic acids complexed with delivery vehicles. Viral vector delivery systems include DNA and RNA viruses and have either episomal or integrated genomes after delivery to the cell. Viral vectors include, for example, retroviruses, lentiviruses, adenoviruses, adeno-associated viruses and herpes simplex virus vectors.
[0179] In certain embodiments, plasmids that are non-replicating or that can be hardened at high temperatures can be used, whereby, under certain conditions, any or all of the necessary components of the composition or system can be removed from the cells. For example, this can enable integration of DNA by transforming the bacterium of interest, but leaving an engineered strain that does not carry the memory of the plasmid or vector used for integration.
[0180] A variety of viral constructs can be used to deliver the present compositions or systems (such as nucleases and one or more gRNAs (plural available)) to target cells and / or subjects. Non-limiting examples of such recombinant viruses include recombinant adeno-associated virus (AAV), recombinant adenovirus, recombinant lentivirus, recombinant retrovirus, recombinant herpes simplex virus, recombinant poxvirus, phage, and the like. The present disclosure provides vectors capable of integrating into the host genome, such as retroviruses or lentiviruses. See, for example, Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1989; Kay, M.A., et al., 2001 Nat. Med. 7(1):33-40; and Walther W. and Stein U., 2000 Drugs, 60(2):249-71, which are hereby incorporated by reference herein.
[0181] In one embodiment, the DNA segment encoding the nuclease is contained within a plasmid vector that allows for the expression of the protein and subsequent isolation and purification of the protein produced by the recombinant vector. Thus, the nucleases disclosed herein can be purified after expression, obtained by chemical synthesis, or obtained by recombinant methods.
[0182] To construct a cell expressing the present system, an expression vector for stable or transient expression of the present system or any of its components can be constructed by the methods described herein or methods known in the art and introduced into cells. For example, the nucleic acid encoding the components of the present system can be cloned into a suitable expression vector such as a plasmid or viral vector operably linked to a suitable promoter. The choice of expression vector / plasmid / viral vector must be suitable for integration and replication in eukaryotic cells. In some embodiments, a single nucleic acid comprises a first promoter operably linked to a nuclease and a second promoter operably linked to a gRNA. In some cases, the single nucleic acid is a vector.
[0183] In certain embodiments, one or more promoters can drive the expression of one or more sequences (e.g., nuclease and / or gRNA) in prokaryotic cells. Promoters that can be used include the T7 RNA polymerase promoter, constitutive E. coli promoters, and promoters that can be widely recognized by the transcriptional machinery of a wide range of bacterial organisms. The compositions or systems can be used in various bacterial hosts.
[0184] In certain embodiments, one or more promoters can drive the expression of one or more sequences (e.g., nucleases and / or gRNAs) in mammalian cells, such as when contained in a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, Nature (1987) 329:840, incorporated herein by reference) and pMT2PC (Kaufman, et al., EMBO J. (1987) 6:187, incorporated herein by reference). When used in mammalian cells, the control functions of the expression vector are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other expression systems suitable for both prokaryotic and eukaryotic cells, see, for example, Chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL. 2nd eds., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989, which are incorporated herein by reference.
[0185] The promoter for use in the expression of nucleases and gRNAs in this specification may include any of a number of promoters known in the art, and the promoter may be constitutive, regulatable or inducible, cell-type specific, tissue-specific, or species-specific. In addition to the sequence sufficient to induce transcription, the promoter sequences of the present invention may also include the sequences of other regulatory elements involved in the regulation of transcription (e.g., enhancers, Kozak sequences, and introns). Many promoter / regulatory sequences useful for driving the constitutive expression of genes are available in the art, for example, CMV (cytomegalovirus promoter), EF1a (human elongation factor 1 alpha promoter), SV40 (simian vacuolating virus 40 promoter), PGK (mammalian phosphoglycerate kinase promoter), Ubc (human ubiquitin C promoter), human beta-actin promoter, rodent beta-actin promoter, CBh (chicken beta-actin promoter), CAG (hybrid promoter includes CMV enhancer, chicken beta-actin promoter, and rabbit beta-globin splice acceptor), TRE (tetracycline response element promoter), H1 (human polymerase III RNA promoter), U6 (human U6 small nuclear promoter), and the like, but are not limited thereto. Additional promoters that can be used for the expression of the components of this system include, but are not limited to, the cytomegalovirus (CMV) immediate early promoter, viral LTRs, such as Rous sarcoma virus LTR, HIV-LTR, HTLV-1 LTR, Moloney murine leukemia virus (MMLV) LTR, myeloproliferative sarcoma virus (MPSV) LTR, spleen focus-forming virus (SFFV) LTR, simian virus 40 (SV40) early promoter, herpes simplex tk virus promoter, elongation factor 1-alpha (EF1-α) promoter with or without the EF1-α intron. The additional promoters include constitutively active promoters. Alternatively, any regulatable promoter capable of regulating expression intracellularly may be used.In embodiments, a polymerase II promoter (e.g., CMV promoter) is used to drive the expression of the nuclease, and a polymerase III promoter (e.g., U6 promoter) is used to drive the expression of the gRNA.
[0186] To achieve an appropriate balance (ratio of expression levels) between the components of the system (e.g., nuclease, at least one gRNA), different promoters and regulatory elements can be used. For example, in some cases, the nucleic acid comprises a promoter and regulatory elements operably linked (and thus controlling / regulating its translation) to the sequence encoding the nuclease. In some cases, the target nucleic acid comprises a promoter and regulatory elements operably linked to encode the gRNA. In some cases, both the sequence encoding the nuclease and the sequence encoding the gRNA are operably linked to the same promoter and regulatory elements.
[0187] Various types of promoters are suitable for use. The promoter may be a constitutively active promoter (e.g., a promoter that is constitutively active / “on”), an inducible promoter (e.g., a promoter whose active / “on” or inactive / “off” state is controlled by an external stimulus, such as the presence of a particular temperature, compound, or protein), a spatially restricted promoter (e.g., a tissue-specific promoter, a cell-type specific promoter, etc.), or a temporally restricted promoter (e.g., a promoter that is “on” or “off” during a particular stage of embryonic development or a particular stage of a biological process, such as the hair follicle cycle in a mouse).
[0188] Furthermore, the inducible and tissue-specific expression of RNA or protein can be achieved by placing the nucleic acid encoding the molecule under the control of an inducible or tissue-specific promoter / regulatory sequence. A promoter can induce the expression of a nucleic acid in a particular cell type (e.g., a tissue-specific regulatory element is used for the expression of the nucleic acid). Such regulatory elements include promoters that can be tissue-specific or cell-type specific. The term "tissue-specific" as applied to a promoter refers to a promoter that can induce the selective expression of a nucleotide sequence of interest in a particular tissue type (e.g., a seed), and in which the expression of the same nucleotide sequence of interest is relatively absent in different tissue types. The term "cell-type specific" as applied to a promoter refers to a promoter that can induce the selective expression of a nucleotide sequence of interest in a particular cell type, and in which the expression of the same nucleotide sequence of interest is relatively absent in different cell types within the same tissue. The term "cell-type specific" when applied to a promoter also means a promoter that can promote the selective expression of a nucleotide sequence of interest in a region within a single tissue. The cell-type specificity of a promoter can be evaluated using methods well known in the art, such as immunohistochemical staining.
[0189] Examples of tissue-specific or inducible promoters / regulatory sequences useful for this purpose include, but are not limited to, the rhodopsin promoter, the MMTV LTR inducible promoter, the SV40 late enhancer / promoter, the synapsin 1 promoter, the ET hepatocyte promoter, the GS glutamine synthetase promoter, and many others. Not only tissue-specific promoters and tumor-specific promoters, but also various commercially available ubiquitous promoters are available, for example, obtainable from InvivoGen. In addition, promoters well known in the art that can be induced in response to inducers such as metals, glucocorticoids, tetracyclines, hormones, etc. are also contemplated for use with the present invention. Thus, it will be understood that the present disclosure includes the use of any promoter / regulatory sequence known in the art that is capable of driving the expression of a desired nuclease or gRNA operably linked thereto.
[0190] Examples of spatially restricted promoters include, but are not limited to, neuron-specific promoters, adipocyte-specific promoters, cardiomyocyte-specific promoters, smooth muscle-specific promoters, photoreceptor-specific promoters, etc. Spatially restricted neuron-specific promoters include the neuron-specific enolase (NSE) promoter (see, e.g., EMBL HSENO2, X51956); aromatic amino acid decarboxylase (AADC) promoter; neurofilament promoter (see, e.g., GenBank HUMNFL, L04147); synapsin promoter (see, e.g., GenBank HUMSYNIB, M55301); thy-1 promoter; serotonin receptor promoter (see, e.g., GenBank S62283); tyrosine hydroxylase promoter (TH); GnRH promoter; L7 promoter; DNMT promoter; enkephalin; myelin basic protein (MBP) promoter; Ca2+-calmodulin-dependent protein kinase II-alpha (CamKIIα) promoter; CMV enhancer / platelet-derived growth factor-β promoter, etc., but are not limited to these. Suitable liver-specific promoters may include, in some cases, but are not limited to, the TTR, albumin, and AAT promoters. Suitable CNS-specific promoters may include, in some cases, but are not limited to, the synapsin 1, BM88, CHNRB2, GFAP, and CAMK2a promoters. Suitable muscle-specific promoters may include, in some cases, but are not limited to, the MYOD1, MYLK2, SPc5-12 (synthetic), α-MHC, MLC-2, MCK, MHCK7, human cardiac troponin C (cTnC), and desmin promoters.Spatially restricted adipocyte-specific promoters include, but are not limited to, the aP2 gene promoter / enhancer, such as the region from -5.4 kb to +21 bp of human aP2; glucose transporter-4 (GLUT4); fatty acid translocase (FAT / CD36) promoter; stearoyl-CoA desaturase-1 (SCD1) promoter; leptin promoter; adiponectin promoter; adipsin promoter; resistin promoter, etc. Spatially restricted cardiomyocyte-specific promoters include, but are not limited to, regulatory sequences derived from genes such as myosin light chain-2, α-myosin heavy chain, AE3, cardiac troponin C, and cardiac actin. Spatially restricted smooth muscle-specific promoters include, but are not limited to, the SM22α promoter; smoothelin promoter; α-smooth muscle actin promoter, etc. For example, the 0.4 kb region of the SM22α promoter has two CArG elements and has been shown to act specifically on vascular smooth muscle cells. Spatially restricted photoreceptor-specific promoters include, but are not limited to, the rhodopsin promoter; rhodopsin kinase promoter; beta phosphodiesterase gene; retinitis pigmentosa gene promoter; interphotoreceptor retinoid-binding protein (IRBP) gene enhancer; IRBP gene promoter, etc.
[0191] Examples of inducible promoters include, but are not limited to, heat shock promoters, tetracycline-regulated promoters, steroid-regulated promoters, metal-regulated promoters, estrogen receptor-regulated promoters, etc. Thus, inducible promoters can be regulated by molecules including, but not limited to, doxycycline; estrogen receptor; estrogen receptor fusion proteins; estrogen analogs; IPTG, etc. Suitable inducible promoters for use include any inducible promoter described herein or known to those skilled in the art. Examples of inducible promoters include, but are not limited to, chemically / biochemically regulated and physically regulated promoters, such as alcohol-regulated promoters, tetracycline-regulated promoters (e.g., anhydrotetracycline (aTc)-responsive promoters and other tetracycline-responsive promoter systems (including tetracycline repressor protein (tetR), tetracycline operator sequence (tetO), and tetracycline transactivator fusion protein (tTA))), steroid-regulated promoters (e.g., promoters based on rat glucocorticoid receptor, human estrogen receptor, moth ecdysone receptor, and promoters derived from the steroid / retinoid / thyroid receptor superfamily), metal-regulated promoters (e.g., promoters derived from yeast, mouse, and human metallothionein (proteins that bind and sequester ester metal ions)) genes), pathogen-regulated promoters (e.g., those induced by salicylic acid, ethylene, or benzothiadiazole (BTH)), temperature / heat-inducible promoters (e.g., heat shock promoters), and light-regulated promoters (e.g., light-responsive promoters derived from plant cells).
[0192] Inducible promoters include sugar-inducible promoters (e.g., lactose-inducible promoter; arabinose-inducible promoter); amino acid-inducible promoters; alcohol-inducible promoters, etc. Suitable promoters include, for example, the lactose regulatory system (e.g., lactose operon system, sugar regulatory system, isopropyl-beta-D-thiogalactopyranoside (IPTG)-inducible system, arabinose regulatory system (e.g., arabinose operon system, e.g., ARA operon promoter, pBAD, pARA, a part thereof, combinations thereof, etc.), synthetic amino acid regulatory system, fructose repressor, tac promoter / operator (pTac), tryptophan promoter, PhoA promoter, recA promoter, proU promoter, cst-1 promoter, tetA promoter, cadA promoter, nar promoter, P LPromoters, such as the cspA promoter, or combinations thereof are included. In certain cases, the promoter includes Lac-Z, or a part thereof. In some cases, the promoter includes the Lac operon, or a part thereof. In some cases, the inducible promoter includes the ARA operon promoter, or a part thereof. In certain embodiments, the inducible promoter includes the arabinose promoter or a part thereof. The arabinose promoter can be obtained from any suitable bacterium. In some cases, the inducible promoter includes the arabinose operon of E. coli or B. subtilis. In some cases, the inducible promoter is activated by the presence of a sugar or sugar analog. Non-limiting examples of sugars and sugar analogs include lactose, arabinose (e.g., L-arabinose), glucose, sucrose, fructose, IPTG, and the like. Suitable promoters include the T7 promoter; the pBAD promoter; the lacIQ promoter, and the like. In some cases, the promoter is the J23119 promoter. Many bacterial promoters are known in the art and bacterial promoters can be found at parts dot igem dot org / promoters on the Internet.
[0193] In some cases, the promoter is a reversible promoter. Suitable reversible promoters include reversible inducible promoters, which are known in the art. Such reversible promoters can be isolated and derived from many organisms. Such reversible promoters can be isolated and derived from many organisms, such as eukaryotes and prokaryotes. Modification of a reversible promoter derived from a first organism for use in a second organism is well known in the art. For example, modification of a reversible promoter derived from a first organism for use in a second organism, such as a first prokaryote and a second eukaryote, or a first eukaryote and a second prokaryote, is well known in the art. Such reversible promoters, and systems based on such reversible promoters but also including additional regulatory proteins, include alcohol-regulated promoters (e.g., alcohol dehydrogenase I (alcA) gene promoter, a promoter responsive to the alcohol trans-activator protein (AlcR)), tetracycline-regulated promoters (e.g., a promoter system including Tet activator, TetON, TetOFF), steroid-regulated promoters (e.g., rat glucocorticoid receptor promoter system, human estrogen receptor promoter system, retinoid promoter system, thyroid promoter system, ecdysone promoter system, mifepristone promoter system), metal-regulated promoters (e.g., metallothionein promoter system), pathogen-related regulated promoters (e.g., salicylic acid-regulated promoter, ethylene-regulated promoter, benzothiadiazole-regulated promoter), temperature-regulated promoters (e.g., heat shock inducible promoters (e.g., HSP-70, HSP-90, soybean heat shock promoter), light-regulated promoters, synthetic inducible promoters, etc., but not limited thereto).
[0194] Thus, it will be understood that the present disclosure includes the use of any promoter / regulatory sequence capable of driving the expression of a desired nuclease or RNA operably linked thereto.
[0195] Furthermore, the vectors described herein for the expression of nuclease and / or gRNA may contain, for example, some or all of the following: a selectable marker gene, such as the neomycin gene for selecting stable or transient transfectants in host cells; an enhancer / promoter sequence derived from the immediate early gene of human CMV for high-level transcription; an SV40-derived transcription termination and RNA processing signal for mRNA stability; 5' and 3' untranslated regions derived from highly expressed genes such as α-globin or β-globin for mRNA stability and translation efficiency; the SV40 polyomavirus origin of replication and ColE1 for proper episomal replication; an internal ribosome entry site (IRES) within the sequence, a versatile multiple cloning site; T7 and SP6 RNA promoters for in vitro transcription of sense and antisense RNAs; a "suicide switch" or "suicide gene" (e.g., an inducible caspase such as HSV thymidine kinase, iCasp9) that kills cells carrying the vector by a trigger, and a reporter gene for evaluating the expression of chimeric receptors. Vectors and methods suitable for constructing vectors containing transgenes are well known and available in the art. Selectable markers also include chloramphenicol resistance, tetracycline resistance, spectinomycin resistance, streptomycin resistance, erythromycin resistance, rifampicin resistance, bleomycin resistance, thermoadapted kanamycin resistance, gentamicin resistance, hygromycin resistance, trimethoprim resistance, dihydrofolate reductase (DHFR), GPT; the URA3, HIS4, LEU2, and TRP1 genes of S. cerevisiae.
[0196] Once introduced into a cell, the vector can be maintained as an autonomously replicating sequence or episomal element, or can be integrated into the host DNA.
[0197] The compositions and systems (e.g., proteins, polynucleotides encoding these proteins, or compositions comprising the proteins and / or polynucleotides described herein) can be delivered by any suitable means. In certain embodiments, the composition or system is delivered in vivo. In other embodiments, the composition or system is delivered in vitro to isolated / cultured cells (e.g., autologous iPS cells).
[0198] The vectors and nucleic acids according to the present disclosure can be introduced into a variety of host cells by transformation, transfection, or other means. Transfection refers to the uptake of nucleic acid by a host cell, regardless of whether the coding sequence is actually expressed. A number of methods of transfection are known to those skilled in the art, for example, Lipofectamine, calcium phosphate coprecipitation, electroporation, DEAE-dextran treatment, microinjection, viral infection, and other methods known in the art. Transduction refers to the entry of a virus into a cell and the expression (e.g., transcription and / or translation) of the sequences delivered by the viral vector genome. In the case of recombinant vectors, "transduction" generally refers to the entry of a recombinant viral vector into a cell and the expression of the nucleic acid of interest delivered by the vector genome.
[0199] Also within the scope of the present disclosure are vectors comprising nucleic acid sequences encoding the components of the compositions and systems. Such vectors can be delivered to host cells by suitable methods. Methods for delivering vectors to cells are well known in the art and can include electroporation of DNA or RNA, transfection reagents such as liposomes or nanoparticles for delivering DNA or RNA, delivery of DNA, RNA, or proteins by mechanical deformation, or viral transduction. In some embodiments, the vector is delivered to the host cell by viral transduction. Nucleic acids can be delivered as part of a larger construct such as a plasmid or viral vector, or, for example, by electroporation, lipid vesicles, viral transporters, microinjection, and biolistics (high-velocity particle bombardment). Similarly, constructs containing one or more transgenes can be delivered by any method suitable for introducing nucleic acids into cells.
[0200] Furthermore, delivery vehicles such as nanoparticles and lipid-based delivery systems for mRNA or protein can also be used. Further examples of delivery vehicles include lentiviral vectors, ribonucleoprotein (RNP) complexes, lipid-based delivery systems, gene guns, hydrodynamic, electroporation or nucleofection microinjection, biolistics, and the like.
[0201] In some embodiments, the vector is a viral construct, such as a recombinant adeno-associated virus construct, a recombinant adenovirus construct, a recombinant lentivirus construct, a recombinant retrovirus construct, etc. Suitable viral vectors include, but are not limited to, vaccinia virus; poliovirus; adenovirus; adeno-associated virus; SV40; herpes simplex virus; virus vectors based on human immunodeficiency virus; retrovirus vectors (e.g., vectors derived from retroviruses such as murine leukemia virus, spleen necrosis virus, and Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus).
[0202] In some embodiments, the vector is an AAV vector. Adeno-associated virus, or "AAV," means the virus itself or a derivative thereof. The term encompasses, unless otherwise required, all subtypes and both naturally occurring and recombinant forms, e.g., AAV type 1 (AAV-1), AAV type 2 (AAV-2), AAV type 3 (AAV-3), AAV type 4 (AAV-4), AAV type 5 (AAV-5), AAV type 6 (AAV-6), AAV type 7 (AAV-7), AAV type 8 (AAV-8), AAV type 9 (AAV-9), AAV type 10 (AAV-10), AAV type 11 (AAV-11), avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, ovine AAV, hybrid AAV (i.e., AAV containing the capsid protein of one AAV subtype and the genomic material of another subtype), AAV containing a mutant AAV capsid protein or a chimeric AAV capsid (i.e., a capsid protein containing regions or domains or individual amino acids from two or more different AAV serotypes, e.g., AAV-DJ, AAV-LK3, AAV-LK19). For example, "primate AAV" refers to AAV that infects primates, "non-primate AAV" refers to AAV that infects non-primate mammals, and "bovine AAV" refers to AAV that infects bovine mammals.
[0203] The term "recombinant AAV vector" or "rAAV vector" refers to an AAV virus or AAV viral chromosomal material that contains a polynucleotide sequence not of AAV origin (e.g., a polynucleotide heterologous to AAV), typically a nucleic acid sequence of interest that is incorporated into a cell according to the method of interest. Generally, the heterologous polynucleotide is flanked by at least one, and generally two, AAV terminal inverted repeat (ITR) sequences. In some cases, the recombinant viral vector also contains viral genes important for the packaging of the recombinant viral vector material. Packaging refers to a series of intracellular events that result in the assembly and encapsulation of viral particles, e.g., AAV viral particles. Examples of nucleic acid sequences important for AAV packaging include the AAV "rep" and "cap" genes, which encode the replication protein and encapsulation protein of adeno-associated virus, respectively. The term rAAV vector encompasses both rAAV vector particles and rAAV vector plasmids.
[0204] "Viral particle" refers to a single unit of a virus that contains a virus-based polynucleotide, e.g., a viral genome (in the case of a wild-type virus), or a capsid that encapsulates, e.g., a targeted vector of interest (in the case of a recombinant virus). An AAV viral particle refers to a viral particle composed of at least one AAV capsid protein (typically, all capsid proteins of wild-type AAV) and an encapsulated polynucleotide AAV vector. When the particle contains a heterologous polynucleotide (e.g., a polynucleotide other than the wild-type AAV genome, a transgene delivered to a mammalian cell, etc.), the particle is typically referred to as an "rAAV vector particle" or simply an "rAAV vector". Thus, the production of rAAV particles necessarily involves the production of an rAAV vector because the vector is contained within the rAAV particles.
[0205] rAAV virions can be constructed in a variety of ways. For example, heterologous sequence(s) can be directly inserted into an AAV genome from which the major AAV open reading frame (“ORF”) has been excised. Other portions of the AAV genome can be deleted as long as sufficient ITR portions remain to allow replication and packaging functions. To produce rAAV virions, known techniques such as transfection can be used to introduce an AAV expression vector into a suitable host cell. Particularly suitable transfection methods include calcium phosphate co-precipitation, direct microinjection into cultured cells, electroporation, gene transfer via liposomes, transduction via lipids, and nucleic acid delivery using a high-speed microprojectile. Cells suitable for the production of rAAV virions include microorganisms, yeast cells, insect cells, and mammalian cells that can be used or are being used as recipients of heterologous DNA molecules.
[0206] The AAV virus produced can be either replication-competent or replication-incompetent. A “replication-competent” virus (e.g., replication-competent AAV) refers to a phenotypically wild-type virus that is infectious and can also replicate within an infected cell (e.g., in the presence of a helper virus or helper virus functions). In the case of AAV, replication ability generally requires the presence of functional AAV packaging genes. Generally, the rAAV vectors described herein are replication-incompetent in mammalian cells (particularly human cells) by virtue of not having one or more AAV packaging genes. Typically, such rAAV vectors do not have any AAV packaging gene sequences in order to minimize the possibility of producing replication-competent AAV by recombination between an AAV packaging gene and the rAAV vector that comes in with it.
[0207] Retroviruses, such as lentiviruses, are suitable for use in the methods of the present disclosure. Commonly used retroviral vectors are unable to produce the viral proteins necessary for productive infection. Rather, replication of the vector requires growth in a packaging cell line. To generate virus particles containing the nucleic acid of interest, the packaging cell line packages the retroviral nucleic acid containing the nucleic acid into the viral capsid. Different packaging cell lines provide different envelope proteins (ecotropic, amphotropic, or xenotropic) that are incorporated into the capsid, and this envelope protein determines the specificity of the virus particle for cells (ecotropic for mice and rats; amphotropic for most mammalian cell types including humans, dogs, and mice; and xenotropic for most mammalian cell types other than mouse cells). By using an appropriate packaging cell line, targeting of cells by the packaged virus particles can be ensured. Methods for introducing the vector expression vector of interest into the packaging cell line and methods for recovering the virus particles produced by the packaging strain are well known in the art. The nucleic acid can also be introduced by direct microinjection (e.g., injection of RNA).
[0208] As otherwise indicated herein, the protein may instead be provided to the cell as RNA (e.g., RNA containing translational control elements otherwise described herein). Methods for introducing RNA into cells can include, for example, direct injection, transfection, or any other method used for introduction of DNA. Nucleases can also be introduced directly into the host cell as a protein. In such cases, the nuclease can be delivered as an RNP (ribonucleoprotein complex) already complexed with an appropriate guide RNA.
[0209] The disclosed nucleic acids (e.g., vectors) and proteins can be delivered into cells using any convenient method. Suitable methods include, for example, viral infection (e.g., AAV, adenovirus, lentivirus), transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, transfection via polyethyleneimine (PEI), transfection via DEAE-dextran, transfection via liposomes, particle gun technology, calcium phosphate precipitation, direct microinjection, nucleic acid delivery via nanoparticles, and the like.
[0210] In some cases, the nuclease is delivered into cells within particles or associated with particles. In some cases, the nuclease is delivered together with a cationic lipid and a hydrophilic polymer. For example, the cationic lipid includes 1,2-dioleoyl-3-trimethylammonium-propane (DOTAP) or 1,2-ditetradecanoyl-sn-glycero-3-phosphocholine (DMPC), and / or the hydrophilic polymer includes ethylene glycol or polyethylene glycol (PEG), and / or the particles further include cholesterol.
[0211] The nuclease can be delivered using particles or a lipid envelope. For example, biodegradable core-shell structured nanoparticles having a poly(β-amino ester) (PBAE) core coated with a phospholipid bilayer shell can be used. In some cases, particles / nanoparticles based on self-assembling biocompatible polymers are used, and such particles / nanoparticles can be applied to oral delivery of peptides, intravenous delivery of peptides, and nasal delivery of peptides, for example, delivery to the brain. Other embodiments such as oral absorption of hydrophobic drugs and ocular delivery are also contemplated. Molecular envelope technology with engineered polymer envelopes that are protected and delivered to the desired cells can be used.
[0212] Lipidoid compounds (e.g., those described in U.S. Patent Application Publication No. 2011 / 0293703) are also useful for the delivery of polynucleotides and can be used to deliver the disclosed nuclease (or RNA or DNA encoding the same). In one aspect, an amino alcohol lipidoid compound is combined with an agent to be delivered to a cell to form microparticles, nanoparticles, liposomes, or micelles. The amino alcohol lipidoid compound can be combined with other amino alcohol lipidoid compounds, polymers (synthetic or natural), surfactants, cholesterol, carbohydrates, proteins, lipids, etc. to form particles. These particles can then optionally be combined with a pharmaceutical excipient to form a pharmaceutical composition.
[0213] Poly(beta-amino alcohol) (PBAA) can be used to deliver a nuclease or nucleic acid encoding the same and a gRNA or nucleic acid encoding the same to a target cell. U.S. Patent Application Publication No. 2013 / 0302401 relates to a type of poly(beta-amino alcohol) (PBAA) prepared using combinatorial polymerization.
[0214] Sugar-based particles, e.g., GalNAc described in International Patent Publication No. WO2014118272 (which is hereby incorporated by reference in its entirety; and Nair, J K et al., 2014, Journal of the American Chemical Society 136(49), 16958-16961), can be used to deliver a nuclease or nucleic acid encoding the same and a gRNA or nucleic acid encoding the same to a target cell.
[0215] In some cases, lipid nanoparticles (LNPs) are used to deliver nucleases or nucleic acids encoding them and gRNAs or nucleic acids encoding them to target cells. Negatively charged polymers such as RNA can be loaded into LNPs at low pH values (e.g., pH 4) where ionizable lipids exhibit a positive charge. However, LNPs exhibit a low surface charge at physiological pH values and are compatible with a longer circulation period. Four ionizable cationic lipids, namely, 1,2-dilinoleoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleoyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleoyloxy-keto-N,N-dimethyl-3-aminopropane (DLinKDMA), and 1,2-dilinoleoyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA) have been noted. The preparation of LNPs is described, for example, in Rosin et al. (2011) Molecular Therapy 19:1286-2200). The cationic lipids 1,2-dilinoleoyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleoyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleoyloxyketo-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dilinoleoyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA), (3-o-[2’’-(methoxypolyethylene glycol 2000) succinoyl]-1,2-dimyristoyl-sn-glycol (PEG-S-DMG), and R-3-[(.omega.-methoxy-poly(ethylene glycol) 2000) carbamoyl]-1,2-dimyristyloxylpropyl-3-amine (PEG-C-DOMG) can be used. Nucleic acids can be encapsulated into LNPs containing DLinDAP, DLinDMA, DLinK-DMA, and DLinKC2-DMA (cationic lipid:DSPC:CHOL:PEGS-DMG or PEG-C-DOMG in a 40:10:40:10 molar ratio). In some cases, 0.2% of SP-DiOC18 is incorporated.
[0216] Using spherical nucleic acids (SNA™) constructs and other nanoparticles (in particular, gold nanoparticles), nucleases or nucleic acids encoding them and gRNAs or nucleic acids encoding them can be delivered to target cells.
[0217] Self-assembling nanoparticles containing RNA can be constructed with polyethyleneimine (PEI) PEGylated with an Arg-Gly-Asp (RGD) peptide ligand attached to the distal end of polyethylene glycol (PEG).
[0218] Nanoparticles suitable for use in delivering nucleases or nucleic acids encoding them and gRNAs or nucleic acids encoding them to target cells can be provided in different forms, such as solid nanoparticles (e.g., metals such as silver, gold, iron, titanium), non-metals, lipid-based solids, polymers), suspensions of nanoparticles, or combinations thereof. Nanoparticles of metals, dielectrics, and semiconductors can be prepared in the same manner as hybrid structures (e.g., core-shell nanoparticles). Nanoparticles made from semiconductor materials may also be called quantum dots when they are small enough (typically less than 10 nm) for quantization of electronic energy levels to occur. Such nanoscale particles are used in biomedical applications such as drug carriers or contrast agents and can be applied to similar purposes in the present disclosure. Generally, "nanoparticle" refers to any particle having a diameter of less than 1000 nm. In some cases, nanoparticles suitable for use in delivering nucleases or nucleic acids to target cells have a diameter of 500 nm or less, e.g., 25 nm - 35 nm, 35 nm - 50 nm, 50 nm - 75 nm, 75 nm - 100 nm, 100 nm - 150 nm, 150 nm - 200 nm, 200 nm - 300 nm, 300 nm - 400 nm, or 400 nm - 500 nm. In some cases, nanoparticles suitable for use in delivering nucleases or nucleic acids to target cells have a diameter of 25 nm - 200 nm.
[0219] In some cases, exosomes are used to deliver a nuclease or a nucleic acid encoding the same and a gRNA or a nucleic acid encoding the same to target cells. Exosomes are endogenous nano-vesicles that transport RNA and proteins and can deliver RNA to the brain and other target organs.
[0220] In some cases, liposomes are used to deliver a nuclease or a nucleic acid encoding the same and a gRNA or a nucleic acid encoding the same to target cells. Liposomes are spherical vesicular structures composed of uni- or multi-lamellar lipid bilayers surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes can be made from several different types of lipids, but phospholipids are most commonly used in the preparation of liposomes. The formation of liposomes occurs naturally when lipid membranes are mixed with an aqueous solution, but it can also be facilitated by applying force in the form of agitation using a homogenizer, sonicator, or extrusion device. To modify the structure and properties of liposomes, several other additives may be added to the liposomes. For example, adding either cholesterol or sphingomyelin to the liposome mixture can help stabilize the liposome structure and prevent leakage of the cargo inside the liposomes. Liposome formulations can mainly consist of natural phospholipids and lipids such as 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine, and monosialoganglioside.
[0221] Using stable nucleic acid lipid particles (SNALPs), a nuclease or a nucleic acid encoding the same and a gRNA or a nucleic acid encoding the same can be delivered to a target cell. The SNALP formulation may contain 3-N-[(methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristyloxy-propylamine (PEG-C-DMA), 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC) and cholesterol in a molar percentage of 2:40:10:48. The SNALP liposome can be prepared by formulating D-Lin-DMA and PEG-C-DMA with distearoyl phosphatidylcholine (DSPC), cholesterol and siRNA using a lipid / siRNA ratio of 25:1 and a molar ratio of cholesterol / D-Lin-DMA / DSPC / PEG-C-DMA of 48 / 40 / 10 / 2. The resulting SNALP liposome can have a size of about 80 to 100 nm. The SNALP may contain synthetic cholesterol (Sigma-Aldrich, St Louis, Mo., USA), dipalmitoyl phosphatidylcholine (Avanti Polar Lipids, Alabaster, Ala., USA), 3-N-[(ω-methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyrestyloxypropylamine, and cationic 1,2-dilinoleyloxy-3-N,N dimethylaminopropane. The SNALP may contain synthetic cholesterol (Sigma-Aldrich), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC; Avanti Polar Lipids Inc.), PEG-cDMA, and 1,2-dilinoleyloxy-3-(N;N-dimethyl)aminopropane (DLinDMA).
[0222] Other cationic lipids, such as the amino lipid 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane (DLin-KC2-DMA), can also be used to deliver nucleases or nucleic acids to target cells. The following lipid composition: amino lipid, distearoylphosphatidylcholine (DSPC), cholesterol and (R)-2,3-bis(octadecyloxy)propyl-1-(methoxypoly(ethylene glycol)2000)propylcarbamate (PEG-lipid) in a molar ratio of 40 / 10 / 40 / 10 respectively, and a ready-made vesicle with an FVII siRNA / total lipid ratio of about 0.05 (w / w) can be contemplated. To ensure a narrow particle size distribution in the range of 70 - 90 nm and a small polydispersity index of 0.11 ± 0.04 (n = 56), the particles can be extruded through an 80 nm membrane up to 3 times before adding the guide RNA. Particles containing the extremely potent amino lipid 16 can also be used, and the molar ratio of the four lipid components 16, DSPC, cholesterol and PEG-lipid (50 / 10 / 38.5 / 1.5) may be further optimized to enhance in vivo activity.
[0223] Lipids can be formulated with a nuclease or a nucleic acid encoding the same and a gRNA or a nucleic acid encoding the same to form lipid nanoparticles (LNPs). Suitable lipids include, but are not limited to, DLin-KC2-DMA4, C12-200 and the co-lipid distearoylphosphatidylcholine, cholesterol, and PEG-DMG can be formulated with a nuclease or a nucleic acid using a natural vesicle formation procedure.
[0224] The nuclease or a nucleic acid encoding the same and the gRNA or a nucleic acid encoding the same can be encapsulated and delivered in PLGA microspheres, such as those further described in U.S. Patent Application Publications 20130252281, 20130245107, and 20130244279.
[0225] Supercharged proteins can be used to deliver nucleases or nucleic acids encoding them and gRNAs or nucleic acids encoding them to target cells. Supercharged proteins are a type of engineered or naturally occurring proteins with unusually high theoretical net positive or negative charges. Both negatively overcharged proteins and positively overcharged proteins exhibit the ability to withstand thermally or chemically induced aggregation. Positively overcharged proteins can also penetrate mammalian cells. By associating these proteins with cargos such as plasmid DNA, RNA, or other proteins, functional delivery of these macromolecules into mammalian cells in both in vitro and in vivo can be facilitated.
[0226] Cell-penetrating peptides (CPPs) can be used to deliver nucleases or nucleic acids encoding them and gRNAs or nucleic acids encoding them to target cells. CPPs typically have an amino acid composition that either contains positively charged amino acids such as lysine or arginine in high relative abundance or has a sequence containing an alternating pattern of polar / charged and nonpolar hydrophobic amino acids.
[0227] Methods The present disclosure also provides methods for modifying a target nucleic acid sequence (e.g., DNA or RNA). As used herein, the expression "modifying a nucleic acid sequence" refers to modifying at least one physical characteristic of the nucleic acid sequence of interest. Nucleic acid modifications include, for example, single-stranded or double-stranded cleavage, deletion or insertion of one or more nucleotides, and other modifications that affect the structural integrity or nucleotide sequence of the nucleic acid sequence. The modification can include one or more of modification of the target nucleic acid, regulation of transcription from the target nucleic acid, and modification of the polypeptide associated with the target nucleic acid. The method includes contacting the target nucleic acid sequence with a composition disclosed herein, a system disclosed herein, or a composition comprising the system.
[0228] In one embodiment, the method introduces a single-stranded or double-stranded break into a target nucleic acid sequence. In this regard, the disclosed system can induce cleavage of one or both strands of a target DNA sequence, such as within a target genomic DNA sequence and / or within the complement of the target sequence.
[0229] In some embodiments, contacting the target nucleic acid sequence includes introducing the compositions or systems described herein into a cell. As described above, the compositions or systems can be introduced into eukaryotic or prokaryotic cells by methods known in the art.
[0230] The cell can be a prokaryotic cell, a plant cell, an insect cell, a vertebrate cell, an invertebrate cell, an animal cell, a mammalian cell, or a human cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a vertebrate cell. In some embodiments, the cell is an invertebrate cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some cases, the cell is ex vivo (e.g., fresh isolate - early passage). In some cases, the cell is in vivo. In some cases, the cell is an in vitro culture (e.g., immortalized cell line).
[0231] The cell can be derived from an established cell line or be a primary cell. Here, "primary cell", "primary cell line", and "primary culture" are used interchangeably herein to refer to cells and cell cultures that are derived from a subject and grown in vitro with a limited number of passages. For example, a primary culture can be passaged 0, 1, 2, 4, 5, 10, or 15 times, but is a culture that has been passaged a number of times that does not reach the crisis stage. Typically, a primary cell line is maintained in culture for less than 10 passages.
[0232] Suitable cells include, but are not limited to, bacterial cells; archaeal cells; eukaryotic cells; cells of unicellular eukaryotes; plant cells; protozoan cells; algal cells such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. agardh, etc.; fungal cells (e.g., yeast cells); animal cells; cells derived from invertebrates (e.g., Drosophila, cnidarians, echinoderms, nematodes, etc.); cells of insects (e.g., mosquitoes; bees; agricultural pests, etc.); cells of arachnids (e.g., spiders; mites, etc.); cells of vertebrates (e.g., fish, amphibians, reptiles, birds, mammals); cells of mammals (e.g., cells of rodents; human cells; cells of non-human mammals; cells of rodents (e.g., mice, rats); cells of lagomorphs (e.g., rabbits); cells of ungulates (e.g., cows, horses, camels, llamas, vicuñas, sheep, goats, etc.); cells of marine mammals (e.g., whales, seals, walruses, dolphins, sea lions, etc.), etc. All types of cells can be targeted (e.g., stem cells such as embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells (e.g., oocytes, sperm, oogonia, spermatogonia, etc.), adult stem cells, somatic cells such as fibroblasts, hematopoietic cells, neurons, muscle cells, bone cells, liver cells, pancreatic cells; in vitro or in vivo embryonic cells at all stages of the embryo, such as zebrafish embryos at the 1-cell stage, 2-cell stage, 4-cell stage, 8-cell stage, etc.). In some cases, the cells are cells not derived from natural organisms (e.g., the cells can be synthetically made cells, also referred to as artificial cells).
[0233] Non-limiting examples of plant cells include cells derived from plant crops, fruits, vegetables, grains, soybeans, sugarcane, corn, wheat, seeds, tomatoes, rice, cassava, sugar beets, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, angiosperms, fern plants, bryophytes, clubmosses, horsetails, mosses, dicots, monocots, seaweeds (e.g., kelp), etc.
[0234] Suitable cells include stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells); germ cells (e.g., oocytes, sperm, oogonia, spermatogonia, etc.); somatic cells such as fibroblasts, oligodendrocytes, glial cells, hematopoietic cells, neurons, muscle cells, bone cells, hepatocytes, spleen cells, and the like.
[0235] Suitable cells include human embryonic stem cells, fetal cardiomyocytes, myofibroblasts, mesenchymal stem cells, autologous expanded cardiomyocytes, adipocytes, totipotent cells, pluripotent cells, hematopoietic stem cells, myoblasts, adult stem cells, bone marrow cells, mesenchymal cells, embryonic stem cells, parenchymal cells, epithelial cells, endothelial cells, mesothelial cells, fibroblasts, osteoblasts, chondrocytes, exogenous cells, endogenous cells, stem cells, hematopoietic stem cells, bone marrow-derived progenitor cells, cardiomyocytes, skeletal cells, fetal cells, undifferentiated cells, pluripotent progenitor cells, unipotent progenitor cells, monocytes, cardiac myoblasts, skeletal myoblasts, macrophages, capillary endothelial cells, heterologous cells, homologous cells, and postnatal stem cells.
[0236] In some cases, the cells are immune cells, neurons, epithelial cells, and endothelial cells, or stem cells. In some cases, the immune cells are T cells, B cells, monocytes, natural killer cells, dendritic cells, or macrophages. In some cases, the immune cells are cytotoxic T cells. In some cases, the immune cells are helper T cells. In some cases, the immune cells are regulatory T cells (Tregs).
[0237] In some cases, the cells are stem cells. Stem cells include adult stem cells. Adult stem cells are also referred to as somatic stem cells.
[0238] Adult stem cells are resident in differentiated tissues, but retain the property of self-renewal and the ability to give rise to multiple cell types, usually the cell types typical of the tissue in which the stem cells reside. Numerous examples of somatic stem cells are known to those skilled in the art, including muscle stem cells; hematopoietic stem cells; epithelial stem cells; neural stem cells; mesenchymal stem cells; mammary stem cells; intestinal stem cells; mesodermal stem cells; endothelial stem cells; olfactory nerve stem cells; neural crest stem cells, and the like.
[0239] The stem cells of interest include mammalian stem cells, and the term "mammal" refers to any animal classified as a mammal, including humans; non-human primates; domesticated animals and livestock; and zoo animals, laboratory animals, sport animals, or pets, such as dogs, horses, cats, cows, mice, rats, rabbits, and the like. In some cases, the stem cells are human stem cells. In some cases, the stem cells are rodent (e.g., mouse; rat) stem cells. In some cases, the stem cells are non-human primate stem cells.
[0240] In some embodiments, the stem cells are hematopoietic stem cells (HSCs). HSCs are cells derived from the mesoderm that can be isolated from bone marrow, blood, umbilical cord blood, fetal liver, and yolk sac. HSCs are characterized by CD34 + and CD3 - . HSCs enable the repopulation of erythrocytes, neutrophils, macrophages, megakaryocytes, and lymphoid hematopoietic cell lineages in vivo. In vitro, HSCs can be induced to undergo at least some self-renewing cell divisions and can be induced to differentiate into the same lineages seen in vivo. Thus, HSCs can be induced to differentiate into one or more of erythrocytes, megakaryocytes, neutrophils, macrophages, and lymphoid cells.
[0241] In other embodiments, the stem cells are neural stem cells (NSCs). Neural stem cells (NSCs) can differentiate into neurons and glia (including oligodendrocytes and astrocytes). Neural stem cells are pluripotent stem cells capable of multiple divisions and, under specific conditions, can produce daughter cells that are neural stem cells, or neural progenitor cells (e.g., cells that have been determined to become one or more neurons and glial cells, respectively) that can become neuroblasts or glioblasts. Methods for obtaining NSCs are known in the art.
[0242] In other embodiments, the stem cells are mesenchymal stem cells (MSCs). MSCs, which are originally derived from embryonic mesoderm and isolated from adult bone marrow, can differentiate to form muscle, bone, cartilage, fat, bone marrow stroma, and tendon. Methods for isolating MSCs are known in the art, and MSCs can be obtained using any known method. See, for example, U.S. Patent No. 5,736,396, which describes the isolation of human MSCs.
[0243] In some embodiments, the cells are T cells. The present invention is not limited by the type of T cell. T cells can be selected from, for example, CD3+ T cells, CD8+ T cells, CD4+ T cells, natural killer (NK) T cells, alpha-beta T cells, gamma-delta T cells, or any combination thereof (e.g., a combination of CD4+ T cells and CD8+ T cells).
[0244] In some embodiments, the T cells are naturally occurring T cells. For example, the T cells can be isolated from a sample of a subject. In some embodiments, the T cells are anti-tumor T cells (e.g., T cells that are active against a tumor (e.g., an autologous tumor), are activated and proliferate in response to an antigen). Anti-tumor T cells include, but are not limited to, T cells obtained from an excised tumor or a tumor biopsy (e.g., tumor infiltrating lymphocytes (TIL)) and polyclonal or monoclonal tumor-reactive T cells (e.g., those obtained by apheresis, those expanded ex vivo against tumor antigens presented by autologous or artificial antigen-presenting cells). In some embodiments, the T cells are expanded ex vivo.
[0245] The cells are, in some cases, plant cells. The plant cells can be monocotyledonous cells. The plant cells can be dicotyledonous cells. The cells can be root cells, leaf cells, xylem cells, phloem cells, cambium cells, apical meristem cells, parenchyma cells, collenchyma cells, sclerenchyma cells, etc. Plant cells include cells of agricultural crops such as wheat, corn, rice, sorghum, millet, soybean, etc. Plant cells include cells of agricultural fruit and nut plants, e.g., plants that produce apricots, oranges, lemons, apples, plums, pears, almonds, etc.
[0246] Plant cells can be cells of major agricultural plants such as barley, beans (dry edible), canola, corn, cotton (Pima), cotton (upland), flax, hay (alfalfa), hay (non-alfalfa), oats, peanuts, rice, sorghum, soybeans, sugar beets, sugarcane, sunflower (oil), sunflower (non-oil), sweet potato, tobacco (burley), tobacco (hot air dried), tomato, wheat (durum), wheat (spring wheat), wheat (winter wheat), etc. As another example, for example, alfalfa sprouts, aloe leaves, arrowroot, Chinese chives, artichokes, asparagus, bamboo shoots, banana flowers, bean sprouts, beans, beet tops, beets, bitter melon, bok choy, broccoli, broccoli rabe (rapini), Brussels sprouts, cabbage, cabbage sprouts, cactus leaf (nopal), calabaza, cardoon, carrot, cauliflower, celery, chayote, Chinese yam (cross n), Chinese cabbage, Chinese kale, chives, Chinese parsley, chrysanthemum leaf (chrysanthemum), collard greens, cornstalk, sweet corn, cucumber, daikon, dandelion greens, elephant yam, edamame (soybean leaves), winter melon, eggplant, endive, chrysanthemum greens, horsetail, kale, gai choy (mustard greens), kai lan, galangal (Thai ginger), garlic, ginger root, burdock, greens, Hanover salad greens, wasabi tre, Jerusalem artichoke, hickama, kale greens, kohlrabi, amaranth, lettuce (bib), lettuce (Boston), lettuce (Boston red), lettuce (green leaf), lettuce (iceberg), lettuce (sunny lettuce), lettuce (oak leaf green), lettuce (oak leaf red), lettuce (processed), lettuce (red leaf), lettuce (romaine), lettuce (ruby romaine), lettuce (Russian red mustard), lincoln, robok, long bean, lotus root, marsh, maguey (agave) leaf, malanga, musk melon mix, mizuna, loofah (sponge gourd), moo, mokua (fuzzy squash), mushroom, mustard, Chinese potato, okra, seaweed, leek, opo (long squash), ornamental corn, ornamental gourd, parsley, parsnip, beans,Cells of vegetable crops such as Capsicum (Bell group), Capsicum, pumpkin, radicchio, radish sprout, radish, rapeseed green, rapeseed green, rhubarb, romaine (baby red), rutabaga, Akebia (Siebold), loofah (Tokado / Luffa cylindrica), spinach, squash, strawberry, sugarcane, sweet potato, Swiss chard, tamarind, taro, taro leaves, taro sprouts, tatsoi, tepguahé (Gymnema), tindora, tomatillo, tomato, tomato (cherry), tomato (grape type), tomato (plum type), turmeric, turnip top green, turnip, shirogwa, yampi, yam (cassava), rape, yucca (cassava), etc., but not limited to these.
[0247] In some cases, the cells are arthropod cells. For example, the cells can be cells of a suborder, family, subfamily, group, subgroup, or species, such as Chelicerata, Myriapodia, Hexipodia, Arachnida, Insecta, Archaeognatha, Thysanura, Palaeoptera, Ephemeroptera, Odonata, Anisoptera, Zygoptera, Neoptera, Exopterygota, Plecoptera, Embioptera, Orthoptera, Zoraptera, Dermaptera, Dictyoptera, Notoptera, Grylloblattidae, Mantophasmatidae, Phasmatodea, Blattaria, Isoptera, Mantodea, Parapneuroptera, Psocoptera, Thysanoptera, Phthiraptera, Hemiptera, Endopterygota or Holometabola, Hymenoptera, Coleoptera, Strepsiptera, Raphidioptera, Megaloptera, Neuroptera, Mecoptera, Siphonaptera, Diptera, Trichoptera, or Lepidoptera.
[0248] The cells are, in some cases, insect cells. For example, in some cases, the cells are cells of mosquitoes, locusts, hemiptera, flies, lice, bees, wasps, ants, bedbugs, moths, or beetles.
[0249] In some embodiments, introducing the system into the cells includes administering the system to a subject. In some embodiments, the subject is a human. Administration can include in vivo administration. In alternative embodiments, the vector is contacted with the cells in vitro or ex vivo, and the treated cells containing the system are transplanted into the subject.
[0250] In some embodiments, the target nucleic acid is a nucleic acid endogenous to the target cell. In some embodiments, the target nucleic acid is a genomic DNA sequence. As used herein, the term "genome" refers to nucleic acid sequences (e.g., genes or loci) located on chromosomes within a cell.
[0251] In some embodiments, the target nucleic acid encodes a gene or gene product. As used herein, the term "gene product" refers to any biochemical product resulting from the expression of a gene. The gene product may be RNA or protein. RNA gene products include non-coding RNAs such as tRNA, rRNA, microRNA (miRNA), and small interfering RNA (siRNA), as well as coding RNAs such as messenger RNA (mRNA). In some embodiments, the target nucleic acid sequence encodes a protein or polypeptide.
[0252] The disclosed methods can modify the target DNA sequence of a host cell so as to regulate the expression of the target DNA sequence, e.g., such that the expression of the target DNA sequence is increased, decreased, or completely removed (e.g., via deletion of the gene).
[0253] In another embodiment, a method of modifying a target sequence can be used to delete a nucleic acid sequence or a portion thereof from a target sequence in a host cell by cleaving the target sequence, such that the host cell repairs the cleaved sequence in the absence of an exogenously provided donor nucleic acid molecule. Deleting a nucleic acid sequence in such a manner can be used for a variety of applications, such as removing trinucleotide repeat sequences in neurons that cause disease, creating gene knockouts or knockdowns, and generating mutations for disease models in research.
[0254] In some embodiments, the systems and methods described herein can be used to insert a gene or a fragment thereof into a cell. In certain embodiments, the disclosed systems can be used to create cells that express a recombinant receptor. In some embodiments, the recombinant receptor is a T cell receptor (TCR) or a chimeric antigen receptor (CAR). Also provided herein are cells, such as T cells, that contain a recombinant receptor and / or a nucleic acid encoding the same and the systems described herein (e.g., a nuclease and at least one gRNA).
[0255] In some embodiments, the systems and methods described herein can be used to genetically modify a plant or a plant cell. As used herein, a genetically modified plant includes a plant into which an exogenous polynucleotide has been introduced. Genetically modified plants also include genetically engineered plants that have been modified to contain mutations such as deletions, insertions, translocations, conversions, or combinations thereof of endogenous nucleotides. For example, an endogenous coding region can be deleted. Such mutations can result in a polypeptide having an amino acid sequence different from the amino acid sequence encoded by the endogenous polynucleotide. Another example of a genetically modified plant is one in which a regulatory sequence, such as a promoter, has been modified such that the expression of an operably linked endogenous coding region is increased or decreased. Genetically modified plants can facilitate plant traits of a desired phenotype or genotype.
[0256] Genetically modified plants may have the potential to improve crop yields, enhance nutritional value, and extend storage life. They may also have resistance to undesirable environmental conditions, insects, and pesticides. The present system and method have a wide range of applications in gene discovery and validation, mutation and cisgenic breeding, as well as hybrid breeding. The present system and method can promote the production of a new generation of genetically modified crops having various improved agricultural traits such as herbicide resistance, herbicide tolerance, drought tolerance, male sterility, insect resistance, abiotic stress tolerance, modification of fatty acid metabolism, modification of carbohydrate metabolism, modification of seed yield, modification of oil percentage, modification of protein percentage, resistance to bacterial diseases, resistance to diseases (e.g., bacterial, fungal, and viral), high yield, and excellent quality. The present system and method can also promote the production of a new generation of genetically modified crops in which aroma, nutritional value, storage life, pigments (e.g., lycopene content), starch content (e.g., low gluten wheat), toxin levels, reproduction and / or breeding and growth times are optimized. See, for example, CRISPR / Cas Genome Editing and Precision Plant Breeding in Agriculture (Chen et al., Annu Rev Plant Biol. 2019 Apr 29;70:667-69), which is incorporated herein by reference.
[0257] The present system and method can confer one or more of the following traits: herbicide tolerance, drought tolerance, male sterility, insect resistance, abiotic stress tolerance, modification of fatty acid metabolism, modification of carbohydrate metabolism, modification of seed yield, modification of oil percentage, modification of protein percentage, resistance to bacterial diseases, resistance to fungal diseases, and resistance to viral diseases to plant cells.
[0258] The present disclosure provides modified plant cells produced by the present system and method, plants containing such plant cells, and seeds, fruits, plant parts, or propagation materials of such plants. The transformed or genetically modified plant cells of the present disclosure can be a population of cells, or a tissue, seed, whole plant, stem, fruit, leaf, root, flower, stalk, tuber, grain, animal feed, plant field, etc. The present disclosure provides transgenic plants. The transgenic plants can be homozygous or heterozygous for the genetic modification. Also provided by the present disclosure are transformed or genetically modified plant cells, tissues, plants, and products containing the transformed or genetically modified plant cells. The present disclosure further encompasses progeny, clones, cell lines or cells of the transgenic plants.
[0259] The present system and method can be used to modify plant stem cells. The present disclosure further provides progeny of the genetically modified cells, where the progeny can contain the same genetic modification as the genetically modified cells from which they are derived. The present disclosure further provides a composition containing the genetically modified cells.
[0260] In one embodiment, the transformed or genetically modified cells, and tissues and products include nucleic acids integrated into the genome, and products produced by the plant cells of the gene products resulting from the transformation or genetic modification.
[0261] Methods for introducing exogenous nucleic acid into plant cells are well known in the art. Such plant cells are considered to be "transformed". DNA constructs can be introduced into plant cells by a variety of methods, including, but not limited to, protoplast transformation via PEG or electroporation, tissue culture or plant tissue transformation by biolistic bombardment, or Agrobacterium-mediated transient and stable transformation. Transformation can be transient or stable transformation. Suitable methods also include viral infection (such as double-stranded DNA viruses), transfection, conjugation, protoplast fusion, electroporation, particle gun technology, calcium phosphate precipitation, direct microinjection, silicon carbide whisker technology, Agrobacterium-mediated transformation, and the like. The choice of method generally depends on the type of cell to be transformed and the context in which the transformation is performed (i.e., in vitro, ex vivo, or in vivo). Transformation methods based on the soil bacterium Agrobacterium tumefaciens are useful for introducing exogenous nucleic acid molecules into vascular plants. The wild type of Agrobacterium contains the Ti (tumor-inducing) plasmid, which induces the formation of tumorigenic crown gall formation in the host plant. Transfer of the tumor-inducing T-DNA region of the Ti plasmid into the plant genome requires a virulence gene encoded by the Ti plasmid and T-DNA borders, a series of direct DNA repeats that define the region to be transferred. Agrobacterium-based vectors are modified forms of the Ti plasmid in which the tumor-inducing function has been replaced with a nucleic acid sequence of interest to introduce into the plant host.
[0262] Transformation via Agrobacterium generally utilizes a co-integrate vector or binary vector system in which the components of the Ti plasmid are split between a helper vector that is stably present in the Agrobacterium host and carries the virulence genes, and a shuttle vector that contains the gene of interest ligated by the T-DNA sequences. A variety of binary vectors are well known in the art and are commercially available, for example, from Clontech (Palo Alto, Calif.). Methods for co-culturing Agrobacterium with cultured plant cells or wounded tissue, such as leaf tissue, root explants, hypocotyls, stem segments or tubers, are also well known in the art. See, for example, Glick and Thompson, (eds.), Methods in Plant Molecular Biology and Biotechnology, Boca Raton, Fla.: CRC Press (1993), which is incorporated herein by reference.
[0263] Transgenic plants can also be produced using transformation via microprojectiles. This method was first described by Klein et al. (Nature 327:70-73 (1987), which is incorporated herein by reference) and involves microprojectiles such as gold or tungsten coated with the desired nucleic acid molecule by precipitation with calcium chloride, spermidine, or polyethylene glycol. The microprojectile particles are accelerated at high speed into angiosperm tissue using a device such as the BIOLISTIC PD-1000 (Biorad; Hercules Calif.).
[0264] In one embodiment, the present system and method can be adapted for use in plants. In one embodiment, a series of plant-specific RNA-guided genome editing vectors (pRGE plasmids) are provided for the expression of the present system in plants. The vectors can be optimized for transient expression of the present system in plant protoplasts, or for stable integration and expression in intact plants by transformation via Agrobacterium. In one aspect, the vector construct comprises a nucleotide sequence comprising a DNA-dependent RNA polymerase III promoter (the promoter being operably linked to a gRNA molecule and a Pol III terminator sequence), and a nucleotide sequence comprising a DNA-dependent RNA polymerase II promoter operably linked to a nucleic acid sequence encoding a nuclease.
[0265] In certain embodiments, the present system and method use a monocotyledonous plant promoter to drive the expression of one or more components of the present system (e.g., gRNA) in monocotyledonous plants. In certain embodiments, the present system and method use a dicotyledonous plant promoter to drive the expression of one or more components of the present system (e.g., gRNA) in dicotyledonous plants. In some embodiments, the present system is transiently expressed in plant protoplasts. Vectors for transient transformation of plants include, but are not limited to, pRGE3, pRGE6, pRGE31, and pRGE32. In some embodiments, the vector may be optimized for use in a particular plant species or variety, such as pStGE3.
[0266] In one embodiment, the system can be stably integrated into the plant genome, for example, by transformation via Agrobacterium. Thereafter, one or more components of the system (e.g., the transgene) can be removed by genetic mating and segregation, resulting in the production of non-transgenic but genetically modified plants or crops. In one embodiment, the vector is optimized for transformation via Agrobacterium. In one embodiment, the vector for stable integration is pRGEB3, pRGEB6, pRGEB31, pRGEB32, or pStGEB3.
[0267] The system can be used in a variety of bacterial hosts, including medically important human pathogens, bacterial pests that are important targets in the agricultural industry, and their antibiotic-resistant forms.
[0268] In other embodiments, the system and method can be designed to target any gene or any set of genes, e.g., pathogenic or metabolic genes, for clinical and industrial applications. For example, the system and method can be used to target pathogenic genes and perform in situ gene knockout to eliminate them from a population, or to stably introduce new genetic elements into the metagenomic pool of a microbiome. The system and method can be used to treat multi-drug resistant bacterial infections in a subject. The system and method can be used for genome engineering within complex bacterial consortia.
[0269] The system and method can be used to inactivate microbial genes. In some embodiments, the gene is an antibiotic resistance gene. For example, the coding sequence of a bacterial resistance gene can be disrupted in vivo by the insertion of a DNA sequence, which can lead to non-selective re-sensitization to drug treatment.
[0270] The components of the composition or system can be administered as a pharmaceutical composition, together with a pharmaceutically acceptable carrier or excipient. In some embodiments, the components of the system can be mixed with a pharmaceutically acceptable carrier, individually or in any combination, to form a pharmaceutical composition, which is also within the scope of the present disclosure.
[0271] In some embodiments, an effective amount of the components of the system or composition described herein can be administered. In the context of the present disclosure, the term "effective amount" refers to the amount of the components of the system such that modification of the target nucleic acid is achieved.
[0272] The methods described herein also provide treatment of a subject's disease or condition. In some embodiments, the systems and methods are used to treat a pathogen or parasite on or within a subject by altering the pathogen or parasite. In some embodiments, the systems and methods target "disease-related" genes. The term "disease-related gene" refers to any gene or polynucleotide whose gene product is expressed at abnormal levels or in abnormal forms in cells obtained from an individual suffering from a disease as compared to tissues or cells obtained from an individual not suffering from the disease. Disease-related genes may be expressed at abnormally high or abnormally low levels, and the change in expression correlates with the development and / or progression of the disease. Disease-related genes also refer to genes whose mutations or genetic variants are directly involved in the etiology of the disease or are in linkage disequilibrium with genes (s) involved in the etiology of the disease. Examples of genes that cause such "single gene" or "monogenic" diseases include adenosine deaminase, alpha-1 antitrypsin, cystic fibrosis transmembrane conductance regulator (CFTR), beta-hemoglobin (HBB), oculocutaneous albinism II (OCA2), huntingtin (HTT), myotonic dystrophy protein kinase (DMPK), low density lipoprotein receptor (LDLR), apolipoprotein B (APOB), neurofibromin 1 (NF1), polycystic kidney 1 (PKD1), polycystic kidney 2 (PKD2), coagulation factor VIII (F8), dystrophin (DMD), phosphate regulating endopeptidase homolog, X-linked (PHEX), methyl CpG binding protein 2 (MECP2), and ubiquitin specific peptidase 9Y, Y-linked (USP9Y), but are not limited thereto.Other single-gene diseases or monogenic disorders are known in the art and are described, for example, in Chial, H. Rare Genetic Disorders: Learning About Genetic Disease Through Gene Mapping, SNPs, and Microarray Data, Nature Education 1(1):192 (2008); Online Mendelian Inheritance in Man (OMIM); and the Human Gene Mutation Database (HGMD). In another embodiment, the target genomic DNA sequence can include genes that, in combination with mutations in other genes, contribute to a particular disease. Diseases caused by the contributions of multiple genes lacking a simple (i.e., Mendelian) inheritance pattern are referred to in the art as "multifactorial" or "polygenic" diseases. Examples of multifactorial or polygenic diseases include, but are not limited to, asthma, diabetes, epilepsy, hypertension, bipolar disorder, and schizophrenia. Certain developmental abnormalities can also be inherited in a multifactorial or polygenic pattern, such as cleft lip and palate, congenital heart disease, and neural tube defects. In another embodiment, the target DNA sequence can include cancer genes.
[0273] The present disclosure provides gene editing methods that can remove disease-related genes (e.g., cancer genes) and can thus be used for in vivo gene therapy in patients. In some embodiments, the gene editing method includes a donor nucleic acid comprising a therapeutic gene.
[0274] When used as a therapy, the effective amount can depend on individual patient parameters including the particular condition being treated, the severity of the condition, age, physical condition, build, gender and weight, the duration of the treatment, the nature of any concurrent therapies (if any), the particular route of administration, and similar factors within the knowledge and expertise of the medical practitioner. In some embodiments, the effective amount results in alleviation, reduction, improvement, amelioration, reduction of symptoms, or delay of progression of any disease or disorder of the subject. In some embodiments, the subject is human.
[0275] In conjunction with the methods of the present disclosure, a wide range of additional therapies can be used. The additional therapy may be administration of an additional therapeutic agent or may be an additional therapy not related to the administration of another agent. Such additional therapies include, but are not limited to, surgery, immunotherapy, and radiation therapy. The additional therapy can be performed concurrently with the above methods. In some embodiments, the additional therapy can be performed before or after treatment by the disclosed methods at time intervals ranging from several hours to several months.
[0276] In some embodiments, a therapeutically effective amount of a system (e.g., nuclease and / or gRNA) or composition described herein is administered alone or in combination with a therapeutically effective amount of at least one additional therapeutic agent. In some embodiments, an effective combination therapy is achieved using a single composition or pharmaceutical formulation, or using two different compositions or formulations administered simultaneously or at intervals. The at least one additional therapeutic agent can include any type of therapeutic agent, including proteins, small molecules, nucleic acids, etc. For example, exemplary additional therapeutic agents include, but are not limited to, immunomodulators, chemotherapeutic agents, nucleic acids (e.g., mRNA, aptamers, antisense oligonucleotides, ribozyme nucleic acids, interfering RNAs, antigen nucleic acids), decongestants, steroids, analgesics, antibacterial agents, immunotherapeutic agents, or any combination thereof.
[0277] In the context of the present disclosure, insofar as any of the disease states recited herein are concerned, terms such as "treating", "treatment" etc. mean reducing or alleviating at least one symptom associated with such a state, or delaying or reversing the progression of such a state. In the meaning of the present disclosure, the term "treating" also refers to preventing a disease, delaying its onset (e.g., the period until the clinical symptoms of the disease appear) and / or reducing the risk of developing or worsening the disease. For example, in the context of cancer, the term "treating" may mean the disappearance or reduction of the tumor mass in a patient, or the prevention, delay, or suppression of metastasis etc.
[0278] The expression "pharmaceutically acceptable" as used with respect to the compositions and / or cells of the present disclosure means that the molecular entities and other ingredients of such compositions are physiologically acceptable and do not typically produce adverse reactions when administered to a subject (e.g., a mammal, a human). Preferably, as used herein, the term "pharmaceutically acceptable" means approved by a regulatory agency of the federal or state government or listed in the United States Pharmacopeia or other generally recognized pharmacopeias for use in mammals, more particularly in humans. "Acceptable" means that the carrier is compatible with the active ingredients (e.g., nucleic acids, vectors, cells, or therapeutic antibodies) of the composition and does not adversely affect the subject to which the composition(s) is administered. Any of the pharmaceutical compositions and / or cells used in the present methods may contain a pharmaceutically acceptable carrier, excipient, or stabilizer, either in lyophilized form or in aqueous solution form.
[0279] Pharmaceutically acceptable carriers are well known in the art, including buffers, such as phosphates, citrates, and other organic acids; antioxidants, including ascorbic acid and methionine; preservatives; low molecular weight polypeptides; proteins, such as serum albumin, gelatin, or immunoglobulins; amino acids; hydrophobic polymers; monosaccharides; disaccharides; and other carbohydrates; metal complexes; and / or nonionic surfactants. See, e.g., Remington: The Science and Practice of Pharmacy 20th Ed. (2000) Lippincott Williams and Wilkins, Ed. K.E. Hoover.
[0280] In some cases, the desired delivery system provides a substantially uniform distribution and has a controllable release rate of its components (e.g., vector, protein, nucleic acid) in vivo. Various different media useful for making the composition delivery system are described below. It is not intended that any one medium limit the invention. Any medium can be combined with another medium or carrier. It should be noted, for example, that in one embodiment, polymeric microparticles conjugated to a compound can be combined with a gel medium. Implantable devices can be used to deliver, for example, nucleases or nucleic acids encoding them and gRNAs or nucleic acids encoding them to target cells in vivo.
[0281] The intended carriers or media include materials such as gelatin, collagen, cellulose ester, dextran sulfate, poly sulfate pentosan, chitin, saccharide, albumin, fibrin sealant, synthetic polyvinyl pyrrolidone, polyethylene oxide, polypropylene oxide, block polymers of polyethylene oxide and polypropylene oxide, polyethylene glycol, acrylate, acrylamide, methacrylate, for example, but not limited to, 2-hydroxyethyl methacrylate, poly(ortho ester), cyanoacrylate, gelatin-resorcinol-aldehyde type bioadhesive, polyacrylic acid, and their copolymers and block copolymers.
[0282] In some cases, the carrier / media may include microparticles. The microparticles may include, but are not limited to, liposomes, nanoparticles, microspheres, nanospheres, microcapsules, and nanocapsules. In some cases, the microparticles may include one or more of the following: poly(lactide-co-glycolide), aliphatic polyesters, for example, but not limited to, polyglycolic acid and polylactic acid, hyaluronic acid, modified polysaccharides, chitosan, cellulose, dextran, polyurethane, polyacrylic acid, pseudo poly(amino acids), polyhydroxybutyrate related copolymers, polyanhydrides, polymethyl methacrylate, poly(ethylene oxide), lecithin and phospholipids, and any combination thereof.
[0283] In some cases, the carrier / media can include liposomes that enable the attachment and release of therapeutic agents (e.g., nucleic acids and / or proteins of interest). Liposomes are tiny spherical lipid bilayers that surround an aqueous core and are made from amphiphilic molecules such as phospholipids. For example, liposomes can encapsulate therapeutic agents between the hydrophobic tails of phospholipid micelles. Water-soluble agents can be encapsulated in the core, and lipid-soluble agents can dissolve in the shell-like bilayer. Liposomes have the special feature of allowing water-soluble and water-insoluble chemicals to be used together in a medium without using surfactants or other emulsifiers. Liposomes can be formed naturally by forcing the mixing of phospholipids in an aqueous medium. Water-soluble compounds are dissolved in an aqueous solution capable of hydrating the phospholipids. Thus, these compounds are trapped in the center of the aqueous liposomes during liposome formation. The liposome wall is a phospholipid membrane that holds lipid-soluble substances such as oils. Liposomes provide controlled release of the incorporated compounds. In addition, liposomes can also be coated with water-soluble polymers such as polyethylene glycol to increase the pharmacokinetic half-life.
[0284] In some embodiments, cationic or anionic liposomes are used as part of the composition or method of interest, or liposomes with neutral lipids can also be used. Cationic liposomes can include negatively charged materials by mixing the materials and fatty acid liposome components and charge associating. The choice of cationic or anionic liposomes depends on the desired pH of the final liposome mixture.
[0285] Any element of any suitable CRISPR / Cas gene editing system known in the art can be appropriately adopted in the systems and methods described herein. The CRISPR / Cas gene editing technology is described in, for example, U.S. Patent Nos. 8,546,553; 8,697,359; 8,771,945; 8,795,965; 8,865,406; 8,871,445; 8,889,356; 8,889,418; 8,895,308; 8,906,616; 8,932,814; 8,945,839; 8,993,233; 8,999,641; 9,115,348; 9,149,049; 9,493,844; 9,567,603; 9,637,739; 9,663,782; 9,404,098; 9,885,026; 9,951,342; 10,087,431; 10,227,610; 10,266,850; 10,601,748; 10,604,771; and 10,760,064; and U.S. Patent Application Publication Nos. US2010 / 0076057; US2014 / 0113376; US2015 / 0050699; US2015 / 0031134; US2014 / 0357530; US2014 / 0349400; US2014 / 0315985; US2014 / 0310830; US2014 / 0310828; US2014 / 0309487; US2014 / 0294773; US2014 / 0287938; US2014 / 0273230; US2014 / 0242699; US2014 / 0242664; US2014 / 0212869; US2014 / 0201857; US2014 / 0199767; US2014 / 0189896; US2014 / 0186919; US2014 / 0186843; and US2014 / 0179770, each of which is incorporated herein by reference.
[0286] Kit Also included within the scope of the present disclosure are kits that include the compositions, systems, or components thereof disclosed herein.
[0287] For example, the kit may contain one or more reagents or other components that are useful, necessary, or sufficient for performing any of the methods described herein, such as editing reagents (nucleases, guide RNAs, vectors, compositions, etc.), transfection or administration reagents, negative and positive control samples (e.g., cells, template DNA), cells, containers for containing one or more components (e.g., microcentrifuge tubes, boxes), detectable labels, detection and analysis instruments, software, instructions, and the like.
[0288] The kit may include instructions for use in any of the methods described herein. The instructions may include instructions for administering the system or composition to a subject to achieve the intended effect. The instructions generally include information regarding dosage, dosing schedule, and route of administration for the intended treatment. The kit may further include instructions for selecting a suitable subject for treatment based on identification of whether the subject is in need of treatment.
[0289] The kits provided herein are contained in suitable packaging. Suitable packaging includes, but is not limited to, vials, bottles, jars, flexible packaging, and the like. The kit may have a sterile access port (e.g., the container may be an intravenous solution bag or vial having a stopper pierceable by a hypodermic needle). The container may have a sterile access port.
[0290] The packaging may be a unit dose, bulk packaging (e.g., multiple dose packaging), or divided unit dose. The instructions included with the kits of the present disclosure are typically instructions written on a label or package insert. The label or package insert indicates that the pharmaceutical composition is used for the treatment, delay in the onset, and / or alleviation of the subject's disease or disorder.
[0291] The kit may optionally include additional components such as buffers and explanatory information. Usually, the kit includes a container and a label or accompanying document(s) on or associated with the container. In some embodiments, the present disclosure provides a manufactured article that includes the contents of the above kit.
[0292] The kit may further include a device for holding or administering the present system or composition. The device may include an infusion device, an intravenous solution bag, a hypodermic needle, a vial, and / or a syringe.
Examples
[0293] The following are examples of the present invention and should not be construed as limiting.
[0294] Example 1 Nuclease and guide RNA vector As a CRISPR type V nuclease candidate having Cas12f-like characteristics of a single guide RNA vector set, nuclease sequences (SEQ ID NOs: 1 to 250) were identified. For the nucleases of SEQ ID NOs: 1 to 54, single guide RNA (sgRNA) vectors were designed based on their predicted binding and folding patterns of crRNA and tracrRNA (Table 5). The designed sgRNA was placed downstream of the U6 promoter having a starting G and then upstream of the spacer sequence (Table 6).
[0295] Nuclease expression vector Condon-optimized genes encoding nuclease candidates (nuclease amino acid sequences of SEQ ID NOs: 20 to 29 and 36) were synthesized and cloned into a mammalian expression vector under the CMV promoter pTwist_CMV (Twist Biosciences). The cloned nuclease was fused with an SV40 nuclear localization sequence (NLS) at the N-terminus, a nucleoplasmin NLS followed by a 3x HA tag at the C-terminus, and placed in the expression vector. A similar vector was prepared using Un1Cas12f1 (SEQ ID NO: 471).
[0296] Example 2 Editing activity in human cells The nucleases of SEQ ID NOs: 21, 24, and 36 were tested in HEK293T cells by plasmid transfection using Mirus Transit X2 reagent. 50,000 cells were plated in each well of a 96-well plate and immediately transfected with 100 ng of the nuclease expression vector and 100 ng of the corresponding sgRNA vector shown in Table 1. [Table 1]
[0297] The sample was incubated for 72 hours and recovered using QuickExtract (Lucigen). Approximately 200 ng of genomic DNA was amplified using KAPA HiFi polymerase and primers specific to the target region on chromosome 3 along with the Illumina adapter ACACTCTTTCCCTACACGACGCTCTTCCGATCTgtaatgagcaaccttgagggatcagg (SEQ ID NO: 506) and GACTGGAGTTCAGACGTGTGCTCTTCCGATCTctcatggcaaaagcagtaatcagaac (SEQ ID NO: 507). 2 μL of this first 25 μL PCR was input into the second PCR using the Illumina P7 barcode - attached primer of New England BioLabs kit number E6609S. The purity of the PCR product was confirmed on a 2% agarose gel and washed with the ZYMO kit number D4034. The sample was then sequenced on the Illumina MiSeq system to obtain 100,000 - 400,000 150 - bp paired - end reads per sample. Editing analysis was performed using CRISPResso2 with the option “--cleavage_offset 1” (Clement, Kendell, et al. “CRISPResso2 provides accurate and rapid genome editing sequence analysis.” Nature biotechnology 37.3 (2019): 224 - 226.). The percentage of nucleotide insertions or deletions (indels) around the cleavage site was calculated for transfected cells and non - transfected (NT) cells without including mutations with substitutions only. The fold - change in editing was calculated by dividing the indel percentage of transfected cells by the indel percentage of non - transfected cells. The results are shown in Figure 1.
[0298] Example 3 Engineered single - guide RNA Single-guide RNA (sgRNA) vectors engineered against nucleases with accession numbers 21, 24, and 36 were designed at various lengths as shown in Table 2. The designed sgRNAs were placed downstream of the U6 promoter with a starting G, then targeted the intergenic region of chromosome 3 of the human genome, and were placed upstream of the spacer sequence CACACACACAGTGGGCTACC (SEQ ID NO: 423) with a 5’ TTTG PAM sequence. Nucleases with accession numbers 21, 24, and 36 were tested in HEK293T cells by plasmid transfection using Mirus Transit X2 reagent. 50,000 cells were plated in each well of a 96-well plate and immediately transfected with 100 ng of the nuclease expression vector and 100 ng of the corresponding sgRNA vector. Samples were incubated for 72 hours and recovered using QuickExtract (Lucigen). Genomic DNA was amplified around the target region on chromosome 3 and sequenced by Sanger sequencing. TIDE (Tracking of Indels by Decomposition) analysis was performed according to the method of Brinkman et al. (Brinkman EK, Chen T, Amendola M, van Steensel B. Nucleic Acids Res. 2014;42(22):e168, which is hereby incorporated by reference in its entirety) and the recommendations at tide.nki.nl. The results are shown in Figure 2. Table 3 shows the corresponding nuclease and guide RNA sequences for each numbered sample. Editing was improved by the use of some cleavage-type sgRNAs.
[0299] Example 4 Editing activity in human cells The editing activities of nucleases with accession numbers 20 - 29 and 36 were tested in HEK293T cells targeting Kim-T1 (SEQ ID NO: 423) using the sgRNA of SEQ ID NO: 346 according to the method described in Example 2. The results shown in Figure 3 indicated that the selected nucleases had editing activity in human cells.
[0300] Example 5 Off-target editing activity The nuclease of SEQ ID NO: 20 was tested as described in Example 3 using either a guide that matches the TCRA gene (SEQ ID NO: 430) or a guide that contains a single mismatch (SEQ ID NOs: 433-452) at a different position of TCRA. The mismatch guides functioned as artificial off-targets and determined the editing tendency of the nuclease when there was a mismatch at each position of the guide. The editing efficiencies of the match guide and the mismatch guides were measured by Sanger sequencing as described in Example 3. The resulting amplicons were Sanger sequenced, and TIDE analysis was performed according to the method of Brinkman et al., 2014 and the recommendations of the TIDE website (tide.nki.nl). Non-transfected cells were also recovered, amplified, and sequenced in the same way to determine the limit of detection (L.O.D.) at which the editing level could not be determined. The results of the editing efficiency using single mismatch guide RNAs are shown in FIG. 4.
[0301] Example 6 Modification of guide RNA A single guide RNA (sgRNA) construct targeting Kim-T1 was designed based on the predicted binding and folding patterns of its crRNA and tracrRNA as described in Example 1 and cloned into a vector. The sgRNAs (Table 8) were tested using the nucleases having SEQ ID NOs: 20, 24, and 26 according to the method described in Example 3. The results for each of SEQ ID NOs: 20, 24, and 26 are shown in FIGS. 5A-5C respectively, and the results for SEQ ID NO: 20 and additional sequences are shown in FIG. 5D. The predicted structures and modifications of the sgRNAs are shown in FIG. 5E. Surprisingly, some of the modifications, such as the modification at SEQ ID NO: 346 where the predicted stem-loop was removed, enabled the sgRNA construct to function well with multiple nucleases. Even more surprisingly, many of the cleavages located within the stem and the upper loop retained functionality when combined with the nuclease of SEQ ID NO: 20.
[0302] Example 7 Modification of guide RNA The editing activities of nucleases having SEQ ID NO: 20, 24, 26, and Un1Cas12f1 (SEQ ID NO: 471) were compared across different target sites using the sgRNA having SEQ ID NO: 346 according to the method described in Example 3. The results are shown in FIG. 6. The results showed that each of the nucleases could edit various genomic target sites at various levels. Surprisingly, Un1Cas12f1 did not show editing above the background level at the Kim-T1 site when combined with the sgRNA having SEQ ID NO: 346, whereas the other three nucleases showed editing activity with this sgRNA.
[0303] Example 8 Modification of tracrRNA The editing activities of the nucleases of SEQ ID NO: 20 and 21 were compared using sgRNAs having small deletions in the tracrRNA sequence according to the method described in Example 3. The deletions in tracrRNA and the editing results are shown in Table 9.
[0304] Next, the nuclease of SEQ ID NO: 20 was tested with several sgRNA modifications that change the predicted structure of the tracrRNA sequence. Two structures with longer repeats or truncated repeats (see FIG. 7A) were tested and compared to the modification with a truncated 5' stem (SEQ ID NO: 346). In particular, having a complete repeat was not favorable for editing activity compared to the other truncated forms (FIG. 7B).
[0305] To further investigate the relationship of the tracrRNA sequence to these nucleases, additional modifications were constructed. Starting from SEQ ID NO: 346, a portion of the 5' stem and 3' tail of tracrRNA were removed to evaluate their importance in editing efficiency (FIG. 7C). Further removal of the 5' stem did not affect editing, but removal of the 3' tail of tracrRNA was very detrimental to editing and was at an efficiency similar to that observed in non-target cells (FIG. 7D).
[0306] To further evaluate the role of the stem bases, the base pairing was strengthened by changing A-T to G-C, which indicates "stem stability", and separately, this sequence was modified by removing the kink inserted by the unpaired A single nucleotide just above (Figure 7C). By improving the stem stability, the predicted ΔG of the structure changed, but the editing efficiency of the nuclease of SEQ ID NO: 20 did not improve. When the A-kink was removed, the editing ability of the nuclease was completely lost (Figure 7E).
[0307] Example 9 Modification of the Spacer The editing activity of the nuclease of SEQ ID NO: 20 was evaluated for the editing activity with sgRNAs having different spacer sequence lengths according to the method described in Example 3. The editing results are shown in Figure 8. A spacer length of 18-20 nucleotides was optimal for the editing activity.
[0308] Example 10 PAM Preference The effect of the PAM sequence on the editing efficiency of the nuclease was tested according to the method using Spacer 3 of Walton et al. (Walton RT, et al., Science. 2020 Apr 17;368(6488):290-296, which is hereby incorporated by reference in its entirety). Briefly, it is a spacer capable of targeting a randomized PAM plasmid library prepared by incorporating a 10-bp randomized PAM downstream of the tracrRNA and repeat region of the gRNA. The PAMs effective for the nuclease were depleted in this process, and the remaining PAMs were identified by next-generation sequencing (NGS). The preferred PAM sequences for nuclease SEQ ID NOs: 20 and 26 are listed in Table 10. The values are calculated based on Walton et al., and the PAM preferences are described in order of priority (the top of each list represents the more preferred sequence).
[0309] The identified PAM sequences were tested for editing activity using nuclease SEQ ID NOs: 20 and 26 in relation to several spacers of the sgRNA. The results for target sequences with high editing levels (X-axis) (Figure 9A) and low editing levels (Figure 9B) in combination with various PAM sequences (PAM sequences are shown in brackets above the bars) are shown in Figures 9A and 9B. Surprisingly, the nuclease has a PAM preference different from that of known Cas12f nucleases such as Un1Cas12f1, AsCas12f, and SpaCas12f1. For the nucleases tested (SEQ ID NOs: 20, 21, and 26), the preferred PAM sequence was DTTR (where D is A, G, or T and R is A or G), and there was a strong bias towards the ATTA PAM. In contrast, for Un1Cas12f1 and AsCas12f, the PAM preference was TTTR, and for SpaCas12f1, the PAM preference was NTTY (where N can be any base).
[0310] Example 11 Design of AAV vectors and editing in mammalian cells A single AAV vector for delivering the nuclease of SEQ ID NO: 20 and the sgRNA (shown as Tracr in Figure 10) to mammalian cells was designed using a CMV promoter and an SV40 nuclear localization sequence at the 5' end of the nuclease, as well as an HA tag and a nucleoplasmin localization sequence at the 3' end, followed by a U6 promoter driving the expression of the sgRNA. The figure of the vector is shown in Figure 10.
[0311] Using this vector design, constructs were made as shown in Table 11 containing different sgRNAs designed against different targets but the same nuclease.
[0312] Constructs against human targets were tested in HEK293T cells and constructs against mouse targets were tested in NIH3T3 cells. Cells were seeded at 3×10 on day 0 5Cells were plated at confluence of cells / m. On day 1, the cells were transduced at 100K MOI. On day 2, etoposide (which promotes AAV delivery) was added to the cells to a final concentration of 60 mM, and the cells were imaged on day 3. The cells were incubated for 72 hours and then harvested according to the method of Example 2. After DNA extraction, samples for NGS were prepared by amplifying each region with the NGS-specific primers listed in Table 12. NGS reads were processed using the CRISPRESSO2 tool (Clement, Kendell, et al. Nature biotechnology 37.3 (2019): 224 - 226, which is hereby incorporated by reference in its entirety). The editing data for each construct are shown in Figure 11.
[0313] The SMN2 and TTR constructs were further tested for editing in HEK293T cells and NIH3T3 cells with and without etoposide treatment. According to the above method, except that the MOI was 10K, etoposide was added on day 1 to treat the cells, the AAV vector was added on day 2, and the cells were harvested on day 7. Samples for NGS were prepared using the primers in Table 9. NGS paired reads were processed using CRISPRESSO2 (Clement et al., 2019). The editing efficiency is shown in Figure 12. NIH3T3 cells were resistant to etoposide treatment, and generally, editing was improved in the treated cells. In contrast, HEK293T cells showed signs of toxicity, and editing decreased in the treated cells compared to the cells not treated with etoposide.
[0314]
Table 2 - 1
Table 2 - 2
Table 2 - 3
Table 2 - 4
Table 2 - 5
Table 2-6
Table 2-7
[0315]
Table 3
[0316]
Table 4-1
Table 4-2
Table 4-3
Table 4-4
Table 4-5
Table 4-6
Table 4-7
Table 4-8
Table 4-9
Table 4-10
Table 4-11
Table 4-12
Table 4-13
Table 4-14
Table 4-15
Table 4-16
Table 4-17
Table 4-18
Table 4-19
Table 4-20
Table 4-21
Table 4-22
Table 4-23
Table 4-24
Table 4-25
Table 4-26
Table 4-27
Table 4-28
Table 4-29
Table 4-30
Table 4-31
Table 4-32
Table 4-33
Table 4-34
Table 4-35
Table 4-36
Table 4-37
Table 4-38
Table 4-39
Table 4-40
Table 4-41
Table 4-42
Table 4-43
Table 4-44
Table 4-45
Table 4-46
Table 4-47
Table 4-48
Table 4-49
Table 4-50
Table 4-51
Table 4-52
Table 4-53
[0317]
Table 5-1
Table 5-2
Table 5-3
Table 5-4
Table 5-5
Table 5-6
Table 5-7
Table 5-8
Table 5-9
[0318]
Table 6
[0319]
Table 7
[0320]
Table 8
[0321]
Table 9
[0322]
Table 10
[0323]
Table 11
[0324]
Table 12
[0325] The scope of the present invention is not limited to what is specifically shown and described herein. Those skilled in the art will understand that there are suitable alternatives to the examples of materials, configurations, structures, and dimensions shown. Variations, modifications, and other implementations of what is described herein will come to the mind of those skilled in the art without departing from the spirit and scope of the present invention.
[0326] In the description of the present invention, a number of references, including patents and various publications, are cited and described. The citation and description of such references are provided solely to clarify the description of the present invention and do not admit that they are prior art of the present invention described herein. All of the references cited and described in this specification are hereby incorporated by reference in their entirety.
Claims
1. A composition comprising a nuclease, wherein the nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or at least 99% identity to any one of SEQ ID NOs: 1 to 250.
2. The composition according to claim 1, wherein the amino acid sequence of the nuclease comprises any one of SEQ ID NOs: 1 to 250.
3. The composition according to claim 1 or 2, wherein the nuclease further comprises a nuclear localization sequence (NLS) at the N-terminus, C-terminus, or both the N-terminus and C-terminus of the nuclease.
4. The composition according to claim 3, wherein the NLS at the N-terminus and the NLS at the C-terminus of the nuclease are different sequences.
5. A nucleic acid comprising a first polynucleotide sequence encoding the nuclease according to any one of claims 1 to 4.
6. A vector comprising the nucleic acid according to claim 5.
7. The vector according to claim 6, further comprising a promoter operably linked to the first polynucleotide.
8. The vector according to claim 6 or 7, further comprising a second polynucleotide sequence encoding a guide RNA (gRNA).
9. The vector according to claim 8, further comprising a promoter operably linked to the second polynucleotide sequence.
10. The vector according to claim 8 or 9, wherein the gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identity to any one of SEQ ID NOs: 251 to 422 and 472 to 482.
11. The vector according to any one of claims 8 to 10, wherein the gRNA comprises any one of SEQ ID NOs: 251 to 343.
12. The vector according to any one of claims 8 to 10, wherein the gRNA comprises any one of SEQ ID NOs: 344 to 422.
13. The vector according to any one of claims 8 to 10, wherein the gRNA comprises any one of SEQ ID NOs: 472 to 482.
14. The vector according to any one of claims 8 to 13, wherein the gRNA comprises a tracr sequence and the gRNA comprises one or more sequence deletions within or in the vicinity of the region encompassing the tracr sequence.
15. The vector according to claim 14, wherein the one or more array deletions include an array predicted to form a stem-loop structure. **Claim 16** The vector according to claim 14 or 15, wherein the one or more array deletions include an array predicted to form a stem-loop structure at or near the 5'-end of the gRNA. **Claim 17** The vector according to any one of claims 14 to 16, wherein the gRNA includes SEQ ID NO:
346. **Claim 18** The vector according to any one of claims 14 to 16, wherein the gRNA includes SEQ ID NO:
420. **Claim 19** The vector according to any one of claims 14 to 16, wherein the gRNA includes SEQ ID NO:
481. **Claim 20** The vector according to any one of claims 14 to 16, wherein the gRNA includes SEQ ID NO:
479. **Claim 21** The vector according to any one of claims 8 to 20, wherein the gRNA includes a spacer sequence having a length of at least 18 nucleotides or a length of 18 to 20 nucleotides. **Claim 22** A system for modifying a target nucleic acid, comprising: a) a nuclease comprising an amino acid sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any of SEQ ID NOs: 1 to 250, or a nucleic acid encoding the nuclease; and b) at least one guide RNA (gRNA) comprising a sequence complementary to at least a portion of the target nucleic acid and a region that associates with the nuclease, or a nucleic acid encoding the at least one gRNA. **Claim 23** The system according to claim 22, wherein the nuclease is capable of recognizing a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. **Claim 24** The system according to claim 22 or 23, wherein the gRNA includes a spacer sequence complementary to the first strand sequence of the target nucleic acid, wherein the first strand sequence is directly adjacent to a protospacer adjacent motif (PAM) sequence selected from the group consisting of ATTA, GTTA, ATTG, GTTG, TTTA, TTTG, CTTA, and CTTG. **Claim 25** The system according to claim 23 or 24, wherein the PAM sequence contains DTTR, where D is A, G, or T, and R is A or G. **Claim 26** The system according to any one of claims 22 to 25, wherein the nuclease can preferentially modify a target nucleic acid containing a PAM sequence of ATTA as compared to a target nucleic acid containing a PAM sequence of TTTTR (R is A or G). **Claim 27** The system according to any one of claims 22 to 25, wherein the nuclease can modify the target nucleic acid with higher efficiency as compared to the efficiency of modification of the target nucleic acid by the nuclease of SEQ ID NO: 471, wherein the target nucleic acid contains a PAM sequence that is ATTA. **Claim 28** The system according to any one of claims 22 to 27, wherein the modification includes nucleic acid cleavage. **Claim 29** The system according to any one of claims 22 to 28, wherein the modification includes one or more of modification of the target nucleic acid, regulation of transcription from the target nucleic acid, and modification of a polypeptide related to the target nucleic acid. **Claim 30** The system according to any one of claims 22 to 29, wherein the nuclease further includes a nuclear localization sequence (NLS) at the N-terminus, C-terminus, or both the N-terminus and C-terminus of the nuclease. **Claim 31** The system according to claim 30, wherein the NLS at the N-terminus and the NLS at the C-terminus of the nuclease have different sequences. **Claim 32** The system according to any one of claims 22 to 31, wherein the nuclease further includes a purification tag. **Claim 33** The system according to any one of claims 22 to 32, wherein the at least one gRNA further includes a sequence complementary to at least a part of a second target nucleic acid. **Claim 34** The system according to any one of claims 22 to 33, wherein the at least one gRNA includes a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 251 to 422. **Claim 35** The system according to claim 34, wherein the at least one gRNA includes any one of SEQ ID NOs: 251 to 343. **Claim 36** The system according to claim 34, wherein the at least one gRNA includes any one of SEQ ID NOs: 344 to 422. **Claim 37** The system according to claim 34, wherein the at least one gRNA comprises any one of SEQ ID NOs: 472 to 482.
38. The system according to claim 34, wherein the at least one gRNA comprises SEQ ID NO:
346.
39. The system according to claim 34, wherein the at least one gRNA comprises SEQ ID NO:
420.
40. The system according to claim 34, wherein the at least one gRNA comprises SEQ ID NO:
481.
41. The system according to claim 34, wherein the at least one gRNA comprises SEQ ID NO:
479.
42. The system according to any one of claims 22 to 41, wherein the at least one gRNA comprises a spacer sequence having a length of at least 18 nucleotides or a length of 18 to 20 nucleotides.
43. The system according to any one of claims 22 to 42, wherein the nuclease comprises SEQ ID NO: 20, and the at least one gRNA comprises any one of SEQ ID NOs: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479, and 481, or any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392, and 417, or any one of SEQ ID NOs: 346 and 362, or any one of SEQ ID NOs: 410 - 419, and comprises a sequence having an identity of at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% with respect thereto.
44. The system according to any one of claims 22 to 43, wherein the nuclease comprises a sequence having an identity of at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% with respect to SEQ ID NO: 20, and the at least one gRNA comprises any one of SEQ ID NOs: 309, 346, 352, 358, 362 - 364, 380, 392 - 395, 410 - 420, 472 - 479, and 481, or any one of SEQ ID NOs: 352, 358, 363, 364, 380, 392, and 417, or any one of SEQ ID NOs: 346 and 362, or any one of SEQ ID NOs: 410 - 419.
45. The nuclease comprises SEQ ID NO: 21, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482. The system according to any one of claims 22-42.
46. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 21, and the at least one gRNA comprises any one of SEQ ID NOs: 310, 344-349, 361-366, 404-422, and 479-482. The system according to any one of claims 22-42.
47. The nuclease comprises SEQ ID NO: 22, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 311, 346, 381, and 398-399. The system according to any one of claims 22-42.
48. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 22, and the at least one gRNA comprises any one of SEQ ID NOs: 311, 346, 381, and 398-399. The system according to any one of claims 22-42.
49. The nuclease comprises SEQ ID NO: 23, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 312, 346, and 382. The system according to any one of claims 22-42.
50. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 23, and the at least one gRNA comprises any one of SEQ ID NO: 312, 346, and 382. The system according to any one of claims 22 to 42.
51. The nuclease comprises SEQ ID NO: 24, and the at least one gRNA comprises any one of SEQ ID NO: 310, 313, 325, 346, 350 to 355, 358, 361 to 363, 367 to 372, and 389 to 392, or a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NO: 346, 352, 358, 361, 362, 368, 369, and 392. The system according to any one of claims 22 to 42.
52. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 24, and the at least one gRNA comprises any one of SEQ ID NO: 310, 313, 325, 346, 350 to 355, 358, 361 to 363, 367 to 372, and 389 to 392, or any one of SEQ ID NO: 346, 352, 358, 361, 362, 368, 369, and 392. The system according to any one of claims 22 to 42.
53. The nuclease comprises SEQ ID NO: 25, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NO: 314, 346, 383, and 400. The system according to any one of claims 22 to 42.
54. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 25, and the at least one gRNA comprises any one of SEQ ID NO: 314, 346, 383, and 400. The system according to any one of claims 22 to 42.
55. The nuclease comprises SEQ ID NO: 26, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481, or any one of SEQ ID NOs: 346, 384, and 392. The system according to any one of claims 22 to 42.
56. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 26, and the at least one gRNA comprises any one of SEQ ID NOs: 315, 346, 384, 392, 396-397, 420, 479, and 481, or any one of SEQ ID NOs: 346, 384, and 392. The system according to any one of claims 22 to 42.
57. The nuclease comprises SEQ ID NO: 27, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 316, 346, 385, and 401. The system according to any one of claims 22 to 42.
58. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 27, and the at least one gRNA comprises any one of SEQ ID NOs: 316, 346, 385, and 401. The system according to any one of claims 22 to 42.
59. The nuclease comprises SEQ ID NO: 28, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 317, 346, 386, and 402. The system according to any one of claims 22 to 42.
60. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 28, and the at least one gRNA comprises any one of SEQ ID NOs: 317, 346, 386, and 402. The system according to any one of claims 22 to 42.
61. The nuclease comprises SEQ ID NO: 29, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 318, 346, 387, and 403. The system according to any one of claims 22 to 42.
62. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 29, and the at least one gRNA comprises any one of SEQ ID NOs: 318, 346, 387, and 403. The system according to any one of claims 22 to 42.
63. The nuclease comprises SEQ ID NO: 36, and the at least one gRNA comprises a sequence having at least 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity or 100% identity to any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378. The system according to any one of claims 22 to 42.
64. The nuclease comprises a sequence having at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO: 36, and the at least one gRNA comprises any one of SEQ ID NOs: 310, 313, 325, 346, 356 - 360, and 373 - 378. The system according to any one of claims 22 to 42.
65. The nucleic acid molecule encoding one or both of the nuclease and the at least one gRNA comprises messenger RNA, a vector, or a combination thereof. The system according to any one of claims 22 to 64.
66. The system according to any one of claims 22 to 65, wherein the nuclease and the at least one gRNA are encoded on one nucleic acid.
67. The system according to claim 66, wherein the nuclease and the at least one gRNA are operably linked to different promoters.
68. The system according to claim 66 or 67, wherein the one nucleic acid is a vector.
69. The system according to claim 68, wherein the vector is a viral vector.
70. The system according to claim 69, wherein the viral vector is an AAV vector.
71. A kit comprising the system according to any one of claims 22 to 70.
72. A cell comprising the system according to any one of claims 22 to 70.
73. The cell according to claim 72, wherein the cell is a prokaryotic cell or a eukaryotic cell.
74. The cell according to claim 72 or 73, wherein the cell is a mammalian cell.
75. The cell according to any one of claims 72 to 74, wherein the cell is a human cell.
76. A method of modifying a selected target nucleic acid sequence, the method comprising contacting the selected target nucleic acid with the composition according to any one of claims 1 to 4, the nucleic acid according to claim 5, the vector according to any one of claims 6 to 21, or the system according to any one of claims 22 to 70.
77. The method according to claim 76, wherein the target nucleic acid sequence is intracellular.
78. The method according to claim 77, wherein the cell is a prokaryotic cell or a eukaryotic cell.
79. The method according to claim 77 or 78, wherein the cell is a mammalian cell.
80. The method according to any one of claims 76 to 78, wherein the cell is a human cell.
81. The contacting comprises introducing the composition according to any one of claims 1 to 4, the nucleic acid according to claim 5, the vector according to any one of claims 6 to 21, or the system according to any one of claims 22 to 69 into the cell, the method according to any one of claims 76 to 80.
82. The method according to any one of claims 75 to 80, wherein said contacting comprises administering and introducing into a subject the composition according to any one of claims 1 to 4, the nucleic acid according to claim 5, the vector according to any one of claims 6 to 21, or the system according to any one of claims 22 to 70.
83. The method according to any one of claims 76 to 82, wherein the selected target nucleic acid sequence encodes a gene product.
84. The composition according to any one of claims 1 to 4, the nucleic acid according to claim 5, the vector according to any one of claims 6 to 21, or the system according to any one of claims 22 to 70 for use in modifying a selected target nucleic acid sequence.
85. A kit comprising the composition according to any one of claims 1 to 4, the nucleic acid according to claim 5, the vector according to any one of claims 6 to 21, or the system according to any one of claims 22 to 70 for use in modifying a selected target nucleic acid sequence in an in vitro assay.