RNA-guided nucleases and active fragments and variants thereof and methods of use
RNA-guided nucleases and associated RNAs provide efficient and cost-effective solutions for precise genome editing and modification by binding and cleaving target nucleic acid molecules, addressing the inefficiencies of traditional methods.
Patent Information
- Application Number
- JP2025507841
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-12
- Filing Date
- 2023-08-12
- Publication Date
- 2025-08-26
AI Technical Summary
Existing genome editing methods, such as meganucleases, zinc finger fusion proteins, and TALENs, require the generation of chimeric nucleases for each target sequence, which is costly and inefficient, while RNA-guided nucleases like CRISPR-Cas systems can be more cost-effective but lack efficient methods for binding and modifying target nucleic acid molecules.
Compositions and methods using RNA-guided nuclease (RGN) polypeptides, CRISPR RNAs, and guide RNAs for binding, cleaving, or modifying target nucleic acid molecules, including vectors and host cells, to enable specific genome editing through non-homologous end joining, homologous recombination repair, or base editing.
Provides efficient and cost-effective methods for binding and modifying target nucleic acid molecules, enabling precise genome editing and expression control by leveraging RNA-guided nucleases and their ribonucleoprotein complexes.
Smart Images

Figure 2025528195000001 
Figure 2025528195000002 
Figure 2025528195000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 371,230, filed August 12, 2022, which is incorporated by reference herein in its entirety.
[0002] Reference to a sequence listing submitted electronically as an XML file This application contains a Sequence Listing that has been submitted in xml format via the USPTO Patent Center and is incorporated herein by reference in its entirety. The xml copy, created on August 11, 2023, is named L103438_1310WO_Seq_List.xml and is 1.49MB in size.
[0003] The present invention relates to the fields of molecular biology and gene editing. [Background technology]
[0004] Targeted genome editing or modification is rapidly becoming an important tool for basic and applied research. Initial methods involved engineering nucleases, such as meganucleases, zinc finger fusion proteins, or TALENs, which required the generation of chimeric nucleases with engineered, programmable, sequence-specific DNA-binding domains specific for each particular target sequence. RNA-guided nucleases, such as the clustered regularly interspaced short palindromic repeats (CRISPR)-associated (Cas) proteins of the CRISPR-Cas bacterial system, enable targeting of specific sequences by complexing the nuclease with a guide RNA that specifically hybridizes to the specific target sequence. Creating target-specific guide RNAs is less expensive and more efficient than creating chimeric nucleases for each target sequence. Such RNA-guided nucleases can be used to edit genomes by introducing sequence-specific double-strand breaks that are repaired by error-prone non-homologous end joining (NHEJ) to introduce mutations at specific genomic locations. Alternatively, heterologous DNA can be introduced into genomic sites via homology-directed repair. RNA-guided nucleases (RGNs) can also be used for base editing when fused with deaminases. Summary of the Invention [Problem to be solved by the invention]
[0005] Compositions and methods for binding to a target sequence of interest in a target nucleic acid molecule are provided. The compositions find use in cleaving or modifying a target nucleic acid molecule of interest, detecting a target sequence of interest, and modifying the expression of a gene of interest containing the target sequence. The compositions include RNA-guided nuclease (RGN) polypeptides, CRISPR RNAs (crRNAs), trans-activating CRISPR RNAs (tracrRNAs), guide RNAs (gRNAs), such as single guide RNAs (sgRNAs), nucleic acid molecules encoding the same, compositions containing the same, and vectors and host cells containing the nucleic acid molecules. RGN systems and ribonucleoprotein complexes for binding to a target sequence of interest are also provided, where the RGN system and ribonucleoprotein complexes comprise an RNA-guided nuclease polypeptide and one or more guide RNAs. Thus, the methods disclosed herein are directed to binding to a target sequence of interest in a target nucleic acid molecule and, in some embodiments, cleaving or modifying the target nucleic acid molecule of interest. The target nucleic acid molecule of interest can be modified, for example, as a result of non-homologous end joining, homologous recombination repair with an introduced donor sequence, or base editing. [Means for solving the problem]
[0006] In one aspect, the present disclosure provides a nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20.
[0007] In some embodiments of the above aspects, the RGN polypeptide is capable of binding to a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand when bound to a guide RNA (gRNA) that can hybridize to a non-target strand of the target sequence, in an RNA guide sequence-specific manner.
[0008] In some embodiments of the above aspects, the polynucleotide encoding the RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide.
[0009] In some embodiments of the above aspects, the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-20. In some embodiments, the RGN polypeptide comprises an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 1-20.
[0010] In some embodiments of the above aspects, the RGN polypeptide is capable of cleaving the target nucleic acid molecule upon binding. In some embodiments, the RGN polypeptide is capable of generating a double-stranded break. In some embodiments, the RGN polypeptide is capable of generating a single-stranded break.
[0011] In some embodiments of the above aspects, the RGN polypeptide is nuclease inactive or is a nickase.
[0012] In some embodiments of the above aspects, the RGN polypeptide is operably fused to a base-editing polypeptide. In some embodiments, the base-editing polypeptide is a deaminase, such as a cytosine deaminase or an adenine deaminase. In some embodiments, the deaminase has at least 90% or 100% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 481-552.
[0013] In some embodiments of the above aspects, the RGN polypeptide comprises one or more nuclear localization signals.
[0014] In some embodiments of the above aspects, the RGN polypeptide is codon-optimized for expression in eukaryotic cells.
[0015] In some embodiments of the above aspects, the target sequence is located adjacent to a protospacer adjacent motif (PAM).
[0016] In another aspect, the present disclosure provides a vector comprising the above-described nucleic acid molecule.
[0017] In some embodiments of the above aspects, the vector further comprises at least one nucleotide sequence encoding a gRNA capable of hybridizing to the non-target strand of the target sequence.
[0018] In some embodiments of the above aspects, the guide RNA comprises a) a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO:21; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:42; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1; and b) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:22. a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:43; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:2; c) a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:44 or nucleotides 19-111 of SEQ ID NO:1040; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3; and d) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24. and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:45; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4; and e) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:46 or nucleotides 22-85 of SEQ ID NO:1041 or 1042;a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 5; and g) a tracrRNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 26; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 24-138 of SEQ ID NO: 47 or SEQ ID NO: 1043; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 6; and g) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045. a guide RNA comprising i) a CRISPR RNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or nucleotides 27-95 of SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:7; and h) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:49; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:8; and RNA; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 50;a guide RNA comprising: an RGN polypeptide comprising an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:9; j) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:30; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:51; l) a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 32; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 53; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 12; and m) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33. a guide RNA comprising: n) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 13; and n) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55;o) a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 15; and p) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 36. q) a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 37; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 57; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 16; and q) a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 37; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 58; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 17; and r) a CRISPR RNA comprising i) a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 38 or 39. a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 40; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 59 or 60; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 18; and s) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 40; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 61;wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 19; and t) a guide RNA comprising i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 41; and ii) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 62; wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 20;
[0019] In some embodiments of the above aspects, the gRNA is a single guide RNA. In some embodiments of the above aspects, the gRNA is a dual guide RNA.
[0020] In another aspect, the disclosure provides a cell comprising the nucleic acid molecule or vector described above. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the human cell is an immune cell. In some embodiments, the immune cell is a stem cell. In some embodiments, the stem cell is an induced pluripotent stem cell. In some embodiments, the eukaryotic cell is an insect cell or an avian cell. In some embodiments, the eukaryotic cell is a fungal cell. In some embodiments, the eukaryotic cell is a plant cell.
[0021] In another aspect, the present disclosure provides a plant or seed comprising the plant cell described above.
[0022] In another aspect, the present disclosure provides a method for producing an RGN polypeptide comprising culturing a cell as described hereinabove under conditions in which the RGN polypeptide is expressed.
[0023] In one aspect, a method for producing an RNA-guided nuclease (RGN) polypeptide is provided, comprising: introducing into a cell a heterologous nucleic acid molecule comprising a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20; and culturing the cell under conditions in which the RGN polypeptide is expressed.
[0024] In some embodiments of the above aspects, the RGN polypeptide is capable of binding to a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand when bound to a guide RNA (gRNA) that can hybridize to a non-target strand of the target sequence, in an RNA guide sequence-specific manner.
[0025] In some embodiments of the above aspects, the RGN polypeptide comprises an amino acid sequence having at least 95% or 100% sequence identity to any one of SEQ ID NOs: 1-20.
[0026] In some embodiments of the above aspects, the method further comprises purifying the RGN polypeptide.
[0027] In some embodiments of the above aspects, the cell further expresses one or more guide RNAs capable of binding to the RGN polypeptide to form an RGN ribonucleoprotein complex. In some embodiments, the method further comprises purifying the RGN ribonucleoprotein complex.
[0028] In another aspect, the present disclosure provides an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20.
[0029] In some embodiments of the above aspects, the RGN polypeptide is capable of binding to a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand when bound to a guide RNA (gRNA) that can hybridize to a non-target strand of the target sequence, in an RNA guide sequence-specific manner.
[0030] In some embodiments of the above aspects, the RGN polypeptide comprises an amino acid sequence having at least 95% or 100% sequence identity to any one of SEQ ID NOs: 1-20.
[0031] In some embodiments of the above aspects, the RGN polypeptide is an isolated RGN polypeptide.
[0032] In some embodiments of the above aspects, the RGN polypeptide is capable of cleaving the target nucleic acid molecule upon binding. In some embodiments, cleavage by the RGN polypeptide generates a double-stranded break. In some embodiments, cleavage by the RGN polypeptide generates a single-stranded break.
[0033] In some embodiments of the above aspects, the RGN polypeptide is nuclease inactive or is a nickase.
[0034] In some embodiments of the above aspects, the RGN polypeptide is operably fused to a base-editing polypeptide. In some embodiments, the base-editing polypeptide is a deaminase. In some embodiments, the deaminase is a cytosine deaminase or an adenine deaminase. In some embodiments, the deaminase has at least 90%, at least 95%, or 100% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 481-552.
[0035] In some embodiments of the above aspects, the target sequence is located adjacent to a protospacer adjacent motif (PAM).
[0036] In some embodiments of the above aspects, the RGN polypeptide comprises one or more nuclear localization signals.
[0037] In another aspect, the present disclosure provides a ribonucleoprotein (RNP) complex comprising an RGN polypeptide as described herein above and a guide RNA bound to the RGN polypeptide.
[0038] In another aspect, a nucleic acid molecule comprising a CRISPR RNA (crRNA) or a polynucleotide encoding the crRNA is provided, wherein the crRNA comprises a spacer and a CRISPR repeat, and the CRISPR repeat comprises a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 21-41, or nucleotides 1-17 of SEQ ID NOs: 1041 or 1042, or nucleotides 1-22 of SEQ ID NOs: 1044 or 1045.
[0039] In some embodiments of the above aspects, a guide RNA comprising a crRNA and a trans-activating CRISPR RNA (tracrRNA) hybridized to the CRISPR repeats of the crRNA can hybridize to a non-target strand of a target sequence in a target nucleic acid molecule in a sequence-specific manner via the spacer of the crRNA when the guide RNA is bound to an RNA-guided nuclease (RGN) polypeptide.
[0040] In some embodiments of the above aspects, the polynucleotide encoding the crRNA is operably linked to a promoter heterologous to the polynucleotide.
[0041] In some embodiments of the above aspects, the CRISPR repeat comprises a nucleotide sequence having at least 95% or 100% sequence identity to any one of SEQ ID NOs:21-41, or nucleotides 1-17 of SEQ ID NO:1041 or 1042, or nucleotides 1-22 of SEQ ID NO:1044 or 1045.
[0042] In another aspect, the present disclosure provides a vector comprising a nucleic acid molecule comprising a polynucleotide encoding the above-described crRNA.
[0043] In some embodiments of the above aspects, the vector further comprises a polynucleotide encoding a tracrRNA. The tracrRNA is: a) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 42, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 21; b) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 43, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 22; c) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 19-111 of SEQ ID NO: 44 or SEQ ID NO: 1040, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 23; d) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 45. e) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 22-85 of SEQ ID NO: 46 or SEQ ID NO: 1041 or 1042, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 25 or nucleotides 1-17 of SEQ ID NO: 1041 or 1042; f) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 24-138 of SEQ ID NO: 47 or SEQ ID NO: 1043, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 26;g) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 27-96 of SEQ ID NO: 48 or SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045; h) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 49. i) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 50, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 29; j) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 51. k) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 52, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 31; l) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 53. m) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 54, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 33;n) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34; o) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35; p) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 57, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 36; q) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 58, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 37. r) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 59 or 60, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 38 or 39; s) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 61, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 40; and t) a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 62, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 41.
[0044] In some embodiments of the above aspects, the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as a single guide RNA.
[0045] In some embodiments of the above aspects, the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.
[0046] In some embodiments of the above aspects, the vector further comprises a polynucleotide encoding an RGN polypeptide, in some embodiments, the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 1, wherein the CRISPR repeats have at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 21, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 42; b) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 2, wherein the CRISPR repeats have at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 22, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 43; c) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:44 or nucleotides 19-111 of SEQ ID NO:1040; d) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:45;e) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:5, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:46 or nucleotides 22-85 of SEQ ID NO:1041 or 1042; f) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:6, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:26, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043. peptide; g) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:7, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:48 or nucleotides 27-96 of SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; h) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:8, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:49;i) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:9, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:29, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:50; j) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:10, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:30, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:51; k) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:11, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:31. l) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 12, wherein the CRISPR repeats have at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 32, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 53; m) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 13, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54;n) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 14, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55; o) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 15, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56; p) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 16, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 36. q) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 17, wherein the CRISPR repeats have at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 37, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 58; r) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 18, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 38 or 39, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 59 or 60;s) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 19, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 40, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 61; and t) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 20, wherein the CRISPR repeat has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 41, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 62.
[0047] In another aspect, the disclosure provides a nucleic acid molecule comprising a trans-activating CRISPR RNA (tracrRNA) or a polynucleotide encoding a tracrRNA comprising a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs:42-62, nucleotides 19-111 of SEQ ID NO:1040, nucleotides 22-85 of SEQ ID NOs:1041 or 1042, nucleotides 24-138 of SEQ ID NO:143, nucleotides 27-96 of SEQ ID NO:1044, or nucleotides 27-95 of SEQ ID NO:1045.
[0048] In some embodiments of the above aspects, the guide RNA comprises a tracrRNA and a crRNA comprising a spacer and CRISPR repeats, wherein the tracrRNA hybridizes to the CRISPR repeats of the crRNA. The guide RNA can hybridize to a non-target strand of a target sequence in a target nucleic acid molecule in a sequence-specific manner via the spacer of the crRNA when the guide RNA is bound to an RNA-guided nuclease (RGN) polypeptide.
[0049] In some embodiments of the above aspects, the polynucleotide encoding the tracrRNA is operably linked to a promoter that is heterologous to the polynucleotide.
[0050] In some embodiments of the above aspects, the tracrRNA comprises a nucleotide sequence having at least 95% or 100% sequence identity to any one of SEQ ID NOs: 42-62, nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NOs: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045.
[0051] In another aspect, the present disclosure provides a vector comprising a nucleic acid molecule comprising a polynucleotide encoding the tracrRNA described above.
[0052] In some embodiments of the above aspects, the vector further comprises a polynucleotide encoding a crRNA. In some embodiments, the crRNA is selected from the group consisting of: a) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:21, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:42; b) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:22, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:43; c) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:44 or nucleotides 19-111 of SEQ ID NO:1040; d) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24. e) CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-17 of SEQ ID NO: 25 or SEQ ID NO: 1041 or 1042, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 46 or SEQ ID NO: 104 f) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 22-85 of SEQ ID NO: 1 or 1042; f) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 26, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to nucleotides 24-138 of SEQ ID NO: 47 or SEQ ID NO: 1043;g) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 48 or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045; h) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 28 i) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:29, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:50; j) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:30. k) CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 31, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 31; l) CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 32. m) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54;n) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55; o) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56. p) a CRISPR repeat having at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 36, wherein the tracrRNA has at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 57; q) a CRISPR repeat having at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 37, wherein the tracrRNA has at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 58. r) a CRISPR repeat having at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 38 or 39, wherein the tracrRNA has at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 59 or 60; s) a CRISPR repeat having at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 40. a) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 61, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 41, wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 62;
[0053] In some embodiments of the above aspects, the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as a single guide RNA.
[0054] In some embodiments of the above aspects, the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.
[0055] In some embodiments of the above aspects, the vector further comprises a polynucleotide encoding an RGN polypeptide. In some embodiments, the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 1, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 21, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 42; b) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 2, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 22, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 43; c) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3, wherein the crRNA comprises a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:44 or nucleotides 19-111 of SEQ ID NO:1040; d) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4, wherein the crRNA comprises a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:45;e) An RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:5, wherein the crRNA comprises a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:25, or nucleotides 1-17 of SEQ ID NO:1041 or 1042, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:46, or nucleotides 22-85 of SEQ ID NO:1041 or 1042. f) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:6, wherein the crRNA comprises a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:26, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:47 or nucleotides 24 to 138 of SEQ ID NO:1043. g) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:7, wherein the crRNA comprises a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:27, or nucleotides 1-22 of SEQ ID NO:1044 or 1045, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:48, or nucleotides 27-96 of SEQ ID NO:1044, or nucleotides 27-95 of SEQ ID NO:1045. h) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:8, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:49;i) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:9, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:29, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:50; j) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:10. k) an RGN polypeptide comprising CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 30 and wherein the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 51; k) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 11 and wherein the crRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 31. l) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 12, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 32, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 53; m) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 13, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54;n) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 14, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55; o) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 15. a) an RGN polypeptide, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56; b) an RGN polypeptide, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 16, and the crRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 36. and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 57; q) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 17, wherein the crRNA comprises a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 37, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 58. r) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 18, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 38 or 39, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 59 or 60;s) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 19, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 40, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 61; and t) an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 20, wherein the crRNA comprises CRISPR repeats having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 41, and the tracrRNA has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 62.
[0056] In another aspect, the present disclosure provides a cell comprising the nucleic acid molecule, vector, single guide RNA, or dual guide RNA described above.
[0057] In some embodiments of the above aspects, the cell is a prokaryotic cell.
[0058] In some embodiments of the above aspects, the cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the human cell is an immune cell. In some embodiments, the immune cell is a stem cell. In some embodiments, the stem cell is an induced pluripotent stem cell. In some embodiments, the eukaryotic cell is an insect cell or an avian cell. In some embodiments, the eukaryotic cell is a fungal cell. In some embodiments, the eukaryotic cell is a plant cell.
[0059] In another aspect, the present disclosure provides a plant or seed comprising the plant cell described above.
[0060] In one aspect, the present disclosure provides a system for binding to a target sequence within a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand, the system comprising: a) one or more guide RNAs capable of hybridizing to the non-target strand of the target sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20, or a polynucleotide comprising a nucleotide sequence encoding the RGN polypeptide; wherein the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to bind the RGN polypeptide to the target sequence.
[0061] In some embodiments of the above aspects, at least one of the nucleotide sequence encoding one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to the nucleotide sequence.
[0062] In one aspect, the present disclosure provides a system for binding to a target sequence within a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand, the system comprising: a) one or more guide RNAs capable of hybridizing to the non-target strand of the target sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding one or more guide RNAs (gRNAs); and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20, wherein the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to bind the RGN polypeptide to the target sequence.
[0063] In some embodiments of the above aspects, at least one of the nucleotide sequences encoding the one or more guide RNAs is operably linked to a promoter that is heterologous to the nucleotide sequence.
[0064] In some embodiments of the above aspects, the RGN polypeptide comprises an amino acid sequence having at least 95% or 100% sequence identity to any one of SEQ ID NOs: 1-20.
[0065] In some embodiments of the above aspects, the RGN polypeptide and the one or more guide RNAs are not found complexed to each other in nature.
[0066] In some embodiments of the above aspects, the target sequence is a eukaryotic target sequence.
[0067] In some embodiments of the above aspects, the gRNA is a single guide RNA (sgRNA).
[0068] In some embodiments of the above aspects, the gRNA is a dual guide RNA.
[0069] In some embodiments of the above aspects, the gRNA comprises: a) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:42, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1; b) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:43, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:2; c) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:44 or nucleotides 19-111 of SEQ ID NO:1040, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3; d) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:45, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4;e) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-17 of SEQ ID NO: 25 or SEQ ID NO: 1041 or 1042 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 22-85 of SEQ ID NO: 46 or SEQ ID NO: 1041 or 1042, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 5. f) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 26 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 24-138 of SEQ ID NO: 47 or SEQ ID NO: 1043, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 6; g) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 48 or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 7 a) a gRNA comprising an amino acid sequence having at least 95% or 100% sequence identity; b) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:49, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:8;i) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:50, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:9; j) a CRISPR having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:30. a gRNA comprising a CRISPR repeat and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 51, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 10; k) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 31 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 52. 1) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 32 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 53, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 12; m) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 13;n) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 14; o) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35. a gRNA comprising a CRISPR repeat and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 15; p) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 36 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 57. q) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 37 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 58, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 17; r) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 38 or 39 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 59 or 60, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 18;s) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:40 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:61, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:19; and t) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:41 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:62, wherein the RGN polypeptide comprises an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:20;
[0070] In some embodiments of the above aspects, the target sequence is located adjacent to a protospacer adjacent motif (PAM). In some embodiments, the target sequence is intracellular.
[0071] In some embodiments of the above aspects, one or more guide RNAs can hybridize to a non-target strand of the target sequence, and the guide RNAs can form a complex with an RGN polypeptide to direct cleavage of the target nucleic acid molecule.
[0072] In some embodiments of the above aspects, the cleavage generates a double-stranded break.
[0073] In some embodiments of the above aspects, the cleavage generates a single-strand break.
[0074] In some embodiments of the above aspects, the RGN polypeptide is nuclease inactive or is a nickase.
[0075] In some embodiments of the above aspects, the RGN polypeptide is operably linked to a base-editing polypeptide. In some embodiments, the base-editing polypeptide is a deaminase. In some embodiments, the deaminase is a cytosine deaminase or an adenine deaminase. In some embodiments, the deaminase has at least 90%, at least 95%, or 100% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 481-552.
[0076] In some embodiments of the above aspects, the RGN polypeptide comprises one or more nuclear localization signals.
[0077] In some embodiments of the above aspects, the RGN polypeptide is codon-optimized for expression in eukaryotic cells.
[0078] In some embodiments of the above aspects, the nucleotide sequence encoding one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide are located on a single vector.
[0079] In some embodiments of the above aspects, the system further comprises one or more donor polynucleotides.
[0080] In another aspect, the present disclosure provides a cell comprising the above-described system.
[0081] In some embodiments of the above aspects, the cell is a prokaryotic cell.
[0082] In some embodiments of the above aspects, the cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the human cell is an immune cell. In some embodiments, the immune cell is a stem cell. In some embodiments, the stem cell is an induced pluripotent stem cell. In some embodiments, the eukaryotic cell is an insect cell or an avian cell. In some embodiments, the eukaryotic cell is a fungal cell. In some embodiments, the eukaryotic cell is a plant cell.
[0083] In another aspect, the present invention provides a plant or seed comprising the plant cell described above.
[0084] In another aspect, the present invention provides a pharmaceutical composition comprising a nucleic acid molecule, a vector, a cell, an RGN polypeptide, an RNP complex, or the above-mentioned system, and a pharmaceutically acceptable carrier.
[0085] In some embodiments of the above aspects, the pharmaceutically acceptable carrier is heterologous to the nucleic acid molecule, vector, cell, RGN polypeptide, or system.
[0086] In some embodiments of the above aspects, the pharmaceutically acceptable carrier is non-naturally occurring.
[0087] In some embodiments of the above aspects, the pharmaceutical composition is lipid-based. In some embodiments, the lipid-based pharmaceutical composition comprises a liposome or lipid nanoparticle (LNP). In some embodiments, the nucleic acid molecule, vector, cell, RGN polypeptide, RGN complex, or system is encapsulated and / or non-covalently or covalently linked to the liposome or LNP.
[0088] In another aspect, the present disclosure provides a method for binding a target sequence in a target nucleic acid molecule, comprising delivering a system described herein above to the target sequence or a cell containing the target sequence.
[0089] In some embodiments of the above aspects, the RGN polypeptide or guide RNA further comprises a detectable label, thereby allowing for detection of the target sequence.
[0090] In some embodiments of the above aspects, the guide RNA or RGN polypeptide further comprises an expression modulator, thereby regulating expression of a target gene comprising the target sequence.
[0091] In another aspect, the present disclosure provides a method for cleaving and / or modifying a target nucleic acid molecule comprising a target sequence, comprising delivering a system as described herein above to the target sequence or a cell comprising the target sequence, wherein cleavage or modification of the target nucleic acid molecule occurs.
[0092] In some embodiments of the above aspects, the modified target nucleic acid molecule comprises an insertion of heterologous DNA into the target DNA sequence.
[0093] In some embodiments of the above aspects, the modified target nucleic acid molecule comprises a deletion or mutation of at least one nucleotide from the target nucleic acid molecule.
[0094] In another aspect, the present disclosure provides a method for binding a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand, the method comprising: a) assembling an RNA-guided nuclease (RGN) ribonucleotide complex by combining: i) one or more guide RNAs capable of hybridizing to the non-target strand of the target sequence; and ii) an RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOS: 1-20; under conditions suitable for forming an RGN ribonucleotide complex; and b) contacting the target nucleic acid molecule or a cell comprising the target nucleic acid molecule with the assembled RGN ribonucleotide complex, wherein the one or more guide RNAs hybridize to the non-target strand of the target sequence, thereby directing the RGN polypeptide to bind to the target sequence.
[0095] In some embodiments of the above aspects, the method is performed in vitro, in vivo, or ex vivo.
[0096] In some embodiments of the above aspects, the RGN polypeptide or guide RNA further comprises a detectable label, thereby allowing for detection of the target sequence.
[0097] In some embodiments of the above aspects, the guide RNA or RGN polypeptide further comprises an expression modulator, thereby allowing for regulation of expression of a target gene comprising the target sequence.
[0098] In some embodiments of the above aspects, the RGN polypeptide further comprises a base-editing polypeptide, thereby enabling modification of the target nucleic acid molecule. In some embodiments, the base-editing polypeptide comprises a deaminase. In some embodiments, the deaminase is a cytosine deaminase or an adenine deaminase. In some embodiments, the deaminase has at least 90%, at least 95%, or 100% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 481-552.
[0099] In some embodiments of the above aspects, the RGN polypeptide is capable of cleaving the target nucleic acid molecule, thereby allowing for the cleavage and / or modification of the target nucleic acid molecule.
[0100] In another aspect, the present disclosure provides a method for cleaving and / or modifying a target nucleic acid molecule comprising a target sequence, wherein the target sequence comprises a target strand and a non-target strand, the method comprising contacting the target nucleic acid molecule with: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20; and b) one or more guide RNAs capable of targeting the RGN of (a) to the target sequence; wherein the one or more guide RNAs hybridize to the non-target strand of the target sequence, thereby directing the RGN polypeptide to bind to the target nucleic acid molecule and causing cleavage and / or modification of the target nucleic acid molecule.
[0101] In some embodiments of the above aspects, cleavage by the RGN polypeptide generates a double-stranded break.
[0102] In some embodiments of the above aspects, cleavage by the RGN polypeptide generates a single-strand break.
[0103] In some embodiments of the above aspects, the RGN polypeptide is a nuclease-inactive or nickase and is operably fused to a base-editing polypeptide. In some embodiments, the base-editing polypeptide is a deaminase. In some embodiments, the deaminase is a cytosine deaminase or an adenine deaminase. In some embodiments, the deaminase has at least 90%, at least 95%, or 100% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 481-552.
[0104] In some embodiments of the above aspects, the modified target nucleic acid molecule comprises an insertion of heterologous DNA into the target nucleic acid molecule.
[0105] In some embodiments of the above aspects, the modified target nucleic acid molecule comprises a deletion or mutation of at least one nucleotide from the target nucleic acid molecule.
[0106] In some embodiments of the above aspects, the target sequence is located adjacent to a protospacer adjacent motif (PAM).
[0107] In some embodiments of the above aspects, the target sequence is a eukaryotic target sequence.
[0108] In some embodiments of the above aspects, the gRNA is a single guide RNA (sgRNA).
[0109] In some embodiments of the above aspects, the gRNA is a dual guide RNA.
[0110] In some embodiments of the above aspects, the RGN comprises an amino acid sequence having at least 95% or 100% sequence identity to any one of SEQ ID NOs: 1-20.
[0111] In some embodiments of the above aspects, a) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:42; b) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:2, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:43; c) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; d) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:45; e) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:5, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042; f) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:6, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:26 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 24 to 138 of SEQ ID NO:47 or SEQ ID NO:1043; g) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 7, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045, and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 48 or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045; h) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:8, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:49; i) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:9, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:50; j) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 10, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 30 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 51; k) the RGN has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 31 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 52; l) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 12, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 32 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 53; m) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 13, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54; n) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 14, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55; o) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 15, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56; p) the RGN has at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 16, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 36 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 57; q) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 17, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 37 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 58; r) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 18, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 38 or 39 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 59 or 60; s) the RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 19, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 40 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 61; and t) The RGN has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:20, and the guide RNA comprises a crRNA repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:41 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:62.
[0112] In some embodiments of the above aspects, the target sequence is in a cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the human cell is an immune cell. In some embodiments, the immune cell is a stem cell. In some embodiments, the stem cell is an induced pluripotent stem cell. In some embodiments, the eukaryotic cell is an insect cell or an avian cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the eukaryotic cell is a fungal cell. In some embodiments, the eukaryotic cell is a plant cell.
[0113] In some embodiments of the above aspects, the method further comprises culturing cells under conditions in which the RGN polypeptide is expressed and cleaves and modifies the target nucleic acid molecule to produce a modified target nucleic acid molecule, and selecting for cells containing the modified target nucleic acid molecule.
[0114] In another aspect, the present disclosure provides a cell comprising a modified target nucleic acid molecule described hereinabove.
[0115] In some embodiments of the above aspects, the cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is a mammalian cell. In some embodiments, the mammalian cell is a human cell. In some embodiments, the human cell is an immune cell. In some embodiments, the immune cell is a stem cell. In some embodiments, the stem cell is an induced pluripotent stem cell. In some embodiments, the eukaryotic cell is an insect cell or an avian cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the eukaryotic cell is a fungal cell. In some embodiments, the eukaryotic cell is a plant cell.
[0116] In another aspect, the present disclosure provides a plant or seed comprising the plant cell described above.
[0117] In another aspect, the present disclosure provides a pharmaceutical composition comprising the cells described hereinabove and a pharmaceutically acceptable carrier.
[0118] In another aspect, the present disclosure provides a method for producing a genetically engineered cell in which a causative mutation of a genetic disease has been corrected, the method comprising: introducing into the cell: (a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20, and wherein a polynucleotide encoding the RGN polypeptide is operably linked to a promoter to enable expression of the RGN polypeptide in the cell; and (b) a guide RNA (gRNA) or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to enable expression of the gRNA in the cell, thereby causing the RGN and gRNA to target the genomic location of the causative mutation and modifying the genomic sequence to remove the causative mutation.
[0119] In some embodiments of the above aspects, the RGN is a nuclease-inactive or nickase and is fused to a polypeptide with base-editing activity. In some embodiments, the base-editing polypeptide is a deaminase. In some embodiments, the polypeptide with base-editing activity is a cytosine deaminase or an adenine deaminase. In some embodiments, the deaminase has at least 90%, at least 9%, or 100% sequence identity to the amino acid sequence of any one of SEQ ID NOs: 481-552.
[0120] In some embodiments of the above aspects, the genetic disease is caused by a single nucleotide polymorphism.
[0121] In some embodiments of the above aspects, the genetic disease is Hurler syndrome.
[0122] In some embodiments of the above aspects, the gRNA further comprises a spacer that targets the proximal region of the causative single nucleotide polymorphism.
[0123] In another aspect, the disclosure provides a method for generating a genetically engineered cell having a deletion in a disease-causing expanded trinucleotide repeat, comprising: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20, or a polynucleotide encoding the RGN polypeptide, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter to allow expression of the RGN polypeptide in the cell; and b) a first guide RNA (gRNA) or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA induces expression of the gRNA in the cell. and c) a polynucleotide encoding a second guide RNA (gRNA) or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to enable expression of the RGN in the cell, and further wherein the gRNA comprises a spacer that targets the 5' flank of the extended trinucleotide repeat; and b) a second guide RNA (gRNA) or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to enable expression of the gRNA in the cell, and further wherein the second gRNA comprises a spacer that targets the 3' flank of the extended trinucleotide repeat; are introduced into the cell, whereby the RGN and the two gRNAs target the extended trinucleotide repeat and at least a portion of the extended trinucleotide repeat is removed.
[0124] In some embodiments of the above aspects, the genetic disease is Friedrich's ataxia or Huntington's disease.
[0125] In some embodiments of the above aspects, the first gRNA further comprises a spacer that targets a region within or proximal to the extended trinucleotide repeat, hi some embodiments, the second gRNA further comprises a spacer that targets a region within or proximal to the extended trinucleotide repeat.
[0126] In some embodiments of the above aspects, the RGN polypeptide has at least 95% or 100% sequence identity to any one of SEQ ID NOs: 1-20.
[0127] In some embodiments of the above aspects, the gRNA, the first gRNA, the second gRNA, or the first gRNA and the second gRNA are selected from the group consisting of: a) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:42, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1; or b) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:43, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:2. c) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:44 or nucleotides 19-111 of SEQ ID NO:1040, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3; d) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:45, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4;e) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-17 of SEQ ID NO: 25 or SEQ ID NO: 1041 or 1042 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 22-85 of SEQ ID NO: 46 or SEQ ID NO: 1041 or 1042, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 5. f) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 26 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 24 to 138 of SEQ ID NO: 47 or SEQ ID NO: 1043, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 6; gRNA; g) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045, and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 48 or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 7. a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:49, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:8;i) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:50, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:9; j) a CRISPR having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:30. a gRNA comprising a CRISPR repeat and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 51, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 10; k) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 31 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 52. 1) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 32 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 53, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 12; m) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 13;n) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 14; o) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35. a gRNA comprising a CRISPR repeat and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 15; p) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 36 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 57. q) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 37 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 58, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 17; a gRNA having an amino acid sequence with at least 95% or 100% sequence identity; r) a gRNA comprising a CRISPR repeat with at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 38 or 39 and a tracrRNA with at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 58 or 60, wherein the RGN polypeptide has an amino acid sequence with at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 18;s) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:40 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:61, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:19; and t) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:41 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:62, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:20;
[0128] In some embodiments of the above aspects, the cell is an animal cell. In some embodiments, the animal cell is a mammalian cell. In some embodiments, the cell is derived from a dog, cat, mouse, rat, rabbit, horse, cow, pig, or human.
[0129] In another aspect, the disclosure provides a method for generating genetically modified mammalian hematopoietic progenitor cells with reduced BCL11A mRNA and protein expression, the method comprising introducing into isolated human hematopoietic progenitor cells: (a) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20, or a polynucleotide encoding the RGN polypeptide, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter to allow expression of the RGN polypeptide in the cell; and (b) a guide RNA (gRNA) or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to allow expression of the gRNA in the cell, whereby the RGN and gRNA are expressed in the cell and are cleaved at the BCL11A enhancer region, resulting in genetic modification of the human hematopoietic progenitor cells and reduced BCL11A mRNA and / or protein expression.
[0130] In some embodiments of the above aspects, the RGN polypeptide has at least 95% or 100% sequence identity to any one of SEQ ID NOs: 1-20.
[0131] In some embodiments of the above aspects, the gRNA is: a) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:42, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1; b) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:43, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:2; c) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:44 or nucleotides 19-111 of SEQ ID NO:1040, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:3; d) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:45, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:4;e) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-17 of SEQ ID NO: 25 or SEQ ID NO: 1041 or 1042 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 22-85 of SEQ ID NO: 46 or SEQ ID NO: 1041 or 1042, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 5. f) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 26 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to nucleotides 24 to 138 of SEQ ID NO: 47 or SEQ ID NO: 1043, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 6; gRNA; g) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045, and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 48 or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045, wherein the RGN polypeptide has at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 7. a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:49, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:8;i) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:50, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:9; j) a CRISPR having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:30. a gRNA comprising a CRISPR repeat and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 51, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 10; k) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 31 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 52. 1) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 32 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 53, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 12; m) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 54, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 13;n) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 55, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 14; o) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 35. a) a gRNA comprising a CRISPR repeat and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 56, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 15; b) a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 36 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 57. a) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 37 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 58, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 17; a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 38 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 59 or 60, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 18;s) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 39 or 40 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 61, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 19; and t) a gRNA comprising a CRISPR repeat having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 41 and a tracrRNA having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 62, wherein the RGN polypeptide has an amino acid sequence having at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 20;
[0132] In some embodiments of the above aspects, the gRNA further comprises a spacer that targets a region within or proximal to the BCL11A enhancer region.
[0133] In another aspect, the present disclosure provides a method of treating a disease, disorder, or condition, comprising administering to a subject in need thereof the pharmaceutical composition described above.
[0134] In some embodiments of the above aspects, the disease, disorder, or condition is associated with a causative mutation, and the pharmaceutical composition corrects the causative mutation.
[0135] In some embodiments of the above aspects, the subject is at risk of developing a disease, disorder, or condition.
[0136] In another aspect, the present disclosure provides the use of a nucleic acid molecule, a vector, a cell, an RGN polypeptide, an RNP complex, or a system as described herein above for treating a disease, disorder, or condition in a subject in need thereof.
[0137] In some embodiments of the above aspects, the disease, disorder, or condition is associated with a causative mutation, and treating comprises correcting the causative mutation.
[0138] In some embodiments of the above aspects, the subject is at risk of developing a disease, disorder, or condition.
[0139] In another aspect, the present disclosure provides the use of a nucleic acid molecule, a vector, a cell, an RGN polypeptide, an RNP complex, or a system described herein above for the manufacture of a medicament useful for treating a disease, disorder, or condition.
[0140] In some embodiments of the above aspects, the disease is associated with a causative mutation and the medicament corrects the causative mutation.
[0141] In another aspect, the present disclosure provides a single guide RNA comprising a nucleic acid molecule comprising a crRNA as described herein above and a nucleic acid molecule comprising a tracrRNA as described herein above.
[0142] In another aspect, the present disclosure provides a dual guide RNA comprising a nucleic acid molecule comprising a crRNA as described herein above and a nucleic acid molecule comprising a tracrRNA as described herein above. DETAILED DESCRIPTION OF THE INVENTION
[0143] Many modifications and other embodiments of the inventions described herein will come to mind to one skilled in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. It is to be understood, therefore, that the invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the accompanying embodiments. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0144] I. Overview RNA-guided nucleases (RGNs) enable targeted manipulation of specific sites within a genome and are useful in the context of gene targeting for therapeutic and research applications. In various organisms, including mammals, RNA-guided nucleases have been used for genome engineering, for example, by stimulating non-homologous end joining and homologous recombination. The compositions and methods described herein are useful for creating single- or double-strand breaks in polynucleotides, modifying polynucleotides, detecting specific sites within polynucleotides, or modifying the expression of specific genes.
[0145] The RNA-guided nucleases disclosed herein can alter gene expression by modifying target nucleic acid molecules containing the target sequence. In certain embodiments, the RNA-guided nuclease is directed to a target sequence (e.g., a target DNA sequence) by a guide RNA (gRNA) as part of a clustered regularly interspaced short palindromic repeats (CRISPR) RNA-guided nuclease system. RGNs are considered "RNA-guided" because the guide RNA forms a complex with the RNA-guided nuclease, directing the RNA-guided nuclease to bind to the target sequence and, in some embodiments, introducing a single- or double-stranded break in the target sequence (e.g., a target DNA sequence). After the target sequence is cleaved, the break can be repaired so that the sequence of the target nucleic acid molecule is modified during the repair process. Thus, provided herein are methods for modifying target nucleic acid molecules in host cells using RNA-guided nucleases. For example, RNA-guided nucleases can be used to modify target sequences in genomic loci of eukaryotic or prokaryotic cells.
[0146] II. RNA-guided nucleases RNA-guided nucleases are provided herein. The term RNA-guided nuclease (RGN) refers to a polypeptide that binds to a specific target sequence (e.g., a target DNA sequence) in a sequence-specific manner and is directed to the target sequence by a guide RNA molecule that forms a complex with the polypeptide and hybridizes with the target sequence. Although RNA-guided nucleases can cleave the target sequence upon binding, the term RNA-guided nuclease also encompasses nuclease-dead RNA-guided nucleases that can bind to but not cleave the target sequence. Cleavage of the target sequence by an RNA-guided nuclease can result in single-strand or double-strand cleavage. An RNA-guided nuclease that can cleave only one strand of a double-stranded target nucleic acid molecule is referred to herein as a nickase.
[0147] RNA-guided nucleases disclosed herein include LPG10165, LPG10166, LPG10167, LPG10168, LPG10169, LPG10171, LPG10186, LPG10190, LPG10191, LPG10194, LPG10195, LPG10196, LPG10197, LPG10198, LPG10200, LPG10203, LPG10204, LPG10205, LPG10207, and LPG10208 RNA-guided nucleases, whose amino acid sequences are set forth as SEQ ID NOs: 1-20, respectively, and active fragments or variants thereof that retain the ability to bind to target sequences in an RNA-guided sequence-specific manner. In some of these embodiments, the active fragment or variant of LPG10165, LPG10166, LPG10167, LPG10168, LPG10169, LPG10171, LPG10186, LPG10190, LPG10191, LPG10194, LPG10195, LPG10196, LPG10197, LPG10198, LPG10200, LPG10203, LPG10204, LPG10205, LPG10207, or LPG10208 RGN is capable of cleaving a single-stranded or double-stranded target sequence. In some embodiments, LPG10165, LPG10166, LPG10167, LPG10168, LPG10169, LPG10171, LPG10186, LPG10190, LPG10191, LPG10194, LPG10195, LPG10196, LPG10197, LPG10198, LPG10200, LPG10203, LPG10204, LPG10205, LPG10207, or LPG10208 Active variants of RGN include amino acid sequences that have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence set forth in any one of SEQ ID NOs: 1 to 20.
[0148] In some specific embodiments, an active variant of LPG10165 RGN comprises an amino acid sequence having at least 85% sequence identity with the amino acid sequence set forth in SEQ ID NO: 1, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10166 RGN comprises an amino acid sequence having at least 83% sequence identity with the amino acid sequence set forth in SEQ ID NO: 2, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10167 RGN comprises an amino acid sequence having at least 92% sequence identity with the amino acid sequence set forth in SEQ ID NO: 3, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10168 RGN comprises an amino acid sequence having at least 93% sequence identity with the amino acid sequence set forth in SEQ ID NO: 4, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10169 RGN comprises an amino acid sequence having at least 95% sequence identity with the amino acid sequence set forth in SEQ ID NO: 5, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10171 RGN comprises an amino acid sequence having at least 94% sequence identity with the amino acid sequence set forth in SEQ ID NO: 6, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10186 RGN comprises an amino acid sequence having at least 75% sequence identity with the amino acid sequence set forth in SEQ ID NO: 7, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10190 RGN comprises an amino acid sequence having at least 84% sequence identity with the amino acid sequence set forth in SEQ ID NO: 8, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10191 RGN comprises an amino acid sequence having at least 73% sequence identity with the amino acid sequence set forth in SEQ ID NO: 9, and retains RNA guide sequence-specific binding activity.In some embodiments, an active variant of LPG10194 RGN comprises an amino acid sequence having at least 98% sequence identity with the amino acid sequence set forth in SEQ ID NO: 10, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10195 RGN comprises an amino acid sequence having at least 88% sequence identity with the amino acid sequence set forth in SEQ ID NO: 11, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10196 RGN comprises an amino acid sequence having at least 75% sequence identity with the amino acid sequence set forth in SEQ ID NO: 12, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10197 RGN comprises an amino acid sequence having at least 92% sequence identity with the amino acid sequence set forth in SEQ ID NO: 13, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10198 RGN comprises an amino acid sequence having at least 97% sequence identity with the amino acid sequence set forth in SEQ ID NO: 14, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10200 RGN comprises an amino acid sequence having at least 98% sequence identity with the amino acid sequence set forth in SEQ ID NO: 15, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10203 RGN comprises an amino acid sequence having at least 78% sequence identity with the amino acid sequence set forth in SEQ ID NO: 16, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LEPG10204 RGN comprises an amino acid sequence having at least 81% sequence identity with the amino acid sequence set forth in SEQ ID NO: 17, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10205 RGN comprises an amino acid sequence having at least 81% sequence identity with the amino acid sequence set forth in SEQ ID NO: 18, and retains RNA guide sequence-specific binding activity.In some embodiments, an active variant of LPG10207 RGN comprises an amino acid sequence having at least 88% sequence identity with the amino acid sequence set forth in SEQ ID NO: 19, and retains RNA guide sequence-specific binding activity. In some embodiments, an active variant of LPG10208 RGN comprises an amino acid sequence having at least 81% sequence identity with the amino acid sequence set forth in SEQ ID NO: 20, and retains RNA guide sequence-specific binding activity.
[0149] In certain embodiments, LPG10165, LPG10166, LPG10167, LPG10168, LPG10169, LPG10171, LPG10186, LPG10190, LPG10191, LPG10194, LPG10195, LPG10196, LPG10197, LPG10198, LPG10200, LPG10203, LPG10204, LPG10205, LPG10207, or LPG10208 An active fragment of RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1100, 1200, 1300, or more consecutive amino acid residues of the amino acid sequence set forth as any one of SEQ ID NOS: 1-20. The RNA-guided nucleases provided herein may comprise at least one nuclease domain (e.g., DNase, RNase domain) and at least one RNA recognition and / or RNA-binding domain that interacts with the guide RNA. Additional domains that may be found in the RNA-guided nucleases provided herein include, but are not limited to, a DNA-binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain. In specific embodiments, the RNA-guided nucleases provided herein may comprise at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to one or more of the DNA-binding domain, the helicase domain, the protein-protein interaction domain, and the dimerization domain.
[0150] The target sequence is bound by the RNA-guided nuclease provided herein. If the target sequence is double-stranded (e.g., double-stranded DNA), the non-target strand of the target sequence hybridizes with a guide RNA associated with the RNA-guided nuclease. The target and / or non-target strands of the target sequence (e.g., a target DNA sequence) can then be cleaved by the RNA-guided nuclease if the polypeptide has nuclease activity. The term "cleaving" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond in the backbone of one or both strands of a double-stranded target sequence (e.g., a target DNA sequence), which can result in either a single-strand break or a double-strand break within the target sequence. The RGN of the present disclosure can cleave nucleotides within a polynucleotide, functioning as an endonuclease, or can be an exonuclease that removes consecutive nucleotides from the ends (5' and / or 3' ends) of a polynucleotide. In other embodiments, the disclosed RGNs can cleave nucleotides of a target polynucleotide within any position of the polynucleotide and thus function as both an endonuclease and an exonuclease. Cleavage of a target polynucleotide by the disclosed RGNs can result in staggered cuts or blunt ends.
[0151] The RNA-guided nuclease of the present disclosure can be a wild-type sequence derived from a bacterial or archaeal species. Alternatively, the RNA-guided nuclease can be a variant or fragment of a wild-type polypeptide. The wild-type RGN can be modified, for example, to alter nuclease activity or PAM specificity. In some embodiments, the RNA-guided nuclease does not exist in nature.
[0152] In certain embodiments, the RNA-guided nuclease functions as a nickase that cleaves only one strand of a double-stranded target sequence (e.g., a target DNA sequence). Such an RNA-guided nuclease has a single functional nuclease domain. In certain embodiments, the nickase can cleave either the target strand or the non-target strand of a double-stranded target sequence (e.g., a target DNA sequence). In some of these embodiments, the additional nuclease domain is mutated to reduce or eliminate nuclease activity. In embodiments in which a nickase is used, two nickases are required, each nickling a single strand within the double-stranded target sequence to achieve double-stranded cleavage of the double-stranded target sequence (e.g., a target DNA sequence).
[0153] In other embodiments, the RNA-guided nuclease lacks nuclease activity altogether and is referred to herein as nuclease-dead or nuclease-inactive. Any method known in the art for introducing mutations into amino acid sequences, such as PCR-mediated mutagenesis and site-directed mutagenesis, can be used to generate nickases or nuclease-dead RGNs. See, for example, U.S. Patent Application Publication No. 2014 / 0068797 and U.S. Patent No. 9,790,490, each of which is incorporated by reference in its entirety.
[0154] RNA-guided nucleases lacking nuclease activity can be used to deliver fusion polypeptides, polynucleotides, or small molecule payloads to specific genomic locations. In some of these embodiments, RGN polypeptides or guide RNAs can be fused to detectable labels to allow for the detection of specific sequences. As a non-limiting example, nuclease-dead RGN can be fused to detectable labels (e.g., fluorescent proteins) and targeted to specific disease-associated sequences to allow for the detection of disease-associated sequences.
[0155] Alternatively, nuclease-dead RGN can be targeted to a specific genomic location to alter the expression of a desired gene (i.e., a target gene). In some embodiments, binding of the nuclease-dead RNA-guided nuclease to the target sequence results in decreased expression of the target gene by interfering with the binding of RNA polymerase or transcription factors within the target genomic region. In other embodiments, the RGN (e.g., nuclease-dead RGN) or its complexed guide RNA further comprises an expression modulator that serves to repress or activate expression of the target gene upon binding to a target sequence within the target gene. In some of these embodiments, the expression modulator regulates expression of the target gene via an epigenetic mechanism.
[0156] In other embodiments, nuclease-dead RGN or RGN with nickase activity can be targeted to specific genomic locations to modify the sequence of a target polynucleotide by fusion to a base-editing polypeptide, such as a deaminase polypeptide or an active variant or fragment thereof, that directly chemically modifies (e.g., deaminates) a nucleobase, resulting in the conversion of one nucleobase to another. The base-editing polypeptide can be fused to the RGN at its N-terminus or C-terminus. Additionally, the base-editing polypeptide can be fused to the RGN via a peptide linker. Non-limiting examples of deaminase polypeptides useful in such compositions and methods include cytosine deaminase or adenine deaminase (e.g., Gaudelli et al. (2017) Nature 102:101-102). 551:464-471, US Publ., adenine deaminase base editors described in U.S. Patent Application Publication Nos. 2017 / 0121693 and 2018 / 0073012, and WO 2018 / 027078, or any of the deaminases disclosed in WO 2020 / 139783 and WO 2022 / 056254, and PCT / US2022 / 021271, filed March 22, 2022, each of which is incorporated by reference in its entirety. In one embodiment, a deaminase polypeptide useful in such compositions and methods is any of SEQ ID NOs: 481-552. In one embodiment, the deaminase polypeptide useful for such compositions and methods is a cytosine deaminase or adenine deaminase having an amino acid sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to any one of the amino acid sequences set forth as SEQ ID NOs: 481-552.In some embodiments, deaminase polypeptides useful in such compositions and methods of the present disclosure are deaminases disclosed in Table 17 of WO 2020 / 139783, which is incorporated herein by reference in its entirety. Furthermore, it is known in the art that certain fusion proteins between RGN and a base editing enzyme (e.g., cytosine deaminase) can also contain at least one uracil-stabilizing polypeptide that increases the rate at which the deaminase mutates cytidine, deoxycytidine, or cytosine in a nucleic acid molecule to thymidine, deoxythymidine, or thymine. Non-limiting examples of uracil-stabilizing polypeptides include those disclosed in WO 2021 / 217002, which is incorporated herein by reference in its entirety, and which contain USP2 (SEQ ID NO: 564) and a uracil glycosylase inhibitor (UGI) domain (SEQ ID NO: 565), which can enhance base editing efficiency. Thus, the fusion protein can comprise an RGN or a variant thereof described herein, a deaminase, and optionally at least one uracil-stabilizing polypeptide, such as UGI or USP2. In certain embodiments, the RGN fused to the base-editing polypeptide is a nickase (e.g., a deaminase) that cleaves DNA strands not acted on by the base-editing polypeptide.
[0157] An RNA-guided nuclease fused to a polypeptide or domain can be separated or linked by a linker. As used herein, the term "linker" refers to a chemical group or molecule that links two molecules or moieties, such as the binding domain and cleavage domain of a nuclease. In some embodiments, a linker connects the gRNA-binding domain of an RNA-guided nuclease to a base-editing polypeptide such as a deaminase. In some embodiments, a linker connects a nuclease-dead RGN to a deaminase. Typically, a linker is positioned between or adjacent to two groups, molecules, or other moieties and is connected to each via a covalent bond, thus connecting the two. In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.
[0158] The RNA-guided nuclease of the present disclosure may comprise at least one nuclear localization signal (NLS) to enhance transport of RGN to the nucleus of a cell. Nuclear localization signals are known in the art and generally comprise a stretch of basic amino acids (e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In some embodiments, RGN comprises two, three, four, five, six, or more nuclear localization signals. The nuclear localization signal(s) may be heterologous NLSs. Non-limiting examples of nuclear localization signals useful for RGN of the present disclosure are the nuclear localization signals of SV40 large T antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al. (2015) Bioconjug Chem 26(6):1004-7). In certain embodiments, RGN comprises the NLS sequence set forth as SEQ ID NO: 168 or 170. RGN may contain one or more NLS sequences at its N-terminus, C-terminus, or both the N-terminus and C-terminus, for example, RGN may contain two NLS sequences in the N-terminal region and four NLS sequences in the C-terminal region.
[0159] Other localization signal sequences known in the art that localize polypeptides to specific subcellular locations, such as, but not limited to, plastid-localization sequences, mitochondrial-localization sequences, and dual-targeting signal sequences that target both plastids and mitochondria, can also be used to target RGN (e.g., Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol dx.doi.org / 10.3389 / fphys.2015.00259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soll (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al. al.(2009)FEBS J 276:1187-1195;Silva-Filho(2003)Curr Opin Plant Biol 6:589-595;Peeters and Small(2001)Biochim Biophys Acta 1541:54-63;Murcha et al.(2014)J Exp Bot 65:6301-6335; Mackenzie (2005) Trends Cell Biol 15:548-554; Glaser et al. (1998) Plant Mol Biol 38:311-338).
[0160] In certain embodiments, the RNA-guided nuclease of the present disclosure comprises at least one cell-penetrating domain that facilitates cellular uptake of RGN. Cell-penetrating domains are known in the art and generally comprise a stretch of positively charged amino acid residues (i.e., a polycationic cell-penetrating domain), an alternating stretch of polar and non-polar amino acid residues (i.e., an amphipathic cell-penetrating domain), or a stretch of hydrophobic amino acid residues (i.e., a hydrophobic cell-penetrating domain) (see, for example, Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell-penetrating domain is the trans-activating transcription activator (TAT) from human immunodeficiency virus 1.
[0161] The nuclear localization signal, plastid localization signal, mitochondrial localization signal, dual-targeting localization signal, and / or cell-penetrating domain can be located at the amino-terminus (N-terminus), carboxyl-terminus (C-terminus), or at an internal position of the RNA-guided nuclease.
[0162] The RGN of the present disclosure can be fused directly or indirectly via a linker peptide to an effector domain, such as a cleavage domain, a deaminase domain, or an expression modulator domain. Such domains can be located at the N-terminus, C-terminus, or internal position of the RNA-guided nuclease. In some of these embodiments, the RGN component of the fusion protein is nuclease-dead RGN or a nickase.
[0163] In some embodiments, the RGN fusion protein comprises a cleavage domain, which can be any domain capable of cleaving polynucleotides (i.e., RNA, DNA, or RNA / DNA hybrids), including, but not limited to, restriction endonucleases and homing endonucleases, such as type IIS endonucleases (e.g., FokI) (e.g., Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388; Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993).
[0164] In other embodiments, the RGN fusion protein comprises a deaminase domain that deaminates a nucleobase, resulting in the conversion of one nucleobase to another, including, but not limited to, a cytosine deaminase or an adenine deaminase (see, e.g., Gaudelli et al. (2017) Nature 551:464-471; U.S. Patent Application Publication Nos. 2017 / 0121693 and 2018 / 0073012; and WO 2018 / 027078; or WO 2020 / 139783 and WO 2022 / 056254; and International Application No. PCT / US2022 / 021271, filed March 22, 2022, each of which is incorporated by reference in its entirety). In some embodiments, the effector domain of an RGN fusion protein can be an expression modulator domain, which is a domain that acts to either upregulate or downregulate transcription. The expression modulator domain can be an epigenetic modification domain, a transcriptional repressor domain, or a transcriptional activation domain.
[0165] In some of these embodiments, the expression modulator of the RGN fusion protein comprises an epigenetic modification domain that covalently modifies DNA or histone proteins to alter histone and / or chromosomal structure (i.e., upregulation or downregulation) without altering the DNA sequence, resulting in altered gene expression. Non-limiting examples of epigenetic modifications include acetylation or methylation of lysine residues, arginine methylation, serine and threonine phosphorylation, lysine ubiquitination and sumoylation of histone proteins, and methylation and hydroxymethylation of cytosine residues in DNA. Non-limiting examples of epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains, and DNA demethylase domains.
[0166] In other embodiments, the expression modulator of the fusion protein comprises a transcriptional repressor domain that interacts with transcriptional control elements and / or transcriptional regulatory proteins, such as RNA polymerases and transcription factors, to reduce or terminate transcription of at least one gene. Transcriptional repressor domains are known in the art and include, but are not limited to, Sp1-like repressor domains, IκB domains, and Kruppel-associated box (KRAB) domains.
[0167] In yet other embodiments, the expression modulator of the fusion protein comprises a transcriptional activation domain that interacts with transcriptional control elements and / or transcriptional regulatory proteins, such as RNA polymerases and transcription factors, to increase or activate transcription of at least one gene. Transcriptional activation domains are known in the art and include, but are not limited to, herpes simplex virus VP16 activation domain and NFAT activation domain.
[0168] The RGN polypeptide of the present disclosure can include a detectable label or purification tag. The detectable label or purification tag can be located at the N-terminus, C-terminus, or internal position of the RNA-guided nuclease, either directly or indirectly via a linker peptide. In some of these embodiments, the RGN component of the fusion protein is nuclease-dead RGN. In other embodiments, the RGN component of the fusion protein is RGN with nickase activity.
[0169] A detectable label is a molecule that can be visualized or otherwise observed. Detectable labels can be fused to RGN as a fusion protein (e.g., a fluorescent protein) or can be small molecules conjugated to an RGN polypeptide that can be detected visually or by other means. Detectable labels that can be fused to RGN of the present disclosure as a fusion protein include, but are not limited to, any detectable protein domain, including fluorescent proteins or protein domains that can be detected with specific antibodies. Non-limiting examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, EGFP, ZsGreen1) and yellow fluorescent proteins (e.g., YFP, EYFP, ZsYellow1). Non-limiting examples of small molecule detectable labels include: 3 H and 35 Examples of radioactive labels include S.
[0170] The RGN polypeptide can also comprise a purification tag, which is any molecule that can be used to isolate a protein or fusion protein from a mixture (e.g., a biological sample, culture medium). Non-limiting examples of purification tags include biotin, myc, maltose-binding protein (MBP), glutathione-S-transferase (GST), and 3X FLAG tags.
[0171] III. Guide RNA The present disclosure provides guide RNAs and polynucleotides encoding the same that target an associated RGN to a target sequence. The term "guide RNA" refers to a nucleotide sequence that hybridizes with a target sequence and has sufficient complementarity with the target nucleotide sequence to direct sequence-specific binding of an associated RNA-guided nuclease to the target nucleotide sequence. More specifically, when the target nucleotide sequence is double-stranded, such as in the case of DNA, the target nucleotide sequence is composed of a target strand (including a PAM sequence) and a non-target strand. In these embodiments, the guide RNA has sufficient complementarity with the non-target strand of a double-stranded target sequence (e.g., a target DNA sequence) such that the guide RNA hybridizes with the non-target strand and directs sequence-specific binding of the associated RNA-guided nuclease (RGN) to the target sequence (e.g., the target DNA sequence). Thus, in some embodiments, the guide RNA includes a spacer that is identical to the sequence of the target strand, except that uracil (U) replaces thymidine (T) in the guide RNA.
[0172] Each guide RNA of an RGN is one or more RNA molecules (typically one or two) that can bind to the RGN and guide it to bind to a specific target sequence, and in embodiments where the RGN has nickase or nuclease activity, also cleave the target strand and / or non-target strand. Generally, guide RNAs include CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA), although some RGNs do not require tracrRNA. Natural guide RNAs, including both crRNA and tracrRNA, generally comprise two separate RNA molecules that hybridize to each other via the repeat sequence of the crRNA and the anti-repeat sequence of the tracrRNA.
[0173] The present invention provides CRISPR RNA (crRNA) or a polynucleotide encoding the CRISPR RNA that, in conjunction with a tracrRNA, targets an associated RGN to a target sequence. The crRNA comprises a spacer and CRISPR repeats. The "spacer" has a nucleotide sequence that directly hybridizes with a non-target strand of a target sequence of interest (e.g., a target DNA sequence). The spacer is engineered to have full or partial complementarity with the non-target strand of the target sequence of interest. In some embodiments, the spacer can comprise from about 8 nucleotides to about 30 nucleotides or more. For example, the spacer can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the spacer is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides in length, hi some embodiments, the spacer is about 10 to about 26 nucleotides in length, or about 12 to about 30 nucleotides in length. In some embodiments, the degree of complementarity between the spacer and the non-target strand of the target sequence (e.g., the target DNA sequence), when optimally aligned using a suitable alignment algorithm, is 94% to 95% or more, including but not limited to, about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 50%, about 99%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 84%, about 85%, about 96%, about 97%, about 98%, about 99% or more. In some embodiments, the degree of complementarity between the spacer and the non-target strand of the target sequence (e.g., the target DNA sequence) is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more when optimally aligned using a suitable alignment algorithm.In some embodiments, the spacer may be identical in sequence to the target strand of the target sequence. In some of those embodiments in which the target sequence is a target DNA sequence, the spacer may be identical in sequence to the target strand of the target DNA sequence, except that thymidines (Ts) in the target strand are replaced by uracils (Us) in the spacer. In certain embodiments, the spacer does not contain secondary structure, which can be predicted using any suitable polynucleotide folding algorithm known in the art, including, but not limited to, mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).
[0174] The crRNA of the present disclosure comprises a spacer capable of targeting a bound RGN polypeptide to a target DNA sequence, wherein the target strand of the target DNA sequence has a nucleotide sequence set forth in any one of SEQ ID NOs: 344-464, 573-641, 667-677, 684-747, 770-817, 826-1039, and 1046-1057.
[0175] The crRNA further comprises a CRISPR RNA repeat, along with a spacer. The CRISPR RNA repeat comprises a nucleotide sequence that, by itself or in conjunction with a hybridized tracrRNA, forms a structure recognized by an RGN molecule. In various embodiments, the CRISPR RNA repeat can comprise from about 8 nucleotides to about 30 nucleotides or more. For example, the CRISPR repeat can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In certain embodiments, the CRISPR repeats are 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA sequence, when optimally aligned using a suitable alignment algorithm, is about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or more. In certain embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA sequence is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more when optimally aligned using a suitable alignment algorithm.
[0176] In certain embodiments, the CRISPR repeat comprises the nucleotide sequence of any one of SEQ ID NOs: 21-41, or nucleotides 1-17 of SEQ ID NOs: 1041 or 1042, or nucleotides 1-22 of SEQ ID NOs: 1044 or 1045, or an active variant or fragment thereof that, when contained within a guide RNA, is capable of directing sequence-specific binding of an associated RNA-guided nuclease provided herein to a target sequence of interest. In certain embodiments, an active CRISPR repeat variant of a wild-type sequence comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of the nucleotide sequences set forth as SEQ ID NOs:21-41, or nucleotides 1-17 of SEQ ID NO:1041 or 1042, or nucleotides 1-22 of SEQ ID NO:1044 or 1045. In certain embodiments, the active CRISPR repeat fragment of the wild-type sequence comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 or 22 contiguous nucleotides of any one of the nucleotide sequences set forth as SEQ ID NOs:21-41, or nucleotides 1-17 of SEQ ID NO:1041 or 1042, or nucleotides 1-22 of SEQ ID NO:1044 or 1045.
[0177] In certain embodiments, the crRNA does not occur in nature.In some of these embodiments, the specific CRISPR repeat is not linked to a naturally engineered spacer, and the CRISPR repeat is considered to be heterologous to the spacer.In certain embodiments, the spacer is an engineered sequence that does not occur in nature.
[0178] Currently disclosed guide RNAs include crRNA and trans-activating CRISPR RNA (tracrRNA). A tracrRNA molecule comprises a nucleotide sequence containing a region of sufficient complementarity to hybridize to the CRISPR repeats of the crRNA, referred to herein as the anti-repeat. In some embodiments, a tracrRNA molecule further comprises a region with secondary structure (e.g., a stem-loop) or forms a secondary structure upon hybridization with its corresponding crRNA. In certain embodiments, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeats is at the 5' end of the molecule, and the 3' end of the tracrRNA contains a secondary structure. This region of secondary structure generally contains several hairpin structures, including a nexus hairpin found adjacent to the anti-repeat. The nexus forms the core of interaction between the guide RNA and RGN and is at the intersection between the guide RNA, RGN, and target DNA. Nexus hairpins often have a conserved nucleotide sequence at the base of the hairpin stem, with the motif UNANNC (SEQ ID NO: 566) found in many nexus hairpins in tracrRNAs. In embodiments, the guide RNA or RGN systems of the present disclosure use tracrRNAs containing non-canonical sequences at the base of the hairpin stem of their nexus hairpins, including UNANNG (SEQ ID NO: 567), CNANNC (SEQ ID NO: 568), CNANNU (SEQ ID NO: 569), UNANNU (SEQ ID NO: 570), CNANNG (SEQ ID NO: 571), and CNCNNU (SEQ ID NO: 572). The 3' end of the tracrRNA often contains a terminal hairpin that can vary in structure and number, but often contains a GC-rich Rho-independent transcription terminator hairpin followed by a string of U's at the 3' end. See, e.g., Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Protoc; doi:10.1101 / pdb.top090902, and U.S. Patent Application Publication No. 2017 / 0275648, each of which is incorporated by reference in its entirety.
[0179] In various embodiments, the anti-repeat region of the tracrRNA, which is fully or partially complementary to a CRISPR repeat, comprises from about 8 nucleotides to about 30 nucleotides or more. For example, the region of base-pairing between the tracrRNA anti-repeat and the CRISPR repeat can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In certain embodiments, the base-paired region between the tracrRNA anti-repeat and the CRISPR repeat is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In some embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA anti-repeat is about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or more when optimally aligned using a suitable alignment algorithm. In certain embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA anti-repeat is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more when optimally aligned using a suitable alignment algorithm.
[0180] In various embodiments, the entire tracrRNA can comprise from about 60 nucleotides to more than about 210 nucleotides. For example, the tracrRNA can be about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, or more nucleotides in length. In certain embodiments, the tracrRNA is 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210 or more nucleotides in length. In certain embodiments, the tracrRNA is about 57 to about 115 nucleotides in length (about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85). , about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114 and about 115 nucleotides in length). In certain embodiments, the tracrRNA is 59 to 115 nucleotides in length, including 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, and 115 nucleotides in length.
[0181] In certain embodiments, the tracrRNA comprises any one of the nucleotide sequences of SEQ ID NOs: 42-62, nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NOs: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045, or an active variant or fragment thereof, which, when contained within a guide RNA, can direct sequence-specific binding of the associated RNA-guided nuclease provided herein to a target DNA sequence of interest. In certain embodiments, an active tracrRNA sequence variant comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to any one of the nucleotide sequences set forth as SEQ ID NOs:42-62, nucleotides 19-111 of SEQ ID NO:1040, nucleotides 22-85 of SEQ ID NOs:1041 or 1042, nucleotides 24-138 of SEQ ID NO:143, nucleotides 27-96 of SEQ ID NO:1044, or nucleotides 27-95 of SEQ ID NO:1045. In certain embodiments, an active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more contiguous nucleotides of any one of the nucleotide sequences set forth as SEQ ID NOs:42-62, nucleotides 19-111 of SEQ ID NO:1040, nucleotides 22-85 of SEQ ID NOs:1041 or 1042, nucleotides 24-138 of SEQ ID NO:143, nucleotides 27-96 of SEQ ID NO:1044, or nucleotides 27-95 of SEQ ID NO:1045.
[0182] Two polynucleotide sequences can be considered substantially complementary if the two sequences hybridize to each other under stringent conditions. Similarly, an RGN is considered to bind to a particular target sequence in a sequence-specific manner if a guide RNA bound to the RGN binds to the target sequence under stringent conditions. "Stringent conditions" or "stringent hybridization conditions" refer to conditions under which two polynucleotide sequences hybridize to each other to a detectably greater extent than other sequences (e.g., at least twice background). Stringent conditions are sequence-dependent and will vary in different circumstances. Typically, stringent conditions are conditions in which the salt concentration is less than about 1.5 M Na ions, typically about 0.01 to 1.0 M Na ion concentration (or other salt), at pH 7.0 to 8.3, and the temperature is at least about 30°C for short sequences (e.g., 10 to 50 nucleotides) and at least about 60°C for long sequences (e.g., more than 50 nucleotides). Stringent conditions can also be achieved by the addition of destabilizing agents such as formamide. Exemplary low stringency conditions include hybridization with a buffer solution of 30-35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate) at 37°C, followed by washing in 1X-2X SSC (20X SSC = 3.0 M NaCl / 0.3 M trisodium citrate) at 50-55°C. Exemplary medium stringency conditions include hybridization in 40-45% formamide, 1.0 M NaCl, and 1% SDS at 37°C, followed by washing in 0.5X-1X SSC at 55-60°C. Exemplary high stringency conditions include hybridization in 50% formamide, 1 M NaCl, and 1% SDS at 37°C, followed by washing in 0.1X SSC at 60-65°C. Optionally, the wash buffer may contain about 0.1% to about 1% SDS. The duration of hybridization is generally less than about 24 hours, usually about 4 to about 12 hours. The duration of the wash period is at least long enough to reach equilibrium.
[0183] Tm is the temperature (under defined ionic strength and pH) at which 50% of a complementary target sequence hybridizes to a perfectly matched sequence. For DNA-DNA hybrids, Tm can be approximated by the formula of Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5°C + 16.6 (log M) + 0.41 (% GC) - 0.61 (% GC) - 500 / L, where M is the molar concentration of monovalent cations, % GC is the percentage of guanosine and cytosine nucleotides in the DNA, % GC is the percentage of formamide in the hybridization solution, and L is the hybrid length in base pairs. Generally, stringent conditions are selected to be about 5°C lower than the thermal melting point (Tm) of the specific sequence and its complement at a defined ionic strength and pH. However, highly stringent conditions can utilize hybridization and / or washing at 1, 2, 3, or 4° C. below the thermal melting point (Tm), moderately stringent conditions can utilize hybridization and / or washing at 6, 7, 8, 9, or 10° C. below the thermal melting point (Tm), and low stringency conditions can utilize hybridization and / or washing at 11, 12, 13, 14, 15, or 20° C. below the thermal melting point (Tm). Using the formula, hybridization and washing composition, and desired Tm, one of skill in the art will understand that variations in stringency of hybridization and / or washing solutions are essentially described.Extensive guides to nucleic acid hybridization can be found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York); and Ausubel et al., eds. (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2d ed., Cold Spring Harbor Laboratory Press, Plainview, New York).
[0184] The term "sequence-specific" can also refer to binding of an RGN polypeptide to a target sequence at a frequency greater than binding to a randomized background sequence.
[0185] The guide RNA can be a single guide RNA (sgRNA) or a dual guide RNA. A single guide RNA comprises a crRNA and a tracrRNA on a single molecule of RNA, whereas a dual guide RNA comprises a crRNA and a tracrRNA present on two different RNA molecules hybridized to each other via at least a portion of the CRISPR repeats of the crRNA and at least a portion of the tracrRNA that can be fully or partially complementary to the CRISPR repeats of the crRNA (i.e., the anti-repeat). In some of these embodiments in which the guide RNA is a single guide RNA, the crRNA and the tracrRNA are separated by a linker nucleotide sequence. Generally, the linker nucleotide sequence does not contain complementary bases to avoid the formation of secondary structures within or involving the nucleotides of the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between the crRNA and the tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or more nucleotides in length. In certain embodiments, the linker nucleotide sequence of the single guide RNA is at least 4 nucleotides in length. In certain embodiments, the linker nucleotide sequence is the nucleotide sequence set forth as SEQ ID NO: 84.
[0186] Single-guide or dual-guide RNAs can be synthesized chemically or via in vitro transcription. Assays for determining sequence-specific binding between RGN and guide RNA are known in the art and include, but are not limited to, in vitro binding assays between expressed RGN and guide RNA, which can be tagged with a detectable label (e.g., biotin) and used in pull-down detection assays in which the guide RNA:RGN complex is captured via the detectable label (e.g., streptavidin beads). A control guide RNA with a sequence or structure unrelated to the guide RNA can be used as a negative control for non-specific binding of RGN to RNA. In certain embodiments, the guide RNA has a backbone sequence set forth in any one of SEQ ID NOs: 63-83, 1040, 1041, 1042, 1043, 1044, or 1045.
[0187] In certain embodiments, the guide RNA can be introduced into a target cell, organelle, or embryo as an RNA molecule. The guide RNA can be transcribed in vitro or chemically synthesized. In other embodiments, a nucleotide sequence encoding the guide RNA is introduced into a cell or embryo. In some of these embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., an RNA polymerase III promoter). The promoter can be heterologous to the native promoter or the nucleotide sequence encoding the guide RNA.
[0188] In various embodiments, the guide RNA can be introduced into a target cell, organelle, or embryo as a ribonucleoprotein complex as described herein, wherein the guide RNA is bound to an RNA-guided nuclease polypeptide.
[0189] The guide RNA directs the associated RGN to a specific target nucleotide sequence of interest through hybridization of the guide RNA to the target sequence of interest. The target sequence can be bound (and in some embodiments, cleaved) by an RNA-guided nuclease in vitro or in a cell. The target sequence is within a target polynucleotide and can comprise DNA, RNA, or a combination of both, and can be single-stranded or double-stranded. The target sequence can be genomic DNA (i.e., chromosomal DNA), plasmid DNA, or an RNA molecule (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). In embodiments where the target sequence is a chromosomal sequence, the chromosomal sequence can be a nuclear, plastid, or mitochondrial chromosomal sequence. In the compositions and methods of the present disclosure, the target sequence is within a target nucleic acid molecule (e.g., a target DNA sequence) that is double-stranded. In embodiments, the target sequence is unique in the target genome.
[0190] A target sequence is adjacent to a protospacer adjacent motif (PAM), and the target strand of the target sequence is the strand containing the PAM. The PAM is immediately adjacent to the target sequence and often includes Ns, which represent any nucleotide. In some embodiments, the protospacer adjacent motif includes about 1 to about 10 Ns, including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides. In certain embodiments, the PAM includes 1 to 10 Ns, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 Ns. The PAM can be 5' or 3' of the target sequence on its target strand. The PAM of the RGN of the present disclosure is immediately 3' flank of the target sequence on its target strand. Generally, the PAM is a consensus sequence of about 3-4 nucleotides, but in certain embodiments, it can be 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length. In various embodiments, the PAM sequence recognized by the RGN of the present disclosure comprises a consensus sequence set forth in any one of SEQ ID NOs: 127-147. In some embodiments of the above aspects, the crRNA is capable of binding to an RGN polypeptide capable of recognizing a complete protospacer adjacent motif (PAM) having a nucleotide sequence set forth in any one of SEQ ID NOs: 127-147.
[0191] In certain embodiments, an RNA-guided nuclease having any one of SEQ ID NOs: 1-20, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence set forth as any one of SEQ ID NOs: 127-147. In some embodiments, the RGN binds to a guide RNA comprising a CRISPR repeat having a nucleotide sequence set forth in any one of SEQ ID NOs: 21-41, or nucleotides 1-17 of SEQ ID NOs: 1041 or 1042, or nucleotides 1-22 of SEQ ID NOs: 1044 or 1045, or an active variant or fragment thereof, and a tracrRNA having a nucleotide sequence set forth in any one of SEQ ID NOs: 42-62, or nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NO: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045, or an active variant or fragment thereof. The RGN system is further described in Examples 1-3 and Tables 1 and 2 herein.
[0192] In some embodiments, an RNA-guided nuclease having SEQ ID NO:1 or an active variant or fragment thereof binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO:127 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO:21 or an active variant or fragment thereof and a tracrRNA sequence shown as SEQ ID NO:42 or an active variant or fragment thereof.
[0193] In some embodiments, an RNA-guided nuclease having SEQ ID NO:2 or an active variant or fragment thereof, when bound to a guide RNA comprising a CRISPR repeat sequence set forth as SEQ ID NO:22 or an active variant or fragment thereof and a tracrRNA sequence set forth as SEQ ID NO:43 or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence set forth as any one of SEQ ID NOs:128-131.
[0194] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 3 or an active variant or fragment thereof binds to a target nucleotide sequence adjacent to the PAM sequence set forth as SEQ ID NO: 132 when bound to a guide RNA comprising a CRISPR repeat sequence set forth as SEQ ID NO: 23 or an active variant or fragment thereof and a tracrRNA sequence set forth as SEQ ID NO: 44 or nucleotides 19-111 of SEQ ID NO: 1040 or an active variant or fragment thereof.
[0195] In some embodiments, an RNA-guided nuclease having SEQ ID NO:4, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO:132 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO:24, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO:45, or an active variant or fragment thereof.
[0196] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 5 or an active variant or fragment thereof binds to a target nucleotide sequence adjacent to the PAM sequence set forth as SEQ ID NO: 133 when bound to a guide RNA comprising the CRISPR repeat sequence set forth as SEQ ID NO: 25, or nucleotides 1-17 of SEQ ID NO: 1041 or 1042, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO: 46, or nucleotides 22-85 of SEQ ID NO: 1041 or 1042, or an active variant or fragment thereof.
[0197] In some embodiments, an RNA-guided nuclease having SEQ ID NO:6, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence set forth as SEQ ID NO:134 when bound to a guide RNA comprising the CRISPR repeat sequence set forth as SEQ ID NO:26, or an active variant or fragment thereof, and the tracrRNA sequence set forth as SEQ ID NO:47, or nucleotides 24-138 of SEQ ID NO:1043, or an active variant or fragment thereof.
[0198] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 7 or an active variant or fragment thereof binds to a target nucleotide sequence adjacent to the PAM sequence set forth as SEQ ID NO: 135 when bound to a guide RNA comprising the CRISPR repeat sequence set forth as SEQ ID NO: 27, or nucleotides 1-22 of SEQ ID NO: 1044 or 1045, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO: 48, or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045, or an active variant or fragment thereof.
[0199] In some embodiments, an RNA-guided nuclease having SEQ ID NO:8, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence set forth as SEQ ID NO:136 or 137 when bound to a guide RNA comprising a CRISPR repeat sequence set forth as SEQ ID NO:28, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO:49, or an active variant or fragment thereof.
[0200] In some embodiments, an RNA-guided nuclease having SEQ ID NO:9, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO:138 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO:29, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO:50, or an active variant or fragment thereof.
[0201] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 10, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO: 139 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 30, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO: 51, or an active variant or fragment thereof.
[0202] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 11, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO: 140 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 31, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO: 52, or an active variant or fragment thereof.
[0203] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 12, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO: 53 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 141, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO: 32, or an active variant or fragment thereof.
[0204] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 13, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO: 142 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 33, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO: 54, or an active variant or fragment thereof.
[0205] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 14, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO: 143 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 34, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO: 55, or an active variant or fragment thereof.
[0206] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 15, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence set forth as SEQ ID NO: 144 when bound to a guide RNA comprising a CRISPR repeat sequence set forth as SEQ ID NO: 35, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO: 56, or an active variant or fragment thereof.
[0207] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 16, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence set forth as SEQ ID NO: 145 when bound to a guide RNA comprising a CRISPR repeat sequence set forth as SEQ ID NO: 36, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO: 57, or an active variant or fragment thereof.
[0208] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 17, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO: 146 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 37, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO: 58, or an active variant or fragment thereof.
[0209] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 18 or an active variant or fragment thereof, when bound to a guide RNA comprising a CRISPR repeat sequence set forth as SEQ ID NO: 38, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO: 59, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence set forth as SEQ ID NO: 146. In some other embodiments, an RNA-guided nuclease having SEQ ID NO: 18 or an active variant or fragment thereof, when bound to a guide RNA comprising a CRISPR repeat sequence set forth as SEQ ID NO: 39, or an active variant or fragment thereof, and a tracrRNA sequence set forth as SEQ ID NO: 60, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence set forth as SEQ ID NO: 146.
[0210] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 19 or an active variant or fragment thereof binds to a target nucleotide sequence adjacent to the PAM sequence shown as SEQ ID NO: 132 when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 40 or an active variant or fragment thereof and a tracrRNA sequence shown as SEQ ID NO: 61 or an active variant or fragment thereof.
[0211] In some embodiments, an RNA-guided nuclease having SEQ ID NO: 20, or an active variant or fragment thereof, when bound to a guide RNA comprising a CRISPR repeat sequence shown as SEQ ID NO: 41, or an active variant or fragment thereof, and a tracrRNA sequence shown as SEQ ID NO: 62, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence shown as SEQ ID NO: 146 or 147. It is well known in the art that PAM sequence specificity for a given nuclease enzyme is influenced by the promoter used to express the RGN or the enzyme concentration, which can be modified by modifying the amount of ribonucleoprotein complex delivered to a cell or embryo (see, e.g., Karvelis et al. (2015) Genome Biol 16:253).
[0212] Upon recognizing its corresponding PAM sequence, RGN can cleave one or both strands of a target DNA sequence at a specific cleavage site. As used herein, a cleavage site is comprised of two specific nucleotides within a target DNA sequence where the strand of the target DNA locus is cleaved by RGN. The cleavage site may include the first and second, second and third, third and fourth, fourth and fifth, fifth and sixth, seventh and eighth, or eighth and ninth nucleotides from the PAM in either the 5' or 3' direction. In some embodiments, the cleavage site may be more than 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the PAM in either the 5' or 3' direction. Because RGN can cleave and shift the ends of a target DNA sequence, in some embodiments, the cleavage site is defined based on a distance of two nucleotides from the PAM on the target strand of the target DNA sequence and a distance of two nucleotides from the complement of the PAM on the non-target strand.
[0213] IV. Nucleotides encoding RNA-guided nucleases, CRISPR RNAs, and / or tracrRNAs The present disclosure provides polynucleotides comprising a CRISPR RNA, tracrRNA, and / or gRNA of the present disclosure, as well as polynucleotides comprising nucleotide sequences encoding an RNA-guided nuclease, CRISPR RNA, tracrRNA, and / or gRNA of the present disclosure. The presently disclosed polynucleotides include those that comprise or encode a crRNA comprising a CRISPR repeat sequence having any one of the nucleotide sequences set forth as SEQ ID NOs: 21-41, or nucleotides 1-17 of SEQ ID NOs: 1041 or 1042, or nucleotides 1-22 of SEQ ID NOs: 1044 or 1045, or an active variant or fragment thereof that, when contained within a guide RNA, is capable of directing sequence-specific binding of an associated RNA-guided nuclease to a target sequence of interest. Also disclosed are polynucleotides comprising or encoding tracrRNAs having any one of the nucleotide sequences set forth as SEQ ID NOs: 42-62, nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NO: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045, or active variants or fragments thereof that, when contained within a guide RNA, are capable of directing sequence-specific binding of an associated RNA-guided nuclease to a target sequence of interest. Also provided are polynucleotides encoding RNA-guided nucleases having any one of the amino acid sequences set forth as SEQ ID NOs: 1-20, and active fragments or variants thereof that retain the ability to bind to target sequences in an RNA-guided sequence-specific manner.
[0214] The use of the terms "polynucleotide" or "nucleic acid molecule" is not intended to limit the present disclosure to polynucleotides comprising DNA. Those skilled in the art will recognize that polynucleotides can comprise ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. These include peptide nucleic acids (PNAs), PNA-DNA chimeras, locked nucleic acids (LNAs), and phosphotylate-linked sequences. The polynucleotides disclosed herein also encompass all forms of sequences, including, but not limited to, single-stranded forms, double-stranded forms, DNA-RNA hybrids, triplex structures, stem-loop structures, etc.
[0215] In some embodiments, the polynucleotide encoding RGN of the present disclosure is an mRNA (messenger RNA) molecule. mRNA refers to any polynucleotide that encodes a polypeptide of interest and can be translated in vitro, in vivo, in situ, or ex vivo to produce the encoded polypeptide of interest. In embodiments, the basic components of an mRNA molecule include at least a coding region, a 5' UTR, a 3' UTR, a 5' cap, and a polyA tail. In embodiments, an mRNA encoding RGN useful in the methods and compositions of the present disclosure can include one or more structural and / or chemical modifications or changes that confer useful properties to the polynucleotide. For example, useful properties of mRNA include the lack of substantial induction of an innate immune response in cells into which the mRNA is introduced. A "structural" feature or modification is one in which two or more linked nucleotides are inserted, deleted, duplicated, inverted, or randomized in an mRNA without significant chemical modification to the nucleotides themselves. Because chemical bonds are necessarily broken and altered to result in structural modifications, structural modifications are chemical in nature and are therefore chemical modifications. However, structural modifications result in different nucleotide sequences. Chemical modifications to mRNA can include inclusion of 5-methylcytosine, N1-methyl-pseudouridine, pseudouridine, 2-thiouridine, 4-thiouridine, 5-methoxyuridine, 2'fluoroguanosine, 2'fluorouridine, 5-bromouridine, 5-(2-carbomethoxyvinyl)uridine, 5-[3(1-E-propenylamino)]uridine, α-thiocytidine, N6-methyladenosine, 5-methylcytidine, N4-acetylcytidine, 5-formylcytidine, or combinations thereof in the mRNA.
[0216] Nucleic acid molecules encoding RGN can be codon-optimized for expression in a desired organism. A "codon-optimized" coding sequence is a polynucleotide coding sequence with codon usage designed to mimic the preferred codon usage or transcription conditions of a particular host cell. Expression in a particular host cell or organism is enhanced as a result of changing one or more codons at the nucleic acid level, such that the translated amino acid sequence is unchanged. Nucleic acid molecules can be codon-optimized in whole or in part. Codon tables and other references providing preference information for a wide range of organisms are available in the art (e.g., Campbell and Gowri (1990) Plant Physiol. 92:1-11 for a discussion of plant-preferred codon usage). Methods for synthesizing plant-preferred genes or mammalian (e.g., human) codon-optimized coding sequences are available in the art. See, e.g., U.S. Patent Nos. 5,380,831 and 5,436,391, which are incorporated herein by reference, and Murray et al. (1989) Nucleic Acids Res. 17:477-498. Non-limiting examples of codon-optimized coding sequences for RGN of the present disclosure are set forth as SEQ ID NOs:148-167.
[0217] The polynucleotides encoding the RGN, crRNA, tracrRNA, and / or gRNA provided herein can be provided in an expression cassette for in vitro expression or expression in a cell, organelle, embryo, or organism of interest. The cassette includes 5' and 3' regulatory sequences operably linked to the polynucleotides encoding the RGN, crRNA, tracrRNA, and / or gRNA provided herein, allowing for expression of the polynucleotide. The cassette may further include at least one additional gene or genetic element that is co-transformed into the organism. When additional genes or elements are included, the components are operably linked. The term "operably linked" is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a promoter and a coding region of interest (e.g., a region encoding the RGN, crRNA, tracrRNA, and / or gRNA) is a functional linkage that allows for expression of the coding region of interest. Operably linked elements may be contiguous or non-contiguous. When used to refer to the linkage of two protein-coding regions, operably linked intends that the coding regions are in the same reading frame. Alternatively, the additional gene(s) or element(s) can be provided on multiple expression cassettes. For example, the nucleotide sequence encoding the presently disclosed RGN can be present on one expression cassette, while the nucleotide sequences encoding the crRNA, tracrRNA, or guide RNA can be present on separate expression cassettes. Such expression cassettes comprise multiple restriction and / or recombination sites for insertion of polynucleotides under the transcriptional control of the regulatory regions. The expression cassette may further comprise a selectable marker gene.
[0218] An expression cassette comprises, in the 5'-3' direction of transcription, a transcriptional (in some embodiments, translational) initiation region (i.e., promoter), a polynucleotide encoding an RGN-, crRNA-, tracrRNA-, and / or sgRNA- of the present invention, and a transcriptional (in some embodiments, translational) termination region (i.e., termination region) functional in the organism of interest. The promoter of the present invention can direct or drive expression of a coding sequence in a host cell. Regulatory regions (e.g., promoter, transcriptional regulatory region, and translational termination region) can be endogenous or heterologous to the host cell or to each other. As used herein, "heterologous" with respect to a sequence refers to a sequence derived from a foreign species, or, if derived from the same species, a sequence that has been substantially modified from its native form in composition and / or genomic locus by deliberate human intervention. As used herein, a chimeric gene comprises a coding sequence operably linked to a transcriptional initiation region that is heterologous to the coding sequence.
[0219] Convenient termination regions are available from the Ti-plasmid of A. tumefaciens, such as the octopine synthase and nopaline synthase termination regions. Guerineau et al.(1991)Mol.Gen.Genet.262:141-144;Proudfoot(1991)Cell 64:671-674;Sanfacon et al.(1991)Genes Dev.5:141-149;Mogen et al.(1990)Plant Cell 2:1261-1272;Munroe et al. (1990) Gene 91:151-158; Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903; and Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.
[0220] Additional regulatory signals include, but are not limited to, a start site for transcription initiation, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See, e.g., U.S. Patent Nos. 5,039,523 and 4,853,331; EPO 0480762A2; Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY), hereinafter "Sambrook 11"; Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and references cited therein.
[0221] In preparing expression cassettes, various DNA fragments can be manipulated to provide the DNA sequences in the proper orientation and, if necessary, in the proper reading frame. To this end, adapters or linkers can be used to join the DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. To this end, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, such as transitions and transversions, can be involved.
[0222] Numerous promoters can be used in the practice of the present invention. The promoter can be selected based on the desired results. The nucleic acid can be combined with a constitutive promoter, an inducible promoter, a developmental stage-specific promoter, a cell type-specific promoter, a tissue-preferred promoter, a tissue-specific promoter, or other promoter for expression in the organism of interest. For example, see WO 99 / 43838 and U.S. Patent Nos. 8,575,425; 7,790,846; 8,147,856; 8,586,832; 7,772,369; 7,534,939; 6,072,050; 5,659,026, which are incorporated herein by reference. See, for example, the promoters described in US Pat. Nos. 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; 5,608,142; and 6,177,611.
[0223] For expression in plants, constitutive promoters include the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812); rice actin (McElroy et al. (1990) Plant Cell 2:163-171); ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et al. (1992) Plant Mol. Biol. 18:675-689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588); and MAS (Velten et al. (1984) EMBO J. 3:2723-2730).
[0224] Examples of inducible promoters are the Adh1 promoter, which is inducible by hypoxia or cold stress, the Hsp70 promoter, which is inducible by heat stress, and the PPDK and peptocarboxylase promoters, both of which are inducible by light. Chemically inducible promoters, such as the safener-inducible In2-2 promoter (U.S. Pat. No. 5,364,780), the Axig1 promoter, which is auxin-inducible and tapetum-specific but also active in callus (PCT US01 / 22169), steroid-responsive promoters (e.g., the ERE promoter, which is estrogen-inducible, and the glucocorticoid-inducible promoters of Schena et al. (1991) Proc. Natl. Acad. Sci. USA 88:10421-10425 and McNellis et al. (1998) Plant J. 14(2):247-257), and tetracycline-inducible and tetracycline-repressible promoters (e.g., Gatz et al. al. (1991) Mol. Gen. Genet. 227:229-237, and U.S. Patent Nos. 5,814,618 and 5,789,156) are also useful and are incorporated herein by reference.
[0225] Tissue-specific or tissue-preferred promoters can be utilized to target expression of an expression construct in a specific tissue. In certain embodiments, tissue-specific or tissue-preferred promoters are active in plant tissues. Examples of promoters under developmental control in plants include promoters that preferentially initiate transcription in specific tissues (e.g., leaves, roots, fruits, seeds, or flowers). A "tissue-specific" promoter is a promoter that initiates transcription only in a specific tissue. Unlike constitutive expression of a gene, tissue-specific expression is the result of several interacting levels of gene regulation. Thus, promoters from homologous or closely related plant species may be preferred to achieve efficient and reliable expression of a transgene in a specific tissue. In some embodiments, expression involves a tissue-preferred promoter. A "preferred tissue" promoter is a promoter that preferentially initiates transcription, but not necessarily entirely or exclusively, in a specific tissue.
[0226] In some embodiments, nucleic acid molecules encoding RGN, crRNA, and / or tracrRNA comprise cell-type-specific promoters. A "cell-type-specific" promoter is a promoter that primarily drives expression in a particular cell type in one or more organs. Some examples of plant cells in which a functional cell-type-specific promoter may be primarily active in plants include, for example, BETL cells, roots, leaves, stalk cells, and vascular cells in stem cells. Nucleic acid molecules may also comprise cell-type preferred promoters. A "preferred cell-type" promoter is a promoter that primarily, but not necessarily exclusively, drives expression in a particular cell type in one or more organs. Some examples of plant cells in which a functional cell-type preferred promoter may be preferentially active in plants include, for example, BETL cells, roots, leaves, stalk cells, and vascular cells in stem cells.
[0227] The nucleic acid sequences encoding RGN, crRNA, tracrRNA, and / or gRNA can be operably linked to a promoter sequence recognized by a phage RNA polymerase, e.g., for in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA can be purified for use in the methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence or a variant of a T7, T3, or SP6 promoter sequence. In such embodiments, the expressed protein and / or RNA can be purified for use in the genome modification methods described herein.
[0228] In certain embodiments, the polynucleotides encoding RGN, crRNA, tracrRNA, and / or gRNA can also be linked to a polyadenylation signal (e.g., SV40 polyA signal and other signals functional in plants) and / or at least one transcription termination sequence. Additionally, the RGN-encoding sequence can also be linked to a sequence encoding at least one nuclear localization signal, at least one cell-penetrating domain, and / or at least one signal peptide capable of transporting the protein to a specific intracellular location, as described elsewhere herein.
[0229] The polynucleotides encoding RGN, crRNA, tracrRNA, and / or gRNA can be present in one vector or multiple vectors. A "vector" refers to a polynucleotide composition for transferring, delivering, or introducing a nucleic acid into a host cell. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, baculoviral vectors). Vectors may contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Further information can be found in "Current Protocols in Molecular Biology" Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual" Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.
[0230] The vector may also contain a selectable marker gene for selecting transformed cells. Selectable marker genes are used to select transformed cells or tissues. Marker genes include genes encoding antibiotic resistance, such as those encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), as well as genes that confer resistance to herbicidal compounds, such as glufosinate ammonium, bromoxynil, imidazolinone, and 2,4-dichlorophenoxyacetate (2,4-D).
[0231] In some embodiments, an expression cassette or vector containing a sequence encoding an RGN polypeptide can further include a sequence encoding a crRNA and / or a tracrRNA, or a sequence that combines the crRNA and tracrRNA to create a gRNA. The sequence encoding the crRNA and / or tracrRNA can be operably linked to at least one transcriptional control sequence for expression of the crRNA and / or tracrRNA in an organism or host cell of interest. For example, a polynucleotide encoding the crRNA and / or tracrRNA can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters and rice U6 and U3 promoters, such as the human U6 promoter set forth as SEQ ID NO: 173, and U.S. Patent Application No. 63 / 209,660, filed June 11, 2021, and PCT / US2022 / 032940, filed June 10, 2022 (each of which is incorporated by reference in its entirety, including as set forth herein as SEQ ID NOs: 553-562).
[0232] As shown, an expression construct containing a nucleotide sequence encoding an RGN, crRNA, tracrRNA, and / or gRNA can be used to transform an organism of interest. Methods for transformation include introducing a nucleotide construct into an organism of interest. By "introducing," we mean introducing a nucleotide construct into a host cell such that the construct gains access to the interior of the host cell. The methods of the present invention do not require a particular method for introducing a nucleotide construct into a host organism; they only require that the nucleotide construct gain access to the interior of at least one cell of the host organism. The host cell can be a eukaryotic or prokaryotic cell. In certain embodiments, the eukaryotic host cell is a plant cell, a mammalian cell, an avian cell, or an insect cell. In some embodiments, the eukaryotic cell containing or expressing an RGN of the present disclosure or modified by an RGN of the present disclosure is a human cell. In some embodiments, the eukaryotic cell comprising or expressing an RGN of the present disclosure or modified by an RGN of the present disclosure is a cell of hematopoietic origin, such as an immune cell (i.e., a cell of the innate or adaptive immune system), including, but not limited to, a B cell, a T cell, a natural killer (NK) cell, a pluripotent stem cell, an induced pluripotent stem cell, a chimeric antigen receptor T (CAR-T) cell, a monocyte, a macrophage, and a dendritic cell. In some embodiments, the eukaryotic cell comprising or expressing an RGN of the present disclosure or modified by an RGN of the present disclosure is an ocular cell, a muscle cell (e.g., a skeletal muscle cell), an epithelial cell (e.g., a lung epithelial cell), or a diseased cell (e.g., a tumor cell).
[0233] Methods for introducing nucleotide constructs into plant and other host cells are known in the art, including, but not limited to, stable transformation methods, transient transformation methods, and virus-mediated methods.
[0234] The methods result in transformed organisms, such as plants, including plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos, and their progeny. The plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, pollen).
[0235] "Transgenic organism" or "transformant" or "stably transformed" organism, cell, or tissue refers to an organism that incorporates or has incorporated a polynucleotide encoding the RGN, crRNA, and / or tracrRNA of the present invention. It is recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments can also be incorporated into host cells. Agrobacterium- and biolistics-mediated transformation remain the two primary techniques used for plant cell transformation. However, host cell transformation can also be achieved by infection, transfection, microinjection, electroporation, microprojection, biolistics or particle bombardment, electroporation, silica / carbon fiber, ultrasound-mediated, PEG-mediated, calcium phosphate co-precipitation, polycation DMSO techniques, DEAE-dextran procedures, and viral, liposome-mediated, etc. Viral-mediated introduction of polynucleotides encoding RGN, crRNA, and / or tracrRNA includes retrovirus-, lentivirus-, adenovirus-, and adeno-associated virus-mediated introduction and expression, as well as the use of caulimoviruses, geminiviruses, and RNA plant viruses.
[0236] Transformation protocols and protocols for introducing polypeptide or polynucleotide sequences into plants can vary depending on the type of host cell (e.g., monocotyledonous or dicotyledonous plant cell) targeted for transformation. Methods for transformation are known in the art, including those described in U.S. Patent Nos. 8,575,425; 7,692,068; 8,802,934; and 7,541,517, each of which is incorporated herein by reference. Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett.7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4:1-12;Bates, GW (1999) Methods in Molecular Biology 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606; Christou, P. (1992) The Plant Journal 2:275-281; Christou, P. (1995) Euphytica 85:13-27;Tzfira et al. (2004) TRENDS in Genetics 20:375-383; Yao et al. (2006) Journal of Experimental Botany 57:3737-3746; Zupan and Zambryski (1995) Plant Physiology 107:1041-1047; Jones et al. (2005) Plant Methods 1:5 Transformation can result in stable or transient integration of a nucleic acid into a cell. "Stable transformation" is intended to mean that a nucleotide construct introduced into a host cell is integrated into the genome of the host cell and can be inherited by its progeny. "Transient transformation" is intended to mean that a polynucleotide is introduced into a host cell and does not integrate into the genome of the host cell.
[0237] Methods for chloroplast transformation are known in the art. See, e.g., Svab et al. (1990) Proc. Natl. Acad. Sci. USA 87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J. 12:601-606. Methods rely on particle gun delivery of DNA containing a selectable marker and targeting the DNA to the plastid genome by homologous recombination. Additionally, plastid transformation can be achieved by transactivation of a silent plastid-derived transgene through tissue-preferential expression of a nuclear-encoded and plastid-directed RNA polymerase. Such a system is reported in McBride et al. (1994) Proc. Natl. Acad. Sci. USA 91:7301-7305.
[0238] Transformed cells can be grown into transgenic organisms, such as plants, according to conventional methods. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants can then be grown and pollinated with the same transformed strain or a different strain, and the resulting hybrids identified for constitutive expression of the desired phenotypic trait. Two or more generations can be grown to ensure that the expression of the desired phenotypic trait is stably maintained and inherited, and seeds can then be harvested to ensure that expression of the desired phenotypic trait is achieved. In this way, the present invention provides "transgenic seeds," which are transformed seeds carrying a nucleotide construct of the present invention, e.g., an expression cassette of the present invention, stably integrated into their genomes.
[0239] Alternatively, transformed cells can be introduced into an organism. These cells can be derived from an organism, and the cells transformed by an ex vivo approach.
[0240] The sequences provided herein can be used to transform any plant species, including, but not limited to, monocotyledons and dicotyledons. Examples of plants of interest include, but are not limited to, corn (maize), sorghum, wheat, sunflower, tomato, rapeseed, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley and rapeseed, Brassica sp., alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus trees, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, oat, vegetables, ornamentals, and conifers.
[0241] Vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and members of the genus Curcumis, such as cucumbers, melons, and cantaloupes. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. In certain embodiments, the plants of the present invention are crop plants (e.g., corn, sorghum, wheat, sunflowers, tomatoes, rapeseed, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).
[0242] As used herein, the term plant includes plant cells, plant protoplasts, plant cell tissue cultures capable of regenerating plants, plant callus, plant mass, and intact plant cells of plants or plant parts, such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, grains, ears, cobs, husks, stems, roots, root tips, anthers, etc. Grain is intended to mean mature seeds produced by commercial growers for purposes other than seed growth or propagation. Progeny, variants, and mutants of regenerated plants are also included within the scope of the present invention, so long as these parts contain the introduced polynucleotide. Further provided are processed plant products or by-products, including, for example, soybean meal, that carry the sequences disclosed herein.
[0243] Polynucleotides encoding RGN, crRNA and / or tracrRNA or comprising crRNA and / or tracrRNA can also be used to transform any prokaryotic species, including, but not limited to, archaea and bacteria (e.g., Bacillus sp., Klebsiella sp., Streptomyces sp., Rhizobium sp., Escherichia sp., Pseudomonas sp., Salmonella sp., Shigella sp., Vibrio sp., Yersinia sp., Mycoplasma sp., Agrobacterium, Lactobacillus sp.).
[0244] Polynucleotides encoding RGN, crRNA, and / or tracrRNA or containing crRNA and / or tracrRNA can be used to transform any eukaryotic species, including, but not limited to, animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoeba, algae, and yeast.
[0245] Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into mammalian, insect, or avian cells or target tissues. Such methods can be used to administer nucleic acids encoding components of the RGN system to cells in culture or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acid, and nucleic acid complexed with a delivery vehicle such as a liposome. Viral vector delivery systems include DNA and RNA viruses that have either episomal or integrated genomes after delivery to cells. For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and See Immunology, Doerfler and Bohm (eds) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).
[0246] Non-viral delivery methods of nucleic acids include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid:nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386; 4,946,787; and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those of Feigner, WO 91 / 17424; WO 91 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vivo administration). The preparation of lipid:nucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known to those of skill in the art (e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); see U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).
[0247] The use of RNA or DNA virus-based systems for nucleic acid delivery takes advantage of highly evolved processes for targeting viruses to specific cells in the body and transporting the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or used to treat cells in vitro, and the modified cells can then be optionally administered to patients (ex vivo). Traditional virus-based systems can include retroviral, lentiviral, adenoviral, adeno-associated, and herpes simplex virus vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated virus gene transfer methods, often resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiencies have been observed in many different cell types and target tissues.
[0248] The tropism of retroviruses can be altered by incorporating foreign envelope proteins to expand the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system will depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeat sequences that have packaging capacity for foreign sequences up to 6-10 kb. The minimal cis-acting LTRs are sufficient for vector replication and packaging, which are then used to integrate therapeutic genes into target cells to provide permanent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (e.g., Buchscher et al., J. Viral. 66:2731-2739 (1992); Johann et al., J. Viral. 66:1635-1640 (1992); Sommnerfelt et al., J. Viral. 176:58-59 (1990); Wilson et al., J. Viral. 63:2374-2378 (1989); Miller et al., J. Viral. 65:2220-2224 (1991); PCT / US94 / 05700).
[0249] For applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors are capable of very high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been achieved with such vectors. This vector can be produced in large quantities using a relatively simple system. Adeno-associated virus ("AAV") vectors can also be used to transduce target nucleic acids into cells, for example, in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures (e.g., West et al., Virology 160:38-47 (1987); US Pat. No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994)). The construction of recombinant AAV vectors has been described in several publications, including U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., I. Viral. 63:03822-3828 (1989). Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells, which package adenovirus, and ψJ2 or PA317 cells, which package retrovirus.
[0250] Viral vectors used in gene therapy are usually produced by engineering cell lines that package nucleic acid vectors into viral particles. The vectors typically contain minimal viral sequences necessary for packaging and subsequent integration into the host, with other viral sequences replaced by an expression cassette for the polynucleotide to be expressed. Missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome necessary for packaging and integration into the host genome. Viral DNA is packaged in a cell line containing a helper plasmid encoding other AAV genes, namely rep and cap, but lacking the ITR sequences.
[0251] The cell line can also be infected with adenovirus as a helper. The helper virus promotes the replication of AAV vectors and the expression of AAV genes from the helper plasmid. The helper plasmid is not packaged in significant amounts due to the lack of ITR sequences. Contamination by adenovirus can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV. Additional methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, U.S. Patent Application Publication No. 20030087817, which is incorporated herein by reference.
[0252] In some embodiments, host cells are transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, cells are transfected as they naturally occur in a subject. In some embodiments, the transfected cells are harvested from a subject. In some embodiments, the cells are derived from cells harvested from a subject, such as a cell line. In some embodiments, the cell line can be a mammalian cell, an insect cell, or an avian cell. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFl, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182, A37 5, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, TIB 55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4.COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2 780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MA C6, MTD-lA, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, O Examples of suitable cell lines include, but are not limited to, PCN / OPCT cell lines, Peer, PNT-1A / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic varieties thereof. Cell lines are available from a variety of sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC) (Manassas, Va.)).
[0253] In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines containing one or more vector-derived sequences. In some embodiments, cells transiently transfected with components of the RGN system described herein (e.g., by transient transfection of one or more vectors or transfection with RNA) and modified via activity of the RGN system are used to establish new cell lines comprising cells containing the modifications but lacking other exogenous sequences. In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used in evaluating one or more test compounds.
[0254] In some embodiments, one or more vectors described herein are used to generate non-human transgenic animals or transgenic plants. In some embodiments, the transgenic animal is a mammal, such as a mouse, rat, hamster, rabbit, cow, or pig. In some embodiments, the transgenic animal is a bird, such as a chicken or duck. In some embodiments, the transgenic animal is an insect, such as a mosquito or tick.
[0255] V. Polypeptide and Polynucleotide Variants and Fragments The present disclosure provides active variants and fragments of naturally occurring (i.e., wild-type) RNA-guided nucleases whose amino acid sequences are set forth as any one of SEQ ID NOs: 1-20, as well as active variants and fragments of naturally occurring CRISPR repeats, such as the sequences set forth as SEQ ID NOs: 21-41, or any one of nucleotides 1-17 of SEQ ID NOs: 1041 or 1042, or nucleotides 1-22 of SEQ ID NOs: 1044 or 1045, and active variants and fragments of naturally occurring tracrRNAs, such as any one of the sequences set forth as SEQ ID NOs: 42-62, nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NO: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045, and polynucleotides encoding the same.
[0256] The activity of the variants or fragments may be altered compared to the polynucleotide or polypeptide of interest, but the variants and fragments should retain the functionality of the polynucleotide or polypeptide of interest. For example, the variants or fragments may have increased activity, decreased activity, a different spectrum of activity, or any other change in activity when compared to the polynucleotide or polypeptide of interest.
[0257] Fragments and variants of naturally occurring RGN polypeptides, such as those disclosed herein, retain sequence-specific RNA-guided DNA binding activity, and in certain embodiments, fragments and variants of naturally occurring RGN polypeptides, such as those disclosed herein, retain nuclease activity (single- or double-stranded).
[0258] Fragments and variants of naturally occurring CRISPR repeats, such as those disclosed herein, will retain the ability of a portion of the guide RNA (including the tracrRNA) to bind to and guide an RNA-guided nuclease (complexed with the guide RNA) to a target sequence (e.g., a target DNA sequence) in a sequence-specific manner.
[0259] Naturally occurring fragments and variants of tracrRNA, such as those disclosed herein, when part of a guide RNA (including a CRISPR RNA), will retain the ability to guide an RNA-guided nuclease (complexed with the guide RNA) to a target sequence (e.g., a target DNA sequence) in a sequence-specific manner.
[0260] The term "fragment" refers to a portion of a polynucleotide or polypeptide sequence of the present invention. A "fragment" or "biologically active portion" includes a polynucleotide comprising a sufficient number of consecutive nucleotides to retain biological activity (i.e., to bind to RGN in a sequence-specific manner and target RGN to a target nucleotide sequence when contained within a guide RNA). A "fragment" or "biologically active portion" includes a polypeptide comprising a sufficient number of consecutive amino acid residues to retain biological activity (i.e., to bind to a target sequence in a sequence-specific manner when complexed with a guide RNA). Fragments of RGN protein include fragments that are shorter than the full-length sequence due to the use of alternative downstream start sites. A biologically active portion of an RGN protein can be, for example, a polypeptide comprising 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, or more consecutive amino acid residues of any one of SEQ ID NOS: 1-20. Such biologically active portions can be prepared by recombinant techniques and evaluated for sequence-specific RNA-guided DNA binding activity. A biologically active fragment of a CRISPR repeat sequence can comprise at least 8 contiguous amino acids of any one of SEQ ID NOs: 21-41, or nucleotides 1-14 of SEQ ID NO: 1040, nucleotides 1-17 of SEQ ID NOs: 1041 or 1042, nucleotides 1-19 of SEQ ID NO: 1043, or nucleotides 1-22 of SEQ ID NOs: 1044 or 1045. A biologically active portion of a CRISPR repeat sequence can be, for example, a polynucleotide comprising 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 contiguous nucleotides of any one of SEQ ID NOs: 21-41, or nucleotides 1-17 of SEQ ID NO: 1041 or 1042, or nucleotides 1-22 of SEQ ID NO: 1044 or 1045.A biologically active portion of a tracrRNA can be a polynucleotide comprising, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more contiguous nucleotides of any one of SEQ ID NOs: 42-62, nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NOs: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045.
[0261] In general, "variant" is intended to mean a substantially similar sequence. In the case of polynucleotides, variants include deletions and / or additions of one or more nucleotides at one or more internal sites within the naturally occurring polynucleotide and / or substitutions of one or more nucleotides at one or more sites within the naturally occurring polynucleotide. As used herein, "native" or "wild-type" polynucleotides or polypeptides include naturally occurring nucleotide sequences or amino acid sequences, respectively. In the case of polynucleotides, conservative variants include sequences that, due to the degeneracy of the genetic code, encode the naturally occurring amino acid sequence of a gene of interest. Naturally occurring allelic variants such as these can be identified using well-known molecular biology techniques, e.g., polymerase chain reaction (PCR) and hybridization techniques outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those generated, e.g., by using site-directed mutagenesis, but still encoding a polypeptide or polynucleotide of interest. Generally, a variant of a particular polynucleotide disclosed herein will have at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the particular polynucleotide as determined by sequence alignment programs and parameters described elsewhere herein.
[0262] Variants of a particular polynucleotide (i.e., a reference polynucleotide) disclosed herein can also be evaluated by comparing the percent sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The percent sequence identity between any two polypeptides can be calculated using sequence alignment programs and parameters described elsewhere herein. When any given pair of polynucleotides disclosed herein is evaluated by comparing the percent sequence identity shared by the two polypeptides they encode, the percent sequence identity between the two encoded polypeptides will be at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity.
[0263] In certain embodiments, a polynucleotide of the present disclosure encodes an RNA-guided nuclease polypeptide comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to any one of the amino acid sequences set forth as SEQ ID NOs: 1-20.
[0264] Biologically active variants of the RGN polypeptides of the present invention may differ by as few as about 1 to 15 amino acid residues, as few as about 1 to 10, e.g., as few as about 6 to 10, as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 amino acid residue. In certain embodiments, the polypeptide may comprise an N-terminal or C-terminal truncation, which may include a deletion of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350 or more amino acids from either the N-terminus or C-terminus of the polypeptide.
[0265] In certain embodiments, the polynucleotides disclosed herein comprise or encode a CRISPR repeat comprising a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to any one of the nucleotide sequences set forth as SEQ ID NOs:21-41, or nucleotides 1-17 of SEQ ID NO:1041 or 1042, or nucleotides 1-22 of SEQ ID NO:1044 or 1045.
[0266] A polynucleotide of the present disclosure can comprise or encode a tracrRNA comprising a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 90%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 88%, 98%, 89%, 99% or more identity to any one of the nucleotide sequences set forth as SEQ ID NOs:42-62, nucleotides 19-111 of SEQ ID NO:1040, nucleotides 22-85 of SEQ ID NO:1041 or 1042, nucleotides 24-138 of SEQ ID NO:143, nucleotides 27-96 of SEQ ID NO:1044, or nucleotides 27-95 of SEQ ID NO:1045.
[0267] Biologically active variants of CRISPR repeats or tracrRNA of the invention may differ by as few as about 1-15 nucleotides, as few as about 1-10 nucleotides, e.g., about 6-10 nucleotides, as few as 5 nucleotides, as few as 4 nucleotides, as few as 3 nucleotides, as few as 2 nucleotides, or as few as 1 nucleotide. In certain embodiments, the polynucleotide may comprise a truncation at the 5' or 3' end, which may include a deletion of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 95, 100, 105, 110 or more nucleotides from either the 5' or 3' end.
[0268] It is recognized that modifications can be made to the RGN polypeptides, CRISPR repeats, and tracrRNA provided herein to generate variant proteins and polynucleotides. Human-designed changes can be introduced by applying site-directed mutagenesis techniques. Alternatively, naturally occurring, unknown, or as yet unidentified polynucleotides and / or polypeptides structurally and / or functionally related to the sequences disclosed herein can also be identified as falling within the scope of the present invention. Conservative amino acid substitutions can be made in non-conserved regions that do not alter the function of the RGN protein. Alternatively, modifications can be made that improve the activity of RGN.
[0269] Variant polynucleotides and proteins also encompass sequences and proteins derived from mutagenesis and recombination procedures, such as DNA shuffling. Using such procedures, one or more of the different RGN proteins disclosed herein (e.g., SEQ ID NOS: 1-20) can be manipulated to generate new RGN proteins with desired properties. In this manner, libraries of recombinant polynucleotides are generated from a population of related sequence polynucleotides that share substantial sequence identity and contain sequence regions that can be homologously recombined in vitro or in vivo. For example, using this approach, sequence motifs encoding domains of interest can be shuffled between the RGN sequences provided herein and other known RGN genes to generate improved properties of interest, such as, in the case of enzymes, increased K mIt is possible to obtain new genes encoding proteins with the desired structure. Strategies for such DNA shuffling are known in the art. See, for example, Stemmer (1994) Proc. Natl. Acad. Sci. USA 91:10747-10751; Stemmer (1994) Nature 370:389-391; Crameri et al. (1997) Nature Biotech. 15:436-438; Moore et al. (1997) J. Mol. Biol. 272:336-347; Zhang et al. (1997) Proc. Natl. Acad. Sci. USA 94:4504-4509; Crameri et al. (1998) Nature 391:288-291; and U.S. Patent Nos. 5,605,793 and 5,837,458. "Shuffled" nucleic acids are nucleic acids produced by a shuffling procedure, such as any of the shuffling procedures described herein. Shuffled nucleic acids are generated by recombining (physically or virtually) two or more nucleic acids (or character strings), e.g., artificially, and optionally recursively. Generally, the shuffling process to identify nucleic acids of interest employs one or more screening steps, which can occur before or after any recombination steps. In some (but not all) shuffling embodiments, it is desirable to perform multiple rounds of recombination before selection to increase the diversity of the pool being screened. The overall process of recombination and selection is optionally repeated recursively. Depending on the context, shuffling can refer to the overall process of recombination and selection, or it can simply refer to the recombination portion of the overall process.
[0270] As used herein, "sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences refers to residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When using percentage sequence identity with respect to proteins, it is recognized that non-identical residue positions often differ by conservative amino acid substitutions, in which an amino acid residue is substituted with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the functional properties of the molecule. When sequences differ in conservative substitutions, the percent sequence identity may be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity." Means for making this adjustment are well known to those of skill in the art. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, if identical amino acids are given a score of 1 and non-conservative substitutions are given a score of 0, conservative substitutions are given a score of 0-1. Scoring of conservative substitutions is calculated, for example, as performed in the program PC / GENE (Intelligenetics, Mountain View, California).
[0271] As used herein, "percentage of sequence identity" refers to a value determined by comparing two optimally aligned sequences over a comparison window, where a portion of a polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (which does not contain additions or deletions) due to optimal alignment of the two sequences. The percentage is calculated by determining the number of positions where the same nucleic acid base or amino acid residue is present in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.
[0272] Unless otherwise specified, the sequence identity / similarity values provided herein refer to values obtained using GAP version 10 with the following parameters: % identity and % similarity for nucleotide sequences using a GAP weight of 50 and a length weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using a GAP weight of 8 and a length weight of 2, and the BLOSUM62 scoring matrix; or any equivalent program thereof. By "equivalent program" is intended any sequence comparison program that produces an alignment that has identical nucleotide or amino acid residue matches and the same percent sequence identity for any two sequences in question when compared to the corresponding alignment produced by GAP version 10.
[0273] Two sequences are "optimally aligned" when aligned for similarity scoring using a defined amino acid substitution matrix (e.g., BLOSUM62), gap existence penalties, and gap extension penalties to achieve the highest possible score for the sequence pair. Amino acid substitution matrices and their use in quantifying similarity between two sequences are well known in the art and are described, for example, in Dayhoff et al. (1978) "A model of evolutionary change in proteins." In "Atlas of Protein Sequence and Structure," Vol. 5, Suppl. 3 (ed. MO Dayhoff), pp. 345-352. Natl. Biomed. Res. Found., Washington, DC, and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA 89:10915-10919. The BLOSUM62 matrix is often used as the default scoring substitution matrix in sequence alignment protocols. A gap existence penalty is imposed for introducing a single amino acid gap into one of the aligned sequences, and a gap extension penalty is imposed for each additional empty amino acid position inserted into an already opened gap. The alignment is defined by the amino acid positions in each sequence where the alignment begins and ends, and optionally by inserting a gap or multiple gaps in one or both sequences to achieve the highest possible score. While optimal alignment and scoring can be achieved manually, this process is facilitated by the use of computer-implemented alignment algorithms, such as gapped BLAST 2.0, described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402 and publicly available at the National Center for Biotechnology Information Website (www.ncbi.nlm.nih.gov).Optimal alignments containing multiple alignments can be prepared using, for example, PSI-BLAST, available at www.ncbi.nlm.nih.gov and described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402.
[0274] For an amino acid sequence that is optimally aligned with a reference sequence, an amino acid residue "corresponds" to the position in the reference sequence with which it is paired in the alignment. The "position" is indicated by a number sequentially identifying each amino acid in the reference sequence based on its relative position relative to the N-terminus. Due to deletions, insertions, truncations, fusions, etc., which must be taken into account when determining optimal alignment, the amino acid residue number in a test sequence, determined by simple counting from the N-terminus, is generally not necessarily the same as the number at its corresponding position in the reference sequence. For example, if there is a deletion in the aligned test sequence, there will be no amino acid corresponding to the position in the reference sequence at the site of the deletion. If there is an insertion in the aligned reference sequence, the insertion will not correspond to any amino acid position in the reference sequence. In the case of a truncation or fusion, there may be a stretch of amino acids in the reference or aligned sequence that does not correspond to any amino acid in the corresponding sequence.
[0275] VI. Antibodies Also encompassed are antibodies against the RGN polypeptides or ribonucleoproteins comprising RGN polypeptides of the present invention, including those having any one of the amino acid sequences set forth as SEQ ID NOS: 1-20 or active variants or fragments thereof. Methods for producing antibodies are well known in the art (see, e.g., Harlow and Lane (1988) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY; and U.S. Pat. No. 4,196,265). These antibodies can be used in kits for detecting and isolating RGN polypeptides or ribonucleoproteins. Thus, the present disclosure provides kits containing antibodies that specifically bind to the polypeptides or ribonucleoproteins described herein, including, for example, polypeptides having any one of the amino acid sequences set forth as SEQ ID NOS: 1-20.
[0276] VII. RGN Systems and Ribonucleoprotein Complexes for Binding Target Sequences of Interest and Methods for Producing the Same The present disclosure provides systems for binding a target sequence of interest (e.g., a target DNA sequence), comprising at least one RNA-guided nuclease or a nucleotide sequence encoding the same and one or more guide RNAs capable of forming a complex with an RGN polypeptide (a ribonucleoprotein complex). The guide RNA hybridizes to a non-target strand of the target sequence of interest and forms a complex with the RGN polypeptide, thereby directing the RGN polypeptide to bind to the target DNA sequence. In some of these embodiments, the RGN comprises any one of the amino acid sequences set forth as SEQ ID NOS: 1-20, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat having a nucleotide sequence set forth as SEQ ID NOS: 21-41, or any one of nucleotides 1-14 of SEQ ID NO: 1040, nucleotides 1-17 of SEQ ID NO: 1041 or 1042, nucleotides 1-19 of SEQ ID NO: 1043, or nucleotides 1-22 of SEQ ID NO: 1044 or 1045, or an active variant or fragment thereof. In certain embodiments, the guide RNA comprises a tracrRNA having any one of the nucleotide sequences set forth as SEQ ID NOs: 42-62, nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NO: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045, or an active variant or fragment thereof. The guide RNA of the system can be a single-guide RNA or a dual-guide RNA. In certain embodiments, the system comprises an RNA-guided nuclease that is heterologous to the guide RNA, and the RGN and guide RNA are not naturally complexed with each other (i.e., bound to each other).
[0277] The system for binding to a target sequence of interest provided herein can be a ribonucleoprotein complex, which is at least one molecule of RNA bound to at least one protein. The ribonucleoprotein complex provided herein includes at least one guide RNA as the RNA component and an RNA-guided nuclease as the protein component. Such ribonucleoprotein complexes can be purified from cells or organisms that naturally express an RGN polypeptide and have been engineered to express a specific guide RNA specific to the target sequence of interest. Alternatively, the ribonucleoprotein complex can be purified from cells or organisms that have been transformed with a polynucleotide encoding an RGN polypeptide and a guide RNA (or a polynucleotide comprising a guide RNA) and cultured under conditions that allow expression of the RGN polypeptide and the guide RNA. Thus, methods for producing an RGN polypeptide or an RGN ribonucleoprotein complex are provided. Such methods include culturing cells containing a nucleotide sequence encoding an RGN polypeptide and, in some embodiments, a nucleotide sequence encoding or comprising a guide RNA under conditions in which the RGN polypeptide (in some embodiments, the guide RNA) is expressed. The RGN polypeptide or RGN ribonucleoprotein can then be purified from a lysate of the cultured cells. In embodiments, the nucleotide sequence encoding the RGN polypeptide comprises mRNA (messenger RNA). In embodiments, the method of assembling an RNP complex comprises combining one or more guide RNAs of the present disclosure with one or more RGN polypeptides of the present disclosure under conditions suitable for the formation of an RNP complex.
[0278] Methods for purifying RGN polypeptides or RGN ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., size exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reverse-phase chromatography, immunoprecipitation). In certain methods, the RGN polypeptide is recombinantly produced and contains a purification tag to aid in its purification, including, but not limited to, glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG (e.g., 3X FLAG tag), HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 10xHis, biotin carboxyl carrier protein (BCCP), and calmodulin. Typically, tagged RGN polypeptides or RGN ribonucleoprotein complexes are purified using immobilized metal affinity chromatography. It will be appreciated that other forms of chromatography or other similar methods known in the art may be used, alone or in combination, including, for example, immunoprecipitation.
[0279] An "isolated" or "purified" polypeptide, or biologically active portion thereof, is substantially or essentially free from components that normally accompany or interact with the polypeptide as found in its naturally occurring environment. Thus, an isolated or purified polypeptide is substantially free of other cellular material or culture medium if produced by recombinant techniques, or substantially free of chemical precursors or other chemicals if chemically synthesized. Proteins that are substantially free of cellular material include preparations of proteins having less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of contaminating proteins. When a protein of the invention or a biologically active portion thereof is recombinantly produced, optimally, the culture medium will represent less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of chemical precursors or non-target protein chemicals. Similarly, an "isolated" polynucleotide or nucleic acid molecule is removed from its naturally occurring environment. An isolated polynucleotide is substantially free of chemical precursors or other chemicals if chemically synthesized or removed from a genomic locus via cleavage of phosphodiester bonds. An isolated polynucleotide may be part of a vector, a composition, or may be contained within a cell, so long as the cell is not the original environment of the polynucleotide.
[0280] Certain methods provided herein for binding and / or cleaving target nucleic acid molecules containing a target sequence of interest involve the use of an in vitro assembled RGN ribonucleoprotein complex. In vitro assembly of an RGN ribonucleoprotein complex can be performed using any method known in the art for contacting an RGN polypeptide with a guide RNA under conditions that allow binding of the RGN polypeptide to the guide RNA. As used herein, "contact," "contacting," or "contacted" refers to placing components of a desired reaction together under conditions suitable for carrying out the desired reaction. The RGN polypeptide can be purified from a biological sample, cell lysate, or culture medium, produced via in vitro translation, or chemically synthesized. The guide RNA can be purified from a biological sample, cell lysate, or culture medium, transcribed in vitro, or chemically synthesized. The RGN polypeptide and guide RNA can be contacted in solution (e.g., buffered saline) to allow in vitro assembly of the RGN ribonucleoprotein complex.
[0281] VIII. Methods of Binding, Cleavaging, or Modifying Target Nucleic Acid Molecules The present disclosure provides methods for binding, cleaving, and / or modifying a target nucleic acid molecule of interest (e.g., target DNA) containing a target sequence. The methods include delivering a system comprising at least one guide RNA or a polynucleotide encoding the same and at least one RGN polypeptide or a polynucleotide encoding the same to the target sequence or a cell, organelle, or embryo containing the target sequence. In some of these embodiments, the RGN comprises any one of the amino acid sequences set forth as SEQ ID NOS: 1-20, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat comprising any one of the nucleotide sequences set forth as SEQ ID NOS: 21-41, or nucleotides 1-17 of SEQ ID NOS: 1041 or 1042, or nucleotides 1-22 of SEQ ID NOS: 1044 or 1045, or an active variant or fragment thereof. In certain embodiments, the guide RNA comprises a tracrRNA comprising any one of the nucleotide sequences set forth as SEQ ID NOs: 42-62, nucleotides 19-111 of SEQ ID NO: 1040, nucleotides 22-85 of SEQ ID NO: 1041 or 1042, nucleotides 24-138 of SEQ ID NO: 143, nucleotides 27-96 of SEQ ID NO: 1044, or nucleotides 27-95 of SEQ ID NO: 1045, or an active variant or fragment thereof. The guide RNA of the system can be a single guide RNA or a dual guide RNA.
[0282] The RGN of the system can be nuclease-dead RGN, can have nickase activity, or can be a fusion polypeptide. In some embodiments, the fusion polypeptide comprises a base-editing polypeptide, such as a cytosine deaminase or an adenine deaminase. In other embodiments, the RGN fusion protein comprises a reverse transcriptase. In other embodiments, the RGN fusion protein comprises a polypeptide that recruits a member of a functional nucleic acid repair complex, such as a member of the nucleotide excision repair (NER) or transcription-coupled nucleotide excision repair (TC-NER) pathway (Wei et al., 2015, PNAS USA 112(27):E3495-504; Troelstra et al., 1992, Cell 71:939-953; Marnef et al., 2017, J Mol Biol 429(9):1277-1288), as described in U.S. Provisional Patent Application No. 63 / 332,486, filed April 19, 2022, and incorporated by reference in its entirety. In some embodiments, the RGN fusion protein comprises CSB, a member of the TC-NER (nucleotide excision repair) pathway and which functions in recruiting other members (van den Boom et al., 2004, J Cell Biol 166(1):27-36; van Gool et al., 1997, EMBO J 16(19):5955-65; an example of which is set forth as SEQ ID NO: 563). In further embodiments, the RGN fusion protein comprises an active domain of CSB, such as the acidic domain of CSB comprising amino acid residues 356-394 of SEQ ID NO: 563 (Teng et al., 2018, Nat Commun 9(1):4115).
[0283] In certain embodiments, the RGN and / or guide RNA are heterologous to the cell, organelle, or embryo into which the RGN and / or guide RNA (or a polynucleotide encoding at least one of the RGN and guide RNA) is introduced.
[0284] In those embodiments in which the method includes delivering a polynucleotide encoding a guide RNA and / or an RGN polypeptide, the cell or embryo can then be cultured under conditions in which the guide RNA and / or RGN polypeptide is expressed. In various embodiments, the method includes contacting the target nucleic acid molecule with an RGN ribonucleoprotein complex. The RGN ribonucleoprotein complex may include RGN that is nuclease-dead or has nickase activity. In some embodiments, the RGN of the ribonucleoprotein complex is a fusion polypeptide that includes a base-editing polypeptide. In certain embodiments, the method includes introducing the RGN ribonucleoprotein complex into a cell, organelle, or embryo that contains the target nucleic acid molecule. The RGN ribonucleoprotein complex can be purified from a biological sample, recombinantly produced and subsequently purified, or assembled in vitro as described herein. In embodiments in which the RGN ribonucleoprotein complex contacted with the target nucleic acid molecule, organelle, or embryo is assembled in vitro, the method may further include in vitro assembly of the complex prior to contacting with the target nucleic acid molecule, cell, organelle, or embryo.
[0285] Purified or in vitro assembled RGN ribonucleoprotein complexes can be introduced into cells, organelles, or embryos using any method known in the art, including, but not limited to, electroporation. Alternatively, RGN polypeptides and / or polynucleotides encoding or comprising guide RNAs can be introduced into cells, organelles, or embryos using any method known in the art (e.g., electroporation).
[0286] Upon delivery to or contact with a target nucleic acid molecule or a cell, organelle, or embryo containing the target nucleic acid molecule, the guide RNA directs RGN to bind to a target sequence within the target nucleic acid molecule in a sequence-specific manner. In those embodiments in which RGN has nuclease activity, the RGN polypeptide cleaves the target sequence of interest upon binding. The target DNA sequence can then be modified via endogenous repair mechanisms, such as non-homologous end joining or homology-directed repair with the provided donor polynucleotide.
[0287] Methods for measuring the binding of RGN polypeptides to target sequences are known in the art and include chromatin immunoprecipitation assays, gel mobility shift assays, DNA pull-down assays, reporter assays, and microplate capture and detection assays. Similarly, methods for measuring the cleavage or modification of target nucleic acid molecules containing target sequences are known in the art and include in vitro or in vivo cleavage assays in which cleavage is confirmed using PCR, sequencing, or gel electrophoresis, with or without attaching an appropriate label (e.g., radioisotope, fluorescent substance) to the target sequence to facilitate detection of degradation products. Alternatively, a nicking-induced exponential amplification reaction (NTEXPAR) assay can be used (see, for example, Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage can be assessed using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).
[0288] In some embodiments, the methods involve the use of a single type of RGN complexed with two or more guide RNAs, which can target different regions of a single gene or can target multiple genes.
[0289] In embodiments where a donor polynucleotide is not provided, the double-strand break introduced by the RGN polypeptide can be repaired by the non-homologous end joining (NHEJ) repair process. Due to the error-prone nature of NHEJ, repair of the double-strand break can result in modification of the target sequence. As used herein, "modification" with respect to a nucleic acid molecule refers to a change in the nucleotide sequence of the nucleic acid molecule, which can be a deletion, insertion, or substitution of one or more nucleotides, or a combination thereof. Modification of a target nucleic acid molecule containing a target sequence can result in the expression of an altered protein product or inactivation of the coding sequence.
[0290] In these embodiments in which a donor polynucleotide is present, the donor sequence in the donor polynucleotide can be incorporated into or exchanged with the target nucleotide sequence during repair of the introduced double-strand break, resulting in the introduction of the exogenous donor sequence. Thus, the donor polynucleotide contains the donor sequence desired to be introduced into the target sequence of interest. In some embodiments, the donor sequence alters the original target nucleotide sequence so that the newly incorporated donor sequence is not recognized and cleaved by RGN. Donor sequence integration can be enhanced by the inclusion of flanking sequences, referred to herein as "homologous arms," within the donor polynucleotide that share substantial sequence identity with sequences flanking the target nucleotide sequence, enabling the homology-directed repair process. In some embodiments, the homologous arms have a length of at least 50 base pairs, at least 100 base pairs, and up to 2000 base pairs or more, and share at least 90%, at least 95%, or more sequence identity with their corresponding sequences in the target nucleotide sequence.
[0291] In those embodiments in which the RGN polypeptide introduces a double-stranded staggered break, the donor polynucleotide can contain a donor sequence flanked by compatible overhangs, allowing for direct ligation of the donor sequence to the cleaved target nucleotide sequence containing the overhangs by a non-homologous repair process during repair of the double-stranded break.
[0292] In embodiments involving the use of RGN, a nickase (i.e., capable of cleaving only one strand of a double-stranded polynucleotide), the method can include introducing two RGN nickases that target the same or overlapping target sequences and cleave different strands of the polynucleotide. For example, an RGN nickase that cleaves only the positive (+) strand of a double-stranded polynucleotide can be introduced along with a second RGN nickase that cleaves only the negative (-) strand of the double-stranded polynucleotide.
[0293] In various embodiments, methods for binding and detecting a target nucleotide sequence are provided, the methods comprising introducing at least one guide RNA or a polynucleotide encoding the same and at least one RGN polypeptide or a polynucleotide encoding the same into a cell, organelle, or embryo, and expressing the guide RNA and / or the RGN polypeptide (if a coding sequence is introduced), wherein the RGN polypeptide is nuclease-dead RGN and further comprises a detectable label, and the method further comprises detecting the detectable label. The detectable label may be fused to RGN as a fusion protein (e.g., a fluorescent protein) or may be a small molecule conjugated to or incorporated within the RGN polypeptide that can be detected visually or by other means.
[0294] Also provided herein are methods for regulating expression of a target gene of interest, including a target sequence or a gene under the control of a target sequence. The methods include introducing at least one guide RNA or a polynucleotide encoding the same and at least one RGN polypeptide or a polynucleotide encoding the same into a cell, organelle, or embryo, and expressing the guide RNA and / or the RGN polypeptide (if a coding sequence is introduced), wherein the RGN polypeptide is nuclease-dead RGN and further comprises a detectable label. In some of these embodiments, the nuclease-dead RGN is a fusion protein comprising an expression modulator domain (i.e., an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repressor domain) described herein.
[0295] The present disclosure also provides a method for binding and / or modifying a target nucleic acid molecule of interest comprising a target sequence, the method comprising delivering a system comprising at least one guide RNA or a polynucleotide encoding the same, and at least one fusion polypeptide comprising an RGN and a base-editing polypeptide of the present invention, such as a cytosine deaminase or an adenine deaminase, or a polynucleotide encoding the fusion polypeptide, to the target sequence or a cell, organelle, or embryo comprising the target sequence.
[0296] In some embodiments in which a fusion polypeptide comprising an RGN and a base-editing polypeptide is utilized, binding of the fusion protein to the target sequence results in modification of nucleotides adjacent to the target sequence. The nucleobases adjacent to the target sequence that are modified by the deaminase can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 base pairs from the 5' or 3' end of the target sequence.
[0297] Those skilled in the art will understand that any of the methods disclosed herein can be used to target a single target sequence or multiple target sequences. Thus, the methods include the use of a single RGN polypeptide in combination with multiple different guide RNAs that can target multiple different sequences within a single gene and / or multiple genes. Also encompassed herein are methods of introducing multiple different guide RNAs in combination with multiple different RGN polypeptides. These guide RNAs and guide RNA / RGN polypeptide systems can target multiple different sequences within a single gene and / or multiple genes.
[0298] In one aspect, the present invention provides kits comprising any one or more of the elements disclosed in the above methods and compositions, including crRNA, tracrRNA, guide RNA, RGN, and / or polynucleotides encoding them, cells, and an RGN system. In some embodiments, the kit comprises a vector system and instructions for using the kit. In some embodiments, the vector system comprises: (a) a first regulatory element operably linked to a DNA sequence encoding a guide RNA, wherein, when expressed, the guide RNA directs sequence-specific binding of an RGN complex to a target sequence in a eukaryotic cell, the RGN complex comprising an RGN enzyme complexed with the guide RNA polynucleotide, and one or more insertion sites for inserting the guide sequence upstream of the encoded guide RNA, and / or (b) a second regulatory element operably linked to an enzyme-coding sequence encoding the RGN enzyme, the enzyme comprising a nuclear localization sequence. In some embodiments, the kit further comprises a homologous recombination template polynucleotide. The elements may be provided individually or in combination and in any suitable container, such as a vial, bottle, or tube.
[0299] In some embodiments, the kit includes instructions in one or more languages. In some embodiments, the kit includes one or more reagents for use in a process utilizing one or more of the elements described herein. The reagents may be provided in any suitable container. For example, the kit may provide one or more reaction or storage buffers. The reagents may be provided in a form ready for use in a particular assay or in a form that requires the addition of one or more other components (e.g., concentrates or lyophilized forms) prior to use. The buffer may be any buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10.
[0300] In one aspect, the present invention provides a method for using one or more elements of the RGN system. The RGN system of the present invention provides an effective means for modifying target polynucleotides. The RGN system of the present invention has a wide variety of utilities, including modifying (e.g., deleting, inserting, moving, inactivating, activating, base editing) target polynucleotides in multiple cell types. Thus, the RGN system of the present invention has a wide range of applications, for example, in gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary RGN system or RGN complex comprises an RGN enzyme complexed with a guide sequence capable of binding to a target sequence.
[0301] IX. Target Polynucleotides In one aspect, the present invention provides a method for modifying a target polynucleotide comprising a target sequence, or a method for modifying expression of a target polynucleotide in a eukaryotic cell, which may be in vivo, ex vivo, or in vitro. In some embodiments, the method comprises sampling a cell or a population of cells from a human or non-human animal or plant (including microalgae) and modifying one or more cells. Culturing can be performed at any stage ex vivo. One or more cells may be reintroduced into the non-human animal or plant (including microalgae).
[0302] Using natural variability, plant breeders combine the most useful genes for desirable qualities such as yield, quality, uniformity, hardiness, and pest resistance. These desirable traits also include growth, photoperiod preference, temperature requirements, the onset of flowering or reproductive development, fatty acid content, insect resistance, disease resistance, nematode resistance, fungus resistance, herbicide resistance, and tolerance to various environmental factors, including adverse soil conditions, including drought, heat, humidity, cold, wind, and high salinity. Sources of these useful genes include native or exotic varieties, feather loom varieties, wild plant relatives, and induced mutations, such as treating plant material with mutagens. Using the present invention, plant breeders are provided with new tools for inducing mutations. Thus, those skilled in the art can analyze genomes for sources of useful genes and, in varieties with desired characteristics or traits, use the present invention to induce the elevation of useful genes with greater precision than previous mutagenic agents, thereby accelerating and improving plant breeding programs.
[0303] The target polynucleotide of the RGN system can be any polynucleotide, endogenous or exogenous to a eukaryotic cell. For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). Without wishing to be bound by theory, the target strand of the target sequence must be adjacent to a PAM (protospacer adjacent motif), i.e., a short sequence recognized by the RGN system. While the exact sequence and length requirements of the PAM vary depending on the RGN used, the PAM is typically a 2-7 base pair sequence adjacent to the protospacer (i.e., the target sequence).
[0304] Target polynucleotides of the RGN system can include numerous disease-associated genes and polynucleotides and signal transduction biochemical pathway-related genes and polynucleotides. Examples of target polynucleotides include sequences associated with signal transduction biochemical pathways, such as signal transduction biochemical pathway-related genes or polynucleotides. Examples of target polynucleotides include disease-associated genes or polynucleotides. A "disease-associated" gene or polynucleotide refers to any gene or polynucleotide that produces an abnormal level or abnormal form of transcription or translation product in cells derived from disease-affected tissue compared to non-diseased control tissue or cells. It can be a gene that becomes expressed at an abnormally high level or an abnormally low level, where the altered expression correlates with the onset and / or progression of the disease. A disease-associated gene also refers to a gene that has a mutation or genetic variation that is directly involved in or in linkage disequilibrium with a gene involved in the etiology of the disease (e.g., a causative mutation). The transcription or translation product can be known or unknown, and further, can be at normal or abnormal levels. In some embodiments, the disease can be an animal disease. In some embodiments, the disease can be an avian disease. In other embodiments, the disease may be a mammalian disease. In further embodiments, the disease may be a human disease. Examples of human disease-associated genes and polynucleotides are available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.).
[0305] While the RGN system is particularly useful due to its relative ease of targeting to genomic sequences of interest, the question of what RGN can do to address causative mutations remains. One approach is to produce a fusion protein between RGN (e.g., an inactive or nickase variant of RGN) and a base-editing enzyme, such as a cytosine deaminase or adenine deaminase-based editor (U.S. Patent No. 9,840,699, incorporated herein by reference), or the active domain of a base-editing enzyme. In some embodiments, the method includes contacting a DNA molecule containing a target sequence with (a) a fusion protein containing a base-editing polypeptide, such as an RGN or its nickase variant and a deaminase of the present invention; and (b) a gRNA that targets the fusion protein of (a) to the target sequence; the DNA molecule is contacted with the fusion protein and the gRNA in an effective amount under conditions suitable for deamination of nucleobases. In some embodiments, the target DNA sequence includes a sequence associated with a disease or disorder, and deamination of the nucleobases results in a sequence not associated with the disease or disorder. In some embodiments, the target DNA sequence is present in an allele of a crop plant, where a particular allele of a trait of interest results in a plant with less agronomic value, and deamination of the nucleobase results in an allele that improves the trait and increases the agronomic value of the plant.
[0306] In some embodiments, the target DNA sequence contains a T→C or A→G point mutation associated with a disease or disorder, and deamination of the mutant C or G base results in a sequence that is not associated with the disease or disorder. In some embodiments, deamination corrects the point mutation in the disease or disorder-associated sequence.
[0307] In some embodiments, the disease or disorder-associated sequence encodes a protein, and deamination introduces a stop codon into the disease or disorder-associated sequence, resulting in truncation of the encoded protein. In some embodiments, the contacting occurs in vivo in a subject susceptible to, having, or diagnosed with the disease or disorder. In some embodiments, the disease or disorder is a disease associated with a point mutation or single nucleotide mutation in the genome. In some embodiments, the disease is a genetic disease, cancer, metabolic disease, or lysosomal storage disease.
[0308] X. Pharmaceutical Compositions and Methods of Treatment Pharmaceutical compositions are provided comprising the RGN polypeptides disclosed herein and their active variants and fragments, and polynucleotides encoding them; the crRNAs disclosed herein and active variants and fragments thereof or polynucleotides encoding them; the tracrRNAs disclosed herein and active variants and fragments thereof or polynucleotides encoding them; the gRNAs disclosed herein or polynucleotides encoding them; the systems disclosed herein; or cells comprising any of the RGN polypeptides or polynucleotides encoding RGN, gRNAs or polynucleotides encoding gRNAs, or the RGN system; and a pharmaceutically acceptable carrier.
[0309] A pharmaceutical composition is a composition used to prevent, reduce, cure, or otherwise treat a target condition or disease comprising an active ingredient (i.e., an RGN polypeptide, a polynucleotide encoding RGN, a gRNA, a polynucleotide encoding a gRNA, an RGN system, or cells containing any one of these) and a pharmaceutically acceptable carrier.
[0310] As used herein, a "pharmaceutically acceptable carrier" refers to a material that does not cause significant irritation to an organism and does not abolish the activity and properties of the active ingredient (i.e., an RGN polypeptide, a polynucleotide encoding RGN, a gRNA, a polynucleotide encoding a gRNA, an RGN system, or a cell containing any one of these). The carrier must be of sufficiently high purity and sufficiently low toxicity to make it suitable for administration to the subject being treated. The carrier may be inert or may have pharmaceutical benefits. In some embodiments, a pharmaceutically acceptable carrier comprises one or more compatible solid or liquid fillers, diluents, or encapsulating substances suitable for administration to humans or other vertebrates. In some embodiments, a pharmaceutically acceptable carrier is not naturally occurring. In some embodiments, the pharmaceutically acceptable carrier and the active ingredient are not found together in nature.
[0311] Pharmaceutical compositions used in the methods of the present disclosure can be formulated with suitable carriers, excipients, and other agents to provide suitable entry, delivery, tolerance, etc. Numerous suitable formulations are known to those skilled in the art. See, for example, Remington, *The Science and Practice of Pharmacy* (21st ed. 2005). Suitable formulations include, for example, powders, pastes, ointments, jellies, waxes, oils, lipids, lipid (cationic or anionic)-containing vesicles (such as LIPOFECTIN vesicles), lipid nanoparticles, DNA conjugates, anhydrous absorbent pastes, oil-in-water and water-in-oil emulsions, emulsions of carbowax (polyethylene glycols of various molecular weights), semi-solid gels, and semi-solid mixtures containing carbowax. Pharmaceutical compositions for oral or parenteral use can be prepared in unit dosage forms appropriate for the dosage of the active ingredient. Examples of such unit dosage forms include tablets, pills, capsules, injections (ampoules), suppositories, etc.
[0312] The present disclosure provides pharmaceutical compositions comprising lipid-based formulations comprising an active ingredient (i.e., a guide RNA and / or RGN, or a polynucleotide comprising or encoding such). In some embodiments, the lipid-based formulation comprises a liposome. In some embodiments, the lipid-based formulation comprises a lipid nanoparticle (LNP). In some embodiments, the active ingredient is encapsulated in the lipid particle and / or disposed on the surface of the lipid particle. In some embodiments, the active ingredient is covalently bound to the lipid particle. In some embodiments, the active ingredient is non-covalently bound to the lipid particle. Covalent bonding involves the sharing of electrons in a chemical bond. Non-covalent bonding interactions include hydrogen bonding, ionic bonding, van der Waals interactions, and dispersive electromagnetic interactions such as hydrophobic bonding.
[0313] In some embodiments, the active ingredient is encapsulated in the lipid particle. The term "encapsulate" means to enclose, surround, or encase. As it relates to formulations of compounds of the present disclosure, encapsulation can be substantial, complete, or partial. The term "substantially encapsulated" or "substantial encapsulation" means that 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or more of the pharmaceutical composition or active ingredient of the present disclosure can be encapsulated, surrounded, or encased within the delivery agent (e.g., liposome or LNP). The term "partially encapsulated" or "partially encapsulated" means that less than 50%, 40%, 30%, 20%, 10%, or less of the pharmaceutical composition or active ingredient of the present disclosure can be encapsulated, surrounded, or encased within the delivery agent. Encapsulation can be determined by measuring leakage or activity of a pharmaceutical composition or active ingredient of the present disclosure using fluorescence and / or electron microscopy, for example, at least 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9% or more of the pharmaceutical composition or active ingredient of the present disclosure is encapsulated in the delivery agent (e.g., liposome or LNP).
[0314] Liposomes are spherical vesicular structures composed of a single or multiple lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes have attracted considerable attention as drug delivery vehicles because they are biocompatible and nontoxic, can deliver both hydrophilic and lipophilic drug molecules, protect their cargo from degradation by plasma enzymes, and transport their load across biological membranes and the blood-brain barrier (BBB) (see, e.g., Spuch and Navarro (2011) Journal of drug delivery 2011).
[0315] Liposomes can be made from several different types of lipids, but phospholipids are most commonly used to generate liposomes as drug carriers. Liposome formation occurs spontaneously when a lipid membrane is mixed with an aqueous solution, but can also be enhanced by applying force in the form of shaking using a homogenizer, sonicator, or extrusion device (see, e.g., Spuch and Navarro (2011) Journal of drug delivery 2011).
[0316] Conventional liposome formulations are primarily composed of natural phospholipids and phospholipids such as 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine, and monosialogangliosides. In some embodiments, 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) increases the stability of the liposomes.
[0317] Additives may be added to liposomes to modify their structure and properties. In some embodiments, cholesterol and / or sphingomyelin can be added to the liposome mixture to help stabilize the liposome structure and prevent leakage of the liposome internal cargo. In some embodiments, the addition of cholesterol to conventional liposome formulations reduces the rapid release of the encapsulated active ingredient (i.e., guide RNA and / or RGN, or a polynucleotide containing or encoding such) into plasma. In some embodiments, liposomes are prepared from hydrogenated egg phosphatidylcholine or egg phosphatidylcholine, cholesterol, and dicetyl phosphate. In some embodiments, the average liposome vesicle size is adjusted to about 50 or 100 nm.
[0318] In some embodiments, toroidal horse liposomes (also known as molecular Trojan horses or PEGylated immunoliposomes) can be used in pharmaceutical compositions for delivery of active ingredients across the BBB (described on the World Wide Web at cshprotocols.cshlp.org / content / 2010 / 4 / pdb.prot5407.long). Without being bound by any theory, it is believed that neutral lipid particles with specific antibodies conjugated to their surface enable crossing of the BBB via endocytosis. In some embodiments, pharmaceutical compositions containing Trojan horse liposomes can be used to deliver active ingredients (i.e., guide RNA and / or RGN, or polynucleotides containing or encoding such) to the brain via intravascular injection.
[0319] In some embodiments, the liposome comprises a stable nucleic acid-lipid particle (SNALP) (e.g., Morrissey et al. (2005) Nature Biotechnology 23(8):1002-1007; Zimmerman et al. (2006) Nature 441:111-114). SNALPs comprise a mixture of cationic and fusogenic lipids and are coated with polyethylene glycol (PEG), which allows cellular uptake and endosomal release of the active ingredient cargo. In some embodiments, SNALPs are a class of LNPs and comprise an ionizable lipid that is cationic at low pH (e.g., DLinDMA), a neutral helper lipid, cholesterol, and a diffusible polyethylene glycol (PEG)-lipid. In some embodiments, the SNALP formulation comprises the following lipids: 3-N-(-methoxypoly(ethylene glycol)2000)carbamoyl-1,2-dimyristyloxy-propylamine (PEG-cDMA); 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA); 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC); and cholesterol. In some embodiments, the SNALP comprises synthetic cholesterol, dipalmitoylphosphatidylcholine (DOPC), PEG-cDMA, and DLinDMA (see, e.g., Geisbert et al. (2010) Lancet 375:1896-1905). In some embodiments, the SNALP comprises synthetic cholesterol, DSPC, PEG-cDMA, and DLinDMA (see, e.g., Judge et al. (2009) J. Clin. Invest. 119:661-673). In some embodiments, SNALP liposomes are about 80-100 nm in size. SNALP has been used as an effective delivery molecule for highly vascularized HepG2-derived liver tumors (see, e.g., Li et al. (2012) Gene Therapy 19:775-780).
[0320] Without being bound by any theory, in the formulation of SNALP, ionizable lipids serve to condense the lipids with the active ingredient (e.g., nucleic acid molecules) during particle formation. Positively charged under increasingly acidic endosomal conditions, the ionizable lipids may mediate fusion of SNALP with the endosomal membrane, allowing the active ingredient to be released into the cytoplasm. The PEG-lipid may stabilize the particles and reduce aggregation during formulation, and subsequently provide a neutral hydrophilic exterior that improves pharmacokinetic properties. In some embodiments, SNALP liposomes are prepared by formulating DLinDMA and PEG-cDMA with DSPC, cholesterol, and the active ingredient using a lipid:active ingredient ratio of 25:1 and a 48:40:10:2 molar ratio of cholesterol:DLinDMA:DSPC:PEG-cDMA.
[0321] In some embodiments, the pharmaceutical composition of the present disclosure comprises an LNP. In some embodiments, lipids may be formulated with the active ingredients of the present disclosure to form an LNP. An LNP comprises multiple lipid molecules physically associated with each other through intermolecular forces. In some embodiments, an LNP comprises a liposome. In some embodiments, an LNP differs from a liposome in that it does not have a continuous lipid bilayer. In some embodiments, an LNP comprises a solid particle having a mixture of solid and liquid lipids. In some embodiments, an LNP comprises a dendrimer lipid nanoparticle (DLNP), SNALP, and lipid-like nanoparticle (LLNP). Generally, "nanoparticle" refers to any particle having a diameter less than 1000 nanometers (nm). In some embodiments, a nanoparticle has a diameter of 500 nm or less. In some embodiments, a nanoparticle has a diameter ranging from 25 nm to 200 nm, or 100 nm or less. In some embodiments, a nanoparticle has a diameter ranging from 35 nm to 60 nm. In some embodiments, an LNP comprises a lipid particle having a size of about 1 to about 100 nm.
[0322] LNPs comprise four components: an ionizable cationic lipid, a fusogenic zwitterionic phospholipid, cholesterol, and a PEGylated (PEG) lipid. In some embodiments, the ionizable cationic lipid component complexes with negatively charged polynucleotides and promotes endosomal escape. In some embodiments, the phospholipid component functions in modifying the lipid bilayer structure. In some embodiments, the cholesterol component serves to stabilize the LNP. In some embodiments, the PEG lipid component reduces LNP aggregation and nonspecific uptake.
[0323] Ionizable cationic lipids useful in LNPs include 1,2-dilinoleyl-3-dimethylammonium-propane (DLinDAP); DLinDMA; l,2-dilinoleyloxy-keto-N,N-dimethyl-3-aminopropane (DLinK-DMA); 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA); 5A2-SC8 (Zhou et al. (2016) Proc. Natl Acad. Sci. USA 113:520-525); C12-200 (Love et al. (2010) Proc. Natl Acad. Sci. USA 107:1864-1869); 246C10 (Kim et al. (2021) Sci Adv 7(9):eabf4398); cKK-E12 (Fenton et al. (2016) Advanced Materials 28(15):2939-2943); 1,2-distearyloxy-N,N-dimethyl-3-aminopropane (DSDMA); 1,2-dioleyloxy-N,N-dimethyl-3-aminopropane (DODMA); 1,2-dilinolenyloxy-N,N-dimethyl-3-aminopropane (DLenDMA); and dilinoleylmethyl-4-dimethylaminobutyric acid (Dlin-MC3-DMA; Jayaraman et al. (2012) Angew Chem Int Ed Engl. 51(34):8529-8533). Cationic lipids are further described in WO2012040184, WO2011153120, WO2011149733, WO2011090965, WO2011043913, WO2011022460, WO2012061259, WO2012054365, WO2012044638, WO2010080724, WO201021865 and WO2008103276, U.S. Patent Nos. 7,893,302 and 7,404,969, and U.S. Patent Publication No. 20100036115, each of which is incorporated herein by reference in its entirety.
[0324] Zwitterionic phospholipids useful in LNPs include DSPC, DOPE, and DOPC.
[0325] PEG lipids useful in LNPs include 1,2-dimyristoyl-rac-glycero-3-methoxypolyethylene glycol (PEG-DMG); (3-o-[2″-(methoxypolyethylene glycol 2000) succinoyl]-1,2-dimyristoyl-sn-glycol (PEG-S-DMG); R-3-[(ω-methoxy-poly(ethylene glycol) 2000) carbamoyl]-1,2-dimyristyloxylpropyl-3-amine (PEG-C-DOMG); and C16 PEG-ceramide. In some embodiments, LNPs comprise DLinKC2-DMA or C12-200:DSPC:cholesterol:PEG-DMG in a molar ratio of 50:10:38.5:1.5 (see, e.g., Basha et al. (2011) Molecular Therapy 19(12):2186-2200). In some embodiments, the LNPs comprise a 26.5:20:52:1.5 ratio of ionizable lipid:DOPE:cholesterol:PEG lipid (see, e.g., Han et al. (2022) Sci Adv 8(3):eabj6901; Kim et al. (2021) Sci Adv 7(9):eabf4398). PEG lipids are further described in WO 2012099755. In some embodiments, the ratio of PEG in the LNP formulation can be increased or decreased, and / or the carbon chain length of the PEG lipid can be modified from C14 to C18 to alter the pharmacokinetics and / or biodistribution of the LNP formulation.
[0326] In some embodiments, the charge of LNP is taken into consideration. Cationic lipids can be combined with negatively charged lipids to induce a non-bilayer structure that promotes intracellular delivery. Because charged LNPs are rapidly removed from circulation after intravenous injection, ionizable cationic lipids with pKa values less than 7 have been developed (see, for example, Basha et al. (2011) Molecular Therapy 19(12):2186-2200). Negatively charged polymers, such as polynucleotides, can be loaded into LNPs at low pH values (e.g., pH 4), where ionizable lipids exhibit positive charges. However, at physiological pH values, LNPs exhibit low surface charges, which are compatible with longer circulation times.
[0327] The preparation of LNPs and the encapsulation of active ingredients are described, for example, in Basha et al. (2011) Molecular Therapy 19(12):1286-2200; Han et al. (2022) Sci Adv 8(3):eabj6901; Kim et al. (2021) Sci Adv 7(9):eabf4398; Finn et al. (2018) Cell Reports 22:2227-2235; Wei et al. (2020) Nature Communications 11:3232; WO 2011127255; and WO 2008103276. Lipids can be commercially available (e.g., from Tekmira Pharmaceuticals, Vancouver, Canada; Avanti Polar Lipids, Inc., Alabaster, AL) or synthesized (e.g., Kim et al. (2021) Sci Adv 7(9):eabf4398). The synthesis of cationic lipids is also described in WO 2012040184, WO 2011153120, WO 2011149733, WO 2011090965, WO 2011043913, WO 2011022460, WO 2012061259, WO 2012054365, WO 2012044638, WO 2010080724, and WO 201021865. Cholesterol is commercially available (e.g., from Sigma-Aldrich, St Louis, MO).
[0328] In some embodiments, encapsulation can be performed by dissolving a lipid mixture containing a cationic lipid (e.g., Dlin-DMA):phospholipid (e.g., DSPC, DOPE):cholesterol:PEG-lipid (e.g., a molar ratio of 40:10:40:10) in ethanol. The active ingredient (e.g., a polynucleotide containing or encoding a guide RNA or RGN of the present disclosure) can be dissolved in an acidic buffer (e.g., citrate, acetate), pH 3 or 4. In some embodiments, the lipid solution and the active ingredient solution can be mixed using a microfluidics system (Chen et al. (2012) J. Amer. Chem. Soc. 134:6948-6951; e.g., NanoAssemblr from Precision Nanosystems) or by adding the lipid solution dropwise to the active ingredient solution. Removal of ethanol and neutralization of the formulation buffer can be performed by dialysis against phosphate-buffered saline (PBS) for, e.g., 16 hours or overnight using a dialysis cassette (e.g., a 3500 molecular weight cutoff cassette from Life Technologies). Dynamic light scattering can be used to assess LNP size, polydispersity index (PDI), and zeta potential. Encapsulation efficiency of active ingredients, such as RNA, can be determined by assays such as the Quant-it™ Ribogreen Assay (Thermo Fisher). In some embodiments, where the encapsulated active ingredient is a polynucleotide, the polynucleotide can be extracted from the eluted nanoparticles and quantified at 260 nm. LNP pKa can be assessed using the 2-(p-toluidino)-6-naphthalenesulfonic acid (TNS) assay (Zhang et al. (2011) Langmuir 27(5):1907-1914). In some embodiments, the final lipid:active ingredient weight ratio includes 12:1, 11:1, 10:1, 9:1, 8:1, 7:1, 6:1, and 5:1.
[0329] In some embodiments, where the pharmaceutical composition comprises LNPs encapsulated in ribonucleoprotein (RNP) complexes (i.e., RGN and guide RNA), the inclusion of an additional permanent cationic lipid (e.g., 1,2-dioleoyl-3-trimethylammonium-propane (DOTAP)) allows for the formation of LNPs containing RNPs by mixing an ethanolic solution of the lipid with a solution of RNPs at physiological pH (e.g., PBS buffer; Wei et al. (2020) Nature Communications 11:3232). In some embodiments, the permanent cationic lipid comprises 10 to 20 mol% of the total lipids in the LNPs.
[0330] In some embodiments, the LNP formulations described herein can further comprise a permeability enhancer molecule. Non-limiting permeability enhancer molecules are described in U.S. Patent Application Publication No. 2005 / 0222064.
[0331] In some embodiments, LNP compositions are biodegradable in that they do not accumulate to cytotoxic levels in vivo at therapeutically effective doses. LNP formulations can be improved by replacing cationic lipids with biodegradable cationic lipids known as rapidly cleared lipid nanoparticles (reLNPs). In some embodiments, the rapid metabolism of rapidly cleared lipids can improve the tolerability and therapeutic index of LNPs by an order of magnitude from 1 mg / kg to 10 mg / kg doses in rats. The inclusion of enzymatically degradable ester linkages can improve the degradation and metabolic profile of the cationic component while still maintaining the activity of reLNP formulations. The ester linkages can be located internally within the lipid chain or at the terminus of the lipid chain. Internal ester linkages can replace any carbon in the lipid chain.
[0332] In some embodiments, the LNP compositions do not elicit an innate immune response that results in substantial adverse effects at therapeutic dose levels, hi some embodiments, the LNP compositions provided herein do not elicit toxicity at therapeutic dose levels.
[0333] In some embodiments, the active ingredient (i.e., guide RNA and / or RGN, or a polynucleotide containing or encoding such) is formulated as a solid lipid nanoparticle. Solid lipid nanoparticles (SLNs) can be spherical with an average diameter between 10 and 1000 nm. SLNs have a solid lipid core matrix that can solubilize lipophilic molecules and can be stabilized with surfactants and / or emulsifiers. In further embodiments, the lipid nanoparticles can be self-assembling lipid-polymer nanoparticles (see, e.g., Zhang et al. (2008) ACS Nano 2(8):1696-1702).
[0334] In some embodiments, lipid-based formulations comprising an active ingredient (i.e., a guide RNA and / or RGN, or a polynucleotide comprising or encoding such) can be formulated for controlled release and / or targeted delivery. As used herein, "controlled release" refers to a pharmaceutical composition or compound release profile that conforms to a specific release pattern to produce a therapeutic result.
[0335] In some embodiments, the lipid-based formulation comprising the active ingredient (i.e., guide RNA and / or RGN, or a polynucleotide containing or encoding such) comprises at least one controlled-release coating. Controlled-release coatings include OPADRY® (Colorcon Inc., Harleysville, PA); polyvinylpyrrolidone / vinyl acetate copolymer; polyvinylpyrrolidone; hydroxypropyl methylcellulose; hydroxypropyl cellulose; hydroxyethyl cellulose; EUDRAGIT RL® (Evonik, Essen, Germany); EUDRAGIT RS® (Evonik, Essen, Germany); cellulose derivatives such as ethyl cellulose aqueous dispersions (AQUACOAT® and SURELEASE®, Colorcon Inc., Harleysville, PA). In some embodiments, the controlled-release and / or targeted-delivery formulation may comprise at least one degradable polyester, which may comprise polycationic side chains. Degradable polyesters include poly(serine ester), poly(L-lactide-co-L-lysine), poly(4-hydroxy-L-proline ester), and combinations thereof. In some embodiments, the degradable polyesters may include PEG conjugation to form PEGylated polymers.
[0336] In some embodiments, LNP formulations can be prepared to be passively or actively targeted to different cell types in vivo, including hepatocytes, immune cells, tumor cells, endothelial cells, antigen-presenting cells, and leukocytes (Akinc et al. (2010) Mol Ther. 18:1357-1364; Song et al. (2005) Nat Biotechnol. 23:709-717; Judge et al. (2009) J Clin Invest. 119:661-673; Kaufmann et al. (2010) Microvasc Res 80:286-293; Santel et al. (2006) Gene Ther 13:1222-1234; Santel et al. (2006) Gene Ther 13:1360-1370; Gutbier ... (See, e.g.,
[0004] et al. (2010) Pulm Pharmacol. Ther. 23:334-344; Basha et al. (2011) Mol. Ther. 19:2186-2200; Fenske and Cullis (2008) Expert Opin Drug Deliv. 5:25-44; Peer et al. (2008) Science 319:627-630; Peer and Lieberman (2011) Gene Ther. 18:1127-1133). An example of passive targeting of pharmaceuticals to hepatocytes is DLin-DMA, DLin-KC2-DMA, and MC3-based lipid nanoparticle formulations, which have been shown to bind to apolipoprotein E and promote hepatocyte binding and uptake of these pharmaceuticals in vivo (Akinc et al. (2010) Mol Ther. 18:1357-1364).
[0337] LNP formulations can also be selectively targeted through the expression of different ligands on their surface, such as folate, transferrin, N-acetylgalactosamine (GalNAc), and antibody targeting approaches (Kolhatkar et al. (2011) Curr Drug Discov Technol. 8:197-206; Musacchio and Torchilin (2011) Front Biosci. 16:1388-1412; Yu et al. (2010) Mol Membr Biol. 27:286-298; Patil et al. (2008) Crit Rev Ther Drug Carrier Syst. 25:1-61; Benoit et al. (2011) Biomacromolecules. 12:2708-2714; Zhao et al. (2008) Expert Opin Drug Deliv. 5:309-319; Akinc et al. al.(2010)Mol Ther.18:1357-1364;Srinivasan et al.(2012)Methods Mol Biol.820:105-116;Ben-Arie et al.(2012)Methods Mol Biol.757:497-507;Peer,D(2010)J of controlled release 148(1):63-68;Peer et al.(2007)Proc Natl Acad Sci USA.104:4095-4100;Kim et al.(2011)Methods Mol Biol.721:339-353;Subramanya et al.(2010)Mol Ther.18:2028-2037;Song et al. al.(2005)Nat Biotechnol. 23:709-717; Peer et al. (2008) Science 319:627-630; Peer and Lieberman (2011) Gene Ther. 18:1127-1133; all of which are incorporated herein by reference in their entireties).
[0338] In some embodiments, the active ingredient (i.e., guide RNA and / or RGN, or a polynucleotide containing or encoding such) can be encapsulated in an LNP, and the LNP can then be encapsulated in a polymer, polymer matrix, hydrogel, and / or surgical sealant described herein and / or known in the art. In some embodiments, the polymer, hydrogel, or surgical sealant includes poly(lactic-co-glycolic acid) (PLGA); ethylene vinyl acetate (EVAc); poloxamer; surgical sealants such as GELSITE® (Nanotherapeutics, Inc., Alachua, FL); HYLENEX® (Halozyme Therapeutics, San Diego, CA); fibrinogen polymers (Ethicon Inc., Cornelia, GA) and TISSELL® (Baxter International, Inc. Deerfield, IL); PEG-based sealants; and COSEAL® (Baxter International, Inc. Deerfield, IL).
[0339] LNPs and LNP formulations are further described in, for example, U.S. Patent Nos. 7,982,027; 7,799,565; 8,058,069; 8,283,333; 7,901,708; 7,745,651; 7,803,397; 8,101,741; 8,188,263; 7,915,399; 8,236,943 and 7,838,658; European Patent Nos. 1766035; 1519714; 1781593; and 1664316.
[0340] In some embodiments in which cells containing or modified with the RGN, gRNA, RGN system, or polynucleotides encoding the same disclosed herein are administered to a subject, the cells are administered as a suspension containing a pharmaceutically acceptable carrier. Those skilled in the art will recognize that pharmaceutically acceptable carriers used in cell compositions do not contain amounts of buffers, compounds, cryopreservatives, preservatives, or other agents that substantially interfere with the viability of the cells delivered to a subject. Cell-containing formulations may include, for example, an osmotic buffer that allows the integrity of the cell membrane to be maintained, and optionally, nutrients to maintain cell viability or enhance engraftment upon administration. Such formulations and suspensions are known to those skilled in the art and / or can be adapted for use with the cells described herein using routine experimentation.
[0341] The cell compositions may also be emulsified or presented as liposomal compositions, so long as the emulsification procedure does not adversely affect cell viability. The cells and any other active ingredients may be mixed with excipients that are pharmaceutically acceptable and compatible with the active ingredients, in amounts suitable for use in the therapeutic methods described herein.
[0342] Additional agents included in the cell compositions may include pharmaceutically acceptable salts of the components therein. Pharmaceutically acceptable salts include acid addition salts (formed with the free amino groups of the polypeptide) formed with inorganic acids, such as hydrochloric or phosphoric acids, or organic acids such as acetic, tartaric, mandelic, and the like. Salts formed with the free carboxyl groups can also be derived from inorganic bases, such as sodium hydroxide, potassium hydroxide, ammonium hydroxide, calcium hydroxide, or ferric hydroxide, and organic bases such as isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, procaine, and the like.
[0343] Physiologically acceptable and pharmaceutically acceptable carriers are well known in the art. Exemplary liquid carriers are sterile aqueous solutions containing either no or both materials in addition to the active ingredient and water, buffers such as sodium phosphate, saline, or phosphate-buffered saline at physiological pH values. Additionally, aqueous carriers can contain two or more buffer salts, as well as salts such as sodium and potassium chloride, dextrose, polyethylene glycol, and other solutes. Liquid compositions can also contain liquid phases in addition to and in addition to water. Examples of such additional liquid phases are glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of active compound used in a cell composition effective in treating a particular disorder or condition will depend on the nature of the disorder or condition and can be determined by standard clinical techniques.
[0344] The RGN polypeptides, guide RNAs, RGN systems, or polynucleotides encoding them of the present disclosure can be formulated with pharmaceutically acceptable excipients, such as carriers, solvents, stabilizers, adjuvants, and diluents, depending on the particular mode of administration and dosage form. In some embodiments, these pharmaceutical compositions are formulated to achieve a physiologically compatible pH, ranging from about pH 3 to about pH 11, or from about pH 3 to about pH 7, depending on the formulation and route of administration. In some embodiments, the pH can be adjusted to a range of about pH 5.0 to about pH 8. In some embodiments, the compositions can comprise a therapeutically effective amount of at least one compound described herein along with one or more pharmaceutically acceptable excipients. In some embodiments, the compositions comprise a combination of compounds described herein, or a second active ingredient useful for the treatment or prevention of bacterial growth (e.g., but not limited to, an antibacterial or antimicrobial agent), or a combination of reagents of the present disclosure.
[0345] Suitable excipients include, for example, carrier molecules comprising large, slowly metabolized macromolecules such as proteins, polysaccharides, polylactic acids, polyglycolic acids, polymeric amino acids, amino acid copolymers, and inactive virus particles. Other exemplary excipients may include antioxidants (for example, but not limited to, ascorbic acid), chelating agents (for example, but not limited to, EDTA), carbohydrates (for example, but not limited to, dextrin, hydroxyalkylcellulose, and hydroxyalkylmethylcellulose), stearic acid, liquids (for example, but not limited to, oils, water, saline, glycerol, and ethanol), wetting or emulsifying agents, pH buffering substances, etc.
[0346] In some embodiments, the formulations are presented in unit-dose or multi-dose containers, such as sealed ampoules and vials, and may be stored in a lyophilized (lyophilized) condition requiring the addition of a sterile liquid carrier, such as saline, water for injection, a semi-liquid foam, or a gel, immediately prior to use. Extemporaneous injection solutions and suspensions may be prepared from sterile powders, granules, and tablets of the kind described above. In some embodiments, the active ingredient is frozen in a unit-dose or multi-dose container and then thawed for injection or dissolved in a buffer solution that is kept / stabilized under refrigeration until use.
[0347] The therapeutic agent may be contained in a controlled release system. To prolong the effect of a drug, it is often desirable to slow the absorption of the drug from subcutaneous, intrathecal, or intramuscular injection. This can be achieved by using a liquid suspension of crystalline or amorphous material with poor water solubility. The rate of absorption of the drug depends on its dissolution rate, which may depend on crystal size and crystalline form. Alternatively, delayed absorption of a parenterally administered drug form can be achieved by dissolving or suspending the drug in an oil vehicle. In some embodiments, the use of a long-term sustained-release implant may be particularly suitable for treating chronic conditions. Long-term sustained-release implants are well known to those skilled in the art.
[0348] Provided herein are methods for treating a disease in a subject in need thereof, the method comprising administering to the subject in need thereof an effective amount of an RGN polypeptide of the present disclosure or an active variant or fragment thereof, or a polynucleotide encoding same, a gRNA of the present disclosure or a polynucleotide encoding same, an RGN system of the present disclosure, or a cell modified by or comprising any one of these compositions.
[0349] In some embodiments, treatment involves in vivo gene editing by administering the presently disclosed RGN polypeptide, gRNA, or RGN system, or a polynucleotide encoding same. In some embodiments, treatment involves ex vivo gene editing, in which cells are genetically modified ex vivo with the presently disclosed RGN polypeptide, gRNA, or RGN system, or a polynucleotide encoding same, and then the modified cells are administered to a subject. In some embodiments, the genetically modified cells are derived from the subject to whom the modified cells are then administered, and the transplanted cells are referred to herein as autologous. In some embodiments, the genetically modified cells are derived from a different subject (i.e., donor) within the same species as the subject to whom the modified cells (i.e., recipient) are administered, and the transplanted cells are referred to herein as allogeneic. In some examples described herein, the cells can be expanded in culture before administration to a subject in need thereof.
[0350] In some embodiments, the disease treated with the compositions of the present disclosure is a disease that can be treated with immunotherapy, such as chimeric antigen receptor (CAR) T cells. Such diseases include, but are not limited to, cancer. In some embodiments, the disease treated with the compositions of the present disclosure is associated with a causative mutation. As used herein, a "causative mutation" refers to a specific nucleotide, nucleotide, or nucleotide sequence in a genome that contributes to the severity or presence of a disease or disorder in a subject. Correction of the causative mutation results in amelioration of at least one symptom caused by the disease or disorder. In some embodiments, the causative mutation is adjacent to a PAM site recognized by a RGN disclosed herein. The causative mutation can be corrected with a presently disclosed RGN or a fusion polypeptide comprising a presently disclosed RGN and a base-editing polypeptide (i.e., a base editor). Non-limiting examples of diseases associated with causative mutations include cystic fibrosis, Hurler syndrome, Friedreich's ataxia, Huntington's disease, and sickle cell disease. Further non-limiting examples of disease-associated genes and mutations are shown in Table 6 and are also available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.).
[0351] In some embodiments, a method of treating a disease in a subject in need thereof includes generating induced pluripotent stem cells (iPSCs) or isolating mesenchymal stem cells from the subject, contacting the iPSCs or mesenchymal stem cells with any one of the RGN polypeptides, systems, compositions comprising same, or pharmaceutical compositions disclosed herein to genetically modify a target nucleic acid molecule in the cells, differentiating the modified iPSCs or mesenchymal stem cells into genetically modified mature cells or precursors thereof, and administering the genetically modified mature cells or precursors thereof to the subject. In some embodiments, the iPSCs or mesenchymal stem cells are autologous or allogeneic. In some embodiments, the iPSCs or mesenchymal stem cells are derived from a donor that is a full human leukocyte antigen (HLA) match for the subject. In some embodiments, the subject is subjected to myeloablative therapy prior to administration of the modified cells.
[0352] Any method known in the art for generating patient-specific iPS cells can be used, including, but not limited to, those described in Takahashi and Yamanaka, Cell 126(4):663-76, 2006. For example, the generation process includes: a) isolating somatic cells, such as skin cells or fibroblasts, from a subject; and b) introducing a set of pluripotency-associated genes into the somatic cells to induce the cells into pluripotent stem cells. The set of pluripotency-associated genes can be one or more genes selected from the group consisting of OCT4, SOX1, SOX2, SOX3, SOX15, SOX18, NANOG, KLF1, KLF2, KLF4, KLF5, c-MYC, n-MYC, REM2, TERT, and LIN28. Mesenchymal stem cells can be isolated from the patient's bone marrow or peripheral blood according to any method known in the art. For example, bone marrow aspirate can be collected in a syringe containing heparin. Cells can be washed and centrifuged with Percoll. Cells can be cultured in Dulbecco's Modified Eagle's Medium (DMEM) (low glucose) containing 10% fetal bovine serum (FBS) (Pittinger MF, Mackay AM, Beck SC et al., Science 1999;284:143-147).
[0353] The genetically modified cells of the present disclosure administered to a subject include autologous and allogeneic cells. Allogeneic cells refer to cells derived from one or more donors (i.e., one or more individuals from whom the genetically modified cells are derived). Autologous cells refer to cells derived from the subject undergoing treatment (i.e., the recipient of the genetically modified cells). Due to the risk of transplant rejection, efforts are made to optimize the degree of major histocompatibility complex (MHC) / human leukocyte antigen (HLA) compatibility between donor tissue and recipient. HLA is found on the surface of cells and helps the body distinguish between self and non-self, allowing the body to attack foreign substances such as bacteria and viruses. HLA typing of donor tissue and recipient involves determining the genotypes of the six HLA antigens, or alleles, between the donor and recipient to assess the degree of matching of the six HLAs. HLA alleles typically refer to two each at the HLA-A, HLA-B, and HLA-DR loci, or one each at the HLA-A, HLA-B, and HLA-C loci, and one each at the HLA-DRB1, HLA-DQB1, and HLA-DPB1 loci (see, e.g., Kawase et al., 2007, Blood 110:2235-2241). In some embodiments, four of six HLA matches between a donor and a recipient are sufficient for administration of donor-derived cells to a recipient. In some embodiments, five of six HLA matches between a donor and a recipient are sufficient for administration of donor-derived cells to a recipient. In some embodiments, six of six HLA matches are matched between a donor and a recipient for administration of donor-derived cells to a recipient. Generally, a 4 / 6, 5 / 6, or 6 / 6 HLA match is the standard of clinical care. If all six HLAs match between the donor and recipient, the match is called a perfect match.
[0354] As used herein, "treatment" or "treating" or "palliative" or "amelioration" are used interchangeably. These terms refer to an approach for obtaining beneficial or desired results, including, but not limited to, therapeutic benefit and / or prophylactic benefit. Therapeutic benefit refers to a therapeutically relevant improvement or effect of one or more diseases, conditions, or symptoms being treated. For prophylactic benefit, the composition may be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject who may not yet have the disease, condition, or symptom, but who reports one or more physiological symptoms of the disease. In some embodiments, treatment may be administered after one or more symptoms have developed and / or after a disease has been diagnosed. In certain embodiments, treatment may be administered in the absence of symptoms, e.g., to prevent or delay the onset of symptoms or inhibit the onset or progression of a disease. For example, treatment may be administered to a susceptible individual prior to the onset of symptoms (e.g., in light of a history of symptoms and / or in light of genetic or other susceptibility factors). Treatment may also be continued after symptoms have resolved, for example to prevent or delay their recurrence.
[0355] The term "effective amount" or "therapeutically effective amount" refers to an amount of an agent sufficient to produce a beneficial or desired result. A therapeutically effective amount may vary depending on one or more of the subject and disease state being treated, the subject's weight and age, the severity of the disease state, the mode of administration, etc., and these can be readily determined by one of ordinary skill in the art. The specific dose may vary depending on one or more of the particular agent selected, the dosing regimen to be followed, whether it is administered in combination with other compounds, the timing of administration, and the delivery system in which it is delivered.
[0356] The term "administration" refers to the placement of an active ingredient in a subject by a method or route that results in at least partial localization of the introduced active ingredient at a desired site, such as a site of injury or repair, so that a desired effect occurs. In some embodiments, the present disclosure provides methods comprising delivering any of the RGN polypeptides, nucleic acid molecules, ribonucleoprotein complexes, vectors, pharmaceutical compositions, and / or gRNAs described herein. In some embodiments, the present disclosure further provides cells produced by such methods, and organisms (such as animals or plants) comprising or produced from such cells. In some embodiments, the RGN polypeptides and / or nucleic acid molecules described herein, combined (and optionally complexed) with a guide sequence, are delivered to cells.
[0357] In embodiments in which cells are administered, the cells can be administered by any suitable route that results in delivery to the desired location in the subject, where at least a portion of the transplanted cells or components of the cells remain viable. The survival period of the cells after administration to the subject can be a few hours, e.g., 24 hours, several days, several years, or even the lifespan of the patient, i.e., long-term engraftment. For example, in some aspects described herein, an effective amount of photoreceptor cells or retinal progenitor cells is administered via a systemic route, such as intraperitoneally or intravenously.
[0358] In some embodiments, administering includes administering by viral delivery. A viral vector or vector containing a nucleic acid encoding an RGN polypeptide or ribonucleoprotein complex disclosed herein may be administered directly to a patient (i.e., in vivo) or used to treat cells in vitro, and the modified cells may optionally be administered to a patient (i.e., ex vivo). Conventional viral-based systems may include, but are not limited to, retroviral, lentiviral, adenoviral, adeno-associated, and herpes simplex viral vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated viral gene transfer methods, often resulting in long-term expression of the inserted transgene. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. For applications where transient expression is preferred, an adenoviral-based system may be used. Adenoviral-based vectors are capable of very high transduction efficiency in many cell types and do not require cell division.
[0359] In some embodiments, administering includes administering via other non-viral delivery of nucleic acids. Exemplary non-viral delivery methods include, but are not limited to, RNP complexes, lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, LNPs, immunoliposomes, polycation or lipid nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those described by Feigner, WO 1991 / 17424; WO 1991 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or to target tissues (e.g., in vivo administration). In some embodiments, administration of the pharmaceutical compositions of the present disclosure comprises daily intravenous injection of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 mg / kg / day or more of the active ingredient in the pharmaceutical composition comprising liposomes or LNPs. In some embodiments, administration of the pharmaceutical composition comprising liposomes or LNPs comprises a dose of about 0.01 to 1 mg per kg of body weight. In some embodiments, administration of the pharmaceutical composition comprising liposomes or LNPs comprises a dose of about 1 to 10 mg per kg of body weight.
[0360] Suitable routes of administration of the pharmaceutical compositions described herein include, but are not limited to, topical, subcutaneous, transdermal, intradermal, intralesional, intraarticular, intraperitoneal, intravesical, transmucosal, gingival, intradermal, intracochlear, intratympanic, intravisceral, epidural, intrathecal, intramuscular, intravenous, intravascular, intraosseous, periocular, intratumoral, intracerebral, and intraventricular administration.
[0361] In embodiments, the pharmaceutical compositions described herein are administered to a subject by injection, inhalation (e.g., aerosol), catheter, suppository, or implant, where the implant is a porous, non-porous, or gelatinous material, including a membrane, such as a sialastic membrane, or a fiber. In embodiments, the pharmaceutical composition is formulated for delivery to a subject, e.g., for gene editing.
[0362] In embodiments, pharmaceutical compositions are formulated according to routine procedures as compositions suitable for intravenous or subcutaneous administration to a subject, e.g., a human. In embodiments, pharmaceutical compositions for administration by injection are solutions in sterile isotonic aqueous buffer. Where necessary, the pharmaceutical product may also include a solubilizing agent and a local anesthetic, such as lignocaine, to ease pain at the injection site. Generally, the ingredients are supplied separately or mixed together, for example, as a dry lyophilized powder or water-free concentrate in a hermetically sealed container, such as an ampoule or sachet, indicating the quantity of active agent. When the agent is administered by injection, it can be dispensed in an infusion bottle containing sterile pharmaceutical-grade water or saline. When the pharmaceutical composition is administered by injection, an ampoule of sterile water for injection or saline can be provided so that the ingredients can be mixed prior to administration.
[0363] In embodiments, the pharmaceutical composition may be contained within a lipid particle or vesicle, such as a liposome or microcrystal, which may also be suitable for parenteral administration.
[0364] Although the description of pharmaceutical compositions provided herein is primarily directed to pharmaceutical compositions suitable for administration to humans, it will be understood by those skilled in the art that such compositions are generally suitable for administration to any type of animal or organism.
[0365] As used herein, the term "subject" refers to any individual for whom diagnosis, treatment, or therapy is desired. In some embodiments, the subject is an animal. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.
[0366] The effectiveness of treatment can be determined by one of ordinary skill in the art. However, treatment is considered "effective treatment" if any one or all of the signs or symptoms of a disease or disorder are altered in a beneficial manner (e.g., reduced by at least 10%), or if other clinically acceptable symptoms or markers of the disease are improved or ameliorated. Efficacy can also be measured by the individual's lack of deterioration, as assessed by the need for hospitalization or medical intervention (e.g., the progression of the disease is halted or at least slowed). Methods for measuring these indicators are known to those of ordinary skill in the art. Treatment includes (1) inhibiting the disease, e.g., halting symptoms or slowing the progression of symptoms; or (2) alleviating the disease, e.g., causing regression of symptoms; or (3) preventing or reducing the likelihood of symptoms occurring.
[0367] A. Correcting the causative mutation using base editing An example of a genetically inherited disease that can be corrected using an approach that relies on the RGN-base editor fusion protein of the present invention is Hurler syndrome. Hurler syndrome, also known as MPS-1, is the result of a deficiency in α-L-iduronidase (IDUA), resulting in a lysosomal storage disease characterized at the molecular level by the accumulation of dermatan sulfate and heparan sulfate in lysosomes. This disease is generally a hereditary genetic disorder caused by mutations in the IDUA gene, which encodes α-L-iduronidase. Common IDUA mutations are W402X and Q70X, both of which are nonsense mutations that result in premature termination of translation. Such mutations are well addressed by precise genome editing (PGE) approaches, as, for example, reversion of a single nucleotide by a base editing approach restores the wild-type coding sequence, resulting in protein expression controlled by the endogenous regulatory mechanisms of the gene locus. Furthermore, because heterozygotes are known to be asymptomatic, PGE therapy targeting one of these mutations would be useful for the majority of patients with this disease, as only one of the mutant alleles needs to be corrected (Bunge et al. (1994) Hum. Mol. Genet. 3(6):861-866, incorporated herein by reference).
[0368] Current treatments for Hurler syndrome include enzyme replacement therapy and bone marrow transplantation (Vellodi et al. (1997) Arch. Dis. Child. 76(2):92-99; Peters et al. (1998) Blood 91(7):2601-2608, incorporated herein by reference). While enzyme replacement therapy has dramatically improved the survival and quality of life of patients with Hurler syndrome, this approach requires expensive and time-consuming weekly infusions. Additional approaches include delivering the IDUA gene onto an expression vector or inserting the gene into a highly expressed locus, such as the serum albumin locus (U.S. Pat. No. 9,956,247, incorporated herein by reference). However, these approaches do not restore the original IDUA locus to the correct coding sequence. Genome editing strategies would have numerous advantages. Most notably, regulation of gene expression would be controlled by natural mechanisms present in healthy individuals. Furthermore, the use of base editing does not require the generation of double-stranded DNA breaks, which may result in large chromosomal rearrangements, cell death, or oncogene formation through the disruption of tumor suppressor mechanisms. A general strategy may be directed toward targeting and correcting specific disease-causing mutations in the human genome using a fusion protein comprising the RGN-base editor fusion protein of the invention, such as LPG10165, LPG10167, LPG10168, LPG10171, LPG10186, LPG10190, LPG10194, LPG10195, LPG10200, LPG10203, or LPG10207. It will be understood that similar approaches can also be pursued to target diseases that can be corrected by base editing. It will further be understood that similar approaches targeting disease-causing mutations in other species, particularly common household pets or livestock, can also be deployed using the RGNs of the invention. Common household pets and livestock include dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, and fish, including salmon and shrimp.
[0369] B. Correction of the causative mutation by targeted deletion The RGNs of the present invention may also be useful in human therapeutic approaches where the causative mutations are more complex. For example, some diseases, such as Friedreich's ataxia and Huntington's disease, are the result of a significant increase in the repeat of a three-nucleotide motif (i.e., "expanded trinucleotide repeat") in specific regions of the gene, affecting the ability of the expressed protein to function or be expressed. Friedreich's ataxia (FRDA) is an autosomal recessive disease that results in progressive degeneration of nerve tissue in the spinal cord. Reduced levels of frataxin (FXN) protein in mitochondria cause oxidative damage and iron deficiency at the cellular level. Reduced FXN expression is associated with a GAA triplet expansion within intron 1 of the somatic and germline FXN gene. In FRDA patients, GAA repeats often consist of more than 70 and sometimes more than 1000 (most commonly 600-900) triplets, whereas unaffected individuals have approximately 40 repeats or fewer (Pandolfo et al. (2012) Handbook of Clinical Neurology 103:275-294; Campuzano et al. (1996) Science 271:1423-1427; Pandolfo (2002) Adv. Exp. Med. Biol. 516:99-118; all incorporated herein by reference).
[0370] The expansion of the trinucleotide repeat sequence that causes Friedreich's ataxia (FRDA) occurs at a defined locus within the FXN gene, termed the FRDA instability region. RNA-guided nucleases (RGNs) can be used to excise the expanded trinucleotide repeat in FRDA patient cells. This approach requires 1) an RGN and guide RNA sequence that can be programmed to target an allele in the human genome; and 2) a delivery method for the RGN and guide sequence. Many nucleases used for genome editing, such as the commonly used Cas9 nuclease (SpCas9) from Streptococcus pyogenes (S. pyogenes), are too large to be packaged into adeno-associated virus (AAV) vectors, especially considering the length of the SpCas9 gene and guide RNA in addition to other genetic elements required for a functional expression cassette. This makes the SpCas9 approach more challenging.
[0371] Certain RNA-guided nucleases of the present invention, such as LPG10165, LPG10166, LPG10167, LPG10168, LPG10169, LPG10171, LPG10186, LPG10190, LPG10191, LPG10194, LPG10195, LPG10196, LPG10197, LPG10198, LPG10200, LPG10203, LPG10204, LPG10205, LPG10207, and LPG10208, are well suited for packaging into AAV vectors along with guide RNAs. The present invention encompasses strategies using RGNs of the present invention to remove expanded trinucleotide repeats. Such strategies are applicable to other diseases and disorders with similar genetic underpinnings, such as Huntington's disease. Similar strategies using the RGNs of the present invention may be applicable to similar diseases and disorders in agriculturally or economically important non-human animals, including dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, and fish, including salmon and shrimp.
[0372] C. Correction of the causative mutation by targeted mutagenesis The RGNs of the present invention can also introduce disruptive mutations that can have beneficial effects. Genetic defects in genes encoding hemoglobin, particularly the beta globin chain (HBB gene), can cause several diseases known as hemoglobinopathies, including sickle cell anemia and thalassemia.
[0373] In adults, hemoglobin is a heterotetramer containing two alpha (α)-like globin chains, two beta (β)-like globin chains, ...
Claims
1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20.
2. 2. The nucleic acid molecule of claim 1, wherein the RGN polypeptide is capable of binding to a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand when bound to a guide RNA (gRNA) that can hybridize to a non-target strand of the target sequence in an RNA guide sequence-specific manner.
3. 3. The nucleic acid molecule of claim 1, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide.
4. The nucleic acid molecule of any one of claims 1 to 3, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NOs: 1 to 20.
5. The nucleic acid molecule of any one of claims 1 to 4, wherein the RGN polypeptide is capable of cleaving the target nucleic acid molecule upon binding.
6. The nucleic acid molecule of claim 5 , wherein the RGN polypeptide is capable of generating a double-stranded break.
7. The nucleic acid molecule of claim 5 , wherein the RGN polypeptide is capable of generating a single-strand break.
8. The nucleic acid molecule of any one of claims 1 to 4, wherein the RGN polypeptide is nuclease inactive or a nickase.
9. The nucleic acid molecule of any one of claims 1 to 8, wherein the RGN polypeptide is operably fused to a base-editing polypeptide.
10. 10. The nucleic acid molecule of Claim 9, wherein the base-editing polypeptide is a deaminase.
11. The nucleic acid molecule of any one of claims 1 to 10, wherein the RGN polypeptide comprises one or more nuclear localization signals.
12. The nucleic acid molecule of any one of claims 1 to 11, wherein the RGN polypeptide is codon-optimized for expression in a eukaryotic cell.
13. The nucleic acid molecule of any one of claims 1 to 12, wherein the target sequence is located adjacent to a protospacer adjacent motif (PAM).
14. A vector comprising the nucleic acid molecule of any one of claims 1 to 13.
15. 15. The vector of claim 14, further comprising at least one nucleotide sequence encoding the gRNA capable of hybridizing to the non-target strand of the target sequence.
16. The guide RNA is a) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 21; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 42; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; b) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 22; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 43; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 2; c) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 23; and ii) a tracrRNA having at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 3; d) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 24; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 45; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 4; e) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO: 25 or SEQ ID NO: 1041 or 1042; and ii) a tracrRNA having at least 90% sequence identity to nucleotides 22-85 of SEQ ID NO: 46 or SEQ ID NO: 1041 or 1042; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:5; f) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 26; and ii) a tracrRNA having at least 90% sequence identity to nucleotides 24 to 138 of SEQ ID NO: 47 or SEQ ID NO: 1043; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:6; g) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045; and ii) a tracrRNA having at least 90% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 7; h) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 28; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 49; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 8; i) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 29; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 50; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 9; j) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 30; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 51; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 10; k) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 31; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 52; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 11; l) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 32; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 53; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 12; and m) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 33; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 54; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 13; n) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 34; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 55; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 14; o) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 35; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 56; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 15; p) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 36; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 57; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 16; and q) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 37; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 58; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 17; r) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 38 or 39; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 59 or 60; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 18; and s) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 40; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 61; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 19; and t) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 41; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 62; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 20; and 16. The vector of claim 15, selected from the group consisting of:
17. The guide RNA is a) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 21; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 42; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 1; b) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 22; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 43; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 2; c) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 23; and ii) a tracrRNA having at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 3; d) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 24; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 45; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 4; e) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO: 25 or SEQ ID NO: 1041 or 1042; and ii) a tracrRNA having at least 100% sequence identity to nucleotides 22-85 of SEQ ID NO: 46 or SEQ ID NO: 1041 or 1042; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO:5; f) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 26; and ii) a tracrRNA having at least 100% sequence identity to nucleotides 24 to 138 of SEQ ID NO: 47 or SEQ ID NO: 1043; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 6; g) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045; and ii) a tracrRNA having at least 100% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 7; h) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 28; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 49; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 8; i) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 29; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 50; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 9; j) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 30; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 51; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 10; k) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 31; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 52; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 11; l) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 32; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 53; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 12; m) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 33; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 54; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 13; n) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 34; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 55; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 14; o) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 35; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 56; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 15; p) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 36; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 57; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 16; q) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 37; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 58; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 17; r) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 38 or 39; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 59 or 60; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 18; s) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 40; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 61; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 19; and t) i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 100% sequence identity to SEQ ID NO: 41; and ii) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 62; wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity with SEQ ID NO: 20; The vector of claim 13, selected from the group consisting of:
18. The vector of any one of claims 14 to 17, wherein the gRNA is a single guide RNA.
19. The vector of any one of claims 15 to 17, wherein the gRNA is a dual guide RNA.
20. A cell comprising a nucleic acid molecule according to any one of claims 1 to 14 or a vector according to any one of claims 14 to 19.
21. A plant comprising the cell of claim 20.
22. A seed comprising the cells of claim 20.
23. A method for producing an RGN polypeptide, comprising culturing the cell of claim 20 under conditions in which the RGN polypeptide is expressed.
24. Introducing into a cell a heterologous nucleic acid molecule comprising a nucleotide sequence encoding an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20; culturing the cells under conditions in which the RGN polypeptide is expressed; A method for producing an RGN polypeptide, comprising:
25. 25. The method of claim 24, wherein the RGN polypeptide is capable of binding to a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand when bound to a guide RNA (gRNA) that is capable of hybridizing to the non-target strand of the target sequence, in an RNA guide sequence-specific manner.
26. The method of claim 24 or 25, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity to any one of SEQ ID NOs: 1 to 20.
27. The method of any one of claims 24 to 26, further comprising purifying the RGN polypeptide.
28. 28. The method of any one of claims 24 to 27, wherein the cell further expresses one or more guide RNAs capable of binding to the RGN polypeptide to form an RGN ribonucleoprotein complex.
29. 29. The method of claim 28, further comprising purifying the RGN ribonucleoprotein complex.
30. 1. An RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20.
31. 32. The RGN polypeptide of claim 31 , wherein the RGN polypeptide is capable of binding to a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand when bound to a guide RNA (gRNA) that can hybridize to the non-target strand of the target sequence in an RNA guide sequence-specific manner.
32. The RGN polypeptide of claim 30 or 31, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NOs: 1 to 20.
33. 33. The RGN polypeptide of any one of claims 30 to 32, wherein the RGN polypeptide is capable of cleaving the target nucleic acid molecule upon binding.
34. 34. The RGN polypeptide of claim 33, wherein cleavage by the RGN polypeptide generates a double-stranded break.
35. 34. The RGN polypeptide of claim 33, wherein cleavage by the RGN polypeptide generates a single-strand break.
36. The RGN polypeptide of any one of claims 30 to 32, wherein the RGN polypeptide is nuclease inactive or a nickase.
37. The RGN polypeptide of any one of claims 30 to 36, wherein the RGN polypeptide is operably fused to a base-editing polypeptide.
38. The RGN polypeptide of Claim 37, wherein the base-editing polypeptide is a deaminase.
39. The RGN polypeptide of any one of claims 30 to 38, wherein the target sequence is located adjacent to a protospacer adjacent motif (PAM).
40. The RGN polypeptide of any one of claims 28 to 36, wherein the RGN polypeptide comprises one or more nuclear localization signals.
41. A ribonucleoprotein (RNP) complex comprising an RGN polypeptide according to any one of claims 30 to 40 and a guide RNA bound to the RGN polypeptide.
42. 1. A nucleic acid molecule comprising a CRISPR RNA (crRNA) or a polynucleotide encoding the crRNA, wherein the crRNA comprises a spacer and a CRISPR repeat, and the CRISPR repeat comprises a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 21-41, or nucleotides 1-17 of SEQ ID NO: 1041 or 1042, or nucleotides 1-22 of SEQ ID NO: 1044 or 1045.
43. a) the crRNA; and b) a trans-activating CRISPR RNA (tracrRNA) hybridized to the CRISPR repeats of the crRNA; can hybridize to a non-target strand of a target sequence in a target nucleic acid molecule in a sequence-specific manner via the spacer of the crRNA when the guide RNA is bound to an RNA-guided nuclease (RGN) polypeptide, 43. The nucleic acid molecule of claim 42.
44. 44. The nucleic acid molecule of claim 42 or 43, wherein the polynucleotide encoding the crRNA is operably linked to a promoter heterologous to the polynucleotide.
45. 45. The nucleic acid molecule of any one of claims 42 to 44, wherein the CRISPR repeat comprises a nucleotide sequence having 100% sequence identity to any one of SEQ ID NOs: 21 to 41, or nucleotides 1 to 17 of SEQ ID NO: 1041 or 1042, or nucleotides 1 to 22 of SEQ ID NO: 1044 or 1045.
46. A vector comprising a nucleic acid molecule comprising the polynucleotide encoding the crRNA of any one of claims 42 to 45.
47. 47. The vector of claim 46, wherein the vector further comprises a polynucleotide encoding the tracrRNA.
48. The tracrRNA is a) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 42, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 21; b) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 43, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 22; c) a tracrRNA having at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:23; d) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 45, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 24; e) a tracrRNA having at least 90% sequence identity to nucleotides 22-85 of SEQ ID NO: 46 or SEQ ID NO: 1041 or 1042, wherein the CRISPR repeat has at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO: 25 or SEQ ID NO: 1041 or 1042; f) a tracrRNA having at least 90% sequence identity to nucleotides 24-138 of SEQ ID NO: 47 or SEQ ID NO: 1043, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 26; g) a tracrRNA having at least 90% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045, wherein the CRISPR repeat has at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045; h) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 49, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 28; i) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 50, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 29; j) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 51, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 30; k) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 52, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 31; l) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 53, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 32; m) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 54, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 33; n) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 55, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 34; o) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 56, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 35; p) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 57, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 36; q) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 58, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 37; r) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 59 or 60, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 38 or 39; s) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 61, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 40; and t) A tracrRNA having at least 90% sequence identity to SEQ ID NO: 62, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:
41.
48. The vector of claim 47, selected from the group consisting of:
49. The tracrRNA is a) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 42, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 21; b) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 43, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 22; c) a tracrRNA having at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:23; d) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 45, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 24; e) a tracrRNA having at least 100% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042, wherein the CRISPR repeat has at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042; f) a tracrRNA having at least 100% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:26; g) a tracrRNA having at least 100% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045, wherein the CRISPR repeat has at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045; h) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 49, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 28; i) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 50, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 29; j) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 51, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 30; k) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 52, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 31; l) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 53, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 32; m) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 54, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 33; n) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 55, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 34; o) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 56, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 35; p) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 57, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 36; q) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 58, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 37; r) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 59 or 60, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 38 or 39; s) a tracrRNA having at least 100% sequence identity to SEQ ID NO: 61, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 40; and t) A tracrRNA having at least 100% sequence identity with SEQ ID NO: 62, wherein the CRISPR repeat has at least 100% sequence identity with SEQ ID NO:
41.
48. The vector of claim 47, selected from the group consisting of:
50. The vector of any one of claims 42 to 49, wherein the vector further comprises a polynucleotide encoding the RGN polypeptide.
51. The RGN polypeptide is a) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 1, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 21, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 42; b) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:2, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:22, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:43; c) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:3, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:23, and the tracrRNA has at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; d) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:4, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:24, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:45; e) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:5, wherein the CRISPR repeat has at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:46 or nucleotides 22-85 of SEQ ID NO:1041 or 1042; f) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:6, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:26, and the tracrRNA has at least 90% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043; g) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:7, wherein the CRISPR repeat has at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:48 or nucleotides 27-96 of SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; h) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:8, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:28, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:49; i) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:9, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO:29, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:50; j) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 10, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 30, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 51; k) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 11, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 31, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 52; l) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 12, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 32, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 53; m) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 13, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 33, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 54; n) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 14, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 34, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 55; o) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 15, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 35, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 56; p) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 16, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 36, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 57; q) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 17, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 37, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 58; r) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 18, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 38 or 39, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 59 or 60; s) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 19, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 40, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 61; and t) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 20, wherein the CRISPR repeat has at least 90% sequence identity to SEQ ID NO: 41, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 62; 51. The vector of claim 50, selected from the group consisting of:
52. The RGN polypeptide is a) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 1, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 21, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 42; b) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:2, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:22, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:43; c) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:3, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:23, and the tracrRNA has at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; d) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:4, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:24, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:45; e) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:5, wherein the CRISPR repeat has at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:46 or nucleotides 22-85 of SEQ ID NO:1041 or 1042; f) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:6, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:26, and the tracrRNA has at least 100% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043; g) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:7, wherein the CRISPR repeat has at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:48 or nucleotides 27-96 of SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; h) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:8, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:28, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:49; i) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:9, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO:29 and the tracrRNA has at least 100% sequence identity to SEQ ID NO:50; j) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 10, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 30, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 51; k) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 11, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 31, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 52; l) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 12, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 32, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 53; m) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 13, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 33, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 54; n) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 14, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 34, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 55; o) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 15, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 35, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 56; p) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 16, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 36, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 57; q) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 17, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 37, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 58; r) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 18, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 38 or 39, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 59 or 60; s) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 19, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 40, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 61; and t) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 20, wherein the CRISPR repeat has at least 100% sequence identity to SEQ ID NO: 41, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 62; 51. The vector of claim 50, selected from the group consisting of:
53. A nucleic acid molecule comprising a polynucleotide encoding a trans-activating CRISPR RNA (tracrRNA) or a tracrRNA comprising a nucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs:42-62, nucleotides 19-111 of SEQ ID NO:1040, nucleotides 22-85 of SEQ ID NOs:1041 or 1042, nucleotides 24-138 of SEQ ID NO:143, nucleotides 27-96 of SEQ ID NO:1044, or nucleotides 27-95 of SEQ ID NO:1045.
54. a) the tracrRNA; and b) a crRNA comprising a spacer and a CRISPR repeat, wherein the tracrRNA hybridizes to the CRISPR repeat of the crRNA; can hybridize to a non-target strand of a target sequence in a target nucleic acid molecule in a sequence-specific manner via the spacer of the crRNA when the guide RNA is bound to an RNA-guided nuclease (RGN) polypeptide, 54. The nucleic acid molecule of claim 53.
55. 55. The nucleic acid molecule of Claim 53 or 54, wherein the polynucleotide encoding the tracrRNA is operably linked to a promoter heterologous to the polynucleotide.
56. 56. The nucleic acid molecule of any one of claims 53-55, wherein the tracrRNA comprises a nucleotide sequence having 100% sequence identity to any one of SEQ ID NOs:42-62, nucleotides 19-111 of SEQ ID NO:1040, nucleotides 22-85 of SEQ ID NOs:1041 or 1042, nucleotides 24-138 of SEQ ID NO:143, nucleotides 27-96 of SEQ ID NO:1044, or nucleotides 27-95 of SEQ ID NO:1045.
57. 57. A vector comprising a nucleic acid molecule comprising the polynucleotide encoding the tracrRNA of any one of claims 53 to 56.
58. 58. The vector of claim 57, wherein the vector further comprises a polynucleotide encoding the crRNA.
59. The crRNA a) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 21, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 42; b) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 22, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 43; c) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 23, wherein the tracrRNA has at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO: 44 or SEQ ID NO: 1040; d) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 24, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 45; e) a CRISPR repeat having at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, wherein the tracrRNA has at least 90% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042; f) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 26, wherein the tracrRNA has at least 90% sequence identity to nucleotides 24 to 138 of SEQ ID NO: 47 or SEQ ID NO: 1043; g) a CRISPR repeat having at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, wherein the tracrRNA has at least 90% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; h) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 28, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 49; i) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 29, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 50; j) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 30, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 51; k) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 31, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 52; l) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 32, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 53; m) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 33, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 54; n) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 34, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 55; o) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 35, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 56; p) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 36, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 57; q) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 37, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 58; r) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 38 or 39, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 59 or 60; s) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 40, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 61; and t) a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 41, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 62; 59. The vector of claim 58, comprising a CRISPR repeat selected from the group consisting of:
60. The crRNA a) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 21, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 42; b) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 22, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 43; c) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 23, wherein the tracrRNA has at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO: 44 or SEQ ID NO: 1040; d) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 24, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 45; e) a CRISPR repeat having at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, wherein the tracrRNA has at least 100% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042; f) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 26, wherein the tracrRNA has at least 100% sequence identity to nucleotides 24 to 138 of SEQ ID NO: 47 or SEQ ID NO: 1043; g) a CRISPR repeat having at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, wherein the tracrRNA has at least 100% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; h) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 28, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 49; i) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 29, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 50; j) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 30, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 51; k) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 31, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 52; l) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 32, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 53; m) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 33, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 54; n) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 34, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 55; o) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 35, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 56; p) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 36, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 57; q) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 37, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 58; r) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 38 or 39, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 59 or 60; s) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 40, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 61; and t) a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 41, wherein the tracrRNA has at least 100% sequence identity to SEQ ID NO: 62; 59. The vector of claim 58, comprising a CRISPR repeat selected from the group consisting of:
61. The vector of any one of claims 53 to 60, wherein the vector further comprises a polynucleotide encoding the RGN polypeptide.
62. The RGN polypeptide is a) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 1, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 21, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 42; b) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:2, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:22, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:43; c) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:3, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:23, and the tracrRNA has at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; d) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:4, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:24, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:45; e) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:5, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:46 or nucleotides 22-85 of SEQ ID NO:1041 or 1042; f) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:6, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:26, and the tracrRNA has at least 90% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043; g) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:7, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:48 or nucleotides 27-96 of SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; h) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:8, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:28, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:49; i) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:9, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:29, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:50; j) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 10, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 30, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 51; k) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 11, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 31, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 52; l) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 12, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 32, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 53; m) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 13, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 33, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 54; n) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 14, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 34, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 55; o) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 15, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 35, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 56; p) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 16, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 36, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 57; q) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 17, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 37, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 58; r) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 18, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 38 or 39, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 59 or 60; s) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 19, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 40, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 61; and t) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 20, wherein the crRNA comprises a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 41, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 62; 62. The vector of claim 61, selected from the group consisting of:
63. The RGN polypeptide is a) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 1, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 21, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 42; b) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:2, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:22, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:43; c) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:3, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:23, and the tracrRNA has at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; d) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:4, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:24, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:45; e) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:5, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:46 or nucleotides 22-85 of SEQ ID NO:1041 or 1042; f) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:6, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:26, and the tracrRNA has at least 100% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043; g) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:7, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:48 or nucleotides 27-96 of SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045; h) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:8, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:28, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:49; i) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO:9, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:29, and the tracrRNA has at least 100% sequence identity to SEQ ID NO:50; j) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 10, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 30, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 51; k) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 11, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 31, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 52; l) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 12, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 32, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 53; m) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 13, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 33, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 54; n) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 14, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 34, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 55; o) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 15, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 35, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 56; p) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 16, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 36, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 57; q) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 17, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 37, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 58; r) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 18, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 38 or 39, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 59 or 60; s) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 19, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 40, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 61; and t) an RGN polypeptide having at least 100% sequence identity to SEQ ID NO: 20, wherein the crRNA comprises a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 41, and the tracrRNA has at least 100% sequence identity to SEQ ID NO: 62; 62. The vector of claim 61, selected from the group consisting of:
64. 142. A cell comprising the nucleic acid molecule of any one of claims 42 to 45 and 53 to 56, the vector of any one of claims 46 to 52 and 57 to 63, the single guide RNA of claim 142, or the dual guide RNA of claim 143.
65. 65. A plant comprising the cell of claim 64.
66. 65. A seed comprising the cell of claim 64.
67. 1. A system for binding to a target sequence within a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand, the system comprising: a) one or more guide RNAs capable of hybridizing to the non-target strand of the target sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1 to 20, or a polynucleotide comprising a nucleotide sequence encoding said RGN polypeptide; Including, the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to bind the RGN polypeptide to the target sequence; A system for binding to a target sequence within a target nucleic acid molecule.
68. 68. The system of Claim 67, wherein at least one of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to the nucleotide sequence.
69. 1. A system for binding to a target sequence within a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand, the system comprising: a) one or more guide RNAs capable of hybridizing to the non-target strand of the target sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20; Including, the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to bind the RGN polypeptide to the target sequence; A system for binding to a target sequence within a target nucleic acid molecule.
70. 70. The system of any one of claims 67 to 69, wherein at least one of the nucleotide sequences encoding the one or more guide RNAs is operably linked to a promoter heterologous to said nucleotide sequence.
71. The system of any one of claims 67 to 70, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NOs: 1 to 20.
72. 72. The system of any one of claims 67 to 71, wherein the RGN polypeptide and the one or more guide RNAs are not found complexed to each other in nature.
73. 73. The system of any one of claims 67 to 72, wherein the target sequence is a eukaryotic target sequence.
74. 74. The system of any one of claims 67 to 73, wherein the gRNA is a single guide RNA (sgRNA).
75. 74. The system of any one of claims 67 to 73, wherein the gRNA is a dual guide RNA.
76. The gRNA a) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:42, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:1; b) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:43, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2; c) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:3; d) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:45, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:4; e) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:46 or nucleotides 22-85 of SEQ ID NO:1041 or 1042, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:5; f) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:26 and a tracrRNA having at least 90% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:6; g) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and a tracrRNA having at least 90% sequence identity to SEQ ID NO:48 or nucleotides 27-96 of SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:7; h) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:49, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:8; i) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:50, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:9; j) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 30 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 51, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 10; k) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 31 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 11; l) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 32 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 53, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 12; m) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 54, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 13; n) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 55, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 14; o) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 35 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 56, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 15; p) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 36 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 57, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 16; q) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 37 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 58, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 17; r) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 38 or 39 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 59 or 60, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 18; s) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 40 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 61, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 19; and t) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 41 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 62, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 20; 76. The system of any one of claims 67 to 75, selected from the group consisting of:
77. The gRNA a) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:42, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:1; b) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:43, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:2; c) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:3; d) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:45, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:4; e) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and a tracrRNA having at least 100% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:5; f) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:26 and a tracrRNA having at least 100% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:6; g) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and a tracrRNA having at least 100% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:7; h) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:49, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:8; i) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:50, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO:9; j) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 30 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 51, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 10; k) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 31 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 11; l) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 32 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 53, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 12; m) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 54, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 13; n) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 55, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 14; o) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 35 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 56, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 15; p) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 36 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 57, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 16; q) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 37 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 58, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 17; r) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 38 or 39 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 59 or 60, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 18; s) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 40 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 61, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 19; and t) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 41 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 62, wherein the RGN polypeptide comprises an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 20; 76. The system of any one of claims 67 to 75, selected from the group consisting of:
78. 78. The system of any one of claims 67 to 77, wherein the target sequence is located adjacent to a protospacer adjacent motif (PAM).
79. The system of any one of claims 67 to 78, wherein the target sequence is intracellular.
80. 82. The system of any one of claims 67-81, wherein the one or more guide RNAs are capable of hybridizing to the non-target strand of the target sequence, and the guide RNAs are capable of forming a complex with the RGN polypeptide to direct cleavage of the target nucleic acid molecule.
81. 81. The system of claim 80, wherein the cleavage generates a double-stranded break.
82. 81. The system of claim 80, wherein the cleavage generates a single-stranded break.
83. The system of any one of claims 67 to 79, wherein the RGN polypeptide is nuclease inactive or is a nickase.
84. The system of any one of claims 67 to 83, wherein the RGN polypeptide is operably linked to a base-editing polypeptide.
85. 85. The system of Claim 84, wherein the base-editing polypeptide is a deaminase.
86. The system of any one of claims 67 to 85, wherein the RGN polypeptide comprises one or more nuclear localization signals.
87. The system of any one of claims 67 to 86, wherein the RGN polypeptide is codon-optimized for expression in eukaryotic cells.
88. The system of any one of claims 67 to 87, wherein the system further comprises one or more donor polynucleotides.
89. A cell comprising the system according to any one of claims 67 to 88.
90. 90. A plant comprising the cell of claim 89.
91. 90. A seed comprising the cell of claim 89.
92. A pharmaceutical composition comprising a nucleic acid molecule according to any one of claims 1 to 13, 42 to 45, and 53 to 56, a vector according to any one of claims 14 to 19, 46 to 52, and 57 to 63, a cell according to any one of claims 20, 64, and 89, an RGN polypeptide according to any one of claims 30 to 40, an RNP complex according to claim 41, or a system according to any one of claims 67 to 88, and a pharmaceutically acceptable carrier.
93. 90. A method for binding to a target sequence in a target nucleic acid molecule, the method comprising delivering a system according to any one of claims 67 to 88 to the target sequence or to a cell containing the target sequence.
94. 94. The method of claim 93, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby allowing detection of the target sequence.
95. 94. The method of claim 93, wherein the guide RNA or the RGN polypeptide further comprises an expression modulator, thereby regulating expression of a target gene comprising the target sequence.
96. 90. A method for cleaving and / or modifying a target nucleic acid molecule comprising a target sequence, said method comprising delivering a system according to any one of claims 67 to 88 to said target sequence or to a cell comprising said target sequence, wherein cleavage or modification of said target nucleic acid molecule occurs.
97. 97. The method of claim 96, wherein the modified target nucleic acid molecule comprises an insertion of heterologous DNA into the target nucleic acid molecule.
98. 97. The method of claim 96, wherein the modified target nucleic acid molecule comprises a deletion of at least one nucleotide from the target nucleic acid molecule.
99. 97. The method of claim 96, wherein the modified target nucleic acid molecule comprises a mutation of at least one nucleotide in the target nucleic acid molecule.
100. 1. A method for binding a target sequence in a target nucleic acid molecule, the target sequence comprising a target strand and a non-target strand, the method comprising: a) RNA-guided nuclease (RGN) ribonucleotide complexes, i) one or more guide RNAs capable of hybridizing to the non-target strand of the target sequence; and ii) an RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20; under conditions suitable for the formation of the RGN ribonucleotide complex; b) contacting the target nucleic acid molecule or a cell containing the target nucleic acid molecule with the assembled RGN ribonucleotide complex; wherein the one or more guide RNAs hybridize to the non-target strand of the target sequence, thereby directing the RGN polypeptide to bind to the target sequence. A method for binding a target sequence in a target nucleic acid molecule.
101. 101. The method of claim 100, wherein the method is performed in vitro, in vivo or ex vivo.
102. 102. The method of claim 100 or 101, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby allowing detection of the target sequence.
103. 102. The method of claim 100 or 101, wherein the guide RNA or the RGN polypeptide further comprises an expression modulator, thereby allowing regulation of expression of a target gene comprising the target sequence.
104. 102. The method of Claim 100 or 101, wherein the RGN polypeptide further comprises a base-editing polypeptide, thereby enabling modification of the target nucleic acid molecule.
105. 105. The method of Claim 104, wherein the base-editing polypeptide comprises a deaminase.
106. 102. The method of claim 100 or 101, wherein the RGN polypeptide is capable of cleaving the target nucleic acid molecule, thereby allowing for cleavage and / or modification of the target nucleic acid molecule.
107. 1. A method of cleaving and / or modifying a target nucleic acid molecule comprising a target sequence, the target sequence comprising a target strand and a non-target strand, the method comprising: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20; and b) one or more guide RNAs capable of targeting the RGN of (a) to the target sequence; contacting the the one or more guide RNAs hybridize to the non-target strand of the target sequence, thereby directing the RGN polypeptide to bind to the target nucleic acid molecule and causing cleavage and / or modification of the target nucleic acid molecule. A method for cleaving and / or modifying a target nucleic acid molecule containing a target sequence.
108. The method of claim 107, wherein cleavage by the RGN polypeptide generates a double-stranded break.
109. The method of claim 107, wherein cleavage by the RGN polypeptide results in a single-strand break.
110. 108. The method of Claim 107, wherein the RGN polypeptide is nuclease inactive or nickase and is operably fused to a base-editing polypeptide.
111. 111. The method of Claim 110, wherein the base-editing polypeptide is a deaminase.
112. 108. The method of claim 107, wherein the modified target nucleic acid molecule comprises an insertion of heterologous DNA into the target nucleic acid molecule.
113. 108. The method of Claim 107, wherein the modified target nucleic acid molecule comprises a deletion of at least one nucleotide from the target nucleic acid molecule.
114. 108. The method of Claim 107, wherein the modified target nucleic acid molecule comprises a mutation of at least one nucleotide in the target nucleic acid molecule.
115. 114. The method of any one of claims 106 to 113, wherein the target sequence is located adjacent to a protospacer adjacent motif (PAM).
116. 116. The method of any one of claims 107 to 115, wherein the target sequence is a eukaryotic target sequence.
117. 117. The method of any one of claims 107 to 116, wherein the gRNA is a single guide RNA (sgRNA).
118. 117. The method of any one of claims 107 to 116, wherein the gRNA is a dual guide RNA.
119. The method of any one of claims 107 to 118, wherein the RGN comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NOs: 1 to 20.
120. a) the RGN has at least 90% sequence identity to SEQ ID NO: 1, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 21 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 42; b) the RGN has at least 90% sequence identity to SEQ ID NO: 2, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 22 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 43; c) the RGN has at least 90% sequence identity to SEQ ID NO:3, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; d) the RGN has at least 90% sequence identity to SEQ ID NO: 4, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 24 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 45; e) the RGN has at least 90% sequence identity to SEQ ID NO:5, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and a tracrRNA having at least 90% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042; f) the RGN has at least 90% sequence identity to SEQ ID NO: 6, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 26 and a tracrRNA having at least 90% sequence identity to nucleotides 24-138 of SEQ ID NO: 47 or SEQ ID NO: 1043; g) the RGN has at least 90% sequence identity to SEQ ID NO: 7, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045, and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 48 or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045; h) the RGN has at least 90% sequence identity to SEQ ID NO: 8, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 28 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 49; i) the RGN has at least 90% sequence identity to SEQ ID NO: 9, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 29 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 50; j) the RGN has at least 90% sequence identity to SEQ ID NO: 10, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 30 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 51; k) the RGN has at least 90% sequence identity with SEQ ID NO: 11, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 31 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 52; l) the RGN has at least 90% sequence identity with SEQ ID NO: 12, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 32 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 53; m) the RGN has at least 90% sequence identity with SEQ ID NO: 13, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 33 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 54; n) the RGN has at least 90% sequence identity with SEQ ID NO: 14, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 34 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 55; o) the RGN has at least 90% sequence identity with SEQ ID NO: 15, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 35 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 56; p) the RGN has at least 90% sequence identity with SEQ ID NO: 16, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 36 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 57; q) the RGN has at least 90% sequence identity to SEQ ID NO: 17, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 37 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 58; r) the RGN has at least 90% sequence identity with SEQ ID NO: 18, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 38 or 39 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 59 or 60; s) the RGN has at least 90% sequence identity to SEQ ID NO: 19, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity to SEQ ID NO: 40 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 61; and t) the RGN has at least 90% sequence identity with SEQ ID NO: 20, and the guide RNA comprises a crRNA repeat having at least 90% sequence identity with SEQ ID NO: 41 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 62; 119. The method of any one of claims 107 to 118.
121. a) the RGN has at least 100% sequence identity to SEQ ID NO: 1, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 21 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 42; b) the RGN has at least 100% sequence identity to SEQ ID NO: 2, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 22 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 43; c) the RGN has at least 100% sequence identity to SEQ ID NO:3, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040; d) the RGN has at least 100% sequence identity to SEQ ID NO: 4, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 24 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 45; e) the RGN has at least 100% sequence identity to SEQ ID NO:5, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042, and a tracrRNA having at least 100% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042; f) the RGN has at least 100% sequence identity to SEQ ID NO: 6, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 26 and a tracrRNA having at least 100% sequence identity to nucleotides 24-138 of SEQ ID NO: 47 or SEQ ID NO: 1043; g) the RGN has at least 100% sequence identity to SEQ ID NO: 7, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO: 27 or SEQ ID NO: 1044 or 1045, and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 48 or nucleotides 27-96 of SEQ ID NO: 1044 or nucleotides 27-95 of SEQ ID NO: 1045; h) the RGN has at least 100% sequence identity to SEQ ID NO: 8, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 28 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 49; i) the RGN has at least 100% sequence identity to SEQ ID NO: 9, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 29 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 50; j) the RGN has at least 100% sequence identity to SEQ ID NO: 10, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 30 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 51; k) the RGN has at least 100% sequence identity with SEQ ID NO: 11, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 31 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 52; l) the RGN has at least 100% sequence identity with SEQ ID NO: 12, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 32 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 53; m) the RGN has at least 100% sequence identity with SEQ ID NO: 13, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 33 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 54; n) the RGN has at least 100% sequence identity with SEQ ID NO: 14, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 34 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 55; o) the RGN has at least 100% sequence identity with SEQ ID NO: 15, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 35 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 56; p) the RGN has at least 100% sequence identity with SEQ ID NO: 16, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 36 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 57; q) the RGN has at least 100% sequence identity with SEQ ID NO: 17, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 37 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 58; r) the RGN has at least 100% sequence identity with SEQ ID NO: 18, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 38 or 39 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 59 or 60; s) the RGN has at least 100% sequence identity to SEQ ID NO: 19, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity to SEQ ID NO: 40 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 61; and t) the RGN has at least 100% sequence identity with SEQ ID NO: 20, and the guide RNA comprises a crRNA repeat having at least 100% sequence identity with SEQ ID NO: 41 and a tracrRNA having at least 100% sequence identity with SEQ ID NO: 62; 119. The method of any one of claims 107 to 118.
122. 122. The method of any one of claims 107 to 121, wherein the target sequence is intracellular.
123. 123. The method of claim 122, further comprising culturing the cells under conditions in which the RGN polypeptide is expressed and cleaves and modifies the target nucleic acid molecule to produce a DNA molecule comprising the modified target nucleic acid molecule, and selecting for cells comprising the modified target nucleic acid molecule.
124. A cell comprising a modified target nucleic acid molecule according to the method of claim 123.
125. 125. A plant comprising the cell of claim 124.
126. 125. A seed comprising the cell of claim 124.
127. A pharmaceutical composition comprising the cells of claim 124 and a pharmaceutically acceptable carrier.
128. A method for producing a genetically modified cell in which a causative mutation of a genetic disease has been corrected, comprising the steps of: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20, or a polynucleotide encoding the RGN polypeptide, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter to allow expression of the RGN polypeptide in the cell; and b) a guide RNA (gRNA) or a polynucleotide encoding said gRNA, wherein said polynucleotide encoding said gRNA is operably linked to a promoter to allow expression of said gRNA in said cell; into the cells, A method for producing a genetically modified cell in which a causative mutation of a genetic disease has been corrected, whereby the RGN and gRNA target the genomic location of the causative mutation and modify the genomic sequence to remove the causative mutation.
129. The method of claim 128, wherein the RGN is fused to a polypeptide that is nuclease inactive or nickase and has base editing activity.
130. 130. The method of Claim 129, wherein the base-editing polypeptide is a deaminase.
131. The method of any one of claims 128 to 129, wherein the genetic disease is caused by a single nucleotide polymorphism.
132. 132. The method of Claim 131, wherein said gRNA further comprises a spacer that targets a region proximal to said causative single nucleotide polymorphism.
133. 1. A method for generating a genetically engineered cell having a deletion in a disease-causing expanded trinucleotide repeat, comprising: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-20, or a polynucleotide encoding the RGN polypeptide, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter to allow expression of the RGN polypeptide in the cell; and b) a first guide RNA (gRNA) or a polynucleotide encoding said gRNA, wherein said polynucleotide encoding said gRNA is operably linked to a promoter to allow expression of said gRNA in said cell, and further wherein said gRNA comprises a spacer that targets the 5' flank of said extended trinucleotide repeat; and c) a second guide RNA (gRNA) or a polynucleotide encoding said gRNA, wherein said polynucleotide encoding said gRNA is operably linked to a promoter to allow expression of said gRNA in said cell, and further wherein said second gRNA comprises a spacer that targets the 3' flank of said extended trinucleotide repeat; into the cells, Thereby, the RGN and the two gRNAs target the extended trinucleotide repeat, and at least a portion of the extended trinucleotide repeat is removed. A method for generating genetically engineered cells that have a deletion in a disease-causing expanded trinucleotide repeat.
134. 134. The method of Claim 133, wherein the first gRNA further comprises a spacer that targets a region within or proximal to the extended trinucleotide repeat.
135. 135. The method of Claim 134, wherein the second gRNA further comprises a spacer that targets a region within or proximal to the extended trinucleotide repeat.
136. The method of any one of claims 133 to 135, wherein the RGN polypeptide has 100% sequence identity to any one of SEQ ID NOs: 1 to 20.
137. The gRNA, the first gRNA, the second gRNA, or the first gRNA and the second gRNA, a) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:42, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:1; b) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:43, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:2; c) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 90% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:3; d) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:45, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:4; e) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042 and a tracrRNA having at least 90% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:5; f) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:26 and a tracrRNA having at least 90% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:6; g) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and a tracrRNA having at least 90% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:7; h) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:49, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:8; i) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 90% sequence identity to SEQ ID NO:50, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO:9; j) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 30 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 51, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 10; k) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 31 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 52, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 11; l) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 32 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 53, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 12; m) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 54, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 13; n) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 55, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 14; o) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 35 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 56, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 15; p) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 36 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 57, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 16; q) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 37 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 58, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 17; r) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 38 or 39 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 58 or 60, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 18; s) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 40 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 61, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 19; and t) a gRNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 41 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 62, wherein the RGN polypeptide has an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 20; The method of any one of claims 133 to 136, wherein the gRNA is selected from the group consisting of:
138. The gRNA, the first gRNA, the second gRNA, or the first gRNA and the second gRNA, a) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:21 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:42, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:1; b) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:22 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:43, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:2; c) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:23 and a tracrRNA having at least 100% sequence identity to nucleotides 19-111 of SEQ ID NO:44 or SEQ ID NO:1040, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:3; d) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:24 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:45, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:4; e) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to nucleotides 1-17 of SEQ ID NO:25 or SEQ ID NO:1041 or 1042 and a tracrRNA having at least 100% sequence identity to nucleotides 22-85 of SEQ ID NO:46 or SEQ ID NO:1041 or 1042, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:5; f) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:26 and a tracrRNA having at least 100% sequence identity to nucleotides 24-138 of SEQ ID NO:47 or SEQ ID NO:1043, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:6; g) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to nucleotides 1-22 of SEQ ID NO:27 or SEQ ID NO:1044 or 1045, and a tracrRNA having at least 100% sequence identity to nucleotides 27-96 of SEQ ID NO:48 or SEQ ID NO:1044 or nucleotides 27-95 of SEQ ID NO:1045, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:7; h) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:28 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:49, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:8; i) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO:29 and a tracrRNA having at least 100% sequence identity to SEQ ID NO:50, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO:9; j) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 30 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 51, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 10; k) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 31 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 52, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 11; l) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 32 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 53, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 12; m) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 33 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 54, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 13; n) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 34 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 55, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 14; o) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 35 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 56, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 15; p) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 36 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 57, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 16; q) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 37 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 58, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 17; r) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 38 or 39 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 58 or 60, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 18; s) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 40 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 61, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 19; and t) a gRNA comprising a CRISPR repeat having at least 100% sequence identity to SEQ ID NO: 41 and a tracrRNA having at least 100% sequence identity to SEQ ID NO: 62, wherein the RGN polypeptide has an amino acid sequence having at least 100% sequence identity to SEQ ID NO: 20; The method of any one of claims 133 to 136, wherein the gRNA is selected from the group consisting of:
139. 130. A method of treating a disease, disorder or condition, comprising administering to a subject in need thereof the pharmaceutical composition of claim 92 or 127.
140. 140. The method of claim 139, wherein the disease, disorder or condition is associated with a causative mutation and the pharmaceutical composition corrects the causative mutation.
141. 141. The method of claim 139 or 140, wherein the subject is at risk of developing the disease, disorder or condition.
142. A single guide RNA comprising the nucleic acid molecule comprising the crRNA of any one of claims 42 to 45 and the nucleic acid molecule comprising the tracrRNA of any one of claims 53 to 57.
143. A dual guide RNA comprising a nucleic acid molecule comprising the crRNA of any one of claims 42 to 45 and a nucleic acid molecule comprising the tracrRNA of any one of claims 53 to 57.