RNA-guided nucleases and active fragments and variants thereof, and methods of use

JP2025086912A5Inactive Publication Date: 2025-10-30LIFEEDIT THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025025892
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-26
Filing Date
2025-02-20
Publication Date
2025-10-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current gene editing technologies, such as CRISPR-Cas systems, require complex nucleases and guide RNAs to target specific genomic sequences, which can be costly and inefficient for each target sequence.

Method used

The development of RNA-guided nuclease (RGN) compositions and methods that include RGN polypeptides, CRISPR RNAs (crRNAs), trans-activating CRISPR RNAs (tracrRNAs), guide RNAs (gRNAs), and vectors/host cells encoding these molecules, enabling targeted cleavage, modification, or detection of specific DNA sequences.

Benefits of technology

These RGN systems allow for efficient and cost-effective targeting of specific genomic sequences, enabling precise modifications or detections through mechanisms like non-homologous end joining or homology-directed repair.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

To provide RNA-guided nucleases and active fragments and variants thereof, and methods of use.SOLUTION: Compositions and methods for binding to a target sequence of interest are provided. The compositions find use in cleaving or modifying a target sequence of interest, visualizing the target sequence of interest, and modifying the expression of the sequence of interest. Compositions comprise RNA-guided nuclease polypeptides, CRISPR RNAs, trans-activating CRISPR RNAs, guide RNAs, and nucleic acid molecules encoding the same. Vectors and host cells comprising the nucleic acid molecules are also provided. Further there are provided CRISPR systems for binding a target sequence of interest, wherein the CRISPR system comprises an RNA-guided nuclease polypeptide and one or more guide RNAs. Methods and kits for detecting a target DNA sequence are also provided.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of molecular biology and gene editing.

[0002] Reference to a sequence listing submitted as a text file via EFS-WEB This application includes a sequence listing submitted in ASCII format via EFS-Web, which is hereby incorporated by reference in its entirety. This ASCII copy, created on August 11, 2020, is L103438 1170WO 0049 2 Seq named List.txt, and is 489,326 bytes in size.

Background Art

[0003] Editing or modifying a target genome is becoming an important tool for basic and applied research. The first method is to engineer nucleases such as meganucleases, zinc finger fusion proteins or TALENs, which required the generation of chimeric nucleases with programmable sequence-specific DNA-binding domains that are specific for each particular target sequence. RNA-guided nucleases such as the Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) proteins of the CRISPR-Cas bacterial system enable targeting of specific sequences by complexing a guide RNA that specifically hybridizes with a particular target sequence with a nuclease. Generating a target-specific guide RNA is less costly and more efficient than generating a chimeric nuclease for each target sequence. Such RNA-guided nucleases can be used to optionally edit the genome through the introduction of sequence-specific double-strand breaks that are repaired via error-prone non-homologous end joining (NHEJ) to introduce mutations at specific genomic positions. Alternatively, heterologous DNA may be introduced into a genomic site via homology-directed repair.

Summary of the Invention

[0004] Provided are compositions and methods for joining target sequences of interest. The compositions are found to be useful for cleaving or modifying target sequences of interest, detecting target sequences of interest, and modifying the expression of the sequences of interest. The compositions include RNA-guided nuclease (RGN) polypeptides, CRISPR RNAs (crRNAs), trans-activating CRISPR RNAs (tracrRNAs), guide RNAs (gRNAs), nucleic acid molecules encoding them, and vectors and host cells containing the nucleic acid molecules. Also provided is a CRISPR system for joining a target sequence of interest, where the CRISPR system includes an RNA-guided nuclease polypeptide and one or more guide RNAs. Accordingly, the methods disclosed herein are depicted for joining a target sequence of interest and, in some embodiments, for cleaving or modifying a target sequence of interest. The target sequence of interest can be modified, for example, as a result of non-homologous end joining or homologous directed repair with an introduced donor sequence. Further provided are methods and kits for detecting a target DNA sequence of a DNA molecule using a detection single-stranded DNA.

BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Many modifications and other embodiments of the inventions described herein will come to mind to those of ordinary skill in the art to which these inventions pertain, having the benefit of the teachings presented in the foregoing description and the related drawings. Therefore, the inventions are not to be limited to the specific embodiments disclosed, and modifications and other embodiments are intended to be included within the scope of the appended claims. Specific terms are used herein, but they are used for the purpose of illustration only and not for purposes of limitation.

[0006] I. SUMMARY RNA-guided nucleases (RGNs) enable targeted manipulation of a single site within the genome and are useful in the context of gene targeting for therapeutic and research applications. In various organisms, including mammals, RNA-guided nucleases have been used for genome engineering, for example, by stimulating non-homologous end joining and homologous recombination. The compositions and methods described herein are useful for creating single- or double-strand breaks in polynucleotides, for modifying polynucleotides, for detecting specific sites within polynucleotides, or for modifying the expression of a specific gene.

[0007] The RNA-guided nucleases disclosed herein can alter gene expression by modifying a target sequence. In certain embodiments, the RNA-guided nuclease is directed to a target sequence by a guide RNA (gRNA) as part of a clustered regularly interspaced short palindromic repeat (CRISPR) RNA-guided nuclease system. The guide RNA forms a complex with the RNA-guided nuclease and binds the RNA-guided nuclease to the target sequence and, in some embodiments, introduces a single- or double-strand break into the target sequence, such that the RGN is considered to be “RNA-guided.” After the target sequence is cleaved, the cleavage is repaired and the DNA sequence of the target sequence is modified during the repair process. Accordingly, provided herein are methods of using an RNA-guided nuclease to modify a target sequence in the DNA of a host cell. For example, an RNA-guided nuclease can be used to modify a target sequence at a genomic locus in a eukaryotic or prokaryotic cell.

[0008] II. RNA-guided nuclease Provided herein is an RNA-guided nuclease. The term "RNA-guided nuclease" refers to a polypeptide that binds to a specific target nucleotide sequence in a sequence-specific manner, complexes with a polypeptide, and is directed to the target nucleotide sequence by a guide RNA molecule that hybridizes to the target sequence. An RNA-guided nuclease can cleave the target sequence upon binding, but the term "RNA-guided nuclease" also includes nuclease-dead RNA-guided nucleases that bind to the target sequence but do not cleave it. Cleavage of the target sequence by an RNA-guided nuclease can result in a single-stranded or double-stranded break. An RNA-guided nuclease capable of cleaving only one strand of a double-stranded nucleic acid molecule is referred to herein as a nickase.

[0009] The RNA-guided nucleases disclosed in this specification include APG05733.1, APG06207.1, APG01647.1, APG08032.1, APG05712.1, APG01658.1, APG06498.1, APG09106.1, APG09882.1, APG02675.1, APG01405.1, APG06250.1, APG06877.1, APG09053.1, APG04293.1, APG01308.1, APG06646.1, APG09748, and APG07433.1 RNA-guided nucleases, and their amino acid sequences are described as SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235, respectively. Their active fragments or variants retain the ability to bind to a target nucleotide sequence in an RNA-guided sequence-specific manner. In some of these embodiments, an active fragment or variant of an APG05733.1, APG06207.1, APG01647.1, APG08032.1, APG05712.1, APG01658.1, APG06498.1, APG09106.1, APG09882.1, APG02675.1, APG01405.1, APG06250.1, APG06877.1, APG09053.1, APG04293.1, APG01308.1, APG06646.1, APG09748, or APG07433.1 RGN can cleave single-stranded or double-stranded target sequences.In some embodiments, an active variant of an APG05733.1, APG06207.1, APG01647.1, APG08032.1, APG05712.1, APG01658.1, APG06498.1, APG09106.1, APG09882.1, APG02675.1, APG01405.1, APG06250.1, APG06877.1, APG09053.1, APG04293.1, APG01308.1, APG06646.1, APG09748, or APG07433.1 RGN comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the amino acid sequence shown as SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137 or 235. In certain embodiments, an active fragment of an APG05733.1, APG06207.1, APG01647.1, APG08032.1, APG05712.1, APG01658.1, APG06498.1, APG09106.1, APG09882.1, APG02675.1, APG01405.1, APG06250.1, APG06877.1, APG09053.1, APG04293.1, APG01308.1, APG06646.1, APG09748, or APG07433.1 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues with the amino acid sequence shown in SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235. The RNA-guided nucleases provided herein can comprise at least one nuclease domain (e.g., a DNase, RNase domain) and at least one RNA recognition and / or RNA binding domain to interact with a guide RNA.Additional domains that may be found in the RNA-guided nucleases provided herein include, but are not limited to, DNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. In certain embodiments, the RNA-guided nucleases provided herein comprise one or more DNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains from at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%.

[0010] The target nucleotide sequence is bound by an RNA-guided nuclease provided herein and hybridizes with a guide RNA associated with the RNA-guided nuclease. Next, if the polypeptide has nuclease activity, the target sequence can be cleaved by the RNA-guided nuclease. The terms "cleavage" or "cleaving" refer to the hydrolysis of at least one phosphodiester bond within the backbone of the target nucleotide sequence that can result in either a single-stranded or double-stranded break within the target sequence. The presently disclosed RGNs can cleave nucleotides within a polynucleotide that functions as an endonuclease, or may be an exonuclease that removes consecutive nucleotides from the ends (5' and / or 3' ends) of the polynucleotide. In other embodiments, the disclosed RGNs can cleave nucleotides of the target sequence within any position of the polynucleotide and thus function as both an endonuclease and an exonuclease. Cleavage of the target polynucleotide by the presently disclosed RGNs can result in a staggered cleavage or blunt ends.

[0011] The presently disclosed RNA-guided nucleases can be wild-type sequences derived from bacterial or archaeal species. Alternatively, the RNA-guided nuclease may be a variant or fragment of a wild-type polypeptide. The wild-type RGN can be modified, for example, to change nuclease activity or PAM specificity. In some embodiments, the RNA-guided nuclease is not naturally occurring.

[0012] In some embodiments, the RNA-guided nuclease functions as a nickase and only cleaves one strand of the target nucleotide sequence. Such RNA-guided nucleases have a single functional nuclease domain. In some of these embodiments, additional nuclease domains have been mutated such that nuclease activity is reduced or eliminated.

[0013] In other embodiments, the RNA-guided nuclease is completely deficient in nuclease activity or exhibits reduced nuclease activity, and is referred to herein as nuclease dead or nuclease-inactive. Any method known in the art for introducing mutations into an amino acid sequence, such as PCR-mediated mutagenesis and site-directed mutagenesis, can be used to generate a nickase or nuclease-dead RGN. See, for example, U.S. Patent Application Publication No. 2014 / 0068797 and U.S. Patent No. 9,790,490, each of which is incorporated by reference in its entirety.

[0014] An RNA-guided nuclease lacking nuclease activity can be used to deliver a fusion polypeptide, polynucleotide, or small molecule payload to a specific genomic location. In some of these embodiments, the RGN polypeptide or guide RNA can be fused to a detectable label to enable detection of a specific sequence. As a non-limiting example, a nuclease-dead RGN can be fused to a detectable label (e.g., a fluorescent protein) and targeted to a specific sequence associated with a disease to enable detection of the disease-associated sequence.

[0015] Alternatively, the nuclease-dead RGN can target specific genomic locations to alter the expression of a desired sequence. In some embodiments, binding of the nuclease-dead RNA-guided nuclease to the target sequence results in suppression of the expression of the target sequence or a gene under the transcriptional control of the target sequence by interfering with the binding of RNA polymerase or a transcription factor within the target genomic region. In other embodiments, the RGN (e.g., nuclease-dead RGN) or its complexed guide RNA further comprises an expression modulator that, upon binding to the target sequence, acts to either suppress or activate the expression of the target sequence or a gene under the transcriptional control of the target sequence. In some of these aspects, the expression modulator modulates the expression of the target sequence or regulated gene via an epigenetic mechanism.

[0016] In other embodiments, nuclease-dead RGNs or RGNs having only nickase activity can target specific genomic positions to modify the sequence of a target polynucleotide via fusion to a base editing polypeptide, such as a deaminase polypeptide or an active variant or fragment thereof that deaminates nucleotide bases and results in the conversion of one nucleotide base to another. The base editing polypeptide can be fused to the RGN at its N-terminus or C-terminus. Additionally, the base editing polypeptide can be fused to the RGN via a peptide linker. Non-limiting examples of deaminase polypeptides useful in such compositions and methods include cytidine deaminase or adenosine deaminase (e.g., as described in Gaudelli et al. (2017) Nature 551:464-471, U.S. Patent Application Publication Nos. 2017 / 0121693 and 2018 / 0073012, and WO2018 / 027078, which are hereby incorporated by reference in their entirety, adenosine deaminase base editing, or any of the deaminases disclosed in PCT / US2019 / 068079). Further, it is known in the art that a specific fusion protein between an RGN and a base editing enzyme may include at least one uracil-stabilizing polypeptide that increases the mutation rate of cytosine, deoxycytosine, or cytidine to thymidine, deoxythymidine, or thymine in a nucleic acid molecule by a deaminase. Non-limiting examples of uracil-stabilizing polypeptides include those described in U.S. Provisional Application No. 63 / 052,175, filed July 15, 2020, and uracil glycosylase inhibitor (UGI) domains (SEQ ID NO: 261) that can enhance base editing efficiency (U.S. Patent No. 10,167,547, which is hereby incorporated by reference). Thus, the fusion protein may include an RGN as described herein, or a variant thereof, a deaminase, and optionally at least one uracil-stabilizing polypeptide such as UGI.

[0017] An RNA-guided nuclease fused to a polypeptide or domain can be separated or linked by a linker. As used herein, the term "linker" refers to a chemical group or molecule that links two molecules or moieties, such as a binding domain and a cleavage domain of a nuclease. In certain embodiments, the linker binds the gRNA-binding domain of an RNA-guided nuclease to a base-editing polypeptide such as a deaminase. In some embodiments, the linker binds a nuclease-dead RGN and a deaminase. Typically, the linker is positioned between or adjacent to two groups, molecules, or other moieties and is linked to each via a covalent bond, thus linking the two. In certain embodiments, the linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In certain embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is an amino acid of length 1-5, e.g., the length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids. Longer or shorter linkers are also contemplated.

[0018] The presently disclosed RNA-guided nucleases can include at least one nuclear localization signal (NLS) to enhance the transport of the RGN to the cell nucleus. Nuclear localization signals are known in the art and generally include a series of basic amino acids (e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In certain embodiments, the RGN includes 2, 3, 4, 5, 6, or more nuclear localization signals. The nuclear localization signals can be heterologous NLSs. Non-limiting examples of nuclear localization signals useful for the presently disclosed RGNs are the nuclear localization signals of SV40 large T antigen, nucleopasmin, and c-Myc (see, e.g., Ray et al. (2015) Bioconjug Chem 26(6):1004-7). In certain embodiments, the RGN includes the NLS sequence set forth in SEQ ID NO: 125 or 127. The RGN can include one or more NLS sequences at its N-terminus, C-terminus, or both the N-terminus and C-terminus. For example, the RGN can include two NLS sequences in the N-terminal region and four NLS sequences in the C-terminal region.

[0019] Other localization signal sequences known in the art for localizing polypeptides to specific intracellular locations can also be used to target RGN, including, but not limited to, plastid localization sequences, mitochondrial localization sequences, and dual-targeting signal sequences that target both plastids and mitochondria (see, for example, Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol dx.doi.org / 10.3389 / fphys.2015.00259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soll (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al. (2009) FEBS J 276:1187-1195; Silva-Filho (2003) Curr Opin Plant Biol 6:589-595; Peeters and Small (2001) Biochim Biophys Acta 1541:54-63; Murcha et al. (2014) J Exp Bot 65:6301-6335; Mackenzie (2005) Trends Cell Biol 15:548-554; Glaser et al. (1998) Plant Mol Biol 38:311-338).

[0020] In certain embodiments, the presently disclosed RNA-guided nucleases include at least one cell-penetrating domain that facilitates cellular uptake of the RGN. Cell-penetrating domains are known in the art and generally include stretches of positively charged amino acid residues (i.e., polycationic cell-penetrating domains), alternating polar and nonpolar amino acid residues (i.e., amphipathic cell-penetrating domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domains) (see, e.g., Milletti F. (2012) Drug Discov Today 17:850-860). Non-limiting examples of cell-penetrating domains are the trans-activating transcriptional activator (TAT) from human immunodeficiency virus 1.

[0021] The nuclear localization signal, plastid localization signal, mitochondrial localization signal, dual-targeting localization signal, and / or cell-penetrating domain can be located at the amino terminus (N-terminus), carboxyl terminus (C-terminus), or an internal position of the RNA-guided nuclease.

[0022] The presently disclosed RGNs can be fused, directly or indirectly via a linker peptide, to an effector domain such as a cleavage domain, deaminase domain, or expression modulator domain. Such domains can be located at the N-terminus, C-terminus, or an internal position of the RNA-guided nuclease. In some of these embodiments, the RGN component of the fusion protein is a nuclease-dead RGN.

[0023] In some embodiments, the RGN fusion protein is any domain capable of cleaving a polynucleotide (i.e., RNA, DNA, or an RNA / DNA hybrid), including but not limited to cleavage domains including restriction endonucleases and homing nucleases, such as type IIS endonucleases (e.g., FokI) (Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388; Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993).

[0024] In other embodiments, the RGN fusion protein includes a deaminase domain that deaminates nucleotide bases, resulting in a conversion from one nucleotide base to another, including but not limited to cytidine deaminase or adenosine deaminase base editing (e.g., Gaudelli et al. (2017) Nature 551:464-471; US Patent Application Publication Nos. 2017 / 0121693 and 2018 / 0073012, US Patent No. 9,840,699, and International Publication No. WO2018 / 027078).

[0025] In some embodiments, the effector domain of the RGN fusion protein can be an expression modulator domain that is a domain useful for upregulating or downregulating transcription. The expression modulator domain can be an epigenetic modification domain, a transcriptional repressor domain, or a transcriptional activation domain.

[0026] In some of these embodiments, the expression modulator of the RGN fusion protein covalently modifies DNA or histone proteins to change the histone structure and / or chromosomal structure without changing the DNA sequence, resulting in a change in gene expression (i.e., upregulation or downregulation), and includes an epigenetic modification domain. Non-limiting examples of epigenetic modifications include acetylation or methylation of lysine residues, arginine methylation, serine and threonine phosphorylation, as well as lysine ubiquitination and sumoylation of histone proteins, and methylation and hydroxymethylation of cytosine residues in DNA. Non-limiting examples of epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains, and DNA demethylase domains.

[0027] In other embodiments, the expression modulator of the fusion protein includes a transcriptional repressor domain that interacts with transcriptional control elements and / or transcriptional regulatory proteins, such as RNA polymerase and transcription factors, to reduce or terminate the transcription of at least one gene. Transcriptional repressor domains are known in the art and include, but are not limited to, Sp1-like repressors, IκB, and Krüppel-associated box (KRAB) domains.

[0028] In still other embodiments, the expression modulator of the fusion protein includes a transcriptional activation domain that interacts with transcriptional control elements and / or transcriptional regulatory proteins, such as RNA polymerase and transcription factors, to increase or activate the transcription of at least one gene. Transcriptional activation domains are known in the art and include, but are not limited to, the herpes simplex virus VP16 activation domain and the NFAT activation domain.

[0029] The presently disclosed RGN polypeptides can include a detectable label or a purification tag. The detectable label or purification tag can be located directly or indirectly via a linker peptide at the N-terminus, C-terminus, or an internal position of the RNA-guided nuclease. In some of these embodiments, the RGN component of the fusion protein is a nuclease-dead RGN. In other embodiments, the RGN component of the fusion protein is an RGN having nickase activity.

[0030] A detectable label is a molecule that can be visualized or otherwise observed. A detectable label can be fused to the RGN as a fusion protein (e.g., a fluorescent protein) or can be a small molecule complexed to an RGN polypeptide that can be detected visually or by other means. Detectable labels that can be fused to the presently disclosed RGN as a fusion protein include, but are not limited to, any detectable protein domain including a fluorescent protein or a protein domain that can be detected with a specific antibody. Non-limiting examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, EGFP, ZsGreen1) and yellow fluorescent proteins (e.g., YFP, EYFP, ZsYellow1). Non-limiting examples of small molecule detectable labels include 3 H and 35 radioactive labels such as S.

[0031] The RGN polypeptide can also include a purification tag, which is any molecule that can be used to isolate the protein or fusion protein from a mixture (e.g., a biological sample, a culture medium). Non-limiting examples of purification tags include biotin, myc, maltose-binding protein, and glutathione-S-transferase.

[0032] II. Guide RNA The present disclosure provides guide RNAs and polynucleotides encoding the same. The term "guide RNA" refers to a nucleotide sequence that hybridizes to a target sequence and has complementarity to a target nucleotide sequence sufficient for direct sequence-specific binding of an associated RNA-guided nuclease to the target nucleotide sequence. Thus, each guide RNA of an RGN is one or more RNA molecules (generally one or two) that can bind to the RGN and direct the RGN to bind to a specific target nucleotide sequence, and, if the RGN has nickase or nuclease activity, can also cleave the target nucleotide sequence. Generally, a guide RNA includes a CRISPR RNA (crRNA), and in some embodiments, a trans-activating CRISPR RNA (tracrRNA). A native guide RNA that includes both a crRNA and a tracrRNA generally includes two separate RNA molecules that hybridize to each other via the repeat sequence of the crRNA and the anti-repeat sequence of the tracrRNA.

[0033] The natural direct repeats within a CRISPR array generally range in length from 28 to 37 base pairs, although the length can vary between about 23 bp and about 55 bp. The spacer sequences within a CRISPR array generally range in length from about 32 to about 38 bp, although the length can be between about 21 bp and about 72 bp. Each CRISPR array generally contains less than 50 units of the CRISPR repeat-spacer sequences. CRISPR is transcribed as part of a long transcript called the primary CRISPR transcript that makes up most of the CRISPR array. The primary CRISPR transcript is cleaved by Cas proteins to produce crRNAs or, in some cases, pre-crRNAs. The pre-crRNAs are further processed to mature crRNAs by another Cas protein. The mature crRNAs contain spacer sequences and CRISPR repeat sequences. In some modes in which the pre-crRNA is processed to a mature (or processed) crRNA, the maturation involves removal of about 1 to about 6 or more 5’, 3’, or 5’ and 3’ nucleotides. For the purpose of genome editing or targeting a specific target nucleotide sequence of interest, these nucleotides that are removed during maturation of the pre-crRNA molecule are not required to generate or design the guide RNA.

[0034] CRISPR RNA (crRNA) contains a spacer sequence and a CRISPR repeat sequence. The "spacer sequence" is a nucleotide sequence that directly hybridizes with the target nucleotide sequence of interest. The spacer sequence is engineered to be fully or partially complementary to the target sequence of interest. In various embodiments, the spacer sequence can include from about 8 nucleotides to about 30 nucleotides, or more. For example, the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides in length. In some embodiments, the degree of complementarity between the spacer sequence and its corresponding target sequence, when aligned using a suitable alignment algorithm as appropriate, is about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or more. In certain embodiments, the spacer sequence does not contain secondary structure, which can be predicted using any suitable polynucleotide folding algorithm known in the art, for example, but not limited to, mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).

[0035] The CRISPR RNA repeat sequence comprises a nucleotide sequence that includes a region having sufficient complementarity to hybridize to tracrRNA. In various embodiments, the CRISPR RNA repeat sequence can include from about 8 nucleotides to about 30 nucleotides, or more. For example, the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more in length. In one aspect, the CRISPR repeat sequence is about 21 nucleotides in length. In some embodiments, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA sequence is about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or more when aligned using a suitable alignment algorithm as appropriate. In certain embodiments, the CRISPR repeat sequence, when included within the nucleotide sequence of SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287, or within a guide RNA, includes an active variant or a fragment thereof that can direct the sequence-specific binding of the associated RNA-guided nuclease provided herein to a target sequence of interest. In certain embodiments, an active CRISPR repeat sequence variant of a wild-type sequence includes a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to the nucleotide sequence shown as SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273 or 287.In certain embodiments, the active CRISPR repeat fragment of the wild-type sequence comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or more consecutive nucleotides of the nucleotide sequence shown as SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287.

[0036] In some embodiments, the crRNA is not naturally occurring. In some of these embodiments, the specific CRISPR repeat is not ligated to an engineered spacer sequence per se, and the CRISPR repeat is considered to be heterologous to the spacer sequence. In some embodiments, the spacer sequence is an engineered sequence that is not naturally occurring.

[0037] The trans-activating CRISPR RNA or tracrRNA molecule comprises a nucleotide sequence comprising a region having sufficient complementarity to hybridize to the CRISPR repeat sequence of the crRNA, which is referred to herein as the anti-repeat region. In certain embodiments, the tracrRNA molecule further comprises a region having a secondary structure (e.g., a stem-loop), or a region that forms a secondary structure by hybridizing to the corresponding crRNA. In certain embodiments, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is at the 5' end of the molecule, and the 3' end of the tracrRNA comprises a secondary structure. This secondary structure region generally comprises several hairpin structures including a nexus hairpin found adjacent to the anti-repeat sequence. The nexus hairpin often has a nucleotide sequence conserved at the base of the hairpin stem having the motif UNANNA, CNANNG, CNANNU, UNANNG, UNANNC, or CNANNU (SEQ ID NOs: 8, 37, 45, 53, 68, and 102), respectively, found in many of the nexus hairpins of the tracrRNA. At the 3' end of the tracrRNA, there are often terminal hairpin structures that differ in structure and number, but often include a GC-rich Rho-independent transcription terminator hairpin followed by a string of 3' terminal U's. See, e.g., Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Protoc; doi: 10.1101 / pdb.top090902, and U.S. Publication No. 2017 / 0275648, each of which is incorporated herein by reference in its entirety.

[0038] In various embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence comprises from about 8 nucleotides to about 30 nucleotides, or more. For example, the region of base pair formation between the tracrRNA anti-repeat sequence and the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In certain embodiments, the anti-repeat region of tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is about 20 nucleotides in length. In some embodiments, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA sequence is about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or more when aligned using a suitable alignment algorithm as appropriate.

[0039] In various embodiments, the entire tracrRNA can comprise from about 60 nucleotides to about 140 nucleotides. For example, the tracrRNA can be about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, or more nucleotides in length. In certain embodiments, the tracrRNA is about 80 to about 90 nucleotides in length, and is about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, and about 90 nucleotides in length. In one aspect, the tracrRNA is about 85 nucleotides in length.

[0040] In certain embodiments, the tracrRNA comprises the nucleotide sequence of SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119, or an active variant or fragment thereof that, when included within a guide RNA, can direct sequence-specific binding of a provided RNA-guided nuclease herein to a target sequence of interest. In certain embodiments, an active tracrRNA sequence variant of a wild-type sequence comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity to the nucleotide sequence set forth as SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286. In certain embodiments, an active tracrRNA sequence fragment of a wild-type sequence comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides to the nucleotide sequence set forth as SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286.

[0041] Two polynucleotide sequences can be considered to be substantially complementary if the two sequences hybridize to each other under stringent conditions. Similarly, if the guide RNA bound to the RGN binds to the target sequence under stringent conditions, the RGN is considered to bind to a specific target sequence in a sequence-specific manner. "Stringent conditions" or "stringent hybridization conditions" means conditions under which two polynucleotide sequences hybridize to each other to a detectably high degree compared to other sequences (e.g., at least twice that of the background). Stringent conditions are sequence-dependent and will be different in different circumstances. Typically, stringent conditions are at pH 7.0 - 8.3, with a salt concentration of about 1.5 M sodium ions, typically less than about 0.01 - 1.0 M sodium ion concentration (or other salts), and the temperature is at least about 30 °C for short sequences (e.g., 10 - 50 nucleotides), and for long sequences (e.g., over 50 nucleotides), stringent conditions can also be achieved by adding a destabilizing agent such as formamide. Exemplary low stringency conditions include hybridization in a buffer of 30 - 35% formamide, 1 M NaCl, 1% SDS (sodium dodecyl sulfate) at 37 °C, and washing in 1× - 2× SSC (20× SSC = 3.0 M NaCl / 0.3 M trisodium citrate) (50 - 55 °C). Examples of moderate stringency conditions include hybridization in 40 - 45% formamide, 1.0 M NaCl, 1% SDS (37 °C), and washing in 0.5× - 1× SSC (55 - 60 °C). High stringency conditions include hybridization in 50% formamide, 1 M NaCl, 1% SDS at 37 °C, and washing in 0.1× SSC at 60 - 65 °C. Optionally, the wash buffer may contain about 0.1% - about 1% SDS. The duration of hybridization is generally less than about 24 hours and is usually about 4 - about 12 hours. The duration of the wash time is of a length sufficient to reach at least equilibrium.

[0042] Tm is the temperature at which 50% of the complementary target sequence hybridizes to a perfectly matched sequence (under defined ionic strength and pH). In the case of a DNA-DNA hybrid, Tm can be approximated from the equation of Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5°C + 16.6(log M) + 0.41(%GC) - 0.61(% form) - 0.61(% form) - 500 / L. Here, M is the molarity of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, % form is the percentage of formamide in the hybridization solution, and L is the length of the hybrid in base pairs. Generally, stringent conditions are selected to be about 5°C lower than the thermal melting point (Tm) for a particular sequence and its complementarity at a defined ionic strength and pH. However, under highly stringent conditions, hybridization and / or washing can be utilized at 1, 2, 3, or 4°C lower than the thermal melting point (Tm). Under moderately stringent conditions, hybridization and / or washing can be utilized at 6, 7, 8, 9, or 10°C lower than the thermal melting point (Tm). Under low stringency conditions, hybridization and / or washing can be utilized at 11, 12, 13, 14, 15, or 20°C lower than the thermal melting point (Tm). Using this equation, hybridization and wash compositions, and the desired Tm, one of ordinary skill in the art will understand that variations in the stringency of hybridization and / or wash solutions are essentially described.Extensive guidance regarding nucleic acid hybridization can be found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology - Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York); and Ausubel et al., eds. (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2d ed., Cold Spring Harbor Laboratory Press, Plainview, New York).

[0043] The guide RNA may be a single guide RNA or a dual guide RNA system. A single guide RNA contains a crRNA and a tracrRNA on a single RNA molecule, while a dual guide RNA system exists on two different RNA molecules for the crRNA and contains a tracrRNA that hybridizes to each other via at least a part of the CRISPR repeat sequence of the crRNA and at least a part of the tracrRNA that can be completely or partially complementary to the CRISPR repeat sequence of the crRNA. In some embodiments where the guide RNA is a single guide RNA, the crRNA and the tracrRNA are separated by a linker nucleotide sequence. Generally, the linker nucleotide sequence does not contain complementary bases in order to avoid the formation of secondary structures within or containing the nucleotides of the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between the crRNA and the tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12 nucleotides, or longer. In certain embodiments, the linker nucleotide sequence of the single guide RNA is at least 4 nucleotides in length. In one aspect, the linker nucleotide sequence is the nucleotide sequence shown as SEQ ID NO: 123.

[0044] A single guide RNA or a dual guide RNA can be synthesized chemically or via in vitro transcription. Assays for determining sequence-specific binding between an RGN and a guide RNA are known in the art and include, but are not limited to, in vitro binding assays between an expressed RGN and a guide RNA that can be tagged with a detectable label (e.g., biotin) and used in a pull-down detection assay to capture the guide RNA:RGN complex via a detectable label (e.g., streptavidin beads). A control guide RNA having a sequence or structure unrelated to the guide RNA can be used as a negative control for non-specific binding of the RGN to RNA. In certain embodiments, the guide RNA is SEQ ID NO: 4, 12, 19, 26, 33, 41, 49, 57, 64, 72, 78, 85, 92, 98, 106, 113, or 120, and the spacer sequence can be any sequence, shown as a poly N sequence.

[0045] In some embodiments, the guide RNA can be introduced as an RNA molecule into a target cell, organelle, or embryo. The guide RNA can be transcribed in vitro or chemically synthesized. In other embodiments, a nucleotide sequence encoding the guide RNA is introduced into a cell, organelle, or embryo. In some of these embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., an RNA polymerase III promoter). The promoter can be a native promoter or can be heterologous to the nucleotide sequence encoding the guide RNA.

[0046] In various embodiments, the guide RNA can be introduced into a target cell, organelle, or embryo as a ribonucleoprotein complex as described herein, where the guide RNA is bound to an RNA-guided nuclease polypeptide.

[0047] The guide RNA directs the associated RNA-guided nuclease to a specific target nucleotide sequence via hybridization of the guide RNA to the target nucleotide sequence. The target nucleotide sequence can include DNA, RNA, or a combination of both, and can be single-stranded or double-stranded. The target nucleotide sequence can be genomic DNA (i.e., chromosomal DNA), plasmid DNA, or an RNA molecule (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). The target nucleotide sequence can be bound (and in some embodiments cleaved) by the RNA-guided nuclease in vitro or in a cell. The chromosomal sequence targeted by the RGN can be a chromosomal sequence of the nucleus, plastid, or mitochondrion. In some embodiments, the target nucleotide sequence is unique in the target genome.

[0048] The target nucleotide sequence is adjacent to a protospacer adjacent motif (PAM). The protospacer adjacent motif is generally within about 1 to about 10 nucleotides from the target nucleotide sequence, including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides from the target nucleotide sequence. The PAM can be 5' or 3' of the target sequence. In some embodiments, the PAM is 3' of the target sequence of the presently disclosed RGN. Generally, the PAM is a consensus sequence of about 3-4 nucleotides, but in certain embodiments can be 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length. In various embodiments, the PAM sequences recognized by the presently disclosed RGNs include consensus sequences shown as SEQ ID NO: 7, 15, 22, 29, 36, 44, 52, 60, 67, 81, 88, 101, 109, or 116.

[0049] In certain embodiments, an RNA-guided nuclease having SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117, or an active variant or fragment thereof, binds to a target nucleotide adjacent to a PAM sequence set forth in SEQ ID NO: 7, 15, 22, 29, 36, 44, 52, 60, 67, 81, 88, 101, 109, or 116, respectively. In some embodiments, an RNA-guided nuclease having SEQ ID NO: 54 or 137, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence set forth as SEQ ID NO: 147. In some of these embodiments, the RGN binds to a guide sequence comprising a CRISPR repeat sequence set forth in SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118, or an active variant or fragment thereof, and a tracrRNA sequence set forth in SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119, or an active variant or fragment thereof. The RGN system is further described in Example 1 herein, as well as in Tables 1 and 2.

[0050] It is well known in the art that the PAM sequence specificity for a given nuclease enzyme can be affected by enzyme concentration (see, e.g., Karvelis et al. (2015) Genome Biol 16:253). This can be modified by altering the promoter used to express the amount of RGN, or ribonucleoprotein complex delivered to a cell, organelle, or embryo.

[0051] When the RGN recognizes its corresponding PAM sequence, it can cleave the target nucleotide sequence at a specific cleavage site. As used herein, the cleavage site is composed of two specific nucleotides within the target nucleotide sequence where the nucleotide sequence is cleaved by the RGN. The cleavage site can include the first and second, second and third, third and fourth, fourth and fifth, fifth and sixth, seventh and eighth, or eighth and ninth nucleotides from either the 5’ or 3’ PAM. In some embodiments, the cleavage site can be more than 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the PAM in either the 5’ or 3’ direction. In some aspects, the cleavage site is 4 nucleotides away from the PAM. In other aspects, the cleavage site is at least 15 nucleotides away from the PAM. Since the RGN can cleave the target nucleotide sequence, generating a staggered end, in some embodiments, the cleavage site is defined based on the distance of two nucleotides from the PAM on the positive (+) strand of the polynucleotide and the distance of two nucleotides from the PAM on the negative (-) strand of the polynucleotide.

[0052] PAM is considered a characteristic of the RNA-guided nuclease of type II CRISPR systems (Szczelkun et al., PNAS, 111: 9798-9803, 2014; Sternberg et al., Nature 507: 62-67, 2014). Interestingly, APG06646.1 and APG04293.1 function as RNA-guided nucleases and have many of the same domains as the type II CRISPR Cas9 nuclease, but they lack the typical PAM interaction domain (PID; Interpro: IPR032237; Pfam: PF16595), respectively. Thus, APG06646.1 and APG04293.1 also do not have the typical PAM requirement, which is a 2- to 5-nucleotide motif as described above. Instead, these proteins have unique DNA recognition domains at their C termini (residues 821-1092 of APG06646.1 (the full-length sequence is SEQ ID NO: 117); residues 1064-1401 of APG04293.1 (the full-length sequence is SEQ ID NO: 103)). These unique DNA recognition domains enable the nuclease to cleave at the genomic target site based on a single nucleotide motif near the genomic target sequence (SEQ ID NO: 109; see Table 2).

[0053] APG04293.1 also has a unique characteristic domain of 133 amino acid residues proximal to its N terminus (residues 144-276). The function of this domain is not known in the type II CRISPR Cas9 nuclease or generally in the art.

[0054] III. Nucleotides encoding RNA-guided nuclease, CRISPR RNA, and / or tracrRNA The present disclosure provides polynucleotides comprising currently disclosed CRISPR RNAs, tracrRNAs, and / or sgRNAs, as well as polynucleotides comprising nucleotide sequences encoding currently disclosed RNA-guided nucleases, CRISPR RNAs, tracrRNAs, and / or sgRNAs. Currently disclosed polynucleotides include those comprising a CRISPR repeat sequence comprising the nucleotide sequence of SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287, or those comprising or encoding an active variant or fragment thereof that, when included within a guide RNA, can direct sequence-specific binding of a related RNA-guided nuclease to a target sequence of interest. Also disclosed are polynucleotides comprising a tracrRNA comprising the nucleotide sequence of SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286, or those comprising or encoding an active variant or fragment thereof that, when included within a guide RNA, can direct sequence-specific binding of a related RNA-guided nuclease to a target sequence of interest. Also provided are polynucleotides encoding RNA-guided nucleases comprising the amino acid sequences shown as SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235, and their active fragments or variants that retain the ability to bind to target nucleotide sequences in an RNA-guided sequence-specific manner.

[0055] The use of the term "polynucleotide" is not intended to limit the present disclosure to polynucleotides containing DNA. Those skilled in the art will recognize that polynucleotides can include ribonucleotides, as well as combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. These include peptide nucleic acids (PNAs), PNA-DNA chimeras, locked nucleic acids (LNAs), and phosphorothioate linkage sequences. The polynucleotides disclosed herein also encompass all forms of sequences including, but not limited to, single-stranded forms, double-stranded forms, DNA-RNA hybrids, triple-stranded structures, stem-and-loop structures, and the like.

[0056] The nucleic acid molecule encoding the RGN can be codon-optimized for expression in the organism of interest. A "codon-optimized" coding sequence is a polynucleotide coding sequence having its codon usage frequency designed to mimic the preferred codon usage or transcription condition frequencies of a particular host cell. Expression in a particular host cell or organism is enhanced as a result of one or more codon changes at the nucleic acid level such that the translated amino acid sequence is unchanged. The nucleic acid molecule can be fully or partially codon-optimized. Codon tables and other references providing preference information for a wide range of organisms are available in the art (see, for example, Campbell and Gowri (1990) Plant Physiol. 92:1-11). Methods for synthesizing plant-preferred genes are available in the art. See, for example, U.S. Patent Nos. 5,380,831, 5,436,391, and Murray et al. (1989) Nucleic Acids Res. 17:477-498, which are incorporated herein by reference.

[0057] The polynucleotides encoding the RNA, crRNA, tracrRNA, and / or sgRNA provided herein may be provided in an expression cassette for in vitro expression or expression in a cell, organelle, embryo, or organism of interest. The cassette will include 5' and 3' regulatory sequences linked such that they are functional with respect to the polynucleotide encoding the RGN, crRNA, tracrRNA, and / or sgRNA provided herein, which allows for expression of the polynucleotide. The cassette may further include at least one additional gene or genetic element to be co-transformed into the organism. If additional genes or elements are included, the components are operably linked. The term "operably linked" is intended to mean a functional linkage between two or more elements. For example, a functional linkage between a promoter and a coding region of interest (e.g., a region encoding an RGN, crRNA, tracrRNA, and / or sgRNA) is a functional linkage that allows for expression of the coding region of interest. Operably linked elements may be adjacent or non-adjacent. When used to refer to the linkage of two protein-coding regions, it is intended that the coding regions be in the same reading frame by being operably linked. Alternatively, additional genes or elements can be provided on multiple expression cassettes. For example, the nucleotide sequence encoding the presently disclosed RGN can be present on one expression cassette, while the nucleotide sequence encoding a crRNA, tracrRNA, or full guide RNA can be present on a separate expression cassette. Such expression cassettes are provided with multiple restriction sites and / or recombination sites for insertion of polynucleotides under the transcriptional regulation of the regulatory region. The expression cassette can further include a selectable marker gene.

[0058] The expression cassette includes, in the 5'-3' direction of transcription, the transcription (and, in some embodiments, translation) start region (i.e., promoter) of the present invention, a polynucleotide encoding RGN-, crRNA-, tracrRNA- and / or sgRNA-, and a transcription (and, in some embodiments, translation) termination region that is functional in the organism of interest. The promoter of the present invention can direct or drive the expression of the coding sequence in a host cell. The regulatory regions (e.g., promoter, transcription regulatory region, and translation termination region) may be endogenous or heterologous to the host cell or to each other. As used herein, "heterologous" with respect to a sequence means a sequence derived from a foreign species or, if derived from the same species, a sequence that has been substantially modified from its native form in composition and / or genomic locus by intentional human intervention. As used herein, a chimeric gene includes a coding sequence linked so as to be functional to a transcription start region that is heterologous to the coding sequence.

[0059] Convenient termination regions are available from the Ti plasmids of A. tumefaciens, such as the octopine synthase and nopaline synthase termination regions. See also Guerineau et al. (1991) Mol. Gen. Genet. 262:141-144; Proudfoot (1991) Cell 64:671-674; Sanfacon et al. (1991) Genes Dev. 5:141-149; Mogen et al. (1990) Plant Cell 2:1261-1272; Munroe et al. (1990) Gene 91:151-158; Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903; and Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.

[0060] Additional regulatory signals include, but are not limited to, transcription start sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See, for example, U.S. Pat. Nos. 5,039,523 and 4,853,331; European Patent No. 0480762A2; Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.) (hereinafter, "Sambrook 11"); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, N.Y., and the references cited therein.

[0061] When preparing an expression cassette, various DNA fragments can be manipulated to provide the DNA sequences in the appropriate orientation and, as appropriate, within the appropriate reading frame. For this purpose, adapters or linkers can be used to ligate DNA fragments or other manipulations to provide convenient restriction sites, removal of unwanted DNA, removal of restriction sites, etc. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, replacement, for example, transitions and transversions may be involved.

[0062] In the practice of the present invention, a number of promoters can be used. The promoter can be selected based on the desired result. The nucleic acid can be combined with a constitutive, inducible, growth stage-specific, cell type-specific, tissue-preferred, tissue-specific, or other promoter for expression in the subject organism. For example, see the promoters described in WO99 / 43838, which is incorporated herein by reference, as well as U.S. Patent Nos. 8,575,425; 7,790,846; 8,147,856; 8,586,832; 7,772,369; 7,534,939; 6,072,050; 5,659,026; 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; 5,608,142; and 6,177,611.

[0063] For expression in plants, constitutive promoters also include the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812); rice actin (McElroy et al. (1990) Plant Cell 2:163-171); ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et al. (1992) Plant Mol. Biol. 18:675-689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588); and MAS (Velten et al. (1984) EMBO J. 3:2723-2730).

[0064] Examples of inducible promoters include the Adh1 promoter, which is induced by hypoxia or cold stress, the Hsp70 promoter, which is induced by heat stress, the PPDK promoter, and the pepcarboxylase promoter, which is induced by light. Also useful are promoters that are chemically inducible, such as the safener-inducible In2-2 promoter (U.S. Patent No. 5,364,780), the Axig1 promoter (auxin-inducible and tapetum-specific but also active in callus) (PCT / US01 / 22169), and steroid-responsive promoters (e.g., the estrogen-inducible ERE promoter in Schena et al. (1991) Proc. Natl. Acad. Sci. USA 88:10421-10425 and McNellis et al. (1998) Plant J. 14(2):247-257, and tetracycline-inducible and tetracycline-repressible promoters (e.g., see Gatz et al. (1991) Mol. Gen. Genet. 227:229-237, U.S. Patent Nos. 5,814,618 and 5,789,156). These documents are hereby incorporated by reference into this specification.

[0065] Using a tissue-specific or tissue-preferred promoter, the expression of an expression construct can be targeted within a specific tissue. In certain embodiments, the tissue-specific or tissue-preferred promoter is active in plant tissue. Examples of promoters under the control of plant development include promoters that preferentially initiate transcription in specific tissues such as leaves, roots, fruits, seeds, flowers, etc. A "tissue-specific" promoter is a promoter that initiates transcription only in a specific tissue. Unlike the constitutive expression of a gene, tissue-specific expression is the result of the interaction of several levels of gene regulation. Therefore, promoters from homologous or closely related plant species are preferably used to achieve efficient and reliable expression of the transgene in a specific tissue. In some embodiments, the expression includes a tissue-preferred promoter. A "preferred tissue" promoter is a promoter that preferentially initiates transcription in a specific tissue, not necessarily completely or alone.

[0066] In certain embodiments, the nucleic acid molecule encoding RGN, crRNA, and / or tracrRNA includes a cell-type specific promoter. A "cell-type specific" promoter is a promoter that mainly drives expression in a specific cell type in one or more organs. Some examples of plant cells in which cell-type specific promoters functioning in plants can be mainly active include, for example, BETL cells, root vascular cells, leaves, stem cells, and meristematic cells. The nucleic acid molecule can also include a cell-type preferred promoter. A "cell-type preferred" promoter is a promoter that mainly drives expression (not necessarily completely or alone) in a specific cell type in one or more organs. Some examples of plant cells in which cell-type preferred promoters in plants can be selectively active include, for example, BETL cells, root vascular cells, leaves, stem cells, and meristematic cells.

[0067] Nucleic acid sequences encoding RNA, crRNA, tracrRNA, and / or sgRNA can be ligated to be functional with a promoter sequence recognized by a phage RNA polymerase, such as in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA can be purified for use in the methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence, or a variant of a T7, T3, or SP6 promoter sequence. In such embodiments, the expressed protein and / or RNA can be purified for use in the methods of genome modification described herein.

[0068] In certain embodiments, the polynucleotide encoding RGN, crRNA, tracrRNA, and / or sgRNA can also be ligated to a polyadenylation signal (e.g., the SV40 polyA signal and other functional signals in plants) and / or at least one transcription termination sequence. Further, the sequence encoding RGN can also be ligated to a sequence encoding at least one nuclear localization signal, at least one cell penetrating domain, and / or at least one signal peptide capable of transporting the protein to a specific intracellular location, as described elsewhere herein.

[0069] A polynucleotide encoding RGN, crRNA, tracrRNA, and / or sgRNA can be present in a vector or multiple vectors. A "vector" refers to a polynucleotide composition for transferring, delivering, or introducing nucleic acid into a host cell. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, baculoviral vectors). The vector can further contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Additional information can be found in "Current Protocols in Molecular Biology", Ausubel et al., John Wiley & Sons, New York, 2003 or "Molecular Cloning: A Laboratory Manual", Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, N.Y., 3rd edition, 2001.

[0070] The vector can also contain a selectable marker gene for the selection of transformed cells. The selectable marker gene is used for the selection of transformed cells or tissues. Marker genes include genes encoding antibiotic resistance such as neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), as well as genes conferring resistance to herbicide compounds such as glufosinate ammonium, bromoxynil, imidazolinone, and 2,4-dichlorophenoxyacetate (2,4-D).

[0071] In some embodiments, an expression cassette or vector comprising a sequence encoding an RGN polypeptide can further comprise a sequence encoding a crRNA and / or a tracrRNA, or a crRNA and a tracrRNA combined to generate a guide RNA. The sequence encoding the crRNA and / or tracrRNA can be operably linked to at least one transcriptional control sequence for expression of the crRNA and / or tracrRNA in the organism or host cell of interest. For example, the polynucleotide encoding the crRNA and / or tracrRNA can be operably linked to a promoter sequence recognized by RNA polymerase III (PolIII). Examples of suitable PolIII promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters, and rice U6 and U3 promoters.

[0072] As shown, an organism of interest can be transformed using an expression construct comprising a nucleotide sequence encoding an RGN, a crRNA, a tracrRNA, and / or an sgRNA. Methods for transformation include introducing a nucleotide construct into the organism of interest. By "introducing," it is intended that the nucleotide construct be introduced into the host cell such that the construct gains access to the interior of the host cell. The methods of the present invention do not require a specific method for introducing the nucleotide construct into the host organism, only that the nucleotide construct gain access to the interior of at least one cell of the host organism. The host cell can be a eukaryotic or prokaryotic cell. In certain embodiments, the eukaryotic host cell is a plant cell, a mammalian cell, or an insect cell. Methods for introducing nucleotide constructs into plants and other host cells are known in the art and include, but are not limited to, stable transformation methods, transient transformation methods, and virus-mediated methods.

[0073] This method results in transformed organisms such as plants including the whole plant, as well as plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos, and their progeny. Plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, meristematic cells, pollen).

[0074] "Transgenic organism" or "transformed organism" or "stably transformed" organism or cell or tissue refers to an organism into which a polynucleotide encoding the RGN, crRNA, and / or tracrRNA of the present invention has been incorporated or incorporated. It is recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments can also be taken up by the host cell. Transformation via Agrobacterium and biolistics are two approaches mainly used for the transformation of plant cells. However, the transformation of host cells can be carried out by infection, transfection, microinjection, electroporation, microprojection, particle gun or particle bombardment method, electroporation, silica / carbon fiber, sonication-mediated, PEG-mediated, calcium phosphate coprecipitation, polycation DMSO technology, DEAE, dextran procedure, and virus-mediated, liposome-mediated, etc. Virus-mediated introduction of polynucleotides encoding RGN, crRNA, and / or tractor RNA includes retrovirus, lentivirus, adenovirus, and adeno-associated virus-mediated introduction and expression, as well as the use of caulimovirus, geminivirus, and RNA plant viruses.

[0075] Transformation protocols, as well as protocols for introducing polypeptide or polynucleotide sequences into plants, can vary depending on the type of host cell targeted for transformation (e.g., monocot or dicot plant cells). Methods of transformation are known in the art and include those described in U.S. Patent Nos. 8,575,425, 7,692,068, 8,802,934, 7,541,517. See also Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett. 7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4:1-12; Bates, G.W. (1999) Methods in Molecular Biology 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606; Christou, P. (1992) The Plant Journal 2:275-281; Christou, P. (1995) Euphytica 85:13-27; Tzfira et al. (2004) TRENDS in Genetics 20:375-383; Yao et al. (2006) Journal of Experimental Botany 57:3737-3746; Zupan and Zambryski (1995) Plant Physiology 107:1041-1047; Jones et al. (2005) Plant Methods 1:5.

[0076] Transformation can result in the stable or transient uptake of nucleic acid into cells. "Stable transformation" means that a nucleotide construct introduced into a host cell is integrated into the genome of the host cell and can be inherited by its progeny. "Transient transformation" means that a polynucleotide is introduced into a host cell and is not integrated into the genome of the host cell.

[0077] Methods for chloroplast transformation are known in the art. See, for example, Svab et al. (1990) Proc. Natl. Acad. Sci. USA 87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J. 12:601-606. This method relies on particle gun delivery of DNA containing a selectable marker and targeting of the DNA to the plastid genome via homologous recombination. Additionally, plastid transformation can be achieved by transactivation of silent plastid-mediated transgenes by expression of a tissue-preferred nuclear-encoded and plastid-directed RNA polymerase. Such a system has been reported in McBride et al. (1994) Proc. Natl. Acad. Sci. USA 91:7301-7305.

[0078] The transformed cells can be grown into transgenic organisms such as plants according to conventional methods. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants were then grown, pollinated with the same transformant or a different strain, and the resulting hybrids were identified for the constitutive expression of the desired phenotypic traits. To ensure that the expression of the desired phenotypic traits is stably maintained and inherited, more than two generations were grown, and then harvested seeds can be obtained such that the expression of the desired phenotypic traits is reliably achieved. In this way, the present invention provides transgenic seeds (also referred to as "transformed seeds") having the nucleotide constructs of the present invention, for example, the expression cassettes of the present invention, stably integrated into their genomes.

[0079] Alternatively, the transformed cells may be introduced into an organism. These cells can be of biological origin, where the cells are transformed by an ex vivo approach.

[0080] The sequences provided herein can be used for the transformation of any plant species, including but not limited to monocotyledonous and dicotyledonous plants. Examples of plants of interest include, but are not limited to, maize (corn), sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugarcane, sweet potato, tobacco, barley, and oilseed rape, Brassica, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia, almond, oats, vegetables, ornamentals, and conifers.

[0081] The wild vegetables include, but are not limited to, tomatoes, lettuce, green beans, lima beans, peas, and cucumbers, members of the cucurbit genus such as cantaloupe and muskmelon. Ornamental plants include, but are not limited to, azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, chrysanthemums, etc. Preferably, the plants of the present invention are crop plants (e.g., corn, sorghum, wheat, sunflower, tomato, crucifer, pepper, potato, cotton, rice, soybean, sugarcane, sugar beet, tobacco, barley, oilseed rape, etc.).

[0082] As used herein, the term "plant" includes plant cells, plant protoplasts, plant cell tissue cultures capable of regenerating plants, plant callus, plant clumps, and intact plant cells in plants or parts of plants (embryos, pollen, eggs, seeds, leaves, flowers, branches, fruits, grains, ears, cob, husk, stems, roots, root tips, anthers, etc.). "Grain" means mature seeds produced by commercial producers for purposes other than the growth or reproduction of the species. The progeny, variants, and mutants of the regenerated plants are also included within the scope of the present invention, provided that these parts contain the polynucleotide introduced. Further, processed plant products or by-products that retain the sequences disclosed herein, such as soybeans, are provided.

[0083] Polynucleotides encoding RNA, crRNA, and / or tracrRNA can also be used to transform any prokaryotic species, including but not limited to archaebacteria and bacteria (e.g., Bacillus sp., Klebsiella sp., Streptomyces sp., Rhizobium sp., Escherichia sp., Pseudomonas sp., Salmonella sp., Shigella sp., Vibrio sp., Yersinia sp., Mycoplasma sp., Agrobacterium, Lactobacillus sp.).

[0084] Polynucleotides encoding RGN, crRNA, and / or tracrRNA can be used to transform any eukaryotic species including, but not limited to, animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, amoeba, algae, and yeast.

[0085] Nucleic acids can be introduced into mammalian cells or target tissues using conventional virus- and non-virus-based gene delivery methods. Such methods can be used to administer nucleic acids encoding components of the CRISPR system to cells in culture or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles (e.g., liposomes). Viral vector delivery systems include DNA and RNA viruses that have either an episomal genome or an integrated genome after delivery to the cell. For an overview of gene therapy procedures, see Anderson, Science 256: 808-813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10): 1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).

[0086] Methods of non-viral delivery of nucleic acids include lipofection, nucleofection, microinjection, biolistic, virosome, liposome, immunoliposome, polycation or lipid:nucleic acid conjugate, naked DNA, artificial virion, and uptake of agent-enhanced DNA. Lipofection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787; and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor recognition lipofection of polynucleotides include those of Feigner, WO91 / 17424; WO91 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vivo administration). Preparation of lipid:nucleic acid complexes, including targeted liposomes such as immunolipid complexes, is well known to those of skill in the art (see, for example, Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0087] The use of RNA- or DNA-virus-based systems for nucleic acid delivery exploits highly evolved processes for targeting viruses to specific cells in the body and transporting the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or used to treat cells in vitro, and the modified cells can be administered to patients as needed (ex vivo). Conventional virus-based systems can include retroviral, lentiviral, adenoviral, adeno-associated, and herpes simplex virus vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated virus gene transfer methods, often resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiencies have been observed in many different cell types and target tissues.

[0088] The affinity of retroviruses can be altered by incorporating foreign envelope proteins and expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system depends on the target tissue. Retroviral vectors consist of cis-acting long terminal repeat sequences with a packaging capacity of up to 6-10 kb of foreign sequences. The minimal cis-acting LTR is sufficient for vector replication and packaging and is then used to integrate the therapeutic gene into target cells to provide permanent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, for example, Buchscher et al., J. Viral. 66:2731-2739 (1992); Johann et al., J. Viral. 66:1635-1640 (1992); Sommnerfelt et al., Viral. 176:58-59 (1990); Wilson et al., J. Viral. 63:2374-2378 (1989); Miller et al., 1. Viral. 65:2220-2224 (1991); PCT / US94 / 05700).

[0089] For applications where transient expression is preferred, an adenovirus-based system can be used. Adenovirus-based vectors enable very high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained using such vectors. This vector can be produced in large quantities in a relatively simple system. Adeno-associated virus ("AAV") vectors can also be used to transduce cells with a target nucleic acid, for example, in in vitro production of nucleic acids and peptides, as well as in in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; WO93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, 1. Clin. Invest. 94:1351 (1994)). The construction of recombinant AAV vectors has been described in many publications, including U.S. patents, such as U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin, et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., 1. Viral. 63:03822-3828 (1989). Packaging cells are typically used to form virus particles capable of infecting host cells. Such cells include 293 cells for packaging adenovirus, and ΨJ2 cells or PA317 cells for packaging retrovirus.

[0090] Viral vectors used in gene therapy are typically generated by creating a cell line that packages a nucleic acid vector into viral particles. The vector typically contains the minimal viral sequences necessary for packaging and subsequent integration into the host, and other viral sequences are replaced with an expression cassette for the polynucleotide to be expressed. The defective viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically have only the ITR sequences from the AAV genome necessary for packaging and integration into the host genome. The viral DNA is packaged into a cell line that contains a helper plasmid that encodes the other AAV genes, i.e., rep and cap, but lacks the ITR sequences.

[0091] The cell line can also be infected with adenovirus as a helper. The helper virus facilitates the replication of the AAV vector and the expression of the AAV genes from the helper plasmid. Since the helper plasmid lacks the ITR sequences, a significant amount is not packaged. Contamination with adenovirus can be reduced, for example, by heat treatment, as adenovirus is more sensitive than AAV. Additional methods for delivering nucleic acids to cells are known to those of skill in the art. See, for example, US20003 / 0087817, which is incorporated herein by reference.

[0092] In certain embodiments, the host cell is transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, the cell is transfected because it occurs naturally in a subject. In certain embodiments, the cells to be transfected are obtained from a subject. In some embodiments, the cells are derived from a subject, e.g., cells obtained from a cell line. A variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPa cells, Panel, PC-3, TFl, CTLL-2, CIR, Rat6, CVI, RPTE, A1O, T24, 182, A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, TIB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelial, BALB / 3T3 mouse embryo fibroblast, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblast; 10.1 mouse fibroblast, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / - , COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC 6, MTD-lA, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-lA / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THPl cell lines, U373, U87, U937, VCaP, Vero cells WM39, WT-49, X63, YAC-1, YAR, and their transgenic variants are included. The cell lines are available from various sources known to those skilled in the art (see, for example, American Type Culture Collection (ATCC) (Manasas, Va.)).

[0093] In some embodiments, cells transfected with one or more of the vectors described herein are used to establish a new cell line containing one or more vector-derived sequences. In some embodiments, cells transiently transfected with components of the CRISPR system described herein (e.g., transient transfection of one or more vectors, or transfection with RNA), and modified through the activity of the CRISPR complex, are used to establish a new cell line containing cells that contain the modification but lack any other exogenous sequences. In one aspect, cells transiently or non-transiently transfected with one or more of the vectors described herein, or cell lines derived from such cells, are used to evaluate one or more test compounds.

[0094] In certain embodiments, one or more of the vectors described herein are used to produce non-human transgenic animals or transgenic plants. In certain embodiments, the transgenic animal is a mammal such as a mouse, rat, or rabbit.

[0095] IV. Variants and Fragments of Polypeptides and Polynucleotides The present disclosure provides active variants and fragments of naturally occurring (i.e., wild-type) RNA-guided nucleases, the amino acid sequences of which are set forth as SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235, and as SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287, and active variants and fragments of naturally occurring tracrRNA, e.g., SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286, and polynucleotides encoding the same.

[0096] The activity of the variant or fragment can be altered compared to the subject polynucleotide or polypeptide, but the variant and fragment should retain the functionality of the subject polynucleotide or polypeptide. For example, the variant or fragment can have an increase in activity, a decrease in activity, a different spectrum of activity, or any other change in activity when compared to the polynucleotide or polypeptide of interest.

[0097] Fragments and variants of the native RGN polypeptides as disclosed herein will retain sequence-specific RNA-guided DNA binding activity. In certain embodiments, fragments and variants of the native RGN polypeptides as disclosed herein will retain nuclease activity (single-stranded or double-stranded).

[0098] Fragments and variants of naturally occurring CRISPR repeats, such as those disclosed herein, will retain the ability of a portion of a guide RNA (including tracrRNA) to bind, in a sequence-specific manner, to an RNA-guided nuclease (complexed with the guide RNA) and guide it to a target nucleotide sequence.

[0099] Fragments and variants of naturally occurring tracrRNA, such as those disclosed herein, will retain the ability of a portion of a guide RNA (including CRISPR RNA) to guide an RNA-guided nuclease (complexed with the guide RNA) to a target nucleotide sequence in a sequence-specific manner.

[0100] The term "fragment" refers to a part of a polynucleotide or polypeptide sequence of the present invention. "Fragments" or "biologically active portions" include polynucleotides containing a sufficient number of contiguous nucleotides to retain biological activity (i.e., when included within a guide RNA, bind to an RGN in a sequence-specific manner and direct it to a target nucleotide sequence). "Fragments" or "biologically active portions" include polypeptides containing a sufficient number of contiguous amino acid residues to retain biological activity (i.e., when complexed with a guide RNA, bind to a target nucleotide sequence in a sequence-specific manner). Fragments of the RGN protein include those that are shorter than the full-length sequence by use of an alternative downstream start site. Biologically active portions of the RGN protein can be, for example, peptides containing 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more contiguous amino acid residues of SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235. Such biologically active portions can be prepared by recombinant techniques and can be evaluated for sequence-specific RNA-guided DNA binding activity. Biologically active fragments of the CRISPR repeat sequence can contain at least 8 contiguous amino acids of SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287. Biologically active portions of the CRISPR repeat sequence can be, for example, polynucleotides containing 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 contiguous nucleotides of SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287.The biologically active portion of the tracrRNA can be a polynucleotide comprising, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides of SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286.

[0101] In general, "variant" is intended to mean a substantially similar sequence. For polynucleotides, variants include deletions and / or additions of one or more nucleotides at one or more internal sites within the native polynucleotide, and / or substitutions of one or more nucleotides at one or more sites within the native polynucleotide. As used herein, "native" or "wild-type" polynucleotide or polypeptide includes a naturally occurring nucleotide sequence or amino acid sequence, respectively. For polynucleotides, conservative variants include sequences that encode the native amino acid sequence of the gene of interest due to the degeneracy of the genetic code. Such naturally occurring allelic variants can be identified using well-known molecular biology techniques, such as polymerase chain reaction and hybridization techniques as outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those generated, for example, by site-directed mutagenesis, that still encode the polypeptide or polynucleotide of interest. In general, variants of a particular polynucleotide disclosed herein have at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to that particular polynucleotide as determined by the sequence alignment programs and parameters described elsewhere herein.

[0102] Variants of a particular polynucleotide (i.e., a reference polynucleotide) disclosed herein can also be evaluated by comparison of the % sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The percent sequence identity between any two polypeptides can be calculated using the sequence alignment programs and parameters described elsewhere herein. If any given pair of polynucleotides disclosed herein is evaluated by comparison of the percent sequence identity shared by the two polypeptides they encode, the percent sequence identity between the two encoded polypeptides is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity.

[0103] In certain embodiments, the presently disclosed polynucleotides encode an RNA-guided nuclease polynucleotide having an amino acid sequence that has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity to the amino acid sequence of SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235.

[0104] Biologically active variants of the RGN polypeptide of the present invention differ by about 1 to 15 amino acid residues, about 1 to 10 amino acid residues, about 6 to 10 amino acids, just 5 amino acids, just 4 amino acids, just 3 amino acids, just 2 amino acids, or just 1 amino acid residue. In certain embodiments, the polypeptide can include N-terminal or C-terminal truncations, which can be deletions of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 amino acids or more from either the N- or C-terminus of the polypeptide.

[0105] In certain embodiments, the presently disclosed polynucleotides include or encode a CRISPR repeat having a nucleotide sequence that has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity to the nucleotide sequence shown as SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287.

[0106] The presently disclosed polynucleotides can include or encode a tracrRNA having a nucleotide sequence that has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater identity to the nucleotide sequence shown as SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286.

[0107] Biologically active variants of the CRISPR repeats or tracrRNAs of the present invention may differ by about 1 to 15 nucleotides, about 1 to 10, about 6 to 10, just 5, just 4, just 3, just 2, or just 1 nucleotide. In certain embodiments, the polynucleotide can include a 5' or 3' truncation, which can include a deletion of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 nucleotides or more from the 5' or 3' end of the polynucleotide.

[0108] Modifications can be made to the RGN polypeptide, CRISPR repeats, and tracrRNAs provided herein, and it is recognized that modified proteins and polynucleotides can be made. Changes designed by humans can be introduced by applying site-directed mutagenesis techniques. Alternatively, natural, unknown or unidentifed polynucleotides and / or polypeptides that are structurally and / or functionally related to the sequences disclosed herein can also be identified within the scope of the present invention. Conservative amino acid substitutions can be made in non-conserved regions that do not alter the function of the RGN protein. Alternatively, modifications that improve the activity of the RGN may be made.

[0109] Variant polynucleotides and proteins also include sequences and proteins derived from mutagenic and recombinant techniques such as DNA shuffling. In such procedures, one or more different RGN proteins disclosed herein (e.g., SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235) are manipulated to create new RGN proteins with desired properties. In this way, a library of recombinant polynucleotides is generated from a population of related sequence polynucleotides that have substantial sequence identity and contain sequence regions that can be recombinantly paired in vitro or in vivo. For example, using this method, sequence motifs encoding domains of interest are shuffled between the RGN sequences provided herein and other known RGN genes to obtain a novel gene encoding a protein with improved properties such as increased K m of the protein of interest. Strategies for such DNA shuffling are known in the art. See Crameri et al. (1998) Nature 391:288-291; U.S. Patent Nos. 5,605,793 and 5,837,458. A "shuffled" nucleic acid is a nucleic acid generated by a shuffling procedure such as any of the shuffling procedures described herein. Shuffled nucleic acids are generated by recombining (physically or virtually) two or more nucleic acids (or characteristic strings), e.g., in an artificial and optionally recursive manner. Generally, one or more screening steps are used in the shuffling procedure to identify the nucleic acid of interest; this screening step can be performed before or after any recombination step. In some (but not all) shuffling embodiments, it is desirable to perform multiple recombinations prior to selection to increase the diversity of the pool to be screened. The overall process of recombination and selection is optionally repeated recursively. Depending on the context, shuffling can refer to the overall process of recombination and selection, or alternatively, can simply refer to the recombination portion of the overall process.

[0110] As used herein, "sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences refers to residues in two sequences that are the same when aligned for maximum match over a specified comparison window. When percentage of sequence identity is used in reference to a protein, it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity), and thus do not change the functional properties of the molecule. When sequences differ in conservative substitutions, the percent sequence identity can be adjusted upwards to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity". Means for making this adjustment are well known to those of skill in the art. Typically, this involves scoring conservative substitutions as partial matches rather than perfect mismatches, thereby increasing the percentage sequence identity. Thus, for example, if identical amino acids are given a score of 1 and non-conservative substitutions are given a score of zero, conservative substitutions are given a score between zero and 1. Scoring of conservative substitutions is calculated, e.g., as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).

[0111] As used herein, "percentage of sequence identity" means a value determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide sequence in the comparison window may include additions or deletions (i.e., gaps) as compared to the reference sequence for optimal alignment of the two sequences (which does not include additions or deletions). The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to yield the percentage of sequence identity.

[0112] Unless otherwise specified, the sequence identity / similarity values provided in this specification refer to values obtained using GAP version 10 with the following parameters: GAP weight 50 and length weight 3, and % identity and % similarity for nucleotide sequences using the nwsgapdna.cmp scoring matrix; % identity and % similarity for amino acid sequences using GAP weight 8 and length weight 2; and the BLOSUM62 scoring matrix; or equivalent programs. By "equivalent programs" is intended any sequence comparison program that, when compared to the corresponding alignment generated by GAP version 10 for the two sequences in question, generates an alignment having the same nucleotide or amino acid residue matches and the same percent sequence identity.

[0113] Two arrays are "optimally aligned" if they are aligned for similarity scoring using a defined amino acid substitution matrix (e.g., BLOSUM62), a gap existence penalty, and a gap extension penalty, and reach the highest possible score for that pair of arrays. The use of amino acid substitution matrices and their use in quantifying the similarity between two arrays are well known in the art and are described, for example, in Dayhoff et al. (1978) "A model of evolutionary change in proteins", In "Atlas of Protein Sequence and Structure", Vol. 5, Suppl. 3 (ed. M. O. Dayhoff), pp. 345-352. Natl. Biomed. Res. Found., Washington, D.C. and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA 89:10915-1091. The BLOSUM62 matrix is often used as the default scoring substitution matrix in sequence alignment protocols. The gap existence penalty is imposed to introduce a single amino acid gap into one of the aligned arrays, and the gap extension penalty is imposed for additional empty amino acid positions inserted into an already open gap. An alignment is defined by the amino acid positions of each array where the alignment starts and ends, and optionally, by inserting gaps or multiple gaps into one or both arrays, and reaches the highest possible score. Optimal alignment and scoring can be done manually, but this process is facilitated by the use of computer-implemented alignment algorithms such as gapped BLAST 2.0, which is disclosed, for example, in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402 and is publicly available at the National Center for Biotechnology Information Website (www.ncbi.nlm.nih.gov).Optimal alignments that include multiple alignments are available, for example, via www.ncbi.nlm.nih.gov and can be prepared using PSI-BLAST as described by Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402.

[0114] With respect to an amino acid sequence optimally aligned with a reference sequence, an amino acid residue "corresponds" to the position in the reference sequence where the residue is paired in the alignment. "Position" is indicated by a number that sequentially identifies each amino acid in the reference sequence based on its position relative to the N-terminus. Due to deletions, insertions, cleavages, fusions, etc. that must be considered when determining the optimal alignment, generally, the number of amino acid residues in a test sequence determined simply by counting from the N-terminus will not necessarily be the same as the number of its corresponding positions in the reference sequence. For example, if there is a deletion in the aligned test sequence, there will be no amino acid corresponding to the position in the reference sequence at the deletion site. If there is an insertion in the aligned reference sequence, the insertion does not correspond to any amino acid position in the reference sequence. In the case of a cleavage or fusion, there may be a stretch of amino acids in the reference or aligned sequence that does not correspond to any amino acid in the corresponding sequence.

[0115] V. Antibodies Antibodies against the RGN polypeptide or ribonucleoprotein comprising the RGN polypeptide of the present invention (SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117, or active variants or fragments thereof) are also included. Methods for producing antibodies are well known in the art (see, e.g., Harlow and Lane (1988) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y.; and U.S. Patent No. 4,196,265). These antibodies can be used in kits for the detection and isolation of RGN polypeptides or ribonucleoproteins. Accordingly, the present disclosure provides a kit comprising an antibody that specifically binds to a polypeptide or ribonucleoprotein described herein, including, for example, a polypeptide having the sequence of SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117.

[0116] VI. Systems that Bind to a Sequence of Interest and Ribonucleoprotein Complexes and Methods for Synthesizing Them The present disclosure provides a system for binding a target sequence of interest, where the system includes at least one guide RNA or nucleotide sequence encoding the same, and at least one RNA-guided nuclease or nucleotide sequence encoding the same. The guide RNA hybridizes to the target sequence of interest and forms a complex with the RGN polypeptide, thereby binding the RGN polypeptide to the target sequence. In some of these embodiments, the RGN includes SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235, or an active variant or fragment thereof. In various embodiments, the guide RNA includes a nucleotide sequence of a CRISPR repeat sequence including SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287, or an active variant or fragment thereof. In certain embodiments, the guide RNA includes a tracrRNA having a nucleotide sequence of SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286, or an active variant or fragment thereof. The guide RNA of the system may be a single guide RNA or a dual guide RNA. In certain embodiments, the system includes an RNA-guided nuclease that is heterologous to the guide RNA, where the RGN and the guide RNA are not naturally complexed.

[0117] The system for binding to a target sequence of interest provided herein can be a ribonucleoprotein complex that is at least one molecule of RNA bound to at least one protein. The ribonucleoprotein complex provided herein includes at least one guide RNA as an RNA component and an RNA-guided nuclease as a protein component. Such a ribonucleoprotein complex can be purified from a cell or organism that naturally expresses the RGN polypeptide and has been engineered to express a specific guide RNA specific for the target sequence of interest. Alternatively, the ribonucleoprotein complex can be purified from a cell or organism that has been transformed with a polynucleotide encoding the RGN polypeptide and the guide RNA and cultured under conditions that permit expression of the RGN polypeptide and the guide RNA. Accordingly, a method for producing an RGN polypeptide or an RGN ribonucleoprotein complex is provided. Such a method includes culturing a cell comprising a nucleotide sequence encoding the RGN polypeptide and, in some embodiments, a cell comprising a nucleotide sequence encoding the guide RNA under conditions in which the RGN polypeptide (and in some embodiments, the guide RNA) is expressed. Next, the RGN polypeptide or the RGN ribonucleoprotein can be purified from the lysate of the cultured cells.

[0118] Methods for purifying RGN polypeptides or RGN ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., size exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reverse phase chromatography, immunoprecipitation). In certain methods, the RGN polypeptide is recombinantly produced and contains a purification tag to aid in its purification, including, but not limited to, glutathione-S-transferase (GST), chitin binding protein (CBP), maltose binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 10xHis, biotin carboxyl carrier protein (BCCP), and calmodulin. Generally, labeled RGN polypeptides or RGN ribonucleoprotein complexes are purified using immobilized metal affinity chromatography. It will be understood that other forms of chromatography or other similar methods known in the art, including, for example, immunoprecipitation, can be used alone or in combination.

[0119] An "isolated" or "purified" polypeptide, or a biologically active portion thereof, is substantially or essentially free of components that are normally associated with or interact with the polypeptide as found in its natural environment. Thus, an isolated or purified polypeptide is substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized. A protein that is substantially free of cellular material includes preparations of the protein having less than about 30%, 20%, 10%, 5%, or 1% (dry weight) of contaminating protein. When the protein of the invention or a biologically active portion thereof is produced recombinantly, the optimal culture medium represents less than about 30%, 20%, 10%, 5%, or 1% (dry weight) of chemical precursors or non-protein chemicals of interest.

[0120] The specific methods provided herein for binding and / or cleaving a target sequence of interest involve the use of an in vitro assembled RGN ribonucleoprotein complex. In vitro assembly of the RGN ribonucleoprotein complex can be performed using any method known in the art for contacting an RGN polypeptide with a guide RNA under conditions that allow the RGN polypeptide to bind to the guide RNA. As used herein, "contacting," "contact," and "contacted" refer to placing the components of a desired reaction together under conditions suitable for performing the desired reaction. The RGN polypeptide can be purified from a biological sample, cell lysate, or culture medium, produced by in vitro translation, or chemically synthesized. The guide RNA can be purified from a biological sample, cell lysate, or culture medium, transcribed in vitro, or chemically synthesized. The RGN polypeptide and the guide RNA can be contacted in solution (e.g., buffered aqueous saline solution) to allow for in vitro assembly of the RGN ribonucleoprotein complex.

[0121] VII. Methods for Binding, Removing, or Altering a Target Sequence The present disclosure provides methods for binding, cleaving, and / or modifying a target nucleotide sequence of interest. The method includes delivering to the target sequence, or a cell, organelle, or embryo containing the target sequence, a system comprising at least one guide RNA or polynucleotide encoding the same, and at least one RGN polypeptide or polynucleotide encoding the same. In some of these embodiments, the RGN comprises the amino acid sequence of SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence of SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287, or an active variant or fragment thereof. In certain embodiments, the guide RNA comprises a tracrRNA comprising the nucleotide sequence of SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286, or an active variant or fragment thereof. The guide RNA of the system may be a single guide RNA or a dual guide RNA. The RGN of the system may be a nuclease-dead RGN, may have nickase activity, or may be a fusion polypeptide. In some aspects, the fusion polypeptide comprises a base editing polypeptide, for example, a cytidine deaminase or an adenosine deaminase. In other embodiments, the RGN fusion protein comprises a reverse transcriptase.In other embodiments, the RGN fusion protein comprises a polypeptide that rescues a member of a functional nucleic acid repair complex, such as a member of the nucleotide excision repair (NER) or transcription-coupled nucleotide excision repair (TC-NER) pathway (Wei et al., 2015, PNAS USA 112(27):E3495-504; Troelstra et al., 1992, Cell 71:939-953; Marnef et al., 2017, J Mol Biol 429(9):1277-1288), as described in U.S. Provisional Application 62 / 966,203, filed January 27, 2020, which is hereby incorporated by reference in its entirety. In some embodiments, the RGN fusion protein comprises CSB (van den Boom et al., 2004, J Cell Biol 166(1):27-36; van Gool et al., 1997, EMBO J 16(19):5955-65; an example of which is shown as SEQ ID NO: 268) (which is the TC-NER (nucleotide excision repair) pathway and functions as the rescue of other members). In further embodiments, the RNAGNABP fusion protein comprises an active domain of CSB, such as the acidic domain of CSB comprising amino acid residues 356-394 of SEQ ID NO: 268 (Teng et al., 2018, Nat Commun 9(1):4115).

[0122] In certain embodiments, the RGN and / or guide RNA is heterologous to the cell, organelle, or embryo into which the RGN and / or guide RNA (or polynucleotide encoding at least one of the RGN and guide RNA) is introduced.

[0123] In embodiments where the method involves delivering a polynucleotide encoding a guide RNA and / or an RGN polypeptide, the cell or embryo can then be cultured under conditions in which the guide RNA and / or RGN polypeptide are expressed. In various aspects, the method involves contacting a target sequence with an RGN ribonucleoprotein complex. The RGN ribonucleoprotein complex can include an RGN in which the nuclease is dead or has nickase activity. In some embodiments, the RGN of the ribonucleoprotein complex is a fusion polypeptide that includes a base editing polypeptide. In one aspect, the method involves introducing an RGN ribonucleoprotein complex into a cell, organelle, or embryo that contains the target sequence. The RGN ribonucleoprotein complex can be purified from a biological sample, recombinantly produced and then purified, or assembled in vitro as described herein. In embodiments where the RGN ribonucleoprotein complex that contacts the target sequence, cell, organelle, or embryo is assembled in vitro, the method can further include in vitro assembly of the complex prior to contacting the target sequence, cell, cell organelle, or embryo.

[0124] The purified or in vitro assembled RGN ribonucleoprotein complex can be introduced into a cell, organelle, or embryo using any method known in the art, including but not limited to electroporation. Alternatively, an RGN polypeptide and / or polynucleotide encoding or including the guide RNA can be introduced into a cell, organelle, or embryo using any method known in the art (e.g., electroporation).

[0125] Upon delivery or contact to a target array, cell, organelle, or embryo containing the target array, the guide RNA directs the RGN to bind to the target array in a sequence-specific manner. In embodiments where the RGN has nuclease activity, the RGN polypeptide cleaves the target array of interest upon binding. The target array can then be modified via endogenous repair mechanisms such as non-homologous end joining or homology-directed repair by a provided donor polynucleotide.

[0126] Methods for measuring the binding of an RGN polypeptide to a target array are known in the art and include chromatin immunoprecipitation assays, gel mobility shift assays, DNA pull-down assays, reporter assays, microplate capture and detection assays. Similarly, methods for measuring cleavage or modification of a target array are known in the art and include in vitro or in vivo cleavage assays where cleavage is confirmed using PCR, sequencing, or gel electrophoresis, with or without the attachment of an appropriate label (e.g., radioisotope, fluorophore) to the target array to facilitate detection of the cleavage products. Alternatively, a nick-induced exponential amplification reaction (NTEXPAR) assay can be used (see, e.g., Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage can be evaluated using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).

[0127] In some embodiments, the method includes the use of a single type of RGN complexed with two or more guide RNAs. The multiple guide RNAs can target different regions of one gene or multiple genes.

[0128] In embodiments where no donor polynucleotide is provided, the double-strand breaks introduced by the RGN polypeptide can be repaired by the non-homologous end joining (NHEJ) repair process. Due to the error-prone nature of NHEJ, the repair of double-strand breaks can result in modification of the target sequence. As used herein, "modification" with respect to a nucleic acid molecule refers to a change in the nucleotide sequence of the nucleic acid molecule that can be a deletion, insertion, or substitution of one or more nucleotides, or a combination thereof. Modification of the target sequence can result in the expression of an altered protein product or inactivation of the coding sequence.

[0129] In embodiments where a donor polynucleotide is present, during the process of repair of the introduced double-strand breaks, the donor sequence in the donor polynucleotide can be incorporated into or exchanged with the target nucleotide sequence, resulting in the introduction of an exogenous donor sequence. Thus, the donor polynucleotide contains a donor sequence that is desired to be introduced into the target sequence of interest. In some embodiments, the donor sequence modifies the original target nucleotide sequence such that the newly incorporated donor sequence is not recognized and cleaved by the RGN. Integration of the donor sequence can be enhanced by inclusion within the donor polynucleotide of a flanking sequence having substantial sequence identity to the sequence adjacent to the target nucleotide sequence, enabling a homology-directed repair process. In embodiments where the RGN polypeptide introduces a double-strand twist break, the donor polynucleotide contains a donor sequence adjacent to a compatible overhang, allowing for direct ligation of the donor sequence to the cleaved target nucleotide sequence containing the overhang by a non-homologous repair process during repair of the double-strand break.

[0130] In embodiments where the method involves the use of nickases (i.e., RNPs that can only cleave one strand of a double-stranded polynucleotide), the method can include introducing two nickase RNPs that target the same or overlapping target sequences and cleave different strands of the polynucleotide. For example, an RNP nickase that cleaves only the positive (+) strand of a double-stranded polynucleotide can be introduced together with a second RNP nickase that cleaves only the negative (-) strand of the double-stranded polynucleotide.

[0131] In various embodiments, methods are provided for binding to a target nucleotide sequence and detecting the target sequence, the method comprising introducing into a cell, organelle, or embryo at least one guide RNA or polynucleotide encoding the same, and (when a coding sequence is introduced) at least one RNP polypeptide that expresses the guide RNA and / or RNP polypeptide, where the RNP polypeptide is a nuclease-inactive RNP and further comprises a detectable label. The detectable label can be fused to the RNP as a fusion protein (e.g., a fluorescent protein), or conjugated to the RNP polypeptide that can be detected visually or by other means, or can be a small molecule incorporated within the RNP polypeptide.

[0132] Also provided herein are methods for modulating the expression of a target sequence or a target gene under the regulation of the target sequence. The method comprises introducing into a cell, organelle, or embryo at least one guide RNA or polynucleotide encoding the same, and (when a coding sequence is introduced) at least one RNP polypeptide or polynucleotide encoding the same that expresses the guide RNA and / or RNP polypeptide, where the RNP polypeptide is a nuclease-dead RNP. In some of these embodiments, the nuclease-dead RNP is a fusion protein comprising an expression modulator domain (i.e., an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repressor domain) described herein.

[0133] The present disclosure also provides methods for binding and / or modifying a target nucleotide sequence of interest. The methods include at least one guide RNA or a polynucleotide encoding the same, and delivering to the target sequence, or a cell, organelle, or embryo comprising the target sequence, a system comprising at least one fusion polypeptide comprising an RGN and a base editing polypeptide of the invention, such as a cytidine deaminase or an adenosine deaminase, or a polynucleotide encoding the fusion polypeptide.

[0134] One of ordinary skill in the art will understand that any of the presently disclosed methods can be used to target a single target sequence or multiple target sequences. Thus, the methods include the use of a single RGN polypeptide in combination with multiple distinct guide RNAs that can target multiple different sequences within a single gene and / or multiple genes. Also included herein are methods in which multiple different guide RNAs are introduced in combination with multiple different RGN polypeptides. These guide RNAs and guide RNA / RGN polypeptide systems can target multiple different sequences within a single gene and / or multiple genes.

[0135] In one aspect, the present invention provides a kit comprising any one or more of the elements disclosed in the above methods and compositions. In some aspects, the kit comprises a vector system and instructions for using the kit. In some embodiments, the vector system comprises: (a) a first regulatory element operably linked to a DNA sequence encoding a crRNA sequence, and one or more insertion sites for inserting a guide sequence upstream of the encoded crRNA sequence, wherein when expressed, the guide sequence directs sequence-specific binding of a CRISPR complex to a target sequence in a eukaryotic cell, and the CRISPR complex comprises a CRISPR enzyme complexed with the guide RNA polynucleotide; and / or (b) an enzyme encoding the CRISPR enzyme comprising a nuclear localization sequence. The elements can be provided individually or in combination and can be provided in any suitable container, such as a vial, bottle, or tube.

[0136] In some embodiments, the kit comprises instructions in one or more languages. In some aspects, the kit comprises one or more reagents for use in a process that utilizes one or more of the elements described herein. The reagents can be provided in any suitable container. For example, the kit can provide one or more reaction or storage buffers. The reagents can be provided in a form that is ready for use in a particular assay or in a form that requires the addition of one or more other components prior to use (e.g., concentrated or lyophilized form). The buffer can be any buffer including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, boric acid buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some aspects, the buffer is alkaline. In some aspects, the buffer has a pH of from about 7 to about 10.

[0137] In some embodiments, the kit comprises one or more oligonucleotides corresponding to a guide sequence for insertion into a vector so as to operably link the guide sequence and the regulatory element. In some embodiments, the kit comprises a homologous recombination template polynucleotide. In one aspect, the invention provides a method of using one or more elements of a CRISPR system. The CRISPR complex of the invention provides an effective means for modifying a target polynucleotide. The CRISPR complex of the invention has broad utility including modification of target polynucleotides (e.g., deletions, insertions, translocations, inactivation, activation, base editing) in a number of cell types. Thus, the CRISPR complex of the invention has broad applications in, for example, gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence within the target polynucleotide.

[0138] VIII. Target Polynucleotide In one aspect, the invention provides a method of modifying a target polynucleotide in a eukaryotic cell, which can be in vivo, ex vivo or in vitro. In certain aspects, the method comprises sampling a cell or cell population from a human or non-human animal or plant (including microalgae), and modifying the cell or cells. Culturing can occur at any stage ex vivo. The cell or cells can be reintroduced into a non-human animal or plant (including microalgae).

[0139] Plant breeders utilize natural diversity to combine the most useful genes for desirable traits such as yield, quality, uniformity, tolerance, and resistance to pests. These desirable qualities include growth, day-length preference, temperature requirements, the starting date of flower or reproductive development, fatty acid content, insect resistance, disease tolerance, nematode resistance, fungal resistance, herbicide resistance, tolerance to drought, heat, humidity, cold, wind, and adverse soil conditions including high salinity. Sources of these useful genes include native or exotic varieties, heirloom varieties, wild plant relatives, and induced mutations, for example, treating plant material with a mutagen. Using the present invention, plant breeders are provided with new tools for inducing mutations. Thus, one skilled in the art can analyze the genome for sources of useful genes and use the present invention to induce an increase in useful genes in varieties having desired characteristics or traits, with higher precision than previous mutagenic agents, and thus accelerate and improve plant breeding programs.

[0140] The RGN-based target polynucleotide can be any polynucleotide that is endogenous or exogenous to a eukaryotic cell. For example, the target polynucleotide can be a polynucleotide that is present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence that encodes a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). Without wishing to be bound by theory, it is believed that the target sequence should bind to a PAM (protospacer adjacent motif), i.e., a short sequence recognized by the CRISPR complex. The exact sequence and length requirements of the PAM vary depending on the CRISPR enzyme used, but the PAM is typically a 2- to 5-base pair sequence adjacent to the protospacer (i.e., the target sequence).

[0141] The target polynucleotides of the CRISPR complex can include a number of disease-related genes and polynucleotides, as well as genes and polynucleotides related to signal transduction biochemical pathways. Examples of target polynucleotides include sequences related to signal transduction biochemical pathways, such as genes or polynucleotides related to signal transduction biochemical pathways. Examples of target polynucleotides include disease-related genes or polynucleotides. A "disease-related" gene or polynucleotide refers to any gene or polynucleotide that produces a transcript or translation product at abnormal levels or in abnormal forms in cells derived from diseased tissue compared to non-diseased control tissue or cells. This can be a gene that comes to be expressed at abnormally high levels or a gene that comes to be expressed at abnormally low levels, and the change in its expression correlates with the occurrence and / or progression of the disease. A disease-related gene also refers to a gene that is directly involved in or has a mutation or genetic variation in linkage disequilibrium with a gene involved in the etiology of the disease (e.g., a causative mutation). The transcribed or translated product may be known or unknown, and may further be at normal or abnormal levels. Examples of disease-related genes and polynucleotides are available from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.) (available on the World Wide Web).

[0142] The CRISPR system is particularly useful due to the relative ease of targeting genomic sequences of interest, but the question of what the RGN can do to address the causative mutation still remains. One approach is to generate a fusion protein between an RGN (preferably an inactive or nickase mutant of the RGN) and the active domain of a base editing enzyme such as a base editing enzyme or cytidine deaminase or adenosine deaminase base editing (U.S. Patent No. 9,840,699, incorporated herein by reference). In some embodiments, the method comprises (a) a fusion protein comprising an RGN of the invention and a base editing polypeptide such as a deaminase, (b) a gRNA targeting the fusion protein of (a), and contacting a target nucleotide sequence of a DNA strand, wherein the DNA molecule is contacted with the fusion protein and the gRNA under conditions and in an amount suitable for deamination of nucleotide bases. In one aspect, the target DNA sequence comprises a sequence associated with a disease or disorder, and deamination of the nucleotide base results in a sequence not associated with the disease or disorder. In some embodiments, the target DNA sequence is present in an allele of a crop plant, wherein a particular allele of the trait of interest results in a plant of low agricultural value. Deamination of the nucleotide base results in an allele that improves the trait and increases the agricultural value of the plant.

[0143] In one aspect, the DNA sequence comprises a T→C or A→G point mutation associated with a disease or disorder, and deamination of the mutant C or G base results in a sequence not associated with the disease or disorder. In some embodiments, the deamination corrects a point mutation in a sequence associated with a disease or disorder.

[0144] In some embodiments, the sequence associated with the disease or disorder encodes a protein, where deamination introduces a stop codon into the sequence associated with the disease or disorder, resulting in truncation of the encoded protein. In some embodiments, the contacting is performed in vivo in a subject having, or likely to be diagnosed with, the disease or disorder. In some embodiments, the disease or disorder is a disease associated with a point mutation or single nucleotide mutation in the genome. In some embodiments, the disease is a genetic disease, cancer, metabolic disease, or lysosomal storage disease.

[0145] Modification of Causative Mutations Using Base Editing An example of a genetically inherited disease that can be corrected using the approach relying on the RGN-base editing fusion protein of the present invention is Hurler syndrome. Hurler syndrome, also known as MPS-1, is a lysosomal storage disorder characterized at the molecular level by the accumulation of dermatan sulfate and heparan sulfate in lysosomes as a result of a deficiency in α-L-iduronidase (IDUA). This disease is generally a genetic disease caused by mutations in the IDUA gene that encodes α-L-iduronidase. Common IDUA mutations are W402X and Q70X, both of which are nonsense mutations that result in premature termination of translation. Such mutations are adequately addressed by precise genome editing (PGE) approaches. This is because a single base revertant, for example, by a base editing approach, restores the wild-type coding sequence, resulting in protein expression controlled by the endogenous regulatory mechanisms at the locus. Furthermore, since heterozygotes are known to be asymptomatic, PGE therapies targeting one of these mutations will be useful for most patients with this disease as they only need to correct one of the mutated alleles (Bunge et al. (1994) Hum. Mol. Genet. 3(6): 861-866, which is incorporated herein by reference).

[0146] Current treatments for Hurler syndrome include enzyme replacement therapy and bone marrow transplantation (Vellodi et al. (1997) Arch. Dis. Child. 76(2): 92-99; Peters et al. (1998) Blood 91(7): 2601-2608, which are incorporated herein by reference). Enzyme replacement therapy has had a dramatic impact on the survival and quality of life of Hurler syndrome patients, but this approach requires expensive and time-consuming weekly infusions. Additional approaches include delivery of the IDUA gene on an expression vector or insertion of the gene into a highly expressed locus such as serum albumin (U.S. Patent No. 9,956,247, which is incorporated herein by reference). However, these approaches cannot restore the original IDUA locus to the correct coding sequence. Genome editing strategies would have many advantages. In particular, regulation of gene expression would be controlled by natural mechanisms present in healthy individuals. Furthermore, using base editing does not require causing double-stranded DNA breaks, which can lead to large chromosomal rearrangements, cell death, or oncogenicity due to disruption of tumor suppressor mechanisms. A general strategy could be directed to targeting and correcting mutations that cause specific diseases in the human genome using the RGN-base editing fusion proteins of the present invention. It will be understood that similar approaches can also be pursued for target diseases that can be corrected by base editing. It will further be recognized that similar approaches targeting mutations that cause diseases in other species, particularly common household pets or livestock, can also be deployed using the RGNs of the present invention. Common household pets and livestock include dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, fish such as salmon, and shrimp.

[0147] Modification of the causative mutation by targeted deletion The RGN of the present invention may also be useful in human therapeutic approaches where the causative mutations are more complex. For example, several diseases such as Friedreich's ataxia and Huntington's disease are the result of a significant increase in the repetition of a 3-nucleotide motif in a specific region of a gene, which affects the function or expression ability of the expressed protein. Friedreich's ataxia (FRDA) is an autosomal recessive disorder that results in the progressive degeneration of nerve tissue in the spinal cord. A decrease in the mitochondrial frataxin (FXN) protein level causes oxidative damage and iron deficiency at the cellular level. The decrease in FXN expression is associated with the expansion of the GAA triplet within intron 1 of the somatic and germline FXN genes. In FRDA patients, the GAA repeats often exceed 70 and sometimes exceed 1000 (most commonly 600 - 900) triplets, while in non-affected individuals, they are about 40 repeats or less (all of which are incorporated herein by reference: Pandolfo et al. (2012) Handbook of Clinical Neurology 103: 275 - 294; Campuzano et al. (1996) Science 271: 1423 - 1427; Pandolfo (2002) Adv. Exp. Med. Biol. 516: 99 - 118).

[0148] The expansion of the 3-base repeat sequence that causes Friedreich's ataxia (FRDA) occurs at a specific locus within the FXN gene, called the FRDA unstable region. RNA-guided nuclease (RGN) can be used to excise the unstable region in FRDA patient cells. This approach requires 1) an RGN and a guide RNA sequence that can be programmed to target alleles in the human genome; and 2) a delivery approach for the RGN and the guide sequence. Many nucleases used in genome editing, such as the commonly used Cas9 nuclease (SpCas9) derived from S. pyogenes, are too large to be packaged into adeno-associated virus (AAV) vectors. This makes approaches using SpCas9 more difficult.

[0149] Certain RNA-guided nucleases of the present invention are suitable for packaging into an AAV vector together with a guide RNA. Although a second vector would be required to pack two guide RNAs, this approach is advantageous compared to what is required for larger nucleases such as SpCas9. SpCas9 may need to split the protein sequence between two vectors. The present invention encompasses strategies using the RGNs of the present invention in which regions of genomic instability are removed. Such strategies are applicable to other diseases and disorders having a similar genetic basis such as Huntington's disease. Similar strategies using the RGNs of the present invention are also applicable to similar diseases and disorders in agriculturally or economically important non-human animals including fish such as dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, salmon, and shrimp.

[0150] Modification of Causative Mutations by Targeted Mutagenesis The RGNs of the present invention can also introduce disruptive mutations that result in beneficial effects. Genetic deficiencies in genes encoding hemoglobin, particularly the β-globin chain (HBB gene), can cause many diseases known as abnormal hemoglobinopathies such as sickle cell anemia and thalassemia.

[0151] In adult humans, hemoglobin is a heterotetramer consisting of two alpha (α)-like globin chains, two beta (β)-like globin chains, and four heme groups. In adults, the α2β2 tetramer is called hemoglobin A (HbA) or adult hemoglobin. Typically, the α and β globin chains are synthesized in an approximate 1:1 ratio, which appears to be important for the stabilization of hemoglobin and red blood cells (RBCs). In the developing fetus, another form of hemoglobin, fetal hemoglobin (HbF), is produced, which has a higher binding affinity for oxygen than hemoglobin A and can deliver oxygen to the fetal system via the mother's bloodstream. Fetal hemoglobin also contains two α globin chains but has two fetal γ globin chains instead of the adult β globin chains (i.e., fetal hemoglobin is α2γ2). The regulation of the switch from γ to β globin production is very complex and involves mainly the downregulation of γ globin transcription accompanied by the simultaneous upregulation of β globin transcription. At about 30 weeks of gestation, the synthesis of gamma globin in the fetus begins to decline, while the production of β globin increases. By about 10 months after birth, almost all of the newborn's hemoglobin is α2β2, although some HbF persists into adulthood (about 1-3% of total hemoglobin). In the majority of patients with abnormal hemoglobinopathy, the gene encoding gamma globin is present, but its expression is relatively low due to the normal gene silencing that occurs around the time of birth as described above.

[0152] Sickle cell disease is caused by a V6E mutation (GAG to GTG at the DNA level) in the beta globin gene (HBB), and the resulting hemoglobin is called "hemoglobin S" or "HbS". Under low oxygen conditions, HbS molecules aggregate and form fibrous precipitates. These aggregates cause abnormalities or "sickling" of red blood cells, resulting in the loss of cell flexibility. Sickled red blood cells can no longer be pushed into the capillary bed and can cause vaso-occlusive crises in sickle cell patients. Additionally, sickled red blood cells are more fragile than normal red blood cells and tend to undergo hemolysis, ultimately causing anemia in the patient.

[0153] The treatment and management of sickle cell disease is a lifelong proposition that includes antibiotic treatment, pain management, and blood transfusions during acute episodes. One approach is the use of hydroxyurea, which exerts its effect in part by increasing the production of gamma globin. However, the long-term side effects of chronic hydroxyurea therapy are still unknown, the treatment can result in undesirable side effects, and there can be varying effectiveness from patient to patient. Despite the increasing effectiveness of sickle cell disease treatment, the average life expectancy of patients is only in the mid- to late 50s, and the morbidity associated with this disease has a significant impact on the quality of life of patients.

[0154] Thalassemia (α-thalassemia and β-thalassemia) is also a hemoglobin-related disorder, typically associated with reduced expression of globin chains. This can occur due to mutations in the regulatory regions of the genes, or mutations in the globin coding sequences that result in reduced or decreased levels or functional globin proteins. Treatment of thalassemia usually involves blood transfusions and iron chelation therapy. Bone marrow transplantation is also used in the treatment of severe thalassemia patients, but it is associated with significant risks if a suitable donor can be identified.

[0155] One approach that has been proposed for the treatment of SCD and β-thalassemia is to increase the expression of γ-globin so that abnormal adult hemoglobin is functionally replaced by HbF. As noted above, the treatment of SCD patients with hydroxyurea is thought to be successful in part due to its effect on increasing gamma-globin expression (DeSimone (1982) Proc Nat’l Acad Sci USA 79(14):4428-31; Ley, et al., (1982) N. Engl. J. Medicine, 307: 1469-1475; Ley, et al., (1983) Blood 62: 370-380; Constantoulakis et al., (1988) Blood 72(6):1961-1967, all of which are incorporated herein by reference). To increase the expression of HbF, it is necessary to identify genes that have products involved in the regulation of γ-globin expression. One such gene is BCL11A. BCL11A encodes a zinc finger protein that is expressed in adult erythroid progenitor cells, and down-regulation of its expression leads to an increase in gamma-globin expression (Sankaran et at (2008) Science 322: 1839, which is incorporated herein by reference). The use of inhibitory RNA targeting the BCL11A gene has been proposed (e.g., U.S. Patent Publication No. 2011 / 0182867, which is incorporated herein by reference), but this technique has several potential drawbacks, including the possibility that complete knockdown may not be achieved, the possibility that there are problems with the delivery of such RNA, and the need for the RNA to be continuously present and multiple treatments over a lifetime.

[0156] The RGNs of the present invention can be used to target the BCL11A enhancer region to disrupt the expression of BCL11A, thereby increasing the expression of gamma globin. This targeted disruption can be achieved by non-homologous end joining (NHEJ), whereby the RGNs of the present invention target specific sequences within the BCL11A enhancer region, create double-strand breaks, and the cellular machinery repairs the breaks, typically introducing deleterious mutations simultaneously. Similar to those described for other disease targets, the RGNs of the present invention may have advantages over other known RGNs due to their relatively small size, which allows the packaging of the expression cassettes for the RGNs and their guide RNAs into a single AAV vector for in vivo delivery. Similar strategies using the RGNs of the present invention are applicable to similar diseases and disorders in both humans and agriculturally or economically important non-human animals.

[0157] IX. Cells Comprising Genetically Modified Polynucleotides Provided herein are cells and organisms comprising a target sequence modified using a process mediated by an RGN, a crRNA, and / or a tracrRNA as described herein. In some of these embodiments, the RGN comprises the amino acid sequence of SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, 137, or 235, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence of SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, 240, 273, or 287, or an active variant or fragment thereof. In certain embodiments, the guide RNA comprises a tracrRNA comprising the nucleotide sequence of SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, 241, 274, or 286, or an active variant or fragment thereof. The guide RNA of the system may be a single guide RNA or a dual guide RNA.

[0158] The modified cell can be a eukaryotic cell (e.g., mammalian, plant, insect cell) or a prokaryotic cell. Also provided are organelles and embryos comprising at least one nucleotide sequence modified by a method utilizing RGN, crRNA, and / or tracrRNA as described herein. Genetically engineered cells, organisms, organelles, embryos can be heterozygous or homozygous with respect to the modified nucleotide sequence.

[0159] Chromosomal modification of a cell, organism, organelle, or embryo can result in a change in expression (upregulation or downregulation), inactivation, or expression of a modified protein product, or an integrated sequence. When the chromosomal modification results in either gene inactivation or expression of a non-functional protein product, the genetically modified cell, organism, organelle, or embryo is referred to as a "knockout". A knockout phenotype can be the result of a deletion mutation (i.e., deletion of at least one nucleotide), an insertion mutation (i.e., insertion of at least one nucleotide), or a nonsense mutation (i.e., substitution of at least one nucleotide such that a stop codon is introduced).

[0160] Alternatively, chromosomal modification of a cell, organism, organelle, or embryo can generate a "knock-in" resulting from chromosomal integration of a nucleotide sequence encoding a protein. In some of these embodiments, the coding sequence is integrated into the chromosome such that the chromosomal sequence encoding the wild-type protein is inactivated, but the exogenously introduced protein is expressed.

[0161] In other embodiments, the chromosomal modification results in the production of a mutant protein product. The expressed mutant protein product can have at least one amino acid substitution and / or at least one amino acid addition or deletion. The modified protein product encoded by the modified chromosomal sequence can exhibit modified characteristics or activities as compared to the wild-type protein, including, but not limited to, modified enzyme activity or substrate specificity.

[0162] In yet other embodiments, the chromosomal modification can result in an altered expression pattern of a protein. By way of non-limiting example, chromosomal changes in regulatory regions that control the expression of a protein product can result in overexpression or down-regulation of the protein product, or an altered tissue or temporal expression pattern.

[0163] X. Kits and Methods for Detecting Target DNA or Single-Stranded DNA Some RGNs (e.g., APG09106.1 and APG09748 shown as SEQ ID NO: 54 and 137) can cleave non-target single-stranded DNA (ssDNA) indiscriminately once activated by the detection of target DNA. Accordingly, provided herein are compositions and methods for detecting target DNA (double-stranded or single-stranded) in a sample. In some embodiments, the desired target can be RNA, such as a genome, or a portion of the genome of an RNA virus, such as may be present as a coronavirus. In one aspect, the coronavirus can be a SARS-like coronavirus. In a further aspect, the coronavirus can be SARS-CoV-2, SARS-CoV, or a bat SARS-like coronavirus, such as, for example, bat-SL-CoVZC45 (accession MG772933). In embodiments where the target is present as RNA, the target can be reverse transcribed into a DNA molecule that can be effectively targeted by an RGN. Following reverse transcription, an amplification step such as RT-PCR methods known in the art that include thermal cycling, or an isothermal method such as RT-LAMP (reverse transcription loop-mediated isothermal amplification) can be used (Notomi et al., Nucleic Acids Res 28: E63, (2000)).

[0164] These compositions and methods involve the use of a detection ssDNA that does not hybridize to the guide RNA and is non-target ssDNA. In some embodiments, the detection ssDNA includes a detectable label that provides a detectable signal after cleavage of the detection ssDNA. Non-limiting examples are detection ssDNAs that include a fluorophore / quencher pair, where the detection ssDNA does not fluoresce when its signal is suppressed by the presence of a very proximal quencher and is intact (i.e., not left). Cleavage of the detection ssDNA results in removal of the quencher, and then the fluorescent label can be detected. Non-limiting examples of fluorescent labels or dyes include Cy5, fluorescein (e.g., FAM, 6 FAM, 5(6)FAM, FITC), Cy3, Alexa Fluor® dyes, and Texas Red. Non-limiting examples of quenchers include Iowa Black® FQ, Iowa Black® RQ, Qx1 quencher, ATT0 quencher, and QSY dyes. In some embodiments, the detection ssDNA includes a second quencher, such as an internal quencher like ZEN™, TAO™, and Black Hole Quencher®, which can reduce background and increase signal detection.

[0165] In other embodiments, the detected ssDNA comprises a detectable label that provides a detectable signal before cleavage of the detected ssDNA and before cleavage of the ssDNA inhibits or prevents detection of the signal. Non-limiting examples of such scenarios are detected ssDNAs that include a fluorescence resonance energy transfer (FRET) pair. FRET is a process in which non-radiative transfer of energy from the excited state of a first (donor) fluorophore to a second (acceptor) fluorophore occurs when they are in very close proximity. The emission spectrum of the donor phosphor overlaps with the excitation spectrum of the acceptor phosphor. Thus, the acceptor phosphor will fluoresce when the detected ssDNA is intact (i.e., non-cleaved), and the acceptor phosphor will no longer fluoresce when the detected ssDNA is cleaved because the donor and acceptor phosphors are no longer in close proximity to each other. FRET donor and acceptor fluorophores are known in the art and include, but are not limited to, cyan fluorescent protein / green fluorescent protein (GFP), Cy3 / Cy5, and GFP / yellow fluorescent protein (YFP).

[0166] In some embodiments, the detector ssDNA has a length of from about 2 nucleotides to about 30 nucleotides and includes, but is not limited to, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25 nucleotides, about 26 nucleotides, about 27 nucleotides, about 28 nucleotides, about 29 nucleotides, and about 30 nucleotides.

[0167] A method for detecting a target DNA of a DNA molecule includes contacting a sample with an RGN, a guide RNA capable of hybridizing the sample with the target DNA sequence in the RGN and the DNA molecule, and a detectable single-stranded DNA (ssDNA) that does not hybridize with the guide RNA, and then measuring a detectable signal generated by cleavage of the ssDNA by the RGN, thereby detecting the target DNA sequence of the DNA molecule. In some embodiments, the method can include an amplification step of nucleic acid molecules in the sample before or simultaneously with the contact with the RGN and the guide RNA. In some of these embodiments, the specific sequence to which the guide RNA hybridizes can be amplified to enhance the sensitivity of the detection method.

[0168] Samples in which the target DNA can be detected using these compositions, and methods comprising the detection of ssDNA, include any sample that contains, or is considered to contain, a nucleic acid (e.g., a DNA or RNA molecule). The sample can be derived from any source including a combination of purified nucleic acid synthesis, or a biological sample such as a respiratory swab (e.g., nasopharyngeal swab) extract, cell lysate, patient sample, cell, tissue, saliva, blood, serum, plasma, urine, aspirate, biopsy sample, cerebrospinal fluid, or an organism (such as bacteria, virus, etc.).

[0169] The contact of the sample with the RGN, the guide RNA, and the detectable ssDNA can include contact in vitro, ex vivo, or in vivo. In some embodiments, the detectable ssDNA and / or the RGN and / or the guide RNA are immobilized, for example, on a lateral flow device, where the sample contacts the immobilized detectable ssDNA and / or the RGN and / or the guide RNA. In one aspect, an antibody against an antigenic portion on the detectable ssDNA is immobilized, for example, on a lateral flow device, in a manner that enables discrimination between the cleaved detectable ssDNA and the intact detectable ssDNA.

[0170] In some embodiments, the method can further include determining the amount of target DNA present in the sample. The measurement of the detectable signal in the test sample can be compared to a reference measurement (e.g., measurements of a reference sample or a series thereof containing a known amount of target DNA).

[0171] Non-limiting examples of the applications of the compositions and methods include single nucleotide polymorphism (SNP) detection, cancer screening, detection of bacterial infections, detection of antibiotic resistance, and detection of viral infections.

[0172] The detectable signal generated by cleavage of ssDNA by the RGN can be measured using any suitable method known in the art, including, but not limited to, measurement of a fluorescent signal, visual analysis of a band on a gel, colorimetric change, and the presence or absence of an electrical signal.

[0173] The present invention provides a kit for detecting a target DNA of a DNA molecule in a sample, the kit comprising an RGN polypeptide, a guide RNA capable of hybridizing to the RGN, a target DNA sequence in the DNA molecule, and a detection ssDNA that does not hybridize to the guide RNA.

[0174] Also provided herein is a method of cleaving single-stranded DNA by contacting a population of nucleic acids, wherein the population comprises a target DNA sequence of a DNA molecule and a plurality of non-target ssDNAs, an RGN, and a guide RNA capable of hybridizing to the RGN and the target DNA sequence.

[0175] The articles "a" and "an" are used herein to refer to one or more (i.e., at least one) of the grammatical objects of the article. By way of example, "polypeptide" means one or more polypeptides.

[0176] All publications and patent applications cited in the specification are indicative of the level of those skilled in the art to which the present disclosure pertains. All publications and patent applications are incorporated herein by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.

[0177] The foregoing invention has been described in some detail for purposes of clarity of understanding by way of illustration and example, but it will be apparent that certain changes and modifications may be practiced within the scope of the appended embodiments.

[0178] Non-limiting embodiments include the following: 1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117, wherein the RGN polypeptide can bind to a target DNA sequence in an RNA guide sequence-specific manner when bound to a guide RNA (gRNA) that can hybridize to the target DNA sequence, and the polynucleotide encoding the RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide in the nucleic acid molecule. 2. The nucleic acid molecule according to embodiment 1, wherein the RGN polypeptide can cleave the target DNA sequence upon binding. 3. The nucleic acid molecule according to embodiment 2, wherein the RGN polypeptide can generate a double-strand break. 4. The nucleic acid molecule according to embodiment 2, wherein the RGN polypeptide can generate a single-strand break. 5. The nucleic acid molecule according to embodiment 2, wherein the RGN polypeptide is nuclease-inactive.

[0179] The nucleic acid molecule according to any one of embodiments 1 to 5, wherein the RGN polypeptide is functionally fused to a base editing polypeptide. 7. The nucleic acid molecule according to embodiment 6, wherein the base editing polypeptide is a deaminase. 8. The nucleic acid molecule according to embodiment 7, wherein the deaminase is cytidine deaminase or adenosine deaminase. 9. The nucleic acid molecule according to any one of embodiments 1 to 8, wherein the RGN polypeptide comprises one or more nuclear localization signals. 10. The nucleic acid molecule according to any one of embodiments 1 to 9, wherein the RGN polypeptide has codons optimized for expression in eukaryotic cells.

[0180] 11. The nucleic acid molecule according to any one of embodiments 1 to 10, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM). 12. A vector comprising the nucleic acid molecule according to any one of embodiments 1 to 11. 13. The vector according to embodiment 12, further comprising at least one nucleotide sequence encoding the gRNA capable of hybridizing to the target DNA sequence. 14. The vector according to embodiment 13, wherein the guide RNA comprises a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118. 15. The vector according to embodiment 13 or 14, wherein the gRNA comprises a tracrRNA.

[0181] 16. The vector according to embodiment 15, wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119. 17. The vector according to embodiment 15 or 16, wherein the gRNA is a single guide RNA. 18. The vector according to embodiment 15 or 16, wherein the gRNA is a dual guide RNA. 19. A cell comprising the nucleic acid molecule according to any one of embodiments 1 to 11, or the vector according to any one of embodiments 12 to 18. 20. A method for producing an RGN polypeptide, comprising culturing the cell according to embodiment 18 under conditions in which the RGN polypeptide is expressed.

[0182] 21. Introducing into a cell a heterologous nucleic acid molecule comprising a nucleotide sequence encoding an RNA-guided nuclease (RGN) polypeptide having an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117, wherein the RGN polypeptide can bind to a target DNA sequence of a DNA molecule in an RNA guide sequence-specific manner when bound to a guide RNA (gRNA) capable of hybridizing to the target DNA sequence, culturing the cell under conditions in which the RGN polypeptide is expressed A method for producing an RGN polypeptide, comprising the steps of: 22. The method according to embodiment 20 or 21, further comprising purifying the RGN polypeptide. 23. The method according to embodiment 20 or 21, wherein the cell further expresses one or more guide RNAs that bind to the RGN polypeptide to form an RGN ribonucleoprotein complex. 24. The method according to embodiment 23, further comprising a step of purifying the RGN ribonucleoprotein complex. 25. A nucleic acid molecule comprising a polynucleotide encoding a CRISPR RNA (crRNA), wherein the crRNA comprises a spacer sequence and a CRISPR repeat sequence, and the CRISPR repeat sequence comprises a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118, a) the crRNA; and, optionally, b) A trans-activating CRISPR RNA (tracrRNA) capable of hybridizing with the CRISPR repeat sequence of the crRNA When the guide RNA containing the same binds to an RNA-guided nuclease (RGN) polypeptide, it can hybridize sequence-specifically to the target DNA sequence of a DNA molecule via the spacer sequence of the crRNA, The polynucleotide encoding the crRNA is the above nucleic acid molecule operably linked to a promoter heterologous to the polynucleotide.

[0183] 26. A vector comprising the nucleic acid molecule according to embodiment 25. 27. The vector according to embodiment 26, wherein the vector further comprises a polynucleotide encoding the tracrRNA. 28. The vector according to embodiment 27, wherein the tracrRNA comprises a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119. 29. The vector according to embodiment 27 or 28, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and encoded as a single guide RNA. 30. The vector according to embodiment 27 or 28, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.

[0184] 31. The vector according to any one of embodiments 26 to 30, wherein the vector further comprises a polynucleotide encoding the RGN polypeptide, and the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117. A nucleic acid molecule comprising a polynucleotide encoding a trans-activating CRISPR RNA (tracrRNA) having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119, a) said tracrRNA; and b) a crRNA comprising a spacer sequence and a CRISPR repeat sequence, wherein the crRNA is capable of hybridizing to the CRISPR repeat sequence of said crRNA A guide RNA comprising, when bound to an RNA-guided nuclease (RGN) polypeptide, is capable of hybridizing sequence-specifically to a target DNA sequence of a DNA molecule via the spacer sequence of the crRNA, The polynucleotide encoding the tracrRNA is the above nucleic acid molecule operably linked to a promoter heterologous to the polynucleotide. The above nucleic acid molecule comprising. 33. A vector comprising the nucleic acid molecule according to embodiment 32. 34. The vector according to embodiment 33, wherein the vector further comprises a polynucleotide encoding the crRNA. 35. The vector according to embodiment 34, wherein the CRISPR repeat sequence of the crRNA comprises a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118.

[0185] 36. The vector according to embodiment 34 or 35, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and encoded as a single guide RNA. 37. The vector according to embodiment 34 or 35, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters. 38. The vector according to any one of embodiments 33 to 37, further comprising a polynucleotide encoding the RGN polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110 or 117. 39. A system for binding to a target DNA sequence of a DNA molecule, the system comprising: a) one or more polynucleotides comprising a nucleotide sequence encoding one or more guide RNAs (gRNAs) or one or more guide RNAs capable of hybridizing to the target DNA sequence; and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117; or a polynucleotide comprising a nucleotide sequence encoding the RGN polypeptide comprising, the nucleotide sequences encoding one or more guide RNAs and encoding the RNAGN polypeptide are each ligated to a promoter heterologous to the nucleotide sequence so as to be functional; the above system, wherein one or more guide RNAs can form a complex with the RGN polypeptide to bind the RGN polypeptide to the target DNA sequence of the DNA molecule. 40. The system according to embodiment 39, wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118.

[0186] 41. The system according to embodiment 39 or 40, wherein the gRNA comprises a tracrRNA. 42. The system according to embodiment 41, wherein the tracrRNA comprises a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119. 43. The system according to embodiment 41 or 42, wherein the gRNA is a single guide RNA (sgRNA). 44. The system according to embodiment 41 or 42, wherein the gRNA is a dual guide RNA. 45. The system according to any one of embodiments 39 to 44, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM).

[0187] 46. The system according to any one of embodiments 39 to 45, wherein the target DNA sequence is intracellular. 47. The system according to embodiment 46, wherein the cell is a eukaryotic cell. 48. The system according to embodiment 47, wherein the eukaryotic cell is a plant cell. 49. The system according to embodiment 47, wherein the eukaryotic cell is a mammalian cell. 50. The system according to embodiment 47, wherein the eukaryotic cell is an insect cell.

[0188] 51. The system according to embodiment 46, wherein the cell is a prokaryotic cell. 52. The system according to any one of embodiments 39 to 51, wherein when one or more guide RNAs are transcribed, they can hybridize to the target DNA sequence, and the guide RNA can form a complex with the RGN polypeptide to directly cleave the target DNA sequence. 53. The system according to embodiment 52, wherein the RGN polypeptide can generate a double-strand break. 54. The system according to embodiment 52, wherein the RGN polypeptide can generate a single-strand break. 55. The system according to embodiment 52, wherein the RGN polypeptide is nuclease-inactive.

[0189] The system according to any one of embodiments 39 to 55, wherein the RGN polypeptide is linked so as to be able to function as a base editing polypeptide. 57. The system according to embodiment 56, wherein the base editing polypeptide is a deaminase. 58. The system according to embodiment 57, wherein the deaminase is cytidine deaminase or adenosine deaminase. 59. The system according to any one of embodiments 39 to 58, wherein the RGN polypeptide comprises one or more nuclear localization signals. 60. The system according to any one of embodiments 39 to 59, wherein the RGN polypeptide has codons optimized for expression in eukaryotic cells.

[0190] 61. The system according to any one of embodiments 39 to 60, wherein a polynucleotide comprising a nucleotide sequence encoding one or more guide RNAs and a polynucleotide comprising a nucleotide sequence encoding an RGN polypeptide are located on one vector. 62. The system according to any one of embodiments 39 to 61, wherein the system further comprises one or more donor polynucleotides, or one or more nucleotide sequences encoding one or more donor polynucleotides. 63. A method of binding to a target DNA sequence of a DNA molecule, comprising delivering the system according to any one of embodiments 39 to 62 to a target DNA sequence or a cell comprising the target DNA sequence. 64. The method according to embodiment 63, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby enabling detection of the target DNA sequence. 65. The method according to embodiment 63, wherein the guide RNA or the RGN polypeptide further comprises an expression modulator, thereby regulating the expression of the target DNA sequence or a gene under the transcriptional control of the target DNA sequence.

[0191] 66. A method for cleaving or modifying a target DNA sequence of a DNA molecule, comprising delivering the system according to any one of Embodiments 39 to 62 to a target DNA sequence or a cell containing the DNA molecule, and causing cleavage or modification of the target DNA sequence. 67. The method according to Embodiment 66, wherein the modified target DNA sequence comprises insertion of a heterologous DNA into the target DNA sequence. 68. The method according to Embodiment 66, wherein the modified target DNA sequence comprises deletion of at least one nucleotide from the target DNA sequence. 69. The method according to Embodiment 66, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence. 70. A method for binding to a target DNA sequence of a DNA molecule, the method comprising: a) under conditions suitable for the formation of an RGN ribonucleotide complex, i) one or more guide RNAs capable of hybridizing to the target DNA sequence; and ii) an RGN polypeptide comprising an amino acid having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117 are combined to assemble an RNA-guided nuclease (RGN) ribonucleotide complex in vitro; and b) contacting the target DNA sequence or a cell containing the target DNA sequence with the RGN ribonucleotide complex assembled in vitro comprising wherein one or more guide RNAs hybridize to the target DNA sequence, thereby binding the RGN polypeptide to the target DNA sequence.

[0192] 71. The method according to Embodiment 70, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby enabling detection of the target DNA sequence. 72. The method according to embodiment 70, wherein the guide RNA or the RGN polypeptide further comprises an expression modulator, thereby enabling regulation of the expression of the target DNA sequence. 73. A method for cleaving and / or modifying a target DNA sequence of a DNA molecule, the method comprising contacting the DNA molecule with a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN comprises an amino acid having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117; and b) one or more guide RNAs capable of targeting the RGN of (a) to the target DNA sequence; comprising contacting with the following, wherein one or more guide RNAs hybridize to the target DNA sequence, thereby binding the RGN polypeptide to the target DNA sequence, and cleavage and / or modification of the target DNA sequence occurs. 74. The method according to embodiment 73, wherein cleavage by the RGN polypeptide generates a double-strand break. 75. The method according to embodiment 73, wherein cleavage by the RGN polypeptide generates a single-strand break.

[0193] 76. The method according to embodiment 73, wherein the RGN polypeptide is nuclease-inactive. 77. The method according to any one of embodiments 73 to 76, wherein the RGN polypeptide is operably linked to a base editing polypeptide. 78. The method according to embodiment 77, wherein the base editing polypeptide comprises a deaminase. 79. The method according to embodiment 78, wherein the deaminase is cytidine deaminase or adenosine deaminase. 80. The method according to any one of embodiments 73 to 79, wherein the modified target DNA sequence comprises insertion of a heterologous DNA into the target DNA sequence.

[0194] 81. The method according to any one of embodiments 73 to 79, wherein the modified target DNA sequence comprises a deletion of at least one nucleotide from the target DNA sequence. 82. The method according to any one of embodiments 73 to 79, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence. 83. The method according to any one of embodiments 70 to 82, wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118. 84. The method according to any one of embodiments 70 to 83, wherein the gRNA comprises a tracrRNA. 85. The method according to embodiment 84, wherein the tracrRNA comprises a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119.

[0195] 86. The method according to embodiment 84 or 85, wherein the gRNA is a single guide RNA (sgRNA). 87. The method according to embodiment 84 or 85, wherein the gRNA is a dual guide RNA. 88. The method according to any one of aspects 63 to 87, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM). 89. The method according to any one of embodiments 63 to 88, wherein the target DNA sequence is intracellular. 90. The method according to embodiment 89, wherein the cell is a eukaryotic cell.

[0196] 91. The method according to embodiment 90, wherein the eukaryotic cell is a plant cell. 92. The method according to embodiment 90, wherein the eukaryotic cell is a mammalian cell. 93. The method according to embodiment 90, wherein the eukaryotic cell is an insect cell. 94. The method according to embodiment 89, wherein the cell is a prokaryotic cell. 95. Culturing cells under conditions where a 95.RGN polypeptide is expressed and cleaves a target DNA sequence to generate a DNA molecule containing a modified DNA sequence; and further comprising selecting cells containing the modified target DNA sequence, the method according to any one of embodiments 89 to 94.

[0197] 96. A cell containing a target DNA sequence modified according to the method of embodiment 95. 97. The cell according to embodiment 96, wherein the cell is a eukaryotic cell. 98. The cell according to embodiment 97, wherein the eukaryotic cell is a plant cell. 99. A plant comprising the cell according to embodiment 98. 100. A seed comprising the cell according to embodiment 98.

[0198] 101. The cell according to embodiment 97, wherein the eukaryotic cell is a mammalian cell. 102. The cell according to embodiment 98, wherein the eukaryotic cell is an insect cell. 103. The cell according to embodiment 96, wherein the cell is a prokaryotic cell. 104. A method for producing a genetically modified cell in which a mutation causing a genetic disease is corrected, the method comprising introducing into the cell a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117; or a polynucleotide encoding said RGN polypeptide, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter to enable expression of the RGN polypeptide in the cell; and b) A guide RNA (gRNA), wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118; or a polynucleotide encoding said gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to enable expression of the gRNA in a cell comprising introducing Thereby, the method as described above, wherein the RGN and the gRNA target the genomic position of the causative mutation and modify the genomic sequence to remove the causative mutation 105. The method according to embodiment 104, wherein the RGN is operably linked to a base editing polypeptide

[0199] 106. The method according to embodiment 105, wherein the base editing polypeptide is a deaminase 107. The method according to embodiment 106, wherein the deaminase is a cytidine deaminase or an adenosine deaminase 108. The method according to any one of embodiments 104 to 107, wherein the cell is an animal cell 109. The method according to any one of embodiments 104 to 107, wherein the cell is a mammalian cell 110. The method according to embodiment 108, wherein the cell is derived from a dog, a cat, a mouse, a rat, a rabbit, a horse, a cow, a pig, or a human

[0200] 111. The method according to embodiment 108, wherein the genetic disease is caused by a single nucleotide polymorphism 112. The method according to embodiment 111, wherein the genetic disease is Haller syndrome 113. The method according to embodiment 112, wherein the gRNA further comprises a spacer sequence targeting a region proximal to the causative single nucleotide polymorphism 114. A method for generating a recombinant cell having a deletion in an unstable genomic region causing a disease, the method comprising introducing into a cell a) An RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117; or a polynucleotide encoding the RGN polypeptide, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter to enable expression of the RGN polypeptide in a cell; and b) A guide RNA (gRNA), wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118; or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to enable expression of the gRNA in a cell, and further, the gRNA comprises a spacer sequence targeting the 5' flank of an unstable genomic region; and c) A second guide RNA (gRNA), wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118; or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to enable expression of the second gRNA in a cell, and further, the second gRNA comprises a spacer sequence targeting the 3' flank of an unstable genomic region comprising introducing the same, whereby the RGN and the two gRNAs target an unstable genomic region and at least a part of the unstable genomic region. 115. The method according to embodiment 114, wherein the cell is an animal cell.

[0201] 116. The method according to embodiment 114, wherein the cell is a mammalian cell. 117. The method according to embodiment 115, wherein the cell is derived from a dog, cat, mouse, rat, rabbit, horse, cow, pig, or human. 118. The method according to embodiment 115, wherein the genetically inherited disease is Friedreich's ataxia or Huntington's disease. 119. The method according to embodiment 118, wherein the first gRNA further comprises a spacer sequence targeting a region within or near an unstable genomic region. 120. The method according to embodiment 118, wherein the second gRNA further comprises a spacer sequence targeting a region within or proximal to an unstable genomic region.

[0202] 121. A method for producing a genetically modified mammalian hematopoietic progenitor cell with reduced BCL11A mRNA and protein expression, the method comprising introducing into isolated human hematopoietic progenitor cells: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117; or a polynucleotide encoding the RGN polypeptide, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter to enable expression of the RGN polypeptide in the cell; and b) a guide RNA (gRNA), wherein the gRNA comprises a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, or 118; or a polynucleotide encoding the gRNA, wherein the polynucleotide encoding the gRNA is operably linked to a promoter to enable expression of the gRNA in the cell. comprising introducing, whereby the RGN and gRNA are expressed intracellularly and cleaved at the BCL11A enhancer region, resulting in genetic modification of human hematopoietic progenitor cells and reduced BCL11A mRNA and / or protein expression, the method as described above. 122. The method according to embodiment 121, wherein the gRNA further comprises a spacer sequence targeting a region within or proximal to the BCL11A enhancer region. 123. The method according to any one of embodiments 104 to 122, wherein the guide RNA comprises a tracrRNA comprising a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, or 119. 124. A system for binding to a target DNA sequence of a DNA molecule, the system comprising a) one or more polynucleotides comprising one or more nucleotide sequences encoding one or more guide RNAs (gRNAs) or one or more guide RNAs capable of hybridizing to the target DNA sequence; and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117 comprising, wherein the one or more guide RNAs are capable of hybridizing to the target DNA sequence, the system as described above, wherein the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to bind the RGN polypeptide to the target DNA sequence of the DNA molecule. 125. The system according to embodiment 124, wherein the RGN polypeptide is nuclease-inactive or can function as a nickase.

[0203] 126. The system according to embodiment 124 or 125, wherein the RGN polypeptide is functionally fused to a base editing polypeptide. 127. The system according to embodiment 126, wherein the base editing polypeptide is a deaminase. 128. The system according to embodiment 127, wherein the deaminase is cytidine deaminase or adenosine deaminase. 129. A method for detecting a target DNA sequence of a DNA molecule in a sample, the method comprising: a) contacting the sample with i) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, or 137, wherein the RGN polypeptide can bind to a target DNA sequence of a DNA molecule in an RNA-guided sequence-specific manner when bound to a guide RNA that can hybridize to the target DNA sequence; ii) the guide RNA iii) a detection single-stranded DNA (ssDNA) that does not hybridize to the guide RNA ; and b) measuring a detectable signal generated by cleavage of the detection ssDNA by the RGN, thereby detecting the target DNA. 130. The method according to embodiment 129, wherein the sample comprises DNA molecules derived from a cell lysate.

[0204] 131. The method according to embodiment 129, wherein the sample comprises cells. 132. The method according to embodiment 131, wherein the cells are eukaryotic cells. 133. The method according to embodiment 129, wherein the DNA molecule comprising the target DNA sequence is generated by reverse transcription of an RNA template molecule present in a sample containing RNA. 134. The method according to embodiment 133, wherein the RNA template molecule is an RNA virus. 135. The method according to embodiment 134, wherein the RNA virus is a coronavirus.

[0205] 136. The method according to embodiment 135, wherein the coronavirus is a bat SARS-like coronavirus, SARS-CoV, or SARS-CoV-2. 137. The method according to any one of embodiments 133 to 136, wherein the sample containing RNA is derived from a sample containing cells. 138. The method according to any one of embodiments 129 to 137, wherein the detected ssDNA contains a fluorophore / quencher pair. 139. The method according to any one of embodiments 129 to 137, wherein the detected ssDNA contains a fluorescence resonance energy transfer (FRET) pair. 140. The method according to any one of embodiments 129 to 139, wherein the guide RNA contains a CRISPR repeat sequence having a nucleotide sequence with at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, or 273.

[0206] 141. The method according to any one of embodiments 129 to 140, wherein the guide RNA contains a tracrRNA. 142. The method according to embodiment 141, wherein the tracrRNA contains a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, or 274. 143. The method according to embodiment 141 or 142, wherein the guide RNA is a single guide RNA. 144. The method according to embodiment 141 or 142, wherein the guide RNA is a dual guide RNA. 145. The method according to any one of embodiments 129 to 144, further comprising amplifying the nucleic acid in the sample before or together with the contact in step a.

[0207] 146. The method according to embodiment 145, wherein the amplification is a DNA molecule generated by reverse transcription of an RNA molecule. 147. A kit for detecting a target DNA sequence of a DNA molecule in a sample, the kit comprising a) An RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, or 137, wherein the RGN polypeptide can bind to a target DNA sequence of a DNA molecule in an RNA-guided sequence-specific manner when bound to a guide RNA that can hybridize to the target DNA sequence; b) The guide RNA; and c) A detection single-stranded DNA (ssDNA) that does not hybridize to the guide RNA The kit comprising the above. 148. The kit according to embodiment 147, wherein the detection ssDNA comprises a fluorophore / quencher pair. 149. The kit according to embodiment 147, wherein the detection ssDNA comprises a fluorescence resonance energy transfer (FRET) pair. 150. The kit according to any one of embodiments 147 to 149, wherein the guide RNA comprises a CRISPR repeat sequence having a nucleotide sequence with at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, or 273.

[0208] 151. The kit according to any one of embodiments 147 to 150, wherein the guide RNA comprises a tracrRNA. 152. The kit according to embodiment 151, wherein the tracrRNA comprises a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, or 274. 153. The kit according to embodiment 151 or 152, wherein the guide RNA is a single guide RNA. 154. The kit according to embodiment 151 or 152, wherein the guide RNA is a dual guide RNA. 155. A method for cleaving single-stranded DNA, the method comprising a DNA molecule comprising a target DNA sequence and a nucleic acid population comprising a plurality of non-target ssDNAs a) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, 117, or 137, wherein the RGN polypeptide is capable of binding to a guide RNA that can hybridize to the target DNA sequence and, when bound to the guide RNA, is capable of binding to the target DNA sequence in an RNA-guided sequence-specific manner; and b) contacting with the guide RNA The method of claim 1, wherein the RGN polypeptide cleaves the plurality of non-target ssDNAs.

[0209] 156. The method of embodiment 155, wherein the nucleic acid population is within a cell lysate. 157. The method of embodiment 155 or 156, wherein the guide RNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 2, 10, 17, 24, 31, 39, 47, 55, 62, 70, 76, 83, 90, 96, 104, 111, 118, or 273. 158. The method of any one of embodiments 155-157, wherein the guide RNA comprises a tracrRNA. 159. The method of embodiment 158, wherein the tracrRNA comprises a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 3, 11, 18, 25, 32, 40, 48, 56, 63, 71, 77, 84, 91, 97, 105, 112, 119, or 274. 160. The method of embodiment 158 or 159, wherein the guide RNA is a single guide RNA.

[0210] 161. The method of embodiment 158 or 159, wherein the guide RNA is a dual guide RNA. A nucleic acid molecule comprising a polynucleotide encoding a CRISPR RNA (crRNA), wherein the crRNA comprises a spacer sequence and a CRISPR repeat sequence, and the CRISPR repeat sequence comprises a nucleotide sequence having at least 95% sequence identity with SEQ ID NO: 240, a) the crRNA; and optionally b) a trans-activating CRISPR RNA (tracrRNA) that can hybridize with the CRISPR repeat sequence of the crRNA A guide RNA comprising the same, when bound to an RNA-guided nuclease (RGN) polypeptide, can hybridize sequence-specifically to a target DNA sequence of a DNA molecule via the spacer sequence of the crRNA, The polynucleotide encoding the crRNA is a nucleic acid molecule operably linked to a promoter heterologous to the polynucleotide. 163. A vector comprising the nucleic acid molecule according to embodiment 162. 164. The vector according to embodiment 163, wherein the vector further comprises a polynucleotide encoding the tracrRNA. 165. The vector according to embodiment 164, wherein the tracrRNA comprises a nucleotide sequence having at least 95% sequence identity with SEQ ID NO: 241.

[0211] 166. The vector according to embodiment 164 or 165, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and encoded as a single guide RNA. 167. The vector according to embodiment 164 or 165, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters. 168. The vector according to any one of embodiments 163 to 167, further comprising a polynucleotide encoding the RGN polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 235. 169. A nucleic acid molecule comprising a polynucleotide encoding a trans-activating CRISPR RNA (tracrRNA) comprising a nucleotide sequence having at least 95% sequence identity with SEQ ID NO: 241, a) the tracrRNA; and b) a crRNA comprising a spacer sequence and a CRISPR repeat sequence, wherein the tracrRNA is capable of hybridizing with the CRISPR repeat sequence of the crRNA The guide RNA comprising the above is capable of specifically hybridizing to a target DNA sequence of a DNA molecule via the spacer sequence of the crRNA when bound to an RNA-guided nuclease (RGN) polypeptide, and the polynucleotide encoding the tracrRNA is functionally linked to a promoter heterologous to the polynucleotide. 170. A vector comprising the nucleic acid molecule according to embodiment 169.

[0212] 171. The vector according to embodiment 170, wherein the vector further comprises a polynucleotide encoding the crRNA. 172. The vector according to embodiment 171, wherein the CRISPR repeat sequence of the crRNA comprises a nucleotide sequence having at least 95% sequence identity with SEQ ID NO: 240. 173. The vector according to embodiment 171 or 172, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and encoded as a single guide RNA. 174. The vector according to embodiment 171 or 172, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters. 175. The vector according to any one of embodiments 170 to 174, wherein the vector further comprises a polynucleotide encoding the RGN polypeptide, and the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 235.

[0213] The following examples are provided by way of illustration and not limitation.

Example

[0214] Example 1: Identification of RNA-guided nucleases 17 types of CRISPR-related RNA-guided nucleases (RGNs) were identified and are listed in Table 1 below. Table 1 shows the name, amino acid sequence, origin, processed crRNA and tracrRNA sequences of each RGN (see Example 2 for the identification method). Table 1 further provides a general single-guide RNA (sgRNA) sequence, where poly-N indicates the position of the spacer sequence that determines the nucleic acid target sequence of the sgRNA. The RGN systems APG05733.1, APG06207.1, APG01647.1, APG08032.1, APG02675.1, APG01405.1, APG06250.1, APG04293.1, and APG01308.1 had a conserved sequence at the base of the hairpin stem of the UNA tracrRNA (SEQ ID NO: 8). In APG05712.1, the sequence at the same position is CNANNG (SEQ ID NO: 37). In APG01658.1, the sequence at the same position is CNANU (SEQ ID NO: 45). For the RGN systems APG06498.1 and APG06877.1, the conserved sequence at the base of the tracrRNA hairpin stem is UNANG (SEQ ID NO: 53). In APG09882.1 and APG06646.1, the sequence at the same position is UNANC (SEQ ID NO: 68). In APG09053.1, the sequence at the same position is CNANU (SEQ ID NO: 102).

[0215]

Table 1

[0216] Example 2: Identification of guide RNAs and construction of sgRNAs Cultures of bacteria that naturally express the RNA-guided nuclease system under study were grown to mid-log phase (OD600 ~ 0.600), pelleted, and flash-frozen. RNA was isolated from the pellet using the mirVANA miRNA Isolation Kit (Life Technologies, Carlsbad, CA), and the NEBNext Small RNA Library Prep Kit (NEB, Beverly, MA) was fractionated on a 6% polyacrylamide gel with the library preparation solution to capture RNA species less than 200 nt, and crRNA and tracrRNA were detected respectively. Deep sequencing (75 bp paired-end) was performed on the NextSeq 500 (High Output kit) by a service provider (MoGene, St. Louis, MO). Reads were quality-trimmed using Cutadapt and mapped to the reference genome using Bowtie2. A custom RNAseq pipeline was created in Python to detect crRNA and tracrRNA transcripts. The processed crRNA boundaries were determined by the sequence coverage of the native repeat spacer array. The anti-repeat portion of tracrRNA was identified using permissive BLASTn parameters. The boundaries of processed tracrRNA were confirmed by identifying transcripts containing anti-repeats from the depth of the RNA sequence. Manual curation of the RNA was performed using secondary structure prediction by the RNA folding software NUPACK. The sgRNA cassette was prepared by DNA synthesis and was generally designed as follows (5’->3’): operably linked to the processed repeat portion of crRNA at the 3’ end, operably linked to a 4-bp non-complementary linker (AAAG; SEQ ID NO: 123), and operably linked to the processed tracrRNA at its 3’ end. Other 4-bp non-complementary linkers can also be used for the 20-30 bp spacer.

[0217] For in vitro assays, sgRNAs were synthesized by in vitro transcription of the sgRNA cassette using the GeneArt™ Precision gRNA Synthesis Kit (ThermoFisher). The processed crRNA and tracrRNA sequences for each RGN polypeptide were identified and are shown in Table 1. See below for sgRNAs constructed for PAM libraries 1 and 2.

[0218] Example 3: Determination of the PAM requirements for each RGN The PAM requirements for each RGN were determined essentially using the PAM depletion assay adapted from Kleinstiver et al. (2015) Nature 523:481-485 and Zetsche et al. (2015) Cell 163:759-771. Briefly, two plasmid libraries (L1 and L2) were generated in the pUC18 backbone (ampR), each containing a different 30 bp protospacer (target) sequence flanked by eight random nucleotides (i.e., the PAM region). The target sequences and flanking PAM regions for library 1 and library 2 for each RGN are shown in Table 2.

[0219] The libraries were separately electroporated into E. coli BL21(DE3) cells harboring the pRSF-1b expression vector containing the RGN of the invention (codons optimized for E. coli) together with cognate sgRNAs containing spacer sequences corresponding to the protospacers of L1 or L2. The transformation reactions used >10 6A library plasmid sufficient to obtain CFU was used. Both the RGN and sgRNA in the pRSF-1b backbone were under the control of the T7 promoter. After the transformation reaction was incubated for 1 hour, it was diluted into LB medium containing carbenicillin and kanamycin and grown overnight. The next day, the mixture was diluted into self-inducing Overnight Express™ Instant TB Medium (Millipore Sigma) to allow expression of the RGN and sgRNA, and after further growth for 4 hours or 20 hours, the cells were centrifuged and the plasmid DNA was isolated using a Mini-prep kit (Qiagen, German town, ¥32). In the presence of the appropriate sgRNA, plasmids containing the PAM recognized by the RGN are cleaved, and as a result, they are removed from the population. Plasmids containing PAM that are not recognized by the RGN or that are transformed into bacteria that do not contain the appropriate sgRNA survive and replicate. The PAM and protospacer regions of the uncut plasmids were PCR amplified and prepared for sequencing according to a published protocol (16s-metagenomic library prep guide 15044223B, Illumina, San Diego, CA). Deep sequencing (75bp single-end reads) was performed on a MiSeq (Illumina) by a service provider (MoGene, St. Louis, MO). Typically, 1-4M reads were obtained per amplicon. The PAM regions were extracted, counted, and normalized to the total reads of each sample. PAMs leading to plasmid cleavage were identified by being underrepresented compared to a control (i.e., when the library was transformed into E. coli that contains the RGN but lacks the appropriate sgRNA). To represent the PAM requirements for novel RGNs, the depletion ratio (frequency in the sample / frequency in the control) for all sequences in the region of interest was converted to an enrichment value by log base 2 transformation. Sufficient PAMs were defined as those having an enrichment value > 2.3 (which corresponds to a depletion ratio < ~0.2). PAMs exceeding this threshold in both libraries were collected and used to generate a web logo, which can be generated using, for example, a web-based service on the Internet known as "WebLogo".PAM arrays were identified and reported when there was a consistent pattern in the top-enriched PAMs. The consensus PAMs (enrichment factor (EF) > 2.3) for each RGN are shown in Table 2. The PAM orientation is also shown in Table 2. As described elsewhere in this application, the APG06646.1 and APG04293.1 nucleases do not have a PAM interaction domain. The results in Table 2 show that they do not have the typical PAM requirements of 2-5 nucleotides. APG06646.1 and APG04293.1 were shown to have a single base required for cleavage.

[0220]

Table 2

[0221] Example 4: Manipulation of guide RNAs to enhance nuclease activity 4.1 RGN APG09748 and APG09106.1 For the RGNs, APG09748 (SEQ ID NO: 137, and the crRNA repeat sequence, tracrRNA sequence, and general sgRNA sequence of APG09748 are described as SEQ ID NOs: 273, 274, and 275, respectively; all sequences are described in International Application No. PCT / US2019 / 068079, which is incorporated herein by reference), and APG09106.1 has very high sequence identity and the same PAM, but regions in the guide RNA that can be modified to optimize nuclease activity were determined using RNA folding predictions. Repeat: The stability of the crRNA:tracrRNA base pairs in the anti-repeat region was increased by shortening the repeat:anti region, adding G-C base pairs, and removing G-U wobble pairs. The "optimized" guide variants were tested and compared to the wild-type gRNA in an in vitro cleavage assay using RGN APG09748.

[0222] To generate an RGN for RNP formation, an expression plasmid containing an RGN fused to a C-terminal His6 (SEQ ID NO: 276) or His10 (SEQ ID NO: 277) tag was constructed and transformed into the E. coli BL21(DE3) strain. Expression was performed using Magic Media (Thermo Fisher) supplemented with 50 μg / mL kanamycin. After lysis and clarification, the protein was purified by immobilized metal affinity chromatography and quantified by UV-vis using the Qubit Protein Assay Kit (Thermo Fisher) or using the calculated extinction coefficient.

[0223] The purified RGN was incubated with sgRNA at a ratio of approximately 2:1 for 20 minutes at room temperature to prepare ribonucleoprotein (RNP). For the in vitro cleavage reaction, the RNP was incubated with a plasmid or linear dsDNA containing a target protospacer adjacent to the preferred PAM sequence for >30 minutes at room temperature. Two target nucleic acid sequences within the TRAC locus, TRAC11 (SEQ ID NO: 278) and TRAC14 (SEQ ID NO: 279), were tested. The gRNAs were assayed for target activity against both the correct target nucleic acid sequence (e.g., the gRNA has the TRAC11 spacer sequence and the target being assayed is TRAC11) and the correct target nucleic acid sequence (e.g., the gRNA has the TRAC11 spacer sequence and the target being assayed is TRAC14). The activity determined by plasmid cleavage was evaluated by agarose gel electrophoresis. The results are shown in Table 3. The guide variants are listed as SEQ ID NOs: 280 - 283 and have spacer sequences. These guide sequences use a non-complementary nucleotide linker of AAAA (SEQ ID NO: 284). Repeat: The optimized gRNA with increased anti-repeat binding (SEQ ID NO: 285; poly-N indicates the position of the spacer sequence) has optimized tracrRNA (SEQ ID NO: 286) and crRNA (SEQ ID NO: 287) components. The optimized guide variant was able to cleave two loci where cleavage was not previously detected using the wild-type guide RNA. Repeat: Optimization of hybridization in the anti-repeat region increased in vitro cleavage of APG09748 from 0% cleavage to 100% cleavage for multiple targets at the TRAC locus.

[0224]

Table 3

[0225] Further optimized gRNA variants were designed and assayed. Additionally, different lengths of the spacer sequence were also tested to determine how spacer length affects cleavage efficiency. The sgRNA outside the spacer sequence is referred to as the "backbone" in this assay. In Table 4, these are shown as "WT" (SEQ ID NO: 288, wild-type sequence), and three optimized sgRNAs: V1 (SEQ ID NO: 289), V2 (SEQ ID NO: 290), and V3 (SEQ ID NO: 291). All of these sequences have poly-N to indicate the position of the spacer sequence. The guides were expressed as sgRNAs by in vitro transcription (IVT). Compared to the wild-type sgRNA backbone, V1 is 87.8% identical, V2 is 92.4% identical, and V3 is 85.5% identical. Synthetic tracheal rRNA:crRNA duplexes ("synthetic") that represent double-guide RNAs and are otherwise similar to the wild-type and the optimized sgRNAs described above were also generated and tested.

[0226] For this set of assays, RG N APG09106.1 was used; otherwise, the method for the in vitro cleavage reaction was the same as described above. The target nucleic acid sequences were Target 1 (SEQ ID NO: 292) and Target 2 (SEQ ID NO: 293). The results are shown in Table 4.

[0227]

Table 4

[0228] 4.2 RG N APG07433.1 Using RNA folding prediction, regions in the guide RNA for APG07433.1 (described as SEQ ID NO: 235 and described in US Patent Application 2019 / 0367949 and WO2019 / 236566, each incorporated herein by reference in its entirety) were determined. It can be modified to optimize nuclease activity and shorten the guide RNA for packaging into viral vectors. As sites where changes might occur, the pairing of the repeat:anti-repeat between the crRNA and tracrRNA regions and the terminal hairpin of tracrRNA was identified. The repeat:anti-repeat regions were cleaved to lengths of 7 (APG07433.1-7bp, SEQ ID NOs: 238 and 239), 13 (APG07433.1-13bp, SEQ ID NOs: 250 and 251), and 15 base pairs (APG07433.1-15bp, SEQ ID NOs: 242 and 243). Additionally, a fourth variant was tested where the sequence of the repeat:anti-repeat region was changed to reduce the pairing to a length of 11 base pairs (APG07433.1-11bp-syn, SEQ ID NOs: 240 and 241), introducing a smaller RNA bulge compared to the wild-type guide. Cleavage was also performed on the terminal hairpin in tracrRNA, including reduction in length to 40 nucleotides (APG07433.1-40nt THP, SEQ ID NOs: 254 and 255) and 42 nucleotides (APG07433.1-42nt THP, SEQ ID NOs: 244 and 245) of the native sequence. Modified stem-loops were designed to also shorten these to 35 nucleotides (APG07433.1-35ntTHP-syn, SEQ ID NOs: 248 and 249) and 39 nucleotides (APG07433.1-39ntTHP-syn, SEQ ID NOs: 246 and 247). One guide combined a 13-base pair shortened repeat:anti-repeat region with a 42-nucleotide shortened terminal hairpin (APG07433.1-13bp42ntTH, SEQ ID NOs: 252 and 253). The effect of spacer length on cleavage was also tested.Nucleotides with spacer lengths of 25 (APG07433.1 - natural

[25] , SEQ ID NOs: 236 and 237; spacer sequence is shown as SEQ ID NO: 271) and 18 (APG07433.1 - natural

[18] , SEQ ID NOs: 256 and 257; spacer sequence is shown as SEQ ID NO: 272) on the wild - type guide RNA backbone were also tested.

[0229] A solution containing 20 μM crRNA and 10 μM tracrRNA in annealing buffer (Synthego) was prepared, then heated at 78 °C for 10 minutes and then at 37 °C for 30 minutes to anneal the crRNA and tracrRNA, thereby preparing the following double - guide RNA. The RNP was formed by incubating with purified APG07433.1 protein at 0.5 μM and incubating the double - guide RNA at 1 μM in phosphate - buffered saline for 20 minutes. The guide RNAs used in this experiment are shown in Table 5.

[0230]

Table 5

[0231] These RNPs were incubated with a PCR product containing an appropriate target sequence (SEQ ID NO: 260) generated by amplification using FAM and Cy3 - labeled primers to facilitate accurate and sensitive quantification of the cleavage products. The reaction contained 250 nM RNP manufactured as above and 150 nM FAM and Cy3 - labeled PCR products in 1× Cutsmart buffer (New England Biolabs). The reaction was allowed to proceed at 37 °C for 15 minutes and stopped by adding RNase A to 0.1 mg / mL and EDTA to 45 mM. The quenched reaction was heated at 50 °C for 30 minutes and 95 °C for 5 minutes. Next, the samples were analyzed using native acrylamide gel electrophoresis on a 5% TBE gel (Bio - rad). These were imaged on a ChemiDoc MP imager. Each of the two dyes could be used for quantification, and each sample was performed twice. The results of this experiment are shown in Table 6.

[0232]

Table 6

[0233] This analysis demonstrates that the natural backbone outperforms most truncated variants and variants containing truncated target sequences (APG07433.1 - native

[18] , SEQ ID NOs: 256 and 257). Among the truncated backbone sequences, the one with the highest level of cleavage is the sequence called APG07433.1 - 11bp - syn, which contains a 25nt targeting sequence (spacer; SEQ ID NO: 271), SEQ ID NO: 240 for crRNA, and SEQ ID NO: 241 as tracrRNA. This guide variation contained a modified repeat with a bulge engineered into the stem: an anti - repeat stem - loop.

[0234] Example 5: Demonstration of gene editing activity in mammalian cells An RGN expression cassette was prepared and introduced into a vector for mammalian expression. RGN APG09748, APG09106.1, APG05712.1, APG01658.1, APG05733.1, APG06498.1, APG06646.1, APG09882.1, APG01405.1, and APG01308.1 were each codon-optimized for mammalian expression (SEQ ID NOs: 304, 305, and 411-418), and the expressed proteins were operably fused to an SV40 nuclear localization sequence (NLS; SEQ ID NO: 125) and a 3×FLAG tag (SEQ ID NO: 126) at the N-terminus, and a nucleoplasmin NLS sequence (SEQ ID NO: 127) at the C-terminus. Two copies of the NLS sequence were used and operably fused in tandem. Each expression cassette was placed under the control of a cytomegalovirus (CMV) promoter (SEQ ID NO: 306). It is known in the art that a CMV transcriptional enhancer (SEQ ID NO: 307) can also be included in a construct containing the CMV promoter. Guide RNA expression constructs encoding a single gRNA each were prepared under the control of a human RNA polymerase III U6 promoter (SEQ ID NO: 308) and introduced into an expression vector. As shown in Table 7, target regions of selected genes including RelA, AurkB, GAPDH, LINC01509, HBB, CFTR, HPRT1, TRA, EMX1, and VEGFA were targeted. For the RNA-guided nuclease APG09106.1, specific residues were mutated to increase the nuclease activity of the protein, specifically, the T849 residue of APG09106 was mutated to arginine (SEQ ID NO: 309). This point mutation increased the editing rate in mammalian cells.

[0235] The above constructs were introduced into mammalian cells. One day before transfection, 1x10 5HEK293T cells (Sigma) were plated in 24-well plates in Dulbecco's Modified Eagle Medium (DMEM) + 10% (vol / vol) fetal bovine serum (Gibco) and 1% penicillin-streptomycin (Gibco). The day after the cells were 50 - 60% confluent, 500 ng of RGN expression plasmid + 500 ng of single gRNA expression plasmid were co-transfected using 1.5 μL of Lipofectamine 3000 (Thermo Scientific) per well according to the manufacturer's instructions. After 48 hours of growth, total genomic DNA was recovered using a genomic DNA isolation kit (Machery-Nagel) according to the manufacturer's instructions.

[0236] Next, total genomic DNA was analyzed to determine the rate of editing in the target gene. Oligonucleotides were generated (SEQ ID NOs: 310 and 311) for use in PCR amplification and subsequent analysis of the amplified genomic target sites. All PCR reactions were carried out using 10 μL of 2X Master Mix Phusion High-Fidelity DNA polymerase (Thermo Scientific) in a 20 μL reaction containing 0.5 μM of each primer. A large genomic region containing each target gene was first amplified using PCR#1 primers (SEQ ID NOs: 310 and 311) with the following program: 98°C, 1 minute; [98°C, 10 seconds; 62°C, 15 seconds; 72°C, 5 minutes]; 72°C, 5 minutes; 12°C, hold.

[0237] Next, 1 microliter of this PCR reaction was further amplified using primers specific for each guide (PCR#2 primers; SEQ ID NOs: 365 - 370) with the program: 98°C, 1 minute; [98°C, 10 seconds; 67°C, 15 seconds; 72°C, 30 seconds] for 35 cycles; 72°C, 5 minutes; 12°C. The primers for PCR#2 contain the Nextera Read 1 and Read 2 transposase adapter overhang sequences for Illumina sequencing.

[0238] After the second PCR amplification, the DNA was washed using a PCR clean-up kit (Zymo) according to the manufacturer's instructions and eluted in water. 200 - 500 ng of the purified PCR#2 product was combined with 2 μL of 10X NEB buffer 2 and water in a 20 μL reaction and annealed using the following program to form heteroduplex DNA: 95°C for 5 minutes; cooled from 95 - 85°C at a rate of 2°C / second; cooled from 85 - 25°C at a rate of 0.1°C / second; 12°C indefinitely. After annealing, 5 μL of the DNA was removed as an enzyme-free control, 1 μL of T7 endonuclease I (NEB) was added, and the reaction was incubated at 37°C for 1 hour. After incubation, 5×FlashGel loading dye (Lonza) was added, and 5 μL of each reaction and control was analyzed on a 2.2% agarose FlashGel (Lonza) using gel electrophoresis. After visualization of the gel, the percentage of non-homologous end joining (NHEJ) was determined using the following formula: NHEJ event % = 100x[1 - (1 - fraction cleaved)(1 / 2)], where (fraction cleaved) is defined as (density of the digested product) / (density of the digested product + density of the undigested parental band).

[0239] For some samples, SURVEYOR® was used to analyze the post-expression results in mammalian cells. Prior to genomic DNA extraction, the cells were incubated at 37°C for 72 hours after transfection. Genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) according to the manufacturer's protocol. The genomic region adjacent to the RGN target site was PCR amplified and the product was purified using a QiaQuick Spin Column (Qiagen) according to the manufacturer's protocol. A total of 200 - 500 ng of the purified PCR product was mixed with 1 μl of 10× Taq DNA polymerase PCR buffer (enzyme) and ultrapure water to a final volume of 10 μl and subjected to a re-annealing step to enable heteroduplex formation: 95°C for 10 minutes, 95°C - 85°C at -2°C / s, 85°C - 25°C at -0.25°C / s, and hold at 25°C for 1 minute.

[0240] After re-annealing, the products were treated with SURVEYOR® nuclease and SURVEYOR® enhancer S (Integrated DNA Technologies) according to the manufacturer's recommended protocol and analyzed on a 4–20% Novex TBE polyacrylamide gel (Life Technologies). The gel was stained with SYBR Gold DNA stain (Life Technologies) for 10 minutes and imaged on a Gel Doc gel imaging system (Bio-rad). Quantification was performed based on relative band intensity. The indel percentage was determined by the formula, 100×(1-(b + c) / (a + b + c)))1 / 2. Here, a is the integrated intensity of the undigested PCR product, and b and c are the integrated intensities of each cleavage product.

[0241] Furthermore, products from PCR #2 containing the Illumina overhang sequence were library-prepared according to the Illumina 16S Metagenomic Sequencing Library protocol. Deep sequencing was performed on the Illumina Mi-Seq platform by a service provider (MOGene). Typically, 200,000 of 250 bp paired-end reads (2×100,000 reads) are generated per amplicon. The read values were analyzed using CRISPResso (Pinello, et al. 2016 Nature Biotech, 34:695-697) to calculate the editing rate. The output alignment was manually curated to identify microhomology sites at the recombination site and to confirm the insertion and deletion sites. The overall editing rate, deletion rate, and insertion rate for each sample are shown in Table 7. All experiments were performed in human cells. "Target" is the target sequence within the gene target. For each target sequence, the guide RNA contained a complementary RNA spacer sequence and the appropriate sgRNA depending on the RGN used. Selected breakdowns of the experiments with guide RNAs are shown in Tables 8.1 and 8.2.

[0242]

Table 7-1

[0243]

Table 7-2

[0244] Tables 8.1 and 8.2 show the specific insertions and deletions of each guide. In these tables, the target sequence is identified by bold uppercase letters. The 8mer PAM region is double-underlined, and the main recognition nucleotides are shown in bold. Insertions are identified by lowercase letters. Deletions are indicated by dashes (---). The INDEL position is calculated from the PAM-proximal end of the target sequence, with the end being position 0. Positions are positive (+) if on the target side of the edge and negative (-) if on the PAM side of the edge.

[0245]

Table 8

[0246]

Table 9

[0247] Example 6: Demonstration of gene editing activity in plant cells The RNA-guided nuclease activity of the RGNs of the present invention is demonstrated in plant cells using a protocol modified from Li, et al., 2013 (Nat. Biotech. 31:688-691). Briefly, the plant codon-optimized versions of the RGNs of the present invention (SEQ ID NO: 1, 9, 16, 23, 30, 38, 46, 54, 61, 69, 75, 82, 89, 95, 103, 110, or 117) are operably linked to a nucleic acid sequence encoding an N-terminal SV40 nuclear localization signal and cloned behind the strong constitutive 35S promoter of a transient transformation vector. An sgRNA targeting one or more sites of the plant PDS gene adjacent to an appropriate PAM sequence is cloned behind the plant U6 promoter in a second transient expression vector. The expression vectors are introduced into Nicotiana benthamiana mesophyll protoplasts using PEG-mediated transformation. The transformed protoplasts are incubated in the dark for up to 36 hours. Genomic DNA is isolated from the protoplasts using the DNeasy Plant Mini Kit (Qiagen). The genomic region adjacent to the RGN target site is PCR amplified and the product is purified using a QiaQuick Spin Column (Qiagen) according to the manufacturer's protocol. A total of 200 - 500 ng of the purified PCR product is mixed with 1 μl of 10× Taq DNA polymerase PCR buffer (enzyme) and ultrapure water to a final volume of 10 μl and subjected to a re-annealing step to allow heteroduplex formation: 95°C for 10 minutes, a ramp from 95°C to 85°C at -2°C / s, a ramp from 85°C to 25°C at -0.25°C / s, and hold at 25°C for 1 minute.

[0248] After re-annealing, the products are treated with SURVEYOR nuclease and SURVEYOR enhancer S (Integrated DNA Technologies) according to the manufacturer's recommended protocol, stained with SYBR Gold DNA stain (Life Technologies) for 10 minutes on a 4-20% Novex TBE polyacrylamide gel (Life Technologies), and imaged with a Gel Doc gel imaging system (Bio-rad). Quantification is based on relative band intensity. The indel percentage is determined by the formula: 100×(1-(1-(b + c) / (a + b + c))1 / 2), where a is the integrated intensity of the undigested PCR product, and b and c are the integrated intensities of each cleavage product.

[0249] Example 7: Identification of disease targets The database of clinical variants was obtained from the NCBI ClinVar database. The database is available from websites around the world on the NCBI ClinVar website. Pathogenic single nucleotide polymorphisms (SNPs) were identified from this list. Using the genomic locus information, CRISPR targets in the overlapping regions and surrounding regions of each SNP were identified. Table 9.1 shows the selection of SNPs that can be corrected using base editing in combination with an RGN system containing APG09748 or APG09106.1 to target the causative mutation ("Casl Mut."). Table 9.2 shows the selection of SNPs that can be corrected using base editing in combination with an RGN system containing APG06646.1 or APG01658.1 to target the causative mutation ("Casl Mut."). In both Table 9.1 and Table 9.2, only one alias for each disease is listed. "RS#" corresponds to the RS accession number via the SNP database on the NCBI website. The chromosome accession number provides the accession reference information found on the NCBI website. Tables 9.1 and 9.2 also provide genomic target sequence information suitable for the RGN system APG09748 or APG09106.1 (Table 9.1), or APG06646.1 or APG01658.1 (Table 9.2) for each disease. The target sequence information also provides the protospacer sequence for the production of the sgRNA required for the corresponding RGN of the present invention.

[0250]

Table 10-1

[0251]

Table 10-2

[0252]

Table 10-3

[0253]

Table 11-1

[0254]

Table 11-2

[0255]

Table 11-3

[0256]

Table 11-4

[0257]

Table 11-5

[0258]

Table 11-6

[0259]

Table 11-7

[0260] Example 8: Target mutations causing Friedreich's ataxia The expansion of the trinucleotide repeat sequence that causes Friedreich ataxia (FRDA) occurs at specific loci within the FXN gene, called the FRDA unstable region. RNA-guided nucleases (RGNs) can be used to excise the unstable region in FRDA patient cells. This approach requires 1) an RGN and guide RNA sequence that can be programmed to target alleles in the human genome; and 2) a delivery approach for the RGN and guide sequences. Many nucleases used for genome editing, such as the commonly used Cas9 nuclease (SpCas9) from S. pyogenes, are too large to be packaged into adeno-associated virus (AAV) vectors. Thus, a feasible approach using SpCas9 is unlikely.

[0261] The compact RNA-guided nucleases of the present invention, such as APG09748, APG09106.1, and APG06646.1, are specifically suitable for excision of the FRDA unstable region. Each RGN has a PAM requirement near the FRDA unstable region. Further, each of these RGNs can be packaged into an AAV vector together with a guide RNA. A second vector would be required to pack two guide RNAs, but this approach is advantageous compared to what is required for larger nucleases such as SpCas9. SpCas9 needs to split the protein sequence between two vectors.

[0262] Table 10 shows the positions of genomic target sequences suitable for targeting APG09748, APG09106.1, or APG06646.1 to the 5' and 3' flanks of the FRDA unstable region, as well as the sequences of the sgRNAs for the genomic targets. Once at the locus, the RGN excises the FA unstable region. Excision of the region can be confirmed by Illumina sequencing of the locus.

[0263]

Table 12

[0264] Example 9: Targeting the mutation causing sickle cell disease Target sequences within the BCL11A enhancer region (SEQ ID NO: 220) may provide a mechanism for increasing fetal hemoglobin (HbF) to cure or alleviate the symptoms of sickle cell disease. For example, genome-wide association studies have identified a series of genetic mutations in BCL11A associated with elevated HbF levels. These mutations are a set of SNPs found in the non-coding region of BCL11A that function as stage-specific lineage-restricted enhancer regions. Further studies have revealed that this BCL11A enhancer is required in erythroid cells for BCL11A expression (Bauer et al, (2013) Science 343:253-257, incorporated herein by reference). The enhancer region is found within intron 2 of the BCL11A gene, and three regions of DNaseI hypersensitivity in intron 2 (which often indicate chromatin states associated with regulatory potential) have been identified. These three regions were identified as "+62", "+58", and "+55" according to their kilobase distance from the transcription start site of BCL11A. These enhancer regions are approximately 350 (+55); 550 (+58); and 350 (+62) nucleotides in length (Bauer et al., 2013).

[0265] Example 9.1: Identification of Preferred RGN Systems Here, we describe a potential treatment for β-hemoglobinopathies using an RGN system that disrupts the binding of BCL11A to binding sites within the HBB locus, a gene involved in β-globin production of adult hemoglobin. This approach uses NHEJ, which is more efficient in mammalian cells. Furthermore, this approach uses nucleases that are small enough to be packaged into a single AAV vector for in vivo delivery.

[0266] The GATA1 enhancer motif in the human BCL11A enhancer region (SEQ ID NO: 220) is an ideal target for disruption using an RNA-guided nuclease (RGN) to reduce BCL11A expression with concomitant re-expression of HbF in adult human erythrocytes (Wu et al. (2019) Nat Med 387:2554). Several PAM sequences compatible with APG09748 or APG09106.1 are readily apparent at loci surrounding this GATA1 site. These nucleases have a PAM sequence of 5’-DTTN-3’ (SEQ ID NO: 60), are compact in size, and enable their delivery with an appropriate guide RNA in a single AAV or adenoviral vector. APG06646.1, in addition to its size, has minimal PAM requirements (SEQ ID NO: 109) and is suitable for this approach. This delivery approach offers several advantages compared to others, such as access to hematopoietic stem cells and a well-established safety profile and manufacturing technology.

[0267] The commonly used Cas9 nuclease from S. pyogenes (SpyCas9) requires a PAM sequence of 5’-NGG-3’ (SEQ ID NO: 101), some of which are present in the vicinity of the GATA1 motif. However, the size of SpyCas9 precludes packaging into a single AAV or adenoviral vector and thus negates the aforementioned advantages of this approach. A dual delivery strategy could be employed, but it adds significant manufacturing complexity and cost. Furthermore, dual viral vector delivery requires infection by both vectors for editing to succeed in a given cell, significantly reducing the efficiency of gene correction.

[0268] An expression cassette encoding human codon-optimized APG09748 (SEQ ID NO: 409), APG09106.1 (SEQ ID NO: 410), or APG06646.1 (SEQ ID NO: 415) is produced in the same manner as described in Example 5. An expression cassette for expressing guide RNAs for RNAGNAPG09748, APG09106.1, or APG06646.1 is also produced. These guide RNAs contain 1) a protospacer sequence complementary to either the non-coding or coding DNA strand within the BCL11A enhancer locus (target sequence), and 2) an RNA sequence necessary for the association of the guide RNA with the RGN. Since several potential PAM sequences for targeting by each RGN surround the BCL11A GATA1 enhancer motif, several potential guide RNA constructs are generated to determine the best protospacer sequence that results in strong cleavage and disruption of the BCL11A GATA1 enhancer sequence via NHEJ. The target genomic sequences in Table 11 are evaluated using the sgRNAs shown in Table 11.

[0269]

Table 13

[0270] To evaluate the efficiency of APG09748, APG09106.1, or APG06646.1 in generating insertions or deletions that disrupt the BCL11A enhancer region, human cell lines such as human embryonic kidney cells (HEK cells) are used. A DNA vector containing an RGN expression cassette (e.g., as described in Example 5) is prepared. A separate vector containing an expression cassette encoding the guide RNA sequences of Table 11 is also generated. Such expression cassettes can further contain the human RNA polymerase III U6 promoter (SEQ ID NO: 308), as described in Example 5. Alternatively, a single vector containing expression cassettes for both RGN and guide RNA can be used. The vector is introduced into HEK cells using standard techniques such as those described in Example 5, and the cells are cultured for 1 - 3 days. After this culture period, genomic DNA is isolated, and the frequency of insertions or deletions is determined using T7 endonuclease I digestion and / or direct DNA sequencing, as described in Example 5.

[0271] A DNA region containing the target BCL11A region is amplified by PCR using primers containing Illumina Nextera XT overhang sequences. These PCR amplicons are either examined for NHEJ formation using T7 endonuclease I digestion or undergo library preparation after Illumina 16S Metagenomic Sequencing Library protocol or a similar next-generation sequencing (NGS) library preparation. Following deep sequencing, the generated reads are analyzed by CRISPRESSO to calculate the editing rate. The output alignment is manually curated to identify the insertion and deletion sites. This analysis identifies the preferred RGN and the corresponding preferred guide RNA(s) (sgRNA). As a result of this analysis, it is considered whether any of APG09748, APG09106.1, or APG06646.1 are equally preferred or one RGN is most preferred. Further, the analysis can determine whether multiple preferred guide RNAs exist or whether all target genomic sequences in Table 17 are equally preferred.

[0272] Example 9.2: Fetal Hemoglobin Expression Test In this example, APG09748, APG09106.1, or APG06646.1, which have insertions or deletions that disrupt the BCL11A enhancer region, are assayed for fetal hemoglobin expression. Healthy human donor CD34 + hematopoietic stem cells (HSCs) are used. These HSCs are cultured and a vector containing an expression cassette comprising the coding region of a preferred RGN and a preferred sgRNA are introduced using a method similar to that described in Example 5. Alternatively, electroporation may be used. After electroporation, these cells are differentiated into erythrocytes in vitro using an established protocol (e.g., Giarratana et al. (2004) Nat Biotechnology 23:69-74, which is incorporated herein by reference). Next, the expression of HbF is measured using Western blotting with an anti-human HbF antibody or quantified via high performance liquid chromatography (HPLC). If disruption of the BCL11A enhancer locus is successful, HbF production is increased compared to HSCs electroporated with RGN only, as expected without a guide.

[0273] Example 9.3: Sickle Erythropoiesis Reduction Test In this experiment, APG09748, APG09106.1, or APG06646.1, which have insertions or deletions that disrupt the BCL11A enhancer region, are assayed for reduction of sickle erythropoiesis. Donor CD34 from patients with sickle cell disease +Hematopoietic stem cells (HSCs) are used. These HSCs are cultured, and a vector containing an expression cassette comprising the coding region of a preferred RGN is introduced, and a preferred sgRNA is introduced using a method similar to that described in Example 5. Alternatively, electroporation may be used. After electroporation, these cells are differentiated into erythrocytes in vitro using an established protocol (Giarratana et al. (2004) Nat Biotechnology 23:69-74). Next, the expression of HbF is measured using Western blotting with an anti-human HbF antibody or quantified via high performance liquid chromatography (HPLC). If disruption of the BCL11A enhancer locus is successful, HbF production is expected to increase compared to HSCs electroporated with RGN alone, but no guide is expected.

[0274] Sickle erythropoiesis is induced by adding metabisulfite to these differentiated erythrocytes. The number of sickle and normal erythrocytes is counted microscopically. The number of sickle erythrocytes is expected to be less in cells treated with APG09748, APG09106.1, or APG06646.1+sgRNA than in untreated or cells treated with RGN alone.

[0275] Example 9.4: Validation of disease treatment in a mouse model To evaluate the efficacy of disruption of the BCL11A locus by APG09748, APG09106.1, or APG06646.1, a suitable humanized mouse model of sickle cell anemia is used. Expression cassettes encoding a preferred RGN and a preferred sgRNA are packaged into an AAV vector or an adenovirus vector. In particular, the adenovirus type Ad5 / 35 is effective in targeting HSCs. A suitable mouse model containing a humanized HBB locus with the sickle cell allele is selected such as B6;FVB-Tg(LCR-HBA2,LCR-HBB*E26K)53Hhb / J or B6.Cg-Hbatm1Paz btm1Tow Tg(HBA-HBBs)41Paz / HhbJ. These mice are treated with granulocyte colony-stimulating factor alone or in combination with plerixafor for mobilization of HSCs into the blood. Next, AAV or adenovirus carrying the RGN and the guide plasmid is injected intravenously and the mice are allowed to recover for 1 week. Blood obtained from these mice is tested in an in vitro sickling assay using sodium metabisulfite and the mice are followed longitudinally to monitor mortality and hematopoietic function. Treatment with AAV or adenovirus having the RGN and guide RNA is expected to improve sickling, death, and hematopoietic function compared to mice treated with a virus lacking both expression cassettes or a virus having the RGN expression cassette alone.

[0276] Example 10: Base Editing Activity in Mammalian Cells An expression cassette producing a cytidine deaminase-RGN fusion protein was constructed as follows. The coding sequence of the RGN (APG06646.1 described herein, and APG08290.1, as described in U.S. Patent Application Publication No. 2019 / 0367949 and International Publication No. WO2019 / 236566, each of which is incorporated herein by reference in its entirety, was codon-optimized for mammalian expression and mutated to function as a nickase (SEQ ID NOs: 128 and 262, respectively). This coding sequence contains an NLS at its N-terminus (SEQ ID NO: 125), is operably linked to a 3xFLAG tag (SEQ ID NO: 126) at its C-terminus, and is introduced into an expression cassette that generates a fusion protein (SEQ ID NOs: 129-132) operably linked to a cytidine deaminase at its C-terminus. PCT / US2019 / 068079 is operably linked at its C-terminus to an amino acid linker (SEQ ID NO: 133) that is operably linked to an RGN nickase at its C-terminus and operably linked to a second NLS (SEQ ID NO: 127) at its C-terminus, and is incorporated herein by reference in its entirety. These expression cassettes were each introduced into a pTwist CMV vector (Twist Bioscience) capable of driving the expression of the fusion protein in mammalian cells. Separate vectors were also generated that express guide RNAs under the control of the human U6 promoter in mammalian cells. These guide RNAs (SEQ ID NOs: 134-136 for nAPG06646.1 and SEQ ID NOs: 263-265 for nAPG08290.1) can each direct the deaminase-nRGN fusion protein, or the RGN itself, to a target genomic sequence for base editing or gene editing.

[0277] 500 ng of cytidine deaminase-RGN expression plasmid or standard RGN expression plasmid, and 500 ng of guide RNA expression plasmid were co-transfected into HEK293FT cells at 75 - 90% confluence in a 24-well plate using Lipofectamine 2000 reagent (Life Technologies). Next, the cells were incubated at 37 °C for 72 hours. Next, genomic DNA was extracted using NucleoSpin 96 Tissue (Macherey-Nagel) according to the manufacturer's protocol. The genomic region adjacent to the guide RNA target site was PCR amplified, and the product was purified using ZR-96 DNA Clean & Concentrator (Zymo Research) according to the manufacturer's protocol. Next, the purified PCR product was sent for next-generation sequencing on Illumina Miq (2×250). The results were analyzed for indel formation or specific cytosine mutations.

[0278] Table 12 below shows the editing of cytosine bases and the actual formation of each construct and guide combination. Interestingly, when comparing the activity of RGN APG06646.1 with that of cytidine deaminase-nAPG06646.1 using the same guide RNA, cytosine base editing up to 20-fold higher than gene editing was observed. These results demonstrate that an RGN with low nuclease activity at a specific target site can be efficient base editing at that site. Furthermore, since indel formation in base editing applications is often an undesirable result, an RGN with low nuclease activity at a site is preferred for base editing applications.

[0279]

Table 14

[0280] Example 11: Trans ssDNA Cleavage 11.1 Determination of Assay Conditions for Trans DNA Cleavage Purified APG09748 (described as SEQ ID NO: 137 and described in International Application No. PCT / US2019 / 068079, which is hereby incorporated by reference in its entirety) was incubated for 10 minutes at a final concentration of either 50 nM nuclease and 100 nM sgRNA or 200 nM nuclease, together with single guide RNA (sgRNA) in Cutsmart buffer (New England Biolabs B7204S). These RNP solutions were added to a solution of ssDNA (target or mismatch negative control ssDNA) and reporter probe at a final concentration of 10 nM and a final concentration of 250 nM in 1.5× Cutsmart buffer (New England Biolabs B7204S). Reporter probes (TB0125 and TB0089, shown as SEQ ID NOs: 138 and 139, respectively) contain a fluorescent dye at the 5′ end (56-FAM for TB0125, Cy5 for TB0089), a quencher at the 3′ end (3IABkFQ for TB0125, 3IAbRQSp for TB0089), and optionally an internal quencher (internal quencher ZEN is present only on TB0125). Cleavage of the reporter probe results in dequenching of the fluorescent dye and thus an increase in the fluorescent signal. To monitor the fluorescence intensity, 10 μl of each reaction was incubated at 30° C. in a Corning low-volume 384-well microplate in a microplate reader (CLARIOSTAR Plus).

[0281] To determine the parameters suitable for this assay, a number of conditions were explored. To determine whether there was an effect of the quenched probe design or fluorophore properties, two such reporters were included as a mixture in each reaction. They were at the same concentration as each other in any given reaction. In all cases, the control or target ssDNA concentrations (LE201 or LE205, shown as SEQ ID NOs: 140 and 141, respectively) were 10 nM. The RNP names mean nuclease and target, as shown in Table 13 below.

[0282]

Table 15

[0283] The results are shown in Table 14 below.

[0284]

Table 16-1

[0285]

Table 16-2

[0286] From this experiment, it was concluded that RNP at a concentration of 100 nM generally results in a higher cleavage rate of the reporter probe than an RNP concentration of 25 nM. Generally, the reporter cleavage rate is higher with high concentrations of reporter oligonucleotides up to a reporter concentration of 250 - 500 nM, and little benefit is observed from further increases in reporter concentration. Notably, for the TB0089 reporter (detected in the Cy5 channel), there is a substantially high level of background activity that interferes with target discrimination, particularly at reporter concentrations above 250 nM. Therefore, it was concluded that reporter concentrations above 250 nM are not beneficial. Both probes appear to be suitable for the observation of target recognition by RNP.

[0287] 11.2 Influence of G09748 Trans DNA Cleavage and Purification on Nonspecific Activity Purified APG09748 was incubated with single guide RNA (sgRNA) in 1× CutSmart buffer (New England Biolabs B7204S) at 37°C for 10 minutes with a nuclease at a final concentration of 200 nM and 400 nM of sgRNA. These RNP solutions were then incubated at a final concentration of 250 nM in 1.5× Cutsmart buffer (New England Biolabs B7204S) with ssDNA (the target or mismatch negative control ss reporter probe contains a fluorophore at the 5' end and a quencher at the 3' end. Cleavage of the reporter probe results in dequenching of the fluorophore and thus an increase in the fluorescence signal. To monitor the fluorescence intensity, 10 μl of each reaction was incubated in a Corning low volume 384 well microplate at 37°C in a microplate reader (CLARIOSTAR Plus).

[0288] Incubation with the target sequence resulted in a substantial increase in fluorescence intensity as a function of time compared to the negative control. The cleavage rate is summarized as the slope of the linear portion of the fluorescence vs. time function as shown in Table 15.

[0289]

Table 17

[0290] These data demonstrate the discrimination of target sequences from negative control by RNP with a clearly detectable reporter probe upon cleavage.

[0291] 11.3 PAM measurement using a parallel plasmid DNA library Oligonucleotides containing a target sequence preceding a 5nt-shrunk PAM sequence (NNNN) (LE00680 and LE00688 shown as SEQ ID NOs: 266 and 267, respectively) were annealed and cloned into a double-digested pUC19 plasmid using NEBuilder HiFi DNA Assembly Master Mix (New England Biolabs). Since each colony resulting from the transformation of this reaction corresponds to a clonal plasmid DNA sequence, a preparation of plasmid DNA from a culture derived from a single colony is a unique plasmid preparation sampled from the original library. Plasmid preparations were obtained from sampling 96 colonies. These preparations were individually subjected to Sanger sequencing to verify their PAM sequences.

[0292] Purified APG09748 was incubated with sgRNA (27sg.2) in 1× CutSmart buffer (New England Biolabs B7204S) at room temperature for 20 minutes with a final concentration of 200 nM nuclease and 400 nM sgRNA.

[0293] These RNP solutions were added at a final concentration of 100 nM to solutions of plasmid DNA targets of final concentrations 8.3 nM and TB0125 and TB0089 reporters (described as SEQ ID NOs: 138 and 139, respectively) at 250 nM and 50 nM, respectively, in 1.5× Cutsmart buffer (New England Biolabs B7204S). To monitor fluorescence intensity, 10 μl of each reaction was incubated in a Corning low-volume 384-well microplate at 37 °C in a microplate reader (CLARIOSTAR Plus).

[0294] The plasmid concentrations in a small number of samples were below the target value for concentration normalization, or the volume of the solution was insufficient to deliver the intended target amount to the reaction wells. These samples were not excluded from the analysis and are thus expected to contribute to errors through inconsistencies between replicates. Samples using plasmids determined to have cross-contamination (multiple traces seen in the PAM region) by Sanger sequencing evaluation, or plasmids with target sequence changes removed from the analysis, the results of which are not shown below. The analysis results are shown in Table 16 below in descending order by gradient in the FAM channel.

[0295]

Table 18-1

[0296]

Table 18-2

[0297]

Table 18-3

[0298]

Table 18-4

[0299] The sequence with the highest gradient seems to match the predicted PAM (DTTN described in SEQ ID NO: 60) measured by the plasmid removal assay already described in International Apple. PCT / US2019 / 068079 is hereby incorporated by reference in its entirety. In particular, it can be seen that the two nucleotides at the 5'-side of the target have a strong selectivity for "T". Surprisingly, the most active PAM site (ATATG, SEQ ID NO: 147) observed in this assay does not exactly match the consensus of DTTN (SEQ ID NO: 60), suggesting a certain level of flexibility at this recognition site.

[0300] 11.4. ssDNA cleavage activity in the presence of PCR-amplified DNA The PCR-amplified targets were generated from genomic DNA (TRAC and VEGF amplicons shown as SEQ ID NO: 225 and SEQ ID NO: 226, respectively) using appropriate primers. The VEGF target (described as SEQ ID NO: 227) was PCR-amplified from HEK293T cell genomic DNA with primers having sequences LE573 and LE578 (described as SEQ ID NO: 228 and SEQ ID NO: 229, respectively). The TRAC target (shown as SEQ ID NO: 230) was PCR-amplified with LE257 and LE258 (shown as SEQ ID NO: 231 and SEQ ID NO: 232, respectively).

[0301] RNPs were formed by incubating the APG09106.1 nuclease and sgRNA described herein at 0.5 μM and 1 μM, respectively, in 1× NEBuffer 2 (New England Biolabs) and incubating at room temperature for 20 minutes.

[0302] [Table 19]

[0303] The cleavage reaction was performed in 1.5× NEBuffer 2 using a 1.5 μM ssDNA oligonucleotide reporter having a 5’ TEX615 label and a 3’ Iowa Black FQ quencher and 100 nM of each PCR product. Cleavage of the reporter probe results in dequenching of the fluorophore and thus an increase in the fluorescence signal. To monitor the fluorescence intensity, 10 μl of each reaction was incubated in a Corning low volume 384 well microplate at 37 °C in a microplate reader (CLARIOSTAR Plus). The results of the kinetic analysis are shown in Table 18.

[0304] [Table 20]

[0305] These results demonstrate the specific activation of trans ssDNA cleavage activity in the presence of PCR-amplified DNA from various sources. The activity is dependent on the concentration of the PCR-amplified substrate.

[0306] 11.5 Use of ssDNA cleavage for diagnosis These nucleases promise utility for implementation into diagnostic devices for the detection of infectious disease pathogens such as bacteria, viruses, or fungi, due to their ability to generate an optically detectable signal in the presence of a target DNA sequence.

[0307] The diagnostic procedure may include isolation or amplification of nucleic acids from the sample being tested. It may also be appropriate to use several samples without performing nucleic acid isolation or purification, because they may be detectable (such as by PCR) without amplification, or may be present in the sample in a sufficient amount to not contain substances that interfere with detection or signal generation.

[0308] Next, the RNP formed as described in the other examples can be exposed to the sample (or the sample processed as described in the previous section) together with a reporter such as a fluorophore- and quencher-modified single-stranded DNA oligonucleotide used in the previous examples, or another type of single-stranded DNA substrate that generates a signal that is visible or otherwise easily detectable upon cleavage. When using fluorophore-quencher conjugated DNA oligonucleotides (as in the previous examples), these can be detected using a fluorometer as described in the previous examples. To simplify detection, an endpoint assay can be performed instead of the above kinetic assay, which means that the assay is performed for a fixed time for positive and negative controls and can be read at the end of this elapsed time.

[0309] These reagents can also be incorporated into lateral flow test devices that enable the detection of a given pathogen or specific nucleic acid sequence (such as a disease allele in an individual) using minimal instrumentation. In this assay, the ssDNA reporter can be conjugated to multiple molecules suitable for antibody or affinity reagent capture, such as fluorescein, biotin, and / or digoxigenin.

Claims

[Claim 1] 1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, the polynucleotide comprising: a) an RGN polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 117, 30, 75, 1, 9, 16, 23, 38, 46, 61, 69, 82, 89, 95, 103, or 110; b) an RGN polypeptide comprising the amino acid sequence set forth in SEQ ID NO:54; The RGN polypeptide comprises a nucleotide sequence encoding an RGN polypeptide selected from the group consisting of: The RGN polypeptide is capable of binding to a target DNA sequence of a DNA molecule in an RNA guide sequence-specific manner when bound to a guide RNA (gRNA) capable of hybridizing to the target DNA sequence; The above nucleic acid molecule, wherein the polynucleotide encoding an RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide.