RNA-guided nucleases, active fragments and variants thereof and use methods

RNA-guided nucleases address the inefficiencies of traditional genome editing by using guide RNAs for precise targeting and modification, enabling efficient and versatile genome manipulation and gene expression control.

JP2025133875APending Publication Date: 2025-09-11LIFEEDIT INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025113206
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-02-13
Filing Date
2025-07-03
Publication Date
2025-09-11

AI Technical Summary

Technical Problem

Existing genome editing methods, such as meganucleases and TALENs, require costly and inefficient creation of chimeric nucleases for each target sequence, and introduce mutations through error-prone non-homologous end joining, limiting their precision and efficiency.

Method used

RNA-guided nucleases (RGNs) that utilize guide RNAs to target specific genomic sequences, enabling precise cleavage or modification through homologous recombination, and can be used with or without nuclease activity to deliver payloads or alter gene expression.

Benefits of technology

RGNs provide efficient and targeted genome manipulation, allowing for precise modifications, gene expression control, and delivery of payloads to specific genomic locations, enhancing the precision and versatility of genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025133875000001
    Figure 2025133875000001
  • Figure 2025133875000002
    Figure 2025133875000002
  • Figure 2025133875000003
    Figure 2025133875000003
Patent Text Reader

Abstract

To provide compositions and methods for binding to a target sequence of interest.SOLUTION: The invention provides a composition and a method for binding a target sequence of interest. The compositions find use in cleaving or modifying a target sequence of interest, visualization of a target sequence of interest, and modifying the expression of a sequence of interest. The composition comprises RNA-guided nuclease polypeptide, CRISPR RNA, trans-activating CRISPR RNA, guide RNA and nucleic acid molecules encoding the same. Vectors and host cells comprising the nucleic acid molecules are also provided. Further provided are a CRISPR system for binding a target sequence of interest, where the CRISPR system comprises an RNA-guided nuclease polypeptide and one or more guide RNAs. Thus, for example, it is possible to modify a target sequence at a genomic locus of eukaryotic cells or prokaryotic cells using the RNA-guided nuclease.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the fields of molecular biology and gene editing. [Background technology]

[0002] Targeted genome editing, or targeted genome modification, is rapidly becoming an important tool for basic and applied research. Early methods involved engineering nucleases (e.g., meganucleases, zinc finger proteins, and TALENs), which required the creation of chimeric nucleases with engineered, programmable sequence-specific DNA-binding domains specific for each particular target sequence. RNA-guided nucleases (e.g., the CRISPR-associated (cas) proteins of the clustered regularly interspaced short palindromic repeats (CRISPR)-cas bacterial system) can target specific sequences by complexing with guide RNAs that specifically hybridize to specific target sequences. Creating target-specific guide RNAs is less costly and more efficient than creating chimeric nucleases for each target sequence. Genome editing can be achieved by using such RNA-guided nucleases to introduce sequence-specific double-strand breaks. Because the repair of these double-strand breaks involves error-prone non-homologous end joining (NHEJ), mutations are introduced at specific genomic locations. Alternatively, heterologous DNA can be introduced into a genomic site through homology-directed repair. Summary of the Invention

[0003] Compositions and methods for binding to a target sequence of interest are provided. The compositions have applications for cleaving or modifying the target sequence of interest, visualizing the target sequence of interest, and altering the expression of the target sequence of interest. The compositions include an RNA-guided nuclease (RGN) polypeptide, a CRISPR RNA (crRNA), a trans-activating CRISPR RNA (tracrRNA), a guide RNA (gRNA), and nucleic acid molecules encoding them. Vectors and host cells contain these nucleic acid molecules. Also provided are CRISPR systems comprising an RNA-guided nuclease polypeptide and one or more guide RNAs for binding to a target sequence of interest. Thus, methods disclosed herein relate to binding to a target sequence of interest, and in some embodiments, the methods cleave or modify the target sequence of interest. The target sequence of interest can be modified, for example, as a result of non-homologous end joining or homologous recombination repair using an introduced donor sequence. DETAILED DESCRIPTION OF THE INVENTION

[0004] Numerous modifications and other embodiments of the inventions described herein will come to mind to one skilled in the art to which this invention pertains having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. It is to be understood, therefore, that the invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the accompanying embodiments. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0005] I. Overview

[0006] RNA-guided nucleases (RGNs) enable targeted manipulation of single sites within the genome, making them useful in the context of gene targeting in therapeutic and research applications. In a variety of organisms, including mammals, RNA-guided nucleases have been used for genome manipulation, for example, by stimulating non-homologous end joining and homologous recombination. The compositions and methods described herein are useful for generating single- or double-strand breaks in polynucleotides, modifying polynucleotides, detecting specific sites within polynucleotides, or altering the expression of specific genes.

[0007] The RNA-guided nucleases disclosed herein can alter gene expression by modifying target sequences. In particular embodiments, the RNA-guided nucleases are directed to the target sequence by a guide RNA (gRNA) as part of a clustered regularly interspaced short palindromic repeats (CRISPR) RNA-guided nuclease system. The guide RNA forms a complex with the RNA-guided nuclease, allowing the RNA-guided nuclease to bind to the target sequence, and in some embodiments, introduces a single-strand or double-strand break at the location of the target sequence. After the target sequence is cleaved, the DNA sequence of the target sequence can be modified during the repair process of the break. Thus, provided herein is a method for modifying a target sequence in the DNA of a host cell using an RNA-guided nuclease. For example, the RNA-guided nuclease can be used to modify a target sequence located at a genomic locus in a eukaryotic or prokaryotic cell.

[0008] II. RNA-guided nucleases

[0009] The present specification provides an RNA-guided nuclease. The term RNA-guided nuclease (RGN) refers to a polypeptide that binds to a specific target nucleotide sequence in a sequence-specific manner, and is directed to the target nucleotide sequence by a guide RNA molecule that forms a complex with the polypeptide and hybridizes with the target sequence. Although the RNA-guided nuclease can cleave the target sequence when bound to the target sequence, the term RNA-guided nuclease also encompasses nuclease-dead RNA-guided nucleases that can bind to the target sequence but cannot cleave the target sequence. The cleavage of the target sequence by the RNA-guided nuclease can result in a single-strand cleavage or a double-strand cleavage. The RNA-guided nuclease that can cleave a single strand of a double-stranded nucleic acid molecule is referred to herein as a nickase.

[0010] The RNA-guided nucleases disclosed herein include the RNA-guided nucleases APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, and APG1688.1 (the amino acid sequences of which are set forth in SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, respectively), and active fragments or variants thereof that retain the ability to bind to target nucleotide sequences in an RNA-guided sequence-specific manner. In some of these embodiments, active fragments or variants of RGN APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, APG1688.1 are capable of cleaving single-stranded or double-stranded target sequences. In some embodiments, the active variants of RGN APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, APG1688.1 comprise an amino acid sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the amino acid sequence set forth in any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54. In some embodiments, an active fragment of RGN APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, APG1688.1 comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more consecutive amino acid residues of the amino acid sequence set forth in any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54.The RNA-guided nucleases provided herein can comprise at least one nuclease domain (e.g., a DNase domain, an RNase domain) and at least one RNA recognition and / or RNA-binding domain that interacts with the guide RNA. Non-limiting examples of additional domains that can be found in the RNA-guided nucleases provided herein include a DNA-binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain. In particular embodiments, the RNA-guided nucleases provided herein can comprise at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of one or more of the DNA-binding domain, the helicase domain, the protein-protein interaction domain, and the dimerization domain.

[0011] The target nucleotide sequence is bound by the RNA-guided nuclease provided herein, and a guide RNA associated with the RNA-guided nuclease is hybridized to it. If the RNA-guided nuclease has nuclease activity, the target sequence can then be cleaved by the RNA-guided nuclease. The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond in the backbone of the target sequence, resulting in a single- or double-stranded break in the target nucleic acid. The RGNs of the present disclosure can function as endonucleases to cleave nucleotides within a polynucleotide, or as exonucleases to sequentially remove nucleotides from the ends (5' and / or 3' ends) of a polynucleotide. In another embodiment, the disclosed RGNs can cleave nucleotides of the target sequence at any position in the polynucleotide, and thus can function as both endonucleases and exonucleases. Cleavage of a target polynucleotide by the RGNs of the present disclosure can result in sticky or blunt ends.

[0012] The RNA-guided nucleases disclosed herein can be wild-type sequences derived from bacterial or archaeal species. Alternatively, the RNA-guided nucleases can be variants or fragments of wild-type polypeptides. Wild-type RGNs can be modified to, for example, alter nuclease activity or PAM specificity. In some embodiments, the RNA-guided nuclease is not naturally occurring.

[0013] In some embodiments, the RNA-guided nuclease functions as a nickase and cleaves only a single strand of the target nucleotide sequence. Such RNA-guided nucleases have a single functional nuclease domain. In some of these embodiments, the additional nuclease domain is mutated to reduce or eliminate nuclease activity.

[0014] In another embodiment, the RNA-guided nuclease lacks or exhibits reduced nuclease activity and is therefore referred to herein as nuclease-dead. Any method known in the art for introducing mutations into an amino acid sequence (e.g., PCR-mediated mutagenesis, site-directed mutagenesis, etc.) can be used to generate a nickase or nuclease-dead RGN. See, e.g., U.S. Patent Application Publication No. 2014 / 0068797 and U.S. Patent No. 9,790,490, each of which is incorporated by reference in its entirety.

[0015] RNA-guided nucleases lacking nuclease activity can be used to deliver fused polypeptides, polynucleotides, or small molecule payloads to specific genomic locations. In some of these embodiments, RGN polypeptides or guide RNAs are fused to detectable labels, allowing for the detection of specific sequences. As a non-limiting example, nuclease-dead RGN can be fused to detectable labels (e.g., fluorescent proteins) to target specific disease-related sequences, allowing for the detection of disease-related sequences.

[0016] Alternatively, nuclease-dead RGN can be targeted to specific genomic locations to alter expression of a desired sequence. In some embodiments, binding of the nuclease-dead RGN-induced nuclease to a target sequence inhibits expression of the target sequence or of a gene under transcriptional control by the target sequence by interfering with the binding of RNA polymerase or transcription factors within the targeted genomic region. In other embodiments, the RGN (e.g., nuclease-dead RGN) or guide RNA complexed with RGN further comprises an expression modulator that, upon binding to the target sequence, inhibits or activates expression of the target sequence or of a gene under transcriptional control by the target sequence. In some of these embodiments, the expression modulator alters expression of the target sequence or of a gene regulated through an epigenetic mechanism.

[0017] In another embodiment, a nuclease-dead RGN or an RGN with only nickase activity can be targeted to a specific genomic location and fused to a base-editing polypeptide (e.g., a deaminase polypeptide, or an active variant or fragment thereof that deaminates nucleotide bases) to modify the sequence of a target polynucleotide, thereby converting one nucleotide into another. The base-editing polypeptide can be fused to the N-terminus or C-terminus of the RGN. Additionally, the base-editing polypeptide can be fused to the RGN via a peptide linker. Non-limiting examples of deaminase polypeptides useful in such compositions and methods include the cytidine deaminase base editors or adenosine deaminase base editors described in Gaudelli et al. (2017) Nature 551:464-471, U.S. Patent Application Publication Nos. 2017 / 0121693 and 2018 / 0073012, and International Application Publication No. WO / 2018 / 027078, each of which is incorporated by reference herein in its entirety.

[0018] The RNA-guided nuclease fused to a polypeptide or domain can be separated or joined by a linker. The term "linker," as used herein, refers to a chemical group or molecule that links two molecules or moieties (e.g., the binding domain and cleavage domain of a nuclease). In some embodiments, a linker joins the gRNA-binding domain of an RNA-guided nuclease to a base-editing polypeptide (such as a deaminase). In some embodiments, a linker links a nuclease-dead RGN to a deaminase. Typically, a linker is located between or adjacent to two groups, molecules, or other moieties and is connected to each other through a covalent bond, thereby connecting the two. In some embodiments, the linker is a single amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker is 5 to 100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers are also contemplated.

[0019] The RNA-guided nuclease of the present disclosure can include at least one nuclear localization signal (NLS) to enhance transport of RGN to the nucleus of a cell. Nuclear localization signals are known in the art and generally comprise a series of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In particular embodiments, RGN comprises two, three, four, five, six, or more nuclear localization signals. The nuclear localization signal can be a heterologous NLS. Non-limiting examples of nuclear localization signals useful for the RGN of the present disclosure include the nuclear localization signals of SV40 large T antigen, nucleopasmin, and c-Myc (see, e.g., Ray et al. (2015) Bioconjug Chem 26(6):1004-1007). In a particular embodiment, RGN comprises the NLS sequence set forth as SEQ ID NO: 67. RGN can comprise one or more NLS sequences at the N-terminus, the C-terminus, or both the N-terminus and the C-terminus. For example, RGN can comprise two NLS sequences in the N-terminal region and four NLS sequences in the C-terminal region.

[0020] Other localization signals known in the art that localize polypeptides to specific locations within a cell can also be used to target RGN, including, but not limited to, plastid localization sequences, mitochondrial localization sequences, and dual targeting signal sequences that target both plastids and mitochondria (e.g., Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol dx.doi.org / 10.3389 / fphys.2015.00259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soll (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al. (2009) FEBS J 276:1187-1195; Silva-Filho (2003) Curr Opin Plant Biol 6:589-595; Peeters and Small (2001) Biochim Biophys Acta 1541:54-63; Murcha et al. (2014) J Exp Bot 65:6301-6335; Mackenzie (2005) Trends Cell Biol 15:548-554; Glaser et al. (1998) Plant Mol Biol 38:311-338).

[0021] In some embodiments, the RNA-guided nuclease of the present disclosure comprises at least one cell entry domain that facilitates the uptake of RGN into cells. Cell entry domains are known in the art and generally comprise a series of positively charged amino acid residues (i.e., polycationic cell entry domains), or alternating polar and non-polar amino acid residues (i.e., amphipathic cell entry domains), or hydrophobic amino acid residues (i.e., hydrophobic cell entry domains) (see, e.g., Milletti F. (2012) Drug Discov Today 17:850-860). One non-limiting example of a cell entry domain is the trans-activating transcription activator (TAT) from human immunodeficiency virus 1.

[0022] The nuclear localization signal, and / or the plastid localization sequence, and / or the mitochondrial localization sequence, and / or the dual targeting signal sequence, and / or the cell entry domain can be located at the amino terminus (N-terminus), or the carboxyl terminus (C-terminus), or at an internal location of the RNA-guided nuclease.

[0023] The RGN of the present disclosure can be fused directly to an effector domain (such as a cleavage domain, a deaminase domain, or an expression modulator domain) or indirectly fused via a linker peptide. Such domains can be located at the N-terminus, C-terminus, or internal position of the RNA-guided nuclease. In some of these embodiments, the RGN component of the fusion protein is nuclease-dead RGN.

[0024] In some embodiments, the RGN fusion protein comprises a cleavage domain, which can be any domain capable of cleaving polynucleotides (i.e., RNA, DNA, or RNA / DNA hybrids), non-limiting examples of which include restriction endonucleases and homing endonucleases, such as Type IIS endonucleases (e.g., FokI) (see, e.g., Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388; Linn et al. (eds.), Nucleases, Cold Spring Harbor Laboratory Press, 1993).

[0025] In another embodiment, the RGN fusion protein comprises a deaminase domain that deaminates a nucleotide base, converting one nucleotide base to another, non-limiting examples of which include a cytidine deaminase base editor or an adenosine deaminase base editor (see, e.g., Gaudelli et al. (2017) Nature 551:464-471; U.S. Patent Application Publication Nos. 2017 / 0121693, 2018 / 0073012; U.S. Patent No. 9,840,699; and International Application Publication No. WO / 2018 / 027078).

[0026] In some embodiments, the effector domain of an RGN fusion protein can be an expression modulator domain. An expression modulator domain is a domain that functions to up- or down-regulate transcription. The expression modulator domain can be an epigenetic modification domain, a transcriptional repression domain, or a transcriptional activation domain.

[0027] In some of these embodiments, the expression modulator of an RGN fusion protein comprises an epigenetic modification domain that modifies covalent bonds in DNA or histone proteins, altering histone and / or chromosomal structure without altering the DNA sequence, resulting in altered (up- or down-regulated) gene expression. Non-limiting examples of epigenetic modifications include acetylation or methylation of lysine residues in DNA, arginine methylation, serine and threonine phosphorylation, lysine ubiquitination, sumoylation of histone proteins, and methylation and hydroxymethylation of cytosine residues. Non-limiting examples of epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains, and DNA demethylase domains.

[0028] In another embodiment, the expression modulator of the fusion protein comprises a transcriptional repression domain that interacts with a transcriptional control element and / or a transcriptional regulatory protein (e.g., RNA polymerase, transcription factor, etc.) to reduce or stop transcription of at least one gene. Transcriptional repression domains are known in the art, and non-limiting examples include Sp1-like repressor, IκB, and Krüppel-associated box (KRAB) domains.

[0029] In yet another embodiment, the expression modulator of the fusion protein comprises a transcriptional activation domain that interacts with a transcriptional control element and / or a transcriptional regulatory protein (such as RNA polymerase, transcription factor, etc.) to increase or activate transcription of at least one gene. Transcriptional activation domains are known in the art, and non-limiting examples include the herpes simplex virus VP16 activation domain and the NFAT activation domain.

[0030] The RGN polypeptide of the present disclosure can include a detectable label or purification tag. The detectable label or purification tag can be located directly at the N-terminus, C-terminus, or internal position of the RNA-guided nuclease, or indirectly via a linker peptide. In some of these embodiments, the RGN component of the fusion protein is nuclease-dead RGN. In other embodiments, the RGN component of the fusion protein is RGN with nickase activity.

[0031] Detectable labels can be visualized or otherwise observed. Detectable labels can be fused to RGN as a fusion protein (e.g., a fluorescent protein) or can be small molecules complexed to the RGN polypeptide and detectable visually or by other means. Detectable labels that can be fused to RGN of the present disclosure as a fusion protein include any detectable protein domain, non-limiting examples of which include fluorescent proteins and protein domains that can be detected using specific antibodies. Non-limiting examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, EGFP, ZsGreen1) and yellow fluorescent proteins (e.g., YFP, EYFP, ZsYellow1). Non-limiting examples of detectable small molecule labels include radioactive labels. 3 H, 35 S, etc.

[0032] The RGN polypeptide can also include a purification tag, which is any molecule that can be used to isolate a protein or fusion protein from a mixture (e.g., a biological sample, a culture medium). Non-limiting examples of purification tags include myc, maltose-binding protein (MBP), and glutathione-S-transferase (GST).

[0033] II. Guide RNA

[0034] The present disclosure provides guide RNAs and polynucleotides encoding them. The term "guide RNA" refers to a nucleotide sequence that is sufficiently complementary to a target nucleotide sequence to hybridize with the target sequence and allow an associated RNA-guided nuclease to bind to the target nucleotide sequence in a sequence-specific manner. Thus, each guide RNA of an RGN is one or more RNA molecules (typically one or two) that can bind to and guide the RGN to bind to a specific target nucleotide sequence, and, if the RGN has nickase or nuclease activity, also cleave the target nucleotide sequence. Generally, guide RNAs include CRISPR RNAs (crRNAs) and trans-activating CRISPR RNAs (tracrRNAs). Natural guide RNAs, including both crRNAs and tracrRNAs, generally comprise two separate RNA molecules that hybridize to each other via the repeat sequence of the crRNA and the anti-repeat sequence of the tracrRNA.

[0035] Naturally occurring direct repeat sequences within a CRISPR array typically range from 28 to 37 base pairs in length, but this length can vary between approximately 23 bp and approximately 55 bp. Spacer sequences within a CRISPR array typically range from 32 to 38 base pairs in length, but this length can vary between approximately 21 bp and approximately 72 bp. Each CRISPR array typically contains fewer than 50 units of CRISPR repeat spacer sequences. CRISPRs are transcribed as part of a long transcript called the primary CRISPR transcript, which contains many of the CRISPR arrays. The primary CRISPR transcript is cleaved by Cas proteins to generate crRNA, and in some cases, pre-crRNA. The pre-crRNA is further processed by additional Cas proteins to form mature crRNA. The mature crRNA contains the spacer sequence and CRISPR repeats. In some embodiments, where pre-crRNA is processed into mature (or processed) crRNA, maturation includes removal of about 1 to about 6 or more 5' nucleotides, or removal of 3' nucleotides, or removal of 5' and 3' nucleotides. For purposes of genome editing or targeting a specific target nucleotide sequence of interest, the nucleotides removed during maturation of the pre-crRNA molecule are not required for the generation or design of the guide RNA.

[0036] CRISPR RNA (crRNA) contains a spacer sequence and a CRISPR repeat sequence. A "spacer sequence" is a nucleotide sequence that directly hybridizes to a target nucleotide sequence of interest. The spacer sequence is engineered to be fully or partially complementary to the target nucleotide sequence of interest. In various embodiments, the spacer sequence can comprise from about 8 to about 30 or more nucleotides. For example, the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the spacer sequence is about 10 to about 26 nucleotides in length, or about 12 to about 30 nucleotides in length. In particular embodiments, the spacer sequence is about 30 nucleotides in length. In some embodiments, the degree of complementarity between the spacer sequence and its corresponding target sequence, when optimally aligned using a suitable alignment program, is about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or more. In particular embodiments, the spacer sequence adopts a free secondary structure, which can be predicted using any suitable polynucleotide folding algorithm known in the art, non-limiting examples of which include mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).

[0037] RGN proteins may vary in their sensitivity to mismatches between the spacer sequence in the gRNA and the gRNA's target sequence, and this sensitivity affects the efficiency of cleavage. As discussed in Example 5, the APG05459.1 RGN is unusually sensitive to mismatches between the spacer sequence and the target sequence spanning at least 15 nucleotides 5' of the PAM site. Thus, APG05459.1 has the ability to target specific sequences more finely (i.e., specifically) than other RGNs that are less sensitive to mismatches between the spacer sequence and the target sequence.

[0038] The CRISPR RNA repeat sequence comprises a nucleotide sequence that includes a region of sufficient complementarity to hybridize to the tracrRNA. In various embodiments, the CRISPR RNA repeat sequence can comprise from about 8 to about 30 or more nucleotides. For example, the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the CRISPR repeat sequence is about 21 nucleotides in length. In some embodiments, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA sequence is about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or more, when optimally aligned using a suitable alignment program. In particular embodiments, the CRISPR repeat comprises the nucleotide sequence of any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, or 55, or an active variant or fragment thereof, which, when included in a guide RNA, enables sequence-specific binding of the associated RNA-guided nuclease provided herein to a target sequence of interest. In some embodiments, an active CRISPR repeat variant of a wild-type sequence comprises a nucleotide sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical in sequence to the nucleotide sequence set forth as any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, or 55.In some embodiments, an active CRISPR repeat fragment of a wild-type sequence comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the nucleotide sequence set forth as SEQ ID NO: 2, 12, 20, 28, 37, 46, or 55.

[0039] In some embodiments, the crRNA does not occur in nature. In some of these embodiments, the specific CRISPR repeat sequence is not naturally linked to the engineered spacer sequence, and the CRISPR repeat sequence is considered heterologous to the spacer sequence. In some embodiments, the spacer sequence is an engineered sequence that does not occur in nature.

[0040] A transactivating CRISPR RNA molecule, or tracrRNA molecule, contains a nucleotide sequence that includes a region of sufficient complementarity to hybridize to the CRISPR repeat sequence of the crRNA (referred to herein as the anti-repeat region). In some embodiments, the tracrRNA molecule further includes a region with a secondary structure (e.g., a stem-loop) or forms a secondary structure when hybridized to the corresponding crRNA. In particular embodiments, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is at the 5' end of the molecule, and the 3' end of the tractRNA contains a secondary structure. This region of secondary structure typically contains several hairpin structures, including a junction hairpin found adjacent to the anti-repeat sequence. The junction hairpins often have conserved nucleotide sequences within the bases of the hairpin stem, and the motifs UNANNG, UNANNU, and UNANNA (SEQ ID NOS: 68, 557, and 558, respectively) are found in many junction hairpins within the tracrRNA. Multiple terminal hairpins are often present at the 3' end of the tracrRNA. Their structure and number can vary, but they often contain a CG-rich, Rho-independent transcription termination hairpin followed by a string of U's at the 3' end. See, e.g., Briner et al. (2014) Molecular Cell 56:333-339; Briner and Barrangou (2016) Cold Spring Harb Protoc; doi: 10.1101 / pdb.top090902; and U.S. Patent Application Publication No. 2017 / 0275648, each of which is incorporated by reference in its entirety.

[0041] In various embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence contains about 8 to about 30 or more nucleotides. For example, the region that base pairs between the tracrRNA anti-repeat region and the CRISPR repeat sequence can be about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more nucleotides in length. In particular embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is about 20 nucleotides in length. In some embodiments, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA anti-repeat sequence is about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or more.

[0042] In various embodiments, the entire tracrRNA can comprise from about 60 nucleotides to more than about 140 nucleotides. For example, the tracrRNA can be about 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, or more nucleotides in length. In particular embodiments, the tracrRNA is about 80 to about 90 nucleotides in length, including lengths of about 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, and 90 nucleotides. In some embodiments, the tracrRNA is about 85 nucleotides in length.

[0043] In particular embodiments, the tracrRNA comprises the nucleotide sequence of any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56, or an active variant or fragment thereof, which, when included in a guide RNA, enables sequence-specific binding of the relevant RNA-guided nucleases provided herein to a target sequence of interest. In some embodiments, an active tracrRNA sequence variant of the wild-type sequence comprises a nucleotide sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical in sequence to the nucleotide sequence set forth as any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56. In some embodiments, an active tracrRNA sequence fragment of a wild-type sequence comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides of the nucleotide sequence set forth as any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56.

[0044] Two polynucleotide sequences can be considered substantially complementary when the two sequences hybridize to each other under stringent conditions. Similarly, an RGN is considered to bind to a particular target sequence in a sequence-specific manner if a guide RNA bound to the RGN binds to that target sequence under stringent conditions. "Stringent conditions" or "stringent hybridization conditions" refer to conditions under which two polynucleotide sequences hybridize to each other with a detectability greater than other sequences (e.g., at least twice the background). Stringent conditions are sequence-dependent and will vary under different circumstances. Typically, stringent conditions are conditions in which the salt concentration is less than about 1.5 M Na ions, typically about 0.01 to 1.0 M Na ions (or other salts), at pH 7.0 to 8.3, and the temperature is at least about 30°C for short sequences (e.g., 10 to 50 nucleotides) and at least about 60°C for long sequences (e.g., more than 50 nucleotides). Stringent conditions can also be achieved by the addition of a destabilizing agent (e.g., formamide). Typical low-stringency conditions include hybridization in a buffer solution of 30-35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate) at 37°C, followed by a wash in 1x-2x SSC (20x SSC = 3.0 M NaCl / 0.3 M trisodium citrate) at 50-55°C. Typical medium-stringency conditions include hybridization in 40-45% formamide, 1.0 M NaCl, and 1% SDS at 37°C, followed by a wash in 0.5x-1x SSC at 55-60°C. Typical high-stringency conditions include hybridization in 50% formamide, 1 M NaCl, and 1% SDS at 37°C, followed by a wash in 0.1x SSC at 60-65°C. In some cases, the wash buffer may contain about 0.1% to about 1% SDS. Hybridization times are generally less than about 24 hours, usually about 4 to about 12 hours. Wash times will be at least long enough to reach equilibrium.

[0045] Tm is the temperature at which 50% of a complementary target sequence hybridizes to a perfectly matched sequence (under a defined ionic strength and pH). For DNA-DNA hybrids, Tm can be roughly determined from the formula (Meinkoth and Wahl, 1984) Anal. Biochem. 138:267-284): Tm = 81.5°C + 16.6 (log M) + 0.41(%GC) - 0.61(%form) - 500 / L, where M is the number of moles of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, %form is the percentage of formamide in the hybridization solution, and L is the hybrid length in base pairs. Generally, stringent conditions are selected to be approximately 5°C lower than the melting point (Tm) of a specific sequence and its complementary sequence at a defined ionic strength and pH. However, highly stringent conditions can utilize hybridization and / or washing at 1°C, or 2°C, or 3°C, or 4°C below the melting temperature (Tm); moderately stringent conditions can utilize hybridization and / or washing at 6°C, or 7°C, or 8°C, or 9°C, or 10°C below the melting temperature (Tm); and low stringency conditions can utilize hybridization and / or washing at 11°C, or 12°C, or 13°C, or 14°C, or 15°C, or 20°C below the melting temperature (Tm). One of skill in the art will understand that variations in the stringency of hybridization and / or washing solutions are essentially described by the above formula, hybridization and washing compositions, and desired Tm.Comprehensive guides to nucleic acid hybridization can be found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology - Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York); Ausubel et al. (eds.) (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2nd ed.), Cold Spring Harbor Laboratory Press, Plainview, NY.

[0046] The guide RNA can be a single guide RNA or a dual guide RNA system. A single guide RNA contains a crRNA and a tracrRNA on a single RNA molecule, whereas a dual guide RNA system contains a crRNA and a tracrRNA on two different RNA molecules, hybridized with each other through at least a portion of the CRISPR repeat sequence of the crRNA and at least a portion of the tracrRNA (which can be fully or partially complementary to the CRISPR repeat sequence of the crRNA). In some embodiments where the guide RNA is a single guide RNA, the crRNA and the tracrRNA are separated by a linker nucleotide sequence. Generally, the linker nucleotide sequence does not contain complementary bases to avoid the formation of secondary structures within or involving the nucleotides of the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between the crRNA and tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or more nucleotides in length. In particular embodiments, the linker nucleotide sequence of the single guide RNA is at least 4 nucleotides in length. In some embodiments, the linker nucleotide sequence is a nucleotide sequence represented as SEQ ID NO: 63, 64, or 65. In other embodiments, the linker nucleotide sequence is at least 6 nucleotides in length. In some embodiments, the linker nucleotide sequence is a nucleotide sequence represented as SEQ ID NO: 65.

[0047] Single-guide RNAs or dual-guide RNAs can be chemically synthesized or synthesized through in vitro transcription. Assays for determining sequence-specific binding between RGN and guide RNA are known in the art, including non-limiting examples of binding assays between expressed RGN and guide RNA. Guide RNAs can be tagged with a detectable label (e.g., biotin) and used in pull-down detection assays, in which guide RNA:RGN complexes are captured via the detectable label (e.g., using streptavidin beads). A control guide RNA with a sequence or structure unrelated to the guide RNA can be used as a negative control for non-specific binding of RGN to RNA. In some embodiments, the guide RNA is SEQ ID NO: 10, 18, 26, 35, 44, 53, or 62, and the spacer sequence therein can be any sequence and is represented as a poly-N sequence.

[0048] As described in Example 8, some RGNs of the present invention can share some guide RNAs. APG05083.1, APG07433.1, APG07513.1, and APG08290.1 ​​can function using a guide RNA comprising a crRNA containing the nucleotide sequence of any of SEQ ID NOs: 2, 12, 20, or 28, respectively, and a corresponding tracrRNA containing the nucleotide sequence of any of SEQ ID NOs: 3, 13, 21, or 29, respectively. Furthermore, APG04583.1 and APG01688.1 can function using a guide RNA comprising a crRNA containing the nucleotide sequence of SEQ ID NO: 46 or 55, respectively, and a corresponding tracrRNA containing the nucleotide sequence of SEQ ID NO: 47 or 56, respectively.

[0049] In some embodiments, the guide RNA can be introduced into a target cell, or organelle, or embryo as an RNA molecule. The guide RNA can be transcribed in vitro or chemically synthesized. In other embodiments, a nucleotide sequence encoding the guide RNA is introduced into a target cell, or organelle, or embryo. In some of these embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., an RNA polymerase III promoter). The promoter can be a native promoter or a heterologous promoter to the nucleotide sequence encoding the guide RNA.

[0050] In various embodiments, the guide RNA can be introduced into a target cell, or organelle, or embryo as a ribonucleoprotein complex, as described herein, wherein the guide RNA binds to an RNA-guided nuclease polypeptide.

[0051] The guide RNA hybridizes to a specific target nucleotide sequence of interest, thereby directing the associated RNA-guided nuclease to its target nucleotide sequence. The target nucleotide sequence can comprise DNA, RNA, or a combination of both, and can be single-stranded or double-stranded. The target nucleotide sequence can be genomic DNA (i.e., chromosomal DNA), plasmid DNA, or an RNA molecule (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). The target nucleotide sequence can be bound by the RNA-guided nuclease in vitro or in cells. The chromosomal sequence targeted by RGN can be a chromosomal sequence in the nucleus, plastid, or mitochondrion. In some embodiments, there is only one target nucleotide sequence in the target genome.

[0052] The target nucleotide sequence is flanked by a protospacer motif (PAM). The protospacer motif is generally located within about 1 to about 10 nucleotides of the target nucleotide sequence, including about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from the target nucleotide sequence. The PAM can be 5' or 3' of the target sequence. In some embodiments, the PAM is 3' of the target sequence for the RGNs of the present disclosure. Generally, the PAM is a consensus sequence of about 3-4 nucleotides, although in particular embodiments, it can be 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length. In various embodiments, the PAM sequence recognized by the RGNs of the present disclosure comprises the consensus sequence set forth as any of SEQ ID NOs: 6, 32, 41, 50, and 59. Non-limiting PAM sequences are the nucleotide sequences represented as SEQ ID NOs: 7, 69, 70, 71, 72.

[0053] In particular embodiments, an RNA-guided nuclease having any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence represented as any of SEQ ID NOs: 6, 32, 41, 50, 59, 7, respectively. In some of these embodiments, the RGN comprises a CRISPR repeat sequence represented as any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, or an active variant or fragment thereof, and binds to a guide sequence comprising a tracrRNA sequence represented as any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56, or an active variant or fragment thereof, respectively. The RGN system is described in further detail in Example 1 and Table 1 herein.

[0054] It is well known that the PAM sequence specificity for a given nuclease enzyme is affected by the concentration of the enzyme (see, e.g., Karvelis et al. (2015) Genome Biol 16:253). This concentration can be altered by changing the promoter used to express RGN or by varying the amount of ribonucleoprotein complex delivered to either the cell, organelle, or embryo.

[0055] Upon recognizing the corresponding PAM sequence, RGN can cleave the target nucleotide sequence at a specific cleavage site. As used herein, a cleavage site is comprised of two specific nucleotides within the target nucleotide sequence between which the nucleotide sequence is cleaved by RGN. The cleavage site can include the first and second nucleotides, or the second and third nucleotides, or the third and fourth nucleotides, or the fourth and fifth nucleotides, or the fifth and sixth nucleotides, or the seventh and eighth nucleotides, or the eighth and ninth nucleotides, 5' or 3' from the PAM. In some embodiments, the cleavage site can be located more than 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides 5' or 3' from the PAM. In some embodiments, the cleavage site is separated by four nucleotides from the PAM. In other embodiments, the cleavage site is separated by at least 15 nucleotides from the PAM. Because RGN can cleave a target nucleotide sequence to produce sticky ends, in some embodiments, the cleavage site is defined based on a distance of two nucleotides from the PAM on the plus (+) strand of the polynucleotide and a distance of two nucleotides from the PAM on the minus (-) strand of the polynucleotide.

[0056] III. Nucleotides encoding RNA-guided nucleases, and / or CRISPR RNAs, and / or tracrRNAs

[0057] The present disclosure provides polynucleotides comprising a CRISPR RNA, and / or tracrRNA, and / or sgRNA of the present disclosure, and polynucleotides comprising nucleotide sequences encoding an RNA-guided nuclease, and / or CRISPR RNA, and / or tracrRNA, and / or sgRNA of the present disclosure. The polynucleotides of the present disclosure include polynucleotides comprising or encoding a CRISPR repeat sequence comprising the nucleotide sequence of any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, or 55, or an active variant or fragment thereof, which, when included in a guide RNA, enables sequence-specific binding of the associated RNA-guided nuclease to a target sequence of interest. Also disclosed are polynucleotides comprising or encoding a tracrRNA comprising the nucleotide sequence of any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56, or an active variant or fragment thereof, which, when included in a guide RNA, enables sequence-specific binding of the associated RNA-guided nuclease to a desired target sequence. Also provided are polynucleotides encoding an RNA-guided nuclease comprising the amino acid sequence set forth as any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, or 54, or an active fragment or variant thereof, which retain the ability to bind to a target nuclease sequence in an RNA-guided sequence-specific manner.

[0058] The term "polynucleotide," as used herein, is not limited to polynucleotides containing DNA. Those skilled in the art will recognize that polynucleotides can include ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. These include peptide nucleic acids (PNAs), PNA-DNA chimeras, bridged nucleic acids (LNAs), and phosphothiolate bridged sequences. Polynucleotides disclosed herein encompass any sequence, including, but not limited to, single-stranded forms, double-stranded forms, DNA-RNA hybrids, triplex structures, and stem-loop structures.

[0059] Nucleic acid molecules encoding RGN can be codon optimized for expression in a desired organism. A "codon-optimized" coding sequence is a polynucleotide encoding a sequence with codon usage designed to mimic the preferred codon usage or transcription conditions of a particular host cell. One or more codons are changed at the nucleic acid level without altering the translated amino acid sequence, resulting in increased expression in that particular host cell or organism. Codon optimization can be performed on all or part of a nucleic acid molecule. Codon tables and other references providing preference information for a wide range of organisms are available in the art (see, e.g., Campbell and Gowri (1990) Plant Physiol. 92:1-11 for plant-preferred codon usage). Methods are available in the art for synthesizing plant-preferred genes. See, e.g., U.S. Patent Nos. 5,380,831, 5,436,391, Murray et al. (1989) Nucleic Acids Res. 17:477-498, incorporated herein by reference.

[0060] Polynucleotides encoding the RGN, crRNA, tracrRNA, and / or sgRNA provided herein can be provided in expression cassettes and expressed in vitro or in a cell, organelle, embryo, or organism of interest. The cassettes can include 5' and 3' regulatory sequences operably linked to the polynucleotide encoding the RGN, crRNA, tracrRNA, and / or sgRNA provided herein, allowing for expression of the polynucleotide. The cassettes can further include at least one additional gene or genetic element for cotransformation into an organism. When additional genes or elements are included, these elements are operably linked. The term "operably linked" refers to a functional linkage between two or more elements. For example, an operably linked linkage between a promoter and a coding region of interest (e.g., a region encoding the RGN, crRNA, tracrRNA, and / or sgRNA) is a functional linkage that allows for expression of the coding region of interest. Operably linked elements can be contiguous or non-contiguous. "Operably linked," when used to refer to the junction of two protein-coding regions, means that the coding regions are in the same reading frame. Alternatively, additional genes or elements can be provided on multiple expression cassettes. For example, the nucleotide sequence encoding the RGN of the present disclosure can be on one expression cassette, while the nucleotide sequence encoding the crRNA, trancrRNA, or complete guide RNA can be on another expression cassette. Such expression cassettes are provided with multiple restriction and / or recombination sites for inserting a polynucleotide under the transcriptional control of the regulatory region. The expression cassette can additionally include a selectable marker gene.

[0061] The expression cassette will include, in the 5' to 3' direction of transcription, a transcriptional (and in some cases translational) initiation region (i.e., promoter) of the invention functional in the organism of interest, a polynucleotide encoding RGN, a polynucleotide encoding the crRNA, a polynucleotide encoding the tracrRNA, a polynucleotide encoding the sgRNA, and a transcriptional (and in some cases translational) termination region (i.e., termination region). The promoter of the invention can direct or command expression of a coding sequence in a host cell. These regulatory regions (e.g., promoter, transcriptional regulatory region, translational termination region) can be endogenous to the host cell, heterologous to the host, or a mixture of endogenous and heterologous. As used herein, "heterologous" with respect to a sequence refers to a sequence derived from a foreign species or, if derived from the same species, substantially altered by human intervention from its natural form in composition and / or genomic locus. As used herein, a chimeric gene comprises a coding sequence operably linked to a transcriptional initiation region heterologous to the coding sequence.

[0062] Convenient termination regions (octopine synthase termination region and nopaline synthase termination region) are available from the Ti plasmid of Agrobacterium tumefaciens. See also Guerineau et al. (1991) Mol. Gen. Genet. 262:141-144; Proudfoot (1991) Cell 64:671-674; Sanfacon et al. (1991) Genes Dev. 5:141-149; Mogen et al. (1990) Plant Cell 2:1261-1272; Munroe et al. (1990) Gene 91:151-158; Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903; Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.

[0063] Non-limiting examples of additional regulatory signals include transcription initiation sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See, e.g., U.S. Patent Nos. 5,039,523 and 4,853,331; EPO 0480762A2; Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, edited by Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY) (hereafter referred to as "Sambrook 11"); Davis et al. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and references cited therein.

[0064] In preparing an expression cassette, various DNA fragments are manipulated to provide DNA sequences in the proper orientation and, if necessary, the proper reading frame. To this end, DNA fragments may be joined using adapters or linkers, and other manipulations may be included to provide convenient restriction sites, remove excess DNA, or remove restriction sites. To this end, in vitro mutagenesis, primer repair, restriction, annealing, and resubstitutions (e.g., transitions and transversions) may be included.

[0065] Numerous promoters can be used to practice the present invention. The promoter can be selected based on the desired result. The nucleic acid can be combined with a constitutive promoter, an inducible promoter, a growth stage-specific promoter, a cell type-specific promoter, a tissue-selective promoter, a tissue-specific promoter, or other promoter for expression in the organism of interest. See, e.g., WO 99 / 43838 and the promoters described in U.S. Patent Nos. 8,575,425; 7,790,846; 8,147,856; 8,586,832; 7,772,369; 7,534,939; 6,072,050; 5,659,026; 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; 5,608,142; and 6,177,611, which are incorporated herein by reference.

[0066] For expression in plants, constitutive promoters include the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812); rice actin (McElroy et al. (1990) Plant Cell 2:163-171); ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et al. (1992) Plant Mol. Biol. 18:675-689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588); and MAS (Velten et al. (1984) EMBO J. 3:2723-2730).

[0067] Examples of inducible promoters are the Adh1 promoter, which is inducible by hypoxic or cold stress, the Hsp70 promoter, which is inducible by heat stress, and the PPDK promoter and peptocarboxylase promoter, which are inducible by light. Chemically inducible promoters, such as the In2-2 promoter, which is induced by safeners (U.S. Patent No. 5,364,780), the Axig1 promoter, which is induced by auxin and is tapetum-specific but active in callus (PCT US01 / 22169), steroid-responsive promoters (see, e.g., Schena et al. (1991) Proc. Natl. Acad. Sci. USA 88:10421-10425 and McNellis et al. (1998) Plant J. 14(2):247-257), and promoters that can be induced and repressed by tetracycline (see, e.g., Gatz et al. (1991) Mol. Gen. Genet. 227:229-237; US Pat. Nos. 5,814,618 and 5,789,156, which are incorporated herein by reference, are also useful.

[0068] Tissue-specific or tissue-selective promoters can be used to express an expression construct in a particular tissue. In some embodiments, tissue-specific or tissue-selective promoters are active in plant tissues. Examples of promoters under developmental control in plants include promoters that selectively initiate transcription in a given tissue (e.g., leaves, roots, fruits, seeds, flowers, etc.). A "tissue-specific" promoter is a promoter that initiates transcription only in a given tissue. Tissue-specific expression, unlike constitutive gene expression, is the result of several levels of gene regulatory interaction. Therefore, promoters from homologous or closely related plant species can be selectively used to achieve efficient and reliable expression of a transgene in a particular tissue. In some embodiments, expression involves a tissue-selective promoter. A "tissue-selective" promoter is a promoter that selectively initiates transcription, but not necessarily exclusively or exclusively, in a given tissue.

[0069] In some embodiments, nucleic acid molecules encoding RGN, crRNA, and / or tracrRNA contain a cell-type-specific promoter. A "cell-type-specific" promoter is a promoter that directs expression primarily in certain cell types within one or more organs. Examples of plant cells in which a functional cell-type-specific promoter can be primarily active in plants include, for example, BELT cells, vascular cells in roots and leaves, stem cells, and stem cells. Nucleic acid molecules can also contain cell-type-selective promoters. A "cell-type-selective" promoter is a promoter that directs expression primarily in certain cell types within one or more organs, but not necessarily exclusively in a given tissue or in a given tissue. Examples of plant cells in which a functional cell-type-selective promoter can be selectively active in plants include, for example, BELT cells, vascular cells in roots and leaves, stem cells, and stem cells.

[0070] Nucleic acid molecules encoding RGN, crRNA, tracrRNA, and / or sgRNA can be operably linked to a promoter sequence recognized by a phage RNA polymerase, e.g., for in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA can be purified and used in the methods described herein. For example, the promoter sequence can be a T7 promoter, a T3 promoter, an SP6 promoter, or variations of the T7, T3, or SP6 promoter sequences. In such embodiments, the expressed protein and / or RNA can be purified and used in the genome modification methods described herein.

[0071] In some embodiments, the polynucleotide encoding RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA can also be linked to a polyadenylation signal (e.g., the SV40 polyA signal or other signals functional in plants) and / or at least one transcription termination sequence. Additionally, as described elsewhere herein, the sequence encoding RGN can also be linked to a sequence encoding at least one nuclear localization signal, and / or at least one cell entry domain, and / or at least one signal peptide capable of directing the protein to a specific subcellular location.

[0072] The polynucleotides encoding RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA can be present in one vector or multiple vectors. "Vector" refers to a polynucleotide composition for transferring, delivering, or introducing a nucleic acid into a host cell. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, and baculoviral vectors). Vectors can include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Additional information can be found in Current Protocols in Molecular Biology, Ausubel et al., John Wiley & Sons, New York, 2003, or Molecular Cloning: A Laboratory Manual, Sambrook and Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd ed., 2001.

[0073] Vectors can also contain selectable marker genes for the selection of transformed cells. Selectable marker genes are used to select transformed cells or tissues. Marker genes include those encoding antibiotic resistance (e.g., genes encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT)) and genes conferring resistance to insecticidal compounds (e.g., glufosinate ammonium, bromoxynil, imidazolinone, 2,4-dichlorophenoxyacetate (2,4-D)).

[0074] In some embodiments, an expression cassette or vector containing a sequence encoding an RGN polypeptide can further include a sequence encoding a crRNA and / or a tracrRNA, or a combination of a crRNA and a tracrRNA to generate a guide RNA. The sequence encoding the crRNA and / or the tracrRNA can be operably linked to at least one transcriptional control sequence to express the crRNA and / or the tracrRNA in an organism or host cell of interest. For example, a polynucleotide encoding the crRNA and / or the tracrRNA can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Non-limiting examples of suitable Pol III promoters include mammalian U6, U3, H1, and 7SL RNA promoters, and rice U6 and U3 promoters.

[0075] As shown, an organism of interest can be transformed with an expression construct encoding RGN, crRNA, tracrRNA, and / or sgRNA. Methods of transformation include introducing a nucleotide construct into the organism of interest. By "introducing," we mean introducing a nucleotide construct into a host cell, allowing the construct to access the interior of the host cell. The methods of the present invention do not require a specific method for introducing a nucleotide construct into a host organism; they only require that the nucleotide construct access the interior of at least one cell of the host organism. The host cell can be a eukaryotic or prokaryotic cell. In particular embodiments, the eukaryotic cell is a plant cell, a mammalian cell, or an insect cell. Methods for introducing nucleotide constructs into plant and other host cells are known in the art, and non-limiting examples include stable transformation, transient transformation, and viral-mediated methods.

[0076] These methods result in transformed organisms (e.g., plants, including whole plants, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos, and progeny thereof). Plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, pollen).

[0077] A "transgenic organism" or "transformed organism" or "stably transformed" organism or cell or tissue refers to an organism into which a polynucleotide encoding the RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA of the present invention has been incorporated or integrated. It is known that other exogenous or endogenous nucleic acid sequences or DNA fragments can also be incorporated into host cells. Agrobacterium- and biolistic-mediated transformation remain the two main approaches utilized to transform plant cells. However, host cell transformation can also be achieved by infection, transfection, microinjection, electroporation, microprojection, biolistic or particle bombardment, silica / carbon fiber, ultrasound-mediated, PEG-mediated, calcium phosphate co-precipitation, polycation DMSO technology, DEAE-dextran procedures, virus-mediated, liposome-mediated, and other methods. Virus-mediated methods for delivery of polynucleotides encoding RGN, and / or crRNA, and / or tracrRNA include retrovirus-, lentivirus-, adenovirus-, and adeno-associated virus-mediated delivery and expression, as well as the use of caulimoviruses, geminiviruses, and RNA plant viruses.

[0078] Transformation protocols, as well as protocols for introducing polypeptide or polynucleotide sequences into plants, can vary depending on the type of host cell targeted for transformation (e.g., monocotyledonous or dicotyledonous plant cells). Transformation methods are known in the art and include those described in U.S. Patent Nos. 8,575,425; 7,692,068; 8,802,934; and 7,541,517 (each of which is incorporated herein by reference). Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett. 7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4:1-12; Bates, G.W. (1999) Methods in Molecular Biology 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology 42:575-606; Christou, P. (1992) The Plant See also Journal 2:275-281; ​​Christou, P. (1995) Euphytica 85:13-27; Tzfira et al. (2004) TRENDS in Genetics 20:375-383; Yao et al. (2006) Journal of Experimental Botany 57:3737-3746; Zupan and Zambryski (1995) Plant Physiology 107:1041-1047.

[0079] Transformation can result in stable or transient integration of a nucleic acid into a cell. "Stable transformation" means that a nucleotide construct introduced into a host cell is integrated into the genome of the host cell and can be inherited by its progeny. "Transient transformation" means that a polynucleotide is introduced into a host cell but is not integrated into the genome of the host cell.

[0080] Methods for transforming chloroplasts are known in the art. See, e.g., Svab et al. (1990) Proc. Nail. Acad. Sci. USA 87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J. 12:601-606. This method relies on particle gun delivery of DNA containing a selectable marker and targeting that DNA into the plastid genome through homologous recombination. Additionally, plastid transformation can be achieved by cross-activating silent plastid-borne transgenes through tissue-selective expression of a nuclear-encoded, plastid-derived RNA polymerase. Such a system is described in McBride et al. (1994) Proc. Natl. Acad. Sci. USA 91:7301-7305.

[0081] The transformed cells can be introduced into transgenic organisms (e.g., plants) according to conventional methods and propagated. See, e.g., McCormick et al. (1986) Plant Cell Reports 5:81-84. The plants are then propagated and pollinated with the same or a different transformed line, and the resulting hybrids are identified that constitutively express the desired phenotypic trait. After propagation for two or more generations to confirm that the expression of the desired phenotypic trait is stably maintained and inherited, seeds are harvested to confirm that the desired phenotypic trait has been achieved. In this way, the present invention provides transformed seeds (also referred to as "transgenic seeds") in which the nucleotide construct of the present invention (e.g., the expression cassette of the present invention) has been stably integrated into the genome.

[0082] Alternatively, transformed cells can be introduced into an organism, and these cells could be derived from an organism, in which case the cells are transformed by an ex vivo approach.

[0083] The sequences provided herein can be used to transform any plant species, non-limiting examples of which include monocotyledons and dicotyledons. Non-limiting examples of plants of interest include corn, sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley, rapeseed, Brassica species, alfalfa, rye, millet, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus fruits, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew nuts, macadamia nuts, almonds, oats, vegetables, ornamentals, and conifers.

[0084] Non-limiting examples of vegetables include tomatoes, lettuce, green peas, lima beans, peas, and members of the Cucumis genus, such as cucumber, cantaloupe, and muskmelon. Non-limiting examples of ornamental plants include azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, poinsettias, and chrysanthemums. Preferably, the plants of the present invention are cultivated crops, such as corn, sorghum, wheat, sunflowers, tomatoes, brassicas, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, and the like.

[0085] As used herein, the term plant includes plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, a population of plants, intact plant cells within a plant, and plant parts (e.g., embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, kernels, ears, cobs, husks, stems, roots, root caps, anthers, etc.). Grain refers to mature seeds produced by agricultural growers for purposes other than propagation or reproduction. Progeny, variants, and mutants of regenerated plants are also within the scope of the present invention, provided they contain the introduced polynucleotide. Additionally, processed plant products or by-products (e.g., soybean meal) containing the sequences disclosed herein are also provided.

[0086] Polynucleotides encoding RGN, and / or crRNA, and / or tracrRNA can also be used to transform any prokaryotic species, including, but not limited to, archaea and bacteria (e.g., Bacillus spp., Klebsiella spp., Streptomyces spp., Rhizobium spp., Escherichia coli spp., Pseudomonas spp., Salmonella spp., Shigella spp., Vibrio spp., Yersinia spp., Mycoplasma spp., Agrobacterium, and Lactobacillus spp.).

[0087] Polynucleotides encoding RGN, and / or crRNA, and / or tracrRNA can be used to transform any eukaryotic species, including, but not limited to, animals (e.g., mammals, insects, fish, birds, reptiles), fungi, amoebas, algae, and yeast.

[0088] Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to administer nucleic acids encoding components of the CRISPR system to cells in culture or host organisms. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles (e.g., liposomes). Viral vector delivery systems include DNA and RNA viruses, which have episomal or integrated genomes after delivery into cells. For reviews of gene therapy, see Anderson, Science 256:808-813 (1992); Nabel and Feigner, TIBTECH 11:211-217 (1993); Mitani and Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer and Perricaudet, British Medical See Bulletin 51(1):31-44 (1995); Haddada et al. (1995) in Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds.); Yu et al., Gene Therapy 1:13-26 (1994).

[0089] Non-viral methods for nucleic acid delivery include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid:nucleic acid complexes, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those described by Felgner in WO 91 / 17424 and WO 91 / 16024. Delivery can be to cells (e.g., in vitro or ex vivo administration) or to target tissues (e.g., in vivo administration). The preparation of lipid:nucleic acid complexes, including the preparation of targeted liposomes (such as immunolipid complexes), is well known to those of skill in the art (e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994)); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); see U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0090] The use of RNA or DNA virus-based systems for nucleic acid delivery takes advantage of highly evolved methods for targeting viruses to specific cells in the body and transporting the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo). Alternatively, viral vectors can be used to treat cells in vitro, and the resulting modified cells can then be administered to patients (ex vivo). Traditional virus-based systems include retroviral, lentiviral, adenoviral, adeno-associated, and herpes simplex virus vectors for gene delivery. Integration into the host genome is possible using retroviral, lentiviral, and adeno-associated virus gene delivery methods, often resulting in long-term expression of the inserted transgene. In addition, high gene transfer efficiencies have been observed in many different cell types and target tissues.

[0091] The tropism of retroviruses can be altered by incorporating foreign envelope proteins to expand the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene delivery system will likely depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats and have the capacity to package foreign sequences up to 6–10 kb. Minimal cis-acting LTRs are sufficient for vector replication and packaging, and these vectors can be used to integrate therapeutic genes into target cells and permanently express the transgene. Widely used retroviral vectors include vectors based on murine leukemia virus (MuLV), gibbon leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g., Buchscher et al., J. Viral. 66:2731-2739 (1992); Johann et al., J. Viral. 66:1635-1640 (1992); Sommnerfelt et al., J. Viral. 176:58-59 (1990); Wilson et al., J. Viral. 63:2374-2378 (1989); Miller et al., J. Viral. 65:2220-2224 (1991); PCT / US94 / 05700).

[0092] For applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors are capable of very high gene transfer efficiency in many cell types and do not require cell division. High titers and high expression levels have been obtained using such vectors. Large quantities of these vectors can be produced in a relatively simple system. Adeno-associated virus ("AAV") vectors can also be used to deliver target nucleic acids into cells, for example, for in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy (see, e.g., West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, J. Clin. Invest. 94:1351 (1994)). The construction of recombinant AAV vectors has been described in numerous publications, including U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin et al., Mol. Cell. Biol. 4:2072-2081 (1984); Hermonat and Muzyczka, PNAS 81:6466-6470 (1984); Samulski et al., J. Viral. 63:3822-3828 (1989). Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells (which package adenovirus) and ΨJ2 or PA317 cells (which package retrovirus).

[0093] Viral vectors used in gene therapy are usually produced by generating cell lines that package nucleic acid vectors into viral particles. The vectors typically contain minimal viral sequences necessary for packaging and subsequent integration into the host; other viral sequences are replaced by an expression cassette for the polynucleotide to be expressed. Missing viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences from the AAV genome necessary for packaging and integration into the host genome. Viral DNA is packaged into cell lines containing helper plasmids encoding other AAV genes (i.e., rep and cap) but lacking ITR sequences.

[0094] Cell lines can also be infected with adenovirus as a helper. The helper virus promoter drives replication of the AAV vector and expression of AAV genes from the helper plasmid. The helper plasmid is not packaged in large quantities due to the lack of ITR sequences. Contamination with adenovirus can be reduced, for example, by heat treatment, to which adenovirus is more sensitive than AAV. Additional methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, U.S. Patent Application Publication No. 2003 / 0087817 (incorporated herein by reference).

[0095] In some embodiments, one or more vectors described herein are transiently or non-transiently transfected into host cells. In some embodiments, cells are transfected in a manner similar to that which occurs naturally in a subject. In some embodiments, the transfected cells are harvested from a subject. In some embodiments, the cells are derived from cells (e.g., cell lines) harvested from a subject. A variety of cell lines for tissue culture are known in the art. Non-limiting examples of cell lines include: C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFl, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182 , A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, T IB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2 780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC Cell lines that can be used include, but are not limited to, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic variants thereof. Cell lines can be used from a variety of sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC), Manassas, VA).

[0096] In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines containing one or more vector-derived sequences.In some embodiments, cells transiently transfected with the components of the CRISPR system described herein (by transient transfection of one or more vectors or transfection of RNA) and modified through the activity of CRISPR complexes are used to establish new cell lines that contain modifications but completely lack other exogenous sequences.In some embodiments, cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines derived from such cells, are used to evaluate one or more test compounds.

[0097] In some embodiments, one or more of the vectors described herein are used to generate non-human transgenic animals or transgenic plants. In some embodiments, the transgenic animals are mammals (such as mice, rats, rabbits, etc.).

[0098] IV. Polypeptide and Polynucleotide Variants and Fragments

[0099] The present disclosure provides active variants and fragments of naturally occurring (i.e., wild-type) RNA-guided nucleases (the amino acid sequences of which are set forth in any of SEQ ID NOS: 1, 11, 19, 27, 36, 45, and 54), as well as active variants and fragments of naturally occurring CRISPR repeats (e.g., sequences set forth in any of SEQ ID NOS: 2, 12, 20, 28, 37, 46, and 55), and naturally occurring tracrRNA (e.g., sequences set forth in any of SEQ ID NOS: 3, 13, 21, 29, 38, 47, and 56), and polynucleotides encoding the same.

[0100] The activity of a variant or fragment may be altered compared to the polynucleotide or polypeptide of interest, but variants and fragments should retain the functionality of the polynucleotide or polypeptide of interest, e.g., a variant or fragment may have increased activity, decreased activity, a different spectrum of activity, or some other altered activity compared to the polynucleotide or polypeptide of interest.

[0101] Fragments and variants of native RGN polypeptides (such as those disclosed herein) will retain sequence-specific RNA-guided DNA binding activity. In particular embodiments, fragments and variants of native RGN polypeptides (such as those disclosed herein) will retain nuclease activity (single- or double-stranded).

[0102] Fragments and variants of naturally occurring CRISPR repeats (such as those disclosed herein), when part of a guide RNA (including tracrRNA), will retain the ability to bind (complex with) an RNA-guided nuclease in a sequence-specific manner and guide it to target a target nucleotide sequence.

[0103] Fragments and variants of naturally occurring tracrRNA (such as those disclosed herein), when part of a guide RNA (including a CRISPR RNA), will retain the ability to bind (complex with) an RNA-guided nuclease in a sequence-specific manner and guide it to target a target nucleotide sequence.

[0104] The term "fragment" refers to a portion of a polynucleotide or polypeptide sequence of the present invention. A "fragment" or "biologically active portion" includes a polynucleotide comprising a sufficient number of consecutive nucleotides to retain biological activity (i.e., when included in a guide RNA, to bind to RGN in a sequence-specific manner and target the RGN to a target nucleotide sequence). A "fragment" or "biologically active portion" includes a polypeptide comprising a sufficient number of consecutive amino acid residues to retain biological activity (i.e., when included in a guide RNA, to bind to RGN in a sequence-specific manner and target the RGN to a target nucleotide sequence). Fragments of RGN proteins include those that are shorter than the full-length sequence due to the use of alternative downstream start sites. A biologically active portion of an RGN protein can be a polypeptide containing, for example, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more consecutive amino acid residues of any of SEQ ID NOS: 1, 11, 19, 27, 36, 45, and 54. Such biologically active portions can be prepared by recombinant techniques and then assayed for sequence-specific RNA-guided DNA binding activity. A biologically active fragment of a CRISPR repeat sequence can contain at least 8 consecutive amino acids of any of SEQ ID NOS: 2, 12, 20, 28, 37, 46, and 55. A biologically active portion of a CRISPR repeat sequence can be, for example, a polynucleotide comprising 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleotides of any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, or 55.A biologically active portion of a tracrRNA can be a polynucleotide that includes, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides of any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, and 56.

[0105] Generally, "variant" refers to a substantially similar sequence. With respect to polynucleotides, variants include deletions and / or additions of one or more nucleotides at one or more internal sites within the naturally occurring polynucleotide and / or substitutions of one or more nucleotides at one or more internal sites within the naturally occurring polynucleotide. As used herein, a "naturally occurring" or "wild-type" polynucleotide or polypeptide comprises a naturally occurring nucleotide sequence or amino acid sequence, respectively. With respect to polynucleotides, conservative variants include sequences that encode the naturally occurring amino acid sequence of a gene of interest due to the degeneracy of the genetic code. Such naturally occurring allelic variants can be identified using well-known techniques of molecular biology, such as polymerase chain reaction (PCR) and hybridization techniques, as outlined below. Variant polynucleotides also include polynucleotides of synthetic origin that still encode a polypeptide or polynucleotide of interest (e.g., polynucleotides generated by site-directed mutagenesis). Generally, a variant of a particular polynucleotide disclosed herein will have at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the particular polynucleotide when examined using sequence alignment programs and parameters described elsewhere herein.

[0106] Variants of a particular polynucleotide (i.e., a reference polynucleotide) disclosed herein can also be assessed by comparing the percent sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The percent sequence identity between any two polypeptides can be calculated using sequence alignment programs and parameters described elsewhere herein. When any pair of polynucleotides disclosed herein is assessed by comparing the percent sequence identity shared by the two polypeptides they encode, the percent sequence identity between the two encoded polypeptides is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more.

[0107] In particular embodiments, a polynucleotide of the present disclosure encodes an RNA-guided nuclease polypeptide comprising an amino acid sequence that is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the amino acid sequence of any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54.

[0108] Biologically active variants of the RGN polypeptides of the present invention may differ by about 1 to 15 amino acid residues, or by about 1 to 10 amino acid residues, or by about 6 to 10 amino acid residues, or by as few as 5, or 4, or 3, or 2, or 1 amino acid residue. In particular embodiments, the polypeptides may comprise N- or C-terminal truncations, which may include deletions of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more amino acids from the N- or C-terminus of the polypeptide.

[0109] In some embodiments, a polynucleotide of the disclosure comprises or encodes a CRISPR repeat comprising a nucleotide sequence at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to a nucleotide sequence set forth as any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.

[0110] The polynucleotides of the present disclosure can comprise or encode a tracrRNA comprising a nucleotide sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the nucleotide sequence set forth as any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, and 56.

[0111] Biologically active variants of CRISPR repeats or tracrRNA of the invention can differ by about 1-15 amino acid residues, or about 1-10 amino acid residues, or about 6-10 amino acid residues, or as few as 5, or 4, or 3, or 2, or 1 amino acid residue. In particular embodiments, polynucleotides can include 5' or 3' truncations, which can include at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more deletions from the 5' or 3' end of the polynucleotide.

[0112] It is understood that the RGN polypeptides, CRISPR repeats, and tracrRNA provided herein can be modified to generate variant proteins and variant polynucleotides. Changes can be designed and introduced using site-directed mutagenesis techniques. Alternatively, naturally occurring, but as yet unknown or unidentified, polynucleotides and / or polypeptides related in structure and / or function to the sequences disclosed herein can be identified and fall within the scope of the present invention. Conservative amino acid substitutions can be made in non-conserved regions that do not alter the function of the RGN protein. Alternatively, modifications can be made that improve the activity of RGN.

[0113] Variant polynucleotides and variant proteins also encompass sequences and proteins resulting from mutagenesis and recombination procedures, such as DNA shuffling. Such procedures can be used to manipulate one or more of the different RGN proteins disclosed herein (e.g., SEQ ID NOS: 1, 11, 19, 27, 36, 45, and 54) to generate novel RGN proteins with desired properties. In this manner, libraries of recombinant polynucleotides are generated from a population of polynucleotides with related sequences, i.e., polynucleotides with substantially identical sequences and containing sequence regions capable of homologous recombination in vitro or in vivo. For example, using this approach, sequence motifs encoding domains of interest can be shuffled between the RGN sequences presented herein and other known RGN genes to generate sequences with improved properties of interest (e.g., enzymes such as K). mNovel genes encoding proteins (with increased activity) can be obtained. Strategies for such DNA shuffling are known in the art. See, e.g., Stemmer (1994) Proc. Natl. Acad. Sci. USA 91:10747-10751; Stemmer (1994) Nature 370:389-391; Crameri et al. (1997) Nature Biotech. 15:436-438; Moore et al. (1997) J. Mol. Biol. 272:336-347; Zhang et al. (1997) Proc. Natl. Acad. Sci. USA 94:4504-4509; Crameri et al. (1998) Nature 391:288-291; U.S. Patent Nos. 5,605,793 and 5,837,458. "Shuffled" nucleic acids are nucleic acids generated by a shuffling procedure (e.g., any of the shuffling procedures described herein). Shuffled nucleic acids are generated by recombining (physically or virtually) two or more nucleic acids (or character strings), e.g., in an artificial, and optionally recursive, manner. Generally, the shuffling process utilizes one or more screening steps to identify nucleic acids of interest. This screening step can be performed before or after any recombination step. In some (but not all) shuffling embodiments, it is desirable to perform multiple rounds of recombination before selection to increase the diversity of the pool to be screened. The entire recombination and selection process is repeated, optionally in a recursive manner. Depending on the context, shuffling can refer to the entire recombination and selection process, or to just the recombination portion of the entire process.

[0114] As used herein, "sequence identity" or "matching" in the context of two polynucleotide or polypeptide sequences refers to residues that are the same in the two sequences when aligned for maximum correspondence over a specified comparison window. When percentage sequence identity is used in the context of proteins, it is recognized that residue positions that are not identical often represent differences in conservative amino acid substitutions, where an amino acid residue is replaced with another amino acid residue of similar chemical properties (e.g., charge or hydrophobicity), thereby not altering the functional properties of the molecule. When sequences differ in conservative substitutions, the percentage sequence identity can be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ in such conservative substitutions are said to have "sequence similarity" or "similarity." Means for making this adjustment are well known to those of skill in the art. Typically, this involves scoring conservative substitutions as partial rather than complete mismatches, thereby increasing the percentage of sequence identity. Thus, for example, identical amino acids are assigned a score of 1, non-conserved substitutions are assigned a score of 0, and conservative substitutions are assigned a score between 0 and 1. Conservative substitution scores are calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, CA).

[0115] As used herein, "percentage of sequence identity" refers to the value obtained by comparing two optimally aligned sequences over a comparison window. In this case, the portion of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to the reference sequence (which does not contain additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions with the same nucleic acid base or amino acid residue in both sequences to obtain the number of identical positions, dividing this number of identical positions by the total number of positions within the comparison window and multiplying by 100 to obtain the percentage of sequence identity.

[0116] Unless otherwise specified, the sequence identity / similarity values ​​presented herein refer to values ​​obtained using GAP Version 10 (using the following parameters: for nucleotide sequences, % identity and % similarity obtained using a GAP Weight of 50, a Length Weight of 3, and the nwsgapdna.cmp scoring matrix; for amino acid sequences, % identity and % similarity obtained using a GAP Weight of 8, a Length Weight of 2, and the BLOSUM62 scoring matrix), or any equivalent program. By "equivalent program" is meant any sequence comparison program that produces alignments with the same nucleotide or amino acid residue matches and the same percentage of sequence identity for any two sequences in question when compared to corresponding alignments produced by GAP Version 10.

[0117] Two sequences are "optimally aligned" when they are aligned using a predetermined amino acid substitution matrix (e.g., BLOSUM62), gap existence penalties, and gap extension penalties to achieve the highest possible score for the pair of sequences for the purposes of determining a similarity score. Amino acid substitution matrices and their use to quantify similarity between two sequences are well known in the art and are described, for example, in Dayhoff et al. (1978) "Models of Evolutionary Change in Proteins," pp. 345-352, in Atlas of Protein Sequence and Structure, Vol. 5, Supplement 3 (M.O. Dayhoff, ed.), Natl. Biomed. Res. Found., Washington, D.C.; and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA 89:10915-10919. The BLOSUM62 matrix is ​​often used as the default scoring substitution matrix in sequence alignment protocols. A gap existence penalty is imposed for introducing a single amino acid gap into one of the aligned sequences, and a gap extension penalty is imposed for introducing each additional empty amino acid position into an existing gap. The alignment is defined by the amino acid positions in each sequence where the alignment begins and ends, and optionally, one or more gaps are inserted in one or both sequences to achieve the highest possible score. While optimal alignment and scoring can be achieved manually, the process can be facilitated by computer-implemented alignment algorithms (e.g., Gapped BLAST 2.0, described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402, and publicly available at the National Center for Biotechnology Information Website (www.ncbi.nlm.nih.gov)).Optimal alignments, including multiple alignments, can be prepared using, for example, PSI-BLAST (described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402 and available at www.ncbi.nlm.nih.gov).

[0118] For an amino acid sequence optimally aligned with a reference sequence, an amino acid residue "corresponds" to the position in the reference sequence that pairs with it in the alignment. The "position" is represented by the number of each amino acid, which is sequentially identified based on its position relative to the N-terminus in the reference sequence. Because of deletions, insertions, truncations, fusions, and the like that must be considered when determining optimal alignment, the number of an amino acid residue in a test sequence, determined by simple counting from the N-terminus, is generally not necessarily the same as the number of the corresponding position in the reference sequence. For example, if there is a deletion in an aligned test sequence, there is no amino acid in the reference sequence that corresponds to the deletion site. If there is an insertion in an aligned test sequence, the insertion does not correspond to any amino acid position in the reference sequence. In the case of a truncation or fusion, there may be a stretch of amino acids in the reference sequence or aligned sequence that does not correspond to any amino acids in the corresponding sequence.

[0119] V. Antibodies

[0120] Antibodies against the RGN polypeptides of the present invention or ribonucleoproteins comprising the RGN polypeptides (including those having the amino acid sequences set forth in SEQ ID NOS: 1, 11, 19, 27, 36, 45, and 54, or active variants or fragments thereof) are also encompassed. Methods for producing antibodies are well known in the art (see, e.g., Harlow and Lane (1988) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY; U.S. Patent No. 4,196,265). These antibodies can be used in kits for detecting and isolating RGN polypeptides or ribonucleoproteins. Thus, the present disclosure provides kits containing antibodies that specifically bind to the polypeptides or ribonucleoproteins described herein (including, for example, polypeptides having the sequences of SEQ ID NOS: 1, 11, 19, 27, 36, 45, and 54).

[0121] VI. SYSTEMS AND RIBONUCLEOPROTEIN COMPLEXES FOR BINDING TO TARGET SEQUENCES OF INTEREST AND METHODS OF MAKING SAME

[0122] The present disclosure provides a system for binding to a target sequence of interest. The system includes at least one guide RNA, or a nucleotide sequence encoding the guide RNA, and at least one RNA-guided nuclease, or a nucleotide sequence encoding the guide RNA. The guide RNA hybridizes to the target sequence of interest and also forms a complex with an RGN polypeptide, thereby allowing the RGN polypeptide to bind to the target sequence. In some of these embodiments, the RGN comprises the amino acid sequence of SEQ ID NO: 1, 11, 19, 27, 36, 45, or 54, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence of SEQ ID NO: 2, 12, 20, 28, 37, 46, or 55, or an active variant or fragment thereof. In particular embodiments, the guide RNA comprises a tracrRNA comprising the nucleotide sequence of SEQ ID NO: 3, 13, 21, 29, 38, 47, or 56, or an active variant or fragment thereof. The guide RNA in this system can be a single guide RNA or a dual guide RNA. In particular embodiments, the system includes an RNA-guided nuclease heterologous to the guide RNA, where the RGN and the guide RNA do not naturally form a complex.

[0123] The system presented herein for binding to a target sequence of interest can be a ribonucleoprotein complex, which is at least one RNA molecule bound to at least one protein. The nucleoprotein complex presented herein includes at least one guide RNA as the RNA component and an RNA-guided nuclease as the protein component. Such nucleoprotein complexes can be purified from cells or organisms that naturally express an RGN polypeptide and have been engineered to express a specific guide RNA specific for a target sequence of interest. Alternatively, nucleoprotein complexes can be purified from cells or organisms that have been transformed with a polynucleotide encoding an RGN polypeptide and a guide RNA and cultured under conditions that allow expression of the RGN polypeptide and the guide RNA. Thus, methods for producing an RGN polypeptide or an RGN ribonucleoprotein complex are provided. Such methods include culturing cells containing a nucleotide sequence encoding an RGN polypeptide under conditions that allow expression of the RGN polypeptide (and, in some embodiments, the guide RNA). The RGN polypeptide or RGN ribonucleoprotein complex can then be purified from a lysate of the cultured cells.

[0124] Methods for purifying RGN polypeptides or ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., size-exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reverse-phase chromatography, immunoprecipitation). In particular methods, the RGN polypeptide is recombinantly produced and contains a purification tag to aid in purification. Non-limiting examples of purification tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tags, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 10xHis, biotin carboxyl carrier protein (BCCP), and calmodulin. Typically, tagged RGN polypeptides or RGN ribonucleoprotein complexes are purified using immobilized metal affinity chromatography. It will be appreciated that other similar methods known in the art, including other forms of chromatography and, for example, immunoprecipitation, can be used alone or in combination.

[0125] An "isolated" or "purified" polypeptide, or biologically active portion thereof, is substantially or essentially free from components that normally accompany or interact with the polypeptide when found in its natural environment. Thus, an isolated or purified polypeptide is substantially free of other cellular material or cell culture medium when produced by recombinant techniques, and substantially free of chemical precursors or other compounds when chemically synthesized. Proteins that are substantially free of cellular material include protein preparations that contain less than about 30%, or less than 20%, or less than 10%, or less than 5%, or less than 1% (by dry weight) of contaminating proteins. When the proteins of the invention, or biologically active portions thereof, are recombinantly produced, optimally, the culture medium contains less than about 30%, or less than 20%, or less than 10%, or less than 5%, or less than 1% of chemical precursors or compound materials without the protein of interest.

[0126] A particular method presented herein for binding to and / or cleaving a target sequence of interest involves the use of an in vitro assembled RGN ribonucleoprotein complex. In vitro assembly of the RGN ribonucleoprotein complex can be carried out using any method known in the art, in which an RGN polypeptide is contacted with a guide RNA under conditions that allow the RGN polypeptide to bind to the guide RNA. As used herein, "contact," "contacting," or "contacted" refers to bringing together components of a desired reaction and placing them under conditions suitable for carrying out the desired reaction. RGN polypeptides can be produced through in vitro translation or chemical synthesis and purified from biological samples, cell lysates, or culture media. Guide RNAs can be produced through in vitro transcription or chemical synthesis and purified from biological samples, cell lysates, or culture media. RGN ribonucleoprotein complexes can be assembled in vitro by contacting the RGN polypeptide and guide RNA in solution (e.g., a buffered saline solution).

[0127] VII. METHODS OF BINDING TO, CLEAVING, AND MODIFYING TARGET SEQUENCES

[0128] The present disclosure provides methods for binding to, and / or cleaving, and / or modifying a target nucleotide of interest. These methods include delivering a system comprising at least one guide RNA or a polynucleotide encoding the same and at least one RGN polypeptide or a polynucleotide encoding the same to a target sequence or to a cell, organelle, or embryo containing the target sequence. In some of these embodiments, the RGN comprises the amino acid sequence of SEQ ID NO: 1, 11, 19, 27, 36, 45, or 54, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat comprising the nucleotide sequence of SEQ ID NO: 2, 12, 20, 28, 37, 46, or 55, or an active variant or fragment thereof. In particular embodiments, the guide RNA comprises a tracrRNA comprising the nucleotide sequence of SEQ ID NO: 3, 13, 21, 29, 38, 47, or 56, or an active variant or fragment thereof. The guide RNA of the system can be a single guide RNA or a dual guide RNA. The RGN of the system can be a nuclease-dead RGN, can have nickase activity, or can be a fusion polypeptide. In some embodiments, the fusion polypeptide comprises a base-editing polypeptide (e.g., cytidine deaminase or adenosine deaminase). In particular embodiments, the RGN and / or guide RNA are heterologous to the cell, organelle, or embryo into which the RGN and / or guide RNA (or a polynucleotide encoding the RGN and / or guide RNA) are introduced.

[0129] In embodiments where the method includes delivery of a polynucleotide encoding a guide RNA and / or an RGN polypeptide, the cell or embryo can then be cultured under conditions in which the guide RNA and / or RGN polypeptide is expressed. In various embodiments, the method includes contacting the target sequence with an RGN ribonucleoprotein complex. The RGN ribonucleoprotein complex can include RGN that is nuclease-dead or has nickase activity. In some embodiments, the RGN ribonucleoprotein complex is a fusion polypeptide that includes a base-editing polypeptide. In some embodiments, the method includes introducing the RGN ribonucleoprotein complex into a cell, organelle, or embryo that contains the target sequence. The RGN ribonucleoprotein complex can be purified from a biological sample, recombinantly produced and then purified, or assembled in vitro as described herein. In embodiments where the in vitro assembled RGN ribonucleoprotein complex is contacted with the target sequence, cell, organelle, or embryo, the method can further include contacting the complex with the target sequence, cell, organelle, or embryo after assembly in vitro.

[0130] Purified or in vitro assembled RGN ribonucleoprotein complexes can be introduced into cells, organelles, or embryos using any method known in the art, including, but not limited to, electroporation, or polynucleotides encoding or comprising RGN polypeptides and / or guide RNAs can be introduced using any method known in the art, such as electroporation.

[0131] Upon delivery or contact with a target sequence or a cell, organelle, or embryo containing the target sequence, the guide RNA causes RGN to bind to the target sequence in a sequence-specific manner. In embodiments in which RGN has nuclease activity, the RGN polypeptide cleaves the target sequence of interest upon binding to the target sequence. The target sequence is then modified through endogenous repair mechanisms (e.g., non-homologous end joining or homology-directed repair using the provided donor polynucleotide).

[0132] Methods for measuring binding of RGN polypeptides to target sequences are known in the art and include chromatin immunoprecipitation assays, gel mobility shift assays, DNA pull-down assays, reporter assays, and microplate capture and detection assays. Similarly, methods for measuring cleavage or modification of target sequences are known in the art and include in vitro or in vivo cleavage assays. These cleavage assays utilize PCR sequencing or gel electrophoresis to confirm cleavage, with or without appropriate target sequence labeling (e.g., radioisotopes, fluorescent materials) to facilitate detection of degradation products. Alternatively, a cleavage-triggered exponential amplification reaction (NTEXPAR) assay can be used (see, e.g., Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage can be assessed using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).

[0133] In some embodiments, the methods involve the use of an RGN complexed with two or more guide RNAs, which can target different regions of a single gene or multiple genes.

[0134] In embodiments where a donor polynucleotide is not provided, the double-strand break introduced by the RGN polypeptide can be repaired by the non-homologous end joining (NHEJ) repair process. Due to the error-prone nature of NHEJ, repair of the double-strand break may result in modification of the target sequence. As used herein with respect to nucleic acid molecules, "modification" refers to a change in the nucleotide sequence of the nucleic acid molecule, which may involve the deletion, insertion, or substitution of one or more nucleotides, or a combination thereof. Modification of the target sequence may result in the expression of an altered protein product or the inactivation of the coding sequence.

[0135] In embodiments in which a donor polynucleotide is present, the donor sequence in the donor polypeptide can be integrated into or exchanged with the target nucleotide sequence during repair of the introduced double-strand break, resulting in the introduction of an exogenous donor sequence. Thus, the donor polynucleotide contains the donor sequence desired to be introduced into the target sequence of interest. In some embodiments, the donor sequence alters the original target nucleotide sequence, so that the newly integrated donor sequence is not recognized by RGN and is not cleaved by RGN. Donor sequence integration can be enhanced by including flanking sequences in the donor polynucleotide that are substantially identical in sequence to sequences adjacent to the target nucleotide sequence, enabling the homology-directed repair process. In embodiments in which an RGN polypeptide introduces sticky ends at the double-strand break, the donor polynucleotide can contain the donor sequence flanked by compatible overhangs, allowing the donor sequence to be directly ligated to the cleaved target nucleotide sequence, including the overhang, by the non-homologous recombination repair process during double-strand break repair.

[0136] In embodiments where the method involves the use of an RGN that is a nickase (i.e., capable of cleaving only a single strand of a double-stranded polynucleotide), the method can include introducing two RGN nickases that target the same or overlapping target sequences and cleave different strands of the polynucleotide. For example, one RGN nickase can be introduced that cleaves only the plus (+) strand of a double-stranded polynucleotide, along with a second RGN nickase that cleaves only the minus (-) strand of a double-stranded polynucleotide.

[0137] In various embodiments, methods for binding to and detecting a target nucleotide sequence are provided, comprising introducing into a cell, organelle, or embryo at least one guide RNA or a polynucleotide encoding the same and at least one RGN polypeptide or a polynucleotide encoding the same, and expressing the guide RNA and / or the RGN polypeptide (if a coding sequence is introduced; the RGN polypeptide is nuclease-dead RGN and further comprises a detectable label), and further comprising detecting the detectable label. The detectable label can be fused to RGN as a fusion protein (e.g., a fluorescent protein). Alternatively, the detectable label can be a small molecule complexed with or incorporated into the RGN polypeptide that can be detected visually or by simple means.

[0138] Also provided herein are methods for altering the expression of a target sequence of interest or the expression of a gene of interest under the control of a target sequence. The methods include introducing into a cell, organelle, or embryo at least one guide RNA or a polynucleotide encoding the same and at least one RGN polypeptide or a polynucleotide encoding the same, and expressing the guide RNA and / or the RGN polypeptide (if a coding sequence is introduced; the RGN polypeptide is a nuclease-dead RGN). In some of these embodiments, the nuclease-dead RGN is a fusion protein containing an expression modulator domain (i.e., an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repression domain) as described herein.

[0139] The present disclosure also provides methods for binding to and / or modifying a target sequence of interest, which methods include delivering a system comprising at least one guide RNA or a polynucleotide encoding the same, and a fusion polypeptide containing an RGN of the present invention and a base-editing polypeptide (e.g., cytidine deaminase or adenosine deaminase), or a polypeptide encoding the fusion polypeptide, to the target sequence or to a cell, organelle, or embryo containing the target sequence.

[0140] Those skilled in the art will appreciate that any of the methods disclosed herein can be used to target a single target sequence or multiple target sequences. For example, this method can involve using a single RGN polypeptide in combination with multiple different guide RNAs to target multiple different sequences within a single gene and / or multiple genes. The present invention also encompasses methods of introducing multiple different guide RNAs in combination with multiple different RGN polypeptides. These guide RNAs and these guide RNA / RGN polypeptide systems can target multiple different sequences within a single gene and / or multiple genes.

[0141] In one aspect, the present invention provides kits comprising any one or more of the elements disclosed in the methods and compositions described above. In some embodiments, the kits comprise a vector system and instructions for using the kit. In some embodiments, the vector system comprises: (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting a guide sequence upstream of the tracr mate sequence (the guide sequence, upon expression, directs a CRISPR complex to a target sequence in a eukaryotic cell in a sequence-specific manner; the CRISPR complex comprises a CRISPR enzyme complexed with (1) the guide sequence hybridizing to the target sequence and (2) the tracr mate sequence hybridizing to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme-coding sequence encoding a CRISPR enzyme comprising a nuclear localization sequence. The elements can be provided individually or in combination and in any suitable container (e.g., vial, bottle, test tube, etc.).

[0142] In some embodiments, the kit includes instructions in one or more languages. In some embodiments, the kit includes one or more reagents for use in a method utilizing one or more elements described herein. The reagents can be provided in any suitable container. For example, the kit can provide one or more reaction or storage buffers. The reagents can be provided in a form that is ready for use in a particular assay or can be provided in a form (e.g., a concentrated solution or lyophilized form) that requires the addition of one or more other components prior to use. The buffer can be any buffer, non-limiting examples of which include sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10.

[0143] In some embodiments, the kit includes one or more oligonucleotides corresponding to a guide sequence for insertion into a vector to operably link the guide sequence and regulatory elements. In some embodiments, the kit includes a homologous recombination template polynucleotide. In one aspect, the invention provides methods for using one or more elements of a CRISPR system. The CRISPR complexes of the invention provide an effective means for modifying target polynucleotides. The CRISPR complexes of the invention have diverse utilities, including modifying (e.g., deleting, inserting, translocating, inactivating, or activating) target polynucleotides in many cell types. As such, the CRISPR complexes of the invention have broad applications in, for example, gene therapy, drug screening, disease diagnosis, and prognosis. An exemplary CRISPR complex includes a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within a target polynucleotide.

[0144] VIII. Target Polynucleotide

[0145] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. The method can be performed in vivo, ex vivo, or in vitro. In some embodiments, the method involves sampling a cell or a population of cells from a human or non-human animal or plant (including microalgae) and modifying the cells. Culturing can be performed ex vivo at any stage. The cells can even be reintroduced into the non-human animal or plant (including microalgae).

[0146] Plant breeders exploit natural variability to combine the most useful genes in search of desirable qualities (e.g., yield, quality, uniformity, hardiness, pathogen resistance, etc.). These desirable qualities include growth, photoperiod preference, temperature requirements, flowering date or reproduction / emergence, fatty acid content, insect resistance, disease resistance, nematode resistance, fungus resistance, herbicide resistance, and tolerance to various environmental factors (drought, heat, humidity, cold, wind, and adverse soil conditions, including high salinity). Sources of these useful genes include native or exotic species, heirloom species, wild plant relatives, and induced mutations (e.g., treating plant material with mutagens). The present invention provides plant breeders with new tools for inducing mutations. Thus, those skilled in the art can analyze genomes to search for useful genes and then use the present invention to induce increases in useful genes in varieties with desired characteristics or traits more precisely than previous mutagen methods, thereby accelerating and improving plant breeding programs.

[0147] The target polynucleotide for the RGN system can be any polynucleotide, endogenous or exogenous to a eukaryotic cell. For example, the target polynucleotide can be a polynucleotide present in the nucleus of a eukaryotic cell. The target polynucleotide can be a sequence encoding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or junk DNA). Without wishing to be bound by theory, it is believed that the target sequence should be associated with a PAM (protospacer adjacent motif), a short sequence recognized by the CRISPR complex. While the exact sequence and length of the PAM vary depending on the CRISPR enzyme used, the PAM is typically a 2-5 base pair sequence adjacent to the protospacer (i.e., the target sequence).

[0148] Target polynucleotides of CRISPR complexes can include genes and polynucleotides associated with disease as well as genes and polynucleotides associated with biochemical signaling pathways. Examples of target polynucleotides include sequences associated with biochemical signaling pathways, such as genes or polynucleotides associated with biochemical signaling pathways. Examples of target polynucleotides include genes or polynucleotides associated with disease. A "disease-associated" gene or polynucleotide refers to any gene or polynucleotide that produces a transcription or translation product at an abnormal level or in an abnormal form in cells derived from disease-affected tissue compared to control tissues or cells without the disease. It can be a gene expressed at an abnormally high level or an abnormally low level, and altered expression correlates with the development and / or progression of the disease. A disease-associated gene can also refer to a gene with a mutation or genetic variation directly responsible for the origin of the disease (e.g., a causative mutation) or a gene with a mutation or genetic variation in linkage equilibrium with a gene responsible for the origin of the disease. The transcription or translation product can be known or unknown and can be at normal or abnormal levels. Examples of disease-associated genes and polynucleotides are available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, MD) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, MD).

[0149] The CRISPR system is particularly useful because it is relatively easy to target genomic sequences of interest, but questions remain about what RGN can do to address causative mutations. One approach is to create a fusion protein between RGN (preferably an inactive or nickase variant of RGN) and a base-editing enzyme (e.g., a cytidine deaminase or adenosine deaminase base editor), or the active domain of a base-editing enzyme (U.S. Patent No. 9,840,699; incorporated herein by reference). In some embodiments, this method includes (a) contacting a DNA molecule with a fusion protein comprising an RGN of the present invention and a base-editing polypeptide (e.g., a deaminase); and (b) contacting the fusion protein of (a) with a gRNA that directs the fusion protein of (a) to a target nucleotide sequence in a DNA strand; the DNA molecule is contacted with the fusion protein and the gRNA in sufficient amounts under conditions suitable for deamination of nucleotide bases. In some embodiments, the target DNA sequence comprises a sequence associated with a disease or disorder, and deamination of the nucleotide bases therein results in a sequence unassociated with the disease or disorder. In some embodiments, the target DNA sequence is present in an allele of a crop plant in which a trait of interest is expressed at a particular allele, resulting in a plant with less agricultural value. Deamination of the nucleotide bases results in an allele that improves the trait and increases the agricultural value of the plant.

[0150] In some embodiments, the DNA sequence contains a T→C point mutation or an A→G point mutation associated with a disease or disorder, and deamination of the mutated C or G base results in a sequence that is not associated with the disease or disorder. In some embodiments, deamination corrects the mutation in the sequence that is associated with the disease or disorder.

[0151] In some embodiments, the disease or disorder-associated sequence encodes a protein, and deamination introduces a stop codon into the disease or disorder-associated sequence, thereby abrogating the encoded protein. In some embodiments, the contacting is performed in vivo in a subject who is likely to have the disease or disorder, or who has the disease or disorder, or who has been diagnosed with the disease or disorder. In some embodiments, the disease or disorder is a disease associated with a point mutation or a single base mutation in the genome. In some embodiments, the disease is a genetic disease, cancer, metabolic disease, or lysosomal storage disease.

[0152] Further examples of loci that cause certain genetic diseases, particularly loci that can be readily targeted by RGN or RGN-base editor fusion proteins of the invention, can be found in Example 9 and corresponding Table 12.

[0153] Hurler syndrome

[0154] One example of a genetic disorder that may be corrected using an approach that relies on the RGN-base editor fusion proteins of the present invention is Hurler syndrome. Hurler syndrome (also known as MPS-1) results from a deficiency in α-L-iduronidase (IDUA), resulting in a lysosomal storage disorder characterized by the molecular accumulation of dermatan sulfate and heparan sulfate in lysosomes. This disease is generally an inherited genetic disorder caused by mutations in the IDUA gene, which encodes α-L-iduronidase. Common IDUA mutations are W402X and Q70X, both of which are nonsense mutations that result in premature translation termination. Such mutations are successfully addressed by precision genome editing (PGE) approaches, as restoring a single nucleotide, for example, via base editing, restores the wild-type coding sequence and allows protein expression to be controlled by the endogenous regulatory mechanisms of the gene locus. Additionally, because heterozygotes are known to be asymptomatic, PGE therapy targeting one of these mutations may be useful for many patients with this disease because only one of the mutated alleles needs to be corrected (Bunge et al. (1994) Hum. Mol. Genet. 3(6):861-866; incorporated herein by reference).

[0155] Current treatments for Hurler syndrome include enzyme replacement therapy and bone marrow transplantation (Vellodi et al. (1997) Arch. Dis. Child. 76(2):92-99; Peters et al. (1998) Blood 91(7):2601-2608; incorporated herein by reference). While enzyme replacement therapy has had a dramatic effect on the survival and quality of life of patients with Hurler syndrome, this approach requires expensive and time-consuming weekly infusions. Additional approaches include delivery of the IDUA gene on an expression vector or insertion of this gene into a highly expressed locus (such as the serum albumin locus) (U.S. Patent No. 9,956,247; incorporated herein by reference). However, these approaches do not restore the original IDUA locus to its correct coding sequence. Genome editing strategies are likely to have numerous advantages, most notably that regulation of gene expression is likely to be controlled by natural mechanisms present in healthy individuals. Additionally, base editing eliminates the need for double-stranded DNA breaks, which can lead to large-scale chromosomal rearrangements, cell death, and cancer development through the disruption of tumor suppressor mechanisms. A description of this enabling method for correcting disease-causing mutations is provided in Example 10. The described method is an example of a general strategy using the RGN-base editor fusion proteins of the present invention to target and correct several disease-causing mutations in the human genome. It will be appreciated that similar approaches can be pursued to target diseases such as those listed in Table 12. Furthermore, it will be appreciated that similar approaches can be deployed using the RGNs of the present invention to target disease-causing mutations in other species, particularly common pets and livestock. Common pets and livestock include dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, fish (including salmon), and shrimp.

[0156] Friedreich's ataxia

[0157] The RGN of the present invention may also be useful in human therapeutics where the causative mutations are more complex. Several diseases, such as Friedreich's ataxia and Huntington's disease, result from a significant increase in the number of repeated trinucleotide motifs in specific regions of genes, affecting the ability of the expressed protein to function or be expressed. Friedreich's ataxia (FRDA) is an autosomal recessive disorder that causes gradual degeneration of neural tissue within the spinal cord. Reduced levels of the mitochondrial frataxin (FXN) protein result in oxidative damage and iron deficiency at the cellular level. Reduced expression of FXN has been linked to a GAA triplet expansion within intron 1 of the FXN gene in somatic and germline cells. In FRDA patients, the number of GAA repeats is often greater than 70, sometimes even exceeding 1000 triplets (600-900 being the most common), whereas individuals without the disease have fewer than about 40 repeats (Pandolfo et al. (2012) Handbook of Clinical Neurology 103:275-294; Campuzano et al. (1996) Science 271:1423-1427; Pandolfo (2002) Adv. Exp. Med. Biol. 516:99-118; all of which are incorporated herein by reference).

[0158] The expansion of the trinucleotide repeat that causes Friedreich's ataxia (FRDA) occurs within a specific locus in the FXN gene, known as the FRDA instability region. RNA-guided nucleases (RGNs) can be used to excise the instability region in FRDA patient cells. This approach requires 1) an RGN-guide RNA sequence that can be programmed to target an allele in the human genome; and 2) a method for delivering this RGN-guide sequence. Many nucleases used for genome editing, including the commonly used Cas9 nuclease from Streptococcus pyogenes (SpCas9), are too large to be packaged into adeno-associated virus (AAV) vectors, especially considering the length of the SpCas9 gene and guide RNA, as well as the other genetic elements required for a functional expression cassette. This makes the SpCas9 approach more challenging.

[0159] The compact RNA-guided nucleases of the present invention (particularly APG07433.1 and APG08290.1) are highly suitable for excising FRDA instability regions. Each RGN requires a PAM near the FRDA instability region. Additionally, each of these RGNs can be packaged into an AAV vector along with a guide RNA. While a second vector will likely be required to package two guide RNAs, this approach is still advantageous over the vectors that would be required for larger nucleases (e.g., SpCas9, whose protein sequence may need to be split across two vectors). A description of how this disease-causing mutation can be corrected is provided in Example 11. The described method encompasses a strategy for using the RGNs of the present invention to remove regions of genomic instability. Such a strategy can be applied to other diseases and disorders with a similar genetic basis, such as Huntington's disease. Similar strategies using the RGNs of the present invention can also be applied to agriculturally or economically important diseases and disorders in non-human animals. Such non-human animals include dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, fish (including salmon), and shrimp.

[0160] Hemoglobinopathy

[0161] It is believed that the RGNs of the present invention are also capable of introducing disruptive mutations that may have beneficial effects. Genetic defects in the genes encoding hemoglobin, particularly the beta globin chain (HBB gene), may be responsible for a number of diseases known as hemoglobinopathies, including sickle cell anemia and thalassemia.

[0162] In adults, hemoglobin is a heterotetramer containing two alpha (α)-like globin chains, two beta (β)-like globin chains, and four heme groups. In adults, this α2β2 tetramer is called hemoglobin A (HbA) or adult hemoglobin. Typically, alpha and beta globin chains are synthesized in an approximately 1:1 ratio, and this ratio appears to be crucial for the stability of hemoglobin and red blood cells (RBCs). In the developing fetus, a different form of hemoglobin (fetal hemoglobin (HbF)) is produced that has a greater binding affinity for oxygen than hemoglobin A, thereby enabling oxygen to be delivered to the fetal system through the maternal bloodstream. Fetal hemoglobin also contains two alpha globin chains, but has two fetal gamma (γ) globin chains instead of the adult beta globin chains (i.e., fetal hemoglobin is α2γ2). The regulation of the switch from gamma globin to beta globin production is extremely complex and primarily involves the downregulation of gamma globin transcription and the concomitant upregulation of beta globin transcription. At approximately 30 weeks of gestation, fetal gamma globin synthesis begins to decline, while beta globin production increases. Neonatal hemoglobin is almost entirely α2β2 until approximately 10 months of age, although some HbF remains into adulthood (approximately 1–3% of total hemoglobin). As explained above, in most patients with hemoglobinopathies, the gene encoding gamma globin remains present, but its expression is relatively low due to the suppression of the normal gene that occurs around the time of birth, as explained above.

[0163] Sickle cell disease is caused by the V6E mutation in the beta-globin gene (HBs) (GAG to GTG at the DNA level), and the resulting hemoglobin is called "hemoglobin S" or "HbS." Under hypoxic conditions, HbS molecules aggregate to form fibrous precipitates. These aggregates cause abnormalities or "sickling" of RBCs, resulting in a loss of cellular flexibility. These sickled RBCs can no longer enter capillary beds, potentially leading to vaso-occlusive crises in sickle cell patients. Additionally, sickled RBCs are more fragile and prone to hemolysis than normal RBCs, ultimately leading to anemia in patients.

[0164] Treatment and management of patients with sickle cell disease is a lifelong challenge that involves antibiotic therapy, pain management, and acute blood transfusions. One approach is the use of hydroxyurea, which works in part by increasing gamma globin production. The long-term side effects of chronic hydroxyurea therapy are unknown, but the treatment may produce unwanted side effects that vary from patient to patient. Despite the increasing effectiveness of sickle cell treatments, the average life expectancy of patients is still only in their mid- to late 50s, and disease-related morbidity severely impacts patients' quality of life.

[0165] Thalassemia (alpha-thalassemia and beta-thalassemia) is also a hemoglobin-related disorder, typically involving reduced expression of globin chains. This occurs through mutations in the regulatory regions of genes or due to mutations within the globin coding sequence that result in reduced expression or levels of functional globin protein. Treatment for thalassemia typically involves blood transfusions and iron chelation therapy. Bone marrow transplants are also used to treat severely ill patients if a suitable donor can be identified, although this procedure can be risky.

[0166] One approach that has been proposed for treating both SCD and beta-thalassemia is to increase the expression of gamma globin and functionally replace abnormal adult hemoglobin A with HbF. As noted above, treatment of SCD patients with hydroxyurea appears to be successful, in part because hydroxyurea is effective in increasing gamma globin expression (DeSimone (1982) Proc Nat'l Acad Sci USA 79(14):4428-4431; Ley et al. (1982) N. Engl. J. Medicine 307:1469-1475; Ley et al. (1983) Blood 62:370-380; Constantoulakis et al. (1988) Blood 72(6):1961-1967; all of which are incorporated herein by reference). Increasing HbF expression implies the identification of genes whose products play a role in regulating gamma globin expression. One such gene is BCL11A. BCL11A encodes a zinc finger protein expressed in adult erythroid progenitor cells, and downregulation of its expression increases gamma globin expression (Sankaran et al. (2008) Science 322:1839; incorporated herein by reference). The use of inhibitory RNA targeting the BCL11A gene has been proposed (e.g., U.S. Patent Application Publication No. 2011 / 0182867; incorporated herein by reference), but this technology has several potential drawbacks. These drawbacks include the possibility of not achieving complete knockdown, potential problems with delivery of such RNA, and the need for continuous presence of the RNA, which could necessitate multiple treatments over a lifetime.

[0167] The RGNs of the present invention are used to target the BCL11A enhancer region and disrupt BCL11A expression, thereby increasing gamma globin expression. This targeted disruption can be achieved by non-homologous end joining (NHEJ). Through NHEJ, the RGNs of the present invention target specific sequences within the BCL11A enhancer region and create a double-strand break, which the cellular machinery repairs, typically concomitantly introducing a deleterious mutation. As described for other disease targets, the RGNs of the present invention have advantages over other known RGNs due to their relatively small size and the ability to package expression cassettes for the RGN and its corresponding guide RNA into a single AAV vector for in vivo delivery. A description of this approach is provided in Example 12. Similar strategies using the RGNs of the present invention can also be applied to similar diseases and disorders in both humans and agriculturally or economically important non-human animals.

[0168] IX. CELLS COMPRISING POLYNUCLEOTIDE GENETIC MODIFICATIONS

[0169] Provided herein are cells and organisms containing target sequences of interest modified using the RGN- and / or crRNA- and / or tracrRNA-mediated methods described herein. In some of these embodiments, the RGN comprises the amino acid sequence of SEQ ID NO: 1, 11, 19, 27, 36, 45, or 54, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence of SEQ ID NO: 2, 12, 20, 28, 37, 46, or 55, or an active variant or fragment thereof. In particular embodiments, the guide RNA comprises a tracrRNA comprising the nucleotide sequence of SEQ ID NO: 3, 13, 21, 29, 38, 47, or 56, or an active variant or fragment thereof. The guide RNAs in this system can be single-guide RNAs or dual-guide RNAs.

[0170] The modified cells can be eukaryotic cells (e.g., mammalian cells, plant cells, insect cells) or prokaryotic cells. Organelles and embryos containing at least one nucleotide sequence modified by the methods using RGN, crRNA, and / or tracrRNA described herein are also provided. The genetically modified cells, organisms, organelles, or embryos can be heterozygous or homozygous for the modified nucleotide sequence.

[0171] The chromosomes of a cell, organism, organelle, or embryo may be modified to result in altered expression (up- or down-regulation), inactivation, or expression of an altered protein product or integrated sequence. When the chromosomal modification results in gene inactivation or expression of a non-functional protein product, the genetically modified cell, organism, organelle, or embryo is referred to as a "knockout." The knockout phenotype may result in a deletion mutation (i.e., deletion of at least one nucleotide), or an insertion mutation (i.e., insertion of at least one nucleotide), or a nonsense mutation (i.e., substitution of at least one nucleotide to introduce a stop codon).

[0172] Alternatively, a "knock-in" can be generated as a result of a chromosomal modification of a cell, organism, organelle, or embryo, resulting in the integration of a protein-encoding nucleotide sequence into the chromosome. In some of these embodiments, the integration of the coding sequence into the chromosome inactivates the chromosomal sequence encoding the wild-type protein, while allowing the expression of the exogenously introduced protein.

[0173] In another embodiment, the chromosome is modified to produce a variant protein product. The expressed variant protein may have at least one amino acid substitution and / or at least one amino acid addition or deletion. The variant protein product encoded by the altered chromosomal sequence may exhibit altered characteristics or activities compared to the wild-type protein, including, but not limited to, altered enzymatic activity or substrate specificity.

[0174] In yet another embodiment, chromosomal modifications may result in altered protein expression patterns. By way of non-limiting example, chromosomal changes within regulatory regions controlling expression of a protein product may result in overexpression or downregulation of the protein product, or an altered tissue expression pattern, or an altered temporal expression pattern.

[0175] The article "a" or "an" is used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, "a polypeptide" means one or more polypeptides.

[0176] All publications and patent applications mentioned in this specification are indicative of the level of those skilled in the art to which this disclosure pertains. Each and every publication and patent application is herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.

[0177] Although the present invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the accompanying embodiments.

[0178] Non-limiting embodiments include the following:

[0179] 1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence that is at least 95% identical in sequence to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, or 54; When the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, it binds to the target DNA sequence in an RNA-guided sequence-specific manner; A nucleic acid molecule, wherein the polynucleotide encoding an RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide.

[0180] 2. The nucleic acid molecule of embodiment 1, wherein the RGN polypeptide is capable of cleaving the target DNA sequence when bound to the target DNA sequence.

[0181] 3. The nucleic acid molecule of embodiment 2, wherein cleavage by the RGN polypeptide results in a double-stranded break.

[0182] 4. The nucleic acid molecule of embodiment 2, wherein cleavage by the RGN polypeptide results in a single-stranded break.

[0183] 5. A nucleic acid molecule described in any one of embodiments 1 to 4, wherein the RGN polypeptide is operably fused to a base-edited polynucleotide.

[0184] 6. The nucleic acid molecule of any one of embodiments 1 to 5, wherein the RGN polypeptide comprises one or more nuclear localization signals.

[0185] 7. The nucleic acid molecule of any one of embodiments 1 to 6, wherein the RGN polypeptide is codon-optimized for expression in a eukaryotic cell.

[0186] 8. The nucleic acid molecule of any one of embodiments 1 to 7, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM).

[0187] 9. A vector comprising the nucleic acid molecule of any one of embodiments 1 to 8.

[0188] 10. The vector of embodiment 9, further comprising at least one nucleotide sequence encoding said gRNA capable of hybridizing to said target DNA sequence.

[0189] 11. The vector of embodiment 10, wherein the gRNA is a single guide RNA.

[0190] 12. The vector of embodiment 10, wherein the gRNA is a dual guide RNA.

[0191] 13. The vector of any one of embodiments 10-12, wherein the guide RNA comprises a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.

[0192] 14. The vector of any one of embodiments 10 to 13, wherein the guide RNA comprises a tracrRNA having a sequence that is at least 95% identical to any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56.

[0193] 15. A cell comprising a nucleic acid molecule according to any one of embodiments 1 to 8 or a vector according to any one of embodiments 9 to 14.

[0194] 16. A method for producing an RGN polypeptide, comprising culturing a cell according to embodiment 15 under conditions in which said RGN polypeptide is expressed.

[0195] 17. A method for producing an RGN polypeptide, comprising introducing into a cell a heterologous nucleic acid molecule comprising a nucleotide sequence encoding an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% sequence identity to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; when the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, the RGN polypeptide binds to the target DNA sequence in an RNA-guided sequence-specific manner; and culturing the cells under conditions in which the RGN polypeptide is expressed.

[0196] 18. The method of embodiment 16 or 17, further comprising purifying the RGN polypeptide.

[0197] 19. The method of embodiment 16 or 17, wherein the cell further expresses one or more guide RNAs that bind to the RGN polypeptide to form an RGN ribonucleoprotein complex.

[0198] 20. The method of embodiment 19, further comprising purifying said RGN ribonucleoprotein complex.

[0199] 21. A nucleic acid molecule comprising a polynucleotide encoding a CRISPR RNA (crRNA), wherein the crRNA comprises a spacer sequence and a CRISPR repeat sequence, and the CRISPR repeat sequence comprises a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, or 55; a) the crRNA; b) a trans-activating CRISPR RNA (tracrRNA) that hybridizes to the CRISPR repeat sequence of the crRNA; A guide RNA comprising capable of hybridizing to a target DNA sequence in a sequence-specific manner via the spacer sequence of said crRNA when bound to an RNA-guided nuclease (RGN) polypeptide; A nucleic acid molecule, wherein the polynucleotide encoding a crRNA is operably linked to a promoter that is heterologous to the polynucleotide.

[0200] 22. A vector comprising the nucleic acid molecule of embodiment 21.

[0201] 23. The vector of embodiment 22, further comprising a polynucleotide encoding the tracrRNA.

[0202] 24. The vector of embodiment 23, wherein the tracrRNA comprises a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56.

[0203] 25. The vector of embodiment 23 or 24, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as a single guide RNA.

[0204] 26. The vector of embodiment 23 or 24, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.

[0205] 27. The vector of any one of embodiments 22 to 26, further comprising a polynucleotide encoding the RGN polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, or 54.

[0206] 28. A nucleic acid molecule comprising a polynucleotide encoding a trans-activating CRISPR RNA (tracrRNA) comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, and 56; a) the tracrRNA; b) The crRNA containing the CRISPR repeat sequence and spacer sequence to which this tracrRNA hybridizes. A guide RNA comprising capable of hybridizing to a target DNA sequence in a sequence-specific manner via the spacer sequence of said crRNA when bound to an RNA-guided nuclease (RGN) polypeptide; A nucleic acid molecule, wherein the polynucleotide encoding tracrRNA is operably linked to a promoter that is heterologous to the polynucleotide.

[0207] 29. A vector comprising the nucleic acid molecule of embodiment 28.

[0208] 30. The vector of embodiment 29, further comprising a polynucleotide encoding said crRNA.

[0209] 31. The vector of embodiment 30, wherein the CRISPR repeat sequence of the crDNA comprises a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.

[0210] 32. The vector of embodiment 30 or 31, wherein the polynucleotide encoding the crDNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as a single guide RNA.

[0211] 33. The vector of embodiment 30 or 31, wherein the polynucleotide encoding the crDNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.

[0212] 34. The vector of any one of embodiments 29 to 33, further comprising a polynucleotide encoding the RGN polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, or 54.

[0213] 35. To bind to a target DNA sequence, a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence that is at least 95% identical to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, or a nucleotide sequence encoding this RGN polypeptide; A system in which each of the nucleotide sequences encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to each of the nucleotide sequences; hybridizing the one or more guide RNAs to the target DNA sequence; The one or more guide RNAs form a complex with the RGN polypeptide, thereby binding the RGN polypeptide to the target DNA sequence.

[0214] 36. The system of embodiment 35, wherein the gRNA is a single guide RNA (sgRNA).

[0215] 37. The system of embodiment 35, wherein the gRNA is a dual guide RNA.

[0216] 38. The system of any one of embodiments 35 to 37, wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, and 55.

[0217] 39. The system of any one of embodiments 35 to 38, wherein the gRNA comprises a tracrRNA comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, and 56.

[0218] 40. A system described in any one of embodiments 35 to 39, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM).

[0219] 41. A system according to any one of embodiments 35 to 40, wherein the target DNA sequence is in a cell.

[0220] 42. The system of embodiment 41, wherein the cell is a eukaryotic cell.

[0221] 43. The system of embodiment 42, wherein the eukaryotic cell is a plant cell.

[0222] 44. The system of embodiment 42, wherein the eukaryotic cells are mammalian cells.

[0223] 45. The system of embodiment 42, wherein the eukaryotic cell is an insect cell.

[0224] 46. ​​The system of embodiment 41, wherein the cell is a prokaryotic cell.

[0225] 47. The system of any one of embodiments 35 to 46, wherein, when transcribed, the one or more guide RNAs hybridize to the target DNA sequence and form a complex with the RGN polypeptide, causing cleavage of the target DNA sequence.

[0226] 48. The system of embodiment 47, wherein the cleavage results in a double-stranded break.

[0227] 49. The system of embodiment 47, wherein cleavage by the RGN polypeptide results in a single-strand break.

[0228] 50. A system described in any one of embodiments 35 to 49, wherein the RGN polypeptide is operably linked to a base-editing polypeptide.

[0229] 51. The system of any one of embodiments 35 to 50, wherein the RGN polypeptide comprises one or more nuclear localization signals.

[0230] 52. The system of any one of embodiments 35 to 51, wherein the RGN polypeptide is codon-optimized for expression in a eukaryotic cell.

[0231] 53. A system according to any one of embodiments 35 to 52, wherein the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide are located on a single vector.

[0232] 54. The system of any one of embodiments 35 to 53, further comprising one or more donor polynucleotides or one or more nucleotide sequences encoding said one or more donor polynucleotides.

[0233] 55. A method for binding to a target DNA sequence, comprising delivering a system described in any one of embodiments 35 to 54 to the target DNA sequence or to a cell containing the target DNA sequence.

[0234] 56. The method of embodiment 55, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby allowing detection of the target DNA sequence.

[0235] 57. The method of embodiment 55, wherein the guide RNA or the RGN polypeptide further comprises an expression modulator, thereby altering the expression of the target DNA sequence or the expression of a gene under the transcriptional control of the target DNA sequence.

[0236] 58. A method for cleaving or modifying a target DNA sequence, comprising delivering a system described in any one of embodiments 35 to 54 to the target DNA sequence or to a cell containing the target DNA sequence.

[0237] 59. The method of embodiment 58, wherein the modified target DNA sequence comprises inserting heterologous DNA into the target DNA sequence.

[0238] 60. The method of embodiment 58, wherein the modified target DNA sequence comprises deleting at least one nucleotide from the target DNA sequence.

[0239] 61. The method of embodiment 58, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence.

[0240] 62. A method for binding to a target DNA sequence, comprising: a) i) one or more guide RNAs capable of hybridizing to said target DNA sequence; ii) an RGN polypeptide comprising an amino acid sequence that is at least 95% identical to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; in vitro to assemble an RNA-guided nuclease (RGN) ribonucleotide complex under conditions suitable for the formation of this RGN ribonucleotide complex; b) contacting the target DNA sequence, or a cell containing the target DNA sequence, with the in vitro assembled RGN ribonucleotide complex; The method wherein the one or more guide RNAs hybridize to the target DNA sequence, thereby causing the RGN polypeptide to bind to the target DNA sequence.

[0241] 63. The method of embodiment 62, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby allowing detection of the target DNA sequence.

[0242] 64. The method of embodiment 62, wherein the guide RNA or the RGN polypeptide further comprises an expression modulator, thereby allowing alteration of expression of the target DNA sequence.

[0243] 65. A method for cleaving and / or modifying a target DNA sequence, comprising: a) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence that is at least 95% identical to any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; b) contacting the RGN of (a) with one or more guide RNAs capable of directing the RGN to the target DNA sequence; wherein the one or more guide RNAs hybridize to the target DNA sequence, thereby allowing the RGN polypeptide to bind to the target DNA sequence and result in cleavage and / or modification of the target DNA sequence.

[0244] 66. The method of embodiment 65, wherein the modified target DNA sequence comprises an insertion of heterologous DNA into the target DNA sequence.

[0245] 67. The method of embodiment 65, wherein the modified target DNA sequence comprises a deletion of at least one nucleotide from the target DNA sequence.

[0246] 68. The method of embodiment 65, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence.

[0247] 69. The method of any one of embodiments 62 to 68, wherein the gRNA is a single guide RNA (sgRNA).

[0248] 70. The method of any one of embodiments 62 to 68, wherein the gRNA is a dual guide RNA.

[0249] 71. The method of any one of embodiments 62 to 70, wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.

[0250] 72. The method of any one of embodiments 62 to 71, wherein the gRNA comprises a tracrRNA comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56.

[0251] 73. The method of any one of embodiments 62 to 72, wherein the target DNA sequence is located adjacent to a protospacer adjacent motif (PAM).

[0252] 74. The method of any one of embodiments 55 to 73, wherein the target DNA sequence is in a cell.

[0253] 75. The method of embodiment 74, wherein the cell is a eukaryotic cell.

[0254] 76. The method of embodiment 75, wherein the eukaryotic cell is a plant cell.

[0255] 77. The method of embodiment 75, wherein the eukaryotic cell is a mammalian cell.

[0256] 78. The method of embodiment 75, wherein the eukaryotic cell is an insect cell.

[0257] 79. The method of embodiment 74, wherein the cell is a prokaryotic cell.

[0258] 80. The method of any one of embodiments 74 to 79, further comprising culturing the cells under conditions in which the RGN polypeptide is expressed and cleaves the target DNA sequence, thereby producing a modified DNA sequence; and selecting cells containing the modified DNA sequence.

[0259] 81. A cell comprising a target DNA sequence modified according to the method of embodiment 80.

[0260] 82. The cell according to embodiment 81, which is a eukaryotic cell.

[0261] 83. The cell according to embodiment 82, wherein the eukaryotic cell is a plant cell.

[0262] 84. A plant comprising a cell according to embodiment 83.

[0263] 85. A seed comprising a cell according to embodiment 83.

[0264] 86. The cell according to embodiment 82, wherein the eukaryotic cell is a mammalian cell.

[0265] 87. The cell according to embodiment 82, wherein the eukaryotic cell is an insect cell.

[0266] 88. The cell according to embodiment 81, which is a prokaryotic cell.

[0267] 89. A method for producing genetically modified cells that correct a causative mutation of a genetic disease, comprising: a) an RNA-guided (RGN) polypeptide comprising an amino acid sequence having at least 95% identity to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, or a polynucleotide encoding the RGN polypeptide, operably linked to a promoter enabling expression of the RGN polypeptide in a cell; b) introducing a guide RNA (gRNA) comprising a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, and 55, or a polynucleotide encoding the gRNA and operably linked to a promoter that allows expression of the gRNA in the cell; and directing the RGN and the gRNA to the genomic location of the causative mutation to modify the sequence of this genomic location and remove the causative mutation.

[0268] 90. The method of embodiment 89, wherein the RGN is fused to a polypeptide having base editing activity.

[0269] 91. The method of embodiment 90, wherein the polypeptide having base editing activity is cytidine deaminase or adenosine deaminase.

[0270] 92. The method of embodiment 89, wherein the cell is an animal cell.

[0271] 93. The method of embodiment 89, wherein the cell is a mammalian cell.

[0272] 94. The method of embodiment 92, wherein the cells are derived from a dog, cat, mouse, rat, rabbit, horse, cow, pig, or human.

[0273] 95. The method of embodiment 92, wherein the genetic disease is a disease listed in Table 12.

[0274] 96. The method of embodiment 92, wherein the genetic disease is Hurler syndrome.

[0275] 97. The method of embodiment 96, wherein the gRNA further comprises a spacer sequence targeting any of SEQ ID NOs: 453, 454, and 455.

[0276] 98. A method for producing a genetically modified cell that has a deletion in a disease-causing region of genomic instability, comprising: a) an RNA-guided (RGN) polypeptide comprising an amino acid sequence having at least 95% identity to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, or a polynucleotide encoding the RGN polypeptide, operably linked to a promoter enabling expression of the RGN polypeptide in a cell; b) a guide RNA (gRNA) comprising a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, and 55, and further comprising a spacer sequence that targets the 5' flanking region of a region of genomic instability, or a polynucleotide encoding this gRNA and operably linked to a promoter that enables expression of this gRNA in a cell; c) introducing a second guide RNA (gRNA) comprising a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, and 55, and further comprising a spacer sequence targeting the 3' flanking region of the region of genomic instability, or a polynucleotide encoding this gRNA and operably linked to a promoter that enables expression of this gRNA in the cell; A method comprising directing the RGN and the two gRNAs to the region of genomic instability and removing at least a portion of the region of genomic instability.

[0277] 99. The method of embodiment 98, wherein the cell is an animal cell.

[0278] 100. The method of embodiment 98, wherein the cell is a mammalian cell.

[0279] 101. The method of embodiment 100, wherein the cells are derived from a dog, cat, mouse, rat, rabbit, horse, cow, pig, or human.

[0280] 102. The method of embodiment 99, wherein the genetic disease is Friedreich's ataxia or Huntington's disease.

[0281] 103. The method of embodiment 102, wherein the first gRNA further comprises a spacer sequence targeting any of SEQ ID NOs: 468, 469, and 470.

[0282] 104. The method of embodiment 103, wherein the second gRNA further comprises a spacer sequence targeting SEQ ID NO: 471.

[0283] 105. A method for producing genetically modified mammalian hematopoietic progenitor cells with reduced expression of BCL11A mRNA and protein, comprising: a) an RNA-guided (RGN) polypeptide comprising an amino acid sequence having at least 95% identity to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, or a polynucleotide encoding the RGN polypeptide, operably linked to a promoter enabling expression of the RGN polypeptide in a cell; b) introducing a guide RNA (gRNA) comprising a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, and 55, or a polynucleotide encoding the gRNA and operably linked to a promoter that allows expression of the gRNA in the cell; A method comprising expressing the RGN and the gRNA in the cells to achieve cleavage at the BCL11A enhancer region, thereby modifying the genes of the human hematopoietic progenitor cells and reducing expression of BCL11A mRNA and / or protein.

[0284] 106. The method of embodiment 105, wherein the gRNA further comprises a spacer sequence targeting any of SEQ ID NOs: 473, 474, 475, 476, 477, 478.

[0285] 107. To bind to a target DNA sequence, a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); b) comprising an RNA-guided nuclease (RGN) polypeptide having an amino acid sequence that is at least 95% identical to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; A system in which each of the nucleotide sequences encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to each of the nucleotide sequences; hybridizing the one or more guide RNAs to the target DNA sequence; the one or more guide RNAs form a complex with the RGN polypeptide, thereby binding the RGN polypeptide to the target DNA sequence.

[0286] 108. The system of embodiment 107, wherein the RGN polypeptide is nuclease-dead or functions as a nickase.

[0287] 109. The system of embodiment 107 or 108, wherein the RGN polypeptide is operably fused to a base-editing polypeptide.

[0288] 110. The system of embodiment 109, wherein the base-editing polypeptide is a deaminase.

[0289] 111. The system of embodiment 110, wherein the deaminase is cytidine deaminase or adenosine deaminase. [Example]

[0290] The following examples are offered by way of illustration and not by way of limitation.

[0291] experiment

[0292] Example 1: Identification of RNA-guided nucleases

[0293] Seven distinct CRISPR-associated RNA-guided nucleases (RGNs) were identified and are listed in Table 1 below. Table 1 lists the name of each RGN, its amino acid sequence, source, and the processed crRNA and tracrRNA sequences. Table 1 also lists common single guide RNA (sgRNA) sequences, in which polyN indicates the position of the spacer sequence that defines the nucleic acid target sequence of the sgRNA. The RNG and APG systems APG05083.1, APG07433.1, APG07513.1, APG08290.1, and APG08290.1 ​​share the conserved sequence UNANNG (SEQ ID NO: 68) in the base of the hairpin stem of the tracrRNA. For the AP05459.1 system, the sequence in the same position is UNANNU (SEQ ID NO: 557). For the APG04583.1 and APG01688.1 systems, the sequence is UNANNA (SEQ ID NO: 558).

[0294] [Table 1]

[0295] Example 2: Identification of guide RNAs and construction of sgRNAs

[0296] Bacterial cultures originally expressing the RNA-guided nuclease system under investigation were grown to mid-logarithmic phase (OD600 ∼0.600), pelleted, and flash-frozen. RNA was isolated from the pellet using the mirVANA miRNA Isolation Kit (Life Technologies, Carlsbad, CA). Sequencing libraries were prepared from the isolated RNA using the NEBNext Small RNA Library Prep Kit (NEB, Beverly, MA). The library preparations were separated on a 6% polyacrylamide gel into two size fractions corresponding to 18–65 nt and 90–200 nt RNA species, respectively, to detect crRNA and tracrRNA. Deep sequencing (40 bp paired-end for the smaller fraction and 80 bp paired-end for the larger fraction) was performed by a service provider (MoGene, St. Louis, MO) on a Next Seq 500 (High Output kit). Reads were quality trimmed using Cutadapt and mapped to the reference genome using Bowtie2. A custom RNAseq pipeline was written in Phyton to detect crRNA and tracrRNA transcripts. The processed crRNA boundary was determined by sequence coverage of natural repeat spacer arrays. The anti-repeat portion of the tracrRNA was identified using acceptable BLASTn parameters. The processed tracrRNA boundary was confirmed by identifying anti-repeat-containing transcripts from RNA sequencing depth. Manual curation of the RNA was performed using secondary structure prediction with the RNA folding software NUPACK. The sgRNA cassette was prepared by DNA synthesis. The sgRNA cassette was generally designed as follows (5' to 3'): 20-30 bp spacer sequence → processed repeat portion of crRNA → 4 bp non-complementary linker (AAAG; SEQ ID NO: 63) → processed tracrRNA. Other 4 bp non-complementary linkers can also be used (e.g., GAAA (SEQ ID NO: 64) or ACUU (SEQ ID NO: 65)).In some cases, a 6-bp nucleotide linker can be used (e.g., CAAAGG (SEQ ID NO: 66)). For in vitro assays, sgRNAs were synthesized by in vitro transcription of the sgRNA cassette using the GeneArt™ Precision gRNA Synthesis Kit (ThermoFisher). The processed crRNA and tracrRNA sequences for each RGN polypeptide were identified and are listed in Table 1. See below for the sgRNAs constructed for PAM Libraries 1 and 2.

[0297] Example 3: Determining the PAM required for each RGN

[0298] The PAM required for each RGN was determined using a PAM reduction assay essentially adapted from Kleinstiver et al. (2015) Nature 523:481-485 and Zetsche et al. (2015) Cell 163:759-771. Briefly, two plasmid libraries (L1 and L2) were generated in a pUC18 backbone (ampR). Each library contains a distinct 30-bp protospacer (target) sequence flanked by eight random nucleotides (i.e., PAM regions). The target sequences and flanking PAM regions for Library 1 and Library 2 for each RGN are shown in Table 2.

[0299] These libraries were separately electroporated into E. coli BL21(DE3) cells containing the pRSF-1b expression vector containing the RGN of the present invention (codon-optimized for E. coli) and a cognate sgRNA containing a spacer sequence corresponding to the protospacer in L1 or L2. Sufficient library plasmids were used in transformation reactions to produce 10 6Over 1000 cfu were obtained. Both RGN and sgRNA in the pRSF-1b backbone were under the control of a T7 promoter. The transformation reaction was allowed to recover for 1 hour, then diluted into LB medium containing carbenicillin and kanamycin and grown overnight. The next day, the mixture was diluted into autoinducing Overnight Express™ Instant TB medium (Millipore Sigma) to express RGN and sgRNA. After growth for an additional 4 or 20 hours, cells were pelleted and plasmid DNA was isolated using a Mini-prep kit (Qiagen, Germantown, MD). In the presence of the appropriate sgRNA, plasmids containing PAMs recognized by RGN are cleaved, eliminating them from the population. Plasmids containing PAMs not recognized by RGN or introduced into bacteria that do not contain the appropriate sgRNA for transformation survive and replicate. The PAM and protospacer regions of uncleaved plasmids were amplified by PCR and prepared for sequencing according to published protocols (16s Metagenomic Library Preparation Guide 15044223B, Illumina, San Diego, CA). Deep sequencing (80 bp single-end reads) was performed on a MiSeq (Illumina) by a service provider (MoGene, St. Louis, MO). Typically, 1–4 million reads were obtained per amplicon. PAM regions were extracted, counted, and normalized to total reads for each sample. PAMs leading to plasmid cleavage were identified by their abundance relative to the control (i.e., when E. coli containing RGN but lacking the appropriate sgRNA were transformed with the library). To represent the PAM requirement for a novel RGN, the reduction ratio (frequency in sample / frequency in control) for all sequences within the region of interest was converted to an enrichment value using a -log2 transformation. Sufficient PAMs were defined as those with an enrichment value greater than 2.3 (corresponding to a reduction ratio of less than approximately 0.2). PAMs greater than this threshold in both libraries were collected and used to generate web logos.This can be generated, for example, using a web-based service on the Internet known as "weblogo." PAM sequences were identified, and certain patterns were reported among the top enriched PAMs. PAMs (with enrichment values ​​(EF) greater than 2.3) for each RGN are presented in Table 2. Non-limiting examples of PAMs (with EF greater than 3.3) were also identified for some RGNs. For APG005083.1, a representative PAM is NNNNCCR (SEQ ID NO: 70). For APG007513.1, a representative PAM is NNRNCC (SEQ ID NO: 71). For APG001688.1, a representative PAM is NNRANC (SEQ ID NO: 72).

[0300] [Table 2]

[0301] Example 4: Determination of cleavage

[0302] Cleavage sites were revealed by in vitro cleavage reactions using RNP (ribonucleoprotein). Expression plasmids containing RGN fused to a His6 or His10 tag were constructed and transformed into the BL21(DE3) strain of Escherichia coli. Expression was achieved using autoinduction or IPTG induction. After lysis and clarification, the proteins were purified by immobilized metal affinity chromatography.

[0303] Ribonucleoprotein complexes (containing the nuclease and the sgRNA or crRNA-tracrRNA duplex) were formed by incubating the nuclease and RNA in a buffered solution at room temperature for 20 minutes. These complexes were transferred to a tube containing digestion buffer and PCR-amplified target (referred to as "Sequence 1"). Sequence 1 contained the nucleotide sequence (SEQ ID NO: 73) for each RGN directly linked to the corresponding PAM sequence at the 3' end. Each RGN in the ribonucleoprotein complex was incubated with each target polynucleotide at 25°C (APG04583.1) or 37°C (all others) for 30 or 60 minutes (APG05459.1 and APG01688.1 only). The digestion reactions were heat-inactivated and then run on an agarose gel. The cleavage product bands were excised from the gel and sequenced using Sanger sequencing. The cleavage sites were identified by aligning the sequencing results with the predicted sequences of the PCR products, and the results are shown in Table 3. As shown in Table 3, RGN APG007433.1 can also generate blunt cuts with different target sequences.

[0304] The cleavage site of sequence 2 (SEQ ID NO:559, operably fused at its 3' end to the PAM sequence for RGN APG0733.1) was revealed using the following approach for nuclease APG07433.1. After digestion, the gel-purified DNA product was treated with a DNA end repair kit (Thermo Scientific K0771) and ligated to a linearized blunt vector. The resulting circular DNA was transformed into competent E. coli cells. A sticky cut with a 5' overhang would allow detection of overlapping sequences in clones from both cleavage products. A 3' overhang would result in missing sequences, while a blunt cut would allow detection of all the original sequence without overlap. This experiment also confirmed the findings from the above method for sequence 1, detecting that the majority of clones were derived from 5' overlapping cuts. Therefore, the blunt cuts observed are unlikely to be artifacts of this method.

[0305] [Table 3]

[0306] Example 5: Mismatch Sensitivity Assay

[0307] Plasmids were designed and obtained containing the target sequence (SEQ ID NO: 73) immediately 5' of the appropriate PAM motif for the nuclease being evaluated. Sequences with single mismatches and altered sequences at the indicated positions were also generated (Table 4). Purified nuclease (APG08290.1 ​​or APG05459.1) and guide RNA RNP complexes were formed and incubated with PCR-amplified linear DNA from the designed plasmids. After incubation for the indicated time to inactivate the nuclease, samples were analyzed by agarose gel electrophoresis to reveal the fraction of the remaining linear PCA product. The percentage of intact cleaved bands for each mismatch position is shown in Table 5.

[0308] [Table 4]

[0309] [Table 5]

[0310] A similar mismatch sensitivity experiment was performed on RGN APG07433.1. This experiment was similar to that described above, except that the alternative bases were introduced into the RNA guide rather than the DNA target. The DNA sequences of the mismatched sgRNA synthesis are shown in Table 6. The results of the mismatch sensitivity assay are shown in Table 7.

[0311] [Table 6]

[0312] [Table 7]

[0313] RGNs APG07433.1 and APG08290.1, with a few exceptions, show significant sensitivity to mismatches located at positions 1–10 5′ to the PAM (Tables 5 and 7). RGN APG05459.1 is also sensitive to mismatches within this region, but its ability to cleave dsDNA is also significantly impaired by mismatches distant from the PAM site (Table 5). The total number of sites that significantly influence whether cleavage occurs is at least 15 positions within the spacer sequence. This is comparable to other genome editing tools, such as the well-studied Cas9 nuclease from Streptococcus pyogenes, which is generally sensitive to 10–13 base pairs (Hsu et al., Nat Biotechnol (2013) 31(9):827–832). Moreover, the critical sites where RGN APG05459.1-mediated cleavage is absent are remarkably far from the PAM sequence, in the range of 13–20 bp. Many other nucleases exhibit little, if any, sensitivity to mismatches there, a property that can be extremely useful in targeting loci that share close sequence similarity with other sites in the organism of interest.

[0314] Example 6: Demonstration of gene editing activity in mammalian cells

[0315] RGN expression cassettes were constructed and introduced into vectors for mammalian expression. Each RGN was codon-optimized for human expression (SEQ ID NOs: 127-133), operably linked at the 5' end to an SV40 nuclear localization sequence (SEQ ID NO: 134) and a 3xFLAG tag (SEQ ID NO: 135), and operably linked at the 3' end to a nuclear plasmin NLS sequence (SEQ ID NO: 136). Each expression cassette was under the control of a cytomegalovirus (CMV) promoter (SEQ ID NO: 137). It is known in the art that a CMV transcriptional enhancer (SEQ ID NO: 138) can also be incorporated into constructs containing a CMV promoter. Guide RNA expression constructs, each encoding a single gRNA under the control of the human RNA polymerase III U6 promoter (SEQ ID NO: 139), were constructed and introduced into the pTwist High Copy Amp vector. The sequences of the target sequences for each guide are shown in Table 9.

[0316] The above constructs were introduced into mammalian cells. One day before transfection, 1 × 10 5 HEK293T cells / well (Sigma) were seeded in 24-well dishes in Dulbecco's Modified Eagle's Medium (DMEM) plus 10% (vol / vol) fetal bovine serum (Gibco) and 1% penicillin-strepromycin (Gibco). The next day, when cells reached 50-60% confluence, they were co-transfected with 500 ng of RGN expression plasmid and 500 ng of single gRNA expression plasmid using 1.5 μl of Lipofectamine 3000 (Thermo Scientific) per well according to the manufacturer's instructions. After 48 hours of growth, total genomic DNA was collected using a genomic DNA isolation kit (Machery-Nagel) according to the manufacturer's instructions.

[0317] We then analyzed total genomic DNA to determine the editing rate of each RGN for each genomic target. First, we generated oligonucleotides for PCR amplification, and then analyzed the amplified genomic target sites. The sequences of the oligonucleotides used are listed in Tables 8.1–8.5.

[0318] All PCR reactions were performed in 20 μl reactions containing 0.5 μM of each primer using 10 μl of 2× Master Mix Phusion High-Fidelity DNA Polymerase (Thermo Scientific). A large genomic region encompassing each target gene was first amplified using PCR#1 primers, using the following program: 98°C, 1 min; 30 cycles of [98°C, 10 s; 62°C, 15 s; 72°C, 5 min]; 72°C, 5 min; 12°C, indefinite. One microliter of this PCR reaction was then further amplified using primers specific to each guide (PCR#2 primers), using the following program: 98°C, 1 min; 35 cycles of [98°C, 10 s; 67°C, 15 s; 72°C, 30 s]; 72°C, 5 min; 12°C, indefinite. The primers for PCR#2 contain the Nextera Read 1 Transposase Adapter overhang sequence and the Nextera Read 2 Transposase Adapter overhang sequence for Illumina sequencing.

[0319] [Table 8-1]

[0320] [Table 8-2]

[0321] [Table 8-3]

[0322] [Table 8-4]

[0323] [Table 8-5]

[0324] PCR#1 and PCR#2 were performed on the purified genomic DNA as described above. After the second PCR amplification, the DNA was cleaned using a PCR cleanup kit (Zymo) according to the manufacturer's instructions and eluted in water. 200–500 mg of purified PCR#2 product was combined with 2 μl of 10× NEB buffer 2 and water in a 20 μl reaction and annealed to form heteroduplex DNA. The program used was 95°C for 5 min; cooling from 95 to 85°C at a rate of 2°C / sec; cooling from 85 to 25°C at a rate of 0.1°C / sec; and 12°C for eternity. After annealing, 5 μl of DNA was removed as a no-enzyme control, 1 μl of T7 endonuclease I (NEB) was added, and the reaction was incubated at 37°C for 1 h. After incubation, 5x FlashGel loading dye (Lonza) was added, and 5 µl of each reaction and control were analyzed by gel electrophoresis on a 2.2% agarose FlashGel (Lonza). After visualization of the gel, the percentage of non-homologous end joining (NHEJ) was calculated using the formula: %NHEJ events = 100 x [1 - (1 - % cleaved)]. 1 / 2 ] (where (fraction of cleavage) is defined as (density of digested product) / (density of digested product + density of undigested parent band)).

[0325] Several samples were analyzed using SURVEYOR® for post-transfection analysis in mammalian cells. After incubation at 37°C for 72 hours, genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) according to the manufacturer's protocol. Genomic regions flanking the RGN target site were PCR amplified, and the products were purified using QiaQuick spin columns (Qiagen) according to the manufacturer's protocol. A total of 200–500 ng of purified PCR product was mixed with 1 μl of 10× Taq DNA polymerase PCR buffer (Enzymatics) and ultrapure water to a final volume of 10 μl. A reannealing process was applied to allow heteroduplex formation: 95°C for 10 min, followed by a -2°C / s ramp from 95°C to 85°C; a -0.25°C / s ramp from 85°C to 25°C; and a 1-min hold at 25°C.

[0326] After reannealing, the products were treated with SURVEYOR® Nuclease and SURVEYOR® Enhancer S (Integrated DNA Technologies) according to the manufacturer's recommended protocol and analyzed on a 4-20% Novex TBE polyacrylamide gel (Life Technologies). The gel was stained with YBR Gold DNA stain (Life Technologies) for 10 minutes and images were acquired using a Gel Doc gel imaging system (Bio-Rad). Quantification was based on relative band intensity. The indel rate was calculated using the formula: 100 × (1-(1-(b+c) / (a+b+c)) 1 / 2 ) (where a is the integrated intensity of the undigested PCR product, and b and c are the integrated intensities of each cleavage product).

[0327] Additionally, the product from PCR#2, which contained Illumina overhang sequences, was used to prepare a library according to the Illumina 16S metagenomic sequencing library protocol. Deep sequencing was performed by a service provider (MOGene) on an Illumina Mi-Seq platform. Typically, 200,000 250-bp paired-end reads (2 × 100,000 reads) were generated per amplicon. These reads were analyzed using CRISPResso (Pinello et al., 2016 Nature Biotech 34:695-697) to calculate the editing rate. Alignment of the output was manually curated to confirm insertion and deletion sites and identify microhomology sites at recombination sites. The editing rates are shown in Table 9. All experiments were performed in human cells. The "target sequence" refers to the sequence targeted within the gene target. For each target sequence, a guide RNA contained the complementary RNA target sequence and the appropriate sgRNA, depending on the RGN used. A selected breakdown of the experiments is shown in Tables 10.1–10.9 for each guide RNA.

[0328] [Table 9]

[0329] For each guide, specific insertions and deletions are shown in Tables 10.1 through 10.7. In these tables, target sequences are identified by bold capital letters. The octamer PAM region is double underlined, and the key nucleotides recognized are bolded. Insertions are identified by lowercase letters. Deletions are indicated by dotted lines (---). INDEL positions are calculated from the PAM-proximal end of the target sequence, with the edge being position 0. Positions are positive (+) if they are on the target side of the edge, and negative (-) if they are on the PAM side of the edge.

[0330] [Table 10-1]

[0331]

Table 10-2

[0332]

Table 10-3

[0333]

Table 10-4

[0334]

Table 10-5A

Table 10-5B

[0335]

Table 10-6

[0336]

Table 10-7

[0337]

Table 10-8

[0338]

Table 10-9

[0339] Example 7: Demonstration of gene editing activity in plant cells

[0340] The RNA-guided nuclease activity of RGNs according to the present invention was demonstrated in plant cells using a protocol modified from Li et al. (2013) Nat. Biotech. 31:688-691. Briefly, plant codon-optimized versions of each RGN (SEQ ID NOs: 169-182) containing an N-terminal SV40 nuclear localization signal were cloned into a transient transformation vector behind a strong constitutive 35S promoter. sgRNAs targeting one or more sites in the plant PDS gene adjacent to appropriate PAM sequences were cloned into a second transient transformation vector behind a plant U6 promoter. These expression vectors were then introduced into Nicotiana benthamiana mesophyll protoplasts using PEG-mediated transformation. Transformed protoplasts were incubated in the dark for up to 36 hours. Genomic DNA was isolated from the protoplasts using the DNeasy Plant Mini Kit (Qiagen). Genomic regions flanking the RGN target site were amplified by PCR, and the products were purified using QiaQuick spin columns (Qiagen) according to the manufacturer's protocol. A total of 200–500 ng of purified PCR product was mixed with 1 μl of 10× Taq DNA polymerase PCR buffer (Enzymatics) and ultrapure water to a final volume of 10 μl, and subjected to a reannealing process to allow heteroduplex formation: 95°C for 10 min, followed by a -2°C / s ramp from 95°C to 85°C; a -0.25°C / s ramp from 85°C to 25°C; and a 1-min hold at 25°C. After reannealing, the products were treated with SURVEYOR® Nuclease and SURVEYOR® Enhancer S (Integrated DNA Technologies) according to the manufacturer's recommended protocol and analyzed on a 4–20% Novex TBE polyacrylamide gel (Life Technologies). The gel was stained with YBR Gold DNA stain (Life Technologies) for 10 minutes and images were acquired using a Gel Doc gel imaging system (Bio-Rad). Quantification was based on relative band intensity. The indel rate was calculated using the formula: 100 × (1 − (1 − (b + c) / (a + b + c)) 1 / 2) (where a is the integrated intensity of the undigested PCR product, and b and c are the integrated intensities of each cleavage product).

[0341] Alternatively, PCR products derived from the targeted genomic sequence can be subjected to PCR similar to that described in Example 6, resulting in PCR products containing Illumina overhang sequences, allowing library preparation and deep sequencing to be performed. This method allows for the determination of editing rates, as shown in Table 9.

[0342] Example 8: Guide Cross-matching

[0343] To investigate the cross-compatibility of guide RNAs between RGNs, we performed a two-plasmid interference experiment (Esvelt et al. (2013) Nat. Methods 10(11):1116-1121). The first plasmid contained RGNs with several targets containing defined PAMs on a kanamycin resistance backbone. These plasmids were transformed into E. coli BL21, and the transformed strain was made chemically competent. A second plasmid containing guide RNAs on an ampicillin resistance backbone was then introduced. The cells were plated on medium containing both antibiotics. When the RGNs could utilize the guides on the second plasmid, the kanamycin resistance plasmid was cleaved and linearized, resulting in few or no colonies forming. When the RGNs could not utilize the guides on the second plasmid, the kanamycin resistance plasmid was not cleaved, resulting in high levels of colony formation. Guide RNAs for Streptococcus pyogenes Cas9 (SpyCas9) and Staphylococcus aureus Cas9 (SauCas9) were also included, and cross-compatibility was examined using these guide RNAs.

[0344] To calculate the reduction probability, the number of colonies for transformation with each guide is compared to the transformation efficiency using the positive control. Based on this comparison, if the RGN can use the guide, no colonies will survive, so the reduction probability should be 0. If the RGN cannot use the guide, all plasmids will remain intact, so the reduction probability should be 1. The results are shown in Table 11 below. "sg" indicates the guide RNA for the listed RGN.

[0345] [Table 11]

[0346] As can be seen in Table 11, there are four groups of orthogonal systems. RGNs can recognize guides from other systems within their own group, but cannot utilize guides from other groups. The first group includes APG05083.1, APG07433.1, APG07513.1, and APG08290.1. The second group includes SpyCas9 and APG05459.1. The third group includes APG04583.1 and APG01688.1. The fourth group includes SauCas9.

[0347] Example 9: Disease Target Identification

[0348] A clinical variant database was obtained from the NCBI ClinVar database, which is available on the worldwide web through the NCBI ClinVar website. Pathogenic single nucleotide polymorphisms (SNPs) were identified from this list. Genomic locus information was used to identify CRISPR targets in the regions overlapping and surrounding each SNP. A selection of SNPs that can be corrected using base editing in combination with the RGNs of the present invention to target causative mutations is listed in Table 12. Table 12 lists only one common name for each disease. "RS#" corresponds to the RS accession number in the SNP database on the NCBI website. The allele ID corresponds to the accession number of the causative allele, and the chromosome accession number also provides accession reference information that can be viewed through the NCBI website. Table 12 also provides information on the genomic target sequence that matches the listed RGN for each disease. The target sequence information also provides protospacer sequences for generating the necessary sgRNAs for the corresponding RGNs of the present invention.

[0349] Table 12: Diseases targeted by the RGN of the present invention [Table 12-1] [Table 12-2] [Table 12-3] [Table 12-4] [Table 12-5] [Table 12-6] [Table 12-7] [Table 12-8] [Table 12-9] [Table 12-10] [Table 12-11] [Table 12-12] [Table 12-13]

[0350] Example 10: Targeting mutations responsible for Hurler syndrome

[0351] Below, we describe a potential treatment for Hurler syndrome (also known as MPS-1) that uses an RNA-guided base editing system to correct the mutation responsible for Hurler syndrome in a large population of patients with this disease. This approach utilizes a base-editing fusion protein that can be RNA-guided and packaged into a single AAV vector for delivery to a wide range of tissue types. Depending on the precise regulatory elements and base-editing domains used, it may be possible to create a single vector that encodes both the base-editing fusion protein and a single guide RNA that targets the diseased locus.

[0352] Example 10.1: Identifying RGNs with ideal PAMs

[0353] The genetic disease MPS-1 is a lysosomal storage disorder characterized by the molecular accumulation of dharmatan sulfate and heparan sulfate in lysosomes. It is an inherited genetic disorder commonly caused by mutations in the IDUA gene (NCBI Reference Sequence NG_008103.1), which encodes α-L-iduronidase. The disease results from a deficiency of α-L-iduronidase. The most common IDUA mutations found in studies of individuals of Northern European descent are W402X and Q70X. Both are nonsense mutations that result in premature translation termination (Bunge et al. (1994) Hum. Mol. Genet. 3(6):861-866; incorporated herein by reference). Restoring a single nucleotide restores the wild-type coding sequence, which likely results in control of protein expression by the locus's intrinsic regulatory mechanisms.

[0354] The W402X mutation in the human Idua gene accounts for a large proportion of MPS-1H cases. Because base editors can target a sequence window narrower than the binding site of the protospacer component of the guide RNA, the presence of a PAM sequence at a specific distance from the target locus is essential for the success of this strategy. Because the target mutation must reside on an exposed non-target strand (NTS) during interaction with the base-editing protein and the footprint of the RGN domain prevents access to regions near the PAM, accessible loci are considered to be 10–30 bp from the PAM. Various linkers are screened to avoid editing and mutagenesis of nearby adenosine bases within this window. The ideal window is 12–16 bp from the PAM.

[0355] The PAM sequences that fit APG07433.1 and APG08290.1 ​​within the locus location and ideal base editing window described above are readily apparent. These nucleases have the PAM sequences NNNNCC (SEQ ID NO: 6) and NNRNCC (SEQ ID NO: 32), respectively, and their compact size allows for potential delivery via a single AAV vector. This delivery approach offers numerous advantages over alternatives, including access to a wide range of tissues (liver, muscle, CNS), a well-established safety profile, and manufacturing capabilities.

[0356] Although Streptococcus pyogenes Cas9 (SpyCas9) requires the PAM sequence NGG (SEQ ID NO: 448), which is present near the W402X locus, the size of SpyCas9 makes it impossible to package a gene encoding a base-editing domain fusion protein and the SpyCas9 nuclease into a single AAV vector, thereby eliminating the aforementioned advantages of this approach. Including a sequence encoding a guide RNA in this vector is even less feasible, even if significant technical improvements exist to reduce the size of gene regulatory elements or increase the packaging limits of AAV vectors. Dual delivery strategies are available (e.g., Ryu et al. (2018) Nat. Biotechnol. 36(6):536-539; incorporated herein by reference), but they would be extremely complex and costly to produce. Additionally, dual viral vector delivery significantly reduces the efficiency of gene correction. This is because successful editing within a given cell requires infection with both vectors and assembly of the fusion protein within the cell.

[0357] The commonly used Cas9 orthologue from Staphylococcus aureus (SauCas9) is significantly smaller in size than SpyCas9, but requires a more complex PAM (NGRRT (SEQ ID NO: 449)). However, this sequence is not within the range expected to be useful for base editing of the causative locus.

[0358] Example 10.2: RGN fusion constructs and sgRNA sequences

[0359] DNA sequences encoding fusion proteins containing 1) an RGN domain with a mutation that inactivates DNA cleavage activity ("dead" or "nickase") and 2) an adenosine deaminase that aids in base editing are generated using standard molecular biology techniques. All constructs listed in the table below contain a fusion protein with a base editing active domain (in this example, ADAT (SEQ ID NO: 450) operably fused to the N-terminus of RGN APG08290.1). Fusion proteins with a base editing enzyme at the C-terminus of RGN are also known in the art. Additionally, the RGN and base editor in the fusion protein are typically separated by a linker amino acid sequence. Typical linker lengths are known in the art to be in the range of 15 to 30 amino acids. Furthermore, it is known in the art that some fusion proteins can also contain at least one uracil glycosylase inhibitor (UGI) domain between RGN and a base-editing enzyme (e.g., cytidine deaminase) that can improve base-editing efficiency (U.S. Patent No. 10,167,457; incorporated herein by reference). Thus, a fusion protein can contain APG08290.1, a base-modifying enzyme, and at least one UGI.

[0360] [Table 13]

[0361] The editing site accessible to RGN is determined by the PAM sequence. When RGN is combined with a base-editing domain, the residue targeted for editing must be located on the non-target strand (NTS), because the NTS is single-stranded while the RGN is associated with the locus. Evaluating multiple nucleases and corresponding guide RNAs will enable us to select the optimal gene-editing tool for this specific locus. Several PAM sequences in the human Idua gene that could potentially be targeted by the above construct are located near the mutated nucleotide responsible for the W402X mutation. We also create sequences encoding guide RNA transcripts that contain 1) a "spacer" complementary to the non-coding DNA strand at the disease locus and 2) the RNA sequence required for the guide RNA to associate with RGN. Useful guide RNA sequences (sgRNAs) are listed in Table 14 below. The efficiency of these guide RNA sequences in targeting the above base editor to the desired locus can be evaluated.

[0362] [Table 14]

[0363] Example 10.3: Assay for activity in cells derived from patients with Hurler's disease

[0364] To confirm the genotyping strategy and evaluate the above constructs, fibroblasts derived from patients with Hurler disease are used. Similar to the vectors described in Example 5, vectors containing an appropriate promoter upstream of the fusion protein coding sequence and an sgRNA coding sequence for their expression in human cells are designed. Promoters and other DNA elements (e.g., enhancers or terminators) known to be highly expressed in human cells or that can be specifically expressed in fibroblasts can also be used. The vectors are transfected into fibroblasts using standard techniques (e.g., transfection similar to that described in Example 6). Alternatively, electroporation can be used. Cells are cultured for 1-3 days. Genomic DNA (gDNA) is isolated using standard techniques. Editing efficiency is determined by qPCR genotyping assays and / or next-generation sequencing on purified gDNA, as described in further detail below.

[0365] Taqman™ qPCR analysis uses a probe specific for the wild-type allele and a probe specific for the mutant allele. These probes carry fluorophores that are separated by their spectral excitation and / or emission characteristics using a qPCR instrument. Genotyping kits containing PCR primers and probes are commercially available (i.e., Fisher Taqman™ SNP Genotyping Assay ID C__27862753_10 for Thermo SNP ID rs121965019) or can be designed. An example of a designed primer and probe set is shown in Table 15.

[0366] [Table 15]

[0367] Following the editing experiment, gDNA will be analyzed by qPCR using standard methods and the primers and probes described above. Expected results are shown in Table 16. This in vitro system can be used to properly evaluate constructs and select constructs with greater editing efficiency for further study. The system will be evaluated in comparison to cells with and without the W402X mutation, preferably in comparison to several cells heterozygous for this mutation. Ct values ​​will be compared to a reference gene or total amplification of that locus using a dye (e.g., Sybr Green).

[0368] [Table 16]

[0369] Tissues can also be analyzed by next-generation sequencing. Primer binding sites such as those shown below (Table 17) or other suitable primer binding sites identifiable by those skilled in the art can be used. After PCR amplification, libraries containing Illumina Nextera XT overhang sequences are prepared from the products according to the Illumina 16S metagenomic sequencing library protocol. Deep sequencing is performed on an Illumina Mi-Seq platform. Typically, 200,000 250-bp paired-end reads (2 × 100,000 reads) are generated per amplicon. These reads are analyzed using CRISPResso (Pinello et al., 2016) and editing rates are calculated. Manual curation of the output alignments is performed to confirm insertion and deletion sites and identify microhomology sites at recombination sites.

[0370] [Table 17]

[0371] Western blotting of cell lysates from transfected and control cells using anti-IDUA antibodies confirms full-length protein expression, and enzyme activity assays on cell lysates using the substrate 4-methylumbelliferyl αL-iduronide confirm that the enzyme is catalytically active (Hopwood et al., Clin. Chim. Acta (1979) 92(2):257-265; incorporated herein by reference). These experiments were performed using the original Idua W402X / W402X Cell lines (untransfected), Idua transfected with base editing constructs and random guide sequences W402X / W402X cell lines, compared to cell lines expressing wild-type IDUA.

[0372] Example 10.4: Validation of Disease Treatment in Mouse Models

[0373] To validate this therapeutic approach, we use a mouse model harboring a nonsense mutation at a similar amino acid. This mouse strain harbors a W392X mutation in the Idua gene (Gene ID: 15932), which corresponds to a mutation homologous to that found in patients with Hurler syndrome (Bunge et al. (1994) Hum. Mol. Genet. 3(6):861-866; incorporated herein by reference). This locus contains a distinctly different nucleotide sequence from that found in humans and lacks the PAM sequence required for correction using the base editor described in the previous example; therefore, a different fusion protein must be designed to correct the nucleotide. Amelioration of disease in this animal confirms the efficacy of this therapeutic approach to correct the mutation in tissues accessible to gene delivery vectors.

[0374] Mice homozygous for this mutation exhibit many phenotypic characteristics similar to those of patients with Hurler syndrome. The base-edited RGN fusion protein (Table 13) and RNA guide sequence described above are incorporated into an expression vector that enables protein expression and RNA transcription in mice. The study design is shown in Table 18 below. The study includes a group treated with a high-dose expression vector containing the base-edited fusion protein and RNA guide sequence, a control group that is a model mouse treated with an expression vector without the base-edited fusion protein or RNA guide sequence, and a second control group that is a wild-type mouse treated with the same empty vector.

[0375] [Table 18]

[0376] Endpoints evaluated include body weight, urinary GAG excretion, serum IDUA enzyme activity, IDUA activity in tissues of interest, tissue pathology, genotyping of tissues of interest to confirm SNP modification, behavioral assessments, and neurological assessments. Because some endpoints are elimination-based, additional groups can be added to assess, for example, tissue pathology and tissue IDUA activity before the end of the study. Further examples of endpoints can be found in published articles establishing Hurler syndrome animal models (Shull et al. (1994) Proc. Natl. Acad. Sci. USA 91(26):12937-12941; Wang et al. (2010) Mol. Genet. Metab. 99(1):62-71; Hartung et al. (2004) Mol. Ther. 9(6):866-875; Liu et al. (2005) Mol. Ther. 11(1):35-47; Clarketa et al. (1997) Hum. Mol. Genet. 6(4):503-511; all of which are incorporated herein by reference).

[0377] One possible delivery vector utilizes adeno-associated virus (AAV). The vector is constructed to include a sequence encoding the base editor-dRGN fusion protein (e.g., SEQ ID NO: 452) preceded by a CMV enhancer (SEQ ID NO: 138) and promoter (SEQ ID NO: 137) combination, or other suitable enhancer and promoter combination, optionally with a Kozak sequence, and a terminator and polyadenylation sequence operably linked at the 3' end (e.g., the minimal sequence described in Levitt, N.; Briggs, D.; Gil, A.; Proudfoot, NJ, "Definition of an Efficient Synthetic Poly(A) Site," Genes Dev. 1989, 3(7), 1019-1025). The vector further comprises an expression cassette encoding a single guide RNA operably linked at its 5' end to the human U6 promoter (SEQ ID NO: 139) or another promoter suitable for producing small non-coding RNAs, and further comprises inverted terminal repeat (ITR) sequences required for packaging into AAV capsids, as known in the art. Vector construction and viral packaging are performed by standard methods (e.g., those described in U.S. Patent No. 9,587,250, incorporated herein by reference).

[0378] Other possible viral vectors include adenosine and lentiviral vectors (which are commonly used and are thought to contain similar elements) with different packaging capabilities and requirements. Non-viral delivery methods are also available, examples of which include mRNA and sgRNA encapsulated in lipid nanoparticles (Cullis, PR and Allen, TM (2013) Adv. Drug Deliv. Rev. 65(1):36-48; Finn et al. (2018) Cell Rep. 22(9):2227-2235; both incorporated herein by reference), or hydrodynamic injection of plasmid DNA (Suda T and Liu D (2007) Mol. Ther. 15(12):2063-2069; incorporated herein by reference), or ribonucleoprotein complexes of sgRNA associated with gold nanoparticles (Lee, K.; Conboy, M.; Park, HM; Jiang, F.; Kim, HJ; Dewitt, MA; Mackley, VA; Chang, K.; Rao, A.; Skinner, C. et al., "In vivo nanoparticle delivery of Cas9 ribonucleoprotein and donor DNA induces homologous recombination DNA repair," Nat. Biomed. Eng. 2017, Vol. 1(11), pp. 889-890.

[0379] Example 10.5: Disease correction in mouse models with humanized loci

[0380] To evaluate the efficacy of the same base editor constructs that could potentially be used in human therapy, mouse models are needed in which the nucleotides surrounding W392 are changed to match the sequence surrounding human W402. This can be achieved by a variety of techniques, including using RGN and HDR templates to cut and replace the locus in mouse embryos.

[0381] Due to the high degree of amino acid conservation, most nucleotides in the mouse locus can be changed to those of the human sequence with silent mutations, as shown in Table 19. The only base changes that result in coding sequence changes in the resulting engineered mouse genome occur after the introduced stop codon.

[0382] [Table 19]

[0383] Once this mouse strain has been engineered, similar experiments will be performed as described in Example 10.4.

[0384] Example 11: Targeting mutations responsible for Friedreich's ataxia

[0385] The expansion of a trinucleotide repeat sequence that causes Friedreich's ataxia (FRDA) occurs at a defined locus (called the FRDA instability region) within the FXN gene. RNA-guided nucleases (RGNs) can be used to excise the instability region in FRDA patient cells. This approach requires 1) an RGN-guide RNA sequence that can be programmed to target an allele within the human genome; and 2) a method for delivering this RGN-guide RNA sequence. Many nucleases used for genome editing, such as the commonly used Cas9 nuclease from Streptococcus pyogenes (SpCas9), are too large to be packaged within adeno-associated virus (AAV) vectors, especially considering the length of the SpCas9 gene and guide RNA, as well as the other genetic elements required for a functional expression cassette. This limits the potential for an SpCas9 approach.

[0386] The compact RNA-guided nucleases of the present invention (particularly APG07433.1 and APG08290.1) are exceptionally well suited for excising FRDA instability regions. Each RGN requires a PAM near the FRDA instability region. Additionally, each of these RGNs can be packaged into an AAV vector along with a guide RNA. While a second vector will likely be required to package two guide RNAs, this approach is still advantageous over the vectors that would be required for larger nucleases (e.g., SpCas9, where the protein sequence may need to be split across two vectors).

[0387] Table 20 shows the location of suitable genomic target sequences for targeting APG07433.1 or APG08290.1 ​​to the 5' and 3' flanks of the FRDA instability region. When RGN is located at this locus, it is thought to excise the FA instability region. Excision of this region can be confirmed by Illumina sequencing of the locus.

[0388] [Table 20]

[0389] Example 12: Targeting mutations responsible for sickle cell disease

[0390] A target sequence (SEQ ID NO: 472) within the BCL11 enhancer region can provide a mechanism for increasing fetal hemoglobin (HbF), thereby curing or alleviating the symptoms of sickle cell disease. For example, genome-wide association studies have identified a group of genetic variations in BCL11A that are associated with elevated levels of HbF. These variations are a collection of SNPs found within a non-coding region of BCL11A that function as a stage-specific, lineage-restricted enhancer region. Further investigation revealed that this BCL11A enhancer is required for BCL11A expression in erythroid cells (Bauer et al. (2013) Science 343:253-257; incorporated herein by reference). This enhancer region is found within intron 2 of the BCL11A gene, and within intron 2, three DNase I-hypersensitive regions (which often indicate chromatin states associated with regulatory competence) have been identified. These three regions were identified as "+62," "+58," and "+55" according to their distance (in kilobases) from the transcription start site of BCL11A. These enhancer regions are roughly 350 nucleotides in length (+55), 550 nucleotides (+58), and 350 nucleotides in length (+62) (Bauer et al., 2013).

[0391] Example 12.1: Identifying a Preferred RGN System

[0392] Here, we describe the potential for treating β-hemoglobinopathies using the RGN system, which blocks BCL11A from binding to a binding site within the HBB locus (the gene responsible for producing β-globin in adult hemoglobin). This approach takes advantage of the more efficient NHEJ pathway in mammalian cells. Additionally, this approach utilizes a nuclease small enough to be packaged into a single AAV vector for in vivo delivery.

[0393] The GATA1 enhancer motif (SEQ ID NO: 472) within the human BCL11 enhancer region is an ideal target for disruption using an RNA-guided nuclease (RGN), which reduces BCL11A expression in adult human erythrocytes and simultaneously re-expresses HbF (Wu et al. (2019) Nat Med 387:2554). Several PAM sequences compatible with APG07433.1 and APG08290.1 ​​have been readily identified in the locus surrounding this GATA1 site. These nucleases share the PAM sequence 5'-NNNNCC-3' (SEQ ID NO: 6) and are compact in size, potentially enabling delivery in a single AAV or adenoviral vector along with an appropriate guide RNA. This delivery approach offers numerous advantages over other approaches, including access to hematopoietic stem cells, a well-established safety profile, and manufacturing technology.

[0394] The commonly used Cas9 nuclease from Streptococcus pyogenes (SpyCas9) requires a PAM sequence of 5'-NGG-3' (SEQ ID NO: 448), some of which are located near the GATA1 motif. However, due to its size, SpyCas9 cannot be packaged into a single AAV or adenoviral vector, eliminating the advantages of this approach. Dual delivery strategies can be employed, but would be prohibitively complex and costly to produce. Additionally, dual viral vector delivery significantly reduces the efficiency of gene correction because successful editing within a given cell requires infection with both vectors.

[0395] Expression cassettes encoding human codon-optimized APG07433.1 (SEQ ID NO: 128) or APG08290.1 ​​(SEQ ID NO: 130) are generated as described in Example 6. Expression cassettes expressing guide RNAs for the RGNs APG07433.1 and APG08290.1 ​​are also generated. These guide RNAs contain 1) a protospacer sequence complementary to the non-coding or coding DNA strand within the BCL11A enhancer locus (target sequence) and 2) an RNA sequence required to associate the guide RNA with the RGN (SEQ ID NO: 18 for APG07433.1 and SEQ ID NO: 35 for APG08290.1). Because several PAM sequences potentially targeted by APG07433.1 or APG08290.1 ​​surround the BCL11A GATA1 enhancer motif, we generated several potential guide RNA constructs to determine the best protospacer sequence for robust cleavage and NHEJ-mediated disruption of the BCL11A GATA1 enhancer sequence. To target RGN to this locus, we evaluated the target genomic sequences in the table below (Table 21).

[0396] [Table 21]

[0397] To evaluate the efficiency of APG07433.1 or APG08290.1 ​​in generating insertions or deletions that disrupt the BCL11A enhancer region, a human cell line (e.g., human embryonic kidney cells (HEK cells)) is used. A DNA vector containing an RGN expression cassette (e.g., as described in Example 6) is prepared. Another vector is also prepared containing an expression cassette containing a sequence encoding a guide RNA sequence from Table 21. Such an expression cassette can further contain a human RNA polymerase III U6 promoter (SEQ ID NO: 139), as described in Example 6. Alternatively, a single vector containing both an RGN expression cassette and a guide RNA expression cassette can be used. The vector is introduced into HEK cells using standard techniques (e.g., as described in Example 6), and the cells are cultured for 1 to 3 days. After this culture period, genomic DNA is isolated and the frequency of insertions or deletions is determined using T7 endonuclease I digestion and / or direct sequencing, as described in Example 6.

[0398] The DNA region encompassing the target BCL11A is amplified by PCR using primers containing Illumina Nextera XT overhang sequences. These PCR amplicons are then examined for NHEJ formation using T7 endonuclease I digestion, or a library is prepared according to the Illumina 16S metagenomic sequencing library protocol or similar next-generation sequencing (NGS) library preparation. After deep sequencing, the resulting reads are analyzed using CRISPResso to calculate the editing rate. Manual curation of the output alignment is performed to confirm insertion and deletion sites. This analysis identifies preferred RGNs and corresponding preferred guide RNAs (sgRNAs). This analysis may result in equally favorable results for both APG07433.1 and APG08290.1. Additionally, this analysis may determine the presence of more than one preferred guide RNA, or may determine that all target genomic sequences in Table 21 are equally favorable.

[0399] Example 12.2: Assay for determining fetal hemoglobin expression

[0400] This example examines fetal hemoglobin expression when insertions or deletions that interrupt the BCL11A enhancer region are generated by APG07433.1 or APG08290.1. + Hematopoietic stem cells (HSCs) are used. These HSCs are cultured using methods similar to those described in Example 11.1 and transfected with a vector containing an expression cassette containing a region encoding a preferred RGN and a region encoding a preferred sgRNA. After electroporation, these cells are differentiated into erythrocytes in vitro using established protocols (e.g., Giarratana et al. (2004) Nat Biotechnology 23:69-74; incorporated herein by reference). HbF expression is then measured using Western blotting with an anti-human HbF antibody or quantified by high-performance liquid chromatography (HPLC). Successful disruption of the BCL11A enhancer locus is expected to result in increased HbF production compared to HSCs electroporated with only RGN without a guide.

[0401] Example 12.3: Assay for the reduction of sickle cell formation

[0402] This example examines the reduction in sickle cell formation when APG07433.1 or APG08290.1 ​​are used to generate insertions or deletions that disrupt the BCL11A enhancer region. Donor CD34 from patients with sickle cell disease +Hematopoietic stem cells (HSCs) are used. These HSCs are cultured using methods similar to those described in Example 11.1 and transfected with vectors containing an expression cassette containing a region encoding a preferred RGN and a region encoding a preferred sgRNA. Following electroporation, these cells are differentiated in vitro into erythrocytes using established protocols (e.g., Giarratana et al. (2004) Nat Biotechnology 23:69-74). HbF expression is then measured using Western blotting with an anti-human HbF antibody or quantified by high-performance liquid chromatography (HPLC). Successful disruption of the BCL11A enhancer locus is expected to result in increased HbF production compared to HSCs electroporated with only RGN without a guide.

[0403] Addition of metabisulfite to these differentiated red blood cells induces the formation of sickle red blood cells. The number of sickle red blood cells and normal red blood cells are counted under a microscope. The number of sickle red blood cells is expected to be lower in cells treated with APG07433.1 or APG08290.1 ​​and sgRNA compared to untreated cells or cells treated with RGN alone.

[0404] Example 12.4: Validation of Disease Treatment in Mouse Models

[0405] To evaluate the effects of disrupting the BCL11A locus with APG07433.1 or APG08290.1, an appropriate humanized mouse model of sickle cell anemia is used. An expression cassette encoding the desired RGN and an expression cassette encoding the desired sgRNA are packaged into an AAV or adenovirus vector. The adenovirus type Ad5 / 35 is particularly effective for targeting HSCs. An appropriate mouse model containing a humanized HBB locus with a sickle cell allele (e.g., B6;FVB-Tg(LCR-HBA2,LCR-HBB*E26K)53Hhb / J or B6.Cg-Hbatm1Paz Hbbtm1Tow Tg(HBA-HBBs)41Paz / HhbJ) is selected. These mice are treated with granulocyte colony-stimulating factor alone or in combination with plerixafor to mobilize HBCs into the circulation. Next, AAV or adenovirus carrying RGN and guide plasmids are injected intravenously, and the mice are allowed to recover for one week. Blood samples from these mice are analyzed in an in vitro sickle cell assay using metabisulfite, and the mice are followed over time to monitor mortality and hematopoietic function. Treatment with AAV or adenovirus carrying RGN and guide RNA reduces sickle cell counts, lowers mortality, and improves hematopoietic function compared to mice treated with viruses lacking both expression cassettes or with viruses carrying only the RGN expression cassette.

[0406] Item 1 A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence that is at least 95% identical in sequence to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; When the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, it binds to the target DNA sequence in an RNA-guided sequence-specific manner; A nucleic acid molecule, wherein the polynucleotide encoding an RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide. Item 2 2. The nucleic acid molecule of item 1, wherein the RGN polypeptide is nuclease-dead or functions as a nickase. Item 3 3. The nucleic acid molecule of claim 2, wherein the RGN polypeptide is operably fused to a base-editing polypeptide. Item 4 A vector comprising the nucleic acid molecule according to any one of items 1 to 3. Item 5 5. The vector of item 4, further comprising at least one nucleotide sequence encoding the guide RNA, wherein the guide RNA comprises a CRISPR RNA comprising a CRISPR repeat sequence that has at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, or 55. Item 6 6. The vector of item 4 or 5, wherein the guide RNA comprises a tracrRNA having at least 95% sequence identity to any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56. Item 7 A cell comprising the nucleic acid molecule according to any one of items 1 to 3 or the vector according to any one of items 4 to 6. Item 8 1. A system for binding to a target DNA sequence, the system comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence that is at least 95% identical to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, or a nucleotide sequence encoding said RGN polypeptide; Including, each of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to the respective nucleotide sequence; the one or more guide RNAs hybridize to the target DNA sequence; The system wherein the one or more guide RNAs form a complex with the RGN polypeptide, thereby allowing the RGN polypeptide to bind to the target DNA sequence. Item 9 9. The system of item 8, wherein the target DNA sequence is in a eukaryotic cell. Item 10 10. The system of claim 8 or 9, wherein the RGN polypeptide is nuclease-dead or functions as a nickase, and the RGN polypeptide is operably linked to a base-edited polynucleotide. Item 11 11. The system of any one of items 8 to 10, further comprising one or more donor polynucleotides, or one or more nucleotide sequences encoding said one or more donor polynucleotides, wherein each of the nucleotide sequences encoding said one or more donor polynucleotides is operably linked to a promoter heterologous to the respective nucleotide sequence. Item 12 12. A method for binding to a target DNA sequence, comprising delivering the system according to any one of items 8 to 11 to the target DNA sequence or to a cell containing the target DNA sequence. Item 13 1. A method for cleaving and / or modifying a target DNA sequence, comprising: a) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence that is at least 95% identical to any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; b) contacting the RGN of (a) with one or more guide RNAs capable of directing the RGN to the target DNA sequence; wherein the one or more guide RNAs hybridize to the target DNA sequence, thereby causing the RGN polypeptide to bind to the target DNA sequence and cleave and / or modify the target DNA sequence. Item 14 14. The method of claim 13, wherein the modified target DNA sequence comprises an insertion of heterologous DNA into the target DNA sequence. Item 15 14. The method of claim 13, wherein the modified target DNA sequence comprises a deletion of at least one nucleotide from the target DNA sequence. Item 16 Item 14. The method of item 13, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence. Item 17 17. The method of any one of items 14 to 16, wherein the target DNA sequence is in a cell. Item 18 18. The method of item 17, wherein the cell is a eukaryotic cell. Item 19 19. The method of claim 17 or 18, further comprising culturing the cells under conditions in which the RGN polypeptide is expressed to cleave the target DNA sequence and generate a modified DNA sequence; and selecting cells containing the modified DNA sequence. Item 20 20. A cell containing a target DNA sequence modified according to the method of item 19.

Claims

1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence that is at least 95% identical in sequence to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; When the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, it binds to the target DNA sequence in an RNA-guided sequence-specific manner; A nucleic acid molecule, wherein the polynucleotide encoding an RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide.

2. The nucleic acid molecule of claim 1, wherein the RGN polypeptide is nuclease-dead or functions as a nickase.

3. 3. The nucleic acid molecule of Claim 2, wherein the RGN polypeptide is operably fused to a base-editing polypeptide.

4. A vector comprising the nucleic acid molecule of any one of claims 1 to 3.

5. 5. The vector of claim 4, further comprising at least one nucleotide sequence encoding the guide RNA, wherein the guide RNA comprises a CRISPR RNA comprising a CRISPR repeat sequence that has at least 95% sequence identity to any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, or 55.

6. 6. The vector of Claim 4 or 5, wherein the guide RNA comprises a tracrRNA that has at least 95% sequence identity to any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, or 56.

7. A cell comprising a nucleic acid molecule according to any one of claims 1 to 3 or a vector according to any one of claims 4 to 6.

8. 1. A system for binding to a target DNA sequence, the system comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence that is at least 95% identical to any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, or a nucleotide sequence encoding said RGN polypeptide; Including, each of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to the respective nucleotide sequence; the one or more guide RNAs hybridize to the target DNA sequence; The system wherein the one or more guide RNAs form a complex with the RGN polypeptide, thereby allowing the RGN polypeptide to bind to the target DNA sequence.

9. The system of claim 8 , wherein the target DNA sequence is in a eukaryotic cell.

10. The system of claim 8 or 9, wherein the RGN polypeptide is nuclease-dead or functions as a nickase, and the RGN polypeptide is operably linked to a base-edited polynucleotide.

11. 11. The system of any one of claims 8 to 10, further comprising one or more donor polynucleotides, or one or more nucleotide sequences encoding said one or more donor polynucleotides, each of said nucleotide sequences encoding said one or more donor polynucleotides operably linked to a promoter heterologous to the respective nucleotide sequence.

12. 12. A method for binding to a target DNA sequence, the method comprising delivering a system according to any one of claims 8 to 11 to the target DNA sequence or to a cell containing the target DNA sequence.

13. 1. A method for cleaving and / or modifying a target DNA sequence, comprising: a) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence that is at least 95% identical to any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54; b) contacting the RGN of (a) with one or more guide RNAs capable of directing the RGN to the target DNA sequence; wherein the one or more guide RNAs hybridize to the target DNA sequence, thereby causing the RGN polypeptide to bind to the target DNA sequence and cleave and / or modify the target DNA sequence.

14. 14. The method of claim 13, wherein the modified target DNA sequence comprises an insertion of heterologous DNA into the target DNA sequence.

15. 14. The method of claim 13, wherein the modified target DNA sequence comprises a deletion of at least one nucleotide from the target DNA sequence.

16. 14. The method of claim 13, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence.

17. The method of any one of claims 14 to 16, wherein the target DNA sequence is in a cell.

18. 18. The method of claim 17, wherein the cell is a eukaryotic cell.

19. 19. The method of claim 17 or 18, further comprising culturing the cells under conditions in which the RGN polypeptide is expressed to cleave the target DNA sequence and generate a modified DNA sequence; and selecting for cells containing the modified DNA sequence.

20. 20. A cell comprising a target DNA sequence modified according to the method of claim 19.

Citation Information

Patent Citations

  • Novel CAS9 systems and methods of use

    WO2017155714A1

  • Novel CAS9 systems and methods of use

    WO2017155717A1