Nucleic acid molecule, vector and cell comprising nucleic acid molecule, rgn polypeptide, system for binding target DNA sequence, pharmaceutical composition, use of such compositions, method for making rgn polypeptide, in vitro or ex vivo method for binding, cleaving, and / or modifying target DNA sequence, and complex comprising rgn polypeptide
Patent Information
- Application Number
- TW110113939
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-11
- Filing Date
- 2021-04-19
- Publication Date
- 2026-07-21
- Estimated Expiration
- 2041-04-18
AI Technical Summary
Existing genome editing technologies, such as meganucleases, zinc finger fusion proteins, and TALENs, are costly and inefficient for targeted sequence-specific modifications, and RNA-guided nucleases like CRISPR-Cas systems offer a more efficient and cost-effective alternative for targeted genome editing, but require improved methods for sequence-specific guide RNA generation and application in gene editing processes.
The development of RNA-guided nuclease compositions and methods that utilize CRISPR-Cas systems, including specific RNA-guided nuclease polypeptides, guide RNAs, and associated vectors, to target and modify DNA sequences through cleavage, detection, or expression regulation, utilizing non-homologous end joining, homology-directed repair, or base editing.
These compositions and methods enable precise and efficient manipulation of genetic sequences, allowing for targeted modifications, detections, and expression control, enhancing research and therapeutic applications in various organisms.
Abstract
Description
Technical field
[0001] Statement on Sequence Listing
[0002] The present invention is related to the fields of molecular biology and gene editing.
Prior technology
[0003] Targeted genome editing or modification is rapidly becoming an important tool in basic and applied research. Initial approaches involved engineered nucleases such as meganucleases, zinc finger fusion proteins, or TALENs, which required the generation of engineered, programmable, sequence-specific A chimeric nuclease with a unique DNA-binding domain. RNA-guided nucleases (e.g., the CRISPR-Cas protein of the Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR)-associated (cas) bacterial system) can be RNAs (guide RNAs that specifically hybridize to specific target sequences) are complexed to target specific sequences. Generating target-specific guide RNAs is less costly and more efficient than generating chimeric nucleases for each target sequence. Alternatively, such RNA-guided nucleases can be used to edit gene bodies by introducing sequence-specific, double-stranded breaks that are repaired via error-prone non-homologous end-joining (NHEJ). Homologous end joining (NHEJ) is repaired to introduce mutations at specific gene body positions. Alternatively, heterologous DNA can be introduced into the gene body site via homologous recombination repair. RNA-guided nucleases (RGNs) can also be used for base editing when fused to deaminase.
Content of invention
[0004] Compositions and methods for binding a target sequence of interest are provided. The composition can be used to cleave or modify a target sequence of interest, detect a target sequence of interest, and modify the expression of a target sequence of interest. Compositions include RNA-guided nuclease (RGN) polypeptides, CRISPR RNA (crRNA), transcription-activating CRISPR RNA (tracrRNA), guide RNA (gRNA), nucleic acid molecules encoding the same, and vectors and host cells comprising nucleic acid molecules . Also provided are RGN systems for binding a target sequence of interest, wherein the RGN system includes an RNA-guiding nuclease polypeptide and one or more guide RNAs. Accordingly, the methods disclosed herein are described for binding a target sequence of interest, and in some embodiments, for cleaving or modifying a target sequence of interest. The target sequence of interest may be modified, for example, as a result of non-homologous end joining, homology-directed repair of an introduced donor sequence, or base editing.
Implementation
[0005] Many modifications and other embodiments of the inventions set forth herein will come to mind to one having ordinary skill in the art to which these inventions pertain having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. Therefore, it is to be understood that the invention is not to be limited to the particular embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended embodiments. Although specific terms are employed herein, these terms are used in a generic and descriptive sense only and not for purposes of limitation. I. overview
[0006] RNA-guided nucleases (RGNs) allow targeted manipulation of specific site(s) within a gene and are useful in therapeutic and research applications in the context of gene targeting. In various organisms, including mammals, RNA-guided nucleases have been used for genome engineering, for example, by stimulating non-homologous end joining and homologous recombination. The compositions and methods described herein are useful for creating single- or double-stranded breaks in polynucleotides, modifying polynucleotides, detecting specific sites within polynucleotides, or modifying the expression of specific genes.
[0007] The RNA-guided nucleases disclosed herein can alter gene expression by modifying target sequences. In certain embodiments, the RNA-guided nuclease is directed to a target sequence by a guide RNA (gRNA) as part of a Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) RNA-guided nuclease system. RGN is considered "RNA guide" because the guide RNA forms a complex with the RNA-guided nuclease to bind the RNA-guided nuclease guide to the target sequence and, in some embodiments, introduce a single strand at the target sequence or double strand break. After the target sequence has been cleaved, the break can be repaired such that the DNA sequence of the target sequence is modified during the repair process. Accordingly, provided herein are methods of using RNA-guided nucleases to modify a target sequence in the DNA of a host cell. For example, RNA-guided nucleases can be used to modify target sequences at gene body loci in eukaryotic or prokaryotic cells. II. RNA-Guided Nucleases
[0008] Provided herein are RNA-guided nucleases. The term "RNA-guided nuclease (RGN)" refers to a guide RNA molecule that binds to a specific target nucleotide sequence in a sequence-specific manner and is guided to the target nucleotide by misaligning with a polypeptide and hybridizing to the target sequence acid sequence. Although an RNA-guiding nuclease is capable of cleaving a target sequence upon binding, the term "RNA-guiding nuclease" also includes nuclease-inactive RNA-guiding nucleases that are capable of binding to a target sequence but not cleaving the target sequence. Cleavage of target sequences by RNA-guided nucleases can result in single- or double-stranded breaks. Single-stranded RNA-guided nucleases capable of cutting only double-stranded nucleic acid molecules are referred to herein as nickases.
[0009] The RNA-guided nucleases disclosed herein include APG06622, APG02787, APG06248, APG06007, APG02874, APG03850, APG07553, APG03031, APG09208, APG05586, APG08770, APG08167, APG01604, APG 03021, APG06015, APG09344, APG07991, APG01868, APG02998, APG09298, APG06251, APG03066, APG01560, APG02777, APG05761, APG02479, APG08385, APG09217 and APG06657 RNA-guided nucleases, the amino acid sequences of which are shown in SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, respectively. 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123, and 570-579 and retain the ability to bind to a target nucleotide sequence in an RNA-guided, sequence-specific manner its active fragments or variants. In some of these embodiments, APG06622, APG02787, APG06248, APG06007, APG02874, APG03850, APG07553, APG03031, APG09208, APG05586, APG08770, APG08167, APG01604, APG0302 1. APG06015, APG09344, APG07991, APG01868, APG02998, APG09298 Active fragments or variants of APG06251, APG03066, APG01560, APG02777, APG05761, APG02479, APG08385, APG09217 or APG06657 RGN are capable of cleaving single- or double-stranded target sequences. In some embodiments, APG06622, APG02787, APG06248, APG06007, APG02874, APG03850, APG07553, APG03031, APG09208, APG05586, APG08770, APG08167, APG01604, APG03021, APG0 6015, APG09344, APG07991, APG01868, APG02998, APG09298, APG06251, APG03066, Active variants of APG01560, APG02777, APG05761, APG02479, APG08385, APG09217 or APG06657 RGN include combinations with such as SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83 , 89, 96, 103, 110, 117, 123 or the amino acid sequence shown in 570-579 has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80% %, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity of amino acid sequences. In certain embodiments, APG06622, APG02787, APG06248, APG06007, APG02874, APG03850, APG07553, APG03031, APG09208, APG05586, APG08770, APG08167, APG01604, APG03021, APG0 6015, APG09344, APG07991, APG01868, APG02998, APG09298, APG06251, APG03066 , APG01560, APG02777, APG05761, APG02479, APG08385, APG09217 or APG06657 RGN active fragments include such as SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 or at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more consecutive amino acid residues. The RNA-guided nucleases provided herein can include at least one nuclease domain (eg, DNase, RNase domain) and at least one RNA recognition and / or RNA binding domain to interact with a guide RNA. Other domains that may be found in the RNA-guided nucleases provided herein include, but are not limited to: DNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. In particular embodiments, the RNA-guided nucleases provided herein can comprise at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity.
[0010] A target nucleotide sequence binds to an RNA-guided nuclease provided herein and hybridizes to a guide RNA associated with the RNA-guided nuclease. The target sequence can then be subsequently cleaved by an RNA-guided nuclease if the polypeptide has nuclease activity. The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond within the backbone of a referenced nucleotide sequence, which results in a single- or double-stranded break within the subject sequence. The presently disclosed RGNs can act as endonucleases, or can be exonucleases (removal of consecutive nucleotides from the ends (5' and / or 3' ends) of polynucleotides) to cleave Nucleotides. In other embodiments, the disclosed RGNs can cleave nucleotides of a target sequence anywhere within a polynucleotide, and thus function as both endonucleases and exonucleases. Cleavage of target polynucleotides by RGNs disclosed so far can result in staggered breaks or blunt ends.
[0011] The presently disclosed RNA-guided nucleases may be wild-type sequences derived from bacterial or archaeal species. Alternatively, the RNA-guided nuclease may be a variant or fragment of a wild-type polypeptide. For example, wild-type RGN can be modified to alter nuclease activity or to alter PAM specificity. In some embodiments, the RNA-guided nuclease is not naturally occurring.
[0012] In certain embodiments, the RNA-guided nuclease acts as a nickase that cleaves only a single strand of a target nucleotide sequence. Such RNA-guided nucleases have a single functional nuclease domain. In certain embodiments, the nicking enzyme is capable of cleaving either the positive or negative strand. In some of these embodiments, the additional nuclease domain has been mutated such that nuclease activity is reduced or eliminated.
[0013] In other embodiments, the RNA-guided nuclease lacks nuclease activity entirely, and is referred to herein as nuclease-dead or nuclease inactive. Any method known in the art for introducing mutations into amino acid sequences, such as PCR-mediated mutagenesis and site-directed mutagenesis, can be used to generate RGNs free of nicking enzyme or nuclease activity. See, eg, US Publication No. 2014 / 0068797 and US Patent No. 9,790,490; the entire contents of each of these patents are hereby incorporated by reference.
[0014] RNA-guided nucleases lacking nuclease activity can be used to deliver fusion polypeptides, polynucleotides, or small molecule payloads to specific gene body locations. In some of these embodiments, the RGN polypeptide or guide RNA can be fused to a detectable label to allow detection of a specific sequence. As a non-limiting example, a nuclease-free RGN can be fused to a detectable label (eg, a fluorescent protein) and targeted to a specific sequence associated with a disease to allow detection of the sequence associated with the disease.
[0015] Alternatively, nuclease-inactive RGNs can be targeted to specific gene body locations to alter the expression of desired sequences. In some embodiments, the binding of a nuclease-inactive RNA-guided nuclease to a target sequence results in a reduction in the target sequence's activity or expression of the target sequence by interfering with the binding of RNA polymerase or transcription factor within the targeted region of the gene body. Expression of genes under the transcriptional control of the target sequence. In other embodiments, the RGN (e.g., a nuclease-inactive RGN) or its mismatched guide RNA further includes an expression regulator that, when bound to a target sequence, acts to repress or activate the target sequence or is subject to the transcriptional control of the target sequence gene expression. In some of these embodiments, the expression regulator regulates the expression of the target sequence or regulated gene through epigenetic mechanisms.
[0016] In other embodiments, the RGN without nuclease activity or the RGN with only nickase activity can be targeted to a specific gene body position, so as to obtain a base editing polypeptide (for example, a deaminase polypeptide, or its counterpart). Nucleotide bases undergo direct chemical modification (e.g., active variants or fragments of deamination) fusion to modify the sequence of the target polynucleotide, resulting in conversion from one nucleobase to another . The base editing polypeptide can be fused to RGN at its N-terminal side or at its C-terminal end. Additionally, base editing polypeptides can be fused to RGN via a peptide linker. Non-limiting examples of deaminase polypeptides useful in such compositions and methods include cytidine deaminase or adenosine deaminase (e.g., Gaudelli et al. (2017) Nature 551:464-471, US Pub. No. 2017 / 0121693 and 2018 / 0073012, the adenine deaminase base editor described in International Publication No. WO / 2018 / 027078, or International Publication No. WO 2020 / 139873 and No. 63 / 077,089 proposed on September 11, 2020, Any of the deaminases disclosed in U.S. Provisional Application Nos. 63 / 146,840, filed February 8, 2021, and U.S. Provisional Application Nos. 63 / 164,273, filed March 22, 2021, the entire contents of each incorporated herein by reference). In addition, certain fusion proteins between RGN and base editing enzymes known in the art may also include at least one uracil stabilizing polypeptide (uracil stabilizing polypeptide), which can increase cytidine, dehydrogenase in nucleic acid molecules by deaminase Mutation rate of oxycytidine or cytosine with thymidine, deoxythymidine or thymine. Non-limiting examples of uracil-stabilizing polypeptides include uracil-stabilizing polypeptides disclosed in U.S. Provisional Application No. 63 / 052,175, filed July 15, 2020, including USP2 (SEQ ID NO: 1089) and uracil-glucosidase inhibitors (UGI) domain (SEQ ID NO: 212), which increases base editing efficiency. Thus, a fusion protein may include RGN or a variant thereof described herein, deaminase, and optionally at least one uracil-stabilizing polypeptide (eg, UGI or USP2). In certain embodiments, the RGN fused to the base editing polypeptide is a nicking enzyme (eg, deaminase) that cleaves DNA strands where the base editing polypeptide does not function.
[0017] RNA-guided nucleases fused to polypeptides or domains can be linked or separated by linkers. The term "linker" as used herein refers to a chemical group or molecule that joins two molecules or moieties (eg, a binding domain and a cleavage domain of a nuclease). In some embodiments, a linker joins the gRNA binding domain of the RNA-guiding nuclease to the base editing polypeptide, eg, deaminase. In some embodiments, a linker joins the nuclease-inactive RGN and deaminase. Typically, a linker is located between or on both sides of two groups, molecules or other moieties and is attached to each group, molecule or other moiety via a covalent bond, thereby linking the two. In some embodiments, a linker is an amino acid or amino acids (eg, a peptide or protein). In some embodiments, a linker is an organic molecule, group, polymer or chemical moiety. In some embodiments, the length of the linker is 5-100 amino acids, for example, the length is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 , 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70 - 80, 80-90, 90-100, 100-150 or 150-200 amino acids. Longer or shorter linkers are also contemplated.
[0018] The presently disclosed RNA-guided nucleases can include at least one nuclear localization signal (NLS) to enhance delivery of the RGN to the nucleus of the cell. Nuclear localization signals are known in the art and typically include a stretch of basic amino acids (see, eg, Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In certain embodiments, the RGN comprises 2, 3, 4, 5, 6 or more nuclear localization signals. The nuclear localization signal(s) can be a heterologous NLS. Non-limiting examples of nuclear localization signals useful for the presently disclosed RGN are those of SV40 large T antigen, nuclein, and c-Myc (see, e.g., Ray et al. (2015) Bioconjug Chem 26(6): 1004 - 7). In a specific embodiment, the RGN comprises an NLS sequence as shown in SEQ ID NO:251 or 253. The RGN can include one or more NLS sequences at its N-terminus, C-terminus, or both. For example, an RGN may include two NLS sequences at the N-terminal region and four NLS sequences at the C-terminal region.
[0019] Other localization signal sequences known in the art to localize a polypeptide to a specific subcellular location(s) may also be used to target RGN, including but not limited to plastid localization sequences, mitochondrial localization sequences, and plastid targeting sequences. Dual targeting signal sequences for both mitochondria and mitochondria (see, e.g., Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physioldx.doi.org / 10.3389 / fphys.2015.00259 ; Herrmann and Neupert (2003) IUBMB Life 55: 219-225; Soll (2002) Curr Opin Plant Biol 5: 529-535; Carrie and Small (2013) Biochim Biophys Acta 1833: 253-259; Carrie et al. (2009) FEBS J276: 1187-1195; Silva-Filho (2003) Curr Opin Plant Biol6:589-595; Peeters and Small (2001) Biochim Biophys Acta1541:54-63; Murcha et al. (2014) J Exp Bot65:6301-6335; Mackenzie (2005) ) Trends Cell Biol 15:548-554; Glase et al. (1998) Plant Mol Biol 38:311-338).
[0020] In certain embodiments, the presently disclosed RNA-guided nucleases include at least one cell-penetrating domain that facilitates cellular uptake of the RGN. Cell penetrating domains are known in the art and typically include stretches of positively charged amino acid residues (i.e., polycationic cell penetrating domains), alternating polar amino acid residues and nonpolar amine amino acid residues (i.e., amphipathic cell-penetrating domain), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domain) (see, e.g., Milletti F. (2012) Drug Discov Today 17 :850-860). A non-limiting example of a cell penetrating domain is the transcription-activating transcriptional activator (TAT) from human immunodeficiency virus-1.
[0021] The nuclear localization signal, the plastid localization signal, the mitochondrial localization signal, the dual target localization signal, and / or the cell penetration domain can be positioned at the amino terminal (N-terminal), carboxyl terminal (C-terminal), or in an internal position.
[0022] The presently disclosed RGNs can be fused directly or indirectly to effector domains such as cleavage domains, deaminase domains, or expression regulator domains via linker peptides. This domain can be located at the N-terminal, C-terminal or internal position of the RNA-guided nuclease. In some of these embodiments, the RGN component of the fusion protein is a nuclease-free RGN.
[0023] In some embodiments, the RGN fusion protein includes a cleavage domain, which is any domain capable of cleaving a polynucleotide (ie, RNA, DNA, or RNA / DNA hybrid), and includes, but is not limited to, Endonucleases and homing endonucleases, such as type IIS endonucleases (e.g., FokI) (see, e.g., Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388; Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press (1993).
[0024] In other embodiments, the RGN fusion protein includes a deaminase domain that deaminates a nucleobase, resulting in conversion from one nucleobase to another, and including, but not Limited to cytidine deaminase or adenine deaminase base editors (see, e.g., Gaudelli et al. (2017) Nature 551:464-471, U.S. Publication Nos. 2017 / 0121693 and 2018 / 0073012, U.S. Patent Nos. 9,840,699 and International Publication No. WO / 2018 / 027078).
[0025] In some embodiments, the effector domain of the RGN fusion protein can be an expression regulator domain, and the expression regulator domain is a domain used to increase or decrease transcription. The expression regulator domain can be an epigenetic modification domain, a transcriptional repressor domain or a transcriptional activation domain.
[0026] In some of these embodiments, the expression regulator of the RGN fusion protein includes an epigenetic modification domain that covalently modifies DNA or histones to alter histone structure and / or chromosome structure , without altering the DNA sequence, resulting in a change in gene expression (ie, up or down). Non-limiting examples of epigenetic modifications include acetylation or methylation of lysine residues, arginine methylation, serine and threonine phosphorylation, and lysine ubiquitination of histones and ubiquitination (sumoylation), and methylation and hydroxymethylation of cytosine residues in DNA. Non-limiting examples of epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains domain and DNA demethylase domain.
[0027] In other embodiments, the expression regulator of the fusion protein includes a transcriptional repressor domain that interacts with transcriptional control elements and / or transcriptional regulatory proteins such as RNA polymerase and transcription factors to reduce or Transcription of at least one gene is terminated. Transcriptional repressor domains are known in the art and include, but are not limited to, Sp1 -like repressors, IκB, and Krüppel-associated box (KRAB) domains.
[0028] In yet other embodiments, the expression regulator of the fusion protein includes a transcriptional activation domain that interacts with transcriptional control elements and / or transcriptional regulatory proteins such as RNA polymerase and transcription factors to increase or Transcription of at least one gene is activated. Transcription activation domains are known in the art and include, but are not limited to, the herpes simplex virus VP16 activation domain and the NFAT activation domain.
[0029] The presently disclosed RGN polypeptides can include a detectable label or a purification tag. A detectable label or purification tag can be positioned at the N-terminal, C-terminal or internal position of the RNA-guided nuclease either directly or indirectly via a linker peptide. In some of these embodiments, the RGN component of the fusion protein is nuclease-inactive RGN. In other embodiments, the RGN component of the fusion protein is RGN with nickase activity.
[0030] A detectable label is a molecule that can be visualized or otherwise observed. The detectable label can be fused to the RGN as a fusion protein (eg, a fluorescent protein), or can be a small molecule coupled to the RGN polypeptide that can be detected visually or otherwise. Detectable labels that can be fused to the presently disclosed RGN as fusion proteins include any detectable protein domain, including but not limited to protein domains detectable with specific antibodies or fluorescent proteins. Non-limiting examples of fluorescent proteins include green fluorescent proteins (eg, GFP, EGFP, ZsGreen1) and yellow fluorescent proteins (eg, YFP, EYFP, ZsYellow1). Non-limiting examples of small molecule detectable labels include radiolabels, eg, 3H and 35S.
[0031] RGN polypeptides may also include a purification tag, which is any molecule useful for isolating a protein or fusion protein from a mixture (eg, biological sample, culture medium). Non-limiting examples of purification tags include biotin, myc, maltose binding protein (MBP), glutathione-S-transferase (GST), and 3X FLAG tags. III. Guide RNA
[0032] The present disclosure provides guide RNAs and polynucleotides encoding guide RNAs. The term "guide RNA" refers to a nucleotide sequence having sufficient complementarity to a target nucleotide sequence to hybridize to the target sequence and direct the sequence-specific binding of the associated RNA-guiding nuclease to the target nucleotide sequence. Therefore, each guide RNA of RGN is one or more RNA molecules (usually one or two), which can bind to RGN and guide RGN to bind to a specific target nucleotide sequence, and have nickase or nuclease activity at RGN In those embodiments, the target nucleotide sequence is also cleaved. In general, guide RNAs include CRISPR RNA (crRNA) and transcriptional activation CRISPR RNA (tracrRNA). Natural guide RNAs, including both crRNA and tracrRNA, typically include two separate RNA molecules that hybridize to each other via the repeat sequence of the crRNA and the anti-repeat sequence of the tracrRNA.
[0033] The length of natural direct repeat sequences within a CRISPR array typically ranges from 28 to 37 base pairs, although the length can vary from about 23 bp to about 55 bp. Spacer sequences within a CRISPR array typically range in length from about 32 to about 38 bp. However, the length can be between about 21 bp and about 72 bp. Each CRISPR array typically includes less than 50 units of CRISPR repeat-spacer sequences. CRISPR is transcribed as part of a long transcript called the primary CRISPR transcript, which includes a large portion of the CRISPR array. Primary CRISPR transcripts are cleaved by Cas proteins to generate crRNAs, or in some cases, precursor crRNAs (pre-crRNAs), which are further processed by other Cas proteins into mature crRNAs. Mature crRNA includes a spacer sequence and a CRISPR repeat sequence. In some embodiments where the precursor crRNA is processed into a mature (or processed) crRNA, maturation involves removal of about 1 to about 6 or more 5', 3', or 5' and 3' nucleotides. These nucleotides, which are removed during the maturation of the precursor crRNA molecule, are not necessary for the generation or design of the guide RNA for the purpose of genome editing or targeting a specific target nucleotide sequence of interest.
[0034] CRISPR RNA (crRNA) includes a spacer sequence and a CRISPR repeat sequence. A "spacer sequence" is a nucleotide sequence that directly hybridizes to a target nucleotide sequence of interest. The spacer sequence is engineered to be fully or partially complementary to the target sequence of interest. In various embodiments, a spacer sequence can comprise from about 8 nucleotides to about 30 nucleotides or more. For example, the length of the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21 , about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In some embodiments, the length of the spacer sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 , 28, 29, 30 or more nucleotides. In some embodiments, the spacer sequence is about 10 to about 26 nucleotides in length, or about 12 to about 30 nucleotides in length. In a specific embodiment, the spacer sequence is about 30 nucleotides in length. In some embodiments, the spacer sequence is 30 nucleotides in length. In some embodiments, the degree of complementarity between the spacer sequence and its corresponding target sequence is between 50% and 99% or higher when optimally aligned using a suitable alignment algorithm, including but not Limited to include about or greater than about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or higher. In certain embodiments, when optimally aligned using a suitable alignment algorithm, the degree of complementarity between the spacer sequence and its corresponding target sequence is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97% , 98%, 99% or higher. In particular embodiments, the spacer sequence is free of secondary structure predictable using any suitable polynucleotide folding algorithm known in the art, including but not limited to mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, eg, Gruber et al. (2008) Cell 106(1):23-24).
[0035] The CRISPR RNA repeat sequence includes a nucleotide sequence that forms a structure recognized by the RGN molecule either alone or in cooperation with a hybridized tracrRNA. In various embodiments, the CRISPR RNA repeat sequence can comprise from about 8 nucleotides to about 30 nucleotides or more. For example, the length of the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21 , about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In specific embodiments, the length of the CRISPR repeat sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides. In some embodiments, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA sequence is about or greater than about 50%, about 60%, about 70% when optimally aligned using a suitable alignment algorithm , about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or higher. In certain embodiments, when optimally aligned using a suitable alignment algorithm, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA sequence is 50%, 60%, 70%, 75%, 80% , 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97 %, 98%, 99% or higher.
[0036] In certain embodiments, the CRISPR repeat sequence comprises SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90, 97, 104, 111 , 118, or 124 nucleotide sequences, or active variants or fragments thereof, capable of directing the sequence specificity of the associated RNA-guided nuclease provided herein to the target sequence of interest when included in the guide RNA combined. In certain embodiments, the active CRISPR repeat sequence variant of the wild-type sequence comprises the same as SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, The nucleotide sequence shown in 90, 97, 104, 111, 118 or 124 has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, Nucleotide sequences with 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity. In certain embodiments, the active CRISPR repeat sequence fragment of the wild-type sequence includes, for example, SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90, At least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 or 22 consecutive nucleotides.
[0037] In certain embodiments, the crRNA is not naturally occurring. In some of these embodiments, the particular CRISPR repeat sequence is not intrinsically linked to the engineered spacer sequence, and the CRISPR repeat sequence is considered heterologous to the spacer sequence. In certain embodiments, the spacer sequence is a non-naturally occurring engineered sequence.
[0038] A transcriptionally active CRISPR RNA or tracrRNA molecule includes a nucleotide sequence that includes a region of sufficient complementarity to hybridize to the CRISPR repeat sequence of the crRNA, referred to herein as an anti-repeat region. In some embodiments, the tracrRNA molecule further includes a region of secondary structure (eg, a stem-loop), or forms a secondary structure when hybridized to its corresponding crRNA. In a specific embodiment, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is at the 5' end of the molecule, and the 3' end of the tracrRNA includes secondary structure. This region of secondary structure typically includes several hairpin structures, including a nexus hairpin that is found adjacent to the anti-repeat sequence. This junction forms the core of the interaction between the guide RNA and RGN, and is at the intersection between the guide RNA, RGN and target DNA. Junction hairpins generally have a conserved nucleotide sequence in the bases of the hairpin stem, and the motif UNANNC (SEQ ID NO: 132) is present in many junction hairpins in tracrRNA. Interestingly, several RGNs of the present invention use tracrRNAs that include atypical sequences in the bases connecting the hairpin stems of the hairpins, including UNANNA, UNANNU, UNANNG, and CNANNC (SEQ ID NOs: 129, 130, 131 and 133). There is often a terminal hairpin at the 3' end of the tracrRNA, which can vary in structure and number, but typically includes a GC-rich Rho-independent transcription terminator hairpin followed by a string of Us at the 3' end. See, eg, Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Protoc; doi: 10.1101 / pdb.top090902, and US Patent Publication No. 2017 / 0275648, the entire contents of each Incorporated herein by reference.
[0039] In various embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to a CRISPR repeat sequence comprises about 8 nucleotides to about 30 nucleotides or more. For example, the length of the base pairing region between the tracrRNA anti-repeat sequence and the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In a specific embodiment, the length of the base pairing region between the tracrRNA anti-repeat sequence and the CRISPR repeat sequence can be 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides. In some embodiments, when optimally aligned using a suitable alignment algorithm, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA anti-repeat sequence is about or greater than about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90% , about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or higher. In certain embodiments, when optimally aligned using a suitable alignment algorithm, the degree of complementarity between the CRISPR repeat sequence and its corresponding tracrRNA anti-repeat sequence is about or greater than 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95% , 96%, 97%, 98%, 99% or higher.
[0040] In various embodiments, the entire tracrRNA can comprise from about 60 nucleotides to more than about 210 nucleotides. For example, the length of the tracrRNA can be about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210 or more nucleotides. In a specific embodiment, the length of tracrRNA is 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210 or more nucleotides. In a particular embodiment, the length of the tracrRNA is about 80 to about 90 nucleotides, including about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, About 89 and about 90 nucleotides. In certain embodiments, the tracrRNA is 80 to 90 nucleotides in length, including 80, 81, 82, 83, 84, 85, 86, 87, 88, 89 and 90 nucleotides in length.
[0041] In specific embodiments, tracrRNA comprises SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, 112, 119 or 125 nucleotide sequences or active variants or fragments thereof, which when included in the guide RNA can guide the sequence-specific binding of the associated RNA-guided nuclease provided herein to the target sequence of interest. In some embodiments, the active tracrRNA sequence variant of the wild-type sequence comprises the same as SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91 , 98, 105, 112, 119 or 125 have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% %, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher sequence identity of nucleotide sequences. In some embodiments, the active tracrRNA sequence fragment of the wild-type sequence includes such as SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98 , 105, 112, 119 or 125 at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more of the nucleotide sequence shown in multiple consecutive nucleotides.
[0042] Two polynucleotide sequences are considered to be substantially complementary when the two sequences hybridize to each other under stringent conditions. Likewise, an RGN is said to bind to a particular target sequence in a sequence-specific manner if the guide RNA bound to the RGN binds to the target sequence under stringent conditions. "Stringent conditions" or "stringent hybridization conditions" mean conditions under which two polynucleotide sequences hybridize to each other to a detectably greater degree than to other sequences (eg, at least 2-fold greater than background). Stringent conditions are sequence-dependent and will be different in different circumstances. Typically, stringent conditions will be those wherein: at pH 7.0 to 8.3, the salt concentration is less than about 1.5 M Na ion concentration, typically about 0.01 to 1.0 M Na ion concentration (or other salts), and for short sequences (eg, 10 to 50 nucleotides), the temperature is at least about 30°C, and for long sequences (eg, greater than 50 nucleotides), the temperature is at least about 60°C. Stringent conditions can also be achieved by the addition of destabilizing agents such as formamide. Exemplary low stringency conditions include hybridization in a buffer solution of 30 to 35% formamide, 1 M NaCl, 1% SDS (sodium dodecyl sulfate) at 37°C and 1X to 2X SSC at 50 to 55°C (20X SSC = 3.0 M NaCl / 0.3 M trisodium citrate) wash. Exemplary moderately stringent conditions include hybridization in 40 to 45% formamide, 1.0 M NaCl, 1% SDS at 37°C and washes in 0.5X to 1X SSC at 55 to 60°C. Exemplary high stringency conditions include hybridization at 37°C in 50% formamide, 1 M NaCl, 1% SDS and washes in 0.1X SSC at 60 to 65°C. Optionally, the wash buffer may include about 0.1% to about 1% SDS. Duration of hybridization is usually less than about 24 hours, usually about 4 to about 12 hours. The duration of the wash time will be at least a length of time sufficient to achieve equilibrium.
[0043] The Tm is the temperature (under defined ionic strength and pH) at which 50% of a complementary target sequence hybridizes to a perfectly matched sequence. For DNA-DNA hybrids, Tm can be obtained from the equation of Meinkoth and Wahl (1984) Anal. Biochem. 138: 267-284: Tm = 81.5°C + 16.6 (log M) + 0.41 (%GC) - 0.61 (% form ) - 500 / L approximate estimate; where M is the molar concentration of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, %Form is the percentage of formamide in the hybridization solution, and L is the base The length of the hybrids in the base pair. Generally, stringent conditions are selected to be about 5°C lower than the thermal melting point (Tm) for the specific sequence and its complement at a defined ionic strength and pH. However, extremely stringent conditions can be hybridized and / or washed at 1, 2, 3, or 4°C lower than the thermal melting point (Tm); moderately stringent conditions can be performed at 6, 7, 8, 9, or Hybridization and / or washing are performed at 10°C; under low stringency conditions, hybridization and / or washing can be performed at 11, 12, 13, 14, 15 or 20°C lower than the thermal melting point (Tm). Those of ordinary skill in the art will appreciate that using this equation, hybridization and wash compositions, and desired Tm, variations in the stringency of hybridization and / or wash solutions are inherently described. An extensive guide to nucleic acid hybridization can be found in Tijssen's (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York) (Laboratory Techniques in Biochemistry and Molecular Biology— Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York)); and Ausubel et al. eds. (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2nd ed., Cold Spring Harbor Laboratory Press, Plainview, New York).
[0044] The term "sequence specificity" can also refer to binding to a target sequence with a higher frequency than binding to a randomized background sequence.
[0045] The guide RNA can be a single guide RNA or a dual guide RNA system. A single guide RNA includes crRNA and tracrRNA on a single RNA molecule, while a dual guide RNA system includes crRNA and tracrRNA present on two different RNA molecules that hybridize to each other via at least a portion of the CRISPR repeat sequence of the crRNA and at least a portion of the tracrRNA , which may be completely or partially complementary to the CRISPR repeat sequence of the crRNA. In some of those embodiments wherein the guide RNA is a single guide RNA, the crRNA and tracrRNA are separated by a linker nucleotide sequence. In general, in order to avoid the formation of secondary structures within the nucleotides of the linker nucleotide sequence or to avoid the formation of secondary structures including the nucleotides, the linker nucleotide sequence is excluding complementary bases. Nucleotide sequence. In some embodiments, the length of the linker nucleotide sequence between crRNA and tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12 one or more nucleotides. In a specific embodiment, the length of the linker nucleotide sequence of the single guide RNA is at least 4 nucleotides. In some embodiments, the nucleotide sequence of the linker is the nucleotide sequence shown in SEQ ID NO:249.
[0046] Single or dual guide RNAs can be synthesized chemically or via in vitro transcription. Assays for determining sequence-specific binding between RGN and guide RNA are known in the art and include, but are not limited to, in vitro binding assays between RGN and guide RNA that can be detected using a detectable label (eg, biotin) and used in a pull-down detection assay, wherein the guide RNA:RGN complex is captured via a detectable label (eg, using streptavidin magnetic beads). A control guide RNA with a sequence or structure unrelated to the guide RNA can be used as a negative control for non-specific binding of RGN to RNA. In certain embodiments, the guide RNA is SEQ ID NO: 4, 11, 18, 25, 32, 39, 46, 53, 59, 66, 73, 79, 86, 92, 99, 106, 113, 120 or 126, wherein the spacer sequence can be any sequence and is represented by a poly-N (poly-N) sequence.
[0047] In certain embodiments, a guide RNA can be introduced into a target cell, organelle, or embryo as an RNA molecule. Guide RNAs can be transcribed in vitro or chemically synthesized. In other embodiments, a nucleotide sequence encoding a guide RNA is introduced into a cell, organelle or embryo. In some of these embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (eg, an RNA polymerase III promoter). The promoter may be a native promoter or heterologous to the nucleotide sequence encoding the guide RNA.
[0048] In various embodiments, a guide RNA can be introduced into a target cell, organelle, or embryo as a ribonucleoprotein complex, as described herein, wherein the guide RNA is associated with an RNA-guiding nuclease polypeptide.
[0049] A guide RNA directs an associated RNA-guided nuclease to a particular target nucleotide sequence of interest via hybridization of the guide RNA to the target nucleotide sequence. The target nucleotide sequence may comprise DNA, RNA, or a combination of both, and may be single- or double-stranded. The target nucleotide sequence can be genomic DNA (ie, chromosomal DNA), plastid DNA, or RNA molecules (eg, message RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). The target nucleotide sequence can be bound (and in some embodiments cleaved) by an RNA-guided nuclease in vitro or in a cell. The chromosomal sequence targeted by the RGN can be a nuclear, plastid or mitochondrial chromosomal sequence. In some embodiments, the target nucleotide sequence is unique within the target gene body.
[0050] The target nucleotide sequence is adjacent to a prospacer adjacent motif (PAM). Prospacer adjacent motifs are typically within about 1 to about 10 nucleotides from the target nucleotide sequence, including about 1, about 2, about 3, about 4, about 5, about 6 nucleotides from the target nucleotide sequence , about 7, about 8, about 9 or about 10 nucleotides. In a specific embodiment, the PAM is within 1 to 10 nucleotides from the target nucleotide sequence, including 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides from the target nucleotide sequence Nucleotides. The PAM can be 5' or 3' to the target sequence. In some embodiments, the PAM is 3' to the target sequence of the presently disclosed RGN. Typically, the PAM is a consensus sequence of about 3-4 nucleotides, but in particular embodiments, the PAM can be 2, 3, 4, 5, 6, 7, 8, 9 or more nucleosides in length acid. In various embodiments, the currently disclosed PAM sequence recognized by RGN includes SEQ ID NO: 7, 14, 21, 28, 35, 42, 49, 62, 69, 79, 82, 95, 102, 109 or 116 The common sequence shown.
[0051] In a specific embodiment, there is SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, The RNA-guided nuclease of 123 or 570-579 or its variant or fragment binds to SEQ ID NO: 7, 14, 21, 28, 35, 42, 49, 62, 69, 79, 82, 95, 102, 109 Or the target nucleotide sequence adjacent to the PAM sequence shown in 116. In some of these embodiments, the RNG is associated with SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90, 97, 104, 111, CRISPR repeat sequence shown in 118 or 124 or its active variant or fragment and SEQ ID NO: 120 to 128, 140, 142, 145, 147 and 148 respectively shown in the guide of tracrRNA sequence or its active variant or fragment sequence binding. The RGN system is further described in Examples 1-3 and Tables 1 and 2 of this specification.
[0052] A variant of RGN APG05586 (SEQ ID NO: 63) was generated and had the amino acid sequence of SEQ ID NO: 570-579. The RGN having any one of SEQ ID NO:63 and 570-579 can bind a target nucleotide sequence adjacent to the PAM sequence shown in SEQ ID NO:79. In some embodiments, the variant of RGN APG05586 binds to a guide sequence comprising the CRISPR repeat sequence set forth in SEQ ID NO:64 and may also comprise the tracrRNA sequence set forth in SEQ ID NO:65. These RGN systems are further described in Example 5 of this specification.
[0053] It is well known in the art that the PAM sequence specificity for a given nuclease enzyme is affected by the concentration of the enzyme (see, e.g., Karvelis et al. (2015) Genome Biol 16:253), which can be expressed by altering The promoter of RGN or the amount of ribonucleoprotein complex delivered to cells, organelles or embryos is modified.
[0054] When recognizing its corresponding PAM sequence, the RGN can cleave the target nucleotide sequence at a specific cleavage site. As used herein, a cleavage site consists of two specific nucleotides within a target nucleotide sequence between which the nucleotide sequence is cleaved by RGN. The cleavage site may include the 1st and 2nd, 2nd and 3rd, 3rd and 4th, 4th and 5th, 5th and 6th, 5th and 6th, 7th and 8th, or 8th and 9th nucleotides. In some embodiments, the cleavage site can be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or more than 20 nucleotides from the PAM in the 5' or 3' direction . Because RGN can cleave the target nucleotide sequence, resulting in staggered ends, in some embodiments, based on the distance between the two nucleotides on the positive (+) strand of the polynucleotide and the PAM and the distance of the polynucleotide The distance of two nucleotides on the minus (-) strand from the PAM defines the cleavage site. III. Nucleotides encoding RNA-guided nucleases, CRISPR RNA and / or tracrRNA
[0055] The disclosure provides polynucleotides comprising the currently disclosed CRISPR RNA, tracrRNA and / or sgRNA and polynucleotides comprising the nucleotide sequences encoding the currently disclosed RNA-guided nucleases, CRISPR RNA, tracrRNA and / or sgRNA glycosides. The polynucleotides disclosed herein include those polynucleotides comprising or encoding CRISPR repeat sequences including SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, The nucleotide sequence of 71, 77, 84, 90, 97, 104, 111, 118 or 124, or an active variant or fragment thereof, which when included in the guide RNA is capable of guiding the associated RNA-guided nuclease with the Sequence-specific binding of a target sequence of interest. Also disclosed are polynucleotides comprising or encoding tracrRNA comprising SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, A nucleotide sequence of 112, 119, or 125, or an active variant or fragment thereof, which when included in a guide RNA is capable of directing sequence-specific binding of an associated RNA-guiding nuclease to a target sequence of interest. Also provided is a polynucleotide encoding an RNA-guided nuclease comprising, for example, SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83 , 89, 96, 103, 110, 117, 123, or the amino acid sequence shown in 570-579 and active fragments or variants thereof, which maintain the ability to bind to the target nucleotide sequence in an RNA-guided sequence-specific manner ability.
[0056] The use of the terms "polynucleotide" or "nucleic acid molecule" is not intended to limit the present disclosure to polynucleotides, including DNA. Those of ordinary skill in the art will recognize that polynucleotides can include ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. These include peptide nucleic acids (PNAs), PNA-DNA chimeras (chimers), locked nucleic acids (LNAs), and phosphorothioate-linked sequences. The polynucleotides disclosed herein also encompass all forms of sequences including, but not limited to, single-stranded forms, double-stranded forms, DNA-RNA hybrids, triplex structures, stem-loop structures, and the like.
[0057] RGN-encoding nucleic acid molecules can be codon-optimized for expression in an organism of interest. A "codon-optimized" coding sequence is a polynucleotide coding sequence whose codon usage frequency is designed to mimic the preferred codon usage frequency or the transcription conditions of a particular host cell. Expression in that particular host cell or organism is enhanced due to changes in one or more codons at the nucleic acid level leaving the translated amino acid sequence unchanged. Nucleic acid molecules can be codon-optimized in whole or in part. Codon tables and other references that provide preference information for a wide range of organisms are available in the art (see, e.g., Campbell and Gowri (1990) Plant Physiol. 92:1-11, on plant preferred codons sub-use discussion). Methods are available in the art for the synthesis of plant-optimized genes or mammalian (eg, human) codon-optimized coding sequences. See, eg, US Patent Nos. 5,380,831 and 5,436,391 and Murray et al. (1989) Nucleic Acids Res. 17:477-498, which are incorporated herein by reference.
[0058] Polynucleotides encoding the RGN, crRNA, tracrRNA and / or sgRNA provided herein can be provided in expression cassettes for expression in vitro or in a cell, organelle, embryo or organism of interest . The cassette can include 5' and 3' regulatory sequences operably linked to polynucleotides encoding the RGN, crRNA, tracrRNA and / or sgRNA provided herein that allow expression of the polynucleotide. The cassette may additionally contain at least one additional gene or genetic element for co-transformation into the organism. The composition is operably linked if additional genes or elements are included. The term "operably linked" is intended to express a functional linkage between two or more elements. For example, an operative linkage between a promoter and a coding region of interest (eg, a region encoding RGN, crRNA, tracrRNA, and / or sgRNA) is a functional linkage that allows the coding region of interest to express. Operably linked elements may be contiguous or noncontiguous. When used to refer to the joining of two protein coding regions, by operably linked means that the coding regions are in the same reading frame. Alternatively, additional gene(s) or elements may be provided on multiple expression cassettes. For example, the nucleotide sequence encoding the RGN disclosed herein can be present on one expression cassette, while the nucleotide sequence encoding crRNA, tracrRNA or complete guide RNA can be on a separate expression cassette. The expression cassette is provided with a plurality of restriction sites and / or recombination sites such that insertion of the polynucleotide is regulated by the transcription of the regulatory region. The expression cassette may additionally contain a selectable marker gene.
[0059] The expression cassette may include transcription (and in some embodiments translation) initiation regions (ie, promoters), RGN-, crRNA-, tracrRNA- And / or sgRNA-encoding polynucleotides, and transcriptional (and in some embodiments translational) termination regions (ie, termination regions) functional in the organism of interest. The promoters of the present invention are capable of directing or driving the expression of coding sequences in a host cell. Regulatory regions (eg, promoters, transcriptional regulatory regions, and translational termination regions) can be endogenous or heterologous to the host cell, or to each other. As used herein, "heterologous" with respect to a sequence is a sequence derived from a foreign species, or, if from the same species, substantially altered from it in composition and / or genomic loci by deliberate human intervention. Modified sequence. As used herein, a chimeric gene includes a coding sequence operably linked to a transcription initiation region that is heterologous to the coding sequence.
[0060] Suitable termination regions can be obtained from the Ti plastid of A. tumefaciens, such as the octopus carnitine synthase and nopaline synthase termination regions. See also Guerineau et al. (1991) Mol. Gen. Genet. 262: 141-144; Proudfoot (1991) Cell 64: 671-674; Sanfacon et al. (1991) Genes Dev. 5: 141-149; Mogen et al. (1990 ) Plant Cell2: 1261-1272; Munroe et al. (1990) Gene91: 151-158; Ballas et al. (1989) Nucleic Acids Res. 17: 7891-7903; and Joshi et al. (1987) Nucleic Acids Res. 15: 9627 -9639.
[0061] Additional regulatory signals include, but are not limited to, transcription initiation start sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, initiation codons, termination signals, and the like. See, eg, U.S. Patent Nos. 5,039,523 and 4,853,331; EPO 0480762A2; Sambrook et al. (1992), Molecular Cloning: A Laboratory Manual, edited by Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York), below Called "Sambrook 11"; Davis et al. eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY) and references cited therein.
[0062] In preparing the expression cassettes, the various DNA segments can be manipulated to provide the DNA sequence in the appropriate reading frame in the appropriate orientation and under the appropriate circumstances. To this end, adapters or linkers may be used to join the DNA fragments, or other manipulations may be involved to provide suitable restriction sites, remove excess DNA, remove restriction sites, etc. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, de novo substitution, eg switching and transversion may be involved.
[0063] A variety of promoters can be used in the practice of the present invention. A promoter can be selected based on the desired outcome. The nucleic acid can be combined with a persistent, inducible, growth stage specific, cell type specific, tissue preferred, tissue specific or other promoter for expression in an organism of interest. See, eg, WO 99 / 43838 and Nos. 8,575,425; 7,790,846; 8,147,856; 8,586832; 7,772,369; 7,534,939; ; No. 5,608,144 ; No. 5,604,121; No. 5,569,597; No. 5,466,785; No. 5,399,680; No. 5,268,463; No. 5,608,142;
[0064] For expression in plants, continuous promoters also include the CaMV 35S promoter (Odell et al. (1985) Nature 313: 810-812); rice actin (McElroy et al. (1990) Plant Cell 2: 163-171 ); Ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619-632 and Christensen et al. (1992) Plant Mol. Biol. 18:675-689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581-588); and MAS (Velten et al. (1984) EMBO J. 3:2723-2730).
[0065] Examples of inducible promoters are the Adh1 promoter inducible by hypoxia or cold stress, the Hsp70 promoter inducible by heat stress, the PPDK promoter inducible by light, and phosphoenolpyruvate carboxylation Enzyme (pepcarboxylase) promoter. Chemically inducible promoters are also useful, such as the protectant-inducible In2-2 promoter (US Patent No. 5,364,780), the Axig1 promoter that is auxin-inducible and tapetum-specific but also active in callus (PCT US01 / 22169), steroid-responsive promoters (see, for example, Schena et al. (1991) Proc. Natl. Acad. Sci. USA88: 10421-10425 and McNellis et al. (1998) Plant J. 14 (2 ): 247-257 estrogen-inducible ERE promoter and glucocorticoid-inducible promoter), and tetracycline-inducible and tetracycline-repressible promoters (see, e.g., Gatz et al. (1991) Mol. Gen. Genet. 227:229-237 and US Patent Nos. 5,814,618 and 5,789,156), which are incorporated herein by reference.
[0066] Tissue-specific or tissue-preferred promoters are used to target the expression of an expression construct within a specific tissue. In certain embodiments, a tissue-specific or tissue-preferred promoter is active in plant tissue. Examples of promoters that are developmentally controlled in plants include promoters that preferentially initiate transcription in certain tissues such as leaves, roots, fruit, seeds or flowers. A "tissue-specific" promoter is a promoter that initiates transcription only in certain tissues. Unlike the persistent expression of genes, tissue-specific expression is the result of several levels of gene regulatory interactions. In this way, promoters from homologous or closely related plant species can be advantageously used to achieve efficient and reliable expression of the transgene in specific tissues. In some embodiments, expression comprises a tissue-preferred promoter. A "tissue-preferred" promoter is a promoter that initiates transcription preferably, but not necessarily exclusively, or only in certain tissues.
[0067] In some embodiments, the nucleic acid molecule encoding RGN, crRNA and / or tracrRNA includes a cell type specific promoter. A "cell type specific" promoter is a promoter that drives expression primarily in certain cell types in one or more organs. For example, some examples of plant cells in which a cell type specific promoter functioning in plants may be predominantly active include BETL cells, roots, vascular cells in leaves, stalk cells, and stem cells. A nucleic acid molecule may also include a cell type-preferred promoter. A "cell type-preferred" promoter is a promoter that primarily drives expression in certain cell types in one or more organs primarily, but not necessarily exclusively, or only. For example, some examples of plant cells in which a cell type-preferred promoter that functions in plants may be preferentially active include BETL cells, roots, vascular cells in leaves, stalk cells, and stem cells.
[0068] The nucleic acid sequence encoding the RGN, crRNA, tracrRNA and / or sgRNA may be operably linked to a promoter sequence recognized, for example, by phage RNA polymerase for in vitro mRNA synthesis. In such embodiments, in vitro transcribed RNA can be purified for use in the methods described herein. For example, the promoter sequence can be a T7, T3 or SP6 promoter sequence, or a variant of a T7, T3 or SP6 promoter sequence. In such embodiments, the expressed protein and / or RNA can be purified for use in the methods of gene body modification described herein.
[0069] In certain embodiments, polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA can also be associated with polyadenylation signals (e.g., SV40 polyA signals and other signals that function in plants) and / or At least one transcription termination sequence is linked. In addition, the sequence encoding RGN may also be linked to sequence(s) encoding at least one nuclear localization signal, at least one cell penetrating domain and / or at least one signaling peptide capable of transporting proteins to specific subcellular parts, As described elsewhere in this article.
[0070] The polynucleotide encoding the RGN, crRNA, tracrRNA and / or sgRNA may be present in a vector or vectors. "Vector" refers to a polynucleotide composition used to transfer, deliver or introduce a nucleic acid into a host cell. Suitable vectors include plastid vectors, phagemids, cohesoplastids, artificial / minichromosomes, transposons, and viral vectors (eg, lentiviral vectors, adeno-associated viral vectors, baculoviral vectors). The vector can include additional expression control sequences (eg, enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selectable marker sequences (eg, antibiotic resistance genes), origins of replication, and the like. Additional information is available in "Current Protocols in Molecular Biology" by Ausubel et al., John Wiley & Sons, New York, 2003; or "Molecular Cloning: A Laboratory Manual" by Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, N.Y., 3rd edition , found in 2001.
[0071] The vector may also include a selectable marker gene for selection of transformed cells. Selectable marker genes are used for selection of transformed cells or tissues. Marker genes include: genes encoding antibiotic resistance, e.g., genes encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT); Genes that confer resistance to herbicidal compounds such as ketones and 2,4-dichlorophenoxyacetate (2,4-D).
[0072] In some embodiments, an expression cassette or vector comprising a sequence encoding an RGN polypeptide may further comprise a sequence encoding crRNA and / or tracrRNA or crRNA and tracrRNA combined to create a guide RNA. The sequence(s) encoding crRNA and / or tracrRNA may be operably linked to at least one transcriptional control sequence for expression of crRNA and / or tracrRNA in an organism or host cell of interest. For example, a polynucleotide encoding crRNA and / or tracrRNA can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1 and 7SL RNA promoters and rice U6 and U3 promoters.
[0073] As noted, expression constructs comprising nucleotide sequences encoding RGN, crRNA, tracrRNA and / or sgRNA can be used to transform organisms of interest. Methods for transformation involve introducing a nucleotide construct into the organism of interest. By "introducing" is intended to introduce a nucleotide construct into a host cell in such a manner that the construct can enter the interior of the host cell. The method of the present invention does not require a specific method for introducing the nucleotide construct into the host organism, only that the nucleotide construct is capable of entering the interior of at least one cell of the host organism. Host cells can be eukaryotic or prokaryotic. In specific embodiments, the eukaryotic host cell is a plant cell, mammalian cell, avian cell, or insect cell. In some embodiments, a eukaryotic cell that includes or expresses a presently disclosed RGN or that has been modified by a presently disclosed RGN is a human cell. In some embodiments, eukaryotic cells comprising or expressing an RGN disclosed herein or that have been modified by an RGN disclosed herein are cells of hematopoietic origin, e.g., immune cells (i.e., cells of the innate or adaptive immune system) , including but not limited to: B cells, T cells, natural killer (NK) cells, pluripotent stem cells, induced pluripotent stem cells, chimeric antigen receptor T (CAR-T) cells, monocytes, macrophages cells and dendritic cells.
[0074] Methods for introducing nucleotide constructs into plant and other host cells are known in the art and include, but are not limited to, stable transformation methods, transient transformation methods, and virus-mediated methods.
[0075] These methods result in transformed organisms, such as plants, including whole plants as well as plant organs (eg, leaves, stems, roots, etc.), seeds, plant cells, propagules, embryos and progeny thereof. Plant cells can be differentiated or undifferentiated (eg, callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, pollen).
[0076] A "transgenic organism" or "transformed organism" or "stably transformed" organism or cell or tissue refers to an organism that has incorporated or integrated a polynucleotide encoding RGN, crRNA and / or tracrRNA of the present invention. It should be recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments may also be incorporated into the host cell. Agrobacterium and biolistic-mediated transformation remain the two predominantly employed methods for plant cell transformation. However, host cell transformation can be achieved by infection, transfection, microinjection, electroporation, microspraying, gene gun or particle bombardment, electroporation, silica / carbon fiber, ultrasound-mediated, PEG-mediated, calcium phosphate Co-precipitation, polycation DMSO technology, DEAE dextran (dextran) program, and virus-mediated, liposome-mediated and the like. Virus-mediated introduction of polynucleotides encoding RGN, crRNA, and / or tracrRNA includes retrovirus, lentivirus, adenovirus, and adeno-associated virus-mediated introduction and expression, as well as cauliflower mosaic virus, geminivirus, and RNA Use of plant viruses.
[0077] Transformation protocols and protocols for introducing polypeptide or polynucleotide sequences into plants can vary depending on the type of host cell (eg, monocot or dicot cell) targeted for transformation. Methods for transformation are known in the art and include those set forth in U.S. Patent Nos. 8,575,425, 7,692,068, 8,802,934, 7,541,517, each of which is incorporated by reference This article. See also Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett. 7:849-858; Jones et al. (2005) Plant Methods 1:5; Rivera et al. (2012) Physics of Life Reviews 9:308-345; Bartlett et al. (2008) Plant Methods 4: 1-12; Bates, G.W. (1999) Methods in Molecular Biology 111: 359-366; Binns and Thomashow (1988), Annual Reviews in Microbiology 42: 575-606; Christou, P. (1992) The Plant Journal2:275-281; Christou, P. (1995) Euphytica85:13-27; Tzfira et al. (2004) TRENDS in Genetics20:375-383; Yao et al. (2006) Journal of Experimental Botany57:3737-3746; Zupan and Zambryski (1995) Plant Physiology 107:1041-1047; Jones et al. (2005) Plant Methods 1:5.
[0078] Transformation can result in stable or transient incorporation of nucleic acid into the cell. "Stable transformation" means that a nucleotide construct introduced into a host cell is integrated into the genome of the host cell and is capable of being inherited by its progeny. "Transient transformation" refers to the introduction of a polynucleotide into a host cell without integration into the genome of the host cell.
[0079] Methods for chloroplast transformation are known in the art, see, e.g., Svab et al. (1990) Proc. Nail. Acad. Sci. USA87:8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA 90:913-917; Svab and Maliga (1993) EMBO J.12:601-606. The method depends on particle gun delivery of DNA containing a selectable marker and targeting of the DNA to the plastid genome via homologous recombination. Alternatively, plastid transformation can be achieved by transcriptionally activating silent plastid-bearing transgenes through the tissue-preferred expression of nuclear-encoded and plastid-directed RNA polymerases. Such a system has been reported by McBride et al. (1994) in Proc. Natl. Acad. Sci. USA 91:7301-7305.
[0080] Transformed cells can be grown into transgenic organisms, eg, plants, according to conventional means. See, eg, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants can then be grown and pollinated with the same transformant or with different strains, and the resulting hybrids with persistent expression of the desired phenotypic characteristics identified. Two or more generations can be grown to ensure stable maintenance and inheritance of expression of the desired phenotypic trait, and then the seeds are harvested to ensure that expression of the desired phenotypic trait has been achieved. In this manner, the invention provides transformed seeds (also referred to as "transgenic seeds") having a nucleotide construct of the invention (eg, an expression cassette of the invention) stably incorporated into its genome.
[0081] Alternatively, transformed cells can be introduced into an organism. These cells may be derived from the organism wherein the cells are transformed ex vivo.
[0082] The sequences provided herein can be used to transform any plant species, including but not limited to monocots and dicots. Examples of plants of interest include, but are not limited to: maize (corn), sorghum, wheat, sunflower, tomato, crucifers, pepper, potato, cotton, rice, soybean, sugar beet, sugar cane, tobacco, barley and canola, cabbage Canola (Brassica sp.), alfalfa, rye, millet, safflower, peanut, sweet potato, tapioca, coffee, coconut, pineapple, citrus, cocoa, tea, banana, avocado, fig, guava, mango, olive, Papayas, cashews, macadamia, almonds, oats, vegetables, ornamentals and conifers.
[0083] Vegetables include, but are not limited to, tomatoes, lettuce, mung beans, king beans, peas, and members of the genus Curcumis such as cucumbers, cantaloupe, and cantaloupe. Ornamental plants include but are not limited to: azaleas, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, Christmas reds and chrysanthemums. Preferably, the plant of the present invention is an agricultural crop (for example, corn, sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugar cane, tobacco, barley, rapeseed, etc.).
[0084] As used herein, the term "plant" includes plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, plant bushes, and intact plants or plant parts. Plant cells (such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, kernels, ears, corn cobs, husks, stems, roots, root tips, anthers, etc.). Grain means mature seed produced by commercial growers for purposes other than growing or reproducing a species. Progeny, variants and mutants of regenerated plants are also included within the scope of the invention, provided that such parts include the introduced polynucleotide. Further provided are processed plant products or by-products retaining the sequences disclosed herein, including, for example, soybean meal.
[0085] Polynucleotides encoding RGN, crRNA and / or tracrRNA can also be used to transform any prokaryotic species, including but not limited to archaea and bacteria (e.g., Bacillus, Klebsiella sp., Streptomyces Genus, Rhizobium, Escherichia sp., Pseudomonas, Salmonella, Shigella, Vibrio, Yersinia, Mycoplasma, Agrobacterium genus, Lactobacillus.
[0086] Polynucleotides encoding RGN, crRNA, and / or tracrRNA can be used to transform any eukaryotic species, including, but not limited to, animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, proteus Insects, algae and yeast.
[0087] Traditional viral and non-viral based gene transfer methods can be used to introduce nucleic acids into mammalian, insect or avian cells or target tissues. Such methods can be used to administer nucleic acids encoding components of the RGN system to cells in culture or to cells in a host organism. Non-viral vector delivery systems include DNA plastids, RNA (eg, transcripts of the vectors described herein), naked nucleic acid, and nucleic acid conjugated to delivery vehicles such as liposomes. Viral vector delivery systems include DNA and RNA viruses that have episomal or integrated genomes after delivery to cells. For a review of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11: 167-175 (1993); Miller, Nature 357: 455-460 (1992); Van Brunt, Biotechnology 6(10): 1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8: 35-36 ( 1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds.) (1995); and Yu et al., Gene Therapy 1:13-26 (1994).
[0088] Non-viral delivery methods for nucleic acids include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycations or lipid:nucleic acid conjugates, Agent-enhanced uptake of naked DNA, artificial virions, and DNA. For example, lipofection is described in US Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (eg, Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor-recognizing liposome transfection of polynucleotides include those in WO 91 / 17424; WO 91 / 16024 by Feigner. Delivery can be to a cell (eg, administration in vitro or ex vivo) or a target tissue (eg, administration in vivo). Preparation of lipid:nucleic acid complexes, including targeted liposomes, such as immunolipid complexes, is well known to those of ordinary skill in the art (see, e.g., Crystal, Science 270:404-410( 1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994) ); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); No. 4,186,183, No. 4,217,344, No. 4,235,871, No. 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).
[0089] The use of RNA or DNA virus-based systems to deliver nucleic acids exploits highly evolved processes to target viruses to specific cells in the body and transport the viral load to the nucleus. The viral vector can be administered directly to the patient (in vivo), or the viral vector can be used to treat cells in vitro and, optionally, the modified cells can be administered to the patient (ex vivo). Traditional virus-based systems can include retroviral, lentiviral, adenoviral, adeno-associated and herpes simplex virus vectors for gene transfer. With retroviral, lentiviral and adeno-associated virus gene transfer methods, integration in the host genome is possible, often resulting in long-term expression of the inserted transgene. Additionally, high transduction efficiencies have been observed in many different cell types and target tissues.
[0090] The tropism of retroviruses can be altered by the incorporation of foreign envelope proteins, thereby expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and generally producing high viral titers. Therefore, the choice of retroviral gene transfer system will depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats with packaging capacity for up to 6-10 kb of foreign sequences. The minimal cis-acting LTR is sufficient to replicate and package the vector, which can then be used to integrate the therapeutic gene into the target cell to provide permanent transgenic expression. Widely used retroviral vectors include those based on mouse leukemia virus (MuLV), gibbon leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, e.g. , Buchscher et al., J. Viral. 66:2731-2739 (1992); Johann et al., J. Viral. 66:1635-1640 (1992); Sommnerfelt et al., Viral. 176:58-59 (1990); Wilson et al., J. Viral. 63:2374-2378 (1989); Miller et al., 1. Viral. 65:2220-2224 (1991); PCT / US94 / 05700).
[0091] In applications where transient presentation is preferred, adenovirus-based systems can be used. Adenovirus-based vectors are capable of very high transduction efficiency in many cell types and do not require cell division. Using such carriers, high potency and high performance levels have been achieved. This vector can be produced in large quantities in a relatively simple system. Adeno-associated virus ("AAV") vectors can also be used, for example, in the in vitro production of nucleic acids and peptides to transduce cells with target nucleic acids, and in in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987); US Patent No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, 1. Clin. Invest. 94:1351 (1994)) . The construction of recombinant AAV vectors is described in numerous publications, including U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985); Tratschin et al., Mol. Cell. Biol. 4: 2072-2081 (1984); Hermonat & Muzyczka, PNAS 81:6466-6470 (1984); and Samulski et al., 1. Viral. 63:03822-3828 (1989). Packaging cells are commonly used to form viral particles capable of infecting host cells. Such cells include 293 cells packaging adenovirus and ψJ2 cells or PA317 cells packaging retrovirus.
[0092] Viral vectors used in gene therapy are typically produced by producing cell lines that package the nucleic acid vectors into viral particles. The vector generally contains the minimal viral sequences required for packaging and subsequent integration into the host, the other viral sequences being replaced by an expression cassette for the polynucleotide to be expressed. The missing viral function is usually provided in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically have only the ITR sequences from the AAV gene body required for packaging and integration into the host genome. Viral DNA is packaged in a cell line that contains helper plastids encoding the other AAV genes, rep and cap, but lacks the ITR sequence.
[0093] Adenoviruses can also be used as helpers to infect cell lines. The helper virus facilitates the replication of the AAV vector and expression of the AAV genes from the helper plastid. Due to the lack of ITR sequences, this helper plastid was not packaged in large quantities. Contamination with adenovirus can be reduced by, for example, heat treatment, to which adenovirus is more sensitive than AAV. Other methods for delivering nucleic acids to cells are known to those of ordinary skill in the art. See, eg, US20030087817, which is incorporated herein by reference.
[0094] In some embodiments, a host cell is transiently or non-transiently transfected with one or more vectors described herein. In some embodiments, cells are transfected as they naturally occur in an individual. In some embodiments, the transfected cells are obtained from an individual. In some embodiments, the cells are obtained from cells, such as cell lines, obtained from an individual. In some embodiments, the cell line can be a mammalian, insect or avian cell. Various cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to: C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFl, CTLL -2, CIR, Rat6, CVI, RPTE, AlO, T24, 182, A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, TIB55 , lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4. COS, COS-1, COS- 6. COS-M6A, BS-C-1 monkey kidney epithelial cells, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293 -T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23 , COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KG1, KYO1, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO- MAC 6, MTD-1A, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines , Peer, PNT-1A / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR and their transgenic varieties. Cell lines are available from a variety of sources known to those of ordinary skill in the art (see, e.g., American Type Culture Collection (ATCC) (Manassas, VA)) .
[0095] In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines comprising one or more vector-derived sequences. In some embodiments, cells transiently transfected (e.g., by transient transfection of one or more vectors, or transfected with RNA) using components of the RGN system as described herein and modified by activity of the RGN system are established New cell lines include cells that contain the modification but lack any other exogenous sequences. In some embodiments, one or more test compounds are evaluated using cells transiently or non-transiently transfected with one or more vectors described herein, or cell lines obtained from such cells.
[0096] In some embodiments, one or more vectors described herein are used to generate non-human transgenic animals or plants. In some embodiments, the transgenic animal is a mammal, eg, a mouse, rat, hamster, rabbit, cow, or pig. In some embodiments, the transgenic animal is a bird, eg, chicken or duck. In some embodiments, the transgenic animal is an insect, eg, a mosquito or a tick. IV. Variants and fragments of polypeptides and polynucleotides
[0097] The present disclosure provides: active variants and fragments of naturally occurring (ie, wild-type) RNA-guided nucleases, the amino acid sequences of which are as SEQ ID NO: 1, 8, 15, 22, 29, 36 , 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123, or 570-579; and active variants and fragments of naturally occurring CRISPR repeats, e.g., SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90, 97, 104, 111, 118 or 124; and naturally occurring tracrRNA Active variants and fragments of, for example, SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, 112, 119 or 125 the listed sequences; and polynucleotides encoding the sequences.
[0098] Although the activity of the variant or fragment may be altered compared to the polynucleotide or polypeptide of interest, the variant or fragment should retain the functionality of the polynucleotide or polypeptide of interest. For example, a variant or fragment may have increased activity, decreased activity, a different activity profile, or any other alteration in activity when compared to the polynucleotide or polypeptide of interest.
[0099] Fragments and variants of naturally occurring RGN polypeptides such as disclosed herein will retain sequence-specific RNA-guided DNA binding activity. In certain embodiments, fragments and variants of naturally occurring RGN polypeptides such as those disclosed herein will retain nuclease activity (single- or double-stranded).
[0100] Fragments and variants of naturally occurring CRISPR repeats such as those disclosed herein, when used as part of a guide RNA (including tracrRNA), will retain nucleases (and guide RNAs) that bind and guide the RNA in a sequence-specific manner. RNA misalignment) ability to guide to a target nucleotide sequence.
[0101] For example, fragments and variants of naturally occurring tracrRNA disclosed herein, when part of a guide RNA (including CRISPR RNA), will retain the nuclease that guides the RNA in a sequence-specific manner (with which the guide RNA complexes ) ability to guide to a target nucleotide sequence.
[0102] The term "fragment" refers to a portion of a polynucleotide or polypeptide sequence of the invention. A "fragment" or "biologically active portion" includes a polynucleotide comprising a sufficient number of contiguous nucleotides to retain that biological activity (i.e., when included in a guide RNA, in a sequence-specific manner). binds RNA to and directs RGN to the target nucleotide sequence). A "fragment" or "biologically active portion" includes a polypeptide that includes a sufficient number of contiguous amino acid residues to retain biological activity (i.e., binds to a target core in a sequence-specific manner when coupled to a guide RNA). Nucleotide sequence combination). Fragments of RGN proteins include those that are shorter than the full-length sequence due to the use of alternative downstream start sites. The biologically active portion of the RGN protein can include, for example, SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 123 or 570-579 A polypeptide of 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, 1700 or more contiguous amino acid residues. Such biologically active portions can be produced by recombinant techniques and assessed for sequence-specific RNA-guided DNA-binding activity. A biologically active fragment of a CRISPR repeat sequence may comprise SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90, 97, 104, 111, 118 or 124 at least 8 consecutive amino acids. A biologically active portion of a CRISPR repeat sequence can include, for example, SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90, 97, 104, 111, 118 or 124 polynucleotides of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 or 22 contiguous nucleotides. The biologically active portion of tracrRNA may comprise, for example, SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, 112, 119 or 125 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40 , 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110 or more contiguous nucleotides.
[0103] In general, "variant" is intended to mean substantially similar sequences. For polynucleotides, variants include deletions and / or additions of one or more nucleotides at one or more internal sites in the native polynucleotide and / or one or more deletions in the native polynucleotide. Substitution of one or more nucleotides at a site. As used herein, a "native" or "wild-type" polynucleotide or polypeptide includes a naturally occurring nucleotide sequence or amino acid sequence, respectively. With respect to polynucleotides, conservative variants include those sequences that, due to degeneracy of the genetic code, encode the native amino acid sequence of the gene of interest. For example, naturally occurring dual variants can be identified using well known molecular biology techniques (eg, using polymerase chain reaction (PCR) and hybridization techniques as outlined below). Variant polynucleotides also include synthetically derived polynucleotides, such as those produced, eg, by using site-directed mutagenesis, but which still encode a polypeptide or polynucleotide of interest. In general, variants of a particular polynucleotide disclosed herein will be identical to that as determined by the sequence alignment programs and parameters described elsewhere herein. A particular polynucleotide has at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94% %, 95%, 96%, 97%, 98%, 99% or greater sequence identity.
[0104] Variants of a particular polynucleotide disclosed herein (i.e., the reference polynucleotide) can also be detected by comparing the polypeptide encoded by the variant polynucleotide with the polypeptide encoded by the reference polynucleotide. The percent sequence identity between them was evaluated. The percent sequence identity between any two polypeptides can be calculated using the sequence alignment programs and parameters described elsewhere herein. When any given polynucleotide pair disclosed herein is assessed by comparing the percent sequence identity common to the two polypeptides encoded by any given polynucleotide pair disclosed herein, the difference between the two encoded polypeptides is The percent sequence identity between is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93% , 94%, 95%, 96%, 97%, 98%, 99% or higher consistency.
[0105] In a particular embodiment, the polynucleotide disclosed in the present invention encodes an RNA-guided nuclease polypeptide, and the RNA-guided nuclease polypeptide includes an amino acid sequence identical to that of SEQ ID NO: 1, 8, 15, 22, The amino acid sequence of 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 or 570-579 has at least 40%, 45%, 50%, 55 %, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher concordance. In some embodiments, the variant of SEQ ID NO: 63 maintains isoleucine at the amino acid position corresponding to 305 of SEQ ID NO: 63, isoleucine at the amino acid position corresponding to 328 of SEQ ID NO: 63 Valine at the amino acid position, leucine at the amino acid position corresponding to 366 of SEQ ID NO: 63, threonine at the amino acid position corresponding to 368 of SEQ ID NO: 63, and valine at the amino acid position corresponding to 405 of SEQ ID NO:63. An amino acid position in a first amino acid sequence that "corresponds" to a particular position in a second amino acid sequence refers to the position of the first amino acid sequence when the first and second amino acid sequences are optimally aligned. The position in which aligns with the specified amino acid residue position in the second sequence. In a specific embodiment, the variant of SEQ ID NO: 63 has the amino acid residues of SEQ ID NO: 63 other than those amino acid residues maintained in SEQ ID NO: 63 (i.e., I305, V328, L366, T368 and V405) At least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88 %, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher consistency.
[0106] Biologically active variants of the RGN polypeptides of the invention may differ by as little as about 1-15 amino acid residues, as little as about 1-10 (e.g., about 6-10), as little as 5, As few as 4, as few as 3, as few as 2, or as few as 1 amino acid residues. In a specific embodiment, the polypeptide may include an N-terminal or C-terminal truncation, which may include at least deletions from the N-terminal or C-terminal of the polypeptide 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, 1700 or more amino acids.
[0107] In certain embodiments, the polynucleotides disclosed in the present invention include or encode CRISPR repeats, which include sequences with SEQ ID NOs: 2, 9, 16, 23, 30, 37, 44, 51, The nucleotide sequence shown at 57, 64, 71, 77, 84, 90, 97, 104, 111, 118 or 124 has at least 40%, 45%, 50%, 55%, 60%, 65%, 70% , 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95 %, 96%, 97%, 98%, 99% or more identical nucleotide sequences.
[0108] The polynucleotides disclosed in the present invention include or encode tracrRNA, which includes sequences with SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91 , 98, 105, 112, 119 or 125 have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82 %, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, Nucleotide sequences with 99% or greater identity.
[0109] Biologically active variants of the CRISPR repeats or tracrRNAs of the invention may differ by as little as about 1-15 amino acid residues, by as little as about 1-10 (e.g., about 6-10), by as little as 5, as few as 4, as few as 3, as few as 2, or as few as 1 nucleotide. In particular embodiments, a polynucleotide may comprise a 5' or 3' truncation, which may comprise at least a deletion of 10, 15, 20, 25, 30, 35, 40, 45 from the 5' or 3' end of the polynucleotide. , 50, 55, 60, 65, 70, 75, 80, 90, 95, 100, 105, 110 or more nucleotides.
[0110] It will be recognized that the RGN polypeptides, CRISPR repeats, and tracrRNA provided herein can be modified to generate variant proteins and polynucleotides. Engineered changes can be introduced via the application of site-directed mutagenesis techniques. Alternatively, native, unknown or as yet unidentified polynucleotides and / or polypeptides that are structurally and / or functionally related to the sequences disclosed herein are also considered to fall within the scope of the present invention. Conserved amino acid substitutions can be made in non-conserved regions that do not alter the function of the RGN protein. Alternatively, modifications that improve RGN activity can be made.
[0111] Variant polynucleotides and proteins also include sequences and proteins derived from, for example, mutagenesis and recombination procedures (eg, DNA shuffling). Using this program, one or more of the different RGN proteins disclosed herein (e.g., SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83 , 89, 96, 103, 110, 117, 123, or 570-579) to generate new RGN proteins with desired properties. In this manner, libraries of recombinant polynucleotides are generated from a population of related sequence polynucleotides comprising sequence regions that have substantial sequence identity and can be homologously recombined in vitro or in vivo. For example, using this approach, sequence motifs encoding domains of interest can be shuffled between the RGN sequences provided herein and other known RGN genes to obtain pairs with improved properties of interest (e.g., in the case of enzymes, Increased Km) of protein-coding novel genes. Strategies for such DNA shuffling are known in the art. See, eg, Stemmer (1994) Proc. Natl. Acad. Sci. USA 91: 10747-10751; Stemmer (1994) Nature 370: 389-391; Crameri et al. (1997) Nature Biotech. 15: 436-438; Moore et al. ( 1997) J. Mol. Biol.272:336-347; Zhang et al. (1997) Proc. Natl. Acad. Sci. USA94:4504-4509; Crameri et al. (1998) Nature391:288-291; 5,605,793 and 5,837,458. A "shuffled" nucleic acid is a nucleic acid produced by a shuffling procedure, such as any of the shuffling procedures described herein. Shuffled nucleic acids are produced by recombining (physically or virtually) two or more nucleic acids (or strings), eg, artificially and optionally recursively. Typically, one or more screening steps are used during shuffling to identify nucleic acids of interest; such screening steps can be performed before or after any recombination steps. In some (but not all) shuffling implementations, it is desirable to perform multiple rounds of recombination prior to selection to increase the diversity of the pool to be screened. The entire process of recombination and selection is optionally repeated recursively. Depending on the context, shuffling can refer to the entire process of recombination and selection, or alternatively, can refer to only the recombination portion of the entire process.
[0112] As used herein, "sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences relates to the sequence identity when aligned for maximum correspondence over a specified comparison window. Residues in the same two sequences. When using percentage sequence identities with respect to proteins, it should be recognized that different residue positions often differ by reserved amino acid substitutions, where amino acid residues are replaced with similar chemical properties (e.g., charge or hydrophobicity). ), and thus do not alter the functional properties of the molecule. When sequences differ in conserved substitutions, the percent sequence identity can be adjusted upwards to correct for the conserved nature of the substitutions. Sequences that differ by such reserved substitutions are said to have "sequence similarity" or "similarity". Means for making such adjustments are well known to those of ordinary skill in the art. Typically, this involves counting reserved substitutions as partial rather than full mismatches, thereby increasing the percent sequence identity. Thus, for example, where the same amino acid has a score of 1, and a non-reserved substitution has a score of zero, a reserved substitution has a score between 0 and 1. Scores for retained substitutions are calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).
[0113] As used herein, "percent sequence identity" refers to a value determined by comparing two optimally aligned sequences over a comparison window, where the positions of the polynucleotide sequences in the comparison window may include Additions or deletions (ie, gaps) compared to a reference sequence (excluding additions or deletions) for optimal alignment of the two sequences. The number of matching positions is found by determining the number of positions at which the same nucleic acid base or amino acid residue occurs in the two sequences, dividing the number of matching positions by the total number of positions in the comparison window, and multiplying the result by 100 to find The percent sequence identity was obtained to calculate the percent.
[0114] Unless otherwise stated, sequence identity / similarity values provided herein refer to values obtained using GAP version 10 with the following parameters: nucleosides using GAP weight 50 and length weight 3 and nwsgapdna.cmp scoring matrix % identity and % similarity for acid sequences; % identity and % similarity for amino acid sequences using GAP weight 8 and length weight 2 and BLOSUM62 scoring matrix; or any equivalent thereof. "Equivalent program" means any sequence comparison program that, for any two sequences involved, produces sequences with identical nucleotide or amino acid residues when compared to a corresponding alignment generated by GAP version 10. Alignment of base matches and percent identical sequence identities.
[0115] When two sequences are aligned using a defined amino acid substitution matrix (eg, BLOSUM62), a gap existence penalty (gap existence penalty) and a gap extension penalty (gap extension penalty), to achieve the possible Two sequences are "best aligned" when the highest score of . Amino acid substitution matrices and their use in quantifying the similarity between two sequences are well known in the art and are described, for example, in Dayhoff et al. (1978) "A model of evolutionary change in proteins Models of Evolutionary Change); "Atlas of Protein Sequence and Structure (Protein Sequence and Structure Atlas)", Volume 5, Suppl. 3 (M. O. Dayhoff edited), pp. 345-352; Natl. Biomed. Res. Found ., Washington, DC; and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA 89:10915-10919. BLOSUM62 matrices are often used as pre-scored substitution matrices in sequence alignment workflows. A gap presence penalty is imposed for the introduction of a single amino acid gap in one of the aligned sequences, while a gap extension penalty is imposed for each additional empty amino acid position inserted into an already opened gap. An alignment is defined by aligning the amino acid positions of each sequence at the beginning and end, and optionally by inserting one or more gaps in one or both sequences, to achieve the highest possible score. Although optimal alignment and scoring can be done manually, the process can be improved by using computer-implemented alignment algorithms such as those described by Altschul et al. in (1997) Nucleic Acids Res. Gapped BLAST 2.0, available to the public at the National Center for Biotechnology Information website (www.ncbi.nlm.nih.gov), facilitates this process. Optimal alignments including multiple alignments can be prepared using, for example, PSI-BLAST available via www.ncbi.nlm.nih.gov and described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402 .
[0116] With respect to an amino acid sequence that is optimally aligned with a reference sequence, an amino acid residue "corresponds to" the position in the reference sequence to which the residue is paired in the alignment. The "position" is represented by a number that sequentially identifies each amino acid in the reference sequence based on its position relative to the N-terminus. Because of the deletions, insertions, truncations, fusions, etc. that must be considered in determining optimal alignment, the number of amino acid residues in the test sequence, which can usually be determined by simply counting from the N-terminus, does not necessarily have to be the same as that in the reference sequence. The number of corresponding positions is the same. For example, where a deletion is present in the aligned test sequences, there will be no amino acid corresponding to the position at the site of the deletion in the reference sequence. Where there is an insertion in the aligned reference sequence, the insertion will not correspond to any amino acid position in the reference sequence. In the case of truncations or fusions, there may be stretches of amino acids in the reference sequence or aligned sequences that do not correspond to any amino acids in the corresponding sequences. V. Antibodies
[0117] Also comprising ribonucleoproteins directed against or comprising RGN polypeptides of the present invention, including those having SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70 , 76, 83, 89, 96, 103, 110, 117, 123 or those RGN polypeptides or ribonucleoproteins of the amino acid sequences shown in 570-579 or active variants or fragments thereof. Methods for generating antibodies are well known in the art (see, e.g., Harlow and Lane (1988) Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y.); and US Patent No. 4,196,265). These antibodies are available in kits for the detection and isolation of RGN polypeptides or ribonucleoproteins. Accordingly, the present disclosure provides sets comprising antibodies that specifically bind to a polypeptide or ribonucleoprotein described herein, including, for example, those having SEQ ID NO: 1, 8, 15, 22, 29, A polypeptide of the sequence 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 or 570-579. VI. Systems and Ribonucleoprotein Complexes for Binding a Target Sequence of Interest and Methods of Making the Same
[0118] The present disclosure provides a system for binding a target sequence of interest, wherein the system includes at least one guide RNA or a nucleotide sequence encoding the at least one guide RNA and at least one RNA-guided nuclease or A nucleotide sequence encoding the at least one RNA-guided nuclease. The guide RNA hybridizes to the target sequence of interest and also forms a complex with the RGN polypeptide, thereby guiding the RGN polypeptide to bind to the target sequence. In some of these embodiments, the RGN comprises SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117 , 123 or the amino acid sequence of 570-579 or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence comprising SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90 , 97, 104, 111, 118 or 124 nucleotide sequence or an active variant or fragment thereof. In some specific embodiments, guide RNA comprises tracrRNA, and this tracrRNA comprises SEQ ID NO:3,10,17,24,31,38,45,52,58,65,72,78,85,91,98,105 , 112, 119 or 125 nucleotide sequences or active variants or fragments thereof. The guide RNA of this system can be single guide RNA or double guide RNA. In some specific embodiments, the system includes an RNA-guided nuclease that is heterologous to the guide RNA, wherein the RGN and guide RNA are not inherently misaligned with each other (ie, bind to each other).
[0119] The systems provided herein for binding a target sequence of interest can be a ribonucleoprotein complex, which is at least one molecule of RNA bound to at least one protein. The ribonucleoprotein complexes provided herein include at least one guide RNA as part of the RNA and an RNA-guided nuclease as part of the protein. Such ribonucleoprotein complexes can be purified from cells or organisms that naturally express RGN polypeptides, and such ribonucleoprotein complexes have been engineered to express a specific guide RNA specific for a target sequence of interest . Alternatively, ribonucleoprotein complexes can be purified from cells or organisms that have been transformed with a polynucleotide encoding an RGN polypeptide and guide RNA and cultured under conditions that permit expression of the RGN polypeptide and guide RNA. Accordingly, methods for making RGN polypeptides or RGN ribonucleoprotein complexes are provided. Such methods include: culturing a cell comprising a nucleotide sequence encoding an RGN polypeptide under conditions in which the RGN polypeptide (and in some embodiments, a guide RNA) is expressed, and in some embodiments, culturing comprises treating Guide RNA to encode the nucleotide sequence of the cell. The RGN polypeptide or RGN ribonucleoprotein can then be purified from a lysate of the cultured cells.
[0120] Methods for purifying RGN polypeptides or RGN ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., particle size screening and / or affinity chromatography, 2D-PAGE, HPLC , reverse phase chromatography, immunoprecipitation). In some specific methods, the RGN polypeptide is recombinantly produced and includes purification tags to facilitate its purification, including but not limited to: glutathione-S-transferase (GST), chitin binding protein (CBP) , maltose binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 10xHis, biotin carboxyl carrier protein (BCCP) and calmodulin. Typically, the tagged RGN polypeptide or RGN ribonucleoprotein complex is purified using immobilized metal affinity chromatography. It will be appreciated that other similar methods known in the art, including other forms of chromatography or eg immunoprecipitation, may be used alone or in combination.
[0121] An "isolated" or "purified" polypeptide, or biologically active portion thereof, is substantially or substantially free of components that normally accompany or interact with the polypeptide found in its naturally occurring environment. Thus, an isolated or purified polypeptide is substantially free of other cellular material or culture medium when produced by recombinant techniques, or chemical precursors or other chemicals when chemically synthesized. Proteins substantially free of cellular material include protein preparations having less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of contaminating protein. When recombinantly producing a protein of the invention, or a biologically active portion thereof, an optimal medium exhibits less than about 30%, 20%, 10%, 5%, or 1% (by dry weight) of the chemical precursor or compound of interest Non-protein chemicals.
[0122] Certain methods provided herein for binding and / or cleaving a target sequence of interest involve the use of in vitro assembled RGN ribonucleoprotein complexes. In vitro assembly of RGN ribonucleoprotein complexes can be performed using methods known in the art, wherein an RGN polypeptide is contacted with a guide RNA under conditions that permit binding of the RGN polypeptide to the guide RNA. As used herein, "contact, contacting," "contacted" refers to bringing together a composition for a desired reaction under conditions suitable for carrying out the desired reaction. RGN polypeptides can be purified from biological samples, cell lysates or culture media, produced via in vitro transformation, or chemically synthesized. Guide RNAs can be purified from biological samples, cell lysates or culture media, transcribed in vitro, or chemically synthesized. The RGN polypeptide and guide RNA can be contacted in solution (eg, buffered saline) to allow in vitro assembly of the RGN ribonucleoprotein complex. VII. Methods of Binding, Cleaving, or Modifying Target Sequences
[0123] The present disclosure provides methods for binding, cleaving and / or modifying a target nucleotide sequence of interest. The method comprises delivering a polynucleotide comprising at least one guide RNA or encoding the at least one guide RNA, and at least one RGN polypeptide or encoding the at least one RGN polypeptide to the target sequence or a cell, organelle or embryo comprising the target sequence polynucleotide systems. In some of these embodiments, the RGN comprises SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117 , 123 or the amino acid sequence of 570-579 or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence comprising SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90 , 97, 104, 111, 118 or 124 nucleotide sequence or an active variant or fragment thereof. In specific embodiments, guide RNA comprises tracrRNA, and this tracrRNA comprises SEQ ID NO:3,10,17,24,31,38,45,52,58,65,72,78,85,91,98,105, The nucleotide sequence of 112, 119 or 125 or an active variant or fragment thereof. The guide RNA of this system can be single guide RNA or double guide RNA. The RGN of the system can be a nuclease-free RGN, have nickase activity, or can be a fusion polypeptide. In some embodiments, the fusion polypeptide includes a base editing polypeptide, eg, cytidine deaminase or adenosine deaminase. In other embodiments, the RGN fusion protein includes reverse transcriptase. In other embodiments, the RGN fusion protein comprises a polypeptide that adds a member of a functional nucleic acid repair complex, e.g., nucleotide excision repair (NER) or transcription-coupled-nucleotide excision repair (TC-NER) members of the pathway (Wei et al., 2015, PNAS USA112(27):E3495-504; Troelstra et al., 1992, Cell71:939-953; Marnef et al., 2017, J Mol Biol429(9):1277-1288), As described in U.S. Provisional Patent Application No. 62 / 966,203, filed January 27, 2020, and incorporated herein by reference in its entirety. In some embodiments, the RGN fusion protein includes CSB (van den Boom et al., 2004, J Cell Biol 166(1): 27-36; van Gool et al., 1997, EMBO J16(19): 5955-65; examples thereof As shown in SEQ ID NO: 608), CSB is a member of the TC-NER (nucleotide excision repair) pathway and functions in joining other members. In additional embodiments, the RGN fusion protein comprises the active domain of CSB, for example, the acidic domain of CSB comprising amino acid residues 356-394 of SEQ ID NO: 608 (Teng et al., 2018, Nat Commun 9(1) :4115).
[0124] In a particular embodiment, the RGN and / or guide RNA is introduced into the cell, organelle or Embryos are allogeneic.
[0125] In those embodiments wherein the method comprises delivering a polynucleotide encoding a guide RNA and / or RGN polypeptide, the cells or embryos can then be cultured under conditions in which the guide RNA and / or RGN polypeptide are expressed. In various embodiments, the method comprises contacting a target sequence with an RGN ribonucleoprotein complex. The RGN ribonucleoprotein complex may comprise RGN that is inactive or has nickase activity. In some embodiments, the RGN of the ribonucleoprotein complex is a fusion polypeptide comprising a base editing polypeptide. In certain embodiments, the method comprises introducing an RGN ribonucleoprotein complex into a cell, organelle or embryo comprising a target sequence. The RGN ribonucleoprotein complex can be one that has been purified from a biological sample, produced recombinantly and subsequently purified, or assembled in vitro as described herein. In those embodiments wherein the RGN ribonucleoprotein complex in contact with the target sequence or cell, organelle or embryo has been assembled in vitro, the method may further comprise the complex in contact with the target sequence, cell, cellular In vitro assembly prior to organ or embryo contact.
[0126] Purified or in vitro assembled RGN ribonucleoprotein complexes can be introduced into cells, organelles or embryos using any method known in the art, including but not limited to electroporation. Alternatively, RGN polypeptides and / or polynucleotides encoding or comprising guide RNAs can be introduced into cells, organelles or embryos using any method known in the art (eg, electroporation).
[0127] When delivered to or in contact with the target sequence or a cell, organelle or embryo comprising the target sequence, the guide RNA directs the RGN to bind the target sequence in a sequence-specific manner. In those embodiments wherein the RGN has nuclease activity, the RGN polypeptide cleaves the target sequence of interest upon binding. The target sequence can then be modified via endogenous repair mechanisms such as non-homologous end joining or homology-mediated repair with a provided donor polynucleotide.
[0128] Methods of measuring binding of RGN polypeptides to target sequences are known in the art and include chromatin immunoprecipitation assays, gel shift shift assays, DNA pull-down assays, reporter assays, microplate capture and detection assays. Likewise, methods of measuring cleavage or modification of a target sequence are known in the art and include in vitro or in vivo cleavage assays, in which cleavage assays are performed with or without appropriate labels (e.g., radioisotopes, fluorescent substances) attached to In the case of target sequences to facilitate detection of degradation products, use PCR, sequencing or gel electrophoresis to confirm cleavage. Alternatively, a nick-triggered exponential amplification reaction (NTEXPAR) assay can be used (see, eg, Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage can be assessed using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).
[0129] In some embodiments, the methods involve the use of a single type of RGN mated to more than one guide RNA. The one or more guide RNAs can target different regions of a single gene, or can target multiple genes.
[0130] In those embodiments in which no donor polynucleotide is provided, the double-stranded break introduced by the RGN polypeptide can be repaired by the non-homologous end joining (NHEJ) repair process. Due to the error-prone nature of NHEJ, repair of double-strand breaks can result in modifications to the targeted sequence. As used herein, "modification" with respect to a nucleic acid molecule refers to a change in the nucleotide sequence of the nucleic acid molecule, which may be a deletion, insertion, or substitution of one or more nucleotides, or a combination thereof. Modification of a target sequence can result in expression of an altered protein product or inactivation of a coding sequence.
[0131] In those embodiments in which a donor polynucleotide is present, the donor sequence in the donor polynucleotide may be integrated into the target nucleotide sequence or Exchange with a target nucleotide sequence, resulting in the introduction of an exogenous donor sequence. Thus, the donor polynucleotide includes the donor sequence desired to be introduced into the target sequence of interest. In some embodiments, the donor sequence alters the original target nucleotide sequence such that the newly integrated donor sequence will not be recognized and cleaved by the RGN. Integration of the donor sequence can be enhanced by including in the donor polynucleotide contiguous sequences having substantial sequence identity with the sequences flanking the target nucleotide sequence, referred to herein as "homology arms". ”, to allow a homology-guided repair process. In some embodiments, the homology arms have a length of at least 50 base pairs, at least 100 base pairs, and up to 2000 base pairs or more, and are at least 50 base pairs in length with respect to the target nucleotide sequence Corresponding sequences within have at least 90%, at least 95% or greater sequence identity.
[0132] In those embodiments wherein the RGN polypeptide introduces a double-stranded staggered break, the donor polynucleotide may include a donor sequence flanked by compatible overhangs to allow for repair of the double-stranded break by non-homologous The repair process directly joins the donor sequence to the cleaved target nucleotide sequence including the overhang.
[0133] In those embodiments where the method involves the use of an RGN that is a nickase (i.e., capable of cutting only a single strand in a double-stranded polynucleotide), the method may include introducing target sequences that target the same or overlap and two RGN nickases that cleave different strands of the polynucleotide. For example, an RGN nickase that cleaves only the positive (+) strand of a double-stranded polynucleotide can be introduced together with a second RGN nickase that cleaves only the negative (-) strand of the double-stranded polynucleotide.
[0134] In various embodiments, a method for binding a target nucleotide sequence and detecting the target sequence is provided, wherein the method comprises at least one guide RNA or a polynucleotide encoding the at least one guide RNA, and at least one RGN polypeptide or a polynucleotide encoding the at least one RGN polypeptide is introduced into the cell, organelle or embryo; expressing the guide RNA and / or the RGN polypeptide (if the coding sequence is introduced), wherein the RGN polypeptide is nuclease-free active RGN and further comprising a detectable label, and the method further comprises detecting the detectable label. The detectable label can be fused to the RGN as a fusion protein (eg, a fluorescent protein), or can be a small molecule conjugated to or incorporated into an RGN polypeptide that can be detected visually or by other means.
[0135] Also provided herein are methods for modulating the expression of a target sequence or gene of interest under the control of the target sequence. The method comprises introducing at least one guide RNA or a polynucleotide encoding the at least one guide RNA, and at least one RGN polypeptide or a polynucleotide encoding the at least one RGN polypeptide into a cell, an organelle or an embryo; expressing the guide RNA and / or or an RGN polypeptide (if a coding sequence is introduced), wherein the RGN polypeptide is RGN without nuclease activity. In some of these embodiments, the nuclease-free RGN is a fusion protein comprising an expression regulator domain (ie, an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repressor domain) as described herein.
[0136] The present disclosure also provides methods for binding and / or modifying a target nucleotide sequence of interest. The method includes delivering a system comprising at least one guide RNA or a polynucleotide encoding the at least one guide RNA, and at least one fusion polypeptide to the target sequence or to a cell, organelle or embryo comprising the target sequence, the at least A fusion polypeptide comprises the RGN of the present invention and a base editing polypeptide (eg, cytidine deaminase or adenosine deaminase) or a polynucleotide encoding the fusion polypeptide.
[0137] One of ordinary skill in the art will appreciate that any of the methods disclosed herein can be used to target a single target sequence or multiple target sequences. Accordingly, these methods include the use of a single RGN polypeptide in combination with multiple different guide RNAs that can target a single gene and / or multiple different sequences within multiple genes. Also included herein are methods wherein multiple different guide RNAs are introduced in combination with multiple different RGN polypeptides. These guide RNAs and guide RNA / RGN polypeptide systems can target a single gene and / or multiple different sequences within multiple genes.
[0138] In one aspect, the invention provides a kit comprising any one or more of the elements disclosed in the methods and compositions described above. In some embodiments, the kit includes a vector system and instructions for using the kit. In some embodiments, the vector system includes (a) a first regulatory element that is associated with a DNA sequence encoding a crRNA sequence and one or more insertion sites for inserting a guide sequence upstream of the encoded crRNA sequence The point is operably linked, wherein when expressed, the guide sequence directs the sequence-specific binding of the RGN complex to the target sequence in a eukaryotic cell, wherein the RGN complex comprises an RGN complex complexed with a guide RNA polynucleotide an enzyme; and / or (b) a second regulatory element, the second regulatory element is operably linked to an enzyme coding sequence encoding the RGN enzyme including a nuclear localization sequence. These elements may be provided individually or in combination and may be provided in any suitable container, eg a vial, bottle or tube.
[0139] In some embodiments, the kit includes instructions in one or more languages. In some embodiments, a kit includes one or more reagents for use in a method utilizing one or more elements described herein. Reagents may be provided in any suitable container. For example, a kit can provide one or more reaction or storage buffers. Reagents may be provided in a form useful for a particular assay, or in a form that requires the addition of one or more additional components prior to use (eg, in concentrated or lyophilized form). The buffer can be any buffer including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is basic. In some embodiments, the buffer has a pH of about 7 to about 10.
[0140] In some embodiments, the kit includes one or more oligonucleotides corresponding to the leader sequence for insertion into the vector so as to operably link the leader sequence and the regulatory elements. In some embodiments, the set includes homologous recombination template polynucleotides. In one aspect, the invention provides a method for using one or more elements of an RGN system. The RGN system of the present invention provides an efficient means for modifying target polynucleotides. The RGN system of the present invention has a wide variety of utilities, including modification (eg, deletion, insertion, translocation, inactivation, activation, base editing) of target polynucleotides in a variety of cell types. Thus, the RGN system of the present invention has wide applications in, for example, gene therapy, drug screening, disease diagnosis and prognosis. Exemplary RGN systems or RGN complexes include an RGN enzyme complexed with a guide sequence that hybridizes to a target sequence within a target polynucleotide. VIII. Target Polynucleotides
[0141] In one aspect, the invention provides a method of modifying a target polynucleotide in a eukaryotic cell, which may be in vivo, ex vivo or in vitro. In some embodiments, the method comprises sampling a cell or population of cells from a human or non-human animal or plant, including microalgae, and modifying the cell or cells. Culturing can occur at any stage ex vivo. The cell or cells can even be reintroduced into non-humanoid animals or plants, including microalgae.
[0142] Using natural variability, plant breeders combine the most useful genes to obtain desired qualities such as yield, quality, uniformity, hardiness, and pest resistance. These desired qualities also include growth, day length preference, temperature requirements, date of initiation of floral or reproductive development, fatty acid content, insect resistance, disease resistance, nematode resistance, fungal resistance, herbicide resistance, Tolerance to various environmental factors including drought, heat, humidity, cold, wind and adverse soil conditions including high salinity. Sources of such useful genes include natural or exotic varieties, heirloom varieties, wild plant relatives, and induced mutations such as treatment of plant material with mutagens. Using the present invention, plant breeders are provided with new tools for inducing mutations. Accordingly, one of ordinary skill in the art can analyze gene bodies for sources of useful genes and employ the present invention to induce increases in useful genes in varieties with desired characteristics or traits with more precision than previous mutagens , and thus accelerate and improve plant breeding programmes.
[0143] The target polynucleotide of the RGN system can be any polynucleotide that is endogenous or exogenous to the eukaryotic cell. For example, the target polynucleotide may be a polynucleotide present in the nucleus of a eukaryotic cell. A target polynucleotide can be a sequence encoding a gene product (eg, a protein) or a non-coding sequence (eg, a regulatory polynucleotide or junk DNA). Without wishing to be bound by theory, the target sequence should be associated with a PAM (prospacer adjacent motif); that is, a short sequence recognized by the RGN system. The exact sequence and length requirements of this PAM vary with the RGN used, but the PAM is typically a 2-5 base pair sequence adjacent to the prospacer (ie, target sequence).
[0144] The target polynucleotides of the RGN system can include a number of disease-associated genes and polynucleotides as well as genes and polynucleotides associated with signaling biochemical pathways. Examples of target polynucleotides include sequences associated with signaling biochemical pathways, eg, genes or polynucleotides associated with signaling biochemical pathways. Examples of target polynucleotides include genes or polynucleotides associated with diseases. A "disease-associated" gene or polynucleotide means any gene that produces a transcriptional or translation product at abnormal levels or in an abnormal form in cells obtained from diseased tissue as compared to non-disease-controlled tissues or cells or polynucleotides. It can be a gene that becomes expressed at an abnormally high level; it can be a gene that becomes expressed at an abnormally low level, where the altered expression correlates with the onset and / or progression of the disease. A disease-associated gene also refers to a gene that has a mutation or genetic variation that is directly responsible for the cause of a disease (eg, a causal mutation) or that is linked to a gene that is responsible for the cause of a disease (eg, a causal mutation) unbalanced. Products of transcription or translation may be known or unknown, and may also be at normal or abnormal levels. In some embodiments, the disease can be an animal disease. In some embodiments, the disease can be an avian disease. In other embodiments, the disease can be a mammalian disease. In a further embodiment, the disease can be a human disease. Examples of human disease-associated genes and polynucleotides are available from the World Wide Web at the National Library of Medicine (Bethesda, Md.) National Center for Biotechnology Information (Maryland Bethesda) and the Johns Hopkins University (Baltimore, MD) McKusick-Nathans Institute of Genetic Medicine.
[0145] While the RGN system is exceptionally useful for its relative ease in targeting gene body sequences of interest, the question of how the RGN resolves causal mutations remains. One method is to combine RGN (preferably, an inactive or nickase variant of RGN) with a base editing enzyme or a base editing enzyme (e.g., cytidine deaminase or adenosine deaminase base editor) A fusion protein was generated between the active domains of ® (US Patent No. 9,840,699, incorporated herein by reference). In some embodiments, the methods comprise contacting a DNA molecule with (a) a fusion protein comprising an RGN of the invention and a base editing polypeptide such as deaminase; and (b) contacting the fusion protein of (a) Contacting a gRNA targeting a target nucleotide sequence of a DNA strand; wherein the DNA molecule is contacted with the fusion protein and the gRNA in an effective amount and under conditions suitable for nucleobase deamination. In some embodiments, the target DNA sequence includes a sequence associated with a disease or disorder, and wherein deamination of the nucleobase results in a sequence not associated with the disease or disorder. In some embodiments, the target DNA sequence is located in an allele of a crop plant, wherein a particular allele for a trait of interest results in a plant of less agronomic value. Deamination of this nucleobase results in alleles that improve plant traits and increase the agronomic value of the plant.
[0146] In some embodiments, the DNA sequence includes a TàC or AàG point mutation associated with a disease or disorder, and wherein deamination of the mutant C or G base results in a sequence not associated with the disease or disorder. In some embodiments, deamination corrects a point mutation in a sequence associated with the disease or disorder.
[0147] In some embodiments, the sequence associated with the disease or disorder encodes a protein, and wherein the deamination introduces a stop codon into the sequence associated with the disease or disorder, resulting in truncation of the encoded protein. In some embodiments, the contacting is in an individual susceptible to, having, or diagnosed with a disease or disorder. In some embodiments, the disease or disorder is a disease associated with a point mutation or a single base mutation in a gene body. In some embodiments, the disease is a genetic disease, cancer, metabolic disease, or lysate storage disease. IX. Pharmaceutical composition and treatment method
[0148] A pharmaceutical composition is provided, the pharmaceutical composition comprising: the RGN polypeptide disclosed in the present invention and its active variant or fragment and the polynucleotide encoding the RGN polypeptide and its active variant or fragment, the gRNA disclosed in the present invention or The polynucleotide encoding the gRNA, the system disclosed in the present invention, or a cell comprising the RGN polypeptide or RGN encoding polynucleotide, gRNA or gRNA encoding polynucleotide, or any of the RGN system and pharmaceutically acceptable carrier.
[0149] A pharmaceutical composition is a composition that is used to prevent, reduce, cure or treat a target disease or disease, and the composition includes an active ingredient (that is, RGN polypeptide, RGN-encoded polynucleotide, gRNA, gRNA-encoded polynucleotide nucleotides, RGN system, or cells comprising any of these) and a pharmaceutically acceptable carrier.
[0150] As used herein, "pharmaceutically acceptable carrier" refers to a carrier that does not cause significant irritation to the organism and does not eliminate the active ingredient (that is, RGN polypeptide, RGN encoding polynucleotide, gRNA, gRNA encoding polynucleotide , RGN systems, or cells comprising any of these) for the activity and properties of materials. The carrier must be of sufficiently high purity and sufficiently low toxicity to render it suitable for administration to the individual being treated. The carrier may be inert, or it may have medicinal benefits. In some embodiments, a pharmaceutically acceptable carrier includes one or more compatible solid or liquid fillers, diluents or encapsulating substances suitable for administration to a human or other vertebrate. In some embodiments, the pharmaceutically acceptable carrier is not naturally occurring. In some embodiments, the pharmaceutically acceptable carrier is not found essentially with the active ingredient.
[0151] The pharmaceutical compositions used in the methods disclosed herein can be formulated with suitable carriers, excipients, and other agents that provide for suitable transfer, delivery, tolerability, and the like. Numerous appropriate formulations are known to those of ordinary skill in the art. See, eg, Remington, The Science and Practice of Pharmacy (21st ed. 2005). Suitable formulations include, for example: powders, pastes, ointments, jellies, waxes, oils, lipids, lipid-containing (cationic or anionic) vesicles (e.g., LIPOFECTIN vesicles), lipid nanoparticles, DNA conjugates, anhydrous absorption paste, water-oil emulsion and water-oil emulsion, emulsions carbowax (polyethylene glycol of various molecular weights), semi-solid gel (semi-solid gel) and containing A semisolid mixture of carbomer waxes. Pharmaceutical compositions for oral or parenteral use may be prepared in unit dosage form adapted to suit the dose of the active ingredient. These unit dosage forms include, for example, tablets, pills, capsules, injections (ampoules), suppositories and the like.
[0152] In some embodiments wherein the cells comprising or using the RGN, gRNA, RGN system disclosed in the present invention or the polynucleotide encoding the RGN, gRNA, RGN system are administered to an individual, these cells are in a pharmaceutically acceptable The carrier is administered together as a suspension. Those of ordinary skill in the art will recognize that a pharmaceutically acceptable carrier to be used in a cellular composition will not include buffers, compounds, Cryopreservatives, preservatives, or other preparations. Formulations including cells can include, for example, osmotic buffers that allow cell membranes to maintain integrity, and optionally nutrients that upon administration maintain cell viability or enhance engraftment. Such formulations and suspensions are known to those of ordinary skill in the art and / or can be adapted for use with the cells disclosed herein using routine experimentation.
[0153] Cellular compositions may also be emulsified or presented as ribosomal compositions, provided that the emulsification procedure does not adversely affect cell viability. The cells and any other active ingredients can be mixed with excipients that are pharmaceutically acceptable and compatible with the active ingredients, and in amounts suitable for use in the methods of treatment disclosed herein.
[0154] Additional agents included in the cell composition may include pharmaceutically acceptable salts of the composition therein. Pharmaceutically acceptable salts include, for example, acid addition salts (formed with free amine groups of the polypeptide) with inorganic acids such as hydrochloric acid or phosphoric acid, or with organic acids such as acetic acid, tartaric acid, mandelic acid and the like. Salts formed with free carboxyl groups may also be derived, for example, from inorganic bases such as sodium, potassium, ammonium, calcium or iron hydroxides, and such as isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, common Organic bases such as lucaine.
[0155] Physiologically tolerable and pharmaceutically acceptable carriers are known in the art. Exemplary liquid carriers are sterile aqueous solutions, which are free of materials other than the active ingredient and water, or which contain a buffered solution such as sodium phosphate, saline, or both at physiological pH (e.g., phosphate-buffered saline ). Still further, aqueous carriers can contain one or more buffer salts, as well as salts such as sodium and potassium chloride, dextrose, polyethylene glycol and other solutes. Liquid compositions may also contain a liquid phase in addition to and excluding water. Examples of such additional liquid phases are glycerol, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of active compound used in a cellular composition effective in the treatment of a particular disorder or condition may depend on the nature of the disorder or condition and can be determined by standard clinical techniques.
[0156] The RGN polypeptides, guide RNAs, RGN systems disclosed in the present invention, or polynucleotides encoding the RGN polypeptides, guide RNAs, RGN systems can be used in the form of, for example, carriers, solvents, stabilizers, etc., depending on the specific mode and dosage form of administration. , adjuvants, diluents and other pharmaceutically acceptable excipients. In some embodiments, these pharmaceutical compositions are formulated to achieve a physiologically compatible pH; and depending on the formulation and route of administration, range from a pH of about 3 to a pH of about 11, from about pH 3 to about pH 7. In some embodiments, the pH can be adjusted to a range of about pH 5.0 to about pH 8. In some embodiments, a composition may include a therapeutically effective amount of at least one compound described herein, together with one or more pharmaceutically acceptable excipients. In some embodiments, a composition includes a combination of compounds described herein, or includes a second active ingredient useful in treating or preventing bacterial growth (such as, without limitation, an antibacterial or antimicrobial agent), or includes an agent of the present disclosure The combination.
[0157] For example, suitable excipients include carrier molecules including large, slowly metabolized macromolecules such as proteins, polysaccharides, polylactic acid, polyglycolic acid, polymeric amino acids, amino acid copolymers and inactive virus particles. Other exemplary excipients may include antioxidants (such as, without limitation, ascorbic acid), chelating agents (such as, without limitation, EDTA), carbohydrates (such as, without limitation, dextrin, hydroxyalkyl cellulose, and hydroxyalkyl methylcellulose), stearic acid, liquids (such as, without limitation, oils, water, saline, glycerin, and ethanol), wetting or emulsifying agents, pH buffering substances, and the like.
[0158] In some embodiments, the formulations are provided in unit-dose or multi-dose containers (e.g., sealed ampoules and vials) and may be stored in a lyophilized (freeze-dried) condition for immediate use. Before, add a sterile liquid carrier (eg, saline, water for injections, semi-liquid foam or gel). Solutions and suspensions for immediate injection can be prepared from sterile powders, granules and tablets of the kind previously described. In some embodiments, the active ingredient is dissolved in a buffered liquid solution that is frozen in unit-dose or multi-dose containers and then thawed for injection or kept / stabilized frozen until use.
[0159] The one or more therapeutic agents can be contained in a controlled release system. In order to prolong the action of a drug, it is often desirable to slow the absorption of the drug by subcutaneous, intrathecal, or intramuscular injection. This can be accomplished through the use of liquid suspensions with poorly water soluble crystalline or non-crystalline materials. The rate of absorption of the drug then depends upon its rate of dissolution which, in turn, may depend upon crystal size and crystalline form. Alternatively, delayed absorption of a parenterally administered drug is accomplished by dissolving or suspending the drug in an oil vehicle. In some embodiments, the use of long-term sustained release implants is particularly suitable for the treatment of chronic conditions. Long term sustained release implants are known to those of ordinary skill in the art. Provided herein are methods for treating disease in an individual in need thereof. The method comprises introducing an effective amount of the RGN polypeptide disclosed in the present invention or its active variant or fragment or the polynucleotide encoding the RGN polypeptide or its active variant or fragment, the gRNA disclosed in the present invention or the polynucleotide encoding the gRNA , the RGN system disclosed herein, or cells modified by or comprising any of these compositions are administered to an individual in need thereof.
[0160] In some embodiments, treatment includes administration of the RGN polypeptide, gRNA, or RGN system disclosed in the present invention, or polynucleotide(s) encoding the RGN polypeptide, gRNA, or RGN system In vivo gene editing. In some embodiments, the treatment includes in vitro gene editing, the cells of which are the RGN polypeptide, gRNA, or RGN system disclosed in the present invention, or the polynucleotide(s) encoding the RGN polypeptide, gRNA, or RGN system The gene is modified in vitro, and the modified cells are then administered to the individual. In some embodiments, the genetically modified cells are derived from the individual to whom the modified cells are later administered, and the transplanted cells are referred to herein as autologous. In some embodiments, the genetically modified cells are derived from a different individual (i.e., the donor) of the same species as the individual (i.e., the recipient) to which the modified cells are administered, and the transplanted cells are described herein are called heterogeneous. In some examples described herein, the cells may be expanded in culture prior to administration to an individual in need thereof.
[0161] In some embodiments, the disease to be treated with the compositions disclosed herein is a disease treatable with immunotherapy (eg, with chimeric antigen receptor (CAR) T cells). Such diseases include, but are not limited to, cancer. In some embodiments, the disease to be treated with the compositions disclosed herein is associated with a causal mutation. As used herein, a "causal mutation" refers to a specific nucleotide, nucleotides or sequence of nucleotides in a gene body that contributes to the severity or presence of a disease or disorder in an individual. Correction of the causal mutation results in amelioration of at least one symptom caused by the disease or disorder. In some embodiments, the causal mutation is adjacent to a PAM site recognized by an RGN disclosed herein. The causal mutation can be corrected with the RGN disclosed herein or a fusion polypeptide comprising the RGN disclosed herein and a base-edited polypeptide (ie, a base editor). Non-limiting examples of diseases associated with causal mutations include cystic fibrosis, Hurler's syndrome, Friedreich's Ataxia, Huntington's Disease, and sickle cell disease . In some embodiments, the diseases to be treated with the RGNs of the present disclosure are the diseases listed in Table 11. Additional non-limiting examples of disease-associated genes and mutations are available on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine at Johns Hopkins University (Baltimore, MA). Medicine, Johns Hopkins University (Baltimore, Md.) and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Md.) obtain.
[0162] As used herein, "treatment" or "treatment", or "alleviation" or "improvement" may be used interchangeably. These terms refer to methods used to obtain a beneficial or desired result, including, but not limited to, a therapeutic benefit and / or a prophylactic benefit. By therapeutic benefit is meant any treatment-related improvement in or effect on one or more diseases, disorders or symptoms under treatment. For prophylactic benefit, the composition may be administered to an individual at risk of developing a particular disease, disorder or symptom, or to an individual reporting one or more physiological symptoms of a disease, even though the disease, disorder or symptom may not have Show signs.
[0163] The term "effective amount" or "therapeutically effective amount" refers to an amount of an agent sufficient to achieve a beneficial or desired result. A therapeutically effective amount may vary depending on one or more of the individual and disease condition being treated, the weight and age of the individual, the severity of the disease condition, the mode of administration, etc., as is generally known in the art This can easily be determined. The particular dosage may vary depending on one or more of the particular agent selected, the dosing regimen to be followed, whether it is administered in combination with other compounds, the timing of administration, and the delivery system in which it is carried.
[0164] The term "administering" refers to placing an active ingredient in an individual by a method or route that results in the introduced active ingredient being at least partially localized at a desired site (e.g., a site of injury or repair), such that produce desired utility(s). In those embodiments in which cells are administered, the cells may be administered by any appropriate route that results in delivery to the desired location in the individual, wherein at least a portion of the transplanted cells, or components of the cells, remain viable. Cell survival after administration to an individual can be as short as a few hours (eg, twenty-four hours), to a few days, to as long as several years, or even the lifetime of the patient, ie, long-term engraftment. For example, in some aspects described herein, an effective amount of photoreceptor cells or retinal precursor cells is administered via a systemic route of administration (eg, intraperitoneal or intravenous).
[0165] In some embodiments, administering comprises administration by viral delivery. In some embodiments, administering comprises administration by electroporation. In some embodiments, administering comprises administration by nanoparticle delivery. In some embodiments, administering comprises administration via ribosomal delivery. Any effective route of administration can be used to administer an effective amount of a pharmaceutical composition described herein. In some embodiments, administering comprises administering by a method selected from the group consisting of: intravenously, subcutaneously, intramuscularly, orally, rectally, by aerosol, parenterally, Ocularly, pulmonary, transdermally, vaginally, aurally, nasally and by external administration, or any combination thereof. In some embodiments, for delivery of cells, administration by injection or perfusion is used.
[0166] As used herein, the term "individual" refers to any individual for whom diagnosis, treatment, or therapy is desired. In some embodiments, the individual is an animal. In some embodiments, the individual is a mammal. In some embodiments, the individual is human.
[0167] The efficacy of treatment can be determined by a skilled clinician. However, if any or all of the signs or symptoms of the disease or disorder are altered in a beneficial manner (e.g., by at least 10%), or other clinically accepted symptoms or markers of the disease are ameliorated or ameliorated, then the treatment is considered considered "effective treatment". Efficacy can also be measured by the fact that the subject does not develop exacerbations, as assessed by hospitalization, or does not require medical intervention (eg, the progression of the disease is stopped or at least slowed). Methods for measuring these indicators are known to those of ordinary skill in the art. Treatment includes: (1) inhibiting the disease, eg, halting or slowing the progression of symptoms; or (2) slowing the disease, eg, causing regression of symptoms; and (3) preventing or reducing the likelihood of developing symptoms. A. Modification of causal mutations using base editing
[0168] An example of a hereditary disease that can be corrected using methods dependent on the RGN-base editor fusion proteins of the invention is Hurler's disease. Hurler's disease (also known as MPS-1) is the result of alpha-L-iduronidase (IDUA) deficiency, leading to a disorder characterized at the molecular level by the accumulation of heparan sulfate and dermatan sulfate in solution Solution storage disease. The disease is usually an inherited genetic disorder caused by mutations in the IDUA gene encoding alpha-L-iduronidase. Common IDUA mutations are W402X and Q70X, both of which are nonsense mutations that lead to premature termination of translation. Such mutations are well resolved by precision genome editing (PGE) methods, since restoration of a single nucleotide (e.g., by base editing methods) will restore the wild-type coding sequence and result in the loss of the inherited locus. Protein expression controlled by endogenous regulatory mechanisms. In addition, since heterozygotes are known to be asymptomatic, PGE therapy targeting one of these mutations would be useful in most patients with the disease since only one of the mutated alleles would need to be corrected (Bunge et al. (1994) Hum. Mol. Genet. 3(6):861-866, incorporated herein by reference).
[0169] Current treatments for Hurler syndrome include enzyme replacement therapy and bone marrow transplantation (Vellodi et al. ((1997)) Arch. Dis. Child. 76(2):92-99; Peters et al. ((1998)) Blood 91(7):2601-2608, incorporated herein by reference). While enzyme replacement therapy has had a dramatic impact on survival and quality of life for patients with Hurler syndrome, the approach requires expensive and time-consuming weekly infusions. Additional methods include delivering the IDUA gene on an expression vector or inserting the gene into a highly expressed locus, eg, the locus of serum albumin (US Patent No. 9,956,247, incorporated herein by reference). However, these methods cannot restore the original IDUA locus to the correct coding sequence. Genome editing strategies may have many advantages, most notably that regulation of gene expression will be controlled by natural mechanisms present in healthy individuals. In addition, the use of base editing does not necessarily cause double-stranded DNA breaks, which may lead to large-scale chromosomal rearrangements, cell death, or carcinogenesis due to disruption of tumor suppressor mechanisms. A general strategy can be directed to use the RGN base editor fusion proteins of the invention to target and correct certain disease-causing mutations in the human genome. It will be appreciated that similar approaches can also be pursued to target diseases correctable by base editing. It should also be further understood that similar methods can be used to target disease-causing mutations in other species (especially common household pets or livestock) using the RGN of the present invention. Common household pets and livestock include dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, fish (including salmon), and shrimp. B. Modification of causal mutations by targeted deletions
[0170] The RGNs of the invention are also useful in the treatment of humans with more complex causal mutations. For example, some diseases such as Friedrich's movement disorder and Huntington's disease are the result of a marked increase in repeats of three nucleotide motifs at specific regions of the gene, which can affect whether the expressed protein functions or is expressed ability. Friedrich's movement disorder (FRDA) is an autosomal recessive disorder that causes progressive degeneration of spinal nerve tissue. Reduced levels of frataxin (FXN) protein in mitochondria cause oxidative damage at the cellular level and iron deficiency. Reduced FXN expression has been associated with GAA triplet expansion within intron 1 of the somatic and germline FXN genes. In FRDA patients, GAA repeats usually consist of more than 70, sometimes more than 1000 (most commonly 600–900) triplets, whereas unaffected individuals have about 40 or fewer repeats (Pandolfo et al (2012) Handbook of Clinical Neurology 103:275-294; Campuzano et al. (1996) Science 271:1423-1427; Pandolfo (2002) Adv. Exp. Med. Biol. 516:99-118; fully incorporated by reference into this article).
[0171] Expansion of the trinucleotide repeat causing Friedrich's movement disorder (FRDA) occurs in a defined genetic locus within the FXN gene, referred to as the region of FRDA instability. RNA-guided nucleases (RGNs) can be used to excise regions of instability in FRDA patient cells. This approach requires: 1) a guide RNA sequence and RGN that can be programmed to target an allele in the human genome; and 2) a delivery method for the RGN and guide sequence. In particular, many nucleases used for genome editing (e.g., the commonly used Cas9 nuclease (SpyCas9)) is too large to be packaged into an adeno-associated virus (AAV) vector. This makes the method using SpCas9 more difficult.
[0172] Certain RNA-guided nucleases of the invention are well suited for packaging into AAV vectors along with guide RNAs. Packaging two guide RNAs may require a second vector, but this approach is still advantageous over methods that may require larger nucleases such as SpCas9, which may require unraveling of the protein sequence between the two vectors. The present invention encompasses strategies using RGNs of the present invention in which regions of gene body instability are removed. This strategy is applicable to other diseases and conditions with a similar genetic basis, such as Huntington's disease. Similar strategies using the RGNs of the present invention can also be applied to non-human animals of agronomic or economic importance including dogs, cats, horses, pigs, cattle, sheep, chickens, donkeys, snakes, ferrets, and fish including salmon. and shrimp). C. Modification of causal mutations by targeted mutagenesis
[0173] The RGNs of the invention can also introduce disruptive mutations that can lead to beneficial effects. Genetic defects in the genes that encode heme, specifically the beta globin chain (HBB gene), can be the cause of many diseases called hemepathies, including sickle cell anemia and thalassemia.
[0174] In adults, heme is a heterotetramer comprising two alpha globoid chains and two beta globoid chains and four heme groups. In adults, the α2β2 tetramer is known as heme A (HbA) or adult heme. Normally, alpha and beta globin chains are synthesized in a ratio of approximately 1:1, and this ratio appears to be critical for heme and red blood cell (RBC) stability. In the developing fetus, a different form of heme (fetal heme (HbF)) is produced that has a higher binding affinity for oxygen than heme A, allowing oxygen to be delivered to the infant's system via the mother's bloodstream. Fetal heme also contains two alpha globin chains, but instead of adult beta-globin chains, it has two fetal gamma-globin chains (ie, fetal heme is α2γ2). The regulation of the switch from γ-globin production to β-globin production is quite complex and mainly involves the downregulation of γ-globin transcription and the simultaneous up-regulation of β-globin transcription. At about 30 weeks of gestation, the synthesis of gamma globulin in the fetus begins to decline, while the production of beta globulin increases. By about 10 months of age, neonatal hemoglobin is nearly all α2β2, although some HbF persists into adulthood (approximately 1–3% of total hemoglobin). In most patients with hemoglobinopathy, the gene encoding gamma globulin is still present, but is relatively underrepresented due to normal gene suppression that occurs close to parturition as described above.
[0175] Sickle cell disease is caused by a V6E mutation (GAG to GTG at the DNA level) in the beta globin gene (HBB), where the heme produced is called "heme S" or "HbS". Under hypoxic conditions, HbS molecules aggregate and form fibrous precipitates. These aggregates cause RBCs to become abnormal or "sickling," resulting in a loss of flexibility in the cells. Sickled RBCs are no longer able to squeeze into the microvascular bed and may lead to a vaso-occlusive crisis in sickle cell patients. In addition, sickled RBCs are more fragile than normal RBCs and are prone to hemolysis, which eventually leads to anemia in patients.
[0176] The treatment and management of sickle cell patients is a lifelong topic involving antibiotic therapy, pain management, and infusions during acute attacks. One approach is the use of hydroxyurea, which exerts its effect in part by increasing the production of gamma globulin. However, the long-term side effects of chronic hydroxyurea therapy are still unknown, and the treatment has adverse side effects and may have variable effects between patients. Despite the improved efficacy of sickle cell therapy, patient life expectancy is still only in the mid to late 50s, and the associated morbidity of the disease has a profound impact on a patient's quality of life.
[0177] The thalassemias (alpha thalassemia and beta thalassemia) are also hemoglobin-related disorders and often involve reduced expression of globin chains. This can occur via mutations in the regulatory regions of the gene or from mutations in the globin coding sequence that result in a reduced or reduced level or functional globin expression. Treatment for thalassemia usually involves blood transfusions and iron chelation therapy. Bone marrow transplants can also be used to treat people with thalassemia major if a suitable donor can be found, but this approach can have significant risks.
[0178] One approach that has been proposed for the treatment of sickle cell disease (SCD) and beta thalassemia is to increase the expression of gamma globulin, allowing HbF to functionally replace the abnormal adult heme. As noted above, treatment of SCD patients with hydroxyurea has been considered partially successful due to its effect on increasing gamma globulin expression (DeSimone (1982) Proc Nat'l Acad Sci USA 79(14):4428-31; Ley et al. (1982) N. Engl. J. Medicine, 307:1469-1475; Ley et al. (1983) Blood 62:370-380; Constantoulakis et al. (1988) Blood 72(6):1961-1967, all borrowed incorporated herein by reference). Increasing HbF expression involved identifying genes whose products play a role in the regulation of gamma globulin expression. One such gene is BCL11A. BCL11A encodes a zinc finger protein expressed in adult human erythroid precursor cells, and downregulation of its expression results in increased gamma globulin expression (Sankaran et al. (2008) Science 322:1839, incorporated herein by reference). The use of inhibitory RNA targeting the BCL11A gene has been proposed (eg, US Patent Publication No. 2011 / 0182867, incorporated herein by reference), but this technique has several potential drawbacks, including that complete knock down may not be achieved , delivery of such RNA can be problematic, and RNA must be present continuously and require multiple treatments throughout life.
[0179] The RGNs of the invention can be used to target the BCL11A enhancer region to disrupt BCL11A expression, thereby increasing gamma globulin expression. This targeted disruption can be achieved by non-homologous end joining (NHEJ), whereby the RGN of the invention targets a specific sequence within the BCL11A enhancer region, causes a double-strand break, and the cell's machinery repairs the break, Deleterious mutations are often introduced simultaneously. Similar to what has been described for other disease targets, the RGN of the present invention has advantages over other established RGNs due to their relatively small size, which enables packaging of the RGN and its guide RNA expression cassette into a single AAV vector for in vivo delivery. Know the advantages of RGN. Similar strategies using the RGNs of the invention can also be applied to similar diseases and conditions in humans and non-human animals of agronomic or economic importance. X. Cells Comprising Polynucleotide Genetic Modifications
[0180] Provided herein are cells and organisms comprising a target sequence of interest that has been modified using RGN, crRNA and / or tracrRNA-mediated processes as described herein. In some of these embodiments, the RGN comprises SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117 , 123 or the amino acid sequence of 570-579, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence comprising SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90 , 97, 104, 111, 118 or 124 nucleotide sequence, or an active variant or fragment thereof. In a particular embodiment, the guide RNA includes tracrRNA, and the tracrRNA includes SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, The nucleotide sequence of 112, 119 or 125, or an active variant or fragment thereof. The guide RNA of this system can be single guide RNA or double guide RNA.
[0181] The modified cells can be eukaryotic (eg, mammalian, plant, insect cells) or prokaryotic. Also provided are organelles and embryos comprising at least one nucleotide sequence that has been modified by methods utilizing RGN, crRNA and / or tracrRNA as described herein. Genetically modified cells, organisms, organelles and embryos can be heterozygous or homozygous for the modified nucleotide sequence.
[0182] The chromosomal modification of the cell, organism, organelle or embryo may result in altered expression (up or down), inactivation, or altered expression of the protein product or integrated sequence. In those embodiments where the chromosomal modification results in gene inactivation or expression of a non-functional protein product, the genetically modified cell, organism, organelle or embryo is said to be "knocked out." The knockout phenotype can be a deletion mutation (that is, a deletion of at least one nucleotide), an insertion mutation (that is, an insertion of at least one nucleotide), or a nonsense mutation (that is, a substitution of at least one nucleotide , resulting in the introduction of a stop codon).
[0183] Alternatively, chromosomal modification of a cell, organism, organelle, or embryo can produce a "knock in," which results from chromosomal integration of a protein-encoding nucleotide sequence. In some of these embodiments, the coding sequence is integrated into the chromosome such that the chromosomal sequence encoding the wild-type protein is inactive but exhibits the exogenously introduced protein.
[0184] In some embodiments, the chromosomal modification results in the production of a variant protein product. The expressed variant protein product may have at least one amino acid substitution and / or at least one amino acid addition or deletion. A variant protein product encoded by an altered chromosomal sequence may exhibit modified characteristics or activities when compared to the wild-type protein, including but not limited to altered enzymatic activity or substrate specificity.
[0185] In yet other embodiments, the chromosomal modification can result in an altered protein expression pattern. As a non-limiting example, chromosomal alterations in regulatory regions that control expression of a protein product can result in overexpression or downregulation or altered tissue or temporal expression patterns of the protein product.
[0186] Cells that have been modified can be grown into organisms, eg, plants, according to conventional means. See, eg, McCormick et al. (1986) Plant Cell Reports 5:81-84. These plants can then be grown and pollinated with the same modified strain or a different strain, and the resulting hybrid has the genetic modification. The present invention provides genetically modified seeds. Progeny, variants and mutants of the regenerated plants are also included within the scope of the invention, provided that these parts include genetic modifications. Further provided are processed plant products or by-products that retain the genetic modification, including, for example, soybean meal.
[0187] The methods provided herein can be used to modify any plant species, including but not limited to monocots and dicots. Examples of plants of interest include, but are not limited to, corn (maize), sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugar cane, tobacco, barley and canola, Brassica oleracea Canola, alfalfa, rye, millet, safflower, peanut, sweet potato, tapioca, coffee, coconut, pineapple, citrus, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew, macadamia , almonds, oats, vegetables, ornamentals and conifers.
[0188] Vegetables include, but are not limited to, tomatoes, lettuce, mung beans, king beans, peas, and members of the genus Melon such as cucumbers, cantaloupe, and cantaloupe. Ornamental plants include but are not limited to rhododendrons, hydrangeas, hibiscus, roses, tulips, daffodils, petunias, carnations, Christmas reds and chrysanthemums. Preferably, the plant of the present invention is an agricultural crop (for example, corn, sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugar cane, tobacco, barley, rapeseed, etc.).
[0189] The methods provided herein can also be used to genetically modify any prokaryotic species, including but not limited to: Archaea and bacteria (e.g., Bacillus, Klebsiella, Streptomyces, Rhizobium, Escherichia, Pseudomonas Salmonella, Shigella, Vibrio, Yersinia, Mycoplasma, Agrobacterium, Lactobacillus.
[0190] The methods provided herein can be used to genetically modify any eukaryotic species or cells derived therefrom, including but not limited to: animals (e.g., mammals, insects, fish, birds, and reptiles), fungi, proteus Insects, algae and yeast. In some embodiments, the cells modified by the methods disclosed in the present invention include cells of hematopoietic origin, for example, cells of the immune system, including but not limited to: B cells, T cells, natural killer (NK) cells, high potential stem cells, Induced pluripotent stem cells, chimeric antigen receptor T (CAR-T) cells, monocytes, macrophages and dendritic cells.
[0191] Modified cells can be introduced into an organism. In the case of autologous cell transplantation, the cells can be derived from the same organism (eg, a human) where the cells have been modified ex vivo. Alternatively, in the case of allogeneic cell transplantation, the cells are derived from another organism in the same species (eg, another human being).
[0192] The articles "a" and "an" are used herein to refer to one or more (ie, at least one) of the grammatical object of the article. By way of example, "polypeptide" means one or more polypeptides.
[0193] All publications and patent applications mentioned in the specification represent the level of ordinary knowledge in the art to which this disclosure belongs. All publications and patent applications are herein incorporated by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference.
[0194] While the foregoing invention has been described in some detail, by way of illustration and example, for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended embodiments. Non-limiting examples include:
[0195] 1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide, and the RGN polypeptide comprises the same sequence as SEQ ID NO: Any of 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123, and 570-579 have at least 90% sequence a consensus amino acid sequence; wherein the RGN polypeptide is capable of binding a target DNA sequence in an RNA-guided, sequence-specific manner when combined with a guide RNA (gRNA) capable of hybridizing to the target DNA sequence, and wherein the RGN polypeptide encodes The polynucleotide of the polypeptide is operably linked to a promoter heterologous to the polynucleotide.
[0196] 2. The nucleic acid molecule according to embodiment 1, wherein the RGN polypeptide comprises a sequence corresponding to SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89 , 96, 103, 110, 117, 123, and any of 570-579 have an amino acid sequence with at least 95% sequence identity.
[0197] 3. The nucleic acid molecule according to embodiment 1, wherein the RGN polypeptide comprises a sequence corresponding to SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89 , 96, 103, 110, 117, 123, and any of 570-579 have an amino acid sequence with 100% sequence identity.
[0198] 4. The nucleic acid molecule according to embodiment 1, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63 and has an amino acid position corresponding to 305 of SEQ ID NO: 63 Isoleucine, valine at the amino acid position corresponding to 328 of SEQ ID NO:63, leucine at the amino acid position corresponding to 366 of SEQ ID NO:63, Threonine at the amino acid position corresponding to 368 of ID NO:63, and valine at the amino acid position corresponding to 405 of SEQ ID NO:63.
[0199] 5. The nucleic acid molecule according to any one of embodiments 1-4, wherein the RGN polypeptide is capable of cleaving the target DNA sequence upon binding.
[0200] 6. The nucleic acid molecule of embodiment 5, wherein the RGN polypeptide is capable of producing a double-strand break.
[0201] 7. The nucleic acid molecule according to embodiment 5, wherein the RGN polypeptide is capable of producing single-strand breaks.
[0202] 8. The nucleic acid molecule according to any one of embodiments 1-4, wherein the RGN polypeptide is nuclease inactive or a nicking enzyme.
[0203] 9. The nucleic acid molecule according to any one of embodiments 1-8, wherein the RGN polypeptide is operably fused to a base editing polypeptide.
[0204] 10. The nucleic acid molecule according to embodiment 9, wherein the base editing polypeptide is deaminase.
[0205] 11. The nucleic acid molecule according to embodiment 10, wherein the deaminase is cytidine deaminase or adenine deaminase.
[0206] 12. The nucleic acid molecule of any one of embodiments 1-11, wherein the RGN polypeptide comprises one or more nuclear localization signals.
[0207] 13. The nucleic acid molecule according to any one of embodiments 1-12, wherein the RGN polypeptide is codon-optimized for expression in eukaryotic cells.
[0208] 14. The nucleic acid molecule of any one of embodiments 1-13, wherein the target DNA sequence is positioned adjacent to a prospacer adjacent motif (PAM).
[0209] 15. A vector comprising the nucleic acid molecule of any one of embodiments 1-14.
[0210] 16. The vector according to embodiment 15, further comprising at least one nucleotide sequence encoding the gRNA, which can hybridize with the target DNA sequence.
[0211] 17. The vector according to embodiment 16, wherein the guide RNA is selected from the group consisting of: a) guide RNA comprising: i) CRISPR RNA comprising at least 90 % sequence identity CRISPR repeat sequence; and ii) tracrRNA having at least 90% sequence identity to SEQ ID NO: 3; wherein the RGN polypeptide includes an amine having at least 90% sequence identity to SEQ ID NO: 1 amino acid sequence; b) a guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 9; and ii) tracrRNA which is identical to SEQ ID NO: 10 has at least 90% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 8; c) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes and SEQ ID NO: 16 has a CRISPR repeat sequence with at least 90% sequence identity; and ii) tracrRNA, the tracrRNA has at least 90% sequence identity with SEQ ID NO: 17; wherein the RGN polypeptide comprises a sequence identity with SEQ ID NO: 15 an amino acid sequence of at least 90% sequence identity; d) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 23; and ii) a tracrRNA , the tracrRNA has at least 90% sequence identity with SEQ ID NO: 24; wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 22; e) guide RNA, including: i) CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 30; and ii) tracrRNA having at least 90% sequence identity to SEQ ID NO: 31; wherein the RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 29; f) a guide RNA comprising: i) a CRISPR RNA comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 37 CRISPR repeat sequence; and ii) tracrRNA, the tracrRNA has at least 90% sequence identity with SEQ ID NO: 38; wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 36; g ) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:44; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO:45 Sequence identity; wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 43; h) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with SEQ ID NO: 51 A CRISPR repeat sequence having at least 90% sequence identity; and ii) tracrRNA having at least 90% sequence identity to SEQ ID NO:52; wherein the RGN polypeptide comprises at least 90% sequence identity to SEQ ID NO:50 i) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 57; and ii) a tracrRNA which is identical to SEQ ID NO: 57 ID NO: 58 has at least 90% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 56; j) guide RNA, including: i) CRISPR RNA, the CRISPR The RNA includes a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 64; and ii) tracrRNA, which has at least 90% sequence identity to SEQ ID NO: 65; wherein the RGN polypeptide includes a sequence identical to SEQ ID NO Any one of: 63 and 570-579 has an amino acid sequence with at least 90% sequence identity; k) a guide RNA comprising: i) a CRISPR RNA comprising at least 90% sequence with SEQ ID NO:71 Consistent CRISPR repeat sequence; and ii) tracrRNA having at least 90% sequence identity to SEQ ID NO: 72; wherein the RGN polypeptide includes amino acids having at least 90% sequence identity to SEQ ID NO: 70 sequence; l) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO:77; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO:78 At least 90% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 76; m) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with SEQ ID NO: 76; NO: 84 has a CRISPR repeat sequence with at least 90% sequence identity; and ii) tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 85; wherein the RGN polypeptide includes at least 90% sequence identity with SEQ ID NO: 83 % amino acid sequence identity; n) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 90; and ii) a tracrRNA comprising tracrRNA has at least 90% sequence identity with SEQ ID NO: 91; wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 89; o) guide RNA, including: i) CRISPR RNA , the CRISPR RNA comprises a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 97; and ii) tracrRNA, the tracrRNA has at least 90% sequence identity to SEQ ID NO: 98; wherein the RGN polypeptide comprises and An amino acid sequence having at least 90% sequence identity to SEQ ID NO: 96; p) a guide RNA comprising: i) a CRISPR RNA comprising CRISPR repeats having at least 90% sequence identity to SEQ ID NO: 104 sequence; and ii) tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 105; wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 103; q) guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO: 111; and ii) tracrRNA having at least 90% sequence identity to SEQ ID NO: 112 wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 110; r) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 118 A CRISPR repeat sequence with 90% sequence identity; and ii) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 119; wherein the RGN polypeptide comprises a tracrRNA having at least 90% sequence identity to SEQ ID NO: 117 Amino acid sequence; s) guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 124; and ii) tracrRNA comprising SEQ ID NO: 124; : 125 has at least 90% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 123; and t) guide RNA, including: i) CRISPR RNA, the CRISPR RNA Including a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO: 84; and ii) tracrRNA, the tracrRNA has at least 90% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide includes a sequence with SEQ ID NO: 83 amino acid sequences having at least 90% sequence identity.
[0212] 18. The vector according to embodiment 16, wherein the guide RNA is selected from the group consisting of: a) guide RNA comprising: i) CRISPR RNA comprising at least 95 % sequence identity CRISPR repeat sequence; and ii) tracrRNA having at least 95% sequence identity to SEQ ID NO: 3; wherein the RGN polypeptide includes an amine having at least 95% sequence identity to SEQ ID NO: 1 amino acid sequence; b) a guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 9; and ii) tracrRNA which is identical to SEQ ID NO: 10 has at least 95% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 8; c) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes and SEQ ID NO: 16 has a CRISPR repeat sequence with at least 95% sequence identity; and ii) tracrRNA, the tracrRNA has at least 95% sequence identity with SEQ ID NO: 17; wherein the RGN polypeptide comprises a sequence with SEQ ID NO: 15 An amino acid sequence having at least 95% sequence identity; d) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 23; and ii) a tracrRNA , the tracrRNA has at least 95% sequence identity with SEQ ID NO: 24; wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 22; e) guide RNA, including: i) CRISPR RNA comprising a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 30; and ii) tracrRNA having at least 95% sequence identity to SEQ ID NO: 31; wherein the RGN polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 29; f) a guide RNA comprising: i) a CRISPR RNA comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 37 CRISPR repeat sequence; and ii) tracrRNA, the tracrRNA has at least 95% sequence identity with SEQ ID NO: 38; wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 36; g ) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat having at least 95% sequence identity to SEQ ID NO:44; and ii) a tracrRNA having at least 95% sequence identity to SEQ ID NO:45 Sequence identity; wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 43; h) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with SEQ ID NO: 51 A CRISPR repeat sequence having at least 95% sequence identity; and ii) tracrRNA having at least 95% sequence identity to SEQ ID NO:52; wherein the RGN polypeptide comprises at least 95% sequence identity to SEQ ID NO:50 i) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 57; and ii) a tracrRNA which is identical to SEQ ID NO: 57 ID NO: 58 has at least 95% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 56; j) guide RNA, including: i) CRISPR RNA, the CRISPR The RNA includes a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 64; and ii) tracrRNA, which has at least 95% sequence identity to SEQ ID NO: 65; wherein the RGN polypeptide includes a sequence identical to SEQ ID NO Any one of: 63 and 570-579 has an amino acid sequence with at least 95% sequence identity; k) a guide RNA comprising: i) a CRISPR RNA comprising at least 95% sequence with SEQ ID NO:71 Consistent CRISPR repeats; and ii) tracrRNA having at least 95% sequence identity to SEQ ID NO: 72; wherein the RGN polypeptide includes amino acids having at least 95% sequence identity to SEQ ID NO: 70 Sequence; l) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO:77; and ii) a tracrRNA having at least 95% sequence identity to SEQ ID NO:78 At least 95% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 76; m) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with SEQ ID NO: 76; NO: 84 has a CRISPR repeat sequence with at least 95% sequence identity; and ii) tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 85; wherein the RGN polypeptide includes at least 95% sequence identity with SEQ ID NO: 83 % amino acid sequence identity; n) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 90; and ii) a tracrRNA comprising tracrRNA has at least 95% sequence identity with SEQ ID NO: 91; wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 89; o) guide RNA, including: i) CRISPR RNA , the CRISPR RNA comprises a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 97; and ii) tracrRNA, the tracrRNA has at least 95% sequence identity to SEQ ID NO: 98; wherein the RGN polypeptide comprises a sequence identical to An amino acid sequence having at least 95% sequence identity to SEQ ID NO: 96; p) a guide RNA comprising: i) a CRISPR RNA comprising CRISPR repeats having at least 95% sequence identity to SEQ ID NO: 104 sequence; and ii) tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 105; wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 103; q) guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat having at least 95% sequence identity to SEQ ID NO: 111; and ii) tracrRNA having at least 95% sequence identity to SEQ ID NO: 112 wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 110; r) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 118 A CRISPR repeat sequence with 95% sequence identity; and ii) a tracrRNA having at least 95% sequence identity to SEQ ID NO: 119; wherein the RGN polypeptide comprises a tracrRNA having at least 95% sequence identity to SEQ ID NO: 117 Amino acid sequence; s) guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat having at least 95% sequence identity to SEQ ID NO: 124; and ii) tracrRNA comprising SEQ ID NO: 124; : 125 has at least 95% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 123; and t) guide RNA, including: i) CRISPR RNA, the CRISPR RNA Including a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 84; and ii) tracrRNA, the tracrRNA has at least 95% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide includes a sequence with SEQ ID NO: 83 amino acid sequences having at least 95% sequence identity.
[0213] 19. The vector according to embodiment 16, wherein the guide RNA is selected from the group consisting of: a) guide RNA comprising: i) CRISPR RNA comprising 100% of SEQ ID NO:2 CRISPR repeat sequences with sequence identity; and ii) tracrRNA having 100% sequence identity to SEQ ID NO: 3; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity to SEQ ID NO: 1 ; b) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity to SEQ ID NO:9; and ii) a tracrRNA having 100% sequence identity to SEQ ID NO:10 Sequence identity; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 8; c) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with SEQ ID NO: 16 A CRISPR repeat sequence with 100% sequence identity; and ii) tracrRNA having 100% sequence identity to SEQ ID NO: 17; wherein the RGN polypeptide includes an amine group having 100% sequence identity to SEQ ID NO: 15 d) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 23; and ii) a tracrRNA comprising a sequence identical to SEQ ID NO: 24 100% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having 100% sequence identity with SEQ ID NO: 22; e) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes and SEQ ID NO: 30 CRISPR repeats with 100% sequence identity; and ii) tracrRNA having 100% sequence identity to SEQ ID NO: 31; wherein the RGN polypeptide comprises 100% sequence identity to SEQ ID NO: 29 an amino acid sequence; f) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 37; and ii) a tracrRNA which is identical to SEQ ID NO: 38 has 100% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having 100% sequence identity with SEQ ID NO: 36; g) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with SEQ ID NO: 36; NO: 44 has a CRISPR repeat sequence with 100% sequence identity; and ii) tracrRNA, which has 100% sequence identity with SEQ ID NO: 45; wherein the RGN polypeptide includes 100% sequence identity with SEQ ID NO: 43 h) a guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 51; and ii) tracrRNA which is identical to SEQ ID NO: 51 NO: 52 has 100% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having 100% sequence identity with SEQ ID NO: 50; i) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes and A CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 57; and ii) tracrRNA having 100% sequence identity to SEQ ID NO: 58; wherein the RGN polypeptide comprises 100% sequence identity to SEQ ID NO: 56 an amino acid sequence of sequence identity; j) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 64; and ii) a tracrRNA that is compatible with SEQ ID NO: 65 has 100% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having 100% sequence identity with any one of SEQ ID NO: 63 and 570-579; k) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 71; and ii) tracrRNA, the tracrRNA has 100% sequence identity to SEQ ID NO: 72; wherein the RGN polypeptide Including an amino acid sequence having 100% sequence identity with SEQ ID NO: 70; 1) a guide RNA, including: i) CRISPR RNA comprising a CRISPR repeat having 100% sequence identity with SEQ ID NO: 77 sequence; and ii) tracrRNA, which has 100% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 76; m) guide RNA, Including: i) CRISPR RNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 84; and ii) tracrRNA having 100% sequence identity to SEQ ID NO: 85; wherein the The RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 83; n) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 90 CRISPR repeat sequence; and ii) tracrRNA, which has 100% sequence identity with SEQ ID NO: 91; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 89; o) guide RNA comprising: i) CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity to SEQ ID NO:97; and ii) tracrRNA having 100% sequence identity to SEQ ID NO:98; Wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 96; p) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 104 and ii) tracrRNA, which has 100% sequence identity with SEQ ID NO: 105; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 103; q ) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity to SEQ ID NO:111; and ii) a tracrRNA having 100% sequence identity to SEQ ID NO:112 wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 110; r) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes 100% sequence identity with SEQ ID NO: 118 CRISPR repeat sequences with sequence identity; and ii) tracrRNA having 100% sequence identity to SEQ ID NO: 119; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity to SEQ ID NO: 117 s) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat having 100% sequence identity to SEQ ID NO:124; and ii) a tracrRNA having 100% sequence identity to SEQ ID NO:125 Sequence identity; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 123; and t) guide RNA, including: i) CRISPR RNA, the CRISPR RNA includes an amino acid sequence with SEQ ID NO: 84 A CRISPR repeat sequence having 100% sequence identity; and ii) a tracrRNA having 100% sequence identity to SEQ ID NO: 78; wherein the RGN polypeptide includes an amine having 100% sequence identity to SEQ ID NO: 83 amino acid sequence.
[0214] 20. The vector of any one of embodiments 16-19, wherein the gRNA is a single guide RNA.
[0215] 21. The vector of any one of embodiments 16-19, wherein the gRNA is a dual guide RNA.
[0216] 22. A cell comprising a nucleic acid molecule according to any one of embodiments 1-14 or a vector according to any one of embodiments 15-21.
[0217] 23. A method for producing an RGN polypeptide, comprising: culturing the cell according to embodiment 22 under the condition that the RGN polypeptide is expressed.
[0218] 24. A method of making an RGN polypeptide, comprising introducing a heterologous nucleic acid molecule into a cell, the heterologous nucleic acid molecule comprising a nucleotide sequence encoding an RNA-guided nuclease (RGN) polypeptide, the RNA-guided nucleic acid Enzyme (RGN) polypeptides include those with SEQ ID NOS: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123, and 570 - an amino acid sequence having at least 90% sequence identity to any of 579; wherein the RGN polypeptide is capable of RNA-guided sequence specificity when bound to a guide RNA (gRNA) capable of hybridizing to the target DNA sequence binding the target DNA sequence in a manner; and culturing the cell under the condition that the RGN polypeptide is expressed.
[0219] 25. The method according to embodiment 24, wherein the RGN polypeptide comprises a sequence with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, Any of 96, 103, 110, 117, 123, and 570-579 have an amino acid sequence with at least 95% sequence identity.
[0220] 26. The method according to embodiment 24, wherein the RGN polypeptide comprises a combination with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, Any of 96, 103, 110, 117, 123, and 570-579 have an amino acid sequence with 100% sequence identity.
[0221] 27. The method according to embodiment 24, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63, and has a difference at the amino acid position corresponding to 305 of SEQ ID NO: 63 Leucine, valine at the amino acid position corresponding to 328 of SEQ ID NO:63, leucine at the amino acid position corresponding to 366 of SEQ ID NO:63, at the amino acid position corresponding to SEQ ID NO:63, Threonine at the amino acid position corresponding to 368 of NO:63, and valine at the amino acid position corresponding to 405 of SEQ ID NO:63.
[0222] 28. The method of any one of embodiments 23-27, further comprising purifying the RGN polypeptide.
[0223] 29. The method of any one of embodiments 23-27, wherein the cell further expresses one or more guide RNAs capable of binding to the RGN polypeptide to form an RGN ribonuclear nucleus protein complexes.
[0224] 30. The method of embodiment 29, further comprising purifying the RGN ribonucleoprotein complex.
[0225] 31. An isolated RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises a sequence with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70 , 76, 83, 89, 96, 103, 110, 117, 123, and any one of 570-579 has an amino acid sequence of at least 90% sequence identity; and wherein when combined with a guide capable of hybridizing to the target DNA sequence When RNA (gRNA) is bound, the RGN polypeptide can bind the target DNA sequence of the DNA molecule in an RNA-guided sequence-specific manner.
[0226] 32. The RGN polypeptide of isolation as in embodiment 31, wherein the RGN polypeptide comprises a sequence with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83 , 89, 96, 103, 110, 117, 123, and any of 570-579 have an amino acid sequence having at least 95% sequence identity.
[0227] 33. The RGN polypeptide of isolation according to embodiment 31, wherein the RGN polypeptide comprises a sequence corresponding to SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83 , 89, 96, 103, 110, 117, 123 and any one of 570-579 has an amino acid sequence with 100% sequence identity.
[0228] 34. The isolated RGN polypeptide of embodiment 31, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63 and has an amino acid position corresponding to 305 of SEQ ID NO: 63 Isoleucine at the position, valine at the amino acid position corresponding to 328 of SEQ ID NO: 63, leucine at the amino acid position corresponding to 366 of SEQ ID NO: 63, at threonine at the amino acid position corresponding to 368 of SEQ ID NO:63, and valine at the amino acid position corresponding to 405 of SEQ ID NO:63.
[0229] 35. The isolated RGN polypeptide of any one of embodiments 31-34, wherein the RGN polypeptide is capable of cleaving the target DNA sequence upon binding.
[0230] 36. The isolated RGN polypeptide of embodiment 35, wherein cleavage by the RGN polypeptide produces a double-strand break.
[0231] 37. The isolated RGN polypeptide of embodiment 35, wherein cleavage by the RGN polypeptide produces a single-strand break.
[0232] 38. The isolated RGN polypeptide of any one of embodiments 31-34, wherein the RGN polypeptide is nuclease inactive or a nicking enzyme.
[0233] 39. The isolated RGN polypeptide of any one of embodiments 31-38, wherein the RGN polypeptide is operably fused to a base editing polypeptide.
[0234] 40. The isolated RGN polypeptide of embodiment 39, wherein the base editing polypeptide is deaminase.
[0235] 41. The isolated RGN polypeptide of any one of embodiments 31-40, wherein the target DNA sequence is positioned adjacent to a prospacer adjacent motif (PAM).
[0236] 42. The isolated RGN polypeptide of any one of embodiments 31-41, wherein the RGN polypeptide comprises one or more nuclear localization signals.
[0237] 43. A nucleic acid molecule comprising a polynucleotide encoding CRISPR RNA (crRNA), wherein the crRNA comprises a spacer sequence and a CRISPR repeat sequence, wherein the CRISPR repeat sequence comprises the same sequence as SEQ ID NO: 2, 9, 16, Any one of 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, 90, 97, 104, 111, 118, and 124 has a nucleotide sequence of at least 90% sequence identity; wherein the guide RNA comprising: a) the crRNA; and b) a transcriptionally activated CRISPR RNA (tracrRNA) that hybridizes to the CRISPR repeat sequence of the crRNA; when the guide RNA is combined with an RNA-guided nuclease ( RGN) polypeptide, through the spacer sequence of the crRNA, can hybridize with the target DNA sequence in a sequence-specific manner, and wherein the polynucleotide encoding the crRNA is operably linked to a promoter heterologous to the polynucleotide son.
[0238] 44. The nucleic acid molecule according to embodiment 43, wherein the CRISPR repeat sequence comprises sequences corresponding to SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, Any of 90, 97, 104, 111, 118, and 124 have a nucleotide sequence with at least 95% sequence identity.
[0239] 45. The nucleic acid molecule according to embodiment 43, wherein the CRISPR repeat sequence comprises sequences corresponding to SEQ ID NO: 2, 9, 16, 23, 30, 37, 44, 51, 57, 64, 71, 77, 84, Any one of 90, 97, 104, 111, 118, and 124 has a nucleotide sequence with 100% sequence identity.
[0240] 46. A vector comprising the nucleic acid molecule of any one of embodiments 43-45.
[0241] 47. The carrier of embodiment 46, wherein the carrier further comprises a polynucleotide encoding the tracrRNA.
[0242] 48. The carrier of embodiment 47, wherein the tracrRNA is selected from the group consisting of: a) a tracrRNA with at least 90% sequence identity to SEQ ID NO: 3, wherein the CRISPR repeat sequence is identical to SEQ ID NO: 2 has at least 90% sequence identity; b) tracrRNA having at least 90% sequence identity with SEQ ID NO: 10, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 9; c) A tracrRNA having at least 90% sequence identity to SEQ ID NO: 17, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 16; d) having at least 90% sequence identity to SEQ ID NO: 24 tracrRNA, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 23; e) a tracrRNA with at least 90% sequence identity with SEQ ID NO: 31, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 30 has at least 90% sequence identity; f) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 38, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 37; g) with SEQ ID NO: 37 has at least 90% sequence identity; ID NO: 45 has a tracrRNA with at least 90% sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 44; h) a tracrRNA with at least 90% sequence identity with SEQ ID NO: 52 , wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 51; i) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 58, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 57 At least 90% sequence identity; j) tracrRNA having at least 90% sequence identity with SEQ ID NO: 65, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 64; k) with SEQ ID NO : 72 a tracrRNA with at least 90% sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 71; l) a tracrRNA with at least 90% sequence identity with SEQ ID NO: 78, wherein The CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO:77; m) a tracrRNA with at least 90% sequence identity with SEQ ID NO:85, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO:84 % sequence identity; n) tracrRNA having at least 90% sequence identity with SEQ ID NO: 91, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 90; o) with SEQ ID NO: 98 A tracrRNA having at least 90% sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 97; p) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 105, wherein the CRISPR The repeat sequence has at least 90% sequence identity to SEQ ID NO: 104; q) a tracrRNA with at least 90% sequence identity to SEQ ID NO: 112, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 111 Identity; r) a tracrRNA having at least 90% sequence identity to SEQ ID NO: 119, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 118; and s) having at least 90% sequence identity to SEQ ID NO: 125 A tracrRNA of at least 90% sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO:124.
[0243] 49. The carrier of embodiment 47, wherein the tracrRNA is selected from the group consisting of: a) a tracrRNA with at least 95% sequence identity to SEQ ID NO: 3, wherein the CRISPR repeat sequence is identical to SEQ ID NO: 2 has at least 95% sequence identity; b) tracrRNA having at least 95% sequence identity with SEQ ID NO: 10, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 9; c) A tracrRNA having at least 95% sequence identity to SEQ ID NO: 17, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 16; d) having at least 95% sequence identity to SEQ ID NO: 24 A tracrRNA, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 23; e) a tracrRNA with at least 95% sequence identity with SEQ ID NO: 31, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 30 has at least 95% sequence identity; f) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 38, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 37; g) with SEQ ID NO: 37 has at least 95% sequence identity; ID NO: 45 has a tracrRNA with at least 95% sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 44; h) a tracrRNA with at least 95% sequence identity with SEQ ID NO: 52 , wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 51; i) a tracrRNA having at least 95% sequence identity to SEQ ID NO: 58, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 57 At least 95% sequence identity; j) tracrRNA having at least 95% sequence identity with SEQ ID NO: 65, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 64; k) with SEQ ID NO : 72 a tracrRNA with at least 95% sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 71; l) a tracrRNA with at least 95% sequence identity with SEQ ID NO: 78, wherein The CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO:77; m) a tracrRNA with at least 95% sequence identity with SEQ ID NO:85, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO:84 % sequence identity; n) tracrRNA having at least 95% sequence identity with SEQ ID NO:91, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO:90; o) with SEQ ID NO:98 A tracrRNA having at least 95% sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 97; p) a tracrRNA having at least 95% sequence identity to SEQ ID NO: 105, wherein the CRISPR repeat A repeat sequence having at least 95% sequence identity to SEQ ID NO: 104; q) a tracrRNA having at least 95% sequence identity to SEQ ID NO: 112, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 111 Identity; r) a tracrRNA having at least 95% sequence identity to SEQ ID NO: 119, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 118; and s) having at least 95% sequence identity to SEQ ID NO: 125 A tracrRNA of at least 95% sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO:124.
[0244] 50. The carrier of embodiment 47, wherein the tracrRNA is selected from the group consisting of: a) a tracrRNA with 100% sequence identity to SEQ ID NO: 3, wherein the CRISPR repeat sequence is identical to SEQ ID NO : 2 with 100% sequence identity; b) tracrRNA with 100% sequence identity with SEQ ID NO: 10, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 9; c) with SEQ ID NO: 100% sequence identity : 17 a tracrRNA with 100% sequence identity, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 16; d) a tracrRNA with 100% sequence identity with SEQ ID NO: 24, wherein the CRISPR repeat The sequence has 100% sequence identity with SEQ ID NO: 23; e) a tracrRNA with 100% sequence identity with SEQ ID NO: 31, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 30; f ) tracrRNA having 100% sequence identity with SEQ ID NO: 38, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 37; g) tracrRNA having 100% sequence identity with SEQ ID NO: 45 , wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 44; h) a tracrRNA with 100% sequence identity with SEQ ID NO: 52, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 51 Sequence identity; i) tracrRNA having 100% sequence identity with SEQ ID NO: 58, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 57; j) 100% with SEQ ID NO: 65 A tracrRNA with sequence identity, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 64; k) a tracrRNA with 100% sequence identity with SEQ ID NO: 72, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO : 71 has 100% sequence identity; l) tracrRNA with 100% sequence identity with SEQ ID NO: 78, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 77; m) with SEQ ID NO: 77 : 85 tracrRNA with 100% sequence identity, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 84; n) tracrRNA with 100% sequence identity with SEQ ID NO: 91, wherein the CRISPR repeat The sequence has 100% sequence identity with SEQ ID NO: 90; o) tracrRNA with 100% sequence identity with SEQ ID NO: 98, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 97; p ) tracrRNA having 100% sequence identity with SEQ ID NO: 105, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 104; q) tracrRNA having 100% sequence identity with SEQ ID NO: 112 , wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 111; r) tracrRNA with 100% sequence identity with SEQ ID NO: 119, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 118 sequence identity; and s) a tracrRNA having 100% sequence identity to SEQ ID NO:125, wherein the CRISPR repeat sequence has 100% sequence identity to SEQ ID NO:124.
[0245] 51. The vector of any one of embodiments 47-50, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as single guide RNA.
[0246] 52. The vector of any one of embodiments 47-50, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.
[0247] 53. The vector according to any one of embodiments 46-52, wherein the vector further comprises a polynucleotide encoding the RGN polypeptide.
[0248] 54. The vector of embodiment 53, wherein the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 1, wherein the CRISPR repeat sequence is identical to SEQ ID NO: 2 has at least 90% sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 3; b) an RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 8, wherein The CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 9, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 10; c) has at least 90% sequence identity with SEQ ID NO: 15 The RGN polypeptide, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 16, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 17; d) has at least 90% sequence identity with SEQ ID NO: 22 An RGN polypeptide with 90% sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 23, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 24; e) with SEQ ID NO: 29 has an RGN polypeptide with at least 90% sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 30, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 31; f) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 36, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 37, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 38 % sequence identity; g) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 43, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 44, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: : 45 has at least 90% sequence identity; h) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 50, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 51, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:52; i) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO:56, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO:57 identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 58; j) an RGN polypeptide with at least 90% sequence identity with any of SEQ ID NO: 63 and 570-579, wherein the CRISPR The repeat sequence has at least 90% sequence identity with SEQ ID NO: 64, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 65; k) RGN with at least 90% sequence identity with SEQ ID NO: 70 A polypeptide, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 71, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 72; l) has at least 90% sequence identity with SEQ ID NO: 76 RGN polypeptides with sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 77, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 78; m) with SEQ ID NO: 83 an RGN polypeptide having at least 90% sequence identity, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 84, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 85; n) An RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 89, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 90, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 91 identity; o) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 96, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 97, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 98 having at least 90% sequence identity; p) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 103, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 104, and the tracrRNA is identical to SEQ ID NO: 105 has at least 90% sequence identity; q) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 110, wherein the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 111 , and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 112; r) an RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 117, wherein the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 118 90% sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 119; s) RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 123, wherein the CRISPR repeat sequence and SEQ ID NO: 124 has at least 90% sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 125; and t) an RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 83, wherein the The CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO:84, and the tracrRNA has at least 90% sequence identity to SEQ ID NO:78.
[0249] 55. The vector of embodiment 53, wherein the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 1, wherein the CRISPR repeat sequence is identical to SEQ ID NO: 2 has at least 95% sequence identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 3; b) an RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 8, wherein The CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 9, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 10; c) has at least 95% sequence identity with SEQ ID NO: 15 The RGN polypeptide, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 16, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 17; d) has at least 95% sequence identity with SEQ ID NO: 22 An RGN polypeptide with 95% sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 23, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 24; e) with SEQ ID NO: 29 has an RGN polypeptide with at least 95% sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 30, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 31; f) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 36, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 37, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 38 % sequence identity; g) an RGN polypeptide having at least 95% sequence identity with SEQ ID NO: 43, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 44, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: : 45 has at least 95% sequence identity; h) an RGN polypeptide having at least 95% sequence identity with SEQ ID NO: 50, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 51, and the tracrRNA has at least 95% sequence identity to SEQ ID NO:52; i) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO:56, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO:57 identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 58; j) an RGN polypeptide with at least 95% sequence identity with any of SEQ ID NO: 63 and 570-579, wherein the CRISPR The repeat sequence has at least 95% sequence identity with SEQ ID NO: 64, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 65; k) RGN with at least 95% sequence identity with SEQ ID NO: 70 A polypeptide, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 71, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 72; l) has at least 95% sequence identity with SEQ ID NO: 76 RGN polypeptides with sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 77, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 78; m) with SEQ ID NO: 83 an RGN polypeptide having at least 95% sequence identity, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 84, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 85; n) An RGN polypeptide having at least 95% sequence identity to SEQ ID NO:89, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO:90, and the tracrRNA has at least 95% sequence identity to SEQ ID NO:91 identity; o) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 96, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 97, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 98 having at least 95% sequence identity; p) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 103, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 104, and the tracrRNA is identical to SEQ ID NO: 105 has at least 95% sequence identity; q) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 110, wherein the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 111 , and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 112; r) an RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 117, wherein the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 118 95% sequence identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 119; s) RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 123, wherein the CRISPR repeat sequence and SEQ ID NO: 124 has at least 95% sequence identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 125; and t) an RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 83, wherein the The CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO:84, and the tracrRNA has at least 95% sequence identity to SEQ ID NO:78.
[0250] 56. The carrier of embodiment 53, wherein the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide with 100% sequence identity to SEQ ID NO: 1, wherein the CRISPR repeat sequence is identical to SEQ ID NO: 1 ID NO: 2 has 100% sequence identity, and the tracrRNA has 100% sequence identity with SEQ ID NO: 3; b) RGN polypeptide with 100% sequence identity with SEQ ID NO: 8, wherein the CRISPR repeat sequence It has 100% sequence identity with SEQ ID NO: 9, and the tracrRNA has 100% sequence identity with SEQ ID NO: 10; c) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 15, wherein the CRISPR The repeat sequence has 100% sequence identity with SEQ ID NO: 16, and the tracrRNA has 100% sequence identity with SEQ ID NO: 17; d) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 22, wherein The CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 23, and the tracrRNA has 100% sequence identity with SEQ ID NO: 24; e) RGN polypeptide with 100% sequence identity with SEQ ID NO: 29 , wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 30, and the tracrRNA has 100% sequence identity with SEQ ID NO: 31; f) has 100% sequence identity with SEQ ID NO: 36 RGN polypeptide, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 37, and the tracrRNA has 100% sequence identity with SEQ ID NO: 38; g) has 100% sequence identity with SEQ ID NO: 43 A specific RGN polypeptide, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 44, and the tracrRNA has 100% sequence identity with SEQ ID NO: 45; h) has 100% sequence identity with SEQ ID NO: 50 RGN polypeptides with sequence identity, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 51, and the tracrRNA has 100% sequence identity with SEQ ID NO: 52; i) has 100% sequence identity with SEQ ID NO: 56 RGN polypeptide with 100% sequence identity, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 57, and the tracrRNA has 100% sequence identity with SEQ ID NO: 58; j) with SEQ ID NO: An RGN polypeptide having 100% sequence identity to any of 63 and 570-579, wherein the CRISPR repeat sequence has 100% sequence identity to SEQ ID NO: 64, and the tracrRNA has 100% sequence identity to SEQ ID NO: 65 Identity; k) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 70, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 71, and the tracrRNA has 100% sequence identity with SEQ ID NO: 72 % sequence identity; l) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 76, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 77, and the tracrRNA has 100% sequence identity with SEQ ID NO: 78 having 100% sequence identity; m) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 83, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 84, and the tracrRNA has 100% sequence identity with SEQ ID NO : 85 has 100% sequence identity; n) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 89, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 90, and the tracrRNA has 100% sequence identity with SEQ ID NO: 90 ID NO: 91 has 100% sequence identity; o) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 96, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 97, and the tracrRNA having 100% sequence identity to SEQ ID NO:98; p) an RGN polypeptide having 100% sequence identity to SEQ ID NO:103, wherein the CRISPR repeat sequence has 100% sequence identity to SEQ ID NO:104, and The tracrRNA has 100% sequence identity with SEQ ID NO: 105; q) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 110, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 111 , and the tracrRNA has 100% sequence identity with SEQ ID NO: 112; r) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 117, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 118 Consistency, and the tracrRNA has 100% sequence identity with SEQ ID NO: 119; s) RGN polypeptide with 100% sequence identity with SEQ ID NO: 123, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 124 % sequence identity, and the tracrRNA has 100% sequence identity with SEQ ID NO: 125; and t) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 83, wherein the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 84 has 100% sequence identity, and the tracrRNA has 100% sequence identity with SEQ ID NO:78.
[0251] 57. A nucleic acid molecule comprising a polynucleotide encoding a transcriptionally activated CRISPR RNA (tracrRNA). 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, 112, 119 or 125 nucleotide sequences having at least 90% sequence identity; wherein the guide RNA includes: a) the tracrRNA; and b) a crRNA comprising a spacer sequence and a CRISPR repeat sequence, wherein the tracrRNA hybridizes to the CRISPR repeat sequence of the crRNA; when the guide RNA binds to an RNA-guided nuclease (RGN) polypeptide, via the crRNA spacer sequence, capable of hybridizing to a target DNA sequence in a sequence-specific manner, and wherein the polynucleotide encoding tracrRNA is operably linked to a promoter heterologous to the polynucleotide.
[0252] 58. The nucleic acid molecule according to embodiment 57, wherein the tracrRNA comprises a sequence with SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, 112, 119 or 125 nucleotide sequences having at least 95% sequence identity.
[0253] 59. The nucleic acid molecule according to embodiment 57, wherein the tracrRNA comprises a sequence with SEQ ID NO: 3, 10, 17, 24, 31, 38, 45, 52, 58, 65, 72, 78, 85, 91, 98, 105, 112, 119 or 125 nucleotide sequences with 100% sequence identity.
[0254] 60. A vector comprising the nucleic acid molecule of any one of embodiments 57-59.
[0255] 61. The carrier of embodiment 60, wherein the carrier further comprises a polynucleotide encoding the crRNA.
[0256] 62. The carrier of embodiment 61, wherein the crRNA comprises a CRISPR repeat sequence selected from the group consisting of: a) a CRISPR having at least 90% sequence identity to SEQ ID NO:2 A repeat sequence, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO:3; b) a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO:9, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO:10 at least 90% sequence identity; c) a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 16, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 17; d) with SEQ ID NO: : 23 a CRISPR repeat sequence having at least 90% sequence identity, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 24; e) a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 30 , wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 31; f) a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 37, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 38 % sequence identity; g) a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 44, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 45; h) with SEQ ID NO: 51 A CRISPR repeat sequence having at least 90% sequence identity, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 52; i) a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 57, wherein The tracrRNA has at least 90% sequence identity with SEQ ID NO: 58; j) a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO: 64, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 65 Identity; k) a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 71, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 72; l) having at least 90% sequence identity with SEQ ID NO: 77 A CRISPR repeat sequence with 90% sequence identity, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 78; m) a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 84, wherein the tracrRNA having at least 90% sequence identity with SEQ ID NO:85; n) a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO:90, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO:91 ; o) a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO:97, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO:98; p) having at least 90% sequence identity to SEQ ID NO:104 A CRISPR repeat sequence with sequence identity, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 105; q) a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO: 111, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 111, wherein the tracrRNA has at least 90% sequence identity with SEQ ID NO: 111 ID NO: 112 has at least 90% sequence identity; r) a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 118, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 119; and s) A CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 124, wherein the tracrRNA has at least 90% sequence identity to SEQ ID NO: 125.
[0257] 63. The vector of embodiment 61, wherein the crRNA comprises a CRISPR repeat sequence selected from the group consisting of: a) a CRISPR having at least 95% sequence identity to SEQ ID NO:2 A repeat sequence, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO:3; b) a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO:9, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO:10 at least 95% sequence identity; c) a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 16, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 17; d) having at least 95% sequence identity with SEQ ID NO: 17; : 23 a CRISPR repeat sequence with at least 95% sequence identity, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 24; e) a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 30 , wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO: 31; f) a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 37, wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO: 38 % sequence identity; g) a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 44, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 45; h) with SEQ ID NO: 51 A CRISPR repeat sequence having at least 95% sequence identity, wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO: 52; i) a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 57, wherein The tracrRNA has at least 95% sequence identity with SEQ ID NO: 58; j) a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 64, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 65 Identity; k) a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 71, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 72; l) having at least 95% sequence identity with SEQ ID NO: 77 A CRISPR repeat sequence with 95% sequence identity, wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO: 78; m) a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 84, wherein the tracrRNA having at least 95% sequence identity with SEQ ID NO: 85; n) a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 90, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 91 ; o) a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO:97, wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO:98; p) having at least 95% sequence identity to SEQ ID NO:104 A CRISPR repeat sequence with sequence identity, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 105; q) a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 111, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 111, wherein the tracrRNA has at least 95% sequence identity with SEQ ID NO: 111 ID NO: 112 has at least 95% sequence identity; r) a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 118, wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO: 119; and s) A CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO:124, wherein the tracrRNA has at least 95% sequence identity to SEQ ID NO:125.
[0258] 64. The carrier of embodiment 61, wherein the crRNA comprises a CRISPR repeat sequence selected from the group consisting of: a) a CRISPR repeat with 100% sequence identity to SEQ ID NO:2 Sequence, wherein the tracrRNA has 100% sequence identity with SEQ ID NO:3; b) CRISPR repeat sequence with 100% sequence identity with SEQ ID NO:9, wherein the tracrRNA has 100% sequence with SEQ ID NO:10 Consistency; c) a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 16, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 17; d) having 100% sequence with SEQ ID NO: 23 A consistent CRISPR repeat sequence, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 24; e) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 30, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 31 has 100% sequence identity; f) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 37, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 38; g) with SEQ ID NO: 44 a CRISPR repeat sequence with 100% sequence identity, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 45; h) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 51, wherein the tracrRNA It has 100% sequence identity with SEQ ID NO: 52; i) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 57, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 58; j) A CRISPR repeat with 100% sequence identity to SEQ ID NO: 64, wherein the tracrRNA has 100% sequence identity to SEQ ID NO: 65; k) a CRISPR repeat with 100% sequence identity to SEQ ID NO: 71 Sequence, wherein the tracrRNA has 100% sequence identity with SEQ ID NO:72; l) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO:77, wherein the tracrRNA has 100% sequence identity with SEQ ID NO:78 Consistency; m) CRISPR repeat sequence having 100% sequence identity with SEQ ID NO:84, wherein the tracrRNA has 100% sequence identity with SEQ ID NO:85; n) having 100% sequence with SEQ ID NO:90 A consistent CRISPR repeat sequence, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 91; o) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 97, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 98 has 100% sequence identity; p) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 104, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 105; q) with SEQ ID NO: 111 a CRISPR repeat sequence with 100% sequence identity, wherein the tracrRNA has 100% sequence identity with SEQ ID NO: 112; r) a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 118, wherein the tracrRNA 100% sequence identity to SEQ ID NO: 119; and s) a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 124, wherein the tracrRNA has 100% sequence identity to SEQ ID NO: 125.
[0259] 65. The vector of any one of embodiments 61-64, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as single guide RNA.
[0260] 66. The vector of any one of embodiments 61-64, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.
[0261] 67. The vector according to any one of embodiments 60-66, wherein the vector further comprises a polynucleotide encoding the RGN polypeptide.
[0262] 68. The carrier of embodiment 67, wherein the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide with at least 90% sequence identity to SEQ ID NO: 1, wherein the crRNA includes CRISPR repeats Sequence, the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 2, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 3; b) has at least 90% sequence identity with SEQ ID NO: 8 Consistent RGN polypeptides, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 9, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 10; c ) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 15, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 16, and the tracrRNA is to SEQ ID NO: 16 NO: 17 has at least 90% sequence identity; d) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 22, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 23 90% sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 24; e) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 29, wherein the crRNA includes a CRISPR repeat sequence, The CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 30, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 31; f) has at least 90% sequence identity with SEQ ID NO: 36 The RGN polypeptide, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 37, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 38; g) with An RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 43, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 44, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 45 has at least 90% sequence identity; h) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 50, wherein the crRNA includes a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 51 Sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 52; i) an RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 56, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR The repeat sequence has at least 90% sequence identity with SEQ ID NO: 57, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 58; j) has any one of SEQ ID NO: 63 and 570-579 An RGN polypeptide with at least 90% sequence identity, wherein the crRNA comprises a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 64, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 65 Identity; k) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 70, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 71, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 72; l) an RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 76, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence is identical to SEQ ID NO : 77 has at least 90% sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 78; m) an RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 83, wherein the crRNA includes CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 84, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 85; n) has at least 90% sequence identity with SEQ ID NO: 89 % sequence identity RGN polypeptide, wherein the crRNA comprises a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO:90, and the tracrRNA has at least 90% sequence identity with SEQ ID NO:91 o) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 96, wherein the crRNA comprises a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 97, and the tracrRNA is identical to SEQ ID NO: 98 has at least 90% sequence identity; p) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 103, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence is identical to SEQ ID NO: 104 has at least 90% sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 105; q) an RGN polypeptide with at least 90% sequence identity with SEQ ID NO: 110, wherein the crRNA includes CRISPR repeats Sequence, the CRISPR repeat sequence has at least 90% sequence identity with SEQ ID NO: 111, and the tracrRNA has at least 90% sequence identity with SEQ ID NO: 112; r) has at least 90% sequence identity with SEQ ID NO: 117 Consistent RGN polypeptides, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 118, and the tracrRNA has at least 90% sequence identity to SEQ ID NO: 119; s ) an RGN polypeptide having at least 90% sequence identity to SEQ ID NO: 123, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 90% sequence identity to SEQ ID NO: 124, and the tracrRNA is to SEQ ID NO: 124 NO: 125 has at least 90% sequence identity; and t) an RGN polypeptide having at least 90% sequence identity with SEQ ID NO: 83, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence has SEQ ID NO: 84 At least 90% sequence identity, and the tracrRNA has at least 90% sequence identity with SEQ ID NO:78.
[0263] 69. The carrier of embodiment 67, wherein the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide with at least 95% sequence identity to SEQ ID NO: 1, wherein the crRNA includes CRISPR repeats Sequence, the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO:2, and the tracrRNA has at least 95% sequence identity with SEQ ID NO:3; b) has at least 95% sequence identity with SEQ ID NO:8 Consistent RGN polypeptides, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 9, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 10; c ) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 15, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 16, and the tracrRNA is to SEQ ID NO: 16 NO: 17 has at least 95% sequence identity; d) an RGN polypeptide having at least 95% sequence identity with SEQ ID NO: 22, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 23 95% sequence identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 24; e) an RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 29, wherein the crRNA includes a CRISPR repeat sequence, The CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 30, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 31; f) has at least 95% sequence identity with SEQ ID NO: 36 The RGN polypeptide, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 37, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 38; g) with An RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 43, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 44, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 45 has at least 95% sequence identity; h) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO:50, wherein the crRNA includes a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO:51 Sequence identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 52; i) RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 56, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR The repeat sequence has at least 95% sequence identity with SEQ ID NO: 57, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 58; j) has any one of SEQ ID NO: 63 and 570-579 An RGN polypeptide with at least 95% sequence identity, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 64, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 65 Identity; k) an RGN polypeptide having at least 95% sequence identity with SEQ ID NO: 70, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 71, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 72; l) an RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 76, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence is identical to SEQ ID NO : 77 has at least 95% sequence identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 78; m) an RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 83, wherein the crRNA includes CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 84, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 85; n) has at least 95% sequence identity with SEQ ID NO: 89 A RGN polypeptide with % sequence identity, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO:90, and the tracrRNA has at least 95% sequence identity to SEQ ID NO:91 o) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 96, wherein the crRNA comprises a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 97, and the tracrRNA is identical to SEQ ID NO: 98 has at least 95% sequence identity; p) an RGN polypeptide having at least 95% sequence identity with SEQ ID NO: 103, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence is identical to SEQ ID NO: 104 has at least 95% sequence identity, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 105; q) an RGN polypeptide with at least 95% sequence identity with SEQ ID NO: 110, wherein the crRNA includes CRISPR repeats Sequence, the CRISPR repeat sequence has at least 95% sequence identity with SEQ ID NO: 111, and the tracrRNA has at least 95% sequence identity with SEQ ID NO: 112; r) has at least 95% sequence identity with SEQ ID NO: 117 Consistent RGN polypeptides, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 118, and the tracrRNA has at least 95% sequence identity to SEQ ID NO: 119; s ) an RGN polypeptide having at least 95% sequence identity to SEQ ID NO: 123, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has at least 95% sequence identity to SEQ ID NO: 124, and the tracrRNA is to SEQ ID NO: 124 NO: 125 has at least 95% sequence identity; and t) an RGN polypeptide having at least 95% sequence identity with SEQ ID NO: 83, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence has SEQ ID NO: 84 At least 95% sequence identity, and the tracrRNA has at least 95% sequence identity to SEQ ID NO:78.
[0264] 70. The carrier of embodiment 67, wherein the RGN polypeptide is selected from the group consisting of: a) an RGN polypeptide with 100% sequence identity to SEQ ID NO: 1, wherein the crRNA includes a CRISPR repeat sequence , the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 2, and the tracrRNA has 100% sequence identity with SEQ ID NO: 3; b) RGN with 100% sequence identity with SEQ ID NO: 8 A polypeptide, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 9, and the tracrRNA has 100% sequence identity with SEQ ID NO: 10; c) with SEQ ID NO: 15 an RGN polypeptide having 100% sequence identity, wherein the crRNA comprises a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity to SEQ ID NO: 16, and the tracrRNA has 100% sequence identity to SEQ ID NO: 17 d) an RGN polypeptide with 100% sequence identity to SEQ ID NO: 22, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity to SEQ ID NO: 23, and the tracrRNA is identical to SEQ ID NO: 23 ID NO: 24 has 100% sequence identity; e) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 29, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 30 Sequence identity, and the tracrRNA has 100% sequence identity with SEQ ID NO: 31; f) RGN polypeptide with 100% sequence identity with SEQ ID NO: 36, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence It has 100% sequence identity with SEQ ID NO: 37, and the tracrRNA has 100% sequence identity with SEQ ID NO: 38; g) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 43, wherein the crRNA comprising a CRISPR repeat sequence having 100% sequence identity to SEQ ID NO: 44, and the tracrRNA having 100% sequence identity to SEQ ID NO: 45; h) having 100% sequence identity to SEQ ID NO: 50 Consistent RGN polypeptides, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 51, and the tracrRNA has 100% sequence identity with SEQ ID NO: 52; i) with SEQ ID NO:56 has an RGN polypeptide with 100% sequence identity, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO:57, and the tracrRNA has 100% sequence identity with SEQ ID NO:58 100% sequence identity; j) an RGN polypeptide having 100% sequence identity with any one of SEQ ID NO: 63 and 570-579, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence is identical to SEQ ID NO: 64 Has 100% sequence identity, and the tracrRNA has 100% sequence identity with SEQ ID NO: 65; k) RGN polypeptide with 100% sequence identity with SEQ ID NO: 70, wherein the crRNA includes a CRISPR repeat sequence, the The CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 71, and the tracrRNA has 100% sequence identity with SEQ ID NO: 72; l) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 76, Wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 77, and the tracrRNA has 100% sequence identity with SEQ ID NO: 78; m) has 100% sequence identity with SEQ ID NO: 83 an RGN polypeptide with 100% sequence identity, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity to SEQ ID NO: 84, and the tracrRNA has 100% sequence identity to SEQ ID NO: 85; n) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 89, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 90, and the tracrRNA has 100% sequence identity with SEQ ID NO : 91 has 100% sequence identity; o) an RGN polypeptide having 100% sequence identity with SEQ ID NO: 96, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 97 and the tracrRNA has 100% sequence identity with SEQ ID NO: 98; p) an RGN polypeptide with 100% sequence identity with SEQ ID NO: 103, wherein the crRNA includes a CRISPR repeat sequence, and the CRISPR repeat sequence is identical to SEQ ID NO: 103 ID NO: 104 has 100% sequence identity, and the tracrRNA has 100% sequence identity with SEQ ID NO: 105; q) RGN polypeptide with 100% sequence identity with SEQ ID NO: 110, wherein the crRNA includes CRISPR Repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 111, and the tracrRNA has 100% sequence identity with SEQ ID NO: 112; r) has 100% sequence identity with SEQ ID NO: 117 The RGN polypeptide, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 118, and the tracrRNA has 100% sequence identity with SEQ ID NO: 119; s) with SEQ ID NO: 123 RGN polypeptide with 100% sequence identity, wherein the crRNA includes a CRISPR repeat sequence, the CRISPR repeat sequence has 100% sequence identity with SEQ ID NO: 124, and the tracrRNA has 100% sequence identity with SEQ ID NO: 125 sequence identity; and t) an RGN polypeptide having 100% sequence identity to SEQ ID NO:83, wherein the crRNA includes a CRISPR repeat sequence having 100% sequence identity to SEQ ID NO:84, and the tracrRNA has 100% sequence identity to SEQ ID NO:78.
[0265] 71. A system for binding a target DNA sequence of a DNA molecule, the system comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or comprising encoding the one or more guide RNAs one or more polynucleotides of one or more nucleotide sequences of (gRNA); Any one of 70, 76, 83, 89, 96, 103, 110, 117, 123, and 570-579 is an RNA-guided nuclease (RGN) polypeptide having an amino acid sequence with at least 90% sequence identity, or comprising A polynucleotide encoding the nucleotide sequence of the RGN polypeptide; wherein at least one of the nucleotide sequence encoding the RGN polypeptide and the nucleotide sequence encoding the one or more guide RNAs is operably linked to A promoter heterologous to the nucleotide sequence; wherein the one or more guide RNAs are capable of hybridizing to the target DNA sequence, and wherein the one or more guide RNAs are capable of forming complexes with the RGN polypeptide in order to guide the The RGN polypeptide binds to the target DNA sequence of the DNA molecule.
[0266] 72. A system for binding a target DNA sequence of a DNA molecule, the system comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or comprising encoding the one or more guide RNAs one or more polynucleotides of one or more nucleotide sequences of (gRNA); Any one of 70, 76, 83, 89, 96, 103, 110, 117, 123, and 570-579 is an RNA-guided nuclease (RGN) polypeptide having an amino acid sequence with at least 90% sequence identity; wherein the One or more guide RNAs are capable of hybridizing to the target DNA sequence, and wherein the one or more guide RNAs are capable of forming complexes with the RGN polypeptide so as to guide the binding of the RGN polypeptide to the target DNA sequence of the DNA molecule.
[0267] 73. The system of embodiment 71 or 72, wherein at least one said nucleotide sequence encoding said one or more guide RNAs is operably linked to a promoter heterologous to said nucleotide sequence.
[0268] 74. The system according to any one of embodiments 71-73, wherein the RGN polypeptide comprises a sequence associated with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70 , 76, 83, 89, 96, 103, 110, 117, 123, and any of 570-579 have an amino acid sequence having at least 95% sequence identity.
[0269] 75. The system according to any one of embodiments 71-73, wherein the RGN polypeptide comprises a sequence associated with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70 , 76, 83, 89, 96, 103, 110, 117, 123, and any of 570-579 have an amino acid sequence with 100% sequence identity.
[0270] 76. The system according to any one of embodiments 71-73, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63, and has a sequence corresponding to 305 of SEQ ID NO: 63 Isoleucine at the amino acid position, valine at the amino acid position corresponding to 328 of SEQ ID NO: 63, leucine at the amino acid position corresponding to 366 of SEQ ID NO: 63 amino acid, threonine at the amino acid position corresponding to 368 of SEQ ID NO:63, and valine at the amino acid position corresponding to 405 of SEQ ID NO:63.
[0271] 77. The system of any one of embodiments 71-76, wherein the RGN polypeptide and the one or more guide RNAs are not found to be mismatched with each other in nature.
[0272] 78. The system of any one of embodiments 71-77, wherein the target DNA sequence is a eukaryotic target DNA sequence.
[0273] 79. The system of any one of embodiments 71-78, wherein the gRNA is a single guide RNA (sgRNA).
[0274] 80. The system of any one of embodiments 71-78, wherein the gRNA is a dual guide RNA.
[0275] 81. The system of any one of embodiments 71-80, wherein the gRNA is selected from the group consisting of: a) comprising a CRISPR repeat having at least 90% sequence identity to SEQ ID NO:2 Sequence and gRNA of tracrRNA having at least 90% sequence identity with SEQ ID NO: 3, wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; b) including an amino acid sequence with SEQ ID NO: 1 NO:9 has a CRISPR repeat sequence with at least 90% sequence identity and a gRNA of tracrRNA with at least 90% sequence identity with SEQ ID NO:10, wherein the RGN polypeptide includes at least 90% sequence identity with SEQ ID NO:8 c) a gRNA comprising a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO: 16 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 17, wherein the RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 15; d) comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 23 and having at least 90% sequence identity to SEQ ID NO: 24 gRNA of tracrRNA with % sequence identity, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 22; e) includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 30 CRISPR repeat sequence and gRNA of tracrRNA having at least 90% sequence identity with SEQ ID NO: 31, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 29; f) including SEQ ID NO:37 has a CRISPR repeat sequence with at least 90% sequence identity and a gRNA of tracrRNA with at least 90% sequence identity with SEQ ID NO:38, wherein the RGN polypeptide comprises at least 90% sequence identity with SEQ ID NO:36 An amino acid sequence with sequence identity; g) a gRNA comprising a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 44 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 45, wherein the The RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 43; h) comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 51 and having an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 52 gRNA of tracrRNA with at least 90% sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:50; i) comprising at least 90% sequence identity with SEQ ID NO:57 A specific CRISPR repeat sequence and a tracrRNA gRNA having at least 90% sequence identity with SEQ ID NO: 58, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 56; j) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 64 and a tracrRNA having at least 90% sequence identity to SEQ ID NO: 65, wherein the RGN polypeptide comprises sequences identical to SEQ ID NO: 63 and 570 - an amino acid sequence having at least 90% sequence identity to any of 579; k) comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 71 and at least 90% to SEQ ID NO: 72 A gRNA of tracrRNA with sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 70; l) comprising a CRISPR having at least 90% sequence identity with SEQ ID NO: 77 Repeat sequence and gRNA of tracrRNA with at least 90% sequence identity with SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 76; m) includes an amino acid sequence with SEQ ID NO: 76; ID NO:84 has a CRISPR repeat sequence with at least 90% sequence identity and a gRNA of tracrRNA with at least 90% sequence identity with SEQ ID NO:85, wherein the RGN polypeptide includes a sequence with at least 90% sequence with SEQ ID NO:83 Consistent amino acid sequence; n) gRNA comprising a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 90 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 91, wherein the RGN The polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 89; o) comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 97 and having at least 90% sequence identity to SEQ ID NO: 98 A gRNA of tracrRNA with 90% sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:96; p) comprising at least 90% sequence identity with SEQ ID NO:104 The CRISPR repeat sequence and the gRNA of tracrRNA having at least 90% sequence identity with SEQ ID NO: 105, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 103; q) includes A CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 111 and a gRNA of tracrRNA having at least 90% sequence identity with SEQ ID NO: 112, wherein the RGN polypeptide comprises at least 90% sequence identity with SEQ ID NO: 110 Amino acid sequence with % sequence identity; r) gRNA comprising a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO: 118 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 119, wherein The RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 117; s) comprising a CRISPR repeat sequence having at least 90% sequence identity to SEQ ID NO: 124 and a sequence identical to SEQ ID NO: 125 A gRNA of a tracrRNA having at least 90% sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 123; and t) comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 84 A CRISPR repeat sequence with sequence identity and a tracrRNA gRNA having at least 90% sequence identity with SEQ ID NO:78, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity with SEQ ID NO:83.
[0276] 82. The system of any one of embodiments 71-80, wherein the gRNA is selected from the group consisting of: a) comprising a CRISPR repeat having at least 95% sequence identity to SEQ ID NO:2 Sequence and gRNA of tracrRNA having at least 95% sequence identity with SEQ ID NO: 3, wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 1; b) including an amino acid sequence with SEQ ID NO: 1 NO:9 has a CRISPR repeat sequence with at least 95% sequence identity and a gRNA of tracrRNA with at least 95% sequence identity with SEQ ID NO:10, wherein the RGN polypeptide includes at least 95% sequence identity with SEQ ID NO:8 c) a gRNA comprising a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 16 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 17, wherein the RGN polypeptide comprising an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 15; d) comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 23 and having at least 95% sequence identity to SEQ ID NO: 24 gRNA of tracrRNA with % sequence identity, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 22; e) includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 30 CRISPR repeat sequence and gRNA of tracrRNA with at least 95% sequence identity with SEQ ID NO: 31, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 29; f) includes an amino acid sequence with SEQ ID NO:37 has a CRISPR repeat sequence with at least 95% sequence identity and a gRNA of tracrRNA with at least 95% sequence identity with SEQ ID NO:38, wherein the RGN polypeptide comprises at least 95% sequence identity with SEQ ID NO:36 An amino acid sequence with sequence identity; g) a gRNA comprising a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 44 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 45, wherein the The RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 43; h) comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 51 and having an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 52 gRNA of tracrRNA with at least 95% sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:50; i) comprising at least 95% sequence identity with SEQ ID NO:57 A specific CRISPR repeat sequence and a gRNA of tracrRNA having at least 95% sequence identity with SEQ ID NO: 58, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 56; j) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 64 and a tracrRNA having at least 95% sequence identity to SEQ ID NO: 65, wherein the RGN polypeptide comprises sequences identical to SEQ ID NO: 63 and 570 - an amino acid sequence having at least 95% sequence identity to any of 579; k) comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 71 and at least 95% to SEQ ID NO: 72 A gRNA of tracrRNA with sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 70; l) comprising a CRISPR having at least 95% sequence identity with SEQ ID NO: 77 Repeat sequence and gRNA of tracrRNA with at least 95% sequence identity with SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 76; m) includes an amino acid sequence with SEQ ID NO: 76; ID NO:84 has a CRISPR repeat sequence with at least 95% sequence identity and a gRNA of tracrRNA with at least 95% sequence identity with SEQ ID NO:85, wherein the RGN polypeptide includes a sequence with at least 95% sequence with SEQ ID NO:83 Consistent amino acid sequence; n) gRNA comprising a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 90 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 91, wherein the RGN The polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 89; o) comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 97 and having at least 95% sequence identity to SEQ ID NO: 98 A gRNA of tracrRNA with 95% sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:96; p) comprising at least 95% sequence identity with SEQ ID NO:104 The CRISPR repeat sequence and the gRNA of tracrRNA having at least 95% sequence identity with SEQ ID NO: 105, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 103; q) includes A CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 111 and a gRNA of tracrRNA having at least 95% sequence identity with SEQ ID NO: 112, wherein the RGN polypeptide comprises at least 95% sequence identity with SEQ ID NO: 110 Amino acid sequence with % sequence identity; r) gRNA comprising a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 118 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 119, wherein The RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 117; s) comprising a CRISPR repeat sequence having at least 95% sequence identity to SEQ ID NO: 124 and a sequence identical to SEQ ID NO: 125 A gRNA of tracrRNA having at least 95% sequence identity, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 123; and t) comprising at least 95% with SEQ ID NO: 84 A CRISPR repeat sequence with sequence identity and a gRNA of tracrRNA having at least 95% sequence identity with SEQ ID NO:78, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity with SEQ ID NO:83.
[0277] 83. The system of any one of embodiments 71-80, wherein the gRNA is selected from the group consisting of: a) comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO:2 And a gRNA of tracrRNA having 100% sequence identity with SEQ ID NO: 3, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 1; b) includes an amino acid sequence with SEQ ID NO: 9 A CRISPR repeat sequence with 100% sequence identity and a gRNA of tracrRNA with 100% sequence identity with SEQ ID NO: 10, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 8 c) a gRNA comprising a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 16 and a tracrRNA with 100% sequence identity with SEQ ID NO: 17, wherein the RGN polypeptide includes a sequence with SEQ ID NO: 15 An amino acid sequence with 100% sequence identity; d) a gRNA including a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 23 and a tracrRNA with 100% sequence identity with SEQ ID NO: 24, wherein the The RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 22; e) includes a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 30 and 100% with SEQ ID NO: 31 gRNA of tracrRNA with sequence identity, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 29; f) includes a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 37 and a gRNA of tracrRNA having 100% sequence identity with SEQ ID NO: 38, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 36; g) includes an amino acid sequence with SEQ ID NO: 44 A CRISPR repeat sequence with 100% sequence identity and a gRNA of tracrRNA with 100% sequence identity with SEQ ID NO: 45, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 43 h) a gRNA comprising a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO:51 and a tracrRNA with 100% sequence identity with SEQ ID NO:52, wherein the RGN polypeptide includes a sequence with SEQ ID NO:50 An amino acid sequence with 100% sequence identity; i) a gRNA comprising a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 57 and a tracrRNA with 100% sequence identity with SEQ ID NO: 58, wherein the The RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 56; j) includes a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 64 and 100% with SEQ ID NO: 65 The gRNA of the tracrRNA with sequence identity, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with any one of SEQ ID NO: 63 and 570-579; k) includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 71 A CRISPR repeat sequence with % sequence identity and a gRNA of tracrRNA with 100% sequence identity with SEQ ID NO: 72, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 70; l ) comprising a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 77 and a gRNA of tracrRNA with 100% sequence identity with SEQ ID NO: 78, wherein the RGN polypeptide includes 100% sequence identity with SEQ ID NO: 76 An amino acid sequence with sequence identity; m) gRNA including a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 84 and a tracrRNA with 100% sequence identity with SEQ ID NO: 85, wherein the RGN polypeptide Including an amino acid sequence having 100% sequence identity with SEQ ID NO:83; n) including a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO:90 and having 100% sequence identity with SEQ ID NO:91 gRNA of a specific tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 89; o) includes a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 97 and SEQ ID NO: 98 is a gRNA of tracrRNA with 100% sequence identity, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 96; p) includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 104 A CRISPR repeat sequence with % sequence identity and a gRNA of tracrRNA with 100% sequence identity with SEQ ID NO: 105, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 103; q ) including a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 111 and a gRNA with tracrRNA with 100% sequence identity with SEQ ID NO: 112, wherein the RGN polypeptide includes 100% with SEQ ID NO: 110 An amino acid sequence with sequence identity; r) gRNA including a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 118 and a tracrRNA with 100% sequence identity with SEQ ID NO: 119, wherein the RGN polypeptide Including an amino acid sequence with 100% sequence identity with SEQ ID NO: 117; s) including a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 124 and 100% sequence identity with SEQ ID NO: 125 gRNA of a specific tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity with SEQ ID NO: 123; and t) includes a CRISPR repeat sequence with 100% sequence identity with SEQ ID NO: 84 and A gRNA of tracrRNA having 100% sequence identity with SEQ ID NO:78, wherein the RGN polypeptide includes an amino acid sequence having 100% sequence identity with SEQ ID NO:83.
[0278] 84. The system of any one of embodiments 71-83, wherein the target DNA sequence is positioned adjacent to a prospacer adjacent motif (PAM).
[0279] 85. The system of any one of embodiments 71-84, wherein the target DNA sequence is intracellular.
[0280] 86. The system of embodiment 85, wherein the cell is a eukaryotic cell.
[0281] 87. The system of embodiment 86, wherein the eukaryotic cell is a plant cell.
[0282] 88. The system of embodiment 86, wherein the eukaryotic cell is a mammalian cell.
[0283] 89. The system of embodiment 88, wherein the mammalian cells are human cells.
[0284] 90. The system of embodiment 89, wherein the human cells are immune cells.
[0285] 91. The system of embodiment 90, wherein the immune cells are stem cells.
[0286] 92. The system of embodiment 91, wherein the stem cells are induced pluripotent stem cells.
[0287] 93. The system of embodiment 86, wherein the eukaryotic cell is an insect cell.
[0288] 94. The system of embodiment 85, wherein the cell is a prokaryotic cell.
[0289] 95. The system of any one of embodiments 71-94, wherein, when transcribed, the one or more guide RNAs are capable of hybridizing to the target DNA sequence, and the guide RNAs are capable of hybridizing to the RGN polypeptide Complexes are formed to direct cleavage of the target DNA sequence.
[0290] 96. The system of embodiment 95, wherein the cutting produces a double strand break.
[0291] 97. The system of embodiment 95, wherein the cutting produces a single strand break.
[0292] 98. The system of any one of embodiments 71-94, wherein the RGN polypeptide is nuclease inactive or a nicking enzyme.
[0293] 99. The system of any one of embodiments 71-98, wherein the RGN polypeptide is operably linked to a base editing polypeptide.
[0294] 100. The system of embodiment 99, wherein the base editing polypeptide is deaminase.
[0295] 101. The system of embodiment 100, wherein the deaminase is cytidine deaminase or adenine deaminase.
[0296] 102. The system of any one of embodiments 71-101, wherein the RGN polypeptide comprises one or more nuclear localization signals.
[0297] 103. The system of any one of embodiments 71-102, wherein the RGN polypeptide is codon-optimized for expression in eukaryotic cells.
[0298] 104. The system of any one of embodiments 71-103, wherein the nucleotide sequence encoding one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide are located on a vector.
[0299] 105. The system of any one of embodiments 71-104, wherein the system further comprises one or more donor polynucleotides or one or more genes encoding the one or more donor polynucleotides Nucleotide sequence.
[0300] 106. A pharmaceutical composition, comprising the nucleic acid molecule of any one of embodiments 1-14, 43-45 and 57-59, any of embodiments 15-21, 46-56 and 60-70 The vector of embodiment, the cell of embodiment 22, the isolated RGN polypeptide of any one of embodiments 31-42, or the system of any one of embodiments 71-105, and a pharmaceutically acceptable carrier.
[0301] 107. A method for binding a target DNA sequence of a DNA molecule, comprising delivering a system according to any one of embodiments 71-105 to the target DNA sequence or a cell comprising the target DNA sequence.
[0302] 108. The method of embodiment 107, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby allowing detection of the target DNA sequence.
[0303] 109. The method of embodiment 107, wherein the guide RNA or the RGN polypeptide further comprises an expression regulator, thereby regulating the expression of the target DNA sequence or the expression of a gene under the transcriptional control of the target DNA sequence.
[0304] 110. A method for cutting or modifying a target DNA sequence of a DNA molecule comprising delivering a system according to any one of embodiments 71-105 to the target DNA sequence or a cell comprising the DNA molecule, And cleavage or modification of the target DNA sequence occurs.
[0305] 111. The method of embodiment 110, wherein the modified target DNA sequence comprises insertion of heterologous DNA in the target DNA sequence.
[0306] 112. The method of embodiment 110, wherein the modified target DNA sequence comprises a deletion of at least one nucleotide from the target DNA sequence.
[0307] 113. The method of embodiment 110, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence.
[0308] 114. A method for binding a target DNA sequence of a DNA molecule, comprising: a) assembling an RNA-guided in vitro by combining the following conditions under conditions suitable for the formation of RGN ribonucleotide complexes Nuclease (RGN) ribonucleotide complexes: i) one or more guide RNAs capable of hybridizing to the target DNA sequence; and ii) an RGN polypeptide comprising the , 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123, and any of 570-579 have at least 90% sequence identity for the amine group and b) contacting the target DNA sequence or a cell comprising the target DNA sequence with the in vitro assembled RGN ribonucleotide complex; wherein the one or more guide RNAs hybridize to the target DNA sequence, Thus, the RGN polypeptide is guided to combine with the target DNA sequence.
[0309] 115. The method of embodiment 114, wherein the RGN polypeptide or the guide RNA further comprises a detectable label, thereby allowing detection of the target DNA sequence.
[0310] 116. The method of embodiment 114, wherein the guide RNA or the RGN polypeptide further comprises an expression regulator, thereby allowing regulation of the expression of the target DNA sequence.
[0311] 117. A method for cutting and / or modifying a target DNA sequence of a DNA molecule, comprising contacting the DNA molecule with: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN comprises the sequence of SEQ ID NO: ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579 have at least an amino acid sequence with 90% sequence identity; and b) one or more guide RNAs capable of targeting the RGN of (a) to the target DNA sequence; wherein the one or more guide RNAs hybridize to the target DNA sequence, Thereby, the RGN polypeptide is guided to combine with the target DNA sequence, and the cleavage and / or modification of the target DNA sequence occurs.
[0312] 118. The method of embodiment 117, wherein cleavage by the RGN polypeptide produces a double-stranded break.
[0313] 119. The method of embodiment 117, wherein cleavage by the RGN polypeptide produces a single-stranded break.
[0314] 120. The method of embodiment 117, wherein the RGN polypeptide is nuclease inactive or a nickase and is operably fused to a base editing polypeptide.
[0315] 121. The method of embodiment 120, wherein the base editing polypeptide is deaminase.
[0316] 122. The method of embodiment 121, wherein the deaminase is cytidine deaminase or adenine deaminase.
[0317] 123. The method of embodiment 117, wherein the modified target DNA sequence comprises insertion of heterologous DNA in the target DNA sequence.
[0318] 124. The method of embodiment 117, wherein the modified target DNA sequence comprises a deletion of at least one nucleotide from the target DNA sequence.
[0319] 125. The method of embodiment 117, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence.
[0320] 126. The method of any one of embodiments 114-125, wherein the target DNA sequence is positioned adjacent to a prospacer adjacent motif (PAM).
[0321] 127. The method of any one of embodiments 114-126, wherein the target DNA sequence is a eukaryotic target DNA sequence.
[0322] 128. The method of any one of embodiments 114-127, wherein the gRNA is a single guide RNA (gRNA).
[0323] 129. The method of any one of embodiments 114-127, wherein the gRNA is a dual guide RNA.
[0324] 130. The method of any one of embodiments 114-129, wherein the RGN comprises a sequence associated with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, Any of 76, 83, 89, 96, 103, 110, 117, 123, and 570-579 have an amino acid sequence with at least 95% sequence identity.
[0325] 131. The method of any one of embodiments 114-129, wherein the RGN comprises a sequence associated with SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, Any of 76, 83, 89, 96, 103, 110, 117, 123, and 570-579 have an amino acid sequence with 100% sequence identity.
[0326] 132. The method according to any one of embodiments 114-129, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63, and has a sequence corresponding to 305 of SEQ ID NO: 63 Isoleucine at the amino acid position, valine at the amino acid position corresponding to 328 of SEQ ID NO: 63, leucine at the amino acid position corresponding to 366 of SEQ ID NO: 63 amino acid, threonine at the amino acid position corresponding to 368 of SEQ ID NO:63, and valine at the amino acid position corresponding to 405 of SEQ ID NO:63.
[0327] 133. The method of any one of embodiments 114-129, wherein: a) the RGN has at least 90% sequence identity with SEQ ID NO: 1, and the guide RNA includes a sequence identity with SEQ ID NO: 2 A crRNA repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 3; b) the RGN has at least 90% sequence identity with SEQ ID NO: 8, and the guide RNA includes SEQ ID NO: 9 has a crRNA repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 10; c) the RGN has at least 90% sequence identity with SEQ ID NO: 15 , the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 16 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 17; d) the RGN has at least 90% sequence identity with SEQ ID NO: 22 At least 90% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 23 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 24; e) the RGN and SEQ ID NO: 29 has at least 90% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 30 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 31 ; f) the RGN has at least 90% sequence identity with SEQ ID NO:36, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO:37 and at least 90% with SEQ ID NO:38 tracrRNA with % sequence identity; g) the RGN has at least 90% sequence identity with SEQ ID NO: 43, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 44 and SEQ ID NO: 44 NO: 45 has a tracrRNA with at least 90% sequence identity; h) the RGN has at least 90% sequence identity with SEQ ID NO: 50, and the guide RNA includes a crRNA with at least 90% sequence identity with SEQ ID NO: 51 Repeat sequence and tracrRNA having at least 90% sequence identity with SEQ ID NO:52; i) the RGN has at least 90% sequence identity with SEQ ID NO:56, and the guide RNA includes at least 90% sequence identity with SEQ ID NO:57 A crRNA repeat sequence with % sequence identity and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 58; j) the RGN has at least 90% sequence identity with any of SEQ ID NO: 63 and 570-579 , the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO:64 and a tracrRNA with at least 90% sequence identity with SEQ ID NO:65; k) the RGN has at least 90% sequence identity with SEQ ID NO:70 At least 90% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 71 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 72; l) the RGN and SEQ ID NO: 76 has at least 90% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 77 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 78 m) the RGN has at least 90% sequence identity with SEQ ID NO:83, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO:84 and at least 90% with SEQ ID NO:85 tracrRNA with % sequence identity; n) the RGN has at least 90% sequence identity with SEQ ID NO: 89, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 90 and with SEQ ID NO: 91 has a tracrRNA with at least 90% sequence identity; o) the RGN has at least 90% sequence identity with SEQ ID NO: 96, and the guide RNA includes a crRNA with at least 90% sequence identity with SEQ ID NO: 97 Repeat sequence and tracrRNA having at least 90% sequence identity with SEQ ID NO:98; p) the RGN has at least 90% sequence identity with SEQ ID NO:103, and the guide RNA includes at least 90% sequence identity with SEQ ID NO:104 A crRNA repeat sequence with % sequence identity and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 105; q) the RGN has at least 90% sequence identity with SEQ ID NO: 110, and the guide RNA includes a sequence identity with SEQ ID NO: 105 NO: 111 has a crRNA repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 112; r) the RGN has at least 90% sequence identity with SEQ ID NO: 117, the The guide RNA includes a crRNA repeat sequence with at least 90% sequence identity to SEQ ID NO: 118 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 119; s) the RGN has at least 90% sequence identity to SEQ ID NO: 123 % sequence identity, the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 124 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 125; or t) the RGN and SEQ ID NO: ID NO:83 has at least 90% sequence identity, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO:84 and a tracrRNA with at least 90% sequence identity with SEQ ID NO:78.
[0328] 134. The method of any one of embodiments 114-129, wherein: a) the RGN has at least 95% sequence identity with SEQ ID NO: 1, and the guide RNA includes a sequence identity with SEQ ID NO: 2 A crRNA repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 3; b) the RGN has at least 95% sequence identity with SEQ ID NO: 8, and the guide RNA includes SEQ ID NO: 9 has a crRNA repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 10; c) the RGN has at least 95% sequence identity with SEQ ID NO: 15 , the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 16 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 17; d) the RGN has at least 95% sequence identity with SEQ ID NO: 22 At least 95% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 23 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 24; e) the RGN and SEQ ID NO: 29 has at least 95% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 30 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 31 ; f) the RGN has at least 95% sequence identity with SEQ ID NO:36, the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO:37 and at least 95% with SEQ ID NO:38 tracrRNA with % sequence identity; g) the RGN has at least 95% sequence identity with SEQ ID NO: 43, the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 44 and SEQ ID NO: 44 NO: 45 has a tracrRNA with at least 95% sequence identity; h) the RGN has at least 95% sequence identity with SEQ ID NO: 50, and the guide RNA includes a crRNA with at least 95% sequence identity with SEQ ID NO: 51 Repeat sequence and tracrRNA having at least 95% sequence identity with SEQ ID NO:52; i) the RGN has at least 95% sequence identity with SEQ ID NO:56, and the guide RNA includes at least 95% sequence identity with SEQ ID NO:57 A crRNA repeat sequence with % sequence identity and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 58; j) the RGN has at least 95% sequence identity with any of SEQ ID NO: 63 and 570-579 , the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 64 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 65; k) the RGN has at least 95% sequence identity with SEQ ID NO: 70 At least 95% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 71 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 72; l) the RGN and SEQ ID NO: 76 has at least 95% sequence identity, the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 77 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 78 m) the RGN has at least 95% sequence identity with SEQ ID NO:83, the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO:84 and at least 95% with SEQ ID NO:85 tracrRNA with % sequence identity; n) the RGN has at least 95% sequence identity with SEQ ID NO: 89, and the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 90 and SEQ ID NO: 90 NO: 91 has a tracrRNA with at least 95% sequence identity; o) the RGN has at least 95% sequence identity with SEQ ID NO: 96, and the guide RNA includes a crRNA with at least 95% sequence identity with SEQ ID NO: 97 Repeat sequence and tracrRNA having at least 95% sequence identity with SEQ ID NO:98; p) the RGN has at least 95% sequence identity with SEQ ID NO:103, and the guide RNA includes at least 95% sequence identity with SEQ ID NO:104 A crRNA repeat sequence with % sequence identity and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 105; q) the RGN has at least 95% sequence identity with SEQ ID NO: 110, and the guide RNA includes a sequence identity with SEQ ID NO: 110 NO: 111 has a crRNA repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 112; r) the RGN has at least 95% sequence identity with SEQ ID NO: 117, the The guide RNA includes a crRNA repeat with at least 95% sequence identity to SEQ ID NO: 118 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 119; s) the RGN has at least 95% sequence identity to SEQ ID NO: 123 % sequence identity, the guide RNA includes a crRNA repeat with at least 95% sequence identity to SEQ ID NO: 124 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 125; or t) the RGN is identical to SEQ ID NO: 125 ID NO:83 has at least 95% sequence identity, and the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO:84 and a tracrRNA with at least 95% sequence identity with SEQ ID NO:78.
[0329] 135. The method of any one of embodiments 114-129, wherein: a) the RGN has 100% sequence identity to SEQ ID NO: 1, and the guide RNA comprises 100% sequence identity to SEQ ID NO: 2 crRNA repeat sequence with % sequence identity and tracrRNA with 100% sequence identity with SEQ ID NO: 3; b) the RGN has 100% sequence identity with SEQ ID NO: 8, and the guide RNA includes SEQ ID NO: 9 crRNA repeats with 100% sequence identity and tracrRNA with 100% sequence identity with SEQ ID NO: 10; c) the RGN has 100% sequence identity with SEQ ID NO: 15, and the guide RNA includes SEQ ID NO: 15 with 100% sequence identity ID NO: 16 has a crRNA repeat sequence with 100% sequence identity and a tracrRNA with 100% sequence identity with SEQ ID NO: 17; d) the RGN has 100% sequence identity with SEQ ID NO: 22, and the guide RNA including a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 23 and a tracrRNA with 100% sequence identity with SEQ ID NO: 24; e) the RGN has 100% sequence identity with SEQ ID NO: 29, The guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 30 and a tracrRNA with 100% sequence identity with SEQ ID NO: 31; f) the RGN has 100% sequence with SEQ ID NO: 36 Consistency, the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 37 and a tracrRNA with 100% sequence identity with SEQ ID NO: 38; g) the RGN and SEQ ID NO: 43 have 100% sequence identity, the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 44 and a tracrRNA with 100% sequence identity with SEQ ID NO: 45; h) the RGN and SEQ ID NO : 50 has 100% sequence identity, and the guide RNA includes a crRNA repe...
Claims
1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide, the RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579; wherein, When bound to a guide RNA (gRNA) capable of hybridizing with a target DNA sequence, the RGN polypeptide can bind to the target DNA sequence in an RNA-guided sequence-specific manner, and the polynucleotide encoding the RGN polypeptide therein is operatively linked to a promoter heterologous to the polynucleotide.
2. The nucleic acid molecule as claimed in claim 1, wherein the RGN polypeptide comprises a monoamino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
3. The nucleic acid molecule as claimed in claim 1, wherein the RGN polypeptide comprises a monoamino acid sequence having 100% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
4. The nucleic acid molecule as claimed in claim 1, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63 and has an isoleucine at the amino acid position corresponding to SEQ ID NO: 63 at position 305, a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 328, a leucine at the amino acid position corresponding to SEQ ID NO: 63 at position 366, a threonine at the amino acid position corresponding to SEQ ID NO: 63 at position 368, and a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 405.
5. The nucleic acid molecule of any one of claims 1 to 4, wherein the RGN polypeptide is capable of cleaving the target DNA sequence upon binding.
6. The nucleic acid molecule as claimed in claim 5, wherein the RGN polypeptide is capable of producing a double-strand break.
7. The nucleic acid molecule as claimed in claim 5, wherein the RGN polypeptide is capable of producing a single strand break.
8. The nucleic acid molecule of any one of claims 1 to 4, wherein the RGN polypeptide is a nuclease inactive or a nuclease.
9. The nucleic acid molecule of any one of claims 1 to 8, wherein the RGN polypeptide is operatively fused with a base-editing polypeptide.
10. The nucleic acid molecule as claimed in claim 9, wherein the base-editing polypeptide is a deaminase.
11. The nucleic acid molecule of any one of claims 1 to 10, wherein the RGN polypeptide includes one or more nuclear localization signals.
12. The nucleic acid molecule of any one of claims 1 to 11, wherein the RGN polypeptide is codon-optimized for expression in a eukaryotic cell.
13. The nucleic acid molecule of any one of claims 1 to 2, wherein the target DNA sequence is located adjacent to a prespacer sequence neighbor motif (PAM).
14. A vector comprising any one of claims 1 to 13.
15. The vector as claimed in claim 14, further comprising at least one nucleotide sequence encoding the gRNA, the gRNA being capable of hybridizing with the target DNA sequence.
16. The vector as claimed in claim 15, wherein the guide RNA is selected from the group consisting of: a) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 2; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 3; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; b) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 9; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 10; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 8; c) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 16; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 2; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; NO: 17 has at least 90% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 15; d) a guide RNA, comprising: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 23; and ii) a tracrRNA, the tracrRNA having at least 90% sequence identity with SEQ ID NO: 24; wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 22; e) a guide RNA, comprising: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 30; and ii) a tracrRNA, the tracrRNA having at least 90% sequence identity with SEQ ID NO: 31; wherein the RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 29; f) a guide RNA, comprising: i) a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 29; The CRISPR RNA comprises a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 37; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 38;The RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 36; g) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 44; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 45; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 43; h) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 51; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 52; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50; i) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 51; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 52; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50; i) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 36; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 36; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50; NO: 57 is a CRISPR repeat sequence with at least 90% sequence identity; and ii) a tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 58; wherein the RGN polypeptide comprises an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 56; j) a guide RNA, comprising: i) a CRISPR RNA, which comprises a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO: 64; and ii) a tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 65; wherein the RGN polypeptide comprises an amino acid sequence with at least 90% sequence identity with any of SEQ ID NO: 63 and 570-579; k) a guide RNA, comprising: i) a CRISPR RNA, which comprises a CRISPR repeat sequence with at least 90% sequence identity with SEQ ID NO: 71; and ii) a tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 58; wherein the RGN polypeptide comprises an amino acid sequence with at least 90% sequence identity with SEQ ID NO: 56; NO: 72 has at least 90% sequence identity; wherein the RGN polypeptide includes a monoamino acid sequence that has at least 90% sequence identity with SEQ ID NO: 70;l) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 77; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 76; m) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 84; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 85; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 83; n) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 77; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 83; NO: 90 A CRISPR repeat sequence having at least 90% sequence identity; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 91; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 89; o) a guide RNA comprising: i) a CRISPR RNA having at least 90% sequence identity with SEQ ID NO: 97; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 98; wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 96; p) a guide RNA comprising: i) a CRISPR RNA having at least 90% sequence identity with SEQ ID NO: 104; and ii) a tracrRNA having at least 90% sequence identity with SEQ ID NO: 105; The RGN polypeptide includes an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 103; q) a guide RNA, including: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 111;and ii) a tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 112; wherein the RGN polypeptide includes an amino acid sequence that has at least 90% sequence identity with SEQ ID NO: 110; r) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence that has at least 90% sequence identity with SEQ ID NO: 118; and ii) a tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 119; wherein the RGN polypeptide includes an amino acid sequence that has at least 90% sequence identity with SEQ ID NO: 117; s) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence that has at least 90% sequence identity with SEQ ID NO: 124; and ii) a tracrRNA, which has at least 90% sequence identity with SEQ ID NO: 125; wherein the RGN polypeptide includes an amino acid sequence that has at least 90% sequence identity with SEQ ID NO: 110; NO: 123 has an amino acid sequence with at least 90% sequence identity; and t) a guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 84; and ii) a tracrRNA comprising at least 90% sequence identity to SEQ ID NO: 78; wherein the RGN polypeptide comprises an amino acid sequence with at least 90% sequence identity to SEQ ID NO:
83.
17. The vector as claimed in claim 15, wherein the guide RNA is selected from the group consisting of: a) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 2; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 3; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 1; b) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 9; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 10; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 8; c) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 16; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 2; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 1; NO: 17 has at least 95% sequence identity; wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 15; d) a guide RNA, comprising: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 23; and ii) a tracrRNA, the tracrRNA having at least 95% sequence identity with SEQ ID NO: 24; wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 22; e) a guide RNA, comprising: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 30; and ii) a tracrRNA, the tracrRNA having at least 95% sequence identity with SEQ ID NO: 31; wherein the RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 29; f) a guide RNA, comprising: i) a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 29; The CRISPR RNA comprises a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 37; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 38;The RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 36; g) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 44; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 45; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 43; h) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 51; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 52; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 50; i) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 51; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 52; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 50; i) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 36; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 36; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 50; NO: 57 is a CRISPR repeat sequence with at least 95% sequence identity; and ii) a tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 58; wherein the RGN polypeptide comprises an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 56; j) a guide RNA, comprising: i) a CRISPR RNA, which comprises a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 64; and ii) a tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 65; wherein the RGN polypeptide comprises an amino acid sequence with at least 95% sequence identity with any of SEQ ID NO: 63 and 570-579; k) a guide RNA, comprising: i) a CRISPR RNA, which comprises a CRISPR repeat sequence with at least 95% sequence identity with SEQ ID NO: 71; and ii) a tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 58; wherein the RGN polypeptide comprises an amino acid sequence with at least 95% sequence identity with SEQ ID NO: 56; NO: 72 has at least 95% sequence identity; wherein the RGN polypeptide includes a monoamino acid sequence that has at least 95% sequence identity with SEQ ID NO: 70;l) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 77; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 76; m) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 84; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 85; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 83; n) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 77; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 83; NO: 90 A CRISPR repeat sequence having at least 95% sequence identity; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 91; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 89; o) a guide RNA comprising: i) a CRISPR RNA having at least 95% sequence identity with SEQ ID NO: 97; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 98; wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 96; p) a guide RNA comprising: i) a CRISPR RNA having at least 95% sequence identity with SEQ ID NO: 104; and ii) a tracrRNA having at least 95% sequence identity with SEQ ID NO: 105; The RGN polypeptide includes an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 103; q) a guide RNA, including: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 111;and ii) a tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 112; wherein the RGN polypeptide includes an amino acid sequence that has at least 95% sequence identity with SEQ ID NO: 110; r) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence that has at least 95% sequence identity with SEQ ID NO: 118; and ii) a tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 119; wherein the RGN polypeptide includes an amino acid sequence that has at least 95% sequence identity with SEQ ID NO: 117; s) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence that has at least 95% sequence identity with SEQ ID NO: 124; and ii) a tracrRNA, which has at least 95% sequence identity with SEQ ID NO: 125; wherein the RGN polypeptide includes an amino acid sequence that has at least 95% sequence identity with SEQ ID NO: 117; NO: 123 has an amino acid sequence with at least 95% sequence identity; and t) a guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 84; and ii) a tracrRNA comprising at least 95% sequence identity to SEQ ID NO: 78; wherein the RGN polypeptide comprises an amino acid sequence with at least 95% sequence identity to SEQ ID NO:
83.
18. The vector as claimed in claim 15, wherein the guide RNA is selected from the group consisting of: a) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 2; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 3; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 1; b) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 9; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 10; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 8; c) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 16; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 3; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 1; NO: 17 has 100% sequence identity; wherein the RGN polypeptide includes an amino acid sequence that has 100% sequence identity with SEQ ID NO: 15; d) a guide RNA, comprising: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence that has 100% sequence identity with SEQ ID NO: 23; and ii) a tracrRNA, the tracrRNA having 100% sequence identity with SEQ ID NO: 24; wherein the RGN polypeptide includes an amino acid sequence that has 100% sequence identity with SEQ ID NO: 22; e) a guide RNA, comprising: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence that has 100% sequence identity with SEQ ID NO: 30; and ii) a tracrRNA, the tracrRNA having 100% sequence identity with SEQ ID NO: 31; wherein the RGN polypeptide includes an amino acid sequence that has 100% sequence identity with SEQ ID NO: 29; f) a guide RNA, comprising: i) a CRISPR repeat sequence that has 100% sequence identity with SEQ ID NO: 29; The CRISPR RNA comprises a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 37; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 38;The RGN polypeptide comprises an amino acid sequence that is 100% sequence identical to SEQ ID NO: 36; g) a guide RNA comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence that is 100% sequence identical to SEQ ID NO: 44; and ii) a tracrRNA comprising ... NO: 57 is a CRISPR repeat sequence with 100% sequence identity; and ii) a tracrRNA, which is 100% sequence identical to SEQ ID NO: 58; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 56; j) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 64; and ii) a tracrRNA, which is 100% sequence identical to SEQ ID NO: 65; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to any one of SEQ ID NO: 63 and 570-579; k) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 71; and ii) a tracrRNA, which is 100% sequence identical to SEQ ID NO: 58; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 56; NO: 72 has 100% sequence identity; wherein the RGN polypeptide includes an amino acid sequence that has 100% sequence identity with SEQ ID NO: 70; l) a guide RNA, including: i) a CRISPR RNA, the CRISPR RNA including a CRISPR repeat sequence that has 100% sequence identity with SEQ ID NO: 77;and ii) a tracrRNA, which has 100% sequence identity with SEQ ID NO: 78; wherein the RGN polypeptide includes an amino acid sequence that has 100% sequence identity with SEQ ID NO: 76; m) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence that has 100% sequence identity with SEQ ID NO: 84; and ii) a tracrRNA, which has 100% sequence identity with SEQ ID NO: 85; wherein the RGN polypeptide includes an amino acid sequence that has 100% sequence identity with SEQ ID NO: 83; n) a guide RNA, comprising: i) a CRISPR RNA, which includes a CRISPR repeat sequence that has 100% sequence identity with SEQ ID NO: 90; and ii) a tracrRNA, which has 100% sequence identity with SEQ ID NO: 91; wherein the RGN polypeptide includes an amino acid sequence that has 100% sequence identity with SEQ ID NO: 89; o) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 97; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 98; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 96; p) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 104; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 105; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 103; q) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 97; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 98; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 96; NO: 111 is a CRISPR repeat sequence with 100% sequence identity; and ii) a tracrRNA with 100% sequence identity to SEQ ID NO: 112; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 110;r) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 118; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 119; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 117; s) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 124; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 125; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 123; and t) A guide RNA, comprising: i) a CRISPR RNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 118; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 119 ... a guide RNA, comprising: i) a CRISPR RNA having 100% sequence identity with SEQ ID NO: 118; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 119; wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 119; and ii) a tracrRNA having 100% sequence identity with SEQ ID NO: 84 is a CRISPR repeat sequence with 100% sequence identity; and ii) a tracrRNA with 100% sequence identity to SEQ ID NO: 78; wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO:
83.
19. The vector as described in any one of claims 15 to 18, wherein the gRNA is a single guide RNA.
20. The vector as claimed in any one of claims 15 to 18, wherein the gRNA is a dual guide RNA.
21. A cell comprising a nucleic acid molecule as described in any one of claims 1 to 13 or a vector as described in any one of claims 14 to 20.
22. A method for manufacturing an RGN polypeptide, comprising: Cells as described in claim 21 were cultured under conditions in which the RGN polypeptide was expressed.
23. A method for manufacturing an RGN polypeptide, comprising introducing a heterologous nucleic acid molecule into a cell, the heterologous nucleic acid molecule comprising a nucleotide sequence encoding an RNA-guided nuclease (RGN) polypeptide, the RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579; wherein, When the RGN polypeptide binds to a guide RNA (gRNA) capable of hybridizing with a target DNA sequence, it can bind to the target DNA sequence in an RNA-guided sequence-specific manner; and the cells are cultured under conditions in which the RGN polypeptide is expressed.
24. The method of claim 23, wherein the RGN polypeptide comprises a monoamino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
25. The method of claim 23, wherein the RGN polypeptide comprises a monoamino acid sequence having 100% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
26. The method of claim 23, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63 and has an isoleucine at the amino acid position corresponding to SEQ ID NO: 63 at position 305, a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 328, a leucine at the amino acid position corresponding to SEQ ID NO: 63 at position 366, a threonine at the amino acid position corresponding to SEQ ID NO: 63 at position 368, and a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 405.
27. The method of any one of claims 22 to 26, further comprising purifying the RGN polypeptide.
28. The method of any one of claims 22 to 26, wherein the cell further expresses one or more guide RNAs capable of binding to the RGN polypeptide to form an RGN ribonucleoprotein complex.
29. The method of claim 28 further includes purifying the RGN ribonucleoprotein complex.
30. An isolated RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide comprises a monoamino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579; and wherein, When bound to a guide RNA (gRNA) capable of hybridizing with a target DNA sequence, the RGN polypeptide can bind to the target DNA sequence of a DNA molecule in an RNA-guided sequence-specific manner.
31. The isolated RGN polypeptide as claimed in claim 30, wherein the RGN polypeptide comprises a monoamino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
32. The isolated RGN polypeptide as claimed in claim 30, wherein the RGN polypeptide comprises a monoamino acid sequence having 100% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
33. The isolated RGN polypeptide as claimed in claim 30, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63 and has an isoleucine at the amino acid position corresponding to SEQ ID NO: 63 at position 305, a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 328, a leucine at the amino acid position corresponding to SEQ ID NO: 63 at position 366, a threonine at the amino acid position corresponding to SEQ ID NO: 63 at position 368, and a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 405.
34. The isolated RGN polypeptide as claimed in any one of claims 30 to 33, wherein the RGN polypeptide is capable of cleaving the target DNA sequence upon binding.
35. The isolated RGN polypeptide as claimed in claim 34, wherein a double strand break is generated by cleavage of the RGN polypeptide.
36. The isolated RGN polypeptide as claimed in claim 34, wherein cleavage of the RGN polypeptide produces a single strand break.
37. The isolated RGN polypeptide as claimed in any one of claims 30 to 33, wherein the RGN polypeptide is either nuclease inactive or an allolytic enzyme.
38. The isolated RGN polypeptide as claimed in any one of claims 30 to 37, wherein the RGN polypeptide is operatively fused with a base-editing polypeptide.
39. The isolated RGN polypeptide as described in claim 38, wherein the base-editing polypeptide is a deaminase.
40. The isolated RGN polypeptide as claimed in any one of claims 30 to 39, wherein the target DNA sequence is located adjacent to a prespacer sequence neighbor motif (PAM).
41. The isolated RGN polypeptide as claimed in any one of claims 30 to 40, wherein the RGN polypeptide includes one or more nuclear localization signals.
42. A system for binding a labeled DNA sequence to a DNA molecule, the system comprising: a) one or more guide RNAs capable of hybridizing with the target DNA sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity with any of SEQ ID NOs: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579, or a polynucleotide comprising a nucleotide sequence encoding the RGN polypeptide; wherein at least one of the nucleotide sequence encoding the RGN polypeptide and the nucleotide sequence encoding the one or more guide RNAs is operatively linked to a promoter heterologous to the nucleotide sequence; wherein the one or more guide RNAs are capable of hybridizing with the target DNA sequence, and wherein the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to direct the RGN polypeptide to bind to the target DNA sequence of the DNA molecule.
43. A system for binding a labeled DNA sequence to a DNA molecule, the system comprising: a) one or more guide RNAs capable of hybridizing with the target DNA sequence, or one or more polynucleotides comprising one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); and b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579; wherein the one or more guide RNAs are capable of hybridizing with the target DNA sequence, and wherein the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to direct the RGN polypeptide to bind to the target DNA sequence of the DNA molecule.
44. The system as described in claim 42 or claim 43, wherein at least one of the nucleotide sequences encoding the one or more guide RNAs is operatively linked to a promoter heterologous to the nucleotide sequence.
45. The system of any one of claims 42 to 44, wherein the RGN polypeptide comprises a monoamino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
46. The system of any one of claims 42 to 44, wherein the RGN polypeptide comprises a monoamino acid sequence having 100% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
47. The system of any one of claims 42 to 44, wherein the RGN polypeptide has at least 95% sequence identity with SEQ ID NO: 63 and has an isoleucine at the amino acid position corresponding to SEQ ID NO: 63 at position 305, a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 328, a leucine at the amino acid position corresponding to SEQ ID NO: 63 at position 366, a threonine at the amino acid position corresponding to SEQ ID NO: 63 at position 368, and a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 405.
48. The system of any one of claims 42 to 47, wherein the RGN polypeptide is not found to be substantially misaligned with the one or more guide RNAs.
49. The system of any one of claims 42 to 48, wherein the target DNA sequence is a eukaryotic target DNA sequence.
50. The system of any one of claims 42 to 49, wherein the gRNA is a single guide RNA (sgRNA).
51. The system of any one of claims 42 to 49, wherein the gRNA is a dual guide RNA.
52. The system of any one of claims 42 to 51, wherein the gRNA is selected from the group consisting of: a) a gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 2 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 3, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 1; b) a gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 9 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 10, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 8; c) a gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 16 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 17, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 15; d) a gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 15 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 16; NO: 23 is a CRISPR repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 24, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 22; e) is a gRNA including a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 30 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 31, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 29; f) is a gRNA including a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 37 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 38, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 36; g) is a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 44 and a tracrRNA with at least 90% sequence identity to SEQ ID NO:
24. NO:45 is a gRNA with at least 90% sequence identity to a tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO:43;h) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 51 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 50; i) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 57 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 58, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 56; j) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 64 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 65, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 63 and 570-579; k) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 51 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 63 and 570-579; NO: 71 is a CRISPR repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 72, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 70; l) is a gRNA including a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 77 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 76; m) is a gRNA including a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 84 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 85, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 83; n) is a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 90 and a tracrRNA with at least 90% sequence identity to SEQ ID NO:
72. NO: 91 is a gRNA with at least 90% sequence identity to a tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 89;o) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 97 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 98, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 96; p) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 104 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 105, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 103; q) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 111 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 112, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 110; r) A gRNA comprising a CRISPR repeat sequence having at least 90% sequence identity with SEQ ID NO: 97 and a tracrRNA having at least 90% sequence identity with SEQ ID NO: 98, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 110; NO: 118 is a CRISPR repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 119, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 117; s) includes a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 124 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 125, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO: 123; and t) includes a CRISPR repeat sequence with at least 90% sequence identity to SEQ ID NO: 84 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with at least 90% sequence identity to SEQ ID NO:
83.
53. The system of any one of claims 42 to 51, wherein the gRNA is selected from the group consisting of: a) a gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 2 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 3, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 1; b) a gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 9 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 10, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 8; c) a gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 16 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 17, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 15; d) a gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 15 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 16; NO: 23 is a CRISPR repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 24, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 22; e) is a gRNA including a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 30 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 31, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 29; f) is a gRNA including a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 37 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 38, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 36; g) is a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 44 and a tracrRNA with at least 95% sequence identity to SEQ ID NO:
24. NO:45 is a gRNA with at least 95% sequence identity to a tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO:43;h) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 51 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 50; i) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 57 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 58, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 56; j) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 64 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 65, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 63 and 570-579; k) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 51 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 63 and 570-579; NO: 71 is a CRISPR repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 72, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 70; l) is a gRNA including a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 77 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 76; m) is a gRNA including a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 84 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 85, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 83; n) is a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 90 and a tracrRNA with at least 95% sequence identity to SEQ ID NO:
72. NO: 91 is a gRNA with at least 95% sequence identity to a tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 89;o) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 97 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 98, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 96; p) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 104 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 105, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 103; q) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 111 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 112, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 110; r) A gRNA comprising a CRISPR repeat sequence having at least 95% sequence identity with SEQ ID NO: 97 and a tracrRNA having at least 95% sequence identity with SEQ ID NO: 98, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 110; NO: 118 is a CRISPR repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 119, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 117; s) includes a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 124 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 125, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO: 123; and t) includes a CRISPR repeat sequence with at least 95% sequence identity to SEQ ID NO: 84 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with at least 95% sequence identity to SEQ ID NO:
83.
54. The system of any one of claims 42 to 51, wherein the gRNA is selected from the group consisting of: a) a gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 2 and a tracrRNA having 100% sequence identity with SEQ ID NO: 3, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 1; b) a gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 9 and a tracrRNA having 100% sequence identity with SEQ ID NO: 10, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 8; c) a gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 16 and a tracrRNA having 100% sequence identity with SEQ ID NO: 17, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 15; d) a gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 15 and a tracrRNA having 100% sequence identity with SEQ ID NO: 16; NO: 23 is a CRISPR repeat sequence with 100% sequence identity and a tracrRNA with 100% sequence identity to SEQ ID NO: 24, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 22; e) is a gRNA including a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 30 and a tracrRNA with 100% sequence identity to SEQ ID NO: 31, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 29; f) is a gRNA including a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 37 and a tracrRNA with 100% sequence identity to SEQ ID NO: 38, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 36; g) is a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 44 and a tracrRNA with SEQ ID NO:
25. NO: 45 is a gRNA with 100% sequence identity to a tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 43;h) A gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 51 and a tracrRNA having 100% sequence identity with SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 50; i) A gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 57 and a tracrRNA having 100% sequence identity with SEQ ID NO: 58, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with SEQ ID NO: 56; j) A gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 64 and a tracrRNA having 100% sequence identity with SEQ ID NO: 65, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NO: 63 and 570-579; k) A gRNA comprising a CRISPR repeat sequence having 100% sequence identity with SEQ ID NO: 51 and a tracrRNA having 100% sequence identity with SEQ ID NO: 52, wherein the RGN polypeptide comprises an amino acid sequence having 100% sequence identity with any one of SEQ ID NO: 63 and 570-579; NO: 71 is a CRISPR repeat sequence with 100% sequence identity and a tracrRNA with 100% sequence identity to SEQ ID NO: 72, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 70; l) is a gRNA including a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 77 and a tracrRNA with 100% sequence identity to SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 76; m) is a gRNA including a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 84 and a tracrRNA with 100% sequence identity to SEQ ID NO: 85, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 83; n) is a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 90 and a tracrRNA with SEQ ID NO:
72. NO: 91 is a gRNA with 100% sequence identity to a tracrRNA, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 89;o) A gRNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 97 and a tracrRNA with 100% sequence identity to SEQ ID NO: 98, wherein the RGN polypeptide comprises an amino acid sequence with 100% sequence identity to SEQ ID NO: 96; p) A gRNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 104 and a tracrRNA with 100% sequence identity to SEQ ID NO: 105, wherein the RGN polypeptide comprises an amino acid sequence with 100% sequence identity to SEQ ID NO: 103; q) A gRNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 111 and a tracrRNA with 100% sequence identity to SEQ ID NO: 112, wherein the RGN polypeptide comprises an amino acid sequence with 100% sequence identity to SEQ ID NO: 110; r) A gRNA comprising a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 97 and a tracrRNA with 100% sequence identity to SEQ ID NO: 98, wherein the RGN polypeptide comprises an amino acid sequence with 100% sequence identity to SEQ ID NO: 110; NO: 118 is a CRISPR repeat sequence with 100% sequence identity and a tracrRNA with 100% sequence identity to SEQ ID NO: 119, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 117; s) includes a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 124 and a tracrRNA with 100% sequence identity to SEQ ID NO: 125, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO: 123; and t) includes a CRISPR repeat sequence with 100% sequence identity to SEQ ID NO: 84 and a tracrRNA with 100% sequence identity to SEQ ID NO: 78, wherein the RGN polypeptide includes an amino acid sequence with 100% sequence identity to SEQ ID NO:
83.
55. The system of any one of claims 42 to 54, wherein the target DNA sequence is located adjacent to a prespacer sequence neighbor motif (PAM).
56. The system of any one of claims 42 to 55, wherein the target DNA sequence is contained within a cell.
57. The system as described in claim 56, wherein the cell is a eukaryotic cell.
58. The system as claimed in claim 57, wherein the eukaryotic cell is a plant cell.
59. The system as claimed in claim 57, wherein the eukaryotic cell is a mammalian cell.
60. The system as claimed in claim 57, wherein the eukaryotic cell is an insect cell.
61. The system as claimed in claim 56, wherein the cell is a prokaryotic cell.
62. The system as described in any one of claims 42 to 61, wherein, When transcribed, the one or more guide RNAs can hybridize with the target DNA sequence and form a complex with the RGN polypeptide to guide the cleavage of the target DNA sequence.
63. The system as described in claim 62, wherein the cutting produces a double-strand fracture.
64. The system as described in claim 62, wherein the cutting produces a single-strand fracture.
65. The system of any one of claims 42 to 61, wherein the RGN polypeptide is a nuclease that is inactive or a nuclease.
66. The system of any one of claims 42 to 65, wherein the RGN polypeptide is operatively linked to a base-editing polypeptide.
67. The system as claimed in claim 66, wherein the base-editing polypeptide is a deaminase.
68. The system of any one of claims 42 to 67, wherein the RGN polypeptide comprises one or more nuclear localization signals.
69. The system of any one of claims 42 to 68, wherein the RGN polypeptide is codon-optimized for performance in a eukaryotic cell.
70. The system of any one of claims 42 to 69, wherein the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding an RGN polypeptide are located on a vector.
71. The system of any one of claims 42 to 70, wherein the system further comprises one or more donor polynucleotides or one or more nucleotide sequences encoding the one or more donor polynucleotides.
72. A pharmaceutical composition comprising a nucleic acid molecule as described in any one of claims 1 to 13, a carrier as described in any one of claims 14 to 20, a cell as described in claim 21, an isolated RGN polypeptide as described in any one of claims 30 to 41, or a system as described in any one of claims 42 to 71, and a pharmaceutically acceptable carrier.
73. A method for binding a target DNA sequence to a DNA molecule, comprising delivering a system according to any one of claims 42 to 31 to the target DNA sequence or a cell comprising the target DNA sequence.
74. The method of claim 73, wherein the RGN polypeptide or the guide RNA further comprises a detectable marker, thereby allowing detection of the target DNA sequence.
75. The method of claim 73, wherein the guide RNA or the RGN polypeptide further comprises an expression regulator that regulates the expression of a gene of the target DNA sequence or under the transcriptional control of the target DNA sequence.
76. A method for cutting or modifying a target DNA sequence of a DNA molecule, comprising delivering a system according to any one of claims 42 to 71 to the target DNA sequence or a cell comprising the DNA molecule, and wherein cutting or modification of the target DNA sequence occurs.
77. The method of claim 76, wherein the modified target DNA sequence includes the insertion of a heterologous DNA within the target DNA sequence.
78. The method of claim 76, wherein the modified target DNA sequence includes at least one nucleotide deletion from the target DNA sequence.
79. The method of claim 76, wherein the modified target DNA sequence includes a mutation of at least one nucleotide in the target DNA sequence.
80. A method for binding a labeled DNA sequence to a DNA molecule, comprising: a) Under conditions suitable for the formation of the RGN ribonucleotide complex, an RNA-guided nuclease (RGN) ribonucleotide complex is assembled in vitro by combining the following: i) one or more guide RNAs capable of hybridizing with the target DNA sequence; and ii) an RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579; and b) contacting the target DNA sequence or a cell comprising the target DNA sequence with the in vitro assembled RGN ribonucleotide complex; wherein the one or more guide RNAs hybridize with the target DNA sequence, thereby directing the RGN polypeptide to bind to the target DNA sequence.
81. The method of claim 80, wherein the RGN polypeptide or the guide RNA further comprises a detectable marker that allows for detection of the target DNA sequence.
82. The method of claim 80, wherein the guide RNA or the RGN polypeptide further comprises an expression regulator that allows for regulation of the expression of the target DNA sequence.
83. A method for cleaving and / or modifying a target DNA sequence of a DNA molecule, comprising contacting the DNA molecule with: a) an RNA-guided nuclease (RGN) polypeptide, wherein the RGN comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579; and b) one or more guide RNAs capable of targeting the RGN of (a) to the target DNA sequence; wherein the one or more guide RNAs hybridize with the target DNA sequence, thereby directing the RGN polypeptide to bind to the target DNA sequence, and causing cleavage and / or modification of the target DNA sequence.
84. The method of claim 83, wherein a double strand break is generated by cleavage of the RGN polypeptide.
85. The method of claim 83, wherein a single strand break is generated by cleavage of the RGN polypeptide.
86. The method of claim 83, wherein the RGN polypeptide is a nuclease-inactive or alloenzyme and is operatively linked to a base-editing polypeptide.
87. The method of claim 86, wherein the base-editing polypeptide comprises a deaminase.
88. The method of claim 83, wherein the modified target DNA sequence includes the insertion of a heterologous DNA within the target DNA sequence.
89. The method of claim 83, wherein the modified target DNA sequence comprises at least one nucleotide deletion from the target DNA sequence.
90. The method of claim 83, wherein the modified target DNA sequence includes a mutation of at least one nucleotide in the target DNA sequence.
91. The method of any one of claims 80 to 90, wherein the target DNA sequence is located adjacent to a prespacer sequence neighbor motif (PAM).
92. The method of any one of claims 80 to 91, wherein the target DNA sequence is a eukaryotic target DNA sequence.
93. The method of any one of claims 80 to 92, wherein the gRNA is a single guide RNA (gRNA).
94. The method of any one of claims 80 to 92, wherein the gRNA is a dual guide RNA.
95. The method of any one of claims 80 to 94, wherein the RGN comprises a monoamino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
96. The method of any one of claims 80 to 94, wherein the RGN comprises a monoamino acid sequence having 100% sequence identity with any one of SEQ ID NO: 1, 8, 15, 22, 29, 36, 43, 50, 56, 63, 70, 76, 83, 89, 96, 103, 110, 117, 123 and 570-579.
97. The method of any one of claims 80 to 94, wherein the RGN polypeptide has at least 90% sequence identity with SEQ ID NO: 63 and has an isoleucine at the amino acid position corresponding to SEQ ID NO: 63 at position 305, a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 328, a leucine at the amino acid position corresponding to SEQ ID NO: 63 at position 366, a threonine at the amino acid position corresponding to SEQ ID NO: 63 at position 368, and a valine at the amino acid position corresponding to SEQ ID NO: 63 at position 405.
98. The method as described in any one of claims 80 to 94, wherein: a) The RGN has at least 90% sequence identity with SEQ ID NO: 1, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 2 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 3; b) The RGN has at least 90% sequence identity with SEQ ID NO: 8, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 9 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 10; c) The RGN has at least 90% sequence identity with SEQ ID NO: 15, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 16 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 17; d) The RGN has at least 90% sequence identity with SEQ ID NO: 22, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 23 and a tracrRNA with at least 90% sequence identity with SEQ ID NO:
14. e) The guide RNA has at least 90% sequence identity with SEQ ID NO: 24 and a tracrRNA with at least 90% sequence identity; f) The guide RNA has at least 90% sequence identity with SEQ ID NO: 39 and includes a tracrRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 30 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 31; g) The guide RNA has at least 90% sequence identity with SEQ ID NO: 36 and includes a tracrRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 37 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 38; g) The guide RNA has at least 90% sequence identity with SEQ ID NO: 43 and includes a tracrRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 44 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 45; h) The guide RNA has at least 90% sequence identity with SEQ ID NO: 50 and includes a tracrRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 39 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 30 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 31; NO: 51 is a crRNA repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 52;i) The RGN has at least 90% sequence identity with SEQ ID NO: 56, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 57 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 58; j) The RGN has at least 90% sequence identity with any one of SEQ ID NO: 63 and 570-579, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 64 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 65; k) The RGN has at least 90% sequence identity with SEQ ID NO: 70, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 71 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 72; l) The RGN has at least 90% sequence identity with SEQ ID NO: 76, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 57 and a tracrRNA with at least 90% sequence identity with SEQ ID NO:
58. NO: 77 has a crRNA repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 78; m) The RGN has at least 90% sequence identity to SEQ ID NO: 83, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity to SEQ ID NO: 84 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 85; n) The RGN has at least 90% sequence identity to SEQ ID NO: 89, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity to SEQ ID NO: 90 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 91; o) The RGN has at least 90% sequence identity to SEQ ID NO: 96, and the guide RNA includes a crRNA repeat sequence with at least 90% sequence identity to SEQ ID NO: 97 and a tracrRNA with at least 90% sequence identity to SEQ ID NO: 98; p) The RGN has at least 90% sequence identity to SEQ ID NO: 77 and a tracrRNA with at least 90% sequence identity to SEQ ID NO:
98. NO: 103 has at least 90% sequence identity, and the guide RNA includes a crRNA repeat sequence that has at least 90% sequence identity with SEQ ID NO: 104 and a tracrRNA that has at least 90% sequence identity with SEQ ID NO: 105;q) The RGN has at least 90% sequence identity with SEQ ID NO: 110, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 111 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 112; r) The RGN has at least 90% sequence identity with SEQ ID NO: 117, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 118 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 119; s) The RGN has at least 90% sequence identity with SEQ ID NO: 123, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 124 and a tracrRNA with at least 90% sequence identity with SEQ ID NO: 125; or t) The RGN has at least 90% sequence identity with SEQ ID NO: 83, and the guide RNA comprises a crRNA repeat sequence with at least 90% sequence identity with SEQ ID NO: 110 and a tracrRNA with at least 90% sequence identity with SEQ ID NO:
112. NO:84 is a crRNA repeat sequence with at least 90% sequence identity and a tracrRNA with at least 90% sequence identity to SEQ ID NO:
78.
99. The method as described in any one of claims 80 to 94, wherein: a) The RGN has at least 95% sequence identity with SEQ ID NO: 1, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 2 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 3; b) The RGN has at least 95% sequence identity with SEQ ID NO: 8, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 9 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 10; c) The RGN has at least 95% sequence identity with SEQ ID NO: 15, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 16 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 17; d) The RGN has at least 95% sequence identity with SEQ ID NO: 22, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 23 and a tracrRNA with at least 95% sequence identity with SEQ ID NO:
10. e) The guide RNA has at least 95% sequence identity with SEQ ID NO: 24 and a tracrRNA with at least 95% sequence identity; f) The guide RNA has at least 95% sequence identity with SEQ ID NO: 39 and includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 30 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 31; g) The guide RNA has at least 95% sequence identity with SEQ ID NO: 36 and includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 37 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 38; g) The guide RNA has at least 95% sequence identity with SEQ ID NO: 43 and includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 44 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 45; h) The guide RNA has at least 95% sequence identity with SEQ ID NO: 50 and includes a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 39 and a tracrRNA with at least 95% sequence identity with SEQ ID NO:
30. NO: 51 is a crRNA repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 52;i) The RGN has at least 95% sequence identity with SEQ ID NO: 56, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 57 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 58; j) The RGN has at least 95% sequence identity with any one of SEQ ID NO: 63 and 570-579, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 64 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 65; k) The RGN has at least 95% sequence identity with SEQ ID NO: 70, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 71 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 72; l) The RGN has at least 95% sequence identity with SEQ ID NO: 76, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 57 and a tracrRNA with at least 95% sequence identity with SEQ ID NO:
58. NO: 77 has a crRNA repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 78; m) The RGN has at least 95% sequence identity to SEQ ID NO: 83, and the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity to SEQ ID NO: 84 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 85; n) The RGN has at least 95% sequence identity to SEQ ID NO: 89, and the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity to SEQ ID NO: 90 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 91; o) The RGN has at least 95% sequence identity to SEQ ID NO: 96, and the guide RNA includes a crRNA repeat sequence with at least 95% sequence identity to SEQ ID NO: 97 and a tracrRNA with at least 95% sequence identity to SEQ ID NO: 98; p) The RGN has at least 95% sequence identity to SEQ ID NO: 77 and a tracrRNA with at least 95% sequence identity to SEQ ID NO:
98. NO: 103 has at least 95% sequence identity, and the guide RNA includes a crRNA repeat sequence that has at least 95% sequence identity with SEQ ID NO: 104 and a tracrRNA that has at least 95% sequence identity with SEQ ID NO: 105;q) The RGN has at least 95% sequence identity with SEQ ID NO: 110, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 111 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 112; r) The RGN has at least 95% sequence identity with SEQ ID NO: 117, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 118 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 119; s) The RGN has at least 95% sequence identity with SEQ ID NO: 123, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 124 and a tracrRNA with at least 95% sequence identity with SEQ ID NO: 125; or t) The RGN has at least 95% sequence identity with SEQ ID NO: 83, and the guide RNA comprises a crRNA repeat sequence with at least 95% sequence identity with SEQ ID NO: 110 and a tracrRNA with at least 95% sequence identity with SEQ ID NO:
112. NO:84 is a crRNA repeat sequence with at least 95% sequence identity and a tracrRNA with at least 95% sequence identity to SEQ ID NO:
78.
100. The method as described in any one of claims 80 to 94, wherein: a) The RGN has 100% sequence identity with SEQ ID NO: 1, and the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 2 and a tracrRNA with 100% sequence identity with SEQ ID NO: 3; b) The RGN has 100% sequence identity with SEQ ID NO: 8, and the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 9 and a tracrRNA with 100% sequence identity with SEQ ID NO: 10; c) The RGN has 100% sequence identity with SEQ ID NO: 15, and the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 16 and a tracrRNA with 100% sequence identity with SEQ ID NO: 17; d) The RGN has 100% sequence identity with SEQ ID NO: 22, and the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 23 and a tracrRNA with SEQ ID NO:
14. e) The RGN is 100% sequence identical to SEQ ID NO: 24, and the guide RNA comprises a crRNA repeat sequence that is 100% sequence identical to SEQ ID NO: 30 and a tracrRNA that is 100% sequence identical to SEQ ID NO: 31; f) The RGN is 100% sequence identical to SEQ ID NO: 36, and the guide RNA comprises a crRNA repeat sequence that is 100% sequence identical to SEQ ID NO: 37 and a tracrRNA that is 100% sequence identical to SEQ ID NO: 38; g) The RGN is 100% sequence identical to SEQ ID NO: 43, and the guide RNA comprises a crRNA repeat sequence that is 100% sequence identical to SEQ ID NO: 44 and a tracrRNA that is 100% sequence identical to SEQ ID NO: 45; h) The RGN is 100% sequence identical to SEQ ID NO: 50, and the guide RNA comprises a crRNA repeat sequence that is 100% sequence identical to SEQ ID NO: 29 and a tracrRNA that is 100% sequence identical to SEQ ID NO: 30 and a tracrRNA that is 100% sequence identical to SEQ ID NO: 31; NO: 51 has a 100% sequence identical crRNA repeat sequence and a tracrRNA with 100% sequence identical SEQ ID NO: 52; i) The RGN has 100% sequence identical SEQ ID NO: 56, and the guide RNA includes a 100% sequence identical crRNA repeat sequence and a tracrRNA with 100% sequence identical SEQ ID NO: 58;j) The RGN has 100% sequence identity with any of SEQ ID NO: 63 and 570-579, and the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 64 and a tracrRNA with 100% sequence identity with SEQ ID NO: 65; k) The RGN has 100% sequence identity with SEQ ID NO: 70, and the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 71 and a tracrRNA with 100% sequence identity with SEQ ID NO: 72; l) The RGN has 100% sequence identity with SEQ ID NO: 76, and the guide RNA includes a crRNA repeat sequence with 100% sequence identity with SEQ ID NO: 77 and a tracrRNA with 100% sequence identity with SEQ ID NO: 78; m) The RGN has 100% sequence identity with SEQ ID NO: 83, and the guide RNA includes a crRNA repeat sequence with SEQ ID NO: 64 and a tracrRNA with 100% sequence identity with SEQ ID NO:
65. NO: 84 has a 100% sequence identical crRNA repeat sequence and a tracrRNA with 100% sequence identical SEQ ID NO: 85; n) The RGN has 100% sequence identical SEQ ID NO: 89, and the guide RNA includes a 100% sequence identical crRNA repeat sequence with SEQ ID NO: 90 and a 100% sequence identical tracrRNA with SEQ ID NO: 91; o) The RGN has 100% sequence identical SEQ ID NO: 96, and the guide RNA includes a 100% sequence identical crRNA repeat sequence with SEQ ID NO: 97 and a 100% sequence identical tracrRNA with SEQ ID NO: 98; p) The RGN has 100% sequence identical SEQ ID NO: 103, and the guide RNA includes a 100% sequence identical crRNA repeat sequence with SEQ ID NO: 104 and a 100% sequence identical tracrRNA with SEQ ID NO: 105; q) The RGN has 100% sequence identical SEQ ID NO: 89 and a tracrRNA with SEQ ID NO:
85. NO: 110 has 100% sequence identity. The guide RNA includes a crRNA repeat sequence that has 100% sequence identity with SEQ ID NO: 111 and a tracrRNA that has 100% sequence identity with SEQ ID NO: 112.r) The RGN is 100% sequence identical to SEQ ID NO: 117, and the guide RNA includes a crRNA repeat sequence that is 100% sequence identical to SEQ ID NO: 118 and a tracrRNA that is 100% sequence identical to SEQ ID NO: 119; s) The RGN is 100% sequence identical to SEQ ID NO: 123, and the guide RNA includes a crRNA repeat sequence that is 100% sequence identical to SEQ ID NO: 124 and a tracrRNA that is 100% sequence identical to SEQ ID NO: 125; or t) The RGN is 100% sequence identical to SEQ ID NO: 83, and the guide RNA includes a crRNA repeat sequence that is 100% sequence identical to SEQ ID NO: 84 and a tracrRNA that is 100% sequence identical to SEQ ID NO:
78.
101. The method of any one of claims 73 to 100, wherein the target DNA sequence is in a cell.
102. The method as described in claim 101, wherein the cell is a eukaryotic cell.
103. The method as described in claim 102, wherein the eukaryotic cell is a plant cell.
104. The method as described in claim 102, wherein the eukaryotic cell is a mammalian cell.
105. The method as described in claim 102, wherein the eukaryotic cell is an insect cell.
106. The method as described in claim 101, wherein the cell is a prokaryotic cell.
107. The method as described in any one of claims 101 to 106, further comprising: The cells are cultured under conditions in which the RGN polypeptide is expressed and the target DNA sequence is cleaved to produce a DNA molecule comprising a modified DNA sequence; and cells comprising the modified target DNA sequence are selected.
108. A cell comprising a modified target DNA sequence as described in claim 107.
109. The cell as claimed in claim 108, wherein the cell is a eukaryotic cell.
110. The cell as claimed in claim 109, wherein the eukaryotic cell is a plant cell.
111. A plant comprising cells as described in claim 110.
112. A seed comprising cells as described in claim 110.
113. The cell as claimed in claim 109, wherein the eukaryotic cell is a mammalian cell.
114. The cell as claimed in claim 113, wherein the mammalian cell is a human cell.
115. The cell as described in claim 114, wherein the human cell is an immune cell.
116. The cell as claimed in claim 115, wherein the immune cell is a stem cell.
117. The cell as claimed in claim 116, wherein the stem cell is an induced pluripotent stem cell.
118. The cell as claimed in claim 109, wherein the eukaryotic cell is an insect cell.
119. The cell as claimed in claim 108, wherein the cell is a prokaryotic cell.
120. A pharmaceutical composition comprising cells as described in any one of claims 108, 109 and 113 to 117, and a pharmaceutically acceptable carrier.
121. A method of treating a disease, the method comprising administering to an individual requiring treatment an effective amount of a pharmaceutical composition as described in claim 72 or claim 120.
122. The method of claim 121, wherein the disease is associated with a causal mutation, and the effective amount of the pharmaceutical composition corrects the gene mutation.