RNA-Induced Nucleases, Active Fragments and Variants Thereof, and Methods of Use
RNA-guided nucleases address the inefficiencies of existing genome editing methods by using guide RNAs for precise sequence targeting and modification, enhancing editing precision and reducing errors.
Patent Information
- Application Number
- JP2021518031
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-13
- Filing Date
- 2019-06-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-06-04
AI Technical Summary
Existing genome editing technologies, such as meganucleases and CRISPR-Cas systems, require costly and inefficient generation of chimeric nucleases for each target sequence, and non-homologous end joining introduces errors during DNA repair.
Development of RNA-guided nucleases (RGNs) that utilize guide RNAs to specifically target and cleave or modify genomic sequences, enabling precise genome editing through homologous recombination and reducing errors by using nuclease-dead variants for expression modulation or payload delivery.
RGNs provide efficient, cost-effective, and precise genome editing with reduced error rates, allowing for targeted gene modification and expression control.
Smart Images

Figure 0007708660000001 
Figure 0007708660000002 
Figure 0007708660000003
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of molecular biology and gene editing.
Background Art
[0002] Target genome editing or target genome modification is rapidly becoming an important tool for basic research and applied research. Initial methods involved engineering nucleases (such as meganucleases, zinc finger proteins, TALENs, etc.), which required creating chimeric nucleases with engineered programmable sequence-specific DNA-binding domains specific to each particular target sequence. RNA-guided nucleases (such as CRISPR-associated (Cas) proteins of the clustered regularly interspaced short palindromic repeats (CRISPR)-Cas bacterial system) can target specific sequences by forming a complex with a guide RNA that specifically hybridizes to the target sequence. The generation of target-specific guide RNAs is less costly and more efficient than generating chimeric nucleases for each target sequence. It is possible to edit the genome by introducing sequence-specific double-strand breaks using such RNA-guided nucleases, and since non-homologous end joining (NHEJ), which repairs these double-strand breaks, is error-prone, mutations are introduced at specific genomic locations. Alternatively, heterologous DNA can be introduced into genomic sites through homologous recombination repair.
Summary of the Invention
[0003] Compositions and methods are provided for binding to a target sequence of interest. The compositions have uses for cleaving or modifying the target sequence of interest, visualizing the target sequence of interest, and altering the expression of the target sequence of interest. The compositions include an RNA-guided nuclease (RGN) polypeptide, a CRISPR RNA (crRNA), a trans-activating CRISPR RNA (tracrRNA), a guide RNA (gRNA), and nucleic acid molecules encoding these. Vectors and host cells contain these nucleic acid molecules. A CRISPR system is also provided that includes an RNA-guided nuclease polypeptide and one or more guide RNAs for binding to a target sequence of interest. Thus, the methods disclosed herein relate to binding to a target sequence of interest, and in some embodiments, the target sequence of interest is cleaved or modified by this method. The target sequence of interest can be modified, for example, as a result of non-homologous end joining or homologous recombination repair using an introduced donor sequence.
DETAILED DESCRIPTION OF THE INVENTION
[0004] Those skilled in the art of the field to which the present invention pertains will likely come up with numerous modifications of the present invention described herein and other embodiments that have the advantages of the teachings shown in the above description and related drawings. Therefore, it should be understood that the present invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are included within the scope of the appended embodiments. Special terms are used herein, but they are used in a general and explanatory sense and not for the purpose of limitation.
[0005] I. Overview
[0006] RNA-guided nucleases (RGNs) enable the manipulation of a single site within a genome, making RGNs useful in the context of gene targeting for therapeutic and research applications. In a variety of organisms, including mammals, RNA-guided nucleases have been used for genome engineering, for example, by stimulating non-homologous end joining and homologous recombination. The compositions and methods described herein are useful for causing single-stranded or double-stranded breaks in polynucleotides, or for modifying polynucleotides, or for detecting specific sites within polynucleotides, or for changing the expression of specific genes.
[0007] The RNA-guided nucleases disclosed herein can change gene expression by modifying the target sequence. In particular embodiments, the RNA-guided nuclease is directed to the target sequence by a guide RNA (gRNA) as part of a clustered regularly interspaced short palindromic repeat (CRISPR) RNA-guided nuclease system. The guide RNA forms a complex with the RNA-guided nuclease to bind the RNA-guided nuclease to the target sequence and, in some embodiments, introduce a single-stranded or double-stranded break at the position of the target sequence. After the target sequence is cleaved, the DNA sequence of the target sequence can be modified during the repair process of the cleavage. Accordingly, provided herein are methods of modifying a target sequence in the DNA of a host cell using an RNA-guided nuclease. For example, an RNA-guided nuclease can be used to modify a target sequence at a genomic locus in a eukaryotic or prokaryotic cell.
[0008] II. RNA-guided nuclease
[0009] This specification presents RNA-guided nucleases. The term "RNA-guided nuclease" (RGN) refers to a polypeptide that binds to a specific target nucleotide sequence in a sequence-specific manner and is directed to the target nucleotide sequence by a guide RNA molecule that forms a complex with this polypeptide and hybridizes with this target sequence. An RNA-guided nuclease can cleave the target sequence when bound to it, but the term "RNA-guided nuclease" also encompasses nuclease-dead RNA-guided nucleases that can bind to the target sequence but cannot cleave the target sequence. As a result of cleavage of the target sequence by an RNA-guided nuclease, it is possible to cleave a single strand or a double strand. An RNA-guided nuclease capable of cleaving a single strand of a double-stranded nucleic acid molecule is referred to herein as a nickase.
[0010] The RNA-guided nucleases disclosed in this specification include RNA-guided nucleases such as APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, APG1688.1 (the amino acid sequences of which are represented by SEQ ID NO: 1, 11, 19, 27, 36, 45, 54, respectively), and those that retain the ability to bind to a target nucleotide sequence in a manner specific to the RNA-guided sequence with an active fragment or variant thereof. In some of these embodiments, the active fragments or variants of the RGNs such as APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, APG1688.1 can cleave a single-stranded target sequence or a double-stranded target sequence. In some embodiments, the active variants of the RGNs such as APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, APG1688.1 include an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity with the amino acid sequence represented by any of SEQ ID NO: 1, 11, 19, 27, 36, 45, 54. In some embodiments, the active fragments of the RGNs such as APG05083.1, APG07433.1, APG07513.1, APG08290.1, APG05459.1, APG04583.1, APG1688.1 include at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more consecutive amino acid residues of the amino acid sequence represented by any of SEQ ID NO: 1, 11, 19, 27, 36, 45, 54.The RNA-guided nucleases presented herein can include at least one nuclease domain (e.g., a DNase domain, an RNase domain) and at least one RNA recognition and / or RNA binding domain that interacts with a guide RNA. Non-limiting examples of additional domains that can be found among the RNA-guided nucleases presented herein include DNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. In particular embodiments, the RNA-guided nucleases presented herein can include at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% of one or more of a DNA binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain.
[0011] The target nucleotide sequence is bound by the RNA-guided nucleases presented herein and the guide RNA related to this RNA-guided nuclease hybridizes thereto. Subsequently, if the RNA-guided nuclease has nuclease activity, the target sequence can be cleaved by this RNA-guided nuclease. The term "cleave" or "cleavage" means that at least one phosphodiester bond within the backbone of the target sequence is hydrolyzed, resulting in the cleavage of a single-stranded or double-stranded region within the target nucleic acid. The RGNs of the present disclosure can function as endonucleases to cleave nucleotides within a polynucleotide or as exonucleases to continuously remove nucleotides from the ends (5' end and / or 3' end) of a polynucleotide. In another embodiment, the disclosed RGNs can function as both endonucleases and exonucleases because they can cleave nucleotides of a target sequence at any position within a polynucleotide. As a result of cleavage of the target polynucleotide by the RGNs of the present disclosure, sticky ends or blunt ends may be generated.
[0012] As the RNA-guided nuclease disclosed in this specification, a wild-type sequence derived from a bacterial or archaeal species is possible. Alternatively, as the RNA-guided nuclease, a variant or fragment of a wild-type polypeptide is possible. The wild-type RGN can be modified to change, for example, nuclease activity or PAM specificity. In some embodiments, the RNA-guided nuclease is not natural.
[0013] In some embodiments, the RNA-guided nuclease functions as a nickase and cleaves only a single strand of the target nucleotide sequence. Such an RNA-guided nuclease has a single nuclease domain that functions. In some of these embodiments, an additional nuclease domain is mutated and the nuclease activity is reduced or eliminated.
[0014] In another embodiment, the RNA-guided nuclease has completely lost nuclease activity or exhibits reduced nuclease activity, and is thus referred to herein as nuclease-dead (inactivated) in this specification. Any method known in the art for introducing mutations into an amino acid sequence (such as PCR-mediated mutagenesis, site-directed mutagenesis, etc.) can be utilized to generate a nickase or a nuclease-dead RGN. See, for example, U.S. Patent Application Publication No. 2014 / 0068797 and U.S. Patent No. 9,790,490 (each of which is incorporated herein by reference in its entirety).
[0015] An RNA-guided nuclease lacking nuclease activity can be used to deliver any of a fused polypeptide, polynucleotide, or small molecule payload to a specific genomic location. In some of these embodiments, the RGN polypeptide or guide RNA is fused to a detectable label to enable detection of a specific sequence. As a non-limiting example, a nuclease-dead RGN is fused to a detectable label (such as a fluorescent protein) and targets a specific sequence related to a disease, enabling detection of the disease-related sequence.
[0016] Alternatively, a nuclease-dead RGN can be directed to a specific genomic location to alter the expression of a desired sequence. In some embodiments, when a nuclease-dead RGN-inducible nuclease binds to a target sequence, it interferes with the binding of RNA polymerase or a transcription factor within the targeted genomic region, thereby suppressing the expression of the target sequence or the expression of a gene under the transcriptional control of the target sequence. In another embodiment, an RGN (e.g., a nuclease-dead RGN), or a guide RNA complexed with an RGN, further comprises an expression modulator that, when bound to a target sequence, suppresses or activates the expression of the target sequence or the expression of a gene under the transcriptional control of the target sequence. In some of these embodiments, the expression modulator alters the expression of the target sequence or a gene regulated through an epigenetic mechanism.
[0017] In another embodiment, the sequence of a target polynucleotide can be modified by directing a nuclease-dead RGN, or an RGN having only nickase activity, to a specific genomic location and fusing it to a base editing polypeptide (e.g., a deaminase polypeptide, or one that deaminates nucleotide bases with an active variant or fragment thereof), resulting in the conversion of one nucleotide to another. The base editing polypeptide can be fused to the N-terminus or C-terminus of the RGN. In addition, the base editing polypeptide can be fused to the RGN via a peptide linker. Non-limiting examples of deaminase polypeptides useful in such compositions and methods include the cytidine deaminase base editors or adenosine deaminase base editors described in Gaudelli et al. (2017) Nature 551:464-471, US Patent Application Publication Nos. 2017 / 0121693 and 2018 / 0073012, and International Application Publication No. WO / 2018 / 027078 (each of which is incorporated herein by reference in its entirety).
[0018] The RNA-guided nuclease fused to a polypeptide or domain can be separated or joined by a linker. As used herein, the term "linker" refers to a chemical group or molecule that connects two molecules or moieties (e.g., the binding domain and cleavage domain of a nuclease). In some embodiments, the linker joins the gRNA binding domain of the RNA-guided nuclease to a base editing polypeptide (such as a deaminase). In some embodiments, the linker connects a nuclease-dead RGN to a deaminase. Typically, the linker is positioned between or adjacent to two groups, molecules, or other moieties and is connected to each through a covalent bond, thereby connecting the two. In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is any of an organic molecule, a group, a polymer, or a chemical moiety. In some embodiments, the linker is an amino acid having a length of 5 to 100 amino acids, such as 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30 - 35, 35 - 40, 40 - 45, 45 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, 90 - 100, 100 - 150, or 150 - 200 amino acids. Longer or shorter linkers are also contemplated.
[0019] The RNA-guided nucleases of the present disclosure can include at least one nuclear localization signal (NLS) to enhance the transport of the RGN to the cell nucleus. Nuclear localization signals are known in the art and generally contain a series of basic amino acids (see, for example, Lange et al., J. Biol. Chem. (2007) Vol. 282: pages 5101-5105). In particular embodiments, the RGN contains two, three, four, five, six, or more nuclear localization signals. Heterologous NLSs are possible as nuclear localization signals. Non-limiting examples of nuclear localization signals useful for the RGNs of the present disclosure are the nuclear localization signals of SV40 large T antigen, nucleopasmin, and c-Myc (see, for example, Ray et al. (2015) Bioconjug Chem Vol. 26(6): pages 1004-1007). In a particular embodiment, the RGN contains the NLS sequence shown as SEQ ID NO: 67. The RGN can contain one or more NLS sequences at the N-terminus, or at the C-terminus, or both at the N-terminus and the C-terminus. For example, the RGN can contain two NLS sequences in the N-terminal region and four NLS sequences in the C-terminal region.
[0020] To target the RGN, other localization signals known in the art for localizing polypeptides to specific intracellular locations can also be used, non-limiting examples of which include plastid (chromoplast) localization sequences, mitochondrial localization sequences, and dual-targeting signal sequences that target both plastids and mitochondria (see, for example, Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol dx.doi.org / 10.3389 / fphys.2015.00259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soll (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al. (2009) FEBS J 276:1187-1195; Silva-Filho (2003) Curr Opin Plant Biol 6:589-595; Peeters and Small (2001) Biochim Biophys Acta 1541:54-63; Murcha et al. (2014) J Exp Bot 65:6301-6335; Mackenzie (2005) Trends Cell Biol 15:548-554; Glaser et al. (1998) Plant Mol Biol 38:311-338).
[0021] In some embodiments, the RNA-guided nucleases of the present disclosure include at least one cell-penetrating domain that facilitates the uptake of the RGN into cells. Cell-penetrating domains are known in the art and generally include a series of positively charged amino acid residues (i.e., polycationic cell-penetrating domains), or alternatively arranged polar and nonpolar amino acid residues (i.e., amphipathic cell-penetrating domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-penetrating domains) (see, for example, Milletti F. (2012) Drug Discov Today Vol. 17: pages 850-860). A non-limiting example of a cell-penetrating domain is the trans-activating transcriptional activator (TAT) from human immunodeficiency virus 1.
[0022] The nuclear localization signal, and / or plastid localization sequence, and / or mitochondrial localization sequence, and / or dual-targeting signal sequence, and / or cell-penetrating domain can be located at the amino terminus (N-terminus), or carboxyl terminus (C-terminus), or internal position of the RNA-guided nuclease.
[0023] The RGNs of the present disclosure can be directly fused to an effector domain (such as a cleavage domain, deaminase domain, expression modulator domain, etc.) or indirectly fused via a linker peptide. Such domains can be located at the N-terminus, or C-terminus, or internal position of the RNA-guided nuclease. In some of these embodiments, the RGN element of the fusion protein is a nuclease-dead RGN.
[0024] In some embodiments, the RGN fusion protein comprises a cleavage domain. The cleavage domain is any domain capable of cleaving a polynucleotide (i.e., either RNA, DNA, or an RNA / DNA hybrid), and non-limiting examples thereof include restriction endonucleases and homing endonucleases such as type IIS endonucleases (e.g., FokI) (see, e.g., Belfort et al. (1997) Nucleic Acids Res. Vol. 25:3379 - 3388; Linn et al. (eds.) "Nucleases", Cold Spring Harbor Laboratory Press, 1993).
[0025] In another embodiment, the RGN fusion protein comprises a deaminase domain that deaminates a nucleotide base to convert one nucleotide base to another nucleotide base, and non-limiting examples thereof include cytidine deaminase base editors or adenosine deaminase base editors (see, e.g., Gaudelli et al. (2017) Nature Vol. 551:464 - 471, U.S. Patent Application Publication Nos. 2017 / 0121693, 2018 / 0073012, U.S. Patent No. 9,840,699, International Application Publication No. WO / 2018 / 027078).
[0026] In some embodiments, an expression modulator domain can be used as the effector domain of the RGN fusion protein. The expression modulator domain is a domain having a function of upregulating or downregulating transcription. As the expression modulator domain, any of an epigenetic modification domain, a transcription repression domain, and a transcription activation domain can be used.
[0027] In some of these embodiments, the expression modulator of the RGN fusion protein comprises an epigenetic modification domain that modifies the covalent bonds of DNA or histone proteins to change the histone structure and / or chromosomal structure without changing the DNA sequence, thereby changing (upregulating or downregulating) gene expression. Non-limiting examples of epigenetic modifications include acetylation or methylation of lysine residues in DNA, methylation of arginine, phosphorylation of serine and threonine, ubiquitination of lysine, SUMOylation of histone proteins, methylation and hydroxymethylation of cytosine residues. Non-limiting examples of epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains, and DNA demethylase domains.
[0028] In another embodiment, the expression modulator of the fusion protein comprises a transcriptional repression domain. The transcriptional repression domain interacts with transcriptional control elements and / or transcriptional regulatory proteins (such as RNA polymerase, transcription factors, etc.) to reduce or stop the transcription of at least one gene. Transcriptional repression domains are known in the art, and non-limiting examples thereof include Sp1-like repressors, IκB, and Kruppel-associated box (KRAB) domains.
[0029] In yet another embodiment, the expression modulator of the fusion protein comprises a transcriptional activation domain. The transcriptional activation domain interacts with transcriptional control elements and / or transcriptional regulatory proteins (such as RNA polymerase, transcription factors, etc.) to increase or activate the transcription of at least one gene. Transcriptional activation domains are known in the art, and non-limiting examples thereof include the herpes simplex virus VP16 activation domain and the NFAT activation domain.
[0030] The RGN polypeptides of the present disclosure can include a detectable label or a purification tag. The detectable label or purification tag can be positioned directly at the N-terminus, or C-terminus, or an internal position of the RNA-guided nuclease, or indirectly via a linker peptide. In some of these embodiments, the RGN element of the fusion protein is a nuclease-dead RGN. In another embodiment, the RGN element of the fusion protein is an RGN having nickase activity.
[0031] The detectable label can be visualized or otherwise observed. The detectable label can be fused to the RGN as a fusion protein (e.g., a fluorescent protein), or made into a small molecule that is complexed with the RGN polypeptide and is detectable visually or by other means. Detectable labels that can be fused to the RGN of the present disclosure as a fusion protein include any protein domain that is detectable, non-limiting examples of which include fluorescent proteins, protein domains detectable using specific antibodies. Non-limiting examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, EGFP, ZsGreen1) and yellow fluorescent proteins (e.g., YFP, EYFP, ZsYellow1). Non-limiting examples of detectable small molecule labels include radioactive labels 3 H, 35 S, etc.
[0032] The RGN polypeptide can also include a purification tag. This is any molecule that can be used to isolate a protein or fusion protein from a mixture (e.g., a biological sample, a medium). Non-limiting examples of purification tags include myc, maltose binding protein (MBP), glutathione-S-transferase (GST).
[0033] II. Guide RNA
[0034] Provided by the present disclosure are guide RNAs and polynucleotides encoding the same. The term "guide RNA" means a nucleotide sequence that hybridizes to a target nucleotide sequence and binds an RNA-guided nuclease to the target nucleotide sequence in a sequence-specific manner because it has sufficient complementarity to the target nucleotide sequence. Thus, each guide RNA of an RGN is one or more RNA molecules (generally one or two) that can bind to the RGN to guide this RGN and bind to a specific target nucleotide sequence, and when the RGN has nickase activity or nuclease activity, the target nucleotide sequence is also cleaved. Generally, a guide RNA contains a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). A natural guide RNA containing both crRNA and tracrRNA generally contains two separate RNA molecules that hybridize to each other through the repeat sequence of crRNA and the anti-repeat sequence of tracrRNA.
[0035] The natural tandem repeat sequences in the CRISPR array generally range in length from 28 to 37 base pairs, although this length can vary between about 23 bp and about 55 bp. The spacer sequences in the CRISPR array generally range in length from 32 to 38 base pairs, although this length can vary between about 21 bp and about 72 bp. Each CRISPR array generally contains less than 50 CRISPR repeat-spacer sequences. CRISPR is transcribed as part of a long transcript called the primary CRISPR transcript, which contains much of the CRISPR array. The primary CRISPR transcript is cleaved by Cas proteins to produce crRNAs, or in some cases pre-crRNAs. The pre-crRNAs are further processed by additional Cas proteins to become mature crRNAs. The mature crRNAs contain spacer sequences and CRISPR repeat sequences. In some embodiments where the pre-crRNA is processed to become a mature (or processed) crRNA, maturation includes the removal of about 1 to about 6 or more 5' nucleotides, or 3' nucleotides, or 5' and 3' nucleotides. For the purpose of genome editing or targeting a specific target nucleotide sequence of interest, the nucleotides removed during the maturation of the pre-crRNA molecule are not required for the generation or design of the guide RNA.
[0036] CRISPR RNA (crRNA) contains a spacer sequence and a CRISPR repeat sequence. The "spacer sequence" is a nucleotide sequence that directly hybridizes to a target nucleotide sequence of interest. The spacer sequence is engineered to be fully or partially complementary to the target nucleotide sequence of interest. In various embodiments, the spacer sequence can contain from about 8 to about 30 or more nucleotides. For example, the spacer sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the spacer sequence is from about 10 to about 26 nucleotides in length, or from about 12 to about 30 nucleotides in length. In a particular embodiment, the spacer sequence is about 30 nucleotides in length. In some embodiments, the degree of complementarity between the spacer sequence and its corresponding target sequence is about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or greater when optimally aligned using an appropriate alignment program. In a particular embodiment, the spacer sequence adopts a free secondary structure, which can be predicted using any suitable polynucleotide folding algorithm known in the art. Non-limiting examples of algorithms include mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).
[0037] RGN proteins may vary in their sensitivity to mismatches between the spacer sequence within the gRNA and the target sequence of the gRNA, and this sensitivity affects the efficiency of cleavage. As discussed in Example 5, APG05459.1 RGN has an unusual sensitivity to mismatches between the spacer sequence and the target sequence spanning at least 15 nucleotides 5' of the PAM site. Thus, APG05459.1 has the ability to more finely (i.e., specifically) target specific sequences compared to other RGNs that are less sensitive to mismatches between the spacer sequence and the target sequence.
[0038] The CRISPR RNA repeat sequence comprises a nucleotide sequence that includes a region having sufficient complementarity to hybridize to tracrRNA. In various embodiments, the CRISPR RNA repeat sequence can include from about 8 to about 30 or more nucleotides. For example, the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides in length. In some embodiments, the CRISPR repeat sequence is about 21 nucleotides in length. In some embodiments, the degree of complementarity between the CRISPR repeat sequence and the corresponding tracrRNA sequence is about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or greater when optimally aligned using an appropriate alignment program. In particular embodiments, the CRISPR repeat sequence comprises a nucleotide sequence of any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, or an active variant or fragment thereof, and when it is included in a guide RNA, it can specifically bind the relevant RNA-guided nuclease presented herein to the target sequence of interest in a sequence-specific manner. In some embodiments, an active CRISPR repeat sequence variant of the wild-type sequence comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity with the nucleotide sequence represented as any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.In some embodiments, the active CRISPR repeat fragment of the wild-type sequence comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleotides of the nucleotide sequence represented as any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.
[0039] In some embodiments, the crRNA does not occur naturally. In some of these embodiments, the specific CRISPR repeat is not naturally linked to the engineered spacer sequence and is considered heterologous to the spacer sequence. In some embodiments, the spacer sequence is an engineered sequence that does not occur naturally.
[0040] The trans-activating CRISPR RNA molecule, i.e., the tracrRNA molecule, contains a nucleotide sequence that includes a region having sufficient complementarity to hybridize to the CRISPR repeat sequence of the crRNA (referred to herein as the anti-repeat region). In some embodiments, the tracrRNA molecule further includes a region having a secondary structure (e.g., a stem-loop) or forms a secondary structure when hybridized to the corresponding crRNA. In a particular embodiment, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is at the 5' end of the molecule, and the 3' end of the tractRNA contains a secondary structure. This region of the secondary structure generally includes several hairpin structures (including the ligation hairpin found adjacent to the anti-repeat sequence). The ligation hairpin often has a conserved nucleotide sequence among the bases of the hairpin stem, and motifs UNANNG, UNANNU, UNANNA (SEQ ID NOs: 68, 557, 558 respectively) are found among many of the ligation hairpins in the tracrRNA. There are often multiple terminal hairpins at the 3' end of the tracrRNA. Although its structure and number can vary, it often includes a CG-rich transcription termination hairpin independent of Rho and a string of U's located at the subsequent 3' end. See, for example, Briner et al. (2014) Molecular Cell Vol. 56: pp. 333 - 339, Briner and Barrangou (2016) Cold Spring Harb Protoc; doi: 10.1101 / pdb.top090902, and U.S. Patent Application Publication No. 2017 / 0275648 (each of which is incorporated herein by reference in its entirety).
[0041] In various embodiments, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence contains from about 8 to about 30 or more nucleotides. For example, as the region forming base pairs between the tracrRNA anti-repeat region and the CRISPR repeat sequence, the length can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides. In a particular embodiment, the anti-repeat region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence is about 20 nucleotides in length. In some embodiments, the degree of complementarity between the CRISPR repeat sequence and the corresponding tracrRNA anti-repeat sequence is about 50% or more, about 60% or more, about 70% or more, about 75% or more, about 80% or more, about 81% or more, about 82% or more, about 83% or more, about 84% or more, about 85% or more, about 86% or more, about 87% or more, about 88% or more, about 89% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or greater.
[0042] In various embodiments, the entire tracrRNA can contain from about 60 nucleotides to more than about 140 nucleotides. For example, as the tracrRNA, the length can be about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, or more nucleotides. In a particular embodiment, the tracrRNA is from about 80 to about 90 nucleotides in length, and includes nucleotides of about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90 nucleotides. In some embodiments, the tracrRNA is about 85 nucleotides in length.
[0043] In certain embodiments, the tracrRNA comprises a nucleotide sequence of any of SEQ ID NO: 3, 13, 21, 29, 38, 47, 56, or an active variant or fragment thereof, and when it is included in the guide RNA, it can specifically bind the relevant RNA-guided nuclease presented herein to the target sequence of interest in a sequence-specific manner. In some embodiments, an active tracrRNA sequence variant of the wild-type sequence comprises a nucleotide sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to the nucleotide sequence represented as any of SEQ ID NO: 3, 13, 21, 29, 38, 47, 56. In some embodiments, an active tracrRNA sequence fragment of the wild-type sequence comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides of the nucleotide sequence represented as any of SEQ ID NO: 3, 13, 21, 29, 38, 47, 56.
[0044] Two polynucleotide sequences can be considered to be substantially complementary when the two sequences hybridize to each other under stringent conditions. Similarly, an RGN is considered to bind to a target sequence in a sequence-specific manner when the guide RNA bound to the RGN binds to the specific target sequence under stringent conditions. "Stringent conditions" or "stringent hybridization conditions" means conditions under which two polynucleotide sequences hybridize to each other with a detectability greater (e.g., at least two-fold greater than background) than other sequences. Stringent conditions are sequence-dependent and will be different in different circumstances. Typically, stringent conditions are those in which the salt concentration is less than about 1.5 M Na ions at pH 7.0 - 8.3, typically about 0.01 - 1.0 M Na ions (or other salts), and the temperature is at least about 30 °C for short sequences (e.g., 10 - 50 nucleotides) and at least about 60 °C for long sequences (e.g., more than 50 nucleotides). Stringent conditions can also be achieved by the addition of destabilizing agents such as formamide. Representative conditions of lower stringency include hybridization at 37 °C in a buffer solution of 30 - 35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate), and washing in 1× - 2× SSC (20× SSC = 3.0 M NaCl / 0.3 M trisodium citrate) at 50 - 55 °C. Representative conditions of moderate stringency include hybridization in 40 - 45% formamide, 1.0 M NaCl, and 1% SDS at 37 °C, and washing in 0.5× - 1× SSC at 55 - 60 °C. Representative conditions of high stringency include hybridization in 50% formamide, 1 M NaCl, and 1% SDS at 37 °C, and washing in 0.1× SSC at 60 - 65 °C. In some cases, the wash buffer can contain about 0.1% - about 1% SDS. The hybridization time is generally less than about 24 hours and is usually about 4 - about 12 hours. The wash time will be long enough to at least reach equilibrium.
[0045] Tm is the temperature at which 50% of the complementary target sequence hybridizes to a perfectly matched sequence (under defined ionic strength and pH). For DNA-DNA hybrids, Tm can be roughly estimated from the equation of Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284, Tm = 81.5°C + 16.6 (log M) + 0.41 (%GC) - 0.61 (%form) - 500 / L, where M is the molarity of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, %form is the percentage of formamide in the hybridization solution, and L is the length of the hybrid in base pairs. Generally, stringent conditions are selected to be about 5°C lower than the melting temperature (Tm) of the specific sequence and its complementary sequence at a defined ionic strength and pH. However, very stringent conditions can utilize hybridization and / or washing at a temperature 1°C, or 2°C, or 3°C, or 4°C lower than the melting temperature (Tm); moderately stringent conditions can utilize hybridization and / or washing at a temperature 6°C, or 7°C, or 8°C, or 9°C, or 10°C lower than the melting temperature (Tm); and low stringency conditions can utilize hybridization and / or washing at a temperature 11°C, or 12°C, or 13°C, or 14°C, or 15°C, or 20°C lower than the melting temperature (Tm). Those skilled in the art will understand that variations in the stringency of the hybridization and / or washing solutions are essentially described using the above equation, hybridization and washing of the compositions, and the desired Tm.Comprehensive guides to nucleic acid hybridization are found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology - Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York); Ausubel et al. (eds.) (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See also Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2nd ed.), Cold Spring Harbor Laboratory Press, Plainview, New York State.
[0046] As the guide RNA, a single guide RNA or a dual guide RNA system is possible. The single guide RNA contains a crRNA and a tracrRNA on a single RNA molecule, whereas the dual guide RNA system contains a crRNA and a tracrRNA on two different RNA molecules, hybridized to each other through at least a part of the CRISPR repeat sequence of the crRNA and at least a part of the tracrRNA (which can be completely or partially complementary to the CRISPR repeat sequence of the crRNA). In some embodiments where the guide RNA is a single guide RNA, the crRNA and the tracrRNA are separated by a linker nucleotide sequence. Generally, the linker nucleotide sequence is a sequence that does not contain complementary bases in order to avoid the formation of secondary structures within the nucleotides of the linker nucleotide sequence or to avoid the formation of secondary structures containing the nucleotides of the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between the crRNA and the tracrRNA is a sequence of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or more nucleotides in length. In a particular embodiment, the linker nucleotide sequence of the single guide RNA is a sequence of at least 4 nucleotides in length. In some embodiments, the linker nucleotide sequence is a nucleotide sequence represented by any of SEQ ID NOs: 63, 64, 65. In another embodiment, the linker nucleotide sequence is a sequence of at least 6 nucleotides in length. In some embodiments, the linker nucleotide sequence is a nucleotide sequence represented by SEQ ID NO: 65.
[0047] A single guide RNA, or a dual guide RNA, can be chemically synthesized or synthesized through in vitro transcription. Assays for revealing the sequence-specific binding between an RGN and a guide RNA are known in the art, and non-limiting examples thereof include binding assays between an expressed RGN and a guide RNA. The guide RNA can be tagged with a detectable label (e.g., biotin) and used in a pull-down detection assay, in which the guide RNA:RGN complex is captured through the detectable label (e.g., using streptavidin beads). A control guide RNA having a sequence or structure unrelated to the guide RNA can be used as a negative control for non-specific binding of the RGN to the RNA. In some embodiments, the guide RNA is any one of SEQ ID NO: 10, 18, 26, 35, 44, 53, 62, and the spacer sequence therein can be any sequence, shown as a poly N sequence.
[0048] As described in Example 8, some RGNs of the present invention can share some guide RNAs. APG05083.1, APG07433.1, APG07513.1, APG08290.1 can each function using a guide RNA containing a crRNA comprising a nucleotide sequence of any one of SEQ ID NO: 2, 12, 20, 28 and a corresponding tracrRNA comprising a nucleotide sequence of any one of SEQ ID NO: 3, 13, 21, 29, respectively. Further, APG04583.1 and APG01688.1 can each function using a guide RNA containing a crRNA comprising a nucleotide sequence of SEQ ID NO: 46 or 55 and a corresponding tracrRNA comprising a nucleotide sequence of SEQ ID NO: 47 or 56, respectively.
[0049] In some embodiments, the guide RNA can be introduced as an RNA molecule into a target cell, or an organelle (subcellular organ), or an embryo. The guide RNA can be transcribed in vitro or chemically synthesized. In another embodiment, the nucleotide sequence encoding the guide RNA is introduced into a target cell, or an organelle, or an embryo. In some of these embodiments, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., an RNA polymerase III promoter). The promoter can be a native promoter or one that is heterologous to the nucleotide sequence encoding the guide RNA.
[0050] In various embodiments, the guide RNA can be introduced into a target cell, or an organelle, or an embryo as a ribonucleoprotein complex as described herein, and the guide RNA binds to an RNA-guided nuclease polypeptide.
[0051] The guide RNA directs the associated RNA-guided nuclease to its target nucleotide sequence by hybridizing to a specific target nucleotide sequence of interest. The target nucleotide sequence can include DNA, or RNA, or a combination of both, and can be single-stranded or double-stranded. The target nucleotide sequence can be genomic DNA (i.e., chromosomal DNA), plasmid DNA, an RNA molecule (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). The RNA-guided nuclease can bind to the target nucleotide sequence in vitro or intracellularly. The chromosomal sequence targeted by the RGN can be a chromosomal sequence of the nucleus, plastid, or mitochondrion. In some embodiments, the target nucleotide sequence is only one in the target genome.
[0052] The target nucleotide sequence is adjacent to a protospacer motif (PAM). The protospacer motif is generally located within about 1 to about 10 nucleotides from the target nucleotide sequence, including about 1, or about 2, or about 3, or about 4, or about 5, or about 6, or about 7, or about 8, or about 9, or about 10 nucleotides from the target nucleotide sequence. The 5' or 3' of the target sequence can be the PAM. In some embodiments, the PAM is 3' of the target sequence for the RGNs of the present disclosure. Generally, the PAM is a consensus sequence consisting of about 3 to 4 nucleotides, but in special embodiments, it can be 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length. In various embodiments, the PAM sequences recognized by the RGNs of the present disclosure include consensus sequences represented as any of SEQ ID NO: 6, 32, 41, 50, 59. Non-limiting PAM sequences are nucleotide sequences represented as SEQ ID NO: 7, 69, 70, 71, 72.
[0053] In special embodiments, an RNA-guided nuclease having any of SEQ ID NO: 1, 11, 19, 27, 36, 45, 54, or an active variant or fragment thereof, binds to a target nucleotide sequence adjacent to a PAM sequence represented as any of SEQ ID NO: 6, 32, 41, 50, 59, 7, respectively. In some of these embodiments, the RGN includes a CRISPR repeat sequence represented as any of SEQ ID NO: 2, 12, 20, 28, 37, 46, 55, or an active variant or fragment thereof, and binds to a guide sequence that includes a tracrRNA sequence represented as any of SEQ ID NO: 3, 13, 21, 29, 38, 47, 56, or an active variant or fragment thereof. The RGN system is described in more detail in Example 1 and Table 1 of this specification.
[0054] It is well known that the PAM sequence specificity for a given nuclease enzyme is affected by the concentration of the enzyme (see, for example, Karvelis et al. (2015) Genome Biol. Vol. 16:253). This concentration can be altered by changing the promoter used to express the RGN, or by changing the amount of ribonucleoprotein complex delivered to any of cells, organelles, or embryos.
[0055] When the RGN recognizes the corresponding PAM sequence, it can cleave the target nucleotide sequence at a specific cleavage site. As used herein, the cleavage site is composed of two specific nucleotides within the target nucleotide sequence, and the nucleotide sequence therebetween is cleaved by the RGN. The cleavage site can include the first and second nucleotides, or the second and third nucleotides, or the third and fourth nucleotides, or the fourth and fifth nucleotides, or the fifth and sixth nucleotides, or the seventh and eighth nucleotides, or the eighth and ninth nucleotides, in the 5' or 3' direction from the PAM. In some embodiments, the cleavage site can be located at a position more than 10, or 11, or 12, or 13, or 14, or 15, or 16, or 17, or 18, or 19, or 20 nucleotides away in the 5' or 3' direction from the PAM. In some embodiments, the cleavage site is separated from the PAM by 4 nucleotides. In another embodiment, the cleavage site is at least 15 nucleotides away from the PAM. Since the RGN can cleave the target nucleotide sequence to produce sticky ends, in some embodiments, the cleavage site is defined based on a distance of 2 nucleotides from the PAM on the plus (+) strand of the polynucleotide and a distance of 2 nucleotides from the PAM on the minus (-) strand of the polynucleotide.
[0056] III. Nucleotides encoding an RNA-guided nuclease, and / or a CRISPR RNA, and / or a tracrRNA
[0057] Provided by the present disclosure are polynucleotides comprising a CRISPR RNA, and / or a tracrRNA, and / or an sgRNA of the present disclosure, and polynucleotides comprising a nucleotide sequence encoding an RNA-guided nuclease of the present disclosure, and / or a CRISPR RNA, and / or a tracrRNA, and / or an sgRNA. The polynucleotides of the present disclosure include polynucleotides comprising or encoding a CRISPR repeat sequence comprising any of the nucleotide sequences of SEQ ID NO: 2, 12, 20, 28, 37, 46, 55, or an active variant or fragment thereof, and when this CRISPR repeat sequence, or an active variant or fragment thereof, is included in a guide RNA, it can specifically bind a related RNA-guided nuclease to a target sequence in a sequence-specific manner. Also disclosed are polynucleotides comprising or encoding a tracrRNA comprising any of the nucleotide sequences of SEQ ID NO: 3, 13, 21, 29, 38, 47, 56, or an active variant or fragment thereof, and when this tracrRNA, or an active variant or fragment thereof, is included in a guide RNA, it can specifically bind a related RNA-guided nuclease to a target sequence in a sequence-specific manner. Also provided are polynucleotides encoding an RNA-guided nuclease comprising an amino acid sequence represented by any of SEQ ID NO: 1, 11, 19, 27, 36, 45, 54, and an active fragment or variant thereof, and this RNA-guided nuclease, and its active fragment or variant retain the ability to bind to a target nuclease sequence in an RNA-guided sequence-specific manner.
[0058] As used herein, the term "polynucleotide" is not limited to polynucleotides containing DNA in the present disclosure. Those skilled in the art will recognize that polynucleotides can include ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both natural molecules and synthetic analogs. Among them are peptide nucleic acids (PNA), PNA-DNA chimeras, locked nucleic acids (LNA), and phosphorothioate bridging sequences. The polynucleotides disclosed herein also encompass all forms of sequences, and non-limiting examples thereof include single-stranded form, double-stranded form, DNA-RNA hybrids, triple-stranded structures, stem-loop structures, and the like.
[0059] As a nucleic acid molecule encoding an RGN, it is possible to optimize the codons for expression in the body of the organism of interest. A "codon-optimized" coding sequence is a polynucleotide encoding a sequence with a codon usage frequency designed to mimic the preferred codon usage frequency or transcription conditions of a particular host cell. At the nucleic acid level, as a result of changing one or more codons so that the translated amino acid sequence does not change, expression in that particular host cell or organism is increased. Codons can be optimized for all or part of the nucleic acid molecule. Codon tables and other references providing preferred information for a wide range of organisms are available in the art (for information on the use of plant-preferred codons, see, for example, Campbell and Gowri (1990) Plant Physiol. Vol. 92: 1-11). In the art, methods for synthesizing plant-preferred genes can be utilized. See, for example, U.S. Patent Nos. 5,380,831, 5,436,391, Murray et al. (1989) Nucleic Acids Res. Vol. 17: 477-498 (incorporated herein by reference).
[0060] Provided is a polynucleotide encoding the RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA presented herein, which is placed in an expression cassette and expressed in vitro, or can be expressed in any of a target cell, organelle, embryo, or organism. The cassette contains a 5' regulatory sequence and a 3' regulatory sequence operably linked to the polynucleotide encoding the RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA presented herein, such that expression of this polynucleotide will be enabled. The cassette can further contain at least one additional gene or gene element for introducing into an organism and co-transforming. When additional genes or elements are included, these elements are operably linked. The expression "operably linked" means a functional linkage between two or more elements. For example, an operable linkage between a promoter and a coding region of interest (e.g., a region encoding RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA) is a functional linkage that enables expression of this coding region of interest. The operably linked elements may or may not be contiguous. When "operably linked" is used to refer to the joining of two protein-coding regions, it means that those coding regions are in the same reading frame. Alternatively, the additional genes or elements can be provided on multiple expression cassettes. For example, the nucleotide sequence encoding the RGN of the present disclosure can be on one expression cassette, while the nucleotide sequence encoding any of crRNA, trancrRNA, or a complete guide RNA can be on a separate expression cassette. Such expression cassettes are provided with multiple restriction sites and / or recombination sites for inserting a polynucleotide under transcriptional regulation of a regulatory region. The expression cassette can further contain a selectable marker gene.
[0061] The expression cassette will comprise, in the 5' to 3' direction of transcription, a transcription (and in some cases translation) initiation region (i.e., promoter) of the invention that functions in the organism of interest, a polynucleotide encoding an RGN, a polynucleotide encoding a crRNA, a polynucleotide encoding a tracrRNA, a polynucleotide encoding an sgRNA, and a transcription (and in some cases translation) termination region (i.e., termination region). The promoter of the invention can direct or instruct the expression of the coding sequence in the host cell. These regulatory regions (e.g., promoter, transcription regulatory region, translation termination region) can be endogenous to the host cell, or heterologous to the host, or a mixture of endogenous and heterologous. As used herein with respect to a sequence, "heterologous" is a sequence that is derived from a foreign species or, if from the same species, is substantially modified from its native form in composition and / or genomic locus by deliberate intervention. As used herein, a chimeric gene comprises, as a coding sequence, a coding sequence operably linked to a transcription initiation region heterologous to the coding sequence.
[0062] Convenient termination regions (octopine synthase termination region and nopaline synthase termination region) can be obtained from the Ti plasmid of Agrobacterium tumefaciens. See also Guerineau et al. (1991) Mol. Gen. Genet. 262:141-144; Proudfoot (1991) Cell 64:671-674; Sanfacon et al. (1991) Genes Dev. 5:141-149; Mogen et al. (1990) Plant Cell 2:1261-1272; Munroe et al. (1990) Gene 91:151-158; Ballas et al. (1989) Nucleic Acids Res. 17:7891-7903; Joshi et al. (1987) Nucleic Acids Res. 15:9627-9639.
[0063] Non-limiting examples of additional regulatory signals include transcription start sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See, for example, U.S. Pat. Nos. 5,039,523 and 4,853,331; EPO 0480762A2; Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, Maniatis et al. eds. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.) (hereinafter referred to as "Sambrook 11"); Davis et al. eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, N.Y., and references cited therein.
[0064] When preparing an expression cassette, various DNA fragments are manipulated so as to provide DNA sequences in the appropriate orientation and, if necessary, in the appropriate reading frame. For this purpose, joining DNA fragments using adapters or linkers, and other manipulations, including the provision of convenient restriction sites, removal of excess DNA, removal of restriction sites, etc., can be included. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, replacement (e.g., transitions and transversions) can be included.
[0065] A number of promoters can be used to practice the present invention. The promoters can be selected based on the desired results. For expressing in a target organism, the nucleic acid can be combined with a constitutive promoter, an inducible promoter, a growth stage-specific promoter, a cell type-specific promoter, a tissue-selective promoter, a tissue-specific promoter, or other promoters. See, for example, the promoters described in WO 99 / 43838 and U.S. Patent Nos. 8,575,425; 7,790,846; 8,147,856; 8,586,832; 7,772,369; 7,534,939; 6,072,050; 5,659,026; 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; 5,608,142; 6,177,611 (which are incorporated herein by reference).
[0066] For expressing in plants, constitutive promoters include the CaMV 35S promoter (Odell et al. (1985) Nature 313:810 - 812); rice actin (McElroy et al. (1990) Plant Cell 2:163 - 171); ubiquitin (Christensen et al. (1989) Plant Mol. Biol. 12:619 - 632 and Christensen et al. (1992) Plant Mol. Biol. 18:675 - 689); pEMU (Last et al. (1991) Theor. Appl. Genet. 81:581 - 588); MAS (Velten et al. (1984) EMBO J. 3:2723 - 2730).
[0067] Examples of inducible promoters include the Adh1 promoter, which can be induced by hypoxia stress or low temperature stress; the Hsp70 promoter, which can be induced by heat stress; the PPDK promoter and the pep carboxylase promoter, which can be induced by light. Chemically inducible promoters, such as the In2-2 promoter (U.S. Patent No. 5,364,780) induced by a phytotoxicity reducing agent, the Axig1 promoter (PCT US01 / 22169) induced by auxin and specific to tapetum tissue but active even in callus, the steroid-responsive promoter (see, for example, Schena et al. (1991) Proc. Natl. Acad. Sci. USA Vol. 88: pp. 10421-10425 and McNellis et al. (1998) Plant J. Vol. 14(2): pp. 247-257), and the promoter that can be induced and repressed by tetracycline (see, for example, Gatz et al. (1991) Mol. Gen. Genet. Vol. 227: pp. 229-237, U.S. Patent Nos. 5,814,618 and 5,789,156; these are incorporated herein by reference) are also useful.
[0068] Tissue-specific or tissue-selective promoters can be used to express an expression construct within a particular tissue. In some embodiments, the tissue-specific or tissue-selective promoter is active in plant tissue. Examples of promoters that are under developmental control in plants include promoters that selectively initiate transcription within a given tissue (such as leaves, roots, fruits, seeds, flowers, etc.). A "tissue-specific" promoter is a promoter that initiates transcription only within a given tissue. Tissue-specific expression is the result of several levels of genetic regulatory interactions, unlike the constitutive expression of a gene. Thus, promoters from homologous or closely related plant species can be selectively used to achieve efficient and reliable expression of a transgene within a particular tissue. In some embodiments, the expression includes a tissue-selective promoter. A "tissue-selective" promoter is a promoter that selectively initiates transcription but does not necessarily initiate transcription exclusively within or only within a given tissue.
[0069] In some embodiments, the nucleic acid molecule encoding RGN, and / or crRNA, and / or tracrRNA includes a cell-type specific promoter. A "cell-type specific" promoter is a promoter that predominantly directs expression in some types of cells within one or more organs. Some examples of plant cells in which cell-type specific promoters that function in plants can be predominantly active include, for example, BETL cells, vascular cells in roots and leaves, stem cells, and meristematic cells. The nucleic acid molecule can also include a cell-type selective promoter. A "cell-type selective" promoter is a promoter that typically directs expression in some types of cells within one or more organs but does not necessarily direct expression exclusively within or only within a given tissue. Some examples of plant cells in which cell-type selective promoters that function in plants can be selectively active include, for example, BELT cells, vascular cells in roots and leaves, stem cells, and meristematic cells.
[0070] Nucleic acid molecules encoding an RGN, and / or a crRNA, and / or a tracrRNA, and / or an sgRNA can be operably linked to a promoter sequence recognized by, for example, a phage RNA polymerase for in vitro mRNA synthesis. In such embodiments, the RNA transcribed in vitro can be purified and used in the methods described herein. For example, as the promoter sequence, a T7 promoter sequence, a T3 promoter sequence, an SP6 promoter sequence, or a variation of a T7 promoter sequence, a T3 promoter sequence, an SP6 promoter sequence is possible. In such embodiments, the expressed protein and / or RNA can be purified and used in the genome modification methods described herein.
[0071] In some embodiments, the polynucleotide encoding an RGN, and / or a crRNA, and / or a tracrRNA, and / or an sgRNA can also be linked to a polyadenylation signal (such as the SV40 polyA signal and other signals that function in plants) and / or at least one transcription termination sequence. In addition, as described elsewhere herein, the sequence encoding the RGN can also be linked to at least one nuclear localization signal capable of moving the protein to a specific intracellular location, and / or at least one cell entry domain, and / or a sequence encoding at least one signal peptide.
[0072] Polynucleotides encoding RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA can be present in one vector, or in multiple vectors. "Vector" means a polynucleotide composition for transferring, delivering, or introducing nucleic acid into a host cell. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, baculoviral vectors). Vectors can include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selection marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Additional information can be found in "Current Protocols in Molecular Biology", Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual", Sambrook and Russell, Cold Spring Harbor Press, Cold Spring Harbor, New York State, 3rd Edition, 2001.
[0073] Vectors can also include a selection marker gene for selecting transformed cells. The selection marker gene is used to select transformed cells or tissues. Marker genes include genes encoding antibiotic resistance (such as genes encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT)), as well as genes conferring resistance to insecticidal compounds (such as glufosinate ammonium, bromoxynil, imidazolinone, 2,4-dichlorophenoxyacetate (2,4-D), etc.).
[0074] In some embodiments, an expression cassette or vector comprising a sequence encoding an RGN polypeptide can further comprise a sequence encoding a crRNA and / or a tracrRNA, or a crRNA and a tracrRNA combined to generate a guide RNA. The sequence encoding the crRNA and / or the tracrRNA can be operably linked to at least one transcriptional control sequence for expressing the crRNA and / or the tracrRNA in the organism or host cell of interest. For example, the polynucleotide encoding the crRNA and / or the tracrRNA can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Non-limiting examples of suitable Pol III promoters include the mammalian U6 promoter, U3 promoter, H1 promoter, 7SL RNA promoter, and the rice U6 promoter, U3 promoter.
[0075] As shown, an organism of interest can be transformed using an expression construct encoding an RGN, and / or a crRNA, and / or a tracrRNA, and / or an sgRNA. The method of transformation includes introducing a nucleotide construct into the organism of interest. "Introducing" means introducing the nucleotide construct into a host cell and the construct gaining access to the interior of the host cell. The methods of the present invention do not require a particular method for introducing the nucleotide construct into the host organism, as long as the nucleotide construct can gain access to the interior of at least one cell of the host organism. The host cell can be a eukaryotic cell or a prokaryotic cell. In a particular embodiment, the eukaryotic cell is any of a plant cell, a mammalian cell, an insect cell. Methods for introducing nucleotide constructs into plants and other host cells are known in the art and non-limiting examples thereof include methods for stable transformation, methods for transient transformation, virus-mediated methods.
[0076] By these methods, transformed organisms (e.g., plants, which include whole plants, as well as plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, embryos, and their progeny) can be obtained. Plant cells can be either differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, sieve tube cells, pollen).
[0077] The term "transgenic organism" or "transformed organism" or "stably transformed" organism or cell or tissue means an organism into which a polynucleotide encoding the RGN, and / or crRNA, and / or tracrRNA, and / or sgRNA of the present invention has been incorporated or integrated. It is known that other exogenous or endogenous nucleic acid sequences or DNA fragments can also be incorporated into host cells. Agrobacterium - and biolistic - mediated transformation are still two major approaches used to transform plant cells. However, transformation of host cells can be achieved by methods mediated by infection, transfection, microinjection, electroporation, microprojection, biolistics or particle bombardment, silica / carbon fiber, ultrasonic - mediated methods, PEG - mediated methods, calcium phosphate coprecipitation, polycation DMSO technology, DEAE - dextran procedure, virus - mediated methods, liposome - mediated methods, etc. Methods for introducing a polynucleotide encoding RGN, and / or crRNA, and / or tracrRNA by virus - mediated means include introduction and expression mediated by retrovirus, lentivirus, adenovirus, adeno - associated virus, as well as the use of caulimovirus, geminivirus, RNA plant virus.
[0078] In addition to transformation protocols, the protocols for introducing polypeptide or polynucleotide sequences into plants can vary depending on the type of host cell targeted for transformation (e.g., cells of monocotyledonous plants or dicotyledonous plants). Methods of transformation are known in the art and include those described in U.S. Patent Nos. 8,575,425; 7,692,068; 8,802,934; 7,541,517 (each incorporated herein by reference). See also Rakoczy-Trojanowska, M. (2002) Cell Mol Biol Lett. Vol. 7:849-858; Jones et al. (2005) Plant Methods Vol. 1:5; Rivera et al. (2012) Physics of Life Reviews Vol. 9:308-345; Bartlett et al. (2008) Plant Methods Vol. 4:1-12; Bates, G.W. (1999) Methods in Molecular Biology Vol. 111:359-366; Binns and Thomashow (1988) Annual Reviews in Microbiology Vol. 42:575-606; Christou, P. (1992) The Plant Journal Vol. 2:275-281; Christou, P. (1995) Euphytica Vol. 85:13-27; Tzfira et al. (2004) TRENDS in Genetics Vol. 20:375-383; Yao et al. (2006) Journal of Experimental Botany Vol. 57:3737-3746; Zupan and Zambryski (1995) Plant Physiology Vol. 107:1041-1047.
[0079] By transformation, nucleic acids can be stably or transiently incorporated into cells. "Stable transformation" means that a nucleotide construct introduced into a host cell is integrated into the genome of that host cell and can be inherited by its progeny. "Transient transformation" means that a polynucleotide is introduced into a host cell but not integrated into the genome of that host cell.
[0080] Methods for transforming chloroplasts are known in the art. See, for example, Svab et al. (1990) Proc. Natl. Acad. Sci. USA Vol. 87: 8526-8530; Svab and Maliga (1993) Proc. Natl. Acad. Sci. USA Vol. 90: 913-917; Svab and Maliga (1993) EMBO J. Vol. 12: 601-606. This method relies on delivering DNA containing a selectable marker by particle gun and targeting that DNA to the plastid genome through homologous recombination. In addition, plastid transformation can be achieved by cross-activating silent transgenes carried by plastids by tissue-specific expression of a plastid-derived RNA polymerase encoded by the nucleus. Such a system is reported in McBride et al. (1994) Proc. Natl. Acad. Sci. USA Vol. 91: 7301-7305.
[0081] The transformed cells can be introduced into transgenic organisms (such as plants) according to conventional methods and propagated. See, for example, McCormick et al. (1986) Plant Cell Reports Vol. 5: pages 81 - 84. The plant is then grown and pollinated using the same transformed strain or a different strain, and those hybrid plants in which the desired phenotypic characteristics are constitutively expressed are identified. After growing for two or more generations to confirm that the expression of the desired phenotypic characteristics is stably maintained and inherited, seeds are collected to confirm that the desired phenotypic characteristics have been achieved. In this way, the present invention provides transformed seeds (also referred to as "transgenic seeds") in which the nucleotide construct of the present invention (for example, the expression cassette of the present invention) is stably integrated into the genome.
[0082] Alternatively, the transformed cells can be introduced into an organism. These cells may be considered to be derived from the organism, in which case the cells are transformed by an in vitro approach.
[0083] The sequences presented herein can be used for the transformation of any plant species, including, by way of non - limiting example, monocotyledonous and dicotyledonous plants. Non - limiting examples of target plants include maize, sorghum, wheat, sunflower, tomato, cruciferous plants, pepper, potato, cotton, rice, soybean, sugar beet, sugarcane, tobacco, barley, rapeseed, species of the genus Brassica, alfalfa, rye, minor cereals, safflower, peanut, sweet potato, cassava, coffee, coconut, pineapple, citrus, cocoa, tea, banana, avocado, fig, guava, mango, olive, papaya, cashew nuts, macadamia nuts, almond, oat, vegetables, ornamental plants, conifers.
[0084] Non-limiting examples of vegetables include tomatoes, lettuce, green peas, lima beans, kidney beans, and members of the cucumber genus such as cucumbers, cantaloupes, and honeydews. Non-limiting examples of ornamental plants include azaleas, hydrangeas, hibiscus, roses, tulips, rhododendrons, petunias, carnations, poinsettias, and chrysanthemums. The plants of the present invention are preferably cultivated crops (e.g., corn, sorghum, wheat, sunflowers, tomatoes, brassicas, peppers, potatoes, cotton, rice, soybeans, sugar beets, sugarcane, tobacco, barley, rapeseed, etc.).
[0085] As used herein, the term "plant" includes plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, plant populations, intact plant cells within plants, and parts of plants (e.g., embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, kernels, ears, ear axes, husks, stems, roots, root caps, anthers, etc.). A grain means a mature seed produced by a crop grower for purposes other than propagation or reproduction. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the present invention, provided that they contain the introduced polynucleotide. Further provided are processed plant products or by-products (e.g., soybean meal) that retain the sequences disclosed herein.
[0086] The polynucleotides encoding RGN, and / or crRNA, and / or tracrRNA can also be used to transform any prokaryotic species. Non-limiting examples of prokaryotic species include archaebacteria and bacteria (e.g., species of the genus Bacillus, species of the genus Klebsiella, species of the genus Streptomyces, species of the genus Rhizobium, species of the genus Escherichia, species of the genus Pseudomonas, species of the genus Salmonella, species of the genus Shigella, species of the genus Vibrio, species of the genus Erwinia, species of the genus Mycoplasma, Agrobacterium, species of the genus Lactobacillus).
[0087] Polynucleotides encoding RGN, and / or crRNA, and / or tracrRNA can also be used to transform any eukaryotic species. Non-limiting examples of eukaryotic species include animals (e.g., mammals, insects, fish, birds, reptiles), fungi, amoeba, algae, and yeast.
[0088] Nucleic acids can be introduced into mammalian cells or target tissues using conventional gene delivery methods based on viruses and non-viruses. Nucleic acids encoding components of the CRISPR system can be administered to cells in culture or host organisms using such methods. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles (such as liposomes). Viral vector delivery systems include DNA viruses and RNA viruses, which have an episomal genome or an integrated genome after being delivered to cells. For reviews of gene therapy, see Anderson, Science 256:808-813 (1992); Nabel and Feigner, TIBTECH 11:211-217 (1993); Mitani and Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer and Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al. in Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds.) (1995); Yu et al., Gene Therapy 1:13-26 (1994).
[0089] Non-viral delivery methods of nucleic acids include lipofection, nucleofection, microinjection, biolistic, virosomes, liposomes, immunoliposomes, polycations or lipid:nucleic acid complexes, naked DNA, artificial virions, and enhanced drug uptake of DNA. Lipofection is described, for example, in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for efficient receptor recognition lipofection of polynucleotides include the lipids of WO 91 / 17424 and WO 91 / 16024 by Felgner. Delivery can be done to cells (e.g., in vitro or ex vivo administration) or target tissues (e.g., in vivo administration). The preparation of lipid:nucleic acid complexes, including the preparation of targeted liposomes (such as immunolipid complexes), is well known to those skilled in the art (see, for example, Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res. 52:4817-4820 (1992); U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, 4,946,787).
[0090] Using RNA virus- or DNA virus-based systems for nucleic acid delivery takes advantage of highly evolved methods for directing the virus to specific cells in the body and transporting the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo). Alternatively, viral vectors can be used to treat cells in vitro, and in some cases, the modified cells can be administered to patients (ex vivo). Conventional virus-based systems have included retroviral vectors, lentiviral vectors, adenoviral vectors, adeno-associated viral vectors, and herpes simplex virus vectors for gene transport. Integration into the host genome is made possible by using gene transfer methods with retroviruses, lentiviruses, and adeno-associated viruses, and as a result, the inserted transgenes often express over the long term. In addition, high gene transfer efficiency has been observed in many different types of cells and target tissues.
[0091] The tropism of retroviruses can be altered by incorporating foreign envelope proteins and expanding the potential target population of the target cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and typically generate high viral titers. Therefore, the choice of retroviral gene delivery system is thought to depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats and have the ability to package foreign sequences up to 6-10 kb. Minimal cis-acting LTRs are sufficient for vector replication and packaging, and this vector is used to integrate a therapeutic gene into target cells and permanently express the transgene. Commonly used retroviral vectors include vectors based on murine leukemia virus (MuLV), gibbon leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, for example, Buchscher et al., J. Viral. Vol. 66:2731-2739 (1992); Johann et al., J. Viral. Vol. 66:1635-1640 (1992); Sommnerfelt et al., J. Viral. Vol. 176:58-59 (1990); Wilson et al., J. Viral. Vol. 63:2374-2378 (1989); Miller et al., J. Viral. Vol. 65:2220-2224 (1991); PCT / US94 / 05700).
[0092] For applications where transient expression is preferred, an adenovirus-based system can be used. Adenovirus-based vectors enable very high gene transfer efficiency in many types of cells and do not require cell division. High titers and high expression levels have been obtained using such vectors. This vector can be produced in large quantities in a relatively simple system. For example, to produce nucleic acids and peptides in vitro, or for in vivo and ex vivo gene therapy, target nucleic acids can also be introduced into cells using adeno-associated virus (「AAV」) vectors (see, e.g., West et al., Virology Vol. 160: 38-47 (1987); U.S. Patent No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy Vol. 5: 793-801 (1994); Muzyczka, J. Clin. Invest. Vol. 94: 1351 (1994)). The construction of recombinant AAV vectors has been described in numerous publications, including U.S. Patent No. 5,173,414; Tratschin et al., Mol. Cell. Biol. Vol. 5: 3251-3260 (1985); Tratschin et al., Mol. Cell. Biol. Vol. 4: 2072-2081 (1984); Hermonat and Muzyczka, PNAS Vol. 81: 6466-6470 (1984); Samulski et al., J. Viral. Vol. 63: 3822-3828 (1989). Typically, packaging cells are used to form virus particles capable of infecting host cells. Such cells include 293 cells (which package adenovirus) and ΨJ2 cells or PA317 cells (which package retrovirus).
[0093] Viral vectors used in gene therapy are typically produced by generating a cell line that packages the nucleic acid vector into viral particles. The vector typically contains the minimal viral sequences necessary for packaging and subsequent integration into the host, and other viral sequences are replaced by an expression cassette for the polynucleotide to be expressed. The lost viral functions are typically supplied in trans by the packaging cell line. For example, AAV vectors used in gene therapy typically have only the ITR sequences from the AAV genome necessary for packaging and integration into the host genome. The viral DNA is packaged into a cell line that contains a helper plasmid that encodes the other AAV genes (i.e., rep and cap) but lacks the ITR sequences.
[0094] The cell line can also be infected with adenovirus as a helper. The helper virus promoter facilitates the replication of the AAV vector and the expression of the AAV genes from the helper plasmid. The helper plasmid is not packaged in large amounts because it lacks the ITR sequences. Contamination with adenovirus can be reduced, for example, by heat treatment to which adenovirus is more sensitive than AAV. Additional methods for delivering nucleic acids to cells are known to those of skill in the art. See, for example, U.S. Patent Application Publication No. 2003 / 0087817 (incorporated herein by reference).
[0095] In some embodiments, one or more of the vectors described herein are transiently or non-transiently transfected into a host cell. In some embodiments, the cells are transfected in the same manner as occurs naturally in a subject. In some embodiments, the transfected cells are harvested from a subject. In some embodiments, the cells are derived from cells (such as cell lines) harvested from a subject. A variety of cell lines for tissue culture are known in the art. Non-limiting examples of cell lines include C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFl, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182, A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, TIB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4, COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelial, BALB / 3T3 mouse embryo fibroblast, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblast; 10.1 mouse fibroblast, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cell, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -、COR-L23、COR-L23 / CPR、COR-L235010、CORL23 / R23、COS-7、COV-434、CML Tl、CMT、CT26、D17、DH82、DU145、DuCaP、EL4、EM2、EM3、EMT6 / AR1、EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC 6, MTD-lA, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-lA / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THPl cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and transgenic variants thereof. Cell lines from a variety of sources known to those skilled in the art can be utilized (see, for example, American Type Culture Collection (ATCC) (Manassas, Virginia)).
[0096] In some embodiments, cells transfected with one or more of the vectors described herein are used to establish a new cell line containing sequences derived from the one or more vectors. In some embodiments, the components of the CRISPR system described herein are transiently transfected (by transient transfection of one or more vectors or transfection of RNA) to use cells modified through the activity of the CRISPR complex to establish a new cell line containing the modification but completely lacking other foreign sequences. In some embodiments, one or more test compounds are evaluated using cells transiently or non-transiently transfected with one or more of the vectors described herein, or cell lines derived from such cells.
[0097] In some embodiments, one or more of the vectors described herein are used to generate non-human transgenic animals or transgenic plants. In some embodiments, the transgenic animal is a mammal (such as a mouse, rat, rabbit, etc.).
[0098] IV. Variants and Fragments of Polypeptides and Polynucleotides
[0099] The present disclosure provides active variants and fragments of natural (i.e., wild-type) RNA-guided nucleases (the amino acid sequences of which are represented as any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54), as well as active variants and fragments of natural CRISPR repeats (such as sequences represented as any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55), and active variants and fragments of natural tracrRNAs (such as sequences represented as any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56), and polynucleotides encoding them.
[0100] The activity of the variant or fragment may vary compared to the polynucleotide or polypeptide of interest, but the variant and fragment must retain the functionality of the polynucleotide or polypeptide of interest. For example, the variant or fragment may have increased activity, decreased activity, a different spectrum of activity, or some other change in activity compared to the polynucleotide or polypeptide of interest.
[0101] Fragments and variants of natural RGN polypeptides (such as those disclosed herein) will retain sequence-specific RNA-guided DNA binding activity. In particular embodiments, fragments and variants of natural RGN polypeptides (such as those disclosed herein) will retain nuclease activity (single-stranded or double-stranded).
[0102] Fragments and variants of native CRISPR repeats, such as those disclosed herein, when part of a guide RNA (including tracrRNA), will retain the ability to bind to an RNA-guided nuclease in a sequence-specific manner (to form a complex with the guide RNA) and guide it to target a target nucleotide sequence.
[0103] Fragments and variants of native tracrRNA, such as those disclosed herein, when part of a guide RNA (including CRISPR RNA), will retain the ability to bind to an RNA-guided nuclease in a sequence-specific manner (to form a complex with the guide RNA) and guide it to target a target nucleotide sequence.
[0104] The term "fragment" means a portion of a polynucleotide or polypeptide sequence of the present invention. "Fragments" or "bioactive portions" include polynucleotides containing a sufficient number of contiguous nucleotides to retain biological activity (i.e., when contained within a guide RNA, bind sequence-specifically to an RGN and direct this RGN to a target nucleotide sequence). "Fragments" or "bioactive portions" include polypeptides containing a sufficient number of contiguous amino acid residues to retain biological activity (i.e., when contained within a guide RNA, bind sequence-specifically to an RGN and direct this RGN to a target nucleotide sequence). Fragments of the RGN protein include those that are shorter than the full-length sequence due to the use of alternative downstream start sites. As bioactive portions of the RGN protein, for example, polypeptides containing 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more amino acid residues of any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54 are possible. Such bioactive portions can be prepared by recombinant techniques and then evaluated for sequence-specific RNA-guided DNA binding activity. Bioactive fragments of the CRISPR repeat sequences can contain at least 8 contiguous amino acids of any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55. As bioactive portions of the CRISPR repeat sequences, for example, polynucleotides containing 8, or 9, or 10, or 11, or 12, or 13, or 14, or 15, or 16, or 17, or 18, or 19, or 20 contiguous nucleotides of any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55 are possible.As the biologically active portion of the tracrRNA, for example, polynucleotides containing 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more nucleotides in a row of any of SEQ ID NO: 3, 13, 21, 29, 38, 47, 56 are possible.
[0105] Generally, "variant" means a substantially similar sequence. For polynucleotides, a variant includes one or more deletions and / or additions of one or more nucleotides at one or more internal sites within a native polynucleotide and / or one or more substitutions of one or more nucleotides at one or more internal sites within a native polynucleotide. As used herein, "native" or "wild-type" polynucleotides or polypeptides contain a native nucleotide sequence or amino acid sequence, respectively. For polynucleotides, conserved variants include sequences that encode the native amino acid sequence of the gene of interest due to the degeneracy of the genetic code. Such natural allelic variants can be identified using well-known techniques of molecular biology such as polymerase chain reaction (PCR) and hybridization techniques outlined below. Variant polynucleotides also include polynucleotides that are synthetic but still encode the polypeptide or polynucleotide of interest (e.g., polynucleotides generated by site-directed mutagenesis). Generally, variants of the particular polynucleotides disclosed herein will have a sequence that is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to that of the particular polynucleotide as determined by the sequence alignment programs and parameters described elsewhere herein.
[0106] Variants of the special polynucleotides (i.e., reference polynucleotides) disclosed herein can also be evaluated by comparing the sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The sequence identity between any two polypeptides can be calculated using the sequence alignment programs and parameters described elsewhere herein. When any pair of polynucleotides disclosed herein is evaluated by comparing the sequence identity common to the two polypeptides encoded by these polynucleotides, the sequence identity between the two encoded polypeptides is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more.
[0107] In a particular embodiment, the polynucleotides of the disclosure encode an RNA-guided nuclease polypeptide comprising an amino acid sequence that is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to the amino acid sequence of any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54.
[0108] Biologically active variants of the RGN polypeptide of the present invention may differ by about 1 to 15 amino acid residues, or about 1 to 10 amino acid residues, or about 6 to 10 amino acid residues, or 5, or 4, or 3, or 2, or 1 few amino acid residues. In particular embodiments, the polypeptide can include N-terminal or C-terminal truncations, which can include deletions of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, or more amino acid residues from the N-terminus or C-terminus of the polypeptide.
[0109] In some embodiments, the polynucleotide of the present disclosure comprises or encodes a CRISPR repeat comprising a nucleotide sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to the nucleotide sequence represented by any of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.
[0110] The polynucleotide of the present disclosure can comprise or encode a tracrRNA comprising a nucleotide sequence that is at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to the nucleotide sequence represented by any of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56.
[0111] Biologically active variants of the CRISPR repeats or tracrRNAs of the present invention may differ by about 1 to 15 amino acid residues, or about 1 to 10 amino acid residues, or about 6 to 10 amino acid residues, or 5, or 4, or 3, or 2, or 1 few amino acid residues. In particular embodiments, the polynucleotide can include 5' or 3' truncations, which can include deletions of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more residues from the 5' or 3' end of the polynucleotide.
[0112] It can be seen that the RGN polypeptides, CRISPR repeats, and tracrRNAs presented herein can be modified to produce variant proteins and variant polynucleotides. Changes can be designed and introduced by applying site-directed mutagenesis techniques. Alternatively, it is also possible to identify polynucleotides and / or polypeptides that are natural but still unknown or not yet identified and are related to the sequences, structures, and / or functions disclosed herein and that fall within the scope of the present invention. Conservative amino acid substitutions that do not change the function of the RGN protein can be made to non-conserved regions. Alternatively, modifications can be made to improve the activity of the RGN.
[0113] Variant polynucleotides and variant proteins also include sequences and proteins obtained from procedures that cause mutations and recombination (such as DNA shuffling). Using such procedures, one or more different RGN proteins disclosed herein (e.g., SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54) are manipulated to create new RGN proteins with desired properties. In this way, a library of recombinant polynucleotides is generated from a population of polynucleotides having related sequences, i.e., a population of polynucleotides having substantially identical sequences and containing sequence regions capable of homologous recombination in vitro or in vivo. For example, using this approach, sequence motifs encoding the domain of interest are shuffled between the RGN sequences presented herein and other known RGN genes, and the desired properties are improved (e.g., for an enzyme, K mIt is possible to obtain a new gene encoding a protein with increased (). Such strategies for DNA shuffling are known in the art. See, for example, Stemmer (1994) Proc. Natl. Acad. Sci. USA Vol. 91: pp. 10747-10751; Stemmer (1994) Nature Vol. 370: pp. 389-391; Crameri et al. (1997) Nature Biotech. Vol. 15: pp. 436-438; Moore et al. (1997) J. Mol. Biol. Vol. 272: pp. 336-347; Zhang et al. (1997) Proc. Natl. Acad. Sci. USA Vol. 94: pp. 4504-4509; Crameri et al. (1998) Nature Vol. 391: pp. 288-291; U.S. Patent Nos. 5,605,793 and 5,837,458. A "shuffled" nucleic acid is a nucleic acid produced by a shuffling procedure (e.g., any shuffling procedure described herein). Shuffled nucleic acids are produced by recombinantly (physically or virtually) combining two or more nucleic acids (or strings), for example, in an artificial and sometimes recursive manner. Generally, the shuffling process utilizes one or more screening steps to identify the nucleic acid of interest. This screening step can be performed before or after any recombination step. In some (but not all) embodiments of shuffling, it is desirable to perform recombination multiple times before selection to increase the diversity of the pool to be screened. The entire process of recombination and selection may optionally be repeated recursively. Depending on the context, shuffling can mean the entire process of recombination and selection or just the recombination part of the entire process.
[0114] As used herein, "sequence identity" or "identity" in the context of two polynucleotide or polypeptide sequences means the residues that are the same between the two sequences when aligned to maximize correspondence over a specified comparison window. When the percent sequence identity is used with respect to a protein, the positions of residues that are not the same are often differences in conserved amino acid substitutions, where an amino acid residue is substituted for another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity) at that position, such that the functional properties of the molecule are recognized as not changing. When sequences differ with respect to conserved substitutions, the percent sequence identity can be adjusted upward to correct for the fact that such substitutions are of a conserved nature. Such sequences that differ by conserved substitutions are said to have "sequence similarity" or "similarity". Means for making this adjustment are well known to those of skill in the art. Typically, such means include scoring a conserved substitution as a partial mismatch rather than a complete mismatch, thereby increasing the percent sequence identity. Thus, for example, if one point is given for identical amino acids and zero points for non-conserved substitutions, a conserved substitution is given a score between zero and one. The score for a conserved substitution is calculated, for example, as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).
[0115] As used herein, "percent sequence identity" means a value determined by comparing two sequences optimally aligned on a comparison window. In such a case, a portion of a polynucleotide sequence within the comparison window may include additions or deletions (i.e., gaps) as compared to a reference sequence for optimal alignment of the two sequences (not including additions or deletions). The percent is calculated by determining the number of positions at which the same nucleic acid base or amino acid residue is present in both sequences to obtain the number of matched positions, dividing this number of matched positions by the total number of positions within the comparison window, and multiplying by 100 to obtain the percent sequence identity.
[0116] Unless otherwise specified, the sequence identity / similarity values presented in this specification are the values obtained using GAP Version 10 (using the following parameters: for nucleotide sequences, GAP Weight of 50, Length Weight of 3, and % identity and % similarity obtained using the nwsgapdna.cmp scoring matrix; for amino acid sequences, GAP Weight of 8, Length Weight of 2, and % identity and % similarity obtained using the BLOSUM62 scoring matrix), or the values obtained using any program equivalent thereto. "Equivalent program" means any sequence comparison program that, when compared to the corresponding alignment generated by GAP Version 10 for any two sequences in question, results in an alignment with the same nucleotide or amino acid residue matches and the same percentage of sequence identity.
[0117] Two arrays are "optimally aligned" when, for the purpose of obtaining a similarity score, they are aligned to reach the highest possible score for the pair of arrays using a given amino acid substitution matrix (e.g., BLOSUM62), a gap existence penalty, and a gap extension penalty. The amino acid substitution matrix and the use thereof to quantify the similarity between two arrays are well known in the art and are described, for example, in 'Atlas of Protein Sequence and Structure', Volume 5, Supplement 3 (edited by M. O. Dayhoff), Natl. Biomed. Res. Found., Washington D.C., by Dayhoff et al. (1978) "Models of Evolutionary Change in Proteins", pages 345 - 352, and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA Vol. 89: 10915 - 10919. The BLOSUM62 matrix is often used as the default scoring substitution matrix in sequence alignment protocols. The gap existence penalty is imposed to introduce a single amino acid gap into one of the aligned sequences, and the gap extension penalty is imposed to introduce additional empty amino acid positions into an already existing gap. The alignment is defined within each sequence by the amino acid positions at which the alignment starts and ends, and in some cases, one or more gaps are inserted into one or both sequences to obtain the highest possible score. The optimal alignment and scoring can be achieved manually, but this process is facilitated by using a computer-implemented alignment algorithm (e.g., the gapped BLAST 2.0 described in Altschul et al. (1997) Nucleic Acids Res. Vol. 25: 3389 - 3402, which is available on the National Center for Biotechnology Information Website (www.ncbi.nlm.nih.gov)).Optimal alignments can be prepared using, for example, PSI-BLAST (described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402, available from www.ncbi.nlm.nih.gov), including a number of alignments.
[0118] With respect to a reference sequence and an amino acid sequence that has been placed in a state of optimal alignment, one amino acid residue "corresponds" to the position in the reference sequence that pairs with this residue in the alignment. "Position" is represented by the number of each amino acid sequentially identified based on its relative position to the N-terminus in the reference sequence. Due to deletions, insertions, truncations, fusions, etc. that should be considered when obtaining the optimal alignment, generally, the number of the amino acid residue in the test sequence when obtained simply by counting from the N-terminus does not necessarily match the number of the corresponding position in the reference sequence. For example, if there is one deletion in the aligned test sequence, there is no amino acid in the reference sequence corresponding to the position of the deletion site. If there is one insertion in the aligned test sequence, the insertion does not correspond to any amino acid position in the reference sequence. In the case of truncation or fusion, there may be a series of amino acids in the reference sequence or the aligned sequence that do not correspond to any amino acid in the corresponding sequence.
[0119] V. Antibody
[0120] Antibodies against the RGN polypeptide of the present invention, or ribonucleoproteins containing this RGN polypeptide (including those having the amino acid sequences represented as SEQ ID NO: 1, 11, 19, 27, 36, 45, 54, or their active variants or fragments) are also included. Methods for producing antibodies are well known in the art (see, for example, Harlow and Lane (1988) "Antibodies: A Laboratory Manual", Cold Spring Harbor Laboratory, Cold Spring Harbor, New York; see U.S. Patent No. 4,196,265). These antibodies can be used in kits for detecting and isolating RGN polypeptides or ribonucleoproteins. Accordingly, the present disclosure provides a kit comprising an antibody that specifically binds to a polypeptide or ribonucleoprotein described herein (including a polypeptide having any of the sequences of SEQ ID NO: 1, 11, 19, 27, 36, 45, 54).
[0121] VI. Systems and ribonucleoprotein complexes for binding to a target sequence of interest, and methods for making the same
[0122] The present disclosure provides a system for binding to a target sequence of interest. This system includes at least one guide RNA, or a nucleotide sequence encoding the same, and at least one RNA-guided nuclease, or a nucleotide sequence encoding the same. The guide RNA hybridizes to the target sequence of interest and also forms a complex with the RGN polypeptide, thereby binding the RGN polypeptide to the target sequence. In some of these embodiments, the RGN includes an amino acid sequence of any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or an active variant or fragment thereof. In various embodiments, the guide RNA includes a CRISPR repeat sequence that includes a nucleotide sequence of any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, or an active variant or fragment thereof. In a particular embodiment, the guide RNA includes a tracrRNA that includes a nucleotide sequence of any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56, or an active variant or fragment thereof. As the guide RNA of this system, a single guide RNA or a dual guide RNA is possible. In a particular embodiment, the system includes an RNA-guided nuclease that is heterologous to the guide RNA, and the RGN and the guide RNA do not naturally form a complex.
[0123] As a system presented herein for binding to a target sequence of interest, a ribonucleoprotein complex is possible, which complex is at least one RNA molecule bound to at least one protein. The ribonucleoprotein complexes presented herein include at least one guide RNA as an RNA component and an RNA-guided nuclease as a protein component. Such ribonucleoprotein complexes can be purified from cells or organisms engineered to naturally express the RGN polypeptide and to express a particular guide RNA specific for the target sequence of interest. Alternatively, the ribonucleoprotein complex can be purified from cells or organisms transformed with polynucleotides encoding the RGN polypeptide and the guide RNA and cultured under conditions that permit expression of the RGN polypeptide and the guide RNA. In this way, a method for producing an RGN polypeptide or an RGN ribonucleoprotein complex is provided. Such a method includes culturing a cell containing a nucleotide sequence encoding the RGN polypeptide under conditions in which the RGN polypeptide (and in some embodiments the guide RNA) is expressed. Thereafter, the RGN polypeptide or the RGN ribonucleoprotein complex can be purified from the lysate of the cultured cells.
[0124] Methods for purifying RGN polypeptides or ribonucleoprotein complexes from the lysate of a biological sample are known in the art (e.g., size exclusion chromatography and / or affinity chromatography, 2D-PAGE, HPLC, reverse phase chromatography, immunoprecipitation). In a particular method, the RGN polypeptide is recombinantly produced and contains a purification tag to assist in purification. Non-limiting examples of purification tags include glutathione-S-transferase (GST), chitin binding protein (CBP), maltose binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6×His, 10×His, biotin carboxyl carrier protein (BCCP), calmodulin. Generally, the tagged RGN polypeptide or RGN ribonucleoprotein complex is purified using immobilized metal affinity chromatography. It will be appreciated that other similar methods known in the art (including other forms of chromatography and, for example, immunoprecipitation) can be used alone or in combination.
[0125] An "isolated" or "purified" polypeptide, or a biologically active portion thereof, is substantially or essentially free of components that are normally associated with the polypeptide as it is found in its natural environment, or that interact with the polypeptide. Thus, an isolated or purified polypeptide is substantially free of other cellular material when produced by recombinant techniques, or of chemical precursors or other chemicals when chemically synthesized. A protein that is substantially free of cellular material includes a protein preparation having less than about 30%, or less than 20%, 10%, 5%, or 1% (by dry weight) of contaminating protein. The proteins of the invention, or biologically active portions thereof, are produced by recombinant means, and optimally, the medium contains less than about 30%, or less than 20%, 10%, 5%, or 1% of chemical precursors or other chemical substances without the protein of interest.
[0126] The special methods presented herein for binding to and / or cleaving a target sequence of interest involve the use of an in vitro assembled RGN ribonucleoprotein complex. In vitro assembly of the RGN ribonucleoprotein complex can be carried out using any method known in the art, in which the RGN polypeptide is contacted with the guide RNA under conditions such that the RGN polypeptide can bind to the guide RNA. As used herein, "contacting", "being in contact", "contacted" mean bringing together the components of the desired reaction and placing them under conditions suitable for carrying out the desired reaction. The RGN polypeptide can be purified from any of a biological sample, cell lysate, or medium, which are produced by in vitro translation or chemical synthesis. The guide RNA can be purified from any of a biological sample, cell lysate, or medium, which are produced by in vitro transcription or chemical synthesis. The RGN ribonucleoprotein complex can be assembled in vitro by contacting the RGN polypeptide and the guide RNA in a solution (e.g., buffered saline).
[0127] VII. Methods for binding to a target sequence, methods for cleaving a target sequence, methods for modifying a target sequence
[0128] The present disclosure provides methods for binding to a target nucleotide of interest, and / or methods for cleaving a target nucleotide of interest, and / or methods for modifying a target nucleotide of interest. These methods involve delivering a system comprising at least one guide RNA or a polynucleotide encoding the same and at least one RGN polypeptide or a polynucleotide encoding the same to a target sequence or to a cell, organelle, or embryo containing the target sequence. In some of these embodiments, the RGN comprises the amino acid sequence of any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or an active variant or fragment thereof. In various embodiments, the guide RNA comprises a CRISPR repeat sequence having the nucleotide sequence of any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, or an active variant or fragment thereof. In a particular embodiment, the guide RNA comprises a tracrRNA having the nucleotide sequence of any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56, or an active variant or fragment thereof. A single guide RNA or a dual guide RNA is possible as the guide RNA of the system. The RGN of the system can be a nuclease-dead RGN, or have nickase activity, or be a fusion polypeptide. In some embodiments, the fusion polypeptide comprises a base editing polypeptide (e.g., cytidine deaminase, or adenosine deaminase). In a particular embodiment, the RGN and / or guide RNA is heterologous to the cell, or organelle, or embryo into which the RGN and / or guide RNA (or a polynucleotide encoding at least one of the RGN and guide RNA) is introduced.
[0129] In embodiments where the method involves delivery of a polynucleotide encoding a guide RNA and / or an RGN polypeptide, the cell or embryo can then be cultured under conditions in which the guide RNA and / or the RGN polypeptide are expressed. In various embodiments, the method involves contacting the target sequence with an RGN ribonucleoprotein complex. The RGN ribonucleoprotein complex can include an RGN that is nuclease-dead or has nickase activity. In some embodiments, the RGN of the ribonucleoprotein complex is a fusion polypeptide that includes a base editing polypeptide. In some embodiments, the method involves introducing the RGN ribonucleoprotein complex into a cell, or an organelle, or an embryo that contains the target sequence. The RGN ribonucleoprotein complex can be one that is purified from a biological sample, or one that is purified after recombinant production, or one that is assembled in vitro as described herein. In embodiments where the RGN ribonucleoprotein complex assembled in vitro is contacted with any of the target sequence, cell, organelle, embryo, the method can further include contacting the complex with any of the target sequence, cell, organelle, embryo after assembling the complex in vitro.
[0130] The RGN ribonucleoprotein complex, which is purified or assembled in vitro, can be introduced into any of a cell, an organelle, an embryo using any method known in the art, non-limiting examples of which include electroporation. Alternatively, the polynucleotide encoding or containing the RGN polypeptide, and / or the guide RNA, can be introduced using any method known in the art (such as electroporation).
[0131] When the guide RNA is delivered to the target sequence or to a cell, organelle, or embryo containing the target sequence, or comes into contact with the target sequence or a cell, organelle, or embryo containing the target sequence, the RGN is bound to the target sequence in a sequence-specific manner. In embodiments where the RGN has nuclease activity, the RGN polypeptide cleaves the target sequence upon binding to the target sequence of interest. The target sequence is then modified through an endogenous repair mechanism (e.g., non-homologous end joining repair or homologous recombination repair using a provided donor polynucleotide).
[0132] Methods for measuring the binding of the RGN polypeptide to the target sequence are known in the art and include chromatin immunoprecipitation assays, gel mobility shift assays, DNA pull-down assays, reporter assays, and microplate capture-detection assays. Similarly, methods for measuring the cleavage or modification of the target sequence are known in the art and include in vitro or in vivo cleavage assays. In this cleavage assay, PCR sequencing or gel electrophoresis is used to confirm cleavage, with or without an appropriate label (e.g., radioisotope, fluorophore) for the target sequence to facilitate detection of the degradation products. Alternatively, a cleavage-triggered exponential amplification reaction (NTEXPAR) assay can be utilized (see, e.g., Zhang et al. (2016) Chem. Sci. Vol. 7:4951-4957). In vivo cleavage can be evaluated using a Surveyor assay (Guschin et al. (2010) Methods Mol Biol Vol. 649:247-256).
[0133] In some embodiments, the method includes the use of an RGN complexed with two or more guide RNAs. The two or more guide RNAs can target different regions of a single gene or multiple genes.
[0134] In embodiments where no donor polynucleotide is provided, the double-strand breaks introduced by the RGN polypeptide can be repaired by the non-homologous end joining (NHEJ) repair process. Since NHEJ is error-prone, the target sequence may be modified by the repair of the double-strand break. As used herein with respect to nucleic acid molecules, "modification" means a change in the nucleotide sequence of the nucleic acid molecule, and deletions, insertions, substitutions of one or more nucleotides, or combinations thereof are possible. Modification of the target sequence may result in the expression of an altered protein product or inactivation of the coding sequence.
[0135] In embodiments where a donor polynucleotide is present, while the introduced double-strand break is being repaired, the donor sequence in the donor polynucleotide can be integrated into or exchanged with the target nucleotide sequence, resulting in the introduction of a foreign donor sequence. Thus, the donor polynucleotide contains a donor sequence that is desired to be introduced into the target sequence of interest. In some embodiments, the donor sequence changes the original target nucleotide sequence so that the newly integrated donor sequence is not recognized by the RGN and is not cleaved by the RGN. Integration of the donor sequence can be enhanced by including in the donor polynucleotide a flanking sequence that is substantially identical in sequence to the sequence adjacent to the target nucleotide sequence, enabling the homologous recombination repair process. In embodiments where the RGN polypeptide introduces sticky ends at the double-strand break, the donor polynucleotide can contain a donor sequence flanked by compatible overhangs, allowing the donor sequence to be directly ligated to the cleaved target nucleotide sequence containing the overhangs by the non-homologous recombination repair process while the double-strand break is being repaired.
[0136] In embodiments where the method involves the use of an RGN that is a nickase (i.e., capable of cleaving only a single strand of a double-stranded polynucleotide), the method can include introducing two RGN nickases that target the same target sequence or overlapping target sequences and cleaving different strands of the polynucleotide. For example, an RGN nickase that cleaves only the plus (+) strand of a double-stranded polynucleotide can be introduced together with a second RGN nickase that cleaves only the minus (−) strand of the double-stranded polynucleotide.
[0137] In various embodiments, methods are provided for detecting a target nucleotide sequence by binding thereto, the method including introducing into a cell, or an organelle, or an embryo at least one guide RNA or a polynucleotide encoding the same, and at least one RGN polypeptide or a polynucleotide encoding the same, and expressing the guide RNA and / or the RGN polypeptide (when introducing the coding sequence; the RGN polypeptide is a nuclease-dead RGN and further includes a detectable label), the method further including detecting the detectable label. The detectable label can be fused to the RGN as a fusion protein (e.g., a fluorescent protein). Alternatively, as the detectable label, a small molecule that is complexed with the RGN polypeptide or incorporated into the RGN polypeptide and can be detected visually or by simple means is possible.
[0138] This specification also provides a method for changing the expression of a target sequence of interest, or the expression of a gene of interest under the regulation of the target sequence. This method involves introducing at least one guide RNA or a polynucleotide encoding the same, and at least one RGN polypeptide or a polynucleotide encoding the same into a cell, or an organelle, or an embryo, and expressing this guide RNA and / or RGN polypeptide (when introducing the coding sequence; this RGN polypeptide is a nuclease-dead RGN). In some of these embodiments, the nuclease-dead RGN is a fusion protein comprising an expression modulator domain as described herein (i.e., an epigenetic modification domain, or a transcriptional activation domain, or a transcriptional repression domain).
[0139] The present disclosure also provides a method for binding to a target sequence of interest and / or modifying a target sequence of interest. These methods involve delivering a system comprising at least one guide RNA or a polynucleotide encoding the same, and a fusion polypeptide containing the RGN of the present invention and a base editing polypeptide (e.g., cytidine deaminase or adenosine deaminase) or a polynucleotide encoding this fusion polypeptide to the target sequence, or to a cell, an organelle, or an embryo containing the target sequence.
[0140] Those skilled in the art will appreciate that any method of the present disclosure can be utilized to target a single target sequence or multiple target sequences. For example, this method involves using a single RGN polypeptide in combination with a number of different guide RNAs, and can target a number of different sequences within a single gene and / or multiple genes. The present invention also encompasses a method of introducing a number of different guide RNAs in combination with a number of different RGN polypeptides. These guide RNAs and these guide RNA / RGN polypeptide systems can target a number of different sequences within a single gene and / or multiple genes.
[0141] In one aspect, the present invention provides a kit comprising any one or more of the elements disclosed in the above methods and compositions. In some embodiments, the kit includes a vector system and instructions for using the kit. In some embodiments, the vector system comprises (a) a first regulatory element operably linked to a tracr mate sequence and one or more insertion sites for inserting a guide sequence upstream of the tracr mate sequence (the guide sequence, when expressed, sequence-specifically binds a CRISPR complex to a target sequence in a eukaryotic cell; the CRISPR complex comprises a CRISPR enzyme complexed with (1) a guide sequence that hybridizes to the target sequence and (2) a tracr mate sequence that hybridizes to the tracr sequence); and / or (b) a second regulatory element operably linked to an enzyme coding sequence encoding a CRISPR enzyme comprising a nuclear localization sequence. The elements can be provided individually or in combination and can be provided in any suitable container (such as vials, bottles, test tubes, etc.).
[0142] In some embodiments, the kit includes instructions in one or more languages. In some embodiments, the kit includes one or more reagents for use in a method of utilizing one or more of the elements described herein. The reagents can be provided in any suitable container. For example, the kit can provide one or more reaction buffers or storage buffers. The reagents can be provided in a form that can be used in a particular assay or in a form that requires the addition of one or more other components prior to use (such as a concentrate or lyophilized form). Any buffer can be used as the buffer, and non-limiting examples thereof include sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of from about 7 to about 10.
[0143] In some embodiments, the kit includes one or more oligonucleotides corresponding to a guide sequence for insertion into a vector to operably link the guide sequence and the regulatory element. In some embodiments, the kit includes a homologous recombination template polynucleotide. In one aspect, the present invention provides a method for using one or more elements of a CRISPR system. The CRISPR complex of the present invention provides an effective means for modifying a target polynucleotide. The CRISPR complex of the present invention has diverse utilities, including modification of target polynucleotides (e.g., deletion, insertion, translocation, inactivation, activation) in many types of cells. Therefore, the CRISPR complex of the present invention has a wide range of applications in, for example, gene therapy, drug screening, disease diagnosis, and prognosis prediction. One representative CRISPR complex includes a CRISPR enzyme complexed with a guide sequence that hybridizes to a target sequence within the target polynucleotide.
[0144] VIII. Target Polynucleotide
[0145] In one aspect, the present invention provides a method for modifying a target polynucleotide in a eukaryotic cell. This method can be carried out in vivo, ex vivo, or in vitro. In some embodiments, this method includes sampling one cell or a population of cells from a human, or a non-human animal, or a plant (including microalgae), and modifying the cells. Culturing can be carried out ex vivo at any stage. The cells can even be reintroduced into a non-human animal or a plant (including microalgae).
[0146] Plant breeders utilize natural variability to combine the most useful genes in order to seek desirable qualities (such as yield, quality, uniformity, robustness, resistance to pathogens, etc.). These desirable qualities also include growth, preference for day length, temperature conditions, onset of flowering or reproduction / development, fatty acid content, insect resistance, disease resistance, nematode resistance, fungal resistance, herbicide resistance, and tolerance to various environmental factors (drought, heat, humidity, cold, wind, adverse soil conditions (including high salinity)). Sources of these useful genes include native or exotic species, heirloom varieties, wild plant relatives, induced mutations (e.g., treating plant material with a mutagen). By utilizing the present invention, plant breeders are provided with a new tool for inducing mutations. Thus, those skilled in the art can analyze the genome to search for useful genes, and utilize the present invention among varieties with desired characteristics or traits to more precisely induce an increase in useful genes than previous mutagens, and thus can achieve the acceleration and improvement of plant breeding programs.
[0147] As the target polynucleotide of the RGN system, any polynucleotide that is endogenous or exogenous to eukaryotic cells is possible. For example, as the target polynucleotide, a polynucleotide existing in the nucleus of eukaryotic cells is possible. As the target polynucleotide, a sequence encoding a gene product (such as a protein), or a non-coding sequence (such as a regulatory polynucleotide or junk DNA) is possible. Although not wishing to be bound by theory, it is considered that the target sequence should be associated with a PAM (protospacer adjacent motif), i.e., a short sequence recognized by the CRISPR complex. The conditions regarding the exact sequence and length of the PAM vary depending on the CRISPR enzyme used, but the PAM is typically a sequence of 2 to 5 base pairs adjacent to the protospacer (i.e., the target sequence).
[0148] The target polynucleotides of the CRISPR complex can include genes and polynucleotides associated with diseases, as well as genes and polynucleotides associated with biochemical signaling pathways. Examples of target polynucleotides include sequences associated with biochemical signaling pathways, such as genes or polynucleotides associated with biochemical signaling pathways. Examples of target polynucleotides include genes or polynucleotides associated with diseases. A "disease-associated" gene or polynucleotide means any gene or polynucleotide that gives rise to abnormal levels, or abnormal forms of transcripts or translation products, in cells derived from diseased tissue as compared to control tissue or cells that are not diseased. It can be a gene that is expressed at abnormally high levels or a gene that is expressed at abnormally low levels, and the altered expression is correlated with the development and / or progression of the disease. A disease-associated gene also means a gene having a mutation or genetic variation that is directly responsible for the origin of the disease (e.g., a causative mutation), or a gene having a mutation or genetic variation that is in linkage equilibrium with a gene responsible for the origin of the disease. The transcribed or translated products can be known or unknown, and can further be at normal or abnormal levels. Examples of disease-associated genes and polynucleotides can be obtained on the World Wide Web from the McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University (Baltimore, Maryland), and the National Center for Biotechnology Information, National Library of Medicine (Bethesda, Maryland).
[0149] The CRISPR system is particularly useful because it is relatively easy to target a desired genomic sequence, but problems remain regarding what RGNs can do to address causative mutations. One approach is to create a fusion protein between an RGN (preferably an inactive variant or nickase variant of the RGN) and a base editing enzyme (such as a cytidine deaminase or an adenosine deaminase base editor), or an active domain of a base editing enzyme (U.S. Patent No. 9,840,699; incorporated herein by reference). In some embodiments, this method involves contacting a DNA molecule (a) with a fusion protein comprising an RGN of the invention and a base editing polypeptide (such as a deaminase), and (b) with a gRNA that directs the fusion protein of (a) towards a target nucleotide sequence of the DNA strand; the DNA molecule is contacted with the fusion protein and the gRNA in a sufficient amount and under conditions suitable for deamination of nucleotide bases. In some embodiments, the target DNA sequence comprises a sequence associated with a disease or disorder, and deamination of the nucleotide bases therein results in a sequence not associated with the disease or disorder. In some embodiments, the target DNA sequence is present in an allele of a crop plant, and the trait of interest being a particular allele results in a plant of lower agricultural value. Deamination of the nucleotide bases results in an allele that improves the trait and increases the agricultural value of the plant.
[0150] In some embodiments, the DNA sequence comprises a T→C point mutation or an A→G point mutation associated with a disease or disorder, and deamination of the mutated C or G base results in a sequence not associated with the disease or disorder. In some embodiments, deamination corrects a mutation within the sequence that is associated with a disease or disorder.
[0151] In some embodiments, the sequence related to the disease or disorder encodes a protein, and a stop codon is introduced into the sequence related to the disease or disorder by deamination, resulting in the termination of the encoded protein. In some embodiments, the contacting operation is performed in vivo on a subject likely to have a disease or disorder, or a subject having a disease or disorder, or a subject diagnosed with a disease or disorder. In some embodiments, the disease or disorder is a disease associated with a point mutation or single nucleotide mutation in the genome. In some embodiments, the disease is any of a genetic disease, cancer, a metabolic disease, or a lysosomal storage disorder.
[0152] Another example of a locus that causes a certain genetic disease, particularly a locus that can be easily targeted by the RGN or RGN-base editor fusion protein of the present invention, can be found in Example 9 and the corresponding Table 12.
[0153] Hurler syndrome
[0154] An example of a genetic disease that can be corrected using the approach relying on the RGN-base editor fusion protein of the present invention is Hurler syndrome. Hurler syndrome (also known as MPS-1) is the result of a deficiency of α-L-iduronidase (IDUA), and is a lysosomal storage disease characterized by the accumulation of dermatan sulfate and heparan sulfate in lysosomes at the molecular level. This disease is generally a hereditary genetic disorder caused by mutations in the IDUA gene encoding α-L-iduronidase. Common IDUA mutations are W402X and Q70X, both of which are nonsense mutations that cause translation to stop midway. Such mutations are successfully addressed by the precision genome editing (PGE) approach. This is because, for example, restoring one nucleotide by a base editing approach restores the wild-type coding sequence, and protein expression is controlled by the endogenous regulatory mechanism at the locus. In addition, since heterozygotes are known to be asymptomatic, PGE therapy targeting one of these mutations is considered useful for many patients with this disease because only one of the mutated alleles needs to be corrected (Bunge et al. (1994) Hum. Mol. Genet. Vol. 3(6):861-866; incorporated herein by reference).
[0155] Current treatments for Hurler syndrome include enzyme replacement therapy and bone marrow transplantation (Vellodi et al. (1997) Arch. Dis. Child. Vol. 76(2):92-99; Peters et al. (1998) Blood Vol. 91(7):2601-2608; incorporated herein by reference). Enzyme replacement therapy has had a dramatic effect on the survival and quality of life of patients with Hurler syndrome, but this approach requires costly and time-consuming weekly infusions. Additional approaches include delivery of the IDUA gene on an expression vector, or insertion of this gene into highly expressed loci (such as the locus of the serum albumin gene) (U.S. Patent No. 9,956,247; incorporated herein by reference). However, with these approaches, the original IDUA locus is not restored to the exact coding sequence. Genome editing strategies are thought to have a number of advantages, the most notable of which is that gene expression regulation is thought to be controlled by natural mechanisms present in healthy individuals. In addition, the use of base editing does not require the cleavage of double-stranded DNA. Cleavage of double-stranded DNA can lead to large-scale chromosomal rearrangements, cell death, and cancer development due to disruption of tumor suppressor mechanisms. A description of a method for correcting the causative mutations of this disease is presented in Example 10. The method described is an example of a general strategy using the RGN-base editor fusion protein of the present invention, which is directed against several mutations in the human genome that cause the disease and corrects those mutations. It will be appreciated that similar approaches targeting diseases such as those listed in Table 12 can be pursued. Furthermore, it will be appreciated that similar approaches targeting mutations that cause diseases in other species (especially common pets and livestock) can be developed using the RGN of the present invention. Common pets and livestock include dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, fish (including salmon), and shrimp.
[0156] Friedreich's ataxia
[0157] The RGN of the present invention may be useful in the treatment of human diseases with more complex causative mutations. For example, several diseases such as Friedreich's ataxia and Huntington's disease are the result of a significantly increased number of repeats of a 3-nucleotide motif in a specific region of a gene, which affects the ability of the expressed protein to function or to be expressed. Friedreich's ataxia (FRDA) is an autosomal recessive disease in which the nerve tissue within the spinal cord degenerates gradually. A decrease in the level of the frataxin (FXN) protein within mitochondria results in oxidative damage and iron deficiency at the cellular level. The decreased expression of FXN has been linked to an expansion of the GAA triplet within intron 1 of the FXN gene in somatic and germline cells. In FRDA patients, the number of GAA repeats is often greater than 70, and sometimes even exceeds 1000 triplets (the most common being 600 - 900), whereas individuals without this disease have less than about 40 repeats (Pandolfo et al. (2012) Handbook of Clinical Neurology Vol. 103: 275 - 294; Campuzano et al. (1996) Science Vol. 271: 1423 - 1427; Pandolfo (2002) Adv. Exp. Med. Biol. Vol. 516: 99 - 118; all of which are incorporated herein by reference).
[0158] The expansion of the three-nucleotide repeat sequence that causes Friedreich's ataxia (FRDA) occurs within a specific locus in the FXN gene, called the FRDA instability region. The instability region in cells of FRDA patients can be excised using an RNA-guided nuclease (RGN). This approach requires: 1) an RGN-guide RNA sequence that can be programmed to target alleles in the human genome; and 2) a method for delivering this RGN-guide sequence. Many nucleases used for genome editing, such as the Cas9 nuclease from Streptococcus pyogenes (SpCas9), which is commonly used, are too large to be packaged into adeno-associated virus (AAV) vectors. This can be seen particularly when considering the length of the SpCas9 gene and the guide RNA, as well as the lengths of other genetic elements required for a functional expression cassette. Therefore, approaches using SpCas9 are more difficult.
[0159] The compact RNA-guided nucleases of the present invention (particularly APG07433.1 and APG08290.1) are highly suitable for the excision of the FRDA instability region. Each RGN requires a PAM in the vicinity of the FRDA instability region. In addition, each of these RGNs can be packaged into an AAV vector together with a guide RNA. Although a second vector is likely required to package two guide RNAs, this approach is still advantageous over the vectors that may be required for larger nucleases (such as SpCas9, which may have to split the protein sequence into two vectors). A description enabling a method for correcting the disease-causing mutations is presented in Example 11. The described methods include strategies for removing regions of genomic instability using the RGNs of the present invention. Such strategies can be applied to other diseases and disorders (such as Huntington's disease) with similar genetic bases. Similar strategies using the RGNs of the present invention can also be applied to agriculturally or economically important diseases and disorders of non-human animals. Such non-human animals include dogs, cats, horses, pigs, cows, sheep, chickens, donkeys, snakes, ferrets, fish (including salmon), and shrimp.
[0160] Hemoglobinopathy
[0161] The RGNs of the present invention are also thought to be able to introduce disruptive mutations that may have advantageous effects. Gene defects in genes encoding hemoglobin, particularly the β-globin chain (HBB gene), may be responsible for a number of diseases known as hemoglobinopathies (including sickle cell anemia and thalassemia).
[0162] In adults, hemoglobin is a heterotetramer that consists of two alpha (α)-like globin chains, two beta (β)-like globin chains, and four heme groups. In adults, this α2β2 tetramer is called hemoglobin A (HbA) or adult hemoglobin. Typically, alpha globin chains and beta globin chains are synthesized in an approximate 1:1 ratio, and this ratio appears to be extremely important for the stability of hemoglobin and red blood cells (RBCs). In the developing fetus, different forms of hemoglobin (fetal hemoglobin (HbF)) with a greater oxygen-binding affinity than HbA are produced so that oxygen can be delivered to the fetal system through the maternal bloodstream. Fetal hemoglobin also contains two alpha globin chains but has two fetal gamma (γ) globin chains instead of the adult beta globin chains (i.e., fetal hemoglobin is α2γ2). The regulation of the switch from gamma globin production to beta globin production is extremely complex and mainly involves the downregulation of gamma globin transcription and the simultaneous upregulation of beta globin transcription. At around 30 weeks of gestation, the synthesis of gamma globin in the fetal body begins to decline, while the production of beta globin increases. Neonatal hemoglobin is almost entirely α2β2 until about 10 months of age, although some HbF remains until adulthood (about 1-3% of total hemoglobin). As described above, most patients with hemoglobinopathies have the gene encoding gamma globin present, but expression is relatively low due to the suppression of the normal gene that occurs around the time of delivery as described above.
[0163] Sickle cell disease is caused by a V6E mutation in the beta globin gene (HBB) (from GAG to GTG at the DNA level), and the resulting hemoglobin is called "hemoglobin S" or "HbS". Under hypoxic conditions, HbS molecules aggregate to form fibrous precipitates. These aggregates cause abnormalities or "sickling" of RBCs, resulting in a loss of cell flexibility. These sickled RBCs can no longer enter the capillary bed, so sickle cell patients are at risk of vaso-occlusive crises. In addition, sickled RBCs are more fragile and prone to hemolysis than normal RBCs, and ultimately the patient becomes anemic.
[0164] The treatment and management of patients with sickle cell disease are lifelong challenges, including antibiotic treatment, pain management, and transfusion during the acute phase. One approach is the use of hydroxyurea. Hydroxyurea exerts some of its effects by increasing the production of gamma globin. The long-term side effects of chronic hydroxyurea therapy are not yet known, but the treatment can cause unwanted side effects, and the effectiveness may vary from patient to patient. Despite the increasing effectiveness of sickle cell treatment, the average life expectancy of patients is still only in the mid- to late 50s, and the pathological conditions associated with this disease have a profound impact on the quality of life of patients.
[0165] Thalassemia (alpha thalassemia and beta thalassemia) is also a disease related to hemoglobin, typically involving a reduced expression of globin chains. This occurs either through mutations in the regulatory regions of the genes or from a decrease in the expression or level of the globin protein that functions due to mutations within the globin coding sequences. The treatment of thalassemia usually includes blood transfusion and iron chelation therapy. Bone marrow transplantation is also used in the treatment of severely affected patients when a suitable donor can be identified, but this method can be risky.
[0166] One approach that has been proposed to treat both SCD and beta thalassemia is to increase the expression of gamma globin and functionally replace abnormal adult hemoglobin A with HbF. As noted above, treatment of SCD patients with hydroxyurea is thought to be successful, in part, because hydroxyurea is effective in increasing the expression of gamma globin (DeSimone (1982) Proc Nat'l Acad Sci USA Vol. 79(14):4428-4431; Ley et al. (1982) N. Engl. J. Medicine, Vol. 307:1469-1475; Ley et al. (1983) Blood Vol. 62:370-380; Constantoulakis et al. (1988) Blood Vol. 72(6):1961-1967; all of which are incorporated herein by reference). Increasing the expression of HbF involves the identification of genes whose products play a role in regulating the expression of gamma globin. One such gene is BCL11A. BCL11A encodes a zinc finger protein that is expressed in adult erythroid progenitor cells and whose expression, when downregulated, results in increased expression of gamma globin (Sankaran et al. (2008) Science Vol. 322:1839; incorporated herein by reference). The use of inhibitory RNAs targeting the BCL11A gene has been proposed (e.g., U.S. Patent Application Publication No. 2011 / 0182867; incorporated herein by reference), but this technique has several potential drawbacks. Such drawbacks include the possibility that a complete knockdown may not be achievable, the potential for problems with the delivery of such RNAs, and the need for multiple treatments over a lifetime because such RNAs must be continuously present.
[0167] Using the RGN of the present invention to target the BCL11A enhancer region and disrupting the expression of BCL11A increases the expression of gamma globin. This target disruption can be achieved by non-homologous end joining (NHEJ). By NHEJ, the RGN of the present invention targets a specific sequence within the BCL11A enhancer region to cleave the double strand, and the cell's mechanism repairs the cleavage, typically introducing a harmful mutation at the same time. Similar to what has been described for other disease targets, the RGN of the present invention is relatively small in size and, for the purpose of in vivo delivery, has advantages compared to other known RGNs because the expression cassettes for the RGN and its corresponding guide RNA can be packaged in a single AAV vector. The description enabling this method is presented in Example 12. A similar strategy using the RGN of the present invention can also be applied to the same diseases and disorders in both humans and agriculturally or economically important non-human animals.
[0168] IX. Cells Containing Polynucleotide Gene Modification
[0169] As described herein, cells and organisms are provided that contain a target sequence of interest modified using the methods mediated by the RGN, and / or crRNA, and / or tracrRNA described herein. In some of these embodiments, the RGN contains the amino acid sequence of any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or an active variant or fragment thereof. In various embodiments, the guide RNA contains a CRISPR repeat sequence that contains the nucleotide sequence of any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, or an active variant or fragment thereof. In a particular embodiment, the guide RNA contains a tracrRNA that contains the nucleotide sequence of any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56, or an active variant or fragment thereof. As the guide RNA of this system, a single guide RNA or a double guide RNA is possible.
[0170] As the modified cell, a eukaryotic cell (e.g., mammalian cell, plant cell, insect cell) or a prokaryotic cell is possible. Also provided are an organelle and an embryo comprising at least one nucleotide sequence modified by the method using the RGN, and / or crRNA, and / or tracrRNA described herein. A cell, or an organism, or an organelle, or an embryo in which a gene is modified can be heterozygous or homozygous with respect to the modified nucleotide sequence.
[0171] As a result of modification of any chromosome of a cell, an organism, an organelle, or an embryo, a changed protein product or a change (upregulation or downregulation) in the expression of the integrated sequence, or inactivation, or expression may occur. When gene inactivation or expression of a non-functional protein product occurs due to chromosomal modification, the cell, or the organism, or the organelle, or the embryo in which the gene is modified is called a "knockout". As a result of the knockout phenotype, a deletion mutation (i.e., deletion of at least one nucleotide), or an insertion mutation (i.e., insertion of at least one nucleotide), or a nonsense mutation (i.e., substitution of at least one nucleotide to introduce a stop codon) may occur.
[0172] Alternatively, as a result of modification of any chromosome of a cell, an organism, an organelle, or an embryo, a "knock-in" may be generated. This is the result of the nucleotide sequence encoding a protein being integrated into the chromosome. In some of these embodiments, the chromosomal sequence encoding the wild-type protein is inactivated by the integration of the coding sequence into the chromosome, but the externally introduced protein is expressed.
[0173] In another embodiment, as a result of the chromosome being modified, a variant protein product is produced. The variant protein expressed may have at least one amino acid substitution and / or at least one amino acid addition or deletion. The variant protein product encoded by the altered chromosomal sequence may exhibit modified characteristics or activities compared to the wild-type protein, including, but not limited to, altered enzyme activity or substrate specificity.
[0174] In yet another embodiment, as a result of the chromosome being modified, the expression pattern of a protein may change. As a non-limiting example, as a result of a change in the chromosome within the regulatory region that controls the expression of the protein product, overexpression or downregulation of the protein product, or a changed tissue expression pattern, or a changed temporal expression pattern may occur.
[0175] As used herein, the article "one" is used to mean one or more (i.e., at least one) objects that are the grammatical subject. For example, "one polypeptide" means one or more polypeptides.
[0176] All publications and patent applications mentioned herein are indicative of the level of those skilled in the art to which the present disclosure pertains. Each publication and patent application is incorporated herein by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference.
[0177] Although the invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be apparent that some changes and modifications can be practiced within the scope of the appended embodiments.
[0178] Non-limiting embodiments include the following.
[0179] 1. A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence having at least 95% identity with any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, and when the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, binds to the target DNA sequence in a manner specific to the RNA-guided sequence, and the polynucleotide encoding the RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide.
[0180] 2. The nucleic acid molecule according to embodiment 1, wherein the RGN polypeptide is capable of cleaving the target DNA sequence when it binds to the target DNA sequence.
[0181] 3. The nucleic acid molecule according to embodiment 2, wherein double-strand cleavage occurs upon cleavage by the RGN polypeptide.
[0182] 4. The nucleic acid molecule according to embodiment 2, wherein single-strand cleavage occurs upon cleavage by the RGN polypeptide.
[0183] 5. The nucleic acid molecule according to any one of embodiments 1 to 4, wherein the RGN polypeptide is operably fused to a base editing polynucleotide.
[0184] 6. The nucleic acid molecule according to any one of embodiments 1 to 5, wherein the RGN polypeptide comprises one or more nuclear localization signals.
[0185] 7. The nucleic acid molecule according to any one of embodiments 1 to 6, wherein the RGN polypeptide is codon-optimized for expression in eukaryotic cells.
[0186] 8. The nucleic acid molecule according to any one of embodiments 1 to 7, wherein the target DNA sequence is located at a position adjacent to a protospacer adjacent motif (PAM).
[0187] 9. A vector comprising the nucleic acid molecule according to any one of embodiments 1 to 8.
[0188] 10. The vector according to embodiment 9, further comprising at least one nucleotide sequence encoding the gRNA capable of hybridizing to the target DNA sequence.
[0189] 11. The vector according to embodiment 10, wherein the gRNA is a single guide RNA.
[0190] 12. The vector according to embodiment 10, wherein the gRNA is a dual guide RNA.
[0191] 13. The vector according to any one of embodiments 10 to 12, wherein the guide RNA comprises a CRISPR RNA comprising a CRISPR repeat sequence having a sequence that is at least 95% identical to any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.
[0192] 14. The vector according to any one of embodiments 10 to 13, wherein the guide RNA comprises a tracrRNA having a sequence that is at least 95% identical to any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56.
[0193] 15. A cell comprising the nucleic acid molecule according to any one of embodiments 1 to 8, or comprising the vector according to any one of embodiments 9 to 14.
[0194] 16. A method for producing an RGN polypeptide, the method comprising culturing the cell according to embodiment 15 under conditions in which the RGN polypeptide is expressed.
[0195] A method for producing an RGN polypeptide, comprising introducing into a cell a heterologous nucleic acid molecule comprising a nucleotide sequence encoding an RNA-guided nuclease (RGN) polypeptide having an amino acid sequence that is at least 95% identical to any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54; when the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, binding the RGN polypeptide to the target DNA sequence in a manner specific to the RNA-guided sequence; culturing the cell under conditions in which the RGN polypeptide is expressed.
[0196] 18. The method according to embodiment 16 or 17, further comprising purifying the RGN polypeptide.
[0197] 19. The method according to embodiment 16 or 17, wherein the cell further expresses one or more guide RNAs that bind to the RGN polypeptide and form an RGN ribonucleoprotein complex.
[0198] 20. The method according to embodiment 19, further comprising purifying the RGN ribonucleoprotein complex.
[0199] 21. A nucleic acid molecule comprising a polynucleotide encoding a CRISPR RNA (crRNA), wherein the crRNA comprises a spacer sequence and a CRISPR repeat sequence, and the CRISPR repeat sequence comprises a nucleotide sequence that is at least 95% identical to any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55; a) the crRNA and; b) a trans-activating CRISPR RNA (tracrRNA) that hybridizes to the CRISPR repeat sequence of the crRNA wherein the guide RNA comprising is capable of hybridizing to a target DNA sequence in a sequence-specific manner via the spacer sequence of the crRNA when bound to an RNA-guided nuclease (RGN) polypeptide, and is such that when the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, the RGN polypeptide can be bound to the target DNA sequence in a manner specific to the RNA-guided sequence; A nucleic acid molecule, wherein the polynucleotide encoding the crRNA is operably linked to a promoter that is heterologous to the polynucleotide.
[0200] 22. A vector comprising the nucleic acid molecule according to embodiment 21.
[0201] 23. The vector according to embodiment 22, further comprising a polynucleotide encoding the tracrRNA.
[0202] 24. The vector according to embodiment 23, wherein the tracrRNA comprises a nucleotide sequence having at least 95% identity with any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56.
[0203] 25. The vector according to embodiment 23 or 24, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as a single guide RNA.
[0204] 26. The vector according to embodiment 23 or 24, wherein the polynucleotide encoding the crRNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.
[0205] 27. The vector according to any one of embodiments 22 to 26, further comprising a polynucleotide encoding the RGN polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54.
[0206] 28. A nucleic acid molecule comprising a polynucleotide encoding a trans-activating CRISPR RNA (tracrRNA) comprising a nucleotide sequence having at least 95% identity with any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56; a) the tracrRNA; b) A crRNA comprising a CRISPR repeat sequence to which the tracrRNA hybridizes and a spacer sequence A guide RNA comprising, When bound to an RNA-guided nuclease (RGN) polypeptide, can hybridize to a target DNA sequence in a sequence-specific manner via the spacer sequence of the crRNA, A nucleic acid molecule, wherein the polynucleotide encoding the tracrRNA is operably linked to a promoter that is heterologous to the polynucleotide.
[0207] 29. A vector comprising the nucleic acid molecule according to embodiment 28.
[0208] 30. The vector according to embodiment 29, further comprising a polynucleotide encoding the crRNA.
[0209] 31. The vector according to embodiment 30, wherein the CRISPR repeat sequence of the crDNA comprises a nucleotide sequence having at least 95% identity with any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.
[0210] 32. The vector according to embodiment 30 or 31, wherein the polynucleotide encoding the crDNA and the polynucleotide encoding the tracrRNA are operably linked to the same promoter and are encoded as a single guide RNA.
[0211] 33. The vector according to embodiment 30 or 31, wherein the polynucleotide encoding the crDNA and the polynucleotide encoding the tracrRNA are operably linked to separate promoters.
[0212] 34. Further comprising a polynucleotide encoding the RGN polypeptide, wherein the RGN polypeptide comprises an amino acid sequence having a sequence that is at least 95% identical to any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, The vector according to any one of embodiments 29 to 33.
[0213] 35. For binding to a target DNA sequence, a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having a sequence that is at least 95% identical to any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or a nucleotide sequence encoding the RGN polypeptide; A system in which each of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to each nucleotide sequence; the one or more guide RNAs hybridize to the target DNA sequence, the one or more guide RNAs form a complex with the RGN polypeptide, thereby binding the RGN polypeptide to the target DNA sequence.
[0214] 36. The system according to embodiment 35, wherein the gRNA is a single guide RNA (sgRNA).
[0215] 37. The system according to embodiment 35, wherein the gRNA is a dual guide RNA.
[0216] 38. The system according to any one of embodiments 35 to 37, wherein the gRNA comprises a CRISPR repeat sequence having a nucleotide sequence that is at least 95% identical to any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.
[0217] 39. The system according to any one of embodiments 35 to 38, wherein the gRNA comprises a tracrRNA comprising a nucleotide sequence having a sequence that is at least 95% identical to any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56.
[0218] 40. The system according to any one of embodiments 35 to 39, wherein the target DNA sequence is located at a position adjacent to a protospacer adjacent motif (PAM).
[0219] 41. The system according to any one of embodiments 35 to 40, wherein the target DNA sequence is in a cell.
[0220] 42. The system according to embodiment 41, wherein the cell is a eukaryotic cell.
[0221] 43. The system according to embodiment 42, wherein the eukaryotic cell is a plant cell.
[0222] 44. The system according to embodiment 42, wherein the eukaryotic cell is a mammalian cell.
[0223] 45. The system according to embodiment 42, wherein the eukaryotic cell is an insect cell.
[0224] 46. The system according to embodiment 41, wherein the cell is a prokaryotic cell.
[0225] 47. The system according to any one of embodiments 35 to 46, wherein when transcribed, the one or more guide RNAs hybridize to the target DNA sequence to form a complex with the RGN polypeptide, causing cleavage of the target DNA sequence.
[0226] 48. The system according to embodiment 47, wherein the cleavage results in a double-stranded break.
[0227] 49. The system according to embodiment 47, wherein the cleavage by the RGN polypeptide results in a single-stranded break.
[0228] 50. The system according to any one of embodiments 35 to 49, wherein the RGN polypeptide is operably linked to a base editing polypeptide.
[0229] 51. The system according to any one of embodiments 35 to 50, wherein the RGN polypeptide comprises one or more nuclear localization signals.
[0230] 52. The system according to any one of embodiments 35 to 51, wherein the RGN polypeptide is codon-optimized for expression in eukaryotic cells.
[0231] 53. The system according to any one of embodiments 35 to 52, wherein the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide are located on one vector.
[0232] 54. The system according to any one of embodiments 35 to 53, further comprising one or more donor polynucleotides, or one or more nucleotide sequences encoding the one or more donor polynucleotides.
[0233] 55. A method of binding to a target DNA sequence, the method comprising delivering the system according to any one of embodiments 35 to 54 to the target DNA sequence or to a cell comprising the target DNA sequence.
[0234] 56. The method according to embodiment 55, wherein detection of the target DNA sequence is enabled by further comprising a detectable label on the RGN polypeptide or the guide RNA.
[0235] 57. The method according to embodiment 55, wherein expression of the target DNA sequence, or expression of a gene under transcriptional control by the target DNA sequence, is altered by further comprising an expression modulator on the guide RNA or the RGN polypeptide.
[0236] 58. A method for cleaving or modifying a target DNA sequence, comprising delivering the system according to any one of Embodiments 35 to 54 to the target DNA sequence or to a cell containing the target DNA sequence.
[0237] 59. The method according to Embodiment 58, wherein the modified target DNA sequence comprises inserting a heterologous DNA into the target DNA sequence.
[0238] 60. The method according to Embodiment 58, wherein the modified target DNA sequence comprises deleting at least one nucleotide from the target DNA sequence.
[0239] 61. The method according to Embodiment 58, wherein the modified target DNA sequence comprises at least one nucleotide mutation in the target DNA sequence.
[0240] 62. A method for binding to a target DNA sequence, a) i) one or more guide RNAs capable of hybridizing to the target DNA sequence; ii) an RGN polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54 are combined in vitro to assemble an RNA-guided nuclease (RGN) ribonucleotide complex under conditions suitable for the formation of this RGN ribonucleotide complex; b) contacting the target DNA sequence, or a cell containing the target DNA sequence, with the RGN ribonucleotide complex assembled in vitro; wherein the one or more guide RNAs hybridize to the target DNA sequence, whereby the RGN polypeptide binds to the target DNA sequence.
[0241] 63. The method according to Embodiment 62, wherein the detection of the target DNA sequence is enabled by the RGN polypeptide or the guide RNA further comprising a detectable label.
[0242] 64. The method according to embodiment 62, wherein the guide RNA or the RGN polypeptide further comprises an expression modulator, enabling a change in the expression of the target DNA sequence.
[0243] 65. A method for cleaving and / or modifying a target DNA sequence, the DNA molecule being a) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% identity with any of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54; b) contacting one or more guide RNAs capable of directing the RGN of (a) to the target DNA sequence, wherein the one or more guide RNAs hybridize to the target DNA sequence, causing the RGN polypeptide to bind to the target DNA sequence and cleavage and / or modification of the target DNA sequence to occur.
[0244] 66. The method according to embodiment 65, wherein the modified target DNA sequence comprises insertion of a heterologous DNA into the target DNA sequence.
[0245] 67. The method according to embodiment 65, wherein the modified target DNA sequence comprises deletion of at least one nucleotide from the target DNA sequence.
[0246] 68. The method according to embodiment 65, wherein the modified target DNA sequence comprises mutation of at least one nucleotide in the target DNA sequence.
[0247] 69. The method according to any one of embodiments 62 to 68, wherein the gRNA is a single guide RNA (sgRNA).
[0248] 70. The method according to any one of embodiments 62 to 68, wherein the gRNA is a dual guide RNA.
[0249] 71. The method according to any one of embodiments 62 to 70, wherein the gRNA comprises a CRISPR repeat sequence comprising a nucleotide sequence having a sequence that is at least 95% identical to any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55.
[0250] 72. The method according to any one of embodiments 62 to 71, wherein the gRNA comprises a tracrRNA comprising a nucleotide sequence having a sequence that is at least 95% identical to any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56.
[0251] 73. The method according to any one of embodiments 62 to 72, wherein the target DNA sequence is located at a position adjacent to a protospacer adjacent motif (PAM).
[0252] 74. The method according to any one of embodiments 55 to 73, wherein the target DNA sequence is in a cell.
[0253] 75. The method according to embodiment 74, wherein the cell is a eukaryotic cell.
[0254] 76. The method according to embodiment 75, wherein the eukaryotic cell is a plant cell.
[0255] 77. The method according to embodiment 75, wherein the eukaryotic cell is a mammalian cell.
[0256] 78. The method according to embodiment 75, wherein the eukaryotic cell is an insect cell.
[0257] 79. The method according to embodiment 74, wherein the cell is a prokaryotic cell.
[0258] 80. The method according to any one of embodiments 74 to 79, wherein the culturing of the cell is performed under conditions such that the RGN polypeptide is expressed to generate a modified DNA sequence by cleaving the target DNA sequence; and further comprising selecting a cell comprising the modified DNA sequence.
[0259] A cell comprising a target DNA sequence modified according to the method described in embodiment 80.
[0260] 82. The cell according to embodiment 81, which is a eukaryotic cell.
[0261] 83. The cell according to embodiment 82, wherein the eukaryotic cell is a plant cell.
[0262] 84. A plant comprising the cell according to embodiment 83.
[0263] 85. A seed comprising the cell according to embodiment 83.
[0264] 86. The cell according to embodiment 82, wherein the eukaryotic cell is a mammalian cell.
[0265] 87. The cell according to embodiment 82, wherein the eukaryotic cell is an insect cell.
[0266] 88. The cell according to embodiment 81, which is a prokaryotic cell.
[0267] 89. A method for producing a genetically modified cell in which a causative mutation of a genetic disease is corrected, comprising introducing into a cell: a) an RNA-guided (RGN) polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or a polynucleotide encoding this RGN polypeptide and operably linked to a promoter that enables the expression of this RGN polypeptide in the cell; and b) a guide RNA (gRNA) comprising a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% identity with any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, or a polynucleotide encoding this gRNA and operably linked to a promoter that enables the expression of this gRNA in the cell, wherein the RGN and the gRNA are directed to the genomic position of the causative mutation to modify the sequence at this genomic position and remove the causative mutation.
[0268] 90. The method according to embodiment 89, wherein the RGN is fused to a polypeptide having base editing activity.
[0269] 91. The method according to embodiment 90, wherein the polypeptide having base editing activity is cytidine deaminase or adenosine deaminase.
[0270] 92. The method according to embodiment 89, wherein the cell is an animal cell.
[0271] 93. The method according to embodiment 89, wherein the cell is a mammalian cell.
[0272] 94. The method according to embodiment 92, wherein the cell is derived from any one of dog, cat, mouse, rat, rabbit, horse, cow, pig, and human.
[0273] 95. The method according to embodiment 92, wherein the genetic disease is a disease listed in Table 12.
[0274] 96. The method according to embodiment 92, wherein the genetic disease is Hurler syndrome.
[0275] 97. The method according to embodiment 96, wherein the gRNA further comprises a spacer sequence targeting any one of SEQ ID NOs: 453, 454, and 455.
[0276] 98. A method for producing a genetically modified cell having a deletion in a genomic instability region causing a disease, wherein in this cell, a) an RNA-guided (RGN) polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, and 54, or a polynucleotide encoding this RGN polypeptide and operably linked to a promoter that enables the expression of this RGN polypeptide in the cell; b) A guide RNA (gRNA) comprising a CRISPR repeat sequence having a nucleotide sequence that is at least 95% identical to any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, and further comprising a spacer sequence targeting the 5' flanking region of the genomic instability region, or a polynucleotide encoding this gRNA and operably linked to a promoter that enables the expression of this gRNA in a cell; c) By introducing a second guide RNA (gRNA) comprising a CRISPR repeat sequence having a nucleotide sequence that is at least 95% identical to any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, and further comprising a spacer sequence targeting the 3' flanking region of the genomic instability region, or a polynucleotide encoding this gRNA and operably linked to a promoter that enables the expression of this gRNA in a cell, A method comprising directing the RGN and the two gRNAs towards the genomic instability region and removing at least a portion of this genomic instability region.
[0277] 99. The method according to embodiment 98, wherein the cell is an animal cell.
[0278] 100. The method according to embodiment 98, wherein the cell is a mammalian cell.
[0279] 101. The method according to embodiment 100, wherein the cell is derived from any one of dog, cat, mouse, rat, rabbit, horse, cow, pig, human.
[0280] 102. The method according to embodiment 99, wherein the genetic disease is Friedreich's ataxia or Huntington's disease.
[0281] 103. The method according to embodiment 102, wherein the first gRNA further comprises a spacer sequence targeting any one of SEQ ID NOs: 468, 469, 470.
[0282] 104. The method according to embodiment 103, wherein the second gRNA further comprises a spacer sequence targeting SEQ ID NO: 471.
[0283] 105. A method for producing genetically modified mammalian hematopoietic progenitor cells with reduced expression of BCL11A mRNA and protein, comprising introducing into isolated human hematopoietic progenitor cells: a) an RNA-guided (RGN) polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or a polynucleotide encoding this RGN polypeptide and operably linked to a promoter capable of enabling the expression of this RGN polypeptide intracellularly; b) a guide RNA (gRNA) comprising a CRISPR repeat sequence comprising a nucleotide sequence having at least 95% identity with any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55, or a polynucleotide encoding this gRNA and operably linked to a promoter capable of enabling the expression of this gRNA intracellularly, such that the RGN and the gRNA are expressed intracellularly to achieve cleavage at the position of the BCL11A enhancer region, resulting in modification of the genes of the human hematopoietic progenitor cells and a decrease in the expression of BCL11A mRNA and / or protein.
[0284] 106. The method according to embodiment 105, wherein the gRNA further comprises a spacer sequence targeting any one of SEQ ID NOs: 473, 474, 475, 476, 477, 478.
[0285] 107. For binding to a target DNA sequence, a) one or more guide RNAs capable of hybridizing to this target DNA sequence, or one or more nucleotide sequences encoding this one or more guide RNAs (gRNAs); b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54; A system in which each of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to each nucleotide sequence; hybridization of the one or more guide RNAs to the target DNA sequence; A system in which the one or more guide RNAs form a complex with the RGN polypeptide, thereby binding the RGN polypeptide to the target DNA sequence
[0286] 108. The system according to embodiment 107, wherein the RGN polypeptide is nuclease-dead or functions as a nickase.
[0287] 109. The system according to embodiment 107 or 108, wherein the RGN polypeptide is operably fused to a base editing polypeptide.
[0288] 110. The system according to embodiment 109, wherein the base editing polypeptide is a deaminase.
[0289] 111. The system according to embodiment 110, wherein the deaminase is cytidine deaminase or adenosine deaminase.
Example
[0290] The following examples are provided for illustrative purposes and not for limitation.
[0291] Experiment
[0292] Example 1: Identification of RNA-guided nucleases
[0293] Seven different CRISPR-related RNA-guided nucleases (RGNs) were identified and are listed in Table 1 below. Table 1 shows the name, amino acid sequence, origin, processed crRNA sequence and tracrRNA sequence of each RGN. Table 1 also shows the general single guide RNA (sgRNA) sequence, where the poly N in it indicates the position of the spacer sequence that determines the nucleic acid target sequence of the sgRNA. The RNG systems APG systems APG05083.1, APG07433.1, APG07513.1, APG08290.1 and APG08290.1 had the conserved sequence UNANNG (SEQ ID NO: 68) in the bases of the hairpin stem of the tracrRNA. For the AP05459.1 system, the sequence within the same position is UNANNU (SEQ ID NO: 557). For the APG04583.1 system and the APG01688.1 system, the sequence is UNANNA (SEQ ID NO: 558).
[0294]
Table 1
[0295] Example 2: Identification of guide RNAs and construction of sgRNAs
[0296] A culture of bacteria that originally expressed the RNA-guided nuclease system was grown to mid-log phase (OD600 of approximately 0.600), pelleted, and snap-frozen. RNA was isolated from the pellet using the mirVANA miRNA Isolation Kit (Life Technologies, Carlsbad, CA), and a sequencing library was prepared from the isolated RNA using the NEBNext Small RNA Library Prep Kit (NEB, Beverly, MA). This library preparation was fractionated on a 6% polyacrylamide gel into two sizes corresponding to 18 - 65 nt RNA species and 90 - 200 nt RNA species, and crRNA and tracrRNA were detected respectively. Deep sequencing (40 bp paired-end for the smaller fraction and 80 bp paired-end for the larger fraction) was performed by a service provider (MoGene, St. Louis, MO) on a Next Seq 500 (High Output kit). Reads were quality-trimmed using Cutadapt and mapped to the reference genome using Bowtie2. A custom RNAseq pipeline was written in phyton to detect crRNA transcripts and tracrRNA transcripts. The boundaries of crRNA processed by sequence coverage of the native repeat spacer array were determined. The anti-repeat portion of tracrRNA was identified using acceptable BLASTn parameters. The boundaries of processed tracrRNA were confirmed by identifying transcripts containing anti-repeats from the depth of RNA sequencing. Manual curation of the RNA was performed using secondary structure prediction by NUPACK, an RNA folding software. The sgRNA cassette was prepared by DNA synthesis. The sgRNA cassette was generally designed as follows (5'→3'): 20 - 30 bp spacer sequence → processed repeat portion of crRNA → 4 bp non-complementary linker (AAAG; SEQ ID NO: 63) → processed tracrRNA. Other 4 bp non-complementary linkers can also be used (e.g., GAAA (SEQ ID NO: 64) or ACUU (SEQ ID NO: 65)).In some cases, a 6 bp nucleotide linker can be used (e.g., CAAAGG (SEQ ID NO: 66)). For in vitro assays, sgRNAs were synthesized by in vitro transcription of the sgRNA cassette using the GeneArt™ Precision gRNA Synthesis Kit (ThermoFisher). The processed crRNA and tracrRNA sequences for each RGN polypeptide were identified and are shown in Table 1. See the following description for the sgRNAs constructed for PAM libraries 1 and 2.
[0297] Example 3: Determination of the PAM required for each RGN
[0298] The PAM required for each RGN was determined using a PAM depletion assay essentially adapted from Kleinstiver et al. (2015) Nature 523:481-485 and Zetsche et al. (2015) Cell 163:759-771. Briefly, two plasmid libraries (L1 and L2) were generated in the pUC18 backbone (ampR). Each library contains a defined 30 bp protospacer (target) sequence flanked by 8 random nucleotides (i.e., the PAM region). The target sequences and flanking PAM regions of libraries 1 and 2 for each RGN are shown in Table 2.
[0299] These libraries were separately electroporated into E. coli BL21(DE3) cells. These E. coli cells contain the RGN of the present invention (codon-optimized for E. coli) and a cognate sgRNA containing a spacer sequence corresponding to the protospacer in L1 or L2 in the pRSF-1b expression vector. Sufficient library plasmid was used in the transformation reaction to obtain 10 6Ultra cfu were obtained. Both the RGN and the sgRNA in the pRSF-1b backbone were under the control of the T7 promoter. After the transformation reaction product was collected over a period of 1 hour, it was diluted in LB medium containing carbenicillin and kanamycin and grown overnight. The next day, this mixture was diluted in self-inducing Overnight Express™ Instant TB medium (Millipore Sigma), the RGN and sgRNA were expressed, and after growing for an additional 4 hours or 20 hours, the cells were pelleted and plasmid DNA was isolated using a Mini-prep kit (Qiagen, Germantown, Maryland). In the presence of the appropriate sgRNA, plasmids containing the PAM that can be recognized by the RGN are cleaved and removed from the population. Plasmids containing a PAM that cannot be recognized by the RGN, or plasmids introduced into bacteria that do not contain the appropriate sgRNA for transformation, survive and replicate. The PAM region and the protospacer region of the non-cleaved plasmids were amplified by PCR and prepared for sequencing according to the published protocol (16s Metagenomic Library Preparation Guide 15044223B, Illumina, San Diego, California). Deep sequencing (80 bp single-end reads) was performed on the MiSeq (Illumina) by a service provider (MoGene, St. Louis, Missouri). Typically, 1 - 4M reads were obtained per amplicon. The PAM region was extracted and counted in each sample and normalized to the total reads. PAMs that lead to plasmid cleavage were identified by being less frequent compared to the control (i.e., when E. coli containing the RGN but lacking the appropriate sgRNA was transformed with the library). To represent the PAM conditions for the novel RGN, the fold change (frequency in the sample / frequency in the control) for all sequences within the region of interest was converted to an enrichment value using a -log2 transformation. A sufficient PAM was defined as a PAM with an enrichment value greater than 2.3 (corresponding to a fold change of less than approximately 0.2). PAMs greater than this threshold in both libraries were recovered and used to generate a web logo.This can be generated by using a service known as "weblogo", which is a web-based service on the Internet, for example. PAM sequences were identified and reported when there was a certain pattern among the top enriched PAMs. For each RGN (with an enrichment factor (EF) greater than 2.3), the PAMs are presented in Table 2. For some RGNs, non-limiting examples of PAMs (with an EF exceeding 3.3) were also identified. For APG005083.1, the representative PAM is NNNNCCR (SEQ ID NO: 70). For APG007513.1, the representative PAM is NNRNCC (SEQ ID NO: 71). For APG001688.1, the representative PAM is NNRANC (SEQ ID NO: 72).
[0300]
Table 2
[0301] Example 4: Determination of cleavage
[0302] The cleavage sites were clarified from in vitro cleavage reactions using ribonucleoprotein (RNP). An expression plasmid containing an RGN fused to a His6 tag or a His10 tag was constructed and used to transform the BL21(DE3) strain of Escherichia coli. Expression was carried out using auto-induction medium or IPTG induction. After lysis and clarification, the protein was purified by immobilized metal affinity chromatography.
[0303] (A ribonucleoprotein complex (including a nuclease and a double strand of sgRNA or crRNA and tracrRNA) was formed by incubating the nuclease and RNA in a buffered solution at room temperature for 20 minutes. This complex was transferred to a test tube containing a digestion buffer and a PCR-amplified target (referred to as "Sequence 1"). Sequence 1 contained, at the 3' end position, a nucleotide sequence (SEQ ID NO: 73) directly linked with the corresponding PAM sequence for each RGN. Each RGN as a ribonucleoprotein complex was incubated with each target polynucleotide at 25 °C (APG04583.1) or 37 °C (all others) for 30 minutes or 60 minutes (only for APG05459.1 and APG01688.1). After inactivating the digestion reaction by heat, it was run on an agarose gel. The bands of the cleavage products were excised from the gel and sequenced using Sanger sequencing. The cleavage sites were identified by aligning the sequencing results with the expected sequence of the PCR product. The results are shown in Table 3. As shown in Table 3, RGN APG007433.1 can also generate smooth cleavage parts with different target sequences.)
[0304] The cleavage site of Sequence 2 (SEQ ID NO: 559, functionally fused to the PAM sequence for RGN APG0733.1 at the 3' end position) was clarified by the following approach for nuclease APG07433.1. After digestion, the gel-purified DNA product was treated with a DNA end repair kit (Thermo Scientific K0771), ligated to a linearized smooth vector, and the obtained circular DNA was used to transform Escherichia coli competent cells. If there are sticky cleavage parts with 5' overhangs, it is considered that overlapping sequences in the clones from both cleavage products will be detected. 3' overhangs will result in sequence deletions, and smooth cleavage parts are considered to have all the original sequences detected without overlap. This experiment also confirmed the findings from the above method regarding Sequence 1, and it was detected that most of the clones originated from 5'-overlapping cleavages. Therefore, the smooth cleavage parts found are not considered artifacts of using this method.)
[0305]
Table 3
[0306] Example 5: Mismatch sensitivity assay
[0307] A plasmid having a target sequence (SEQ ID NO: 73) adjacent to the 5'-side of the appropriate PAM motif for the nuclease to be evaluated was designed and obtained. An array having an array in which the positions shown with one mismatch were changed was also prepared (Table 4). The purified nuclease (APG08290.1 or APG05459.1) was formed into an RNP complex with the guide RNA and incubated with the PCR-amplified linear DNA from the designed plasmid. After incubating for a determined time to inactivate the nuclease, the sample was analyzed by agarose gel electrophoresis to reveal the fraction of the remaining linear PCA product. For the mismatches at each position, the ratio of the completely cleaved bands is shown in Table 5.
[0308]
Table 4
[0309]
Table 5
[0310] A similar mismatch sensitivity experiment was conducted for RGN APG07433.1. This experiment was the same as the experiment described above, except that the alternative bases were introduced into the RNA guide instead of the DNA target. The DNA sequences for the synthesis of the sgRNA with mismatches are shown in Table 6. The results of the mismatch sensitivity assay are shown in Table 7.
[0311]
Table 6
[0312]
Table 7
[0313] RGN APG07433.1 and RGN APG08290.1 show significant sensitivity to mismatches at positions 1 to 10 from the 5' of the PAM, with some exceptions (Tables 5 and 7). RGN APG05459.1 is also sensitive to mismatches within this region, but its ability to cleave dsDNA is also greatly impaired by mismatches distant from the PAM site (Table 5). The total number of sites that have a significant impact on whether cleavage occurs is at least 15 positions within the spacer sequence. This is comparable to other genome editing tools (the well-studied Cas9 nuclease from Streptococcus pyogenes is generally sensitive to 10 - 13 base pairs) (Hsu et al., Nat Biotechnol (2013) Vol. 31(9):827 - 832). In addition, extremely important sites where there is no cleavage mediated by RGN APG05459.1 are very far from the PAM sequence, notably in the range of 13 - 20 bp. There, many other nucleases show little or only slight sensitivity to mismatches. This property may be extremely useful when targeting loci with sequence similarity to other parts of the organism of interest.
[0314] Example 6: Demonstration of gene editing activity in mammalian cells
[0315] An RGN expression cassette was prepared and introduced into a vector for expression in mammals. Each RGN was codon-optimized for expression in humans (SEQ ID NOs: 127-133), functionally linked to an SV40 nuclear localization sequence (SEQ ID NO: 134) and a 3×FLAG tag (SEQ ID NO: 135) at the 5'-end position, and functionally linked to a nuclear plasmin NLS sequence (SEQ ID NO: 136) at the 3'-end position. Each expression cassette was under the control of a cytomegalovirus (CMV) promoter (SEQ ID NO: 137). In the art, it is known that a CMV transcriptional enhancer (SEQ ID NO: 138) can also be introduced into constructs containing the CMV promoter. Guide RNA expression constructs encoding a single gRNA, each under the control of a human RNA polymerase III U6 promoter (SEQ ID NO: 139), were prepared and introduced into a pTwist High Copy Amp vector. The sequences of the target sequences for each guide are shown in Table 9.
[0316] The above constructs were introduced into mammalian cells. One day before transfection, 1×10 5 HEK293T cells / well (Sigma) were seeded in a 24-well plate in Dulbecco's Modified Eagle Medium (DMEM) + 10% (vol / vol) fetal bovine serum (Gibco) and 1% penicillin-streptomycin (Gibco). The next day, when the cells reached a confluence density of 50-60%, 1.5 μl of Lipofectamine 3000 (Thermo Scientific) per well was used according to the manufacturer's instructions, and 500 ng of the RGN expression plasmid and 500 ng of the single gRNA expression plasmid were transfected simultaneously. After growing for 48 hours, total genomic DNA was recovered using a genomic DNA isolation kit (Machery-Nagel) according to the manufacturer's instructions.
[0317] The total genomic DNA was then analyzed to determine the editing rate of each RGN for each genomic target. First, oligonucleotides for use in PCR amplification were generated, and then the amplified genomic target sites were analyzed. The sequences of the oligonucleotides used are listed in Tables 8.1-8.5.
[0318] All PCR reactions were carried out in 20 μl reactions containing 0.5 μM of each primer, using 10 μl of 2× Master Mix Phusion High-Fidelity DNA polymerase (Thermo Scientific). A large genomic region encompassing each target gene was first amplified using the PCR#1 primers. The program used then was: 98°C, 1 minute; [98°C, 10 seconds; 62°C, 15 seconds; 72°C, 5 minutes] for 30 cycles; 72°C, 5 minutes; 12°C, forever. Then 1 microliter of this PCR reaction was further amplified using primers specific to each guide (PCR#2 primers). The program used then was: 98°C, 1 minute; [98°C, 10 seconds; 67°C, 15 seconds; 72°C, 30 seconds] for 35 cycles; 72°C, 5 minutes; 12°C, forever. The primers for PCR#2 contain the Nextera Read 1 Transposase Adapter overhang sequence and the Nextera Read 2 Transposase Adapter overhang sequence for Illumina sequencing.
[0319]
Table 8-1
[0320]
Table 8-2
[0321]
Table 8-3
[0322]
Table 8-4
[0323]
Table 8-5
[0324] PCR#1 and PCR2# were performed on the purified genomic DNA as described above. After the second PCR amplification, the DNA was cleaned using a PCR clean-up kit (Zymo) according to the manufacturer's instructions and eluted in water. 200 - 500 mg of the purified PCR2# product was combined with 2 μl of 10× NEB buffer 2 and water in a 20 μl reaction and annealed to form heteroduplex DNA. The program used at that time was: 95°C, 5 minutes; cooling from 95 to 85°C at a rate of 2°C / second; cooling from 85 to 25°C at a rate of 0.1°C / second; 12°C, forever. After annealing, 5 μl of DNA was taken out as a no-enzyme control, 1 μl of T7 endonuclease I (NEB) was added, and the reaction was incubated at 37°C for 1 hour. After incubation, 5× FlashGel loading dye (Lonza) was added, and 5 μl of each reaction and control were analyzed by 2.2% agarose FlashGel (Lonza) using gel electrophoresis. After visualizing the gel, the non-homologous end joining (NHEJ) ratio was determined using the formula: %NHEJ events = 100 × [1 - (1 - ratio of cleavage) 1 / 2 , where (ratio of cleavage) is defined as (density of digested product) / (density of digested product + density of uncleaved parental band).
[0325] For several samples, the post-expression results in mammalian cells were analyzed using SURVEYOR®. After incubating the cells at 37 °C for 72 hours post-transfection, genomic DNA was extracted. Genomic DNA was extracted using QuickExtract DNA Extraction Solution (Epicentre) according to the manufacturer's protocol. The genomic region adjacent to the RGN target site was PCR amplified and the product was purified using a QiaQuick spin column (Qiagen) according to the manufacturer's protocol. A total of 200 - 500 ng of the purified PCR product was mixed with 1 μl of 10× Taq DNA polymerase PCR buffer (Enzymatics) and ultrapure water to a final volume of 10 μl, and a re-annealing process was applied to enable the formation of heteroduplexes: 10 minutes at 95 °C, from 95 °C to 85 °C at a gradient of -2 °C / second; from 85 °C to 25 °C at a gradient of -0.25 °C / second; hold at 25 °C for 1 minute.
[0326] After re-annealing, the product was treated using SURVEYOR® nuclease and SURVEYOR® enhancer S (Integrated DNA Technologies) according to the manufacturer-recommended protocol and analyzed on a 4 - 20% Novex TBE polyacrylamide gel (Life Technologies). The gel was stained with YBR Gold DNA stain (Life Technologies) for 10 minutes and images were acquired using a Gel Doc gel imaging system (Bio-rad). Quantification was performed based on relative band intensity. The indel rate was determined by the formula: 100×(1-(1-(b + c) / (a + b + c)) 1 / 2 )(where a is the integrated intensity of the undigested PCR product and b and c are the integrated intensities of the respective cleavage products).
[0327] In addition, products from PCR #2 containing the Illumina overhang sequences were used to prepare libraries according to the Illumina 16S metagenomic sequencing library protocol. Deep sequencing was performed by a service provider (MOGene) on the Illumina Mi-Seq platform. Typically, 200,000 (2 × 100,000 reads) of 250 bp paired-end reads were generated per amplicon. These reads were analyzed using CRISPResso (Pinello et al., Nature Biotech 2016, Vol. 34: 695 - 697), and the editing rates were calculated. The output alignments were performed by manual curation to confirm the insertion and deletion sites and to identify the microhomology sites at the recombination sites. The editing rates are shown in Table 9. All experiments were performed in human cells. The "target sequence" is the sequence targeted within the gene target. For each target sequence, the guide RNA contained the complementary RNA target sequence and the appropriate sgRNA depending on the RGN used. The selected breakdown of the experiments is shown in Tables 10.1 - 10.9 for each guide RNA.
[0328]
Table 9
[0329] For each guide, specific insertions and deletions are shown in Tables 10.1 - 10.7. In these tables, the target sequence is identified by bold uppercase letters. The 8-mer PAM region is underlined twice, and the major nucleotides recognized are in bold. Insertions are identified by lowercase letters. Deletions are indicated by a dotted line (---). The INDEL positions are calculated from the PAM-proximal end of the target sequence, with the edge being position 0. The position is positive (+) if on the target side of the edge and negative (-) if on the PAM side of the edge.
[0330]
Table 10-1
[0331]
Table 10-2
[0332]
Table 10-3
[0333]
Table 10-4
[0334]
Table 10-5A
Table 10-5B
[0335]
Table 10-6
[0336]
Table 10-7
[0337]
Table 10-8
[0338]
Table 10-9
[0339] Example 7: Demonstration of gene editing activity in plant cells
[0340] The activity of the RGN RNA-guided nuclease according to the present invention is demonstrated in plant cells using a protocol modified from Li et al. (2013) Nat. Biotech. Vol. 31: pages 688 - 691. Briefly, plant codon-optimized versions of each RGN (SEQ ID NOs: 169 - 182) containing an N-terminal SV40 nuclear localization signal are inserted into a transient transformation vector and cloned behind a strong constitutive 35S promoter. An sgRNA targeting one or more sites within the plant PDS gene adjacent to an appropriate PAM sequence is inserted into a second transient transformation vector and cloned behind a plant U6 promoter. These expression vectors are introduced into Nicotiana benthamiana mesophyll protoplasts using PEG-mediated transformation. The transformed protoplasts are incubated in the dark for up to 36 hours. Genomic DNA is isolated from the protoplasts using the DNeasy Plant Mini Kit (Qiagen). The genomic region adjacent to the RGN target site is amplified by PCR, and the product is purified using a QiaQuick spin column (Qiagen) according to the manufacturer's protocol. A total of 200 - 500 ng of the purified PCR product is mixed with 1 μl of 10× Taq DNA polymerase PCR buffer (Enzymatics) and ultrapure water to a final volume of 10 μl, and a re-annealing process is applied to enable the formation of heteroduplexes: 10 minutes at 95°C, from 95°C to 85°C at a gradient of -2°C / second; from 85°C to 25°C at a gradient of -0.25°C / second; hold at 25°C for 1 minute. After re-annealing, the product is treated with SURVEYOR® nuclease and SURVEYOR® enhancer S (Integrated DNA Technologies) according to the manufacturer's recommended protocol and analyzed on a 4 - 20% Novex TBE polyacrylamide gel (Life Technologies). The gel is stained with YBR Gold DNA stain (Life Technologies) for 10 minutes, and images are acquired using a Gel Doc gel imaging system (Bio-rad). Quantification is performed based on relative band intensity. The indel rate is calculated using the formula: 100×(1-(1-(b + c) / (a + b + c)) 1 / 2) (where a is the integrated intensity of the undigested PCR product, and b and c are the integrated intensities of the respective cleavage products).
[0341] Alternatively, it is possible to perform the same PCR as described in Example 6 on a PCR product derived from a target genomic sequence. As a result, the PCR product contains an Illumina overhang sequence, enabling the preparation of a library and deep sequencing. By this method, as shown in Table 9, it becomes possible to determine the editing rate.
[0342] Example 8: Guide cross-compatibility
[0343] To examine the cross-compatibility of guide RNAs between RGNs, a two-plasmid interference experiment was performed (Esvelt et al. (2013), Nat. Methods Vol. 10(11):1116 - 1121). The first plasmid contained RGNs with several targets containing PAMs defined on a kanamycin resistance backbone. Escherichia coli BL21 was transformed with these plasmids, and the transformed strains were made chemically competent. Subsequently, a second plasmid containing the guide RNA on an ampicillin resistance backbone was introduced. Cells were seeded on a medium containing both antibiotics. If the RGN can utilize the guide on the second plasmid, the kanamycin resistance plasmid is cleaved and linearized, resulting in few or no colonies formed. If the RGN cannot utilize the guide on the second plasmid, the kanamycin resistance plasmid is not cleaved, resulting in high-level colony formation. Guide RNAs for Streptococcus pyogenes Cas9 (SpyCas9) and Staphylococcus aureus Cas9 (SauCas9) were also included, and cross-compatibility was examined using these guide RNAs.
[0344] To calculate the reduction probability, the number of colonies regarding transformation at each guide is compared with the transformation efficiency using a positive control. Based on this comparison, if the RGN can use the guide, no colonies should survive, so the reduction probability should be 0. If the RGN cannot use the guide, all plasmids should remain intact, so the reduction probability should be 1. The results are shown in Table 11 below. "sg" indicates the guide RNA for the described RGN.
[0345]
Table 11
[0346] As can be seen from Table 11, there are four groups of orthogonal systems. The RGN can recognize guides from other systems within its own group but cannot utilize guides from other groups. The first group includes APG05083.1, APG07433.1, APG07513.1, and APG08290.1. The second group includes SpyCas9 and APG05459.1. The third group includes APG04583.1 and APG01688.1. The fourth group includes SauCas9.
[0347] Example 9: Identification of disease targets
[0348] The database of clinical variants was obtained from the NCBI ClinVar database, which is available through the NCBI ClinVar website on the world wide web. Pathogenic single nucleotide polymorphisms (SNPs) were identified from this list. Using the genomic locus information, CRISPR targets in the regions overlapping with each SNP and in the regions surrounding each SNP were identified. The selection of SNPs that can be corrected using base editing in combination with the RGNs of the present invention targeting causative mutations is listed in Table 12. Only one common name for each disease is listed in Table 12. "RS#" corresponds to the RS accession number in the SNP database on the NCBI website. The allele ID corresponds to the accession number of the causative allele, and the chromosome accession number also provides accession reference information that can be viewed through the NCBI website. Table 12 also presents information on the genomic target sequences that match the listed RGNs for each disease. The information on the target sequences also provides the protospacer sequences for generating the sgRNAs required for the corresponding RGNs of the present invention.
[0349] Table 12: Diseases targeted by the RGNs of the present invention
Table 12-1
Table 12-2
Table 12-3
Table 12-4
Table 12-5
Table 12-6
Table 12-7
Table 12-8
[0350] Example 10: Targeting mutations responsible for Hurler syndrome
[0351] The following describes a treatment using an RNA-guided base editing system that corrects mutations responsible for Hurler syndrome in a large population of patients with this disease as a potential treatment for Hurler syndrome (also called MPS-1). This approach utilizes a base editing fusion protein that can be packaged into a single AAV vector that is RNA-guided for delivery to a wide range of tissue types. Depending on the exact regulatory elements and base editing domain used, it may also be possible to generate a single vector encoding both a base editing fusion protein and a single guide RNA that targets the diseased locus.
[0352] Example 10.1: Identification of RGNs with Ideal PAMs
[0353] The genetic disorder MPS-1 is a lysosomal storage disorder characterized by the accumulation of dermatan sulfate and heparan sulfate within lysosomes at the molecular level. This disorder is generally a hereditary genetic disorder caused by mutations in the IDUA gene (NCBI reference sequence NG_008103.1) that encodes α-L-iduronidase. This disorder is the result of a deficiency of α-L-iduronidase. The most common IDUA mutations found in studies of individuals of Nordic descent are W402X and Q70X. Both are nonsense mutations, resulting in premature termination of translation (Bunge et al. (1994), Hum. Mol. Genet, Vol. 3(6):861-866; incorporated herein by reference). Restoring one nucleotide is thought to restore the wild-type coding sequence, and as a result, protein expression is controlled by the endogenous regulatory mechanism of the locus.
[0354] The W402X mutation of the human Idua gene accounts for a large proportion of cases of MPS-1H. Since base editors can target a sequence window narrower than the binding site of the protospacer component of the guide RNA, the presence of a PAM sequence at a specific distance from the target locus is essential for the success of this strategy. The target mutation must be present on the non-target strand (NTS) that is exposed during the interaction with the base editing protein, and there are constraints that the footprint of the RGN domain blocks access to the region near the PAM, so the accessible locus is thought to be at a position 10-30 bp from the PAM. To avoid editing and mutagenesis of another adenosine base in the vicinity within this window, various linkers are screened. The ideal window is 12-16 bp from the PAM.
[0355] The positions of the loci and the PAM sequences that match APG07433.1 and APG08290.1 within the above-mentioned ideal base editing window can be immediately identified. These nucleases each have PAM sequences of NNNNCC (SEQ ID NO: 6) and NNRNCC (SEQ ID NO: 32), respectively, and due to their compact size, they may be deliverable through a single AAV vector. This delivery approach offers many advantages compared to other methods (such as access to a wide range of tissues (liver, muscle, CNS), well-established safety profiles and manufacturing technologies, etc.).
[0356] Cas9 (SpyCas9) from Streptococcus pyogenes requires a PAM sequence of NGG (SEQ ID NO: 448), and this PAM sequence is present near the W402X locus. However, due to the size of SpyCas9, the gene encoding the fusion protein of the base editing domain and the SpyCas9 nuclease cannot be packaged into a single AAV vector, so the above advantages of this approach are lost. Including the sequence encoding the guide RNA in this vector is considered to be even less feasible, even if there are significant technical improvements to reduce the size of the gene regulatory elements or increase the packaging limits of the AAV vector. A dual delivery strategy can be utilized (e.g., Ryu et al. (2018), Nat. Biotechnol., Vol. 36(6):536 - 539; incorporated herein by reference), but it will be extremely complex and costly to produce. In addition, dual viral vector delivery significantly reduces the efficiency of gene correction. This is because successful editing within a given cell requires infecting both vectors and assembling the fusion protein within the cell.
[0357] The commonly used Cas9 ortholog (SauCas9) from Staphylococcus aureus is considerably smaller in size compared to SpyCas9, but the required PAM is more complex (NGRRT (SEQ ID NO: 449)). However, this sequence is not within the range expected to be useful for base editing of the causative locus.
[0358] Example 10.2: RGN Fusion Construct and sgRNA Sequence
[0359] 1) A DNA sequence encoding a fusion protein having an RGN domain with a mutation that inactivates DNA cleavage activity (a "dead" or "nickase") and; 2) an adenosine deaminase useful for base editing is prepared using standard molecular biology techniques. All constructs described in the following table contain a fusion protein having a base editing activity domain (in this example, ADAT (SEQ ID NO: 450) functionally fused to the N-terminus of RGN APG08290.1). It is known in the art that a fusion protein having a base editing enzyme at the C-terminus of RGN can also be prepared. In addition, the RGN and base editor of the fusion protein are typically separated by a linker amino sequence. It is known in the art that the standard linker length ranges from 15 to 30 amino acids. Further, in the art, some fusion proteins can also include at least one uracil glycosylase inhibitor (UGI) domain capable of improving base editing efficiency between the RGN and the base editing enzyme (e.g., cytidine deaminase) (U.S. Patent No. 10,167,457; incorporated herein by reference). Thus, the fusion protein can include APG08290.1, a base modifying enzyme, and at least one UGI.
[0360] [Table 13]
[0361] The editable site accessible by the RGN is determined by the PAM sequence. When combining the RGN with a base editing domain, the residue to be edited must be present on the non-target strand (NTS). This is because the NTS is single-stranded while the RGN is associated with the locus. By evaluating a number of nucleases and corresponding guide RNAs, it becomes possible to select the optimal gene editing tool for this specific locus. Some of the PAM sequences in the human Idua gene that may be targeted by the above construct are in the vicinity of the mutant nucleotide responsible for the W402X mutation. Also prepare 1) a "spacer" complementary to the non-coding DNA strand at the position of the disease locus, and 2) a sequence encoding a guide RNA transcript containing the RNA sequence necessary to associate the guide RNA with the RGN. Useful guide RNA sequences (sgRNAs) are shown in Table 14 below. The efficiency with which these guide RNA sequences direct the above base editor to the target locus can be evaluated.
[0362]
Table 14
[0363] Example 10.3: Assay to Examine Activity in Cells Derived from Patients with Hurler Disease
[0364] To confirm the genotypic strategy and evaluate the above constructs, fibroblasts derived from patients with Hurler disease are used. Similar to the vectors described in Example 5, a vector containing an appropriate promoter upstream of the fusion protein coding sequence and an sgRNA coding sequence for expressing them in human cells is designed. Promoters and other DNA elements (such as enhancers or terminators) that are known to be highly expressed in human cells or can be specifically well-expressed in fibroblasts can also be used. Standard techniques (such as transfection similar to that described in Example 6) are utilized to transfect the vector into fibroblasts. Alternatively, electroporation can be used. The cells are cultured for 1 to 3 days. Genomic DNA (gDNA) is isolated using standard techniques. The editing efficiency is determined by qPCR genotyping assay and / or next-generation sequencing for the purified gDNA. This will be described in more detail below.
[0365] In Taqman™ qPCR analysis, a probe specific for the wild-type allele and a probe specific for the mutant allele are used. These probes carry fluorophores, and these fluorophores are separated by spectral excitation characteristics and / or emission characteristics using a qPCR instrument. Genotyping kits containing primers and probes for PCR can be obtained commercially (i.e., Fisher Taqman™ SNP Genotyping Assay ID C__27862753_10 for Thermo SNP ID rs121965019) or designed. An example of a set of designed primers and probes is shown in Table 15.
[0366]
Table 15
[0367] After the editing experiment, the gDNA is analyzed by qPCR using standard methods and the above-mentioned primers and probes. The expected results are shown in Table 16. This in vitro system can be used to appropriately evaluate the constructs and select constructs with high editing efficiency for further study. This system will be evaluated compared to cells with and without the W402X mutation, and preferably compared to several cells that are heterozygous for this mutation. The Ct value will be compared to the total amplification of the reference gene or its locus using a dye (such as Sybr green).
[0368]
Table 16
[0369] Tissues can also be analyzed by next-generation sequencing. Primer binding sites as shown below (Table 17), or other appropriate primer binding sites identifiable by those skilled in the art, can be used. After PCR amplification, a library is prepared from the product containing the Illumina Nextera XT overhang sequence according to the protocol for the Illumina 16S metagenomic sequencing library. Deep sequencing is performed on the Illumina Mi-Seq platform. Typically, 200,000 250 bp paired-end reads (2 × 100,000 reads) per amplicon are generated. CRISPResso (Pinello et al., 2016) is utilized to analyze these reads and calculate the editing rate. Manual curation of the output alignment is performed to confirm the insertion and deletion sites and to identify the microhomology sites at the recombination sites.
[0370]
Table 17
[0371] The expression of the full-length protein is confirmed by performing Western blotting of cell lysates of the transfected cells and control cells using an anti-IDUA antibody, and the enzyme is confirmed to be catalytically active by an enzyme activity assay on the cell lysates using the substrate 4-methylumbelliferyl a-L-iduronide (Hopwood et al., Clin. Chim. Acta (1979), Vol. 92(2): 257-265; incorporated herein by reference). These experiments are performed on the original Idua W402X / W402X cell line (without transfection), the Idua W402X / W402X cell line transfected with the base editing construct and a random guide sequence, and compared with the cell line expressing wild-type IDUA.
[0372] Example 10.4: Verification of Disease Treatment in a Mouse Model
[0373] To verify the effect of this treatment approach, a mouse model with a nonsense mutation in a similar amino acid is used. This mouse strain has a W392X mutation in the Idua gene (Gene ID: 15932) corresponding to the mutation in patients with Hurler syndrome (Bunge et al. (1994), Hum. Mol. Genet. Vol. 3(6): 861-866; incorporated herein by reference). This locus contains a nucleotide sequence that is clearly different from that in humans and lacks the PAM sequence required for correction using the base editor described in the previous example, so it is necessary to design a different fusion protein to correct the nucleotide. Improvement of the disease in this animal can confirm the effectiveness of this treatment approach for correcting mutations within tissues accessible by the gene delivery vector.
[0374] Mice that are homozygous for this mutation exhibit numerous phenotypic characteristics similar to those of patients with Hurler syndrome. The above base-editing RGN fusion protein (Table 13) and RNA guide sequence are incorporated into an expression vector that enables protein expression and RNA transcription in mice. The study design is shown in Table 18 below. This study includes a group treated with a high-dose expression vector containing the base-editing fusion protein and RNA guide sequence, a control that is a model mouse treated with an expression vector that does not contain the base-editing fusion protein or RNA guide sequence, and a second control that is a wild-type mouse treated with the same empty vector.
[0375]
Table 18
[0376] The endpoints to be evaluated include body weight, urinary GAG excretion, serum IDUA enzyme activity, IDUA activity in the target tissue, tissue pathology, genotype of the target tissue to confirm SNP modification, behavioral assessment, and neurological assessment. Since some endpoints are lost, additional groups can be added before the study ends to evaluate, for example, tissue pathology and tissue IDUA activity. Another example of endpoints can be found in published papers that established Hurler syndrome animal models (Shull et al. (1994), Proc. Natl. Acad. Sci. U.S.A., Vol. 91 (26): 12937-12941; Wang et al. (2010), Mol. Genet. Metab., Vol. 99 (1): 62-71; Hartung et al. (2004), Mol. Ther., Vol. 9 (6): 866-875; Liu et al. (2005), Mol. Ther., Vol. 11 (1): 35-47; Clark et al. (1997), Hum. Mol. Genet. Vol. 6 (4): 503-511; all of which are incorporated herein by reference).
[0377] One possible delivery vector utilizes adeno-associated virus (AAV). What is included when constructing the vector is an array encoding a base editor-dRGN fusion protein (e.g., SEQ ID NO: 452), a combination of the CMV enhancer (SEQ ID NO: 138) and promoter (SEQ ID NO: 137) in front of it, or another suitable combination of enhancer and promoter, and optionally a Kozak sequence, a terminator sequence and a polyadenylation sequence (e.g., the minimal sequence described in Levitt, N.; Briggs, D.; Gil, A.; Proudfoot, N. J. "Definition of an efficient synthetic poly(A) site", Genes Dev., Vol. 3(7), pp. 1019-1025, 1989) that are operably linked at the 3'-end position. This vector further includes an expression cassette encoding a single guide RNA that is operably linked to a human U6 promoter (SEQ ID NO: 139) or another promoter suitable for generating small non-coding RNAs at the 5'-end position, and further includes inverted terminal repeat (ITR) sequences well known in the art necessary for packaging into the AAV capsid. The construction of the vector and the packaging of the virus are carried out by standard methods (e.g., the methods described in U.S. Patent No. 9,587,250; incorporated herein by reference).
[0378] Other possible viral vectors include adenoviral vectors and lentiviral vectors (commonly used and thought to contain similar elements) with various packaging capabilities and conditions. Non-viral delivery methods can also be utilized, examples of which are mRNA and sgRNA encapsulated by lipid nanoparticles (Cullis, P. R. and Allen, T. M. (2013), Adv. Drug Deliv. Rev. Vol. 65(1):36 - 48; Finn et al. (2018), Cell Rep. Vol. 22(9):2227 - 2235; both incorporated herein by reference), or hydrodynamic injection of plasmid DNA (Suda T and Liu D, (2007) Mol. Ther. Vol. 15(12):2063 - 2069; incorporated herein by reference), or ribonucleoprotein complexes of sgRNA associated with gold nanoparticles (Lee, K.; Conboy, M.; Park, H. M.; Jiang, F.; Kim, H. J.; Dewitt, M. A.; Mackley, V. A.; Chang, K.; Rao, A.; Skinner, C.; et al., "In vivo nanoparticle delivery of Cas9 ribonucleoprotein and donor DNA induces homology-directed DNA repair" Nat. Biomed. Eng. 2017, Vol. 1(11), 889 - 890).
[0379] Example 10.5: Correction of Diseases in Mouse Models with Humanized Loci
[0380] To evaluate the effect of the same base editor construct that is considered usable for human treatment, a mouse model with nucleotides near W392 changed to match the sequence around human W402 is required. This can be achieved by a variety of techniques, which include using RGNs and HDR templates to cut and replace the locus in the mouse embryo.
[0381] Because the degree of amino acid conservation is high, as shown in Table 19, most of the nucleotides in the mouse locus can be changed to the nucleotides of the human sequence with silent mutations. Only the base changes that change the coding sequence in the obtained engineered mouse genome occur after the introduced stop codon.
[0382]
Table 19
[0383] When the manipulation of this mouse strain is completed, a similar experiment is carried out as described in Example 10.4.
[0384] Example 11: Targeting mutations responsible for Friedreich's ataxia
[0385] The expansion of the trinucleotide repeat sequence that causes Friedreich's ataxia (FRDA) occurs at a specific locus (referred to as the FRDA instability region) within the FXN gene. The instability region within the cells of FRDA patients can be excised using an RNA-guided nuclease (RGN). This approach requires 1) an RGN-guide RNA sequence that can be programmed to target an allele within the human genome; and 2) a method for delivering this RGN-guide RNA sequence. Many nucleases used for genome editing (such as the Cas9 nuclease (SpCas9) from the commonly used Streptococcus pyogenes) are too large to be packaged within an adeno-associated virus (AAV) vector. This can be seen especially when considering the length of the SpCas9 gene and the guide RNA, as well as the length of other genetic elements required for a functional expression cassette. Therefore, the possibility of an approach using SpCas9 is small.
[0386] The compact RNA-guided nucleases of the present invention (in particular APG07433.1 and APG08290.1) are very well suited for the excision of the FRDA instability region. Each RGN requires a PAM in the vicinity of the FRDA instability region. In addition, each of these RGNs can be packaged into an AAV vector together with a guide RNA. Although a second vector is likely required to package two guide RNAs, this approach is still advantageous over the vectors that may be required for larger nucleases (such as SpCas9, where the protein sequence may have to be split into two vectors).
[0387] Table 20 shows the positions of genomic target sequences suitable for directing APG07433.1 or APG08290.1 towards the 5' and 3' flanks of the FRDA instability region. When the RGNs come to the positions of this locus, they are thought to excise the FA instability region. The excision of this region can be confirmed by Illumina sequencing of this locus.
[0388]
Table 20
[0389] Example 12: Targeting mutations responsible for sickle cell disease
[0390] The target sequence (SEQ ID NO: 472) within the BCL11 enhancer region can provide a mechanism for increasing fetal hemoglobin (HbF), thereby curing or alleviating the symptoms of sickle cell disease. For example, genome-wide association studies have identified a group of genetic variations in BCL11A that are associated with increased HbF levels. These variations are a set of SNPs found within the non-coding region of BCL11A that functions as a stage-specific and lineage-restricted enhancer region. Further exploration has revealed that this BCL11A enhancer is required for expressing BCL11A in erythroid cells (Bauer et al. (2013) Science 343:253-257; incorporated herein by reference). This enhancer region is found within intron 2 of the BCL11A gene, and within intron 2, three DNase I hypersensitive regions (which often indicate the chromatin state related to regulatory capacity) have been identified. These three regions were identified as "+62", "+58", and "+55" according to the distance (in kilobases) from the transcription start site of BCL11A. These enhancer regions are approximately 350 nucleotides in length (+55); 550 nucleotides (+58); and 350 nucleotides (+62) (Bauer et al., 2013).
[0391] Example 12.1: Identifying a Preferred RGN System
[0392] Here, we describe the potential for treating β-hemoglobinopathy using an RGN system that blocks the binding of BCL11A to its binding site within the HBB locus (the gene responsible for making β-globin in adult hemoglobin). This approach utilizes more efficient NHEJ in mammalian cells. In addition, this approach utilizes a nuclease of a sufficiently small size that can be packaged within a single AAV vector for in vivo delivery.
[0393] The GATA1 enhancer motif (SEQ ID NO: 472) within the human BCL11 enhancer region is an ideal target for disruption using an RNA-guided nuclease (RGN), and disruption thereof results in decreased expression of BCL11A and simultaneous re-expression of HbF in adult human erythrocytes (Wu et al. (2019) Nat Med 387:2554). Several PAM sequences compatible with APG07433.1 and APG08290.1 can be readily identified at the locus surrounding this GATA1 site. These nucleases have a PAM sequence of 5'-NNNNCC-3' (SEQ ID NO: 6) and are of compact size, and thus may be delivered within a single AAV vector or adenovirus vector together with an appropriate guide RNA. This delivery approach confers a number of advantages (e.g., access to hematopoietic stem cells, well-established safety profiles and manufacturing techniques) compared to other approaches.
[0394] The commonly used Cas9 nuclease (SpyCas9) from Streptococcus pyogenes requires a PAM sequence of 5'-NGG-3' (SEQ ID NO: 448), some of which are present near the GATA1 motif. However, due to the size of SpyCas9, it cannot be packaged into a single AAV vector or adenovirus vector, eliminating the above advantages of this approach. A dual delivery strategy can be employed, but it is considered extremely complex and costly to produce. In addition, dual viral vector delivery significantly reduces the efficiency of gene modification because both vectors need to infect for editing to succeed in a given cell.
[0395] In the same manner as described in Example 6, an expression cassette encoding APG07433.1 (SEQ ID NO: 128) or APG08290.1 (SEQ ID NO: 130) with human codons optimized is prepared. An expression cassette for expressing guide RNAs for RGN APG07433.1 and APG08290.1 is also prepared. These guide RNAs contain 1) a protospacer sequence complementary to the non-coding DNA strand or the coding DNA strand within the BCL11A enhancer locus (target sequence), and 2) an RNA sequence necessary to associate the guide RNA with RGN (SEQ ID NO: 18 for APG07433.1 and SEQ ID NO: 35 for APG08290.1). Since several PAM sequences that APG07433.1 or APG08290.1 may target surround the BCL11A GATA1 enhancer motif, several possible guide RNA constructs are prepared to determine the best protospacer sequence that causes robust cleavage of the BCL11A GATA1 enhancer sequence and NHEJ-mediated disruption. To direct the RGN to this locus, the target genomic sequences in the following table (Table 21) are evaluated.
[0396]
Table 21
[0397] To evaluate the efficiency of APG07433.1 or APG08290.1 to cause insertions or deletions that disrupt the BCL11A enhancer region, a human cell line (such as human embryonic kidney cells (HEK cells)) is used. A DNA vector containing an RGN expression cassette (such as those described in Example 6) is prepared. Another vector containing an expression cassette encoding the guide RNA sequences of Table 21 is also prepared. Such an expression cassette can further contain a human RNA polymerase III U6 promoter (SEQ ID NO: 139) as described in Example 6. Alternatively, a single vector containing both the RGN expression cassette and the guide RNA expression cassette can be used. The vector is introduced into HEK cells using standard techniques (such as those described in Example 6), and the cells are cultured for 1 to 3 days. After this culture period, genomic DNA is isolated, and the frequency of insertions or deletions is determined using digestion with T7 endonuclease I and / or direct sequencing as described in Example 6.
[0398] The DNA region encompassing the target BCL11A is amplified by PCR using primers containing the Illumina Nextera XT overhang sequences. The formation of NHEJ is examined using digestion with T7 endonuclease I on these PCR amplicons, or library preparation according to the protocol for Illumina 16S metagenomic sequencing libraries, or preparation of a similar next-generation sequencing (NGS) library is performed. After deep sequencing, the generated reads are analyzed by CRISPResso to calculate the editing rate. Manual curation of the output alignment is performed to confirm the insertion and deletion sites. This analysis identifies the preferred RGN and the corresponding preferred guide RNA(s) (sgRNA). This analysis may result in equally favorable results for both APG07433.1 and APG08290.1. In addition, this analysis may determine that there are two or more preferred guide RNAs, or that all of the target genomic sequences in Table 21 are equally preferred.
[0399] Example 12.2: Assay for Examining the Expression of Fetal Hemoglobin
[0400] In this example, the expression of fetal hemoglobin is examined when an insertion or deletion is generated that disrupts the BCL11A enhancer region by APG07433.1 or APG08290.1. Healthy human donor CD34 + hematopoietic stem cells (HSCs) are used. These HSCs are cultured using a method similar to that described in Example 11.1, and a vector containing an expression cassette containing a region encoding a preferred RGN and an expression cassette containing a region encoding a preferred sgRNA is introduced. After electroporation, these cells are differentiated into erythrocytes in vitro using an established protocol (e.g., Giarratana et al. (2004) Nat Biotechnology Vol. 23: pages 69 - 74; incorporated herein by reference). Subsequently, the expression of HbF is measured using Western blotting with an anti - human HbF antibody or quantified through high - performance liquid chromatography (HPLC). If the disruption of the BCL11A enhancer locus is successful, an increase in HbF production is expected compared to HSCs electroporated with only RGN without a guide.
[0401] Example 12.3: Assay for Examining the Reduction of Sickle Erythropoiesis
[0402] In this example, the reduction of sickle erythropoiesis is examined when an insertion or deletion is generated that disrupts the BCL11A enhancer region by APG07433.1 or APG08290.1. Donor CD34 from a patient suffering from sickle cell disease +Hematopoietic stem cells (HSCs) are used. These HSCs are cultured using the same method as described in Example 11.1, and a vector containing an expression cassette comprising a region encoding a preferred RGN and an expression cassette comprising a region encoding a preferred sgRNA is introduced. After electroporation, these cells are differentiated into red blood cells in vitro using an established protocol (e.g., Giarratana et al. (2004) Nat Biotechnology Vol. 23: pages 69 - 74). Then, the expression of HbF is measured using Western blotting with an anti - human HbF antibody or quantified through high - performance liquid chromatography (HPLC). Successful disruption of the BCL11A enhancer locus is expected to increase HbF production compared to HSCs electroporated with only RGN without a guide.
[0403] Induction of sickle red blood cell formation is achieved by adding metabisulfite to these differentiated red blood cells. The number of sickle red blood cells and normal red blood cells is counted using a microscope. The number of sickle red blood cells is expected to be less in cells treated with APG07433.1 or APG08290.1 and sgRNA compared to untreated cells or cells treated with only RGN.
[0404] Example 12.4: Verification of disease treatment in a mouse model
[0405] To evaluate the effect of disruption of the BCL11A locus using APG07433.1 or APG08290.1, an appropriate humanized mouse model of sickle cell anemia is used. Expression cassettes encoding the preferred RGN and the preferred sgRNA are packaged into an AAV vector or an adenovirus vector. In particular, the adenoviral type Ad5 / 35 is effective in targeting HSCs. Select an appropriate mouse model (e.g., B6;FVB-Tg(LCR-HBA2,LCR-HBB*E26K)53Hhb / J or B6.Cg-Hbatm1Paz Hbbtm1Tow Tg(HBA-HBBs)41Paz / HhbJ) containing a humanized HBB locus with the sickle cell allele. These mice are treated with granulocyte colony-stimulating factor alone or in combination with plerixafor to mobilize HBCs into the circulation. Next, AAV or adenovirus carrying the RGN and the guide plasmid is injected intravenously, and the mice are allowed to recover for one week. Blood obtained from these mice is examined in an in vitro sickle cell assay using sodium metabisulfite, and the mice are followed over the long term to monitor mortality and hematopoietic function. Treatment with AAV or adenovirus carrying the RGN and the guide RNA reduces sickle cells, decreases mortality, and improves hematopoietic function compared to mice treated with a virus lacking both expression cassettes or a virus carrying only the RGN expression cassette.
[0406] Claim 1 A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide comprises a nucleotide sequence encoding an RGN polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, wherein the RGN polypeptide binds to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence and binds to the target DNA sequence in a manner specific to the RNA-guided sequence when bound to the gRNA, and the polynucleotide encoding the RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide. A nucleic acid molecule. Claim 2 The nucleic acid molecule according to claim 1, wherein the RGN polypeptide is nuclease-dead or functions as a nickase. Claim 3 The nucleic acid molecule according to claim 2, wherein the RGN polypeptide is operably fused to a base editing polypeptide. Claim 4 A vector comprising the nucleic acid molecule according to any one of claims 1 to 3. Claim 5 Further comprising at least one nucleotide sequence encoding the guide RNA, wherein the guide RNA comprises a CRISPR RNA comprising a CRISPR repeat sequence having at least 95% identity with any one of SEQ ID NOs: 2, 12, 20, 28, 37, 46, 55. The vector according to claim 4. Claim 6 The vector according to claim 4 or 5, wherein the guide RNA comprises a tracrRNA having at least 95% identity with any one of SEQ ID NOs: 3, 13, 21, 29, 38, 47, 56. Claim 7 A cell comprising the nucleic acid molecule according to any one of claims 1 to 3 or the vector according to any one of claims 4 to 6. Claim 8 A system for binding to a target DNA sequence, the system comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); b) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54, or a nucleotide sequence encoding the RGN polypeptide; comprising wherein each of the nucleotide sequence encoding the one or more guide RNAs and the nucleotide sequence encoding the RGN polypeptide is operably linked to a promoter heterologous to each nucleotide sequence; wherein the one or more guide RNAs hybridize to the target DNA sequence, and wherein the one or more guide RNAs form a complex with the RGN polypeptide, thereby causing the RGN polypeptide to bind to the target DNA sequence. Claim 9 The system according to claim 8, wherein the target DNA sequence is in a eukaryotic cell. Claim 10 The system according to claim 8 or 9, wherein the RGN polypeptide is nuclease-dead or functions as a nickase, and the RGN polypeptide is operably linked to a base editing polynucleotide. Claim 11 The system according to any one of claims 8 to 10, further comprising one or more donor polynucleotides, or one or more nucleotide sequences encoding the one or more donor polynucleotides, wherein each of the nucleotide sequences encoding the one or more donor polynucleotides is operably linked to a promoter heterologous to each nucleotide sequence. Claim 12 A method for binding to a target DNA sequence, the method comprising delivering the system according to any one of claims 8 to 11 to the target DNA sequence or to a cell comprising the target DNA sequence. Claim 13 A method for cleaving and / or modifying a target DNA sequence, the method comprising contacting the target DNA sequence with a) an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence having at least 95% identity with any one of SEQ ID NOs: 1, 11, 19, 27, 36, 45, 54 b) contacting the RGN of (a) with one or more guide RNAs capable of directing the RGN to the target DNA sequence, wherein the one or more guide RNAs hybridize to the target DNA sequence, causing the RGN polypeptide to bind to the target DNA sequence and effecting cleavage and / or modification of the target DNA sequence. Claim 14 The method according to claim 13, wherein the modified target DNA sequence comprises insertion of a heterologous DNA into the target DNA sequence. Claim 15 The method according to claim 13, wherein the modified target DNA sequence comprises deletion of at least one nucleotide from the target DNA sequence. Claim 16 The method according to claim 13, wherein the modified target DNA sequence comprises mutation of at least one nucleotide in the target DNA sequence. Claim 17 The method according to any one of claims 14 to 16, wherein the target DNA sequence is in a cell. Claim 18 The method according to claim 17, wherein the cell is a eukaryotic cell. Claim 19 The method according to claim 17 or 18, further comprising culturing the cell under conditions in which the RGN polypeptide is expressed to cleave the target DNA sequence and generate a modified DNA sequence; and selecting a cell comprising the modified DNA sequence. Claim 20 A cell comprising a target DNA sequence modified according to the method of claim 19.
Claims
**Claim 1** A nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the polynucleotide (i) has at least 90% sequence identity with SEQ ID NO: 11; or (ii) has at least 98% sequence identity with SEQ ID NO: 1 or 19, and comprises a nucleotide sequence encoding an RGN polypeptide comprising the amino acid sequence. **Claim 2** The nucleic acid molecule according to claim 1, wherein when the RGN polypeptide is bound to a guide RNA (gRNA) capable of hybridizing to a target DNA sequence, it binds to the target DNA sequence in a manner specific to the RNA-guided sequence. **Claim 3** The nucleic acid molecule according to claim 1 or 2, wherein the polynucleotide encoding the RGN polypeptide is operably linked to a promoter heterologous to the polynucleotide. **Claim 4** The nucleic acid molecule according to claim 1 or 2, wherein the nucleic acid molecule is an RNA polynucleotide. **Claim 5** The nucleic acid molecule according to claim 4, wherein the RNA polynucleotide is mRNA. **Claim 6** The nucleic acid molecule according to any one of claims 1 to 5, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:
11. **Claim 7** The nucleic acid molecule according to any one of claims 1 to 6, wherein the RGN polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 11, 1, or 19. **Claim 8** The nucleic acid molecule according to any one of claims 1 to 6, wherein the RGN polypeptide is nuclease-dead or functions as a nickase. **Claim 9** The nucleic acid molecule according to any one of claims 1 to 8, wherein the RGN polypeptide is operably fused to a base editing polypeptide. **Claim 10** The nucleic acid molecule according to any one of claims 1 to 8, wherein the RGN polypeptide is operably fused to a nuclear localization signal, a plastid localization signal, a mitochondrial localization signal, a dual-target localization signal, and / or a cell-penetrating domain. **Claim 11** The nucleic acid molecule according to claim 10, wherein the nuclear localization signal, the plastid localization signal, the mitochondrial localization signal, the dual-target localization signal, and / or the cell-penetrating domain is operably fused to the N-terminus, C-terminus, or internal position of the RGN polypeptide. **Claim 12** The nucleic acid molecule according to any one of claims 1 to 8, wherein the RGN polypeptide is functionally fused to an effector domain.
13. The nucleic acid molecule according to claim 12, wherein the effector domain is a cleavage domain, a deaminase domain, or an expression modulator domain.
14. The effector domain is functionally fused to the N-terminus, C-terminus, or internal position of the RGN polypeptide. The nucleic acid molecule according to claim 12 or 13.
15. A vector comprising the nucleic acid molecule according to any one of claims 1 to 14.
16. Further comprising at least one nucleotide sequence encoding a guide RNA, wherein the guide RNA is (a) a nucleotide sequence having at least 95% identity with SEQ ID NO: 12, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% identity with SEQ ID NO: 11; (b) a nucleotide sequence having at least 95% identity with SEQ ID NO: 2, wherein the RGN polypeptide comprises an amino acid sequence having at least 98% identity with SEQ ID NO: 1; or (c) a nucleotide sequence having at least 95% identity with SEQ ID NO: 20, wherein the RGN polypeptide comprises an amino acid sequence having at least 98% identity with SEQ ID NO: 19 and comprises a CRISPR RNA comprising a CRISPR repeat sequence. The vector according to claim 15.
17. The guide RNA is (a) a nucleotide sequence having at least 95% identity with SEQ ID NO: 13, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% identity with SEQ ID NO: 11; (b) a nucleotide sequence having at least 95% identity with SEQ ID NO: 3, wherein the RGN polypeptide comprises an amino acid sequence having at least 98% identity with SEQ ID NO: 1; or (c) a nucleotide sequence having at least 95% identity with SEQ ID NO: 21, wherein the RGN polypeptide comprises an amino acid sequence having at least 98% identity with SEQ ID NO: 19 and comprises a tracrRNA. The vector according to claim 15 or 16.
18. A cell that contains the nucleic acid molecule according to any one of claims 1 to 14 or the vector according to any one of claims 15 to 17 and is not a human embryonic cell.
19. A system for binding to a target DNA sequence, the system comprising: a) one or more guide RNAs capable of hybridizing to the target DNA sequence, or one or more nucleotide sequences encoding the one or more guide RNAs (gRNAs); b) (i) having a sequence that is at least 90% identical to SEQ ID NO: 11; or (ii) having a sequence that is at least 98% identical to SEQ ID NO: 1 or 19, an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence, or a nucleotide sequence encoding the RGN polypeptide; A system comprising the same.
20. The system according to claim 19, wherein at least one of the nucleotide sequences encoding the one or more guide RNAs and encoding the RGN polypeptide is operably linked to a promoter heterologous to the nucleotide sequence.
21. The system according to claim 19 or 20, wherein the one or more guide RNAs form a complex with the RGN polypeptide, thereby binding the RGN polypeptide to the target DNA sequence.
22. The system according to any one of claims 19 to 21, wherein the target DNA sequence is in a eukaryotic cell and the eukaryotic cell is not a human embryonic cell.
23. The system according to any one of claims 19 to 22, wherein the RGN polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO:
11.
24. The system according to any one of claims 19 to 23, wherein the RGN polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 11, 1, or 19.
25. The system according to any one of claims 19 to 23, wherein the RGN polypeptide is nuclease-dead or functions as a nickase.
26. The system according to any one of claims 19 to 25, wherein the RGN polypeptide is operably linked to a base editing polynucleotide.
27. The system according to any one of claims 19 to 25, wherein the RGN polypeptide is operably fused to a nuclear localization signal, a plastid localization signal, a mitochondrial localization signal, a dual-targeting localization signal, and / or a cell penetration domain.
28. The system according to claim 27, wherein the nuclear localization signal, plastid localization signal, mitochondrial localization signal, dual-targeting localization signal, and / or cell penetration domain is operably fused to the N-terminus, C-terminus, or internal position of the RGN polypeptide.
29. The system according to any one of claims 19 to 25, wherein the RGN polypeptide is operably fused to an effector domain.
30. The system according to claim 29, wherein the effector domain is a cleavage domain, a deaminase domain, or an expression modulator domain.
31. The system according to claim 29 or 30, wherein the effector domain is operably fused to the N-terminus, C-terminus, or internal position of the RGN polypeptide.
32. The system according to any one of claims 19 to 31, further comprising one or more donor polynucleotides, or one or more nucleotide sequences encoding the one or more donor polynucleotides.
33. A method of binding to a target DNA sequence, comprising delivering the system according to any one of claims 19 to 32 to the target DNA sequence or to a cell comprising the target DNA sequence, wherein the cell is not a human embryonic cell.
34. a) (i) having a sequence that is at least 90% identical to SEQ ID NO: 11; or (ii) having a sequence that is at least 98% identical to SEQ ID NO: 1 or 19, an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence; and b) one or more guide RNAs A composition for use in cleaving and / or modifying a target DNA sequence, wherein the one or more guide RNAs hybridize to the target DNA sequence and bind to the RGN polypeptide, thereby causing the RGN polypeptide to bind to the target DNA sequence and effecting cleavage and / or modification of the target DNA sequence.
35. The composition for use according to claim 34, wherein the modified target DNA sequence comprises insertion of a heterologous DNA into the target DNA sequence.
36. The composition for use according to claim 34, wherein the modified target DNA sequence comprises deletion of at least one nucleotide from the target DNA sequence.
37. The composition for use according to claim 36, wherein the modified target DNA sequence comprises a mutation of at least one nucleotide in the target DNA sequence.
38. The composition for use according to any one of claims 34 to 37, wherein the target DNA sequence is in a cell, and wherein the cell is not a human embryonic cell.
39. The composition for use according to claim 38, wherein the cell is a eukaryotic cell.
40. The composition for use according to claim 38 or 39, further comprising culturing the cell under conditions such that the RGN polypeptide is expressed to cleave the target DNA sequence to generate a modified DNA sequence; and selecting a cell comprising the modified DNA sequence.
41. (i) having a sequence that is at least 90% identical to SEQ ID NO: 11; or (ii) having a sequence that is at least 98% identical to SEQ ID NO: 1 or 19, an RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence.
42. The RGN polypeptide according to claim 41, wherein the RGN polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO:
11.
43. The RGN polypeptide according to claim 41 or 42, wherein the RGN polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 11, 1, or 19.
44. The RGN polypeptide according to claim 41 or 42, wherein the RGN polypeptide is nuclease-dead or functions as a nickase.
45. The RGN polypeptide according to any one of claims 41 to 44, wherein the RGN polypeptide is functionally fused to a base editing polypeptide.
46. The RGN polypeptide according to any one of claims 41 to 44, wherein the RGN polypeptide is functionally fused to a nuclear localization signal, a plastid localization signal, a mitochondrial localization signal, a dual-targeting localization signal, and / or a cell-penetrating domain.
47. The RGN polypeptide according to claim 46, wherein the nuclear localization signal, plastid localization signal, mitochondrial localization signal, dual-target localization signal, and / or cell penetration domain is functionally fused to the N-terminus, C-terminus, or internal position of the RGN polypeptide.
48. The RGN polypeptide according to any one of claims 41 to 44, wherein the RGN polypeptide is functionally fused to an effector domain.
49. The RGN polypeptide according to claim 48, wherein the effector domain is a cleavage domain, deaminase domain, or expression modulator domain.
50. The RGN polypeptide according to claim 48 or 49, wherein the effector domain is functionally fused to the N-terminus, C-terminus, or internal position of the RGN polypeptide.
51. (i) having at least 90% sequence identity with SEQ ID NO: 11; or (ii) having at least 98% sequence identity with SEQ ID NO: 1 or 19, An RNA polynucleotide comprising a nucleotide sequence encoding an RNA-guided nuclease (RGN) polypeptide comprising the amino acid sequence.
52. The RNA polynucleotide according to claim 51, wherein the nucleotide sequence comprises at least one synthetic ribonucleotide analog.
53. The RNA polynucleotide according to claim 51 or 52, wherein the RGN polypeptide comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:
11.
54. The RNA polynucleotide according to any one of claims 51 to 53, wherein the RGN polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 11, 1, or 19.
55. The RNA polynucleotide according to any one of claims 51 to 54, further comprising a nucleotide sequence encoding a base editing polypeptide that is operably linked to the nucleotide sequence encoding the RGN polypeptide.
56. The RNA polynucleotide according to any one of claims 51 to 55, wherein the RNA polynucleotide is mRNA.
57. A eukaryotic cell comprising a nucleic acid molecule comprising a polynucleotide encoding an RNA-guided nuclease (RGN) polypeptide, wherein the RGN polypeptide is (i) having at least 90% sequence identity with SEQ ID NO: 11; or (ii) An amino acid sequence that is at least 98% identical to SEQ ID NO: 1 or 19, A eukaryotic cell comprising an amino acid sequence, wherein the eukaryotic cell is not a human embryonic cell. **Claim 58** The eukaryotic cell according to claim 57, wherein the RGN polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO:
11. **Claim 59** The eukaryotic cell according to claim 57 or 58, wherein the RGN polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 11, 1, or 19. **Claim 60** The eukaryotic cell according to any one of claims 57 to 59, wherein the polynucleotide is mRNA. **Claim 61** The eukaryotic cell according to any one of claims 57 to 60, wherein the RGN polypeptide is operably fused to a base editing polypeptide. **Claim 62** (a) (i) having a sequence that is at least 90% identical to SEQ ID NO: 11; or (ii) having a sequence that is at least 98% identical to SEQ ID NO: 1 or 19, An RNA-guided nuclease (RGN) polypeptide comprising an amino acid sequence, and (b) a base editing polypeptide or effector domain A fusion polypeptide comprising. **Claim 63** The fusion polypeptide according to claim 62, wherein the RGN polypeptide comprises an amino acid sequence that is at least 95% identical to SEQ ID NO:
11. **Claim 64** The fusion polypeptide according to claim 62 or 63, wherein the RGN polypeptide is nuclease-dead or functions as a nickase. **Claim 65** The fusion polypeptide according to claim 62 or 63, wherein the RGN polypeptide comprises the amino acid sequence of SEQ ID NO: 11, 1, or 19. **Claim 66** The fusion polypeptide according to any one of claims 62 to 65, wherein the fusion polypeptide comprises the RGN polypeptide and the base editing polypeptide. **Claim 67** The fusion polypeptide according to any one of claims 62 to 65, wherein the fusion polypeptide comprises the RGN polypeptide and the effector domain. **Claim 68** The fusion polypeptide according to claim 67, wherein the effector domain is a cleavage domain, a deaminase domain, or a transcriptional regulatory domain. **Claim 69** The fusion polypeptide according to claim 67 or 68, wherein the effector domain is operably fused to the N-terminus, C-terminus, or internal position of the RGN polypeptide. **Claim 70** The fusion polypeptide according to any one of claims 62 to 69, further comprising one or more nuclear localization signals.
71. A cell, which is not a human embryonic cell, and contains the fusion polypeptide according to any one of claims 62 to 70.
72. The cell according to claim 71, wherein the cell is a eukaryotic cell.
73. A nucleic acid molecule comprising a polynucleotide encoding the fusion polypeptide according to any one of claims 62 to 70.
74. A vector comprising the nucleic acid molecule according to claim 73.
75. An RNA polynucleotide comprising a nucleotide sequence encoding the fusion polypeptide according to any one of claims 62 to 70.
76. The RNA polynucleotide according to claim 75, wherein the RNA polynucleotide is mRNA.
Citation Information
Patent Citations
Novel CAS9 systems and methods of use
WO2017155714A1
Novel CAS9 systems and methods of use
WO2017155717A1