Compositions and methods for modifying genomes
The Cms1 CRISPR nuclease system addresses the limitations of non-specific genome editing by using guide RNAs to target and introduce precise double-strand breaks, enabling efficient and targeted genomic modifications and gene expression modulation.
Patent Information
- Application Number
- JP2025152643
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-12-15
- Filing Date
- 2025-09-12
- Publication Date
- 2025-12-19
AI Technical Summary
Existing genomic modification methods, such as CRISPR-Cas9, often result in non-specific DNA alterations and off-target effects, limiting their precision and efficacy in site-specific genome editing.
Utilization of the Cms1 CRISPR nuclease system, which includes specific guide RNAs to target and introduce double-strand breaks at predetermined genomic sites, allowing for precise genome editing and modulation of gene expression without altering the DNA sequence.
The Cms1 CRISPR system enables precise and specific genome editing with minimal off-target effects, facilitating the introduction of mutations, deletions, or insertions at desired genomic locations, and modulating gene expression through targeted protein fusion constructs.
Smart Images

Figure 2025185269000014 
Figure 2025185269000015 
Figure 2025185269000016
Abstract
Description
[Technical Field]
[0001] The present invention relates to compositions and methods for editing genomic sequences and modulating gene expression at preselected locations.
[0002] A reference to the sequence list sent as a text file via EFS-WEB The official copy of the sequence listing will be sent, along with the specification, via EFS-Web as an American Standard for Information Interchange (ASCII)-compliant text file with the filename BHP017P5Sequence Listing ST25.txt, submitted with creation date: August 3, 2018, size 1,848 Kb. Sequence listings submitted via EFS-Web are part of the specification and are incorporated herein by reference in their entirety. [Background technology]
[0003] Genomic DNA modification is of great importance for basic and applied research. Genome modification has the potential to elucidate and potentially cure disease causes and provide desirable characteristics to cells and / or individuals containing the modification. Genome modification can include, for example, modifying the genomes of plants, animals, fungi, and / or prokaryotes. While most common methods for modifying genomic DNA tend to alter DNA at random sites within the genome, recent discoveries have made site-specific genome modification possible. Such technologies rely on the creation of DSBs at the desired site. This DSB triggers the recruitment of the host cell's native DNA repair machinery to the DSB. DNA repair mechanisms can be utilized to insert heterologous DNA at a predetermined site, delete native genomic DNA, or create point mutations, insertions, or deletions at a desired site. Of particular interest for site-specific genome modification are clustered regularly interspaced short palindromic repeat (CRISPR) nucleases. CRISPR nucleases use a guide molecule, often a guide RNA molecule, that interacts with the target molecule and base pairs with the nuclease, allowing the nuclease to create a double-stranded break (DSB) at the desired site. Creation of a DSB requires the presence of a protospacer adjacent motif (PAM) sequence. Once the PAM sequence is recognized, the CRISPR nuclease can create the desired DSB. Cms1 CRISPR nuclease is a class of CRISPR nuclease with certain desirable properties compared to other CRISPR nucleases, such as Cas9 nuclease.
[0004] One area in which genome modification is being conducted is the modification of plant genomic DNA. This modification is extremely important for both basic and applied plant research. Transgenic plants whose genomic DNA has been stably modified can have new traits, such as herbicide resistance, insect resistance, and / or the accumulation of valuable proteins, including pharmaceutical proteins and industrial enzymes, attached to them. The expression of native plant genes can be up- or down-regulated or otherwise altered (e.g., by changing the tissue in which the native plant gene is expressed), their expression can be abolished entirely, DNA sequences can be altered (e.g., via point mutations, insertions, or deletions), or new, non-native genes can be inserted into the plant genome to confer new traits to the plant. Summary of the Invention
[0005] Compositions and methods are provided for modifying genomic DNA sequences using the Cms1 CRISPR system. As used herein, genomic DNA refers to linear and / or chromosomal DNA, and / or plasmid or other extrachromosomal DNA sequences present in a cell or cells of interest. The method generates a double-strand break (DSB) at a predetermined target site in the genomic DNA sequence, resulting in a mutation, insertion, and / or deletion of the DNA sequence at the target site in the genome. The composition includes a DNA construct comprising a nucleotide sequence encoding a Cms1 protein operably linked to a promoter operable in the cell of interest. In some embodiments, the Cms1 protein comprises at least one amino acid motif selected from the group consisting of SEQ ID NOs: 177-186. In other embodiments, the Cms1 protein comprises at least one amino acid motif selected from the group consisting of SEQ ID NOs: 288-289 and 187-201. In other embodiments, the Cms1 protein comprises at least one amino acid motif selected from the group consisting of SEQ ID NOs: 290-296. In certain preferred embodiments, the Cms1 protein comprises two or more amino acid motifs selected from the group consisting of SEQ ID NOs: 177-186. In certain preferred embodiments, the Cms1 protein comprises two or more amino acid motifs selected from the group consisting of SEQ ID NOs: 288-289 and 187-201. In certain preferred embodiments, the Cms1 protein comprises two or more amino acid motifs selected from the group consisting of SEQ ID NOs: 290-296. Particular Cms1 protein sequences are set forth in SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, 208-211, and 222-254. Polynucleotide sequences encoding particular Cms1 proteins are set forth in SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287. In certain preferred embodiments, the Cms1 protein has at least about 80% identity to a sequence selected from the group consisting of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287.DNA constructs containing polynucleotide sequences encoding the Cms1 proteins of the present invention, or the Cms1 proteins themselves, can be used to induce modification of genomic DNA at predetermined genomic loci. Methods for using these DNA constructs to modify genomic DNA sequences are described herein. Modified eukaryotes and eukaryotic cells, including yeast, amoeba, insects, fungi, mammals, plants, plant cells, plant parts, and seeds, as well as modified prokaryotes, including bacteria and archaea, are also included. Compositions and methods for modulating gene expression are also provided. The methods involve targeting a protein to a predetermined site in the genome to up- or down-regulate one or more genes whose expression is regulated by the target site in the genome. Compositions include DNA constructs containing nucleotide sequences encoding modified Cms1 proteins with reduced or eliminated nuclease activity, optionally fused to a transcriptional activation or repression domain. Methods for using these DNA constructs to modify gene expression are described herein. [Brief explanation of the drawings]
[0006] [Figure 1] Figure 1 shows a phylogenetic tree derived from RuvC-anchored MUSCLE alignments of the amino acid sequences of the indicated type V nucleases. Sm-, Sulf-, and Unk40-type Cms1 nucleases are shown. [Figure 2] Figure 2 summarizes the amino acid motifs common to Sm-type Cms1 proteins. The blog diagrams in Boxes 1 to 10 correspond to SEQ ID NOS: 177 to 186, respectively, and their locations on the SmCms1 protein (SEQ ID NOS: 10) are indicated. [Figure 3] Figure 3 summarizes the amino acid motifs common to Sulf-type Cms1 proteins. The blog diagrams in Boxes 1 to 17 correspond to SEQ ID NOS: 288 to 289 and 187 to 201, respectively, and their locations on the SulfCms1 protein (SEQ ID NOS: 11) are indicated. [Figure 4]Figure 4 summarizes the amino acid motifs common to Unk40-type Cms1 proteins. The blog diagrams in Boxes 1 to 7 correspond to SEQ ID NOS: 290 to 296, respectively, and their locations on the Unk40Cms1 protein (SEQ ID NOS: 68) are indicated. DETAILED DESCRIPTION OF THE INVENTION
[0007] Provided herein are methods and compositions for controlling gene expression, including sequence targeting, such as genome perturbation or gene editing, involving the CRISPR-Cms system and its components. The CRISPR enzyme of the present invention is selected from Cms enzymes. It is a Cms1 orthologue or mutant Cms1 enzyme. Cms1 is an abbreviation for CRISPR from Microgenomates and Smithella, and is so named because some bacterial species in these groups encode Cms1 nuclease. The terms Csm1 and Cms1 are used interchangeably herein. Cms1 nuclease is also known as Cas12f nuclease. The methods and compositions include nucleic acids for binding to target DNA sequences. Nucleic acids are advantageous because they are much easier and cheaper to produce than, for example, peptides, and specificity can be varied depending on the length of the stretch of homology desired. For example, complex 3D arrangements of multiple fingers are not required.
[0008] Also provided are nucleic acids encoding Cms1 polypeptides, as well as methods for using Cms1 polypeptides to modify chromosomal (i.e., genomic) or organelle DNA sequences in host cells, including plant cells. Cms1 polypeptides interact with specific guide RNAs (gRNAs), which guide the Cms1 endonuclease to specific target sites. At this site, the Cms1 endonuclease introduces a double-strand break that can be repaired by DNA repair processes, resulting in an altered DNA sequence. Because specificity is provided by the guide RNA, the Cms1 polypeptide is universal and can be used with various guide RNAs to target various genomic sequences. Cms1 endonuclease has advantages over Cas nucleases (e.g., Cas9) traditionally used in CRISPR arrays. For example, Cms1-associated CRISPR arrays are processed into mature crRNAs without the need for an additional trans-activating crRNA (tracrRNA). Additionally, the Cms1-crRNA complex can cleave target DNA preceded by a short, often T-rich, protospacer adjacent motif (PAM), in contrast to the G-rich PAM that follows target DNA in many Cas9 systems. Furthermore, Cms1 nuclease can introduce staggered DNA double-strand breaks. The methods disclosed herein can be used to target and modify specific chromosomal sequences and / or introduce exogenous sequences into targeted locations in the genomes of eukaryotic and prokaryotic cells. These methods can also be used to introduce sequences or modify regions within organelles (e.g., chloroplasts and / or mitochondria). Furthermore, targeting is specific, resulting in limited off-target effects.
[0009] I. Cms1 endonuclease Provided herein are Cms1 endonucleases, as well as fragments and variants thereof, for use in modifying genomes, including plant genomes. As used herein, the terms Cms1 endonuclease or Cms1 polypeptide refer to homologs, orthologs, and variants of the Cms1 polypeptide sequences set forth in SEQ ID NOS: 10, 11, 20-23, 30-69, 154-156, 208-211, and 222-254. Typically, Cms1 endonucleases act without the use of tracrRNA and can cause staggered DNA double-strand breaks. Generally, Cms1 polypeptides contain at least one RNA recognition and / or RNA binding domain. The RNA recognition and / or RNA binding domain interacts with the guide RNA. Typically, the guide RNA contains a region with a stem-loop structure that interacts with the Cms1 polypeptide. This stem-loop often has the sequence UCUACN 3-5 It contains GUAGAU (encoded by SEQ ID NOs: 312-314, SEQ ID NOs: 315-317), with the base pairs "UCUAC" and "GUAGA" forming the stem of the stem-loop. 3-5 indicates that any base may be present at this position, and that 3, 4, or 5 nucleotides may be included at this position. A Cms1 polypeptide may also include a nuclease domain (i.e., a DNase or RNase domain), a DNA-binding domain, a helicase domain, an RNAse domain, a protein-protein interaction domain, a dimerization domain, and other domains. In certain embodiments, a Cms1 polypeptide, or a polynucleotide encoding a Cms1 polypeptide, includes: an RNA-binding portion that interacts with a DNA-targeting RNA, and an active portion that exhibits site-specific enzymatic activity, such as a RuvC endonuclease domain.
[0010] The Cms1 polypeptide can be a wild-type Cms1 polypeptide, a modified Cms1 polypeptide, or a fragment of a wild-type or modified Cms1 polypeptide. The Cms1 polypeptide can be modified to increase nucleic acid binding affinity and / or specificity, to modify enzymatic activity, and / or to change another property of the protein. For example, the nuclease (i.e., DNase, RNase) domain of the Cms1 polypeptide can be modified, deleted, or inactivated. Alternatively, the Cms1 polypeptide can be truncated to remove domains that are not essential for protein function.
[0011] In some embodiments, the Cms1 polypeptide may be derived from a wild-type Cms1 polypeptide or a fragment thereof. In other embodiments, the Cms1 polypeptide may be derived from a modified Cms1 polypeptide. For example, the amino acid sequence of the Cms1 polypeptide may be modified to alter one or more properties of the protein (e.g., nuclease activity, affinity, stability, etc.). Alternatively, domains of the Cms1 polypeptide that are not involved in RNA-guided cleavage may be removed from the protein, such that the modified Cms1 polypeptide is smaller than the wild-type Cms1 polypeptide.
[0012] Generally, a Cms1 polypeptide contains at least one nuclease (i.e., DNase) domain, but need not contain an HNH domain such as that found in Cas9 proteins. For example, a Cms1 polypeptide may contain a RuvC or RuvC-like nuclease domain. In some embodiments, a Cms1 polypeptide may be modified to inactivate the nuclease domain so that it is no longer functional. In some embodiments in which one of the nuclease domains is inactive, the Cms1 polypeptide does not cleave double-stranded DNA. In certain embodiments, a mutant Cms1 polypeptide contains one or more mutations at positions corresponding to positions 701 or 922 of SmCms1 (SEQ ID NO: 10) or positions 848 and 1213 of SulfCms1 (SEQ ID NO: 11) when aligned for maximum identity, which reduces or eliminates nuclease activity. Nuclease domains can be modified using well-known methods such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, as well as other methods known in the art. The use of a Cms1 protein with an inactivated nuclease domain (dCms1 protein) allows for the modulation of gene expression without altering the DNA sequence. In certain embodiments, the dCms1 protein can be targeted to specific regions of the genome, such as the promoters of one or more genes of interest, using an appropriate gRNA. The dCms1 protein can bind to a desired region of DNA and potentially interfere with the binding of RNA polymerase or transcription factors to this region of DNA. This technique can be used to up- or down-regulate the expression of one or more genes of interest. In certain other embodiments, the dCms1 protein can be fused to a repressor domain of the gRNA to further down-regulate the expression of one or more genes whose expression is regulated by interaction with regions of chromosomal DNA targeted by RNA polymerase, transcription factors, or other transcriptional regulators.In certain other embodiments, the dCms1 protein can be fused to an activation domain to upregulate one or more genes whose expression is regulated by interaction of the gRNA with a region of chromosomal DNA that targets an RNA polymerase, transcription factor, or other transcriptional regulator.
[0013] The Cms1 polypeptides disclosed herein may further comprise at least one nuclear localization signal (NLS). Generally, an NLS comprises a series of basic amino acids. Nuclear localization signals are known in the art (see, for example, Lange et al., J. Biol. Chem. (2007) 282:5101-5105). The NLS can be located at the N-terminus, C-terminus, or internal position of the Cms1 polypeptide. In some embodiments, the Cms1 polypeptide can further comprise at least one cell-penetrating domain. The cell-penetrating domain can be located at the N-terminus, C-terminus, or internally in the protein.
[0014] The Cms1 polypeptides disclosed herein may further comprise at least one plastid targeting signal peptide, at least one mitochondrial targeting signal peptide, or a signal peptide that targets the Cms1 polypeptide to both plastids and mitochondria. Plastid, mitochondrial, and dual-targeting signal peptide localization signals are known in the art (e.g., Nassoury and Morse (2005) Biochim Biophys Acta 1743:5-19; Kunze and Berger (2015) Front Physiol 6:259; Herrmann, Neupert (2003) IUBMB Life 55:219-225; Soll (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Carrie et al (2009) FEBS J 276:1187-1195; Silva-Filho (2003) Curr Opin Plant Biol 6:589-595; Peeters and Small (2001) Biochim Biophys (See Acta 1541:54-63; Murcha et al (2014) J Exp Bot 65:6301-6335; Mackenzie (2005) Trends Cell Biol 15:548-554; Glaser et al (1998) Plant Mol Biol 38:311-338.) Plastid, mitochondrial, or dual targeting signal peptides can be located at the N-terminus, C-terminus, or internally within the Cms1 polypeptide.
[0015] In still other embodiments, the Cms1 polypeptide may also comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In certain embodiments, the marker domain may be a fluorescent protein. Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, EGFP, emerald, thistle green, monomeric thistle green, CopGFP, AceGFP, ZsGreen1), yellow fluorescent proteins (e.g., YFP, EYFP), citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, azurite, mKalama1, GFPuv, sapphire, T-Sapphire), cyan fluorescent proteins (e.g., Examples of suitable fluorescent proteins include ECFP, Cerulean, CyPet, AmCyan1, and Green Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, and Jred), and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, and tdTomato), or other suitable fluorescent proteins. In other embodiments, the marker domain may be a purification tag and / or an epitope tag. Exemplary tags include glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tag, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, biotin carboxyl carrier protein (BCCP), and calmodulin.
[0016] In certain embodiments, the Cms1 polypeptide may be part of a protein-RNA complex that includes a guide RNA. The guide RNA interacts with the Cms1 polypeptide to guide it to a specific target site. The 5' end of the guide RNA can base pair with a specific protospacer sequence of a nucleotide sequence of interest in the plant genome, such as the nuclear, plastid, or mitochondrial genome. As used herein, the term "DNA-targeting RNA" refers to a guide RNA that interacts with a Cms1 polypeptide and a target site of a nucleotide sequence of interest in the genome of a plant cell. The DNA-targeting RNA, or a DNA polynucleotide encoding the DNA-targeting RNA, may include a first segment comprising a nucleotide sequence complementary to a sequence in the target DNA and a second segment that interacts with the Cms1 polypeptide.
[0017] Polynucleotides encoding the Cms1 polypeptides disclosed herein can be used to isolate corresponding sequences from other prokaryotes or eukaryotes, or from sequences derived from metagenomes whose natural host organisms are unknown or unclear. In this manner, methods such as PCR, hybridization, and the like can be used to identify such sequences based on their sequence homology or identity to the sequences described herein. Sequences isolated based on their sequence identity to the entire Cms1 sequences described herein, or variants and fragments thereof, are encompassed by the present invention. Such sequences include sequences that are orthologs of the disclosed Cms1 sequences. "Ortholog" is intended to mean a gene derived from a common ancestral gene and found in different species as a result of speciation. Genes found in different species are considered orthologs if their nucleotide sequences and / or the protein sequences they encode share at least about 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity. Ortholog function is well conserved between species. Thus, isolated polynucleotides encoding polypeptides with Cms1 endonuclease activity and sharing at least about 75% or more sequence identity with the sequences disclosed herein are included in the present invention. As used herein, Cms1 endonuclease activity refers to CRISPR endonuclease activity, in which a guide RNA (gRNA) associated with a Cms1 polypeptide binds the Cms1-gRNA complex to a predetermined nucleotide sequence complementary to the gRNA. Cms1 activity can then introduce a double-stranded break at or near the site targeted by the gRNA. In certain embodiments, the double-strand break can be a staggered DNA double-strand break. As used herein, a "staggered DNA double-strand break" can result in about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 double-strand breaks. The 3' or 5' overhanging nucleotides after cleavage. In certain embodiments, the Cms1 polypeptide introduces a staggered DNA double-strand break with a 5' overhang.The double-stranded break can occur at or near the sequence to which the DNA-targeting RNA (e.g., guide RNA) sequence is targeted.
[0018] Fragments and variants of Cms1 polynucleotides and the Cms1 amino acid sequences encoded thereby that retain Cms1 nuclease activity are encompassed herein. By "Cms1 nuclease activity" is intended guide RNA-mediated binding of a predetermined DNA sequence. In embodiments in which the Cms1 nuclease retains a functional RuvC domain, Cms1 nuclease activity can further include inducing double-strand breaks. "Fragment" refers to a portion of a polynucleotide or a portion of an amino acid sequence. "Variant" is intended to refer to a substantially similar sequence. In the case of polynucleotides, variants include polynucleotides with deletions (i.e., truncations) at the 5' and / or 3' ends; deletions and / or additions of one or more nucleotides at one or more internal sites of a naturally occurring polynucleotide; and / or substitutions of one or more nucleotides at one or more sites of a naturally occurring polynucleotide. As used herein, a "native" polynucleotide or polypeptide includes a naturally occurring nucleotide sequence or amino acid sequence, respectively. Generally, variants of a particular polynucleotide of the present invention will have at least about 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the particular polynucleotide as determined by sequence alignment programs and parameters described elsewhere herein.
[0019] A "variant" amino acid or protein is intended to mean an amino acid or protein derived from a naturally occurring amino acid or protein by the deletion (so-called truncation) of one or more amino acids at the N-terminus and / or C-terminus of the naturally occurring protein; the deletion and / or addition of one or more amino acids at one or more internal sites of the naturally occurring protein; or the substitution of one or more amino acids at one or more sites of the naturally occurring protein. Mutant proteins encompassed by the present invention are biologically active, i.e., they retain the desired biological activity of the naturally occurring protein. Biologically active variants of naturally occurring polypeptides have at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with the amino acid sequence of the naturally occurring sequence, as determined by sequence alignment programs and parameters described herein. Biologically active variants of the proteins of the present invention may differ from the proteins by as few as 1-15 amino acid residues, as few as 1-10, e.g., 6-10, as few as 5, as few as 4, 3, 2, or 1 amino acid residue.
[0020] Mutant sequences may also be identified by analysis of existing databases of sequenced genomes, in which way corresponding sequences can be identified and used in the methods of the invention.
[0021] Methods of alignment of sequences for comparison are well known in the art. Thus, the determination of percent sequence identity between any two sequences can be accomplished using a mathematical algorithm. Non-limiting examples of such mathematical algorithms include the algorithm of Myers and Miller (1988) CABIOS 4:11-17; the local alignment algorithm of Smith et al. (1981) Adv. Appl. Math. 2:482; the global alignment algorithm of Needleman and Wunsch (1970) J. Mol. Biol. 48:443-453; the search-for-local algorithm method of Pearson and Lipman (1988) Proc. Natl. Acad. Sci. 85:2444-2448; and the algorithm of Karlin and Altschul (1990) modified as Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877.
[0022] Computer implementations of these mathematical algorithms can be used for comparing sequences to determine sequence identity. Such implementations include, but are not limited to, the following: CLUSTAL (PC / Gene program available from Intelligenetics, Mountain View, California); ALIGN program (Version 2.0), GAP, BESTFIT, BLAST, FASTA, and TFASTA (available from the GCG Wisconsin Genetics Software Package, Version 10 (Accelrys Inc., 9685 Scranton Road, San Diego, California, USA)). Genetic analyses using these programs can be performed using default parameters. The CLUSTAL program has been modified from Higgins et al. (1988) Gene 73:237-244; Higgins et al. (1989) CABIOS 5:151-153; Corpet et al. (1988) Nucleic Acids Res. 16:10881-90; Huang et al. (1992) CABIOS 8:155-65; Pearson et al. al. (1994) Meth. Mol. Biol. 24:307-331. The ALIGN program is based on the algorithm of Myers and Miller (1988) supra. When comparing amino acid sequences, the ALIGN program can use a PAM120 weighted residual table, a gap length penalty of 12, and a gap penalty of 4. The MUSCLE algorithm for multiple sequence alignment can be used to compare multiple nucleic acid or protein sequences (Edgar (2004) Nucleic Acids Research 32:1792-1797). The BLAST program of Altschul et al. (1990) J. Mol. 215:403 is based on the algorithm of Karlin and Altschul (1990) supra.BLAST nucleotide searches can be performed with the BLASTN program, score = 100, word length = 12, to obtain nucleotide sequences homologous to the nucleotide sequences encoding the proteins of the present invention. BLAST protein searches can be performed with the BLASTX program, score = 50, word length = 3, to obtain amino acid sequences homologous to the proteins or polypeptides of the present invention. To obtain gapped alignments for comparison, Gapped BLAST (BLAST 2.0) can be used as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389. Alternatively, PSI-BLAST (BLAST 2.0) can be used to perform an iterated search that detects distant relationships between molecules. See Altschul et al. (supra). When using BLAST, Gapped BLAST, or PSI-BLAST, the default parameters of the respective programs (e.g., BLASTN for nucleotide sequences, BLASTX for proteins) can be used. See the web site www.ncbi.nlm.nih.gov. Alignment can also be done manually by inspection.
[0023] Nucleic acid molecules encoding Cms1 polypeptides, or fragments or variants thereof, can be codon-optimized for expression in plants or other cells or organisms of interest. A "codon-optimized gene" is a gene whose codon usage is designed to mimic the preferred codon usage of the host cell. Nucleic acid molecules can be codon-optimized in whole or in part. Because any single amino acid (except methionine and tryptophan) is coded for by several codons, the sequence of a nucleic acid molecule can be altered without changing the encoded amino acid. Codon optimization is when one or more codons are altered at the nucleic acid level so that the amino acid remains unchanged but expression in a particular host organism is increased. Those skilled in the art will recognize that codon tables and other references providing preference information for a wide range of organisms are available in the art (see, e.g., Zhang et al. (1991) Gene 105:61-72; Murray et al. (1989) Nucl. Acids Res. 17:477-508). Methodologies for optimizing nucleotide sequences for expression in plants are provided, for example, in U.S. Patent No. 5,693,047 and U.S. Patent No. 6,015,891, and references cited therein. Examples of codon-optimized polynucleotides for expression in plants are set forth in SEQ ID NOS: 16-19, 110-120, and 174-176.
[0024] II. Fusion Proteins Provided herein are fusion proteins comprising a Cms1 polypeptide, or a fragment or variant thereof, and an effector domain. The Cms1 polypeptide can be directed to a target site by a guide RNA, where the effector domain can modify or affect the targeted nucleic acid sequence. The effector domain can be a cleavage domain, an epigenetic modification domain, a transcriptional activation domain, or a transcriptional repressor domain. The fusion protein can also contain a nuclear localization signal, a plastid signal peptide, a mitochondrial signal peptide, a signal peptide capable of transporting the protein to multiple intracellular locations, a cell penetration domain, or a marker domain, which can be located at the N-terminus, C-terminus, or internal position of the fusion protein. The Cms1 polypeptide can be located at the N-terminus, C-terminus, or internal position of the fusion protein. The Cms1 polypeptide can be fused directly to the effector domain or fused to a linker. In certain embodiments, the linker sequence fusing the Cms1 polypeptide to the effector domain is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, or 50 amino acids in length. For example, the linker can be in the range of 1-5, 1-10, 1-20, 1-50, 2-3, 3-10, 3-20, 5-20, or 10-50 amino acids in length.
[0025] In some embodiments, the Cms1 polypeptide of the fusion protein can be derived from a wild-type Cms1 protein. Cms1-derived proteins can be modified mutants or fragments. In some embodiments, Cms1 polypeptides can be modified to include a nuclease domain (e.g., a RuvC or RuvC-like domain) with reduced or eliminated nuclease activity. For example, Cms1-derived polypeptides can be modified such that the nuclease domain is deleted or mutated so that it is no longer functional (i.e., nuclease activity is absent). In particular, Cms1 polypeptides can have mutations at positions corresponding to positions 701 or 922 of SmCms1 (SEQ ID NO: 10) or positions 848 and 1213 of SulfCms1 (SEQ ID NO: 11) when aligned for maximum identity. The nuclease domain can be inactivated by one or more deletion, insertion, and / or substitution mutations using known methods such as site-directed mutagenesis, PCR-mediated mutagenesis, total gene synthesis, and other methods known in the art. In an exemplary embodiment, the Cms1 polypeptide of the fusion protein is modified by mutating the RuvC-like domain such that the Cms1 polypeptide does not have nuclease activity.
[0026] The fusion protein also includes an effector domain located at the N-terminus, C-terminus, or an internal position of the fusion protein. In some embodiments, the effector domain is a cleavage domain. As used herein, "cleavage domain" refers to a domain that cleaves DNA. Cleavage domains can be obtained from endonucleases or exonucleases. Non-limiting examples of endonucleases from which cleavage domains can be derived include, but are not limited to, restriction endonucleases and homing endonucleases. See, e.g., the New England Biolabs Catalog or Belfort et al. (1997) Nucleic Acids Res. 25:3379-3388. Additional enzymes that cleave DNA are known (e.g., S1 nuclease, mung bean nuclease, pancreatic DNase I, micrococcal nuclease, yeast HO endonuclease, etc.). See also Linn et al. (eds.) Nucleases, Cold Spring Harbour Laboratory Press, 1993. One or more of these enzymes (or functional fragments thereof) can be used as a source of the cleavage domain.
[0027] In some embodiments, the cleavage domain can be derived from a Type II-S endonuclease. Type II-S endonucleases typically cleave DNA several base pairs away from the recognition site and have separable recognition and cleavage domains. These enzymes are monomers that transiently associate to form dimers and cleave each strand of DNA at alternate positions. Non-limiting examples of suitable Type II-S endonucleases include BfiI, BpmI, BsaI, BsgI, BsmBI, BsmI, BspMI, FokI, MbolI, and SapI.
[0028] In certain embodiments, Type II-S cleavage can be modified to promote dimerization of two different cleavage domains, each bound to a Cms1 polypeptide or fragment thereof. In embodiments in which the effector domain is a cleavage domain, the Cms1 polypeptide can be modified as discussed herein to eliminate its endonuclease activity. For example, the Cms1 polypeptide can be modified by mutating the RuvC-like domain so that the polypeptide no longer exhibits endonuclease activity.
[0029] In other embodiments, the effector domain of the fusion protein can be an epigenetic modification domain. Generally, epigenetic modification domains alter histone or chromosomal structure without altering the DNA sequence. Changes in histone and / or chromatin structure can lead to changes in gene expression. Examples of epigenetic modifications include, but are not limited to, acetylation or methylation of lysine residues in histone proteins and methylation of cytosine residues in DNA. Non-limiting examples of suitable epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains, and DNA demethylase domains.
[0030] In embodiments in which the effector domain is a histone acetyltransferase (HAT) domain, the HAT domain can be derived from EP300 (i.e., E1A-binding protein p300), CREBBP (i.e., CREB-binding protein), CDY1, CDY2, CDYL1, CLOCK, ELP3, ESA1, GCN5 (KAT2A), HAT1, KAT2B, KAT5, MYST1, MYST2, MYST3, MYST4, NCOA1, NCOA2, NCOA3, NCOAT, P / CAF, Tip60, TAFII250, or TF3C4. In embodiments in which the effector domain is an epigenetic modification domain, the Cms1 polypeptide can be modified as discussed herein to eliminate its endonuclease activity. For example, the Cms1 polypeptide can be modified by mutating the RuvC-like domain so that the polypeptide no longer has nuclease activity.
[0031] In some embodiments, the effector domain of the fusion protein can be a transcriptional activation domain. Generally, a transcriptional activation domain interacts with a transcriptional control element and / or a transcriptional regulatory protein (i.e., a transcription factor, RNA polymerase, etc.) to increase and / or activate the transcription of one or more genes. In some embodiments, transcriptional activation domains include, but are not limited to, the herpes simplex virus VP16 activation domain, VP64 (a tetrameric derivative of VP16), the NFκB p65 activation domain, p53 activation domains 1 and 2, the CREB (cAMP response element binding protein) activation domain, the E2A activation domain, and the NFAT (nuclear factor of activated T cells) activation domain. In other embodiments, the transcriptional activation domain can be Gal4, Gcn4, MLL, Rtg3, Gln3, Oaf1, Pip2, Pdr1, Pdr3, Pho4, and Leu3. The transcriptional activation domain can be wild-type or a modified version of the original transcriptional activation domain. In some embodiments, the effector domain of the fusion protein is a VP16 or VP64 transcriptional activation domain. In embodiments in which the effector domain is a transcriptional activation domain, the Cms1 polypeptide can be modified as discussed herein to eliminate its endonuclease activity. For example, the Cms1 polypeptide can be modified by mutating the RuvC-like domain so that the polypeptide no longer has nuclease activity.
[0032] In yet other embodiments, the effector domain of the fusion protein can be a transcriptional repressor domain. Generally, a transcriptional repressor domain interacts with a transcriptional control element and / or a transcriptional regulatory protein (i.e., a transcription factor, RNA polymerase, etc.) to reduce and / or terminate transcription of one or more genes. Non-limiting examples of suitable transcriptional repressor domains include the inducible cAMP early repressor (ICER) domain, the Krüppel-associated box A (KRAB-A) repressor domain, the YY1 glycine-rich repressor domain, the Sp1-like repressor, the E(spl) repressor, the I kappa B repressor, and MeCP2. In embodiments where the effector domain is a transcriptional repressor domain, the Cms1 polypeptide can be modified as discussed herein to eliminate its endonuclease activity. For example, the Cms1 polypeptide can be modified by mutating the RuvC-like domain so that the polypeptide no longer has nuclease activity.
[0033] In some embodiments, the fusion protein further comprises at least one additional domain. Non-limiting examples of suitable additional domains include a nuclear localization signal, a cell penetration or translocation domain, and a marker domain.
[0034] When the effector domain of the fusion protein is a cleavage domain, a dimer containing at least one fusion protein can be formed. The dimer can be a homodimer or a heterodimer. In some embodiments, the heterodimer contains two different fusion proteins. In other embodiments, the heterodimer contains one fusion protein and an additional protein.
[0035] The dimer may be a homodimer, in which the two fusion protein monomers are identical in primary amino acid sequence. In one embodiment in which the dimer is a homodimer, the Cms1 polypeptide can be modified to eliminate endonuclease activity. In certain embodiments in which the Cms1 polypeptide is modified to eliminate endonuclease activity, each fusion protein monomer can contain an identical Cms1 polypeptide and an identical cleavage domain. The cleavage domain can be any cleavage domain, such as any of the exemplary cleavage domains provided herein. In such embodiments, a specific guide RNA guides the fusion protein monomers to different but closely adjacent sites, such that upon dimer formation, the nuclease domains of the two monomers create double-stranded breaks in the target DNA.
[0036] The dimer can also be a heterodimer of two different fusion proteins. For example, the Cms1 polypeptide of each fusion protein can be derived from a different Cms1 polypeptide or from an orthologous Cms1 polypeptide. For example, each fusion protein can contain a Cms1 polypeptide from a different source. In these embodiments, each fusion protein will recognize a different target site (i.e., specified by the protospacer and / or PAM sequence). For example, guide RNAs can position the heterodimer at different but closely adjacent sites so that their nuclease domains generate effective double-strand breaks in the target DNA.
[0037] Alternatively, the two fusion proteins of a heterodimer can have different effector domains. In embodiments in which the effector domain is a cleavage domain, each fusion protein can contain a different modified cleavage domain. In these embodiments, the Cms1 polypeptides can be modified to eliminate their endonuclease activity. The two fusion proteins forming the heterodimer can differ in both the Cms1 polypeptide domain and the effector domain.
[0038] In any of the above embodiments, the homodimer or heterodimer can include a nuclear localization signal (NLS), a plastid signal peptide, a mitochondrial signal peptide, a signal peptide capable of transporting the protein to multiple subcellular locations, a cell-penetration, translocation domain, and a marker domain, as described above. In any of the above embodiments, one or both of the Cms1 polypeptides can be modified to eliminate or modify the endonuclease activity of the polypeptide.
[0039] The heterodimer may also comprise one fusion protein and an additional protein. For example, the additional protein may be a nuclease. In one embodiment, the nuclease is a zinc finger nuclease. The zinc finger nuclease comprises a zinc finger DNA-binding domain and a cleavage domain. The zinc fingers recognize and bind to three nucleotides. The zinc finger DNA-binding domain may comprise from about three zinc fingers to about seven zinc fingers. The zinc finger DNA-binding domain may be derived from a naturally occurring protein or may be engineered. For example, Beerli et al. (2002) Nat. Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70: 313-340; Isalan et al. (2001) Nat. Biotechnol. 19:656-660; al.(2001)Curr.Opin.Biotechnol.12:632-637;Choo et al.(2000)Curr.Opin.Struct.Biol.10:411-416;Zhang et al.(2000)J.Biol.Chem.275(43):33850-33860;Doyon et al. al.(2008)Nat.Biotechnol.26:702-708;and Santiago et al. See, e.g., J. et al. (2008) Proc. Natl. Acad. Sci. USA 105:5809-5814. The cleavage domain of the zinc finger nuclease can be any cleavage domain described herein. In some embodiments, the zinc finger nuclease can include at least one additional domain selected from a nuclear localization signal, a plastid signal peptide, a mitochondrial signal peptide, a signal peptide capable of transporting the protein to multiple intracellular locations, a specific cell penetration domain, or a translocation domain, as described herein.
[0040] In certain embodiments, any of the fusion proteins detailed above, or a dimer comprising at least one fusion protein, can be part of a protein-RNA complex comprising at least one guide RNA, which interacts with the Cms1 polypeptide of the fusion protein to target the fusion protein to a specific target site, and the 5' end of the guide RNA base pairs with a specific protospacer sequence.
[0041] III. Nucleic Acids Encoding Cms1 Polypeptides or Fusion Proteins Nucleic acids encoding any of the Cms1 polypeptides or fusion proteins described herein are provided. The nucleic acids can be RNA or DNA. Examples of polynucleotides encoding Cms1 polypeptides are set forth in SEQ ID NOS: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287. In one embodiment, the nucleic acid encoding the Cms1 polypeptide or fusion protein is mRNA. The mRNA can be 5' capped and / or 3' polyadenylated. In another embodiment, the nucleic acid encoding the Cms1 polypeptide or fusion protein is DNA. The DNA can be present in a vector.
[0042] Nucleic acids encoding Cms1 polypeptides or fusion proteins can be codon-optimized for efficient translation into proteins in the plant cells of interest. Programs for codon optimization are available in the art (e.g., OPTIMIZER at genomics.urv.es / OPTIMIZER; www.genscript.com / codon opt.html GenScript's OptimumGene.TM).
[0043] In certain embodiments, DNA encoding a Cms1 polypeptide or fusion protein may be operably linked to at least one promoter sequence. The DNA coding sequence may be operably linked to a promoter control sequence for expression in a host cell of interest. In some embodiments, the host cell is a plant cell. "Operably linked" is intended to mean a functional linkage between two or more elements. For example, an operably linked linkage between a promoter and a coding region of interest (e.g., a region encoding a Cms1 polypeptide or guide RNA) is a functional link that allows for expression of the coding region of interest. Operably linked elements may or may not be contiguous. When used to refer to the joining of two protein coding regions, operably linked means that the coding regions are in the same reading frame.
[0044] The promoter sequence can be constitutive, regulated, developmental stage-specific, or tissue-specific. It is recognized that different applications can be enhanced by using different promoters in the nucleic acid molecule to regulate the timing, location, and / or level of expression of the Cms1 polypeptide and / or guide RNA. Such nucleic acid molecules can also optionally include promoter regulatory regions (e.g., inducible, constitutive, environmentally or developmentally regulated, or those that confer cell- or tissue-specific / selective expression), transcription initiation sites, ribosome binding sites, RNA processing signals, transcription termination sites, and / or polyadenylation signals.
[0045] In some embodiments, the nucleic acid molecules provided herein can be combined with constitutive, tissue-preferred, developmentally-preferred, or other promoters for expression in plants. Examples of constitutive promoters that function in plant cells include the cauliflower mosaic virus (CaMV) 35S transcription initiation region, the 1' or 2' promoter derived from the Agrobacterium tumefaciens T-DNA, the ubiquitin 1 promoter, the Smas promoter, and cinnamyl. These include the alcohol dehydrogenase promoter (U.S. Pat. No. 5,683,439), the Nos promoter, the pEmu promoter, the rubisco promoter, the GRP1-8 promoter, and other transcription initiation regions from various plant genes known to those skilled in the art. If low levels of expression are desired, weak promoters can be used. Weak constitutive promoters include, for example, the core promoter of the Rsyn7 promoter (WO 99 / 43838 and U.S. Pat. No. 6,072,050), the core 35S CaMV promoter, and the like. Other constitutive promoters are described, for example, in U.S. Patent Nos. 4,693,047, 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; and 5,608,142. See also U.S. Patent No. 5,693,047. See also U.S. Patent No. 6,177,611, which is incorporated herein by reference.
[0046] Examples of inducible promoters are the Adh1 promoter, which is inducible by hypoxia or cold stress, the Hsp70 promoter, which is inducible by heat stress, the PPDK promoter and the peptocarboxylase promoter, both of which are inducible by light. Chemically inducible promoters, such as the estrogen-inducible ERE promoter, which is more safely induced (U.S. Patent No. 5,364,780), and the Axig1 promoter, which is auxin-inducible and tapetum-specific but activated in callus (PCT US01 / 22169), are also useful.
[0047] Examples of promoters under developmental control in plants include promoters that preferentially initiate transcription in specific tissues, such as leaves, roots, fruits, seeds, or flowers. A "tissue-specific" promoter is a promoter that initiates transcription only in a specific tissue. Unlike constitutive expression of a gene, tissue-specific expression is the result of several interacting levels of gene regulation. Therefore, promoters from homologous or closely related plant species may be preferred to achieve efficient and reliable expression of a transgene in a specific tissue. In some embodiments, expression involves a tissue-preferred promoter. A "tissue-preferred" promoter is a promoter that preferentially initiates transcription in a specific tissue, but not necessarily entirely or exclusively.
[0048] In some embodiments, nucleic acid molecules encoding Cms1 polypeptides and / or guide RNAs comprise a cell-type-specific promoter. A "cell-type-specific" promoter is a promoter that primarily drives expression in a particular cell type in one or more organs. Some examples of plant cells in which a cell-type-specific promoter functional in a plant is primarily active include, for example, BETL cells, root vascular cells, leaves, stem cells, and stem cells. Nucleic acid molecules can also comprise cell-type-preferred promoters. A "cell-type-preferred" promoter is a promoter that primarily drives expression, although not necessarily predominantly or exclusively, in a particular cell type in one or more organs. Some examples of plant cells in which a cell-type-preferred promoter functional in a plant may be preferentially active include, for example, BETL cells, root vascular cells, leaves, stem cells, and stem cells. The nucleic acid molecules described herein can also comprise seed-preferred promoters. In some embodiments, the seed-preferred promoter has expression in the embryo sac, early embryo, early endosperm, aleurone, and / or basal endosperm explant cell layer (BETL).
[0049] Examples of promoters preferred for seeds include, but are not limited to, the m27kD gamma zein promoter and waxy promoter, as described in Boronat, A. et al. (1986) Plant Sci. 47:95-102; Reina, M. et al. Nucl. Acids Res. 18(21):6426; Kloesgen, R. B. et al. (1986) Mol. Gen. Genet. 203:237-244. Promoters expressed in the embryo, pericarp, and endosperm are disclosed in U.S. Patent No. 6,225,529 and PCT Publication WO 00 / 12733, the disclosures of each of which are incorporated herein by reference in their entireties.
[0050] Promoters capable of driving gene expression in a plant seed-preferred manner, with expression in the embryo sac, early embryo, early endosperm, aleurone, and / or basal endosperm transfer cell layer (BETL), can be used in the compositions and methods disclosed herein. Such promoters include, but are not limited to, Zea mays early endosperm 5 gene, Zea mays early endosperm 1 gene, Zea mays early endosperm 2 gene, GRMZM2G124663, GRMZM2G006585, GRMZM2G120008, GRMZM2G157806, GRMZM2G176390, GRMZM2G472234, GRMZM2G138727, Zea mays CLAVATA1, Zea mays MRP1, Oryza sativa PR602, Oryza sativa PR9a, Zea mays BET1, Zea mays BETL-2, Zea mays BETL-3, Zea mays BETL-4, Zea mays BETL-9, Zea mays BETL-10, Zea mays MEG1, Zea mays TCCR1, Zea mays These include promoters naturally linked to ASP1, Oryza sativa ASP1, Triticum durum PR60, Triticum durum PR91, Triticum durum GL7, AT3G10590, AT4G18870, AT4G21080, AT5G23650, AT3G05860, AT5G42910, AT2G26320, AT3G03260, AT5G26630, AtIPT4, AtIPT8, AtLEC2, and LFAH12.Additional such promoters are described in U.S. Patent Nos. 7,803,990, 8,049,000, 7,745,697, 7,119,251, 7,964,770, 7,847,160, 7,700,836, U.S. Patent Application Publication Nos. 20100313301, 20090049571, 20090089897, 20100281569, 20100281570, 20120 066795, 20040003427; PCT Publication Nos. WO / 1999 / 050427, WO / 2010 / 129999, WO / 2009 / 094704, WO / 2010 / 019996, and WO / 2010 / 147825 (each of which is incorporated by reference in its entirety for all purposes). Functional variants or functional fragments of the promoters described herein can also be operably linked to the nucleic acids disclosed herein.
[0051] Chemically regulated promoters can be used to regulate gene expression through the application of exogenous chemical regulators. Depending on the purpose, the promoter can be a chemically inducible promoter, in which application of a chemical induces gene expression, or a chemically repressible promoter, in which application of a chemical represses gene expression. Chemically inducible promoters are known in the art and include, but are not limited to, the maize In2-2 promoter, which is activated by benzenesulfonamide herbicide safeners, the maize GST promoter, which is activated by hydrophobic electrophilic compounds used as pre-emergence herbicides, and the tobacco PR-1a promoter, which is activated by salicylic acid. Other chemically regulated promoters of interest include steroid-responsive promoters (see, e.g., glucocorticoid-inducible promoters in Schena et al. (1991) Proc. Natl. Acad. Sci. USA 88:10421-10425, and McNellis et al. (1998) Plant J. 14(2):247-257), and tetracycline-inducible and tetracycline-repressible promoters (see, e.g., Gatz et al. (1991) Mol. Gen. Genet. 227:229-237, US Pat. Nos. 5,814,618 and 5,789,156, which are incorporated herein by reference).
[0052] Tissue-preferred promoters can be utilized to target enhanced expression of expression constructs in specific tissues. In certain embodiments, tissue-preferred promoters can be active in plant tissues. Tissue-preferred promoters are known in the art. For example, Yamamoto et al. (1997) Plant J.12(2):255-265; Kawamata et al. (1997) Plant Cell Physiol.38(7):792-803; Hansen et al. (1997) Mol.Gen Genet.254(3):337-343; Russell et al. (1997) Transgenic Res.6(2):157-168;Rinehart et al.(1996)Plant Physiol.112(3):1331-1341;Van Camp et al.(1996)Plant Physiol.112(2):525-535;Canevascini et al.(1996)Plant Physiol.112(2):513-524;Yamamoto et al. See, e.g., (1994) Plant Cell Physiol. 35(5):773-778; Lam (1994) Results Probl. Cell Differ. 20:181-196; Orozco et al. (1993) Plant Mol Biol. 23(6):1129-1138; Matsuoka et al. (1993) Proc Natl. Acad. Sci. USA 90(20):9586-9590; Guevara-Garcia et al. (1993) Plant J. 4(3):495-505. Such promoters can be modified for weak expression, if desired.
[0053] Leaf-preferred promoters are known in the art. See, for example, Yamamoto et al. (1997) Plant J. 12(2):255-265; Kwon et al. (1994) Plant Physiol. 105:357-67; Yamamoto et al. (1994) Plant Cell Physiol. 35(5):773-778; Gotor et al. (1993) Plant J. 3:509-18; Orozco et al. (1993) Plant Mol. Biol. 23(6):1129-1138; Matsuoka et al. (1993) Proc. Natl. Acad. Sci. USA 90(20):9586-9590. In addition, cab and rubisco promoters can also be used. See, for example, Simpson et al. (1958) EMBO J 4:2723-2729 and Timko et al. (1988) Nature 318:57-58.
[0054] Root-preferred promoters are known and can be selected from many available in the literature or isolated de novo from various compatible species (see, e.g., Hire et al. (1992) Plant Mol. Biol. 20(2):207-218 (soybean root-specific glutamine synthetase gene); Keller and Baumgartner (1991) Plant Cell 3(10):1051-1061 (root-specific regulatory element of the Phaseolus vulgaris GRP1.8 gene); Sanger et al. (1990) Plant Mol. Biol. 14(3):433-443 (root-specific promoter of the Agrobacterium tumefaciens mannopine synthase (MAS) gene); Miao et al. (1991) Plant Cell 3(1):11-22 (full-length cDNA clone encoding cytosolic glutamine synthetase (GS) expressed in soybean roots and nodules). See Bogusz et al. (1990) Plant Cell 2(7):633-641. Two root-specific promoters isolated from hemoglobin genes from the nitrogen-fixing non-legume Parasponia andersonii and the related non-nitrogen-fixing non-legume Trema tomentosa are described. The promoters of these genes were linked to a β-glucuronidase reporter gene and introduced into both the non-legume Nicotiana tabacum and the legume Lotus corniculatus; in both cases, root-specific promoter activity was maintained. Leech and Aoyagi (1991) describe the analysis of the promoters of the highly expressed roIC and roID root-inducible genes from Agrobacterium rhizogenes (see Plant Science (Limerick) 79(1):69-76). They concluded that enhancers and DNA determinants of tissue preference were dissociated in those promoters. showed that the encoding Agrobacterium T-DNA gene is particularly active in the root tip epidermis, and that the TR2' gene is root-specific in intact plants and stimulated by wounding of leaf tissue—a particularly desirable combination of properties for use in insecticidal or pesticidal genes (see EMBO J. 8(2):343-350).The TR1' gene fused to nptII (neomycin phosphotransferase II) showed similar properties. Additional root-preferred promoters include the VfENOD-GRP3 gene promoter (Kuster et al. (1995) Plant Mol. Biol. 29(4):759-772); and the roIB promoter (Capana et al. (1994) Plant Mol. Biol. 25(4):681-691. See also U.S. Patent Nos. 5,837,876, 5,750,386, 5,633,363, 5,459,252, 5,401,836, 5,110,732, and 5,023,179). The phaseolin gene (Murai et al. (1983) Science 23:476-482 and Sengopta-Gopalen et al. (1988) PNAS 82:3320-3324). The promoter sequence can be wild-type or modified for more efficient or effective expression.
[0055] The nucleic acid sequence encoding a Cms1 polypeptide or fusion protein can be operably linked to a promoter sequence recognized by phage RNA polymerase for in vitro mRNA synthesis. In such embodiments, the in vitro transcribed RNA can be purified for use in the genome modification methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence, or a variation of a T7, T3, or SP6 promoter sequence. In some embodiments, the sequence encoding a Cms1 polypeptide or fusion protein can be operably linked to a promoter sequence for in vitro expression of the Cms1 polypeptide or fusion protein in plant cells. In such embodiments, the expressed protein can be purified for use in the genome modification methods described herein.
[0056] In certain embodiments, the DNA encoding the Cms1 polypeptide or fusion protein may also be linked to a polyadenylation signal (e.g., the SV40 polyA signal and other signals functional in the cell of interest) and / or at least one transcription termination sequence. Additionally, the sequence encoding the Cms1 polypeptide or fusion protein may also be linked to a sequence encoding at least one nuclear localization signal, at least one plastid signal peptide, at least one mitochondrial signal peptide, at least one signal peptide capable of transporting the protein to multiple subcellular locations, at least one cell-penetrating domain, and / or at least one marker domain described elsewhere herein.
[0057] The DNA encoding the Cms1 polypeptide or fusion protein can be present in a vector. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, etc.). In one embodiment, the DNA encoding the Cms1 polypeptide or fusion protein is present in a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, pCAMBIA, and variants thereof. The vector may include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Additional information can be found in "Current Protocols in Molecular Biology" by Ausubel et al., John Wiley & Sons, New York, 2003, or "Molecular Cloning: A Laboratory Manual" by Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.
[0058] In some embodiments, an expression vector containing a sequence encoding a Cms1 polypeptide or fusion protein may further contain a sequence encoding a guide RNA. The sequence encoding the guide RNA may be operably linked to at least one transcriptional control sequence for expression of the guide RNA in a plant or plant cell of interest. For example, the DNA encoding the guide RNA may be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters and rice U6 and U3 promoters.
[0059] IV. Methods of Modifying the Nucleotide Sequence of a Genome Provided herein are methods for modifying a nucleotide sequence of a genome. Non-limiting examples of genomes include those of a cell, nucleus, organelle, plasmid, and virus. The method involves introducing one or more DNA-targeting polynucleotides, such as a DNA-targeting RNA ("guide RNA," "gRNA," "CRISPR RNA," or "crRNA") or a DNA polynucleotide encoding the DNA-targeting RNA, into a genomic host (e.g., a cell or organelle). The DNA-targeting polynucleotide includes (a) a first segment comprising a nucleotide sequence complementary to a sequence in the target DNA; and (b) a second segment that interacts with a Cms1 polypeptide and introduces a Cms1 polypeptide or a polynucleotide encoding a Cms1 polypeptide into the genomic host. The Cms1 polypeptide includes (a) a polynucleotide-binding portion that interacts with a gRNA or other DNA target polynucleotide; and (b) an activity portion that exhibits site-specific enzymatic activity. The genomic host can then be cultured under conditions in which the Cms1 polypeptide is expressed and cleaves the nucleotide sequence targeted by the gRNA. The system described herein does not require exogenous Mg 2+ Note that the addition of ions such as arginine, arginine, thiamin ...
[0060] The methods disclosed herein involve introducing at least one Cms1 polypeptide or a nucleic acid encoding at least one Cms1 polypeptide into a genomic host, as described herein. In some embodiments, the Cms1 polypeptide can be introduced into the genomic host as an isolated protein. In such embodiments, the Cms1 polypeptide can further comprise at least one cell-penetrating domain that facilitates cellular uptake of the protein. In some embodiments, the Cms1 polypeptide can be introduced into the genomic host as a nucleoprotein complexed with a guide polynucleotide (e.g., as a ribonucleoprotein complexed with a guide RNA). In other embodiments, the Cms1 polypeptide can be introduced into the genomic host as an mRNA molecule encoding the Cms1 polypeptide. In still other embodiments, the Cms1 polypeptide can be introduced into the genomic host as a DNA molecule comprising an open reading frame encoding the Cms1 polypeptide. Generally, a DNA sequence encoding a Cms1 polypeptide or fusion protein described herein is operably linked to a promoter sequence that functions in the genomic host. The DNA sequence can be linear, or the DNA sequence can be part of a vector. In yet other embodiments, the Cms1 polypeptide or fusion protein can be introduced into the genomic host as an RNA-protein complex comprising a guide RNA or a fusion protein and a guide RNA.
[0061] In certain embodiments, the mRNA encoding the Cms1 polypeptide may be targeted to an organelle (e.g., plastids or mitochondria). In certain embodiments, the mRNA encoding one or more guide RNAs may be targeted to an organelle (e.g., plastids or mitochondria). In certain embodiments, the mRNA encoding the Cms1 polypeptide and one or more guide RNAs may be targeted to an organelle (e.g., plastids or mitochondria). Methods for targeting mRNA to organelles are known in the art (see, e.g., U.S. Patent Application No. 2011 / 0296551, U.S. Patent Application No. 2011 / 0321187, Gomez and Pallas (2010) PLoS One 5:e12269) and are incorporated herein by reference.
[0062] In certain embodiments, the DNA encoding the Cms1 polypeptide may further comprise a sequence encoding a guide RNA. Generally, the sequences encoding the Cms1 polypeptide and the guide RNA are operably linked to one or more appropriate promoter control sequences that enable the expression of the Cms1 polypeptide and the guide RNA, respectively, in a genomic host. The DNA sequences encoding the Cms1 polypeptide and the guide RNA may further comprise additional expression control, regulatory, and / or processing sequences. The DNA sequences encoding the Cms1 polypeptide and the guide RNA may be linear or may be part of a vector.
[0063] The methods described herein can further include introducing at least one guide RNA or DNA encoding at least one polynucleotide, such as a guide RNA, into a genomic host. The guide RNA interacts with a Cms1 polypeptide to guide the Cms1 polypeptide to a specific target site, where the guide RNA bases pair with a specific DNA sequence at the target site. The guide RNA is composed of three regions: a first region that is complementary to the target site in the target DNA sequence, a second region that forms a stem-loop structure, and a third region that remains essentially single-stranded. The first region of each guide RNA is different so that each guide RNA guides the Cms1 polypeptide to a specific target site. The second and third regions of each guide RNA can be the same for all guide RNAs.
[0064] One region of the guide RNA is complementary to a sequence at the target site of the targeting DNA (i.e., the protospacer sequence) so that the first region of the guide RNA can base-pair with the target site. In various embodiments, the first region of the guide RNA can comprise from about 8 nucleotides to more than about 30 nucleotides. For example, the region of base-pairing between the first region of the guide RNA and the target site of the nucleotide sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, or about 15 nucleotides in length. It can also be about 16, about 17, about 18, about 19, about 20, about 22, about 23, about 24, about 25, about 27, about 30, or more than 30 nucleotides in length. In exemplary embodiments, the first region of the guide RNA is about 23, 24, or 25 nucleotides in length. The guide RNA can also comprise a second region that forms a secondary structure. In some embodiments, the secondary structure comprises a stem or a hairpin. The length of the stem can vary. For example, the stem can range in length from about 5 to about 6, about 10, about 15, about 20, or about 25 base pairs. The stem can include one or more bulges of 1 to about 10 nucleotides. In some preferred embodiments, the hairpin structure has the sequence UCUACN, where "UCUAC" and "GUAGA" base pairing form the stem. 3-5 GUAGAU (encoded by SEQ ID NOs: 312-314, 315-317).3-5 " indicates 3, 4, or 5 nucleotides. Thus, the total length of the second region can range from about 14 to about 25 nucleotides in length. In certain embodiments, the loop is about 3, 4, or 5 nucleotides in length and the stem comprises about 5, 6, 7, 8, 9, or 10 base pairs.
[0065] The guide RNA may also include a third region that remains essentially single-stranded. Thus, the third region has no complementarity to any nucleotide sequence in the cell of interest and no complementarity to the remainder of the guide RNA. The length of the third region may vary. Typically, the third region is greater than about 4 nucleotides in length. For example, the length of the third region may range from about 5 to about 60 nucleotides. The combined length of the second and third regions (also referred to as universal or scaffold regions) of the guide RNA may range from about 30 to about 120 nucleotides in length. In one embodiment, the combined length of the second and third regions of the guide RNA ranges from about 40 to about 45 nucleotides in length.
[0066] In some embodiments, the guide RNA comprises a single molecule containing all three regions. In other embodiments, the guide RNA may comprise two separate molecules. The first RNA molecule may comprise the first region of the guide RNA and half of the "stem" of the second region of the guide RNA. The second RNA molecule may comprise the second region of the guide RNA and the other half of the "stem" of the third region of the guide RNA. Thus, in this embodiment, the first and second RNA molecules each comprise a sequence of nucleotides complementary to each other. For example, in one embodiment, the first and second RNA molecules each comprise a sequence (about 6 to about 25 nucleotides) that base pairs with the other sequence to form a functional guide RNA. In certain embodiments, the guide RNA is a single molecule (i.e., crRNA) that interacts with a chromosomal target site and a Cms1 polypeptide without the need for a second guide RNA (i.e., tracrRNA).
[0067] In certain embodiments, the guide RNA can be introduced into a genomic host as an RNA molecule. The RNA molecule can be transcribed in vitro. Alternatively, the RNA molecule can be chemically synthesized. In other embodiments, the guide RNA can be introduced into a genomic host as a DNA molecule. In such cases, the DNA encoding the guide RNA can be operably linked to one or more promoter sequences for the expression of the guide RNA in the genomic host. For example, the RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III).
[0068] The DNA molecule encoding the guide RNA can be linear or circular. In some embodiments, the DNA sequence encoding the guide RNA can be part of a vector. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors. In an exemplary embodiment, the DNA encoding the guide RNA is present in a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, pCAMBIA, and variants thereof. The vector may include additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc.
[0069] In embodiments in which both the Cms1 polypeptide and the guide RNA are introduced into the genomic host as DNA molecules, each can be part of separate molecules (e.g., one vector containing the Cms1 polypeptide or fusion protein coding sequence and a second vector containing the guide RNA coding sequence) or both can be part of the same molecule (e.g., one vector containing the coding (and regulatory) sequences for both the Cms1 polypeptide or fusion protein and the guide RNA).
[0070] The Cms1 polypeptide, combined with a guide RNA, is directed to a target site in the host genome, where it introduces a double-strand break in the targeted DNA. There are no sequence restrictions on the target site, except that there must be a consensus sequence immediately preceding (upstream of) the target site. This consensus sequence is also known as a protospacer adjacent motif (PAM). Examples of PAM sequences include TTTN, NTTN, TTTV, and NTTV (where N is defined as any nucleotide and V is defined as A, G, or C). It is well known in the art that an appropriate PAM sequence must be positioned correctly relative to the target DNA sequence to enable Cms1 nuclease to generate the desired double-stranded break. For all Cms1 nucleases characterized to date, the PAM sequence is located immediately 5' of the target DNA sequence. The PAM site requirements for a particular Cms1 nuclease cannot currently be predicted computationally and must instead be determined experimentally using methods available in the art (Zetsche et al. (2015) Cell 163:759-771; Marshall et al. (2018) Mol Cell 69:146-157). It is well known in the art that PAM sequence specificity for a given nuclease enzyme is affected by enzyme concentration (Karvelis et al. (2015) Genome Biol 16:253). Therefore, modulating the concentration of Cms1 protein delivered to a cell or in vitro system of interest represents a method of altering one or more PAM sites associated with that Cms1 enzyme. Adjusting the Cms1 protein concentration in a system of interest can be achieved, for example, by modifying the promoter used to express the gene encoding Cms1, by changing the concentration of the ribonucleoprotein delivered to the cell or in vitro system, or by adding or deleting introns, which may play a role in regulating gene expression levels. As detailed herein, the first region of the guide RNA is complementary to the protospacer of the target sequence. Typically, the first region of the guide RNA is approximately 19 to 21 nucleotides in length.
[0071] The target site may be in the coding region of a gene, an intron of a gene, a regulatory region of a gene, an intergenic non-coding region, etc. The gene may be a protein-coding gene or an RNA-coding gene. The gene may be any gene of interest as described herein.
[0072] In some embodiments, the methods disclosed herein further include introducing at least one donor polynucleotide into the genomic host. The donor polynucleotide comprises at least one donor sequence. In some aspects, the donor sequence of the donor polynucleotide corresponds to an endogenous or native sequence found in the targeted DNA. For example, the donor sequence can be essentially identical to a portion of the DNA sequence at or near the target site, but contain at least one nucleotide change. Thus, the donor sequence can contain a modified version of the wild-type sequence at the target site, such that upon integration or replacement with the native sequence, the sequence at the target location contains at least one nucleotide change. For example, the change can be an insertion of one or more nucleotides, a deletion of one or more nucleotides, a substitution of one or more nucleotides, or a combination thereof. As a result of integration of the modified sequence, the genomic host can produce a modified gene product from the targeted chromosomal sequence.
[0073] Alternatively, the donor sequence of the donor polynucleotide may correspond to an exogenous sequence. As used herein, an "exogenous" sequence refers to a sequence that is not native to the genomic host or that is located in a different location from its natural location in the genomic host. For example, the exogenous sequence may include a protein-coding sequence, which, upon integration into the genome, may be operably linked to an exogenous promoter control sequence such that the genomic host can express the protein encoded by the integrated sequence. For example, the donor sequence may be any gene of interest, such as one encoding an agronomically important trait as described elsewhere herein. Alternatively, the exogenous sequence may be integrated into the target DNA sequence such that its expression is regulated by an endogenous promoter control sequence. In other iterations, the exogenous sequence may be a transcription control sequence, another expression control sequence, or an RNA coding sequence. The integration of an exogenous sequence into a target DNA sequence is referred to as "knock-in." The donor sequence may vary in length from a few nucleotides to hundreds of thousands of nucleotides.
[0074] In some embodiments, the donor sequence of the donor polynucleotide is flanked by upstream and downstream sequences that have substantial sequence identity with sequences located upstream and downstream, respectively, of the target site. Due to this sequence similarity, the upstream and downstream sequences of the donor polynucleotide allow for homologous recombination between the donor polynucleotide and the target sequence to incorporate (or exchange) the donor sequence into the target DNA sequence.
[0075] As used herein, an upstream sequence refers to a nucleic acid sequence that shares substantial sequence identity with the DNA sequence upstream of the target site. Similarly, a downstream sequence refers to a nucleic acid sequence that shares substantial sequence identity with the DNA sequence downstream of the target site. As used herein, the phrase "substantial sequence identity" refers to a sequence that has at least about 75% sequence identity. Thus, the upstream and downstream sequences of a donor polynucleotide may have about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the sequence upstream or downstream of the target site. In exemplary embodiments, the upstream and downstream sequences of the donor polynucleotide can have about 95% or 100% sequence identity with nucleotide sequences upstream or downstream of the target site. In one embodiment, the upstream sequence shares substantial sequence identity with a nucleotide sequence located immediately upstream of the target site (i.e., adjacent to the target site). In other embodiments, the upstream sequence shares substantial sequence identity with a nucleotide sequence located within about 100 nucleotides upstream of the target site. Thus, for example, the upstream sequence can share substantial sequence identity with a nucleotide sequence located about 1 to about 20, about 21 to about 40, about 41 to about 60, about 61 to about 80, or about 81 to about 100 nucleotides upstream from the target site. In one embodiment, the downstream sequence shares substantial sequence identity with a nucleotide sequence located immediately downstream of the target site (i.e., adjacent to the target site). In other embodiments, the downstream sequence shares substantial sequence identity with a nucleotide sequence located within about 100 (100) nucleotides downstream from the target site. Thus, for example, the downstream sequence can share substantial sequence identity with a nucleotide sequence located about 1 to about 20, about 21 to about 40, about 41 to about 60, about 61 to about 80, or about 81 to about 100 nucleotides downstream of the target site.
[0076] Each upstream or downstream sequence can range in length from about 20 nucleotides to about 5000 nucleotides. In some embodiments, the upstream and downstream sequences can comprise about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2800, 3000, 3200, 3400, 3600, 3800, 4000, 4200, 4400, 4600, 4800, or 5000 nucleotides. In exemplary embodiments, the upstream and downstream sequences can range in length from about 50 to about 1500 nucleotides.
[0077] The donor polynucleotide, which includes upstream and downstream sequences having sequence similarity to the target nucleotide sequence, can be linear or circular. In embodiments where the donor polynucleotide is circular, it can be part of a vector. For example, the vector can be a plasmid vector.
[0078] In certain embodiments, the donor polynucleotide may further comprise at least one targeting cleavage site recognized by a Cms1 polypeptide. The targeting cleavage site added to the donor polynucleotide may be located upstream or downstream of the donor sequence, or both upstream and downstream. For example, the donor sequence may be flanked by targeting cleavage sites such that upon cleavage by a Cms1 polypeptide, the donor sequence is flanked by overhangs that are compatible with those in the nucleotide sequence generated upon cleavage by the Cms1 polypeptide. Thus, the donor sequence can be linked to the cleaved nucleotide sequence during repair of the double-strand break by a non-homologous repair process. Generally, the donor polynucleotide containing the targeting cleavage site is circular (e.g., may be part of a plasmid vector).
[0079] The donor polynucleotide can be a linear molecule containing a short donor sequence with any short overhang compatible with the overhang generated by the Cms1 polypeptide. In such embodiments, the donor sequence can be directly ligated to the broken chromosomal sequence during repair of the double-strand break. In some cases, the donor sequence can be less than about 1,000, less than about 500, less than about 250, or less than about 100 nucleotides. In certain cases, the donor polynucleotide can be a linear molecule containing a short donor sequence with blunt ends. In other cases, the donor polynucleotide can be a linear molecule containing a short donor sequence with 5' and / or 3' overhangs. The overhangs can comprise 1, 2, 3, 4, or 5 nucleotides.
[0080] In some embodiments, the donor polynucleotide is DNA. The DNA can be single-stranded or double-stranded and / or linear or circular. The donor polynucleotide can be a DNA plasmid, a bacterial artificial chromosome (BAC), a yeast artificial chromosome (YAC), a viral vector, a linear piece of DNA, a PCR fragment, naked nucleic acid, or a nucleic acid complexed with a delivery vehicle such as a liposome or poloxamer. In certain embodiments, the donor polynucleotide comprising the donor sequence can be part of a plasmid vector. In any of these situations, the donor polynucleotide comprising the donor sequence can further comprise at least one additional sequence.
[0081] In some embodiments, the method can include introducing one Cms1 polypeptide (or encoding nucleic acid) and one guide RNA (or encoding DNA) into a genomic host, where the Cms1 polypeptide introduces a double-stranded break in the target DNA. In embodiments in which no donor polynucleotide is present, the double-stranded break in the nucleotide sequence can be repaired by the non-homologous end joining (NHEJ) repair process. Because NHEJ is error-prone, deletion of at least one nucleotide, insertion of at least one nucleotide, substitution of at least one nucleotide, or a combination thereof, can occur during repair of the break. Thus, the target nucleotide sequence can be modified or inactivated. For example, a single nucleotide change (SNP) can result in an altered protein product, or a shift in the reading frame of the coding sequence can inactivate or "knock out" the sequence so that the protein product is not produced. In the presence of a donor polynucleotide, the donor sequence in the donor polynucleotide can be exchanged or integrated with the nucleotide sequence at the target site during repair of the double-stranded break. With upstream and downstream sequences that share substantial sequence identity with the upstream and downstream sequences, respectively, of a target site in a nucleotide sequence, the donor sequence can be exchanged or incorporated with the nucleotide sequence of the target site during homology-mediated repair. Alternatively, in embodiments in which the donor sequence is flanked by compatible overhangs (or compatible overhangs), the donor sequence can be directly ligated to the cleaved nucleotide sequence by a non-homologous repair process during repair of a double-strand break (where the overhangs are generated in situ by a Cms1 polypeptide). Exchange or incorporation of the donor sequence into the nucleotide sequence modifies the target nucleotide sequence or introduces a foreign nucleotide sequence into the target nucleotide sequence.
[0082] The methods disclosed herein can also include introducing one or more Cms1 polypeptides (or encoding nucleic acids) and two guide polynucleotides (or encoding DNAs) into a genomic host, where the Cms1 polypeptides introduce two double-stranded breaks in a target nucleotide sequence. The two breaks can be separated by a few base pairs, tens of base pairs, or thousands of base pairs. In embodiments without any donor polynucleotides, the resulting double-stranded break can be repaired by a non-homologous repair process, such that the sequence between the two break sites is lost and / or at least one nucleotide deletion, at least one nucleotide insertion, at least one nucleotide substitution, or a combination thereof can occur during repair of the break. In embodiments where any donor polynucleotide is present, the donor sequence in the donor polynucleotide may be exchanged or integrated with the target nucleotide sequence during repair of the double-stranded break by either a homology-based repair process (e.g., in embodiments where the donor sequence is flanked by upstream and downstream sequences that have substantial sequence identity to the upstream and downstream sequences, respectively, of the target site in the nucleotide sequence), or a non-homologous repair process (e.g., in embodiments where the donor sequence is flanked by compatible overhangs).
[0083] A. Methods for modifying the nucleotide sequence of a plant genome Plant cells have nuclear, plastid, and mitochondrial genomes. The compositions and methods of the present invention can be used to modify sequences in the nuclear, plastid, and / or mitochondrial genomes or to regulate expression of genes encoded by the nuclear, plastid, and / or mitochondrial genomes. Thus, "chromosome" or "chromosome" refers to nuclear, plastid, or mitochondrial genomic DNA. "Genome" as applied to plant cells includes not only chromosomal DNA found in the nucleus, but also organelle DNA found within the subcellular components of the cell (e.g., mitochondria or plastids). Nucleotide sequences of interest in plant cells, organelles, or embryos can be modified using the methods described herein. In certain embodiments, the methods disclosed herein are used to modify nucleotide sequences encoding agronomically important traits, such as plant hormones, plant defense proteins, nutrient transport proteins, biorelevant proteins, desirable input traits, desirable output traits, stress tolerance genes, disease / pathogen resistance genes, male sterility, developmental genes, regulatory genes, genes involved in photosynthesis, DNA repair genes, transcriptional regulatory genes, or any other polynucleotides and / or polypeptides of interest. Agronomically important traits, such as oil, starch, and protein content, can also be altered. Modifications include increasing the content of oleic acid, saturated and unsaturated oils, increasing lysine and sulfur levels, providing essential amino acids, and altering starch. Hordothionin protein modifications are described in U.S. Patent Nos. 5,703,049, 5,885,801, 5,885,802, and 5,990,389, which are incorporated herein by reference. Another example is the lysine- and / or sulfur-rich seed protein encoded by soybean 2S albumin, described in U.S. Patent No. 5,850,016, and the barley-derived chymotrypsin inhibitor described in Williamson et al. (1987) Eur. J. Biochem. 165:99-106, the disclosures of which are incorporated herein by reference.
[0084] The Cms1 polypeptide (or encoding nucleic acid), guide RNA (or encoding DNA), and optional donor polynucleotide can be introduced into plant cells, organelles, or plant embryos in a variety of ways, including transformation. Transformation protocols, as well as protocols for introducing polypeptide or polynucleotide sequences into plants, can vary depending on the type of plant or plant cell targeted for transformation, i.e., monocotyledonous or dicotyledonous. Suitable methods for introducing polypeptides and polynucleotides into plant cells include microinjection (Crossway et al. (1986) Biotechniques 4:320-334), electroporation (Riggs et al. (1986) Proc. Natl. Acad. Sci. USA 83:5602-5606), Agrobacterium-mediated transformation (US Patent No. 5,563,055 and US Patent No. 5,981,840), direct gene transfer (Paszkowski et al. (1984) EMBO J. 3:2717-2722), and bombardment particle acceleration (e.g., US Patent Nos. 4,945,050; 5,879,918; 5,886,244; and 5,932,782; Tomes et al. (1995) in Plant Cell, Tissue, and Organ Culture: Fundamental Methods, ed. Gamborg and Phillips (Springer-Verlag, Berlin); McCabe et al. (1988) Biotechnology 6:923-926); and Lec1 transformation (WO00 / 28058). Weissinger et al. (1988) Ann.Rev.Genet.22:421-477; Sanford et al. (1987) Particulate Science and Technology 5:27-37 (onion); Christou et al. (1988) Plant Physiol. 87:671-674 (soybean; McCabe et al.(1988) Bio / Technology 6:923-926 (soybean); Finer and McMullen (1991) In Vitro Cell Dev. Biol. 27P:175-182 (soybean); Singh et al. (1998) Theor. Appl. Genet. 96:319-324 (soybean); Datta et al. (1990) Biotechnology 8:736-740 (rice); Klein et al. (1988) Proc. Natl. Acad. Sci. USA 85:4305-4309 (corn); Klein et al. (1988) Biotechnology 6:559-563 (corn); US Patent Nos. 5,240,855; 5,322,783; 5,324,646; Klein et al. al. (1988) Plant Physiol. 91:440-444 (maize); Fromm et al. (1990) Biotechnology 8:833-839 (maize); Hooykaas-Van Slogteren et al. (1984) Nature (London) 311:763-764; US Patent No. 5,736,369 (barley); Bytebier et al. (1987) Proc. Natl. Acad. Sci. USA 84:5345-5349 (Liliaceae); De Wet et al. (1985) in The Experimental Manipulation of Ovule Tissues, ed. Chapman et al. (Longman, New York), pp. 197-209 (pollen); Kaeppler et al. (1990) Plant Cell Reports 9:415-418 and Kaeppler et al. (1992) Theor. Appl. Genet. 84:560-566 (whisker-mediated transformation); D'Halluin et al. (1992) Plant Cell 4:1495-1505 (electroporation); Li et al.See also (1993) Plant Cell Reports 12:250-255 and Christou and Ford (1995) Annals of Botany 75:407-413 (rice); Osjoda et al. (1996) Nature Biotechnology 14:745-750 (maize via Agrobacterium tumefaciens), all of which are incorporated herein by reference. Site-specific genome editing in plant cells has been demonstrated by biolistic transfer of ribonucleoproteins containing nucleases and appropriate guide RNAs (Svitashev et al (2016) Nat Commun 7:13274); these methods are incorporated herein by reference. "Stable transformation" is intended to mean that a nucleotide construct introduced into a plant is integrated into the plant's genome and can be inherited by its progeny. The nucleotide construct may be integrated into the plant's nuclear, plastid, or mitochondrial genome. Methods for plastid transformation are known in the art (see, e.g., Chloroplast Biotechnology: Methods and Protocols (2014) Pal Maliga, ed., US Patent Application 2011 / 0321187), and methods for plant mitochondrial transformation have been described in the art (see, e.g., US Patent Application 2011 / 0296551, incorporated herein by reference).
[0085] The transformed cells can be grown (i.e., cultured) into plants according to conventional methods. See, e.g., McCormick et al. (1986) Plant Cell Reports 5:81-84. Thus, the present invention provides transformed seeds (also called "transgenic seeds") having nucleic acid modifications stably integrated into the genome.
[0086] "Introduced" in the context of inserting a nucleic acid fragment (e.g., a recombinant DNA construct) into a cell means "transfection" or "transformation" or "transduction," and includes reference to incorporation of the nucleic acid fragment into a plant, where the nucleic acid fragment may be integrated into the cell's genome (e.g., nuclear chromosome, plasmid, plastid chromosome, mitochondrial chromosome), converted into an autonomous replicon, or transiently expressed (e.g., transfected mRNA).
[0087] The present invention may be used to transform any plant species, including but not limited to monocotyledonous and dicotyledonous plants (ie, monocotyledonous and dicotyledonous, respectively). Examples of plant species of interest include, but are not limited to, corn (Zea mays), Brassica sp. (e.g., B. napus, B. rapa, B. juncea), those Brassica species useful as sources of seed oil, alfalfa (Medicago sativa), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), camelina (Camelina sativa), millet (e.g., pearl millet (Pennisetum glaucum), common millet (Panicum miliaceum), foxtail millet (Setaria italica), finger millet (Eleusine coracana)), sunflower (Helianthus annuus), quinoa (Chenopodium quinoa), chicory (Cichorium intybus), lettuce (Lactuca sativa), safflower (Carthamus tinctorius), and the like. tinctorius), wheat (Triticum aestivum), soybean (Glycine max), tobacco (Nicotiana tabacum), potato (Solanum tuberosum), peanut (Arachis hypogaea), cotton (Gossypium barbadense, Gossypium hirsutum), sweet potato (Ipomoea batatus), cassava (Manihot esculenta), coffee (Coffea spp.), coconut (Cocos nucifera), pineapple (Ananas comosus), citrus fruits (Citrus spp.), cocoa (Theobroma cacao), tea (Camellia sinensis), banana (Musa spp.)), avocado (Persea americana), fig (Ficus casica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beet (Beta vulgaris), sugarcane (Saccharum spp.), oil palm (Elaeis guineensis), poplar (Populus spp.), eucalyptus (Eucalyptus spp.), oats (Avena sativa), barley (Hordeum vulgare), vegetables, ornamentals, and conifers.
[0088] The Cms1 polypeptide (or encoding nucleic acid), guide RNA (or DNA encoding the guide RNA), and optional donor polynucleotide can be introduced simultaneously or sequentially into a plant cell, organelle, or plant embryo. The ratio of Cms1 polypeptide (or encoding nucleic acid) to guide RNA (or encoding DNA) is generally approximately stoichiometric, allowing the two components to form an RNA-protein complex with the target DNA. In one embodiment, the DNA encoding the Cms1 polypeptide and the DNA encoding the guide RNA are delivered together within a plasmid vector.
[0089] The compositions and methods disclosed herein can be used to alter the expression of a gene of interest in a plant, such as a gene involved in photosynthesis. Thus, the expression of a gene encoding a protein involved in photosynthesis can be modulated compared to a control plant. A "subject plant or plant cell" is a plant or plant cell in which a genetic modification, such as a mutation, has been made to a gene of interest, or which is derived from a plant or cell so modified and contains an alteration. A "control" or "control plant" or "control plant cell" provides a reference point for measuring phenotypic changes in the subject plant or plant cell. Thus, the expression level can be higher or lower than the expression level in a control plant, depending on the method of the present invention.
[0090] Control plants or plant cells can include, for example: (a) wild-type plants or cells, i.e., of the same genotype as the starting material for the genetic modification that resulted in the plant or cell of interest; (b) plants or plant cells of the same genotype as the starting material but that have been transformed with a null construct (i.e., a construct that has no known effect on the trait of interest, such as a construct containing a marker gene); (c) plants or plant cells that are untransformed segregants among the progeny of the plant or plant cell of interest; (d) plants or plant cells that are genetically identical to the plant or plant cell of interest but that have not been exposed to conditions or stimuli that induce expression of the gene of interest; or (e) the plant or plant cell itself under conditions in which the gene of interest is not expressed.
[0091] Although the present invention is described in terms of transformed plants, it is recognized that the transformed organisms of the present invention also include plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant callus, plant mass, and plant cells that are intact plants or plant parts such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, grains, ears, cobs, husks, stems, roots, root tips, anthers, etc. By grain is meant mature seeds produced by commercial growers for purposes other than seed cultivation or propagation. Progeny, variants, and mutants of the regenerated plants are also included within the scope of the present invention, provided that these parts contain the introduced polynucleotide.
[0092] Derivatives of coding sequences can be made using the methods disclosed herein to increase the level of a preselected amino acid in the encoded polypeptide. For example, the gene encoding barley lysine-rich polypeptide (BHL) is derived from barley chymotrypsin inhibitor; see U.S. application Ser. No. 08 / 740,682, filed Nov. 1, 1996, and WO 98 / 20133, the disclosures of which are incorporated herein by reference. Other proteins include methionine-rich plant proteins from sunflower seeds (Lilley et al. (1989) Proceedings of the World Congress on Vegetable Protein Utilization in Human Foods and Animal Feedstuffs, ed. Applewhite (American Oil Chemists Society, Champaign, Illinois), pp. 497-502; incorporated herein by reference); corn (Pedersen et al. (1986) J. Biol. Chem. 261:6279; Kirihara et al. (1988) Gene 71:359; both incorporated herein by reference); and rice (Musumura et al. (1989) Plant Mol. Biol. 12:123; incorporated herein by reference). Other agronomically important genes encode latex, floury 2, growth factors, seed storage factors, and transcription factors.
[0093] The methods disclosed herein can be used to modify herbicide tolerance traits, including, for example, genes encoding tolerance to herbicides that act by inhibiting the action of acetolactate synthase (ALS), particularly sulfonylurea-type herbicides (e.g., acetolactate synthase (ALS) genes containing mutations that confer such tolerance, particularly S4 and / or Hra mutations), or herbicides that act by inhibiting the action of glutamine synthase, such as phosphinothricin or basta; genes (e.g., the bar gene) encoding tolerance to glyphosate (e.g., the EPSPS and GAT genes; see, e.g., U.S. Publication Nos. 20040082770 and WO 03 / 092360); or other such genes known in the art. The bar gene encodes tolerance to the herbicide basta, the nptII gene encodes tolerance to the antibiotics kanamycin and geneticin, and ALS gene mutants encode tolerance to the herbicide chlorsulfuron. Additional herbicide tolerance traits are described, for example, in U.S. Patent Application 2016 / 0208243, which is incorporated herein by reference.
[0094] Sterile genes can also be modified, providing an alternative to physical degradation. Examples of genes used in such methods include genes with male sterility phenotypes, such as male tissue preference genes and QM, as described in U.S. Patent No. 5,583,210. Other genes include kinases and genes encoding compounds toxic to male or female gametophyte development. Additional sterility properties are described, for example, in U.S. Patent Application Publication No. 2016 / 0208243, which is incorporated herein by reference.
[0095] Grain quality can be altered by modifying genes that encode traits such as oil level and type, saturated and unsaturated, quality and quantity of essential amino acids, cellulose level, etc. In corn, modified hordothionin proteins are described in U.S. Patent Nos. 5,703,049, 5,885,801, 5,885,802, and 5,990,389.
[0096] Commercial traits can also be altered by modifying genes, for example, to increase starch for ethanol production, or to provide protein expression. Another important commercial use of modified plants is the production of polymers and bioplastics, as described in U.S. Patent No. 5,602,321. Genes such as β-ketothiolase, PHBase (polyhydroxyalkanoate synthase), and acetoacetyl-CoA reductase (see Schubert et al. (1988) J. Bacteriol. 170:5837-5847) promote the expression of polyhydroxyalkanoates (PHAs).
[0097] Exogenous products include plant enzymes and products, as well as products from other sources, including prokaryotes and other eukaryotes. Such products include enzymes, cofactors, hormones, etc. Protein levels can be increased, particularly modified proteins with improved amino acid distribution, to improve the nutritional value of plants. This can be achieved by expressing such proteins with enhanced amino acid content.
[0098] The methods disclosed herein can also be used to insert heterologous genes and / or modify native plant gene expression to achieve desirable plant traits. Such traits include, for example, disease resistance, herbicide tolerance, drought tolerance, salt tolerance, insect resistance, resistance to parasitic weeds, improved plant nutritional value, improved feed digestibility, increased grain yield, cytoplasmic male sterility, altered fruit ripening, increased storage life of plants or plant parts, reduced allergen production, and increased or reduced lignin content. Genes that can confer these desirable traits are disclosed in U.S. Patent Application Publication No. 2016 / 0208243, which is incorporated herein by reference.
[0099] B. Methods for modifying the nucleotide sequence of a non-plant eukaryotic genome Provided herein are methods for modifying a nucleotide sequence in a non-plant eukaryotic cell or a non-plant eukaryotic organelle. In some embodiments, the non-plant eukaryotic cell is a mammalian cell. In certain embodiments, the non-plant eukaryotic cell is a non-human mammalian cell. The method includes introducing into the target cell or organelle a DNA-targeting RNA or a DNA polynucleotide encoding the DNA-targeting RNA, wherein the DNA-targeting RNA includes (a) a first segment comprising a nucleotide sequence complementary to the sequence of the target DNA; and (b) a second segment that interacts with a Cms1 polypeptide and introduces the Cms1 polypeptide or a polynucleotide encoding the Cms1 polypeptide into the target cell or organelle, wherein the Cms1 polypeptide comprises (a) an RNA-binding portion that interacts with the DNA-targeting RNA; and (b) an activity portion that exhibits site-specific enzymatic activity. The target cell or organelle can then be cultured under conditions in which the chimeric nuclease polypeptide is expressed and cleaves the nucleotide sequence. The systems described herein can be used in combination with exogenous Mg 2+ Note that the addition of ions such as ATP or any other ions is not required. Finally, non-plant eukaryotic cells or organelles can be selected that contain the modified nucleotide sequence.
[0100] In some embodiments, methods can include introducing a Cms1 polypeptide (or encoding nucleic acid) and a guide RNA (or encoding DNA) into a non-plant eukaryotic cell or organelle, where the Cms1 polypeptide introduces a double-stranded break at a target nucleotide sequence in the nuclear or organelle chromosomal DNA. In some embodiments, methods can include introducing a Cms1 polypeptide (or encoding nucleic acid) and at least one guide RNA (or encoding DNA) into a non-plant eukaryotic cell or organelle, where the Cms1 polypeptide introduces two or more double-stranded breaks (i.e., two, three, or four or more double-stranded breaks) at a target nucleotide sequence in the nuclear or organelle chromosomal DNA. In embodiments in which no donor polynucleotide is present, the double-stranded break in the nucleotide sequence can be repaired by the non-homologous end joining (NHEJ) repair process. Because NHEJ is error-prone, deletion of at least one nucleotide, insertion of at least one nucleotide, substitution of at least one nucleotide, or a combination thereof, can occur during repair of the break. Thus, the target nucleotide sequence can be modified or inactivated. For example, a single nucleotide change (SNP) can result in an altered protein product, or a shift in the reading frame of a coding sequence can inactivate or "knock out" the sequence so that the protein product is not produced. In embodiments in which an optional donor polynucleotide is present, the donor sequence in the donor polynucleotide can be exchanged for or incorporated with the nucleotide sequence of the target site during repair of the double-strand break. For example, in embodiments in which the donor sequence is flanked by upstream and downstream sequences that have substantial sequence identity with the upstream and downstream sequences, respectively, of the target site in the nucleotide sequence of a non-plant eukaryotic cell or organelle, the donor sequence can be exchanged for or integrated with the nucleotide sequence of the target site during repair mediated by a homology-directed repair process.Alternatively, in embodiments in which the donor sequence is flanked by compatible overhangs (or compatible overhangs are generated in situ by a Cms1 polypeptide), the donor sequence can be directly ligated to the cleaved nucleotide sequence by a non-homologous repair process during repair of the double-strand break. Replacement or incorporation of the donor sequence into a nucleotide sequence modifies the target nucleotide sequence or introduces an exogenous sequence into the target nucleotide sequence in a non-plant eukaryotic cell or organelle.
[0101] In some embodiments, the double-stranded break caused by the action of a Cms1 nuclease or nuclease is repaired such that DNA is deleted from the chromosome of the non-plant eukaryotic cell or organelle, ie, one base, a few bases (i.e., 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases), or a large section of DNA (i.e., more than 10, more than 50, more than 100, or more than 500 bases) is deleted from the chromosome of the non-plant eukaryotic cell or organelle.
[0102] In some embodiments, expression of non-plant eukaryotic genes can be regulated as a result of Cms1 nuclease or double-stranded breaks caused by the nuclease. In some embodiments, expression of non-plant eukaryotic genes can be regulated by mutant Cms1 enzymes that contain mutations that prevent the Cms1 nuclease from generating double-stranded breaks. In some preferred embodiments, mutant Cms1 nucleases that contain mutations that prevent the Cms1 nuclease from generating double-stranded breaks can be fused to a transcriptional activation or transcriptional repression domain.
[0103] In some embodiments, eukaryotic cells containing mutations in their nuclear and / or organelle chromosomal DNA caused by the action of one or more Cms1 nucleases are cultured to produce eukaryotic organisms. In some embodiments, eukaryotic cells in which gene expression is modulated as a result of one or more Cms1 nucleases or one or more mutant Cms1 nucleases are cultured to produce eukaryotic organisms. Methods for culturing non-plant eukaryotic cells to produce eukaryotic organisms are known in the art, e.g., U.S. Patent Applications 2016 / 0208243 and 2016 / 0138008, each of which is incorporated herein by reference.
[0104] The present invention can be used to transform any eukaryotic species, including but not limited to animals (including but not limited to mammals, insects, fish, birds, and reptiles), fungi, amoebas, and yeast.
[0105] Methods for introducing nuclease proteins, DNA or RNA molecules encoding nuclease proteins, guide RNAs or DNA molecules encoding guide RNAs, and any donor sequence DNA molecules into non-plant eukaryotic cells or organelles are known in the art, for example, in U.S. Patent Application No. 2016 / 0208243, which is incorporated herein by reference. Exemplary genetic modifications to non-plant eukaryotic cells or organelles may be particularly valuable for industrial applications and are also known in the art, for example, in U.S. Patent Application No. 2016 / 0208243, which is incorporated herein by reference.
[0106] C. Methods for modifying the nucleotide sequence of a prokaryotic genome Provided herein is a method for modifying a nucleotide sequence in a prokaryotic (e.g., bacterial or archaeal) cell. The method comprises introducing a DNA-targeting RNA or a DNA polynucleotide encoding the DNA-targeting RNA into a target cell, wherein the DNA-targeting RNA comprises a first segment comprising a nucleotide sequence complementary to a sequence in (a) the target DNA; and (b) a second segment that interacts with a Cms1 polypeptide and introduces the Cms1 polypeptide or a polynucleotide encoding the Cms1 polypeptide into the target cell. Here, the Cms1 polypeptide comprises: an RNA-binding portion that interacts with (a) the DNA-targeting RNA; and (b) an activity portion that exhibits site-specific enzymatic activity. The target cell can then be cultured under conditions in which the Cms1 polypeptide is expressed and cleaves the nucleotide sequence. The system described herein can be modified by the addition of exogenous Mg 2+ Note that the addition of any ions, such as nucleotides, is not required. Finally, a prokaryotic cell containing the modified nucleotide sequence can be selected. Furthermore, the prokaryotic cell containing the modified nucleotide sequence is not a natural host cell for the polynucleotide encoding the Cms1 polypeptide of interest, but rather uses a non-natural guide RNA to effect the desired change in the prokaryotic nucleotide sequence. Furthermore, note that the target DNA may be present as part of the prokaryotic chromosome, or may be present on one or more plasmids or other non-chromosomal DNA molecules within the prokaryotic cell.
[0107] In some embodiments, the method can include introducing a Cms1 polypeptide (or encoding nucleic acid) and a guide RNA (or encoding DNA) into a prokaryotic cell, where the Cms1 polypeptide introduces a double-stranded break in a target nucleotide sequence in the prokaryotic cell DNA. In some embodiments, the method can include introducing a Cms1 polypeptide (or encoding nucleic acid) and at least one guide RNA (or encoding DNA) into a prokaryotic cell, where the Cms1 polypeptide introduces two or more double-stranded breaks (i.e., two, three, or four or more double-stranded breaks in a target nucleotide sequence in the prokaryotic cell DNA). In embodiments in which no donor polynucleotide is present, the double-stranded break in the nucleotide sequence can be repaired by the non-homologous end joining (NHEJ) repair process. Because NHEJ is error-prone, deletion of at least one nucleotide, insertion of at least one nucleotide, substitution of at least one nucleotide, or a combination thereof can occur during repair of the break. Thus, the target nucleotide sequence can be modified or inactivated. For example, a single nucleotide change (SNP) can result in an altered protein product, or a shift in the reading frame of a coding sequence can inactivate or "knock out" a sequence so that the protein product is not made. In embodiments in which an optional donor polynucleotide is present, the donor sequence in the donor polynucleotide can be exchanged for or incorporated with the nucleotide sequence of the target site during repair of the double-strand break. For example, in embodiments in which the donor sequence is flanked by upstream and downstream sequences that share substantial sequence identity with the upstream and downstream sequences, respectively, of the target site in the nucleotide sequence of the prokaryotic cell, the donor sequence is incorporated into the nucleotide sequence of the target site during repair mediated by a homology-directed repair process. Alternatively, in embodiments in which the donor sequence is flanked by compatible overhangs (or the compatible overhangs are generated in situ by a Cms1 polypeptide), the donor sequence can be directly ligated to the cleaved nucleotide sequence by a non-homologous repair process during repair of the double-strand break.The exchange or incorporation of a donor sequence into a nucleotide sequence modifies the target nucleotide sequence or introduces an exogenous sequence into the target nucleotide sequence of the prokaryotic DNA.
[0108] In some embodiments, the double-stranded break caused by the action of the Cms1 nuclease or nuclease is repaired such that DNA is excised from the DNA of the prokaryotic cell, ie, one base, a few bases (i.e., 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases), or a large section of DNA (i.e., 10 or more, 50 or more, 100 or more, or 500 or more bases) is excised from the DNA of the prokaryotic cell.
[0109] In some embodiments, prokaryotic gene expression can be regulated as a result of a Cms1 nuclease or a double-stranded break caused by the nuclease. In some embodiments, prokaryotic gene expression can be regulated by a mutant Cms1 nuclease that contains a mutation that prevents the Cms1 nuclease from generating a double-stranded break. In some preferred embodiments, a mutant Cms1 nuclease that contains a mutation that prevents the Cms1 nuclease from generating a double-stranded break can be fused to a transcriptional activation or transcriptional repression domain.
[0110] The strain contains a slightly different strain of the bacterium Corynebacterium sp. Bifidobacterium sp. Mycobacterium sp. Streptomyces sp. Thermobifida sp. Chlamydia sp. Prochlorococcus sp sp. Clostridium sp. Geobacillus sp. Lactobacillus sp. Listeria sp. Staphylococcus sp. Streptococcus sp. Fusobacterium sp. Agrobacterium sp. Bradyrhizobium sp. Ehrlichia sp. Mesorhizobium sp sp. Nitrobacter sp. Rickettsia sp. Wolbachia sp. Zymomonas sp. Burkholderia sp. Neisseria sp. Ralstonia sp. Acinetobacter sp. Erwinia sp. Escherichia sp. Haemophilus sp. Legionella sp. Pasteurella sp. Pseudomonas sp. Psychrobacter sp. Salmonella sp. Shewanella sp. Shigella sp. Vibrio sp. Xanthomonas sp. Xylella sp. Yersinia sp. Campylobacter sp. Desulfovibrio sp., Helicobacter sp., Geobacter sp., Leptospira sp., Treponema sp., Mycoplasma sp., and Thermotoga sp.
[0111] Methods for introducing nuclease proteins, DNA or RNA molecules encoding nuclease proteins, guide RNAs or DNA molecules encoding guide RNAs, and any donor sequence DNA molecules into prokaryotic cells or organelles are described, for example, in U.S. Patent Application No. 2016 / 0208243, which is incorporated herein by reference. Exemplary genetic modifications to prokaryotic cells that may be of particular value for industrial applications are also known in the art, for example, in U.S. Patent Application No. 2016 / 0208243, which is incorporated herein by reference.
[0112] D. Methods for modifying the nucleotide sequence of a viral genome Provided herein is a method for modifying the nucleotide sequence of a viral genome. The method involves introducing a DNA-targeting RNA or a DNA polynucleotide encoding the DNA-targeting RNA into a cell containing a virus of interest, where the DNA-targeting RNA comprises: (a) a first segment comprising a nucleotide sequence complementary to a sequence within the target DNA; and (b) a second segment that interacts with a Cms1 polypeptide and introduces the Cms1 polypeptide or a polynucleotide encoding the Cms1 polypeptide into the target cell. The Cms1 polypeptide comprises: (a) an RNA-binding portion that interacts with the DNA-targeting RNA; and (b) an active portion that exhibits site-specific enzymatic activity. Target cells containing the virus of interest can then be cultured under conditions in which the Cms1 polypeptide is expressed and cleaves the viral nucleotide sequence. Alternatively, the viral genome can be manipulated in vitro, where a guide polynucleotide, a Cms1 polypeptide, and any donor polynucleotides are incubated with the viral DNA sequence of interest outside of the cellular host.
[0113] V. Methods of Regulating Gene Expression The methods disclosed herein further encompass modifying a nucleotide sequence or modulating expression of a nucleotide sequence in a genomic host. The method may include introducing into the genomic host (a) at least one fusion protein or a nucleic acid encoding at least one fusion protein, wherein the fusion protein comprises a Cms1 polypeptide or a fragment or modification thereof and an effector domain, and (b) at least one guide RNA or DNA encoding the guide RNA, wherein the guide RNA guides the Cms1 polypeptide of the fusion protein to a target site in the target DNA, and the effector domain of the fusion protein modifies a chromosomal sequence or modulates expression of one or more genes near the target DNA sequence.
[0114] Fusion proteins comprising a Cms1 polypeptide or a fragment or variant thereof and an effector domain are described herein. Generally, the fusion proteins disclosed herein may further comprise at least one nuclear localization signal, plastid signal peptide, mitochondrial signal peptide, or signal peptide capable of transporting the protein to multiple intracellular locations. Nucleic acids encoding the fusion proteins are described herein. In some embodiments, the fusion protein can be introduced into a genomic host as an isolated protein (which may further comprise a cell penetration domain). Furthermore, the isolated fusion protein can be part of a protein-RNA complex that includes a guide RNA. In other embodiments, the fusion protein can be introduced into a genomic host as an RNA molecule (which may be capped and / or polyadenylated). In still other embodiments, the fusion protein can be introduced into a genomic host as a DNA molecule. For example, the fusion protein and guide RNA can be introduced into a genomic host as separate DNA molecules or as part of the same DNA molecule. Such a DNA molecule may be a plasmid vector.
[0115] In some embodiments, the method further comprises introducing at least one donor polynucleotide described elsewhere herein into the genomic host. Means for introducing molecules into genomic hosts, such as cells, as well as means for culturing cells (including cells containing organelles) are described herein.
[0116] In certain embodiments in which the effector domain of the fusion protein is a cleavage domain, the method can include introducing one fusion protein (or a nucleic acid encoding one fusion protein) and two guide RNAs (or DNA encoding two guide RNAs) into a genomic host. The two guide RNAs target the fusion protein to two different target sites in a chromosomal sequence, where the fusion protein dimerizes (e.g., forms a homodimer), allowing the two cleavage domains to introduce double-strand breaks in the target DNA sequence. In embodiments in which no donor polynucleotide is present, the double-strand break in the targeted DNA sequence can be repaired by the non-homologous end joining (NHEJ) repair process. Because NHEJ is error-prone, deletion of at least one nucleotide, insertion of at least one nucleotide, substitution of at least one nucleotide, or a combination thereof can occur during repair of the break. Thus, the targeted chromosomal sequence can be modified or inactivated. For example, a single nucleotide change (SNP) can result in an altered protein product, or a shift in the reading frame of a coding sequence can inactivate or "knock out" a sequence so that the protein product is not made. If any donor polynucleotide is present, the donor sequence in the donor polynucleotide can be exchanged or integrated with the target DNA sequence at the target site during repair of the double-strand break. If the donor sequence is flanked by upstream and downstream sequences that share substantial sequence identity with the upstream and downstream sequences, respectively, of the target site in the target DNA sequence, repair can result in the donor sequence being exchanged or integrated with the target DNA sequence at the target site. Alternatively, in embodiments where the donor sequence is flanked by compatible overhangs (or compatible overhangs generated in situ by a Cms1 polypeptide), the donor sequence can be directly ligated to the target DNA sequence cleaved by a non-homologous repair process during repair of the double-strand break. Exchange or integration of the donor sequence into the target DNA sequence modifies the target DNA sequence or introduces a foreign DNA sequence into the target DNA sequence.
[0117] In other embodiments in which the effector domain of the fusion protein is a cleavage domain, the method may include introducing two different fusion proteins (or nucleic acids encoding two different fusion proteins) and two guide RNAs (or DNA encoding the two guide RNAs) into the host genome. The fusion proteins may be different, as detailed elsewhere herein. Each guide RNA directs the fusion protein to a specific target site in the target DNA sequence. The fusion proteins can dimerize (e.g., form heterodimers) so that the two cleavage domains can introduce a double-strand break in the target DNA sequence. In embodiments in which any donor polynucleotide is not present, the resulting double-strand break may be repaired by a non-homologous repair process, and a deletion of at least one nucleotide, an insertion of at least one nucleotide, a substitution of at least one nucleotide, or a combination thereof, may occur during repair of the break. In embodiments in which any donor polynucleotide is present, the donor sequence in the donor polynucleotide can be exchanged or integrated with the chromosomal sequence during repair of the double-strand break by either a homology-based repair process (e.g., in embodiments in which the donor sequence is flanked by upstream and downstream sequences that have substantial sequence identity to the upstream and downstream sequences, respectively, of the target site in the chromosomal sequence) or a non-homologous repair process (e.g., in embodiments in which the donor sequence is flanked by compatible overhangs).
[0118] In certain embodiments in which the effector domain of the fusion protein is a transcription activation domain or a transcription repressor domain, the method can include introducing one fusion protein (or a nucleic acid encoding one fusion protein) and one guide RNA (or DNA encoding one guide RNA) into a genomic host. The guide RNA targets the fusion protein to a specific target DNA sequence, and the transcription activation domain or transcription repressor domain activates or represses, respectively, the expression of genes located near the target DNA sequence. That is, transcription can affect genes that are close to the targeted DNA sequence, or it can affect genes located further away from the targeted DNA sequence. It is well known in the art that gene transcription can be regulated by distantly located sequences located thousands of bases away from the transcription start site, or even on another chromosome (Harmston and Lenhard (2013) Nucleic Acids Res 41:7185-7199).
[0119] In another embodiment, in which the effector domain of the fusion protein is an epigenetic modification domain, the method may include introducing into a host genome one fusion protein (or a nucleic acid encoding one fusion protein) and one guide RNA (or DNA encoding one guide RNA). The guide RNA directs the fusion protein to a specific target DNA sequence, and the epigenetic modification domain modifies the structure of the target DNA sequence. Epigenetic modifications include acetylation, methylation of histone proteins, and / or methylation of nucleotides. In some instances, the structural modification of a chromosomal sequence results in a change in expression of the chromosomal sequence.
[0120] VI. Genetically Modified Organisms A.Eukaryotes Provided herein are eukaryotic organisms, eukaryotic cells, organelles, and plant embryos containing at least one nucleotide sequence modified using the Cms1 polypeptide-mediated or fusion protein-mediated processes described herein. Also provided are at least one DNA or RNA molecule encoding a Cms1 polypeptide or fusion protein that targets a chromosomal sequence or fusion protein of interest, at least one guide RNA, and optionally one or more eukaryotic organisms, eukaryotic cells, organelles, and plant embryos. Donor polynucleotides. The genetically modified eukaryotic organisms disclosed herein can be heterozygous for the modified nucleotide sequence or homozygous for the modified nucleotide sequence. Eukaryotic cells containing one or more genetic modifications in organelle DNA can be heteroplasmic or homoplasmic.
[0121] Modified chromosomal sequences in eukaryotes, eukaryotic cells, organelles, and plant embryos can be inactivated, have up- or down-regulated expression, produce altered protein products, or contain integrated sequences. Modified chromosomal sequences can be inactivated so that the sequence is not transcribed and / or no functional protein product is produced. Thus, genetically modified eukaryotes containing inactivated chromosomal sequences are sometimes referred to as "knockouts" or "conditional knockouts." Inactivated chromosomal sequences can contain deletion mutations (i.e., deletion of one or more nucleotides), insertion mutations (i.e., insertion of one or more nucleotides), or nonsense mutations (i.e., substitution of a single nucleotide for another nucleotide, such as the introduction of a stop codon). As a result of the mutation, the targeted chromosomal sequence is inactivated and no functional protein is produced. Inactivated chromosomal sequences do not contain exogenously introduced sequences. Also included herein are genetically modified eukaryotic organisms in which 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more chromosomal sequences have been inactivated.
[0122] Modified chromosomal sequences can also be modified to encode modified protein products. For example, genetically modified eukaryotes containing modified chromosomal sequences can contain targeted point mutations or other modifications to produce modified protein products. In one embodiment, the chromosomal sequence can be modified to contain at least one nucleotide change, resulting in the expressed protein containing a single altered amino acid residue (a missense mutation). In another embodiment, the chromosomal sequence can be modified to contain multiple missense mutations, resulting in multiple amino acid changes. Furthermore, the chromosomal sequence can be modified to contain a three-nucleotide deletion or insertion, resulting in the expressed protein containing a single amino acid deletion or insertion. The modified or mutant protein may have altered properties or activities, such as altered substrate specificity, altered enzymatic activity, or altered reaction kinetics, compared to the wild-type protein.
[0123] In some embodiments, a genetically modified eukaryote can include a nucleotide sequence integrated into at least one chromosome. Genetically modified eukaryotes containing an integrated sequence can be referred to as a "knock-in" or "conditional knock-in." The integrated nucleotide sequence can encode, for example, an orthologous protein, an endogenous protein, or a combination of both. In one embodiment, a sequence encoding an orthologous or endogenous protein can be integrated into a nuclear or organelle chromosomal sequence encoding a protein such that the chromosomal sequence is inactivated but the exogenous sequence is expressed. In such cases, the sequence encoding the orthologous or endogenous protein can be operably linked to a promoter control sequence. Alternatively, the sequence encoding the orthologous or endogenous protein can be integrated into a nuclear or organelle chromosomal sequence without affecting the expression of the chromosomal sequence. For example, the protein-encoding sequence can be integrated into a "safe harbor" locus. The present disclosure also encompasses genetically modified eukaryotes in which two, three, four, five, six, seven, eight, nine, or ten or more sequences are integrated into the genome. The sequence, including the protein-encoding sequence, is integrated into the genome. Any gene of interest disclosed herein can be integrated into or introduced into the chromosomal sequences of the eukaryotic nucleus or organelle. In certain embodiments, genes that increase plant growth or yield are integrated into the chromosome.
[0124] The chromosomally integrated protein-encoding sequence can encode a wild-type version of the protein of interest, or can encode a protein that contains at least one modification such that a modified version of the protein is produced. For example, a chromosomally integrated protein-encoding sequence associated with a disease or disorder can contain at least one modification such that the modified version of the protein produced causes or enhances the associated disorder. Alternatively, a chromosomally integrated protein-encoding sequence associated with a disease or disorder can contain at least one modification such that the modified version of the protein protects a eukaryotic organism or cell from developing the associated disease or disorder.
[0125] In certain embodiments, a genetically modified eukaryote can include at least one modified chromosomal sequence encoding a protein such that the expression pattern of the protein is altered. For example, regulatory regions controlling the expression of the protein, such as a promoter or transcription factor binding site, can be altered so that the protein is overexpressed, or so that the tissue-specific or temporal expression of the protein is altered, or a combination thereof. Alternatively, the expression pattern of the protein can be altered using a conditional knockout system. A non-limiting example of a conditional knockout system includes the Cre-lox recombination system, which contains the Cre recombinase enzyme, i.e., a site-specific DNA recombinase that can catalyze the recombination of nucleic acid sequences between specific sites (lox sites) in a nucleic acid molecule. Methods for using this system to generate temporal and tissue-specific expression are known in the art.
[0126] B. Prokaryotes Provided herein are prokaryotic organisms and cells comprising at least one nucleotide sequence modified using a Cms1 polypeptide-mediated or fusion protein-mediated process as described herein. Also provided are prokaryotic organisms and cells comprising at least one DNA or RNA molecule encoding a Cms1 polypeptide or fusion protein that targets a DNA sequence or fusion protein of interest, at least one guide RNA, and optionally one or more donor polynucleotides.
[0127] Modified DNA sequences in prokaryotes and prokaryotic cells can be inactivated, their expression can be up- or down-regulated, or they can be modified to produce an altered protein product or contain an integrated sequence. Modified DNA sequences can be inactivated so that the sequence is not transcribed and / or a functional protein product is not produced. Thus, genetically modified prokaryotes containing inactivated chromosomal sequences are sometimes referred to as "knockouts" or "conditional knockouts." Inactivated DNA sequences can include deletion mutations (i.e., deletion of one or more nucleotides), insertion mutations (i.e., insertion of one or more nucleotides), or nonsense mutations (i.e., substitution of a single nucleotide for another nucleotide, such as the introduction of a stop codon). As a result of the mutation, the target DNA sequence is inactivated and no functional protein is produced. Inactivated DNA sequences do not include exogenously introduced sequences. Also included herein are genetically modified prokaryotes in which two, three, four, five, six, seven, eight, nine, or ten or more DNA sequences are inactivated.
[0128] Modified DNA sequences can also be modified to encode modified protein products. For example, genetically modified prokaryotes containing modified DNA sequences can contain targeted point mutations or other modifications to produce modified protein products. In one embodiment, the DNA sequence can be modified to contain at least one nucleotide change, resulting in the expressed protein containing a single altered amino acid residue (a missense mutation). In another embodiment, the DNA sequence can be modified to contain multiple missense mutations, resulting in multiple amino acid changes. Furthermore, the DNA sequence can be modified to contain a triple nucleotide deletion or insertion, resulting in the expressed protein containing a single amino acid deletion or insertion. Modified or mutant proteins can have altered properties or activities, such as altered substrate specificity, altered enzymatic activity, or altered reaction kinetics, compared to the wild-type protein.
[0129] In some embodiments, a genetically modified prokaryote can contain at least one integrated nucleotide sequence. Genetically modified prokaryotes containing an integrated sequence can be referred to as a "knock-in" or "conditional knock-in." The integrated nucleotide sequence can encode, for example, an orthologous protein, an endogenous protein, or a combination of both. In one embodiment, a sequence encoding an orthologous or endogenous protein can be integrated into a prokaryotic DNA sequence encoding a protein such that the prokaryotic sequence is inactivated but the exogenous sequence is expressed. In such cases, the sequence encoding the orthologous or endogenous protein can be operably linked to a promoter control sequence. Alternatively, the sequence encoding the orthologous or endogenous protein can be integrated into the prokaryotic DNA sequence without affecting the expression of the native prokaryotic sequence. For example, the protein-encoding sequence can be integrated into a "safe harbor" locus. The present disclosure also encompasses genetically modified prokaryotes that are 2, 3, 4, 5, 6, 7, 8, 9, 10, or more sequences. The gene of interest disclosed herein can be introduced into a prokaryotic host by incorporating it into the genome or into a plasmid containing the protein-encoding sequence, or into a DNA sequence in the prokaryotic chromosome, a plasmid, or other extrachromosomal DNA.
[0130] The integrated protein-encoding sequence can encode a wild-type version of the protein of interest, or can encode a protein that includes at least one modification such that a modified version of the protein is produced. For example, an integrated sequence encoding a protein associated with a disease or disorder can include at least one modification such that the modified version of the protein produced causes or enhances the associated disorder. Alternatively, an integrated sequence encoding a protein associated with a disease or disorder can include at least one modification such that the modified version of the protein reduces infectivity in prokaryotes.
[0131] In certain embodiments, a genetically modified prokaryote can include at least one modified DNA sequence encoding a protein such that the expression pattern of the protein is altered. For example, regulatory regions controlling the expression of the protein, such as a promoter or transcription factor binding site, can be altered such that the protein is overexpressed, or such that the temporal expression of the protein is altered, or a combination thereof. Alternatively, the expression pattern of the protein can be altered using a conditional knockout system. A non-limiting example of a conditional knockout system includes the Cre-lox recombination system, which contains the Cre recombinase enzyme, i.e., a site-specific DNA recombinase that can catalyze the recombination of nucleic acid sequences between specific sites (lox sites) in a nucleic acid molecule. Methods for using this system to generate temporal expression are known in the art.
[0132] C. Virus Provided herein are viruses and viral genomes comprising at least one nucleotide sequence modified using the Cms1 polypeptide-mediated or fusion protein-mediated processes described herein. Also provided are viruses and viral genomes comprising at least one DNA or RNA molecule encoding a Cms1 polypeptide or fusion protein that targets a DNA sequence or fusion protein of interest, at least one guide RNA, and optionally one or more donor polynucleotides.
[0133] Modified DNA sequences in viruses and viral genomes can be inactivated, their expression can be up-regulated or down-regulated, or they can be modified to produce an altered protein product or contain integrated sequences. Modified DNA sequences can be inactivated so that the sequence is not transcribed and / or a functional protein product is not produced. Thus, genetically modified viruses containing inactivated chromosomal sequences are sometimes referred to as "knockouts" or "conditional knockouts." Inactivated DNA sequences can include deletion mutations (i.e., deletion of one or more nucleotides), insertion mutations (i.e., insertion of one or more nucleotides), or nonsense mutations (i.e., substitution of a single nucleotide for another, such that a stop codon is introduced). As a result of the mutation, the target DNA sequence is inactivated and no functional protein is produced. Inactivated DNA sequences do not include exogenously introduced sequences. Also included herein are genetically engineered viruses in which two, three, four, five, six, seven, eight, nine, or ten or more viral sequences are inactivated.
[0134] Modified DNA sequences can also be modified to encode modified protein products. For example, genetically modified viruses containing modified DNA sequences can contain targeted point mutations or other modifications to produce modified protein products. In one embodiment, the DNA sequence can be modified to contain at least one nucleotide change, resulting in the expressed protein containing a single altered amino acid residue (a missense mutation). In another embodiment, the DNA sequence can be modified to contain multiple missense mutations, resulting in multiple amino acid changes. Furthermore, the DNA sequence can be modified to contain a triple nucleotide deletion or insertion, resulting in the expressed protein containing a single amino acid deletion or insertion. Modified or mutant proteins can have altered properties or activities, such as altered substrate specificity, altered enzymatic activity, or altered reaction kinetics, compared to the wild-type protein.
[0135] In some embodiments, a genetically modified virus can contain at least one integrated nucleotide sequence. Genetically modified viruses containing an integrated sequence can be referred to as a "knock-in" or "conditional knock-in." The integrated nucleotide sequence can encode, for example, an orthologous protein, an endogenous protein, or a combination of both. In one embodiment, a sequence encoding an orthologous or endogenous protein can be integrated into a viral DNA sequence encoding a protein such that the viral sequence is inactivated but the exogenous sequence is expressed. In such cases, the sequence encoding the orthologous or endogenous protein can be operably linked to a promoter control sequence. Alternatively, the sequence encoding the orthologous or endogenous protein can be integrated into the viral DNA sequence without affecting the expression of the native viral sequence. For example, a protein-encoding sequence can be integrated into a "safe harbor" locus. The present disclosure also encompasses genetically modified viruses with two, three, four, five, six, seven, eight, nine, or ten or more sequences. Any gene of interest, as disclosed herein, can be introduced by integration into a DNA sequence of the viral genome.
[0136] The integrated protein-encoding sequence can encode a wild-type version of the protein of interest, or can encode a protein that contains at least one modification such that a modified version of the protein is produced. For example, an integrated sequence encoding a protein associated with a disease or disorder can contain at least one modification such that the modified version of the protein produced causes or enhances the associated disorder. Alternatively, an integrated sequence encoding a protein associated with a disease or disorder can contain at least one modification such that the modified version of the protein reduces viral infectivity.
[0137] In certain embodiments, the genetically engineered virus can include at least one modified DNA sequence encoding a protein such that the expression pattern of the protein is altered. For example, regulatory regions controlling the expression of the protein, such as a promoter or transcription factor binding site, can be altered so that the protein is overexpressed, or so that the temporal expression of the protein is altered, or a combination thereof. Alternatively, the expression pattern of the protein can be altered using a conditional knockout system. A non-limiting example of a conditional knockout system includes the Cre-lox recombination system, which contains the Cre recombinase enzyme, i.e., a site-specific DNA recombinase that can catalyze the recombination of nucleic acid sequences between specific sites (lox sites) in a nucleic acid molecule. Methods for using this system to generate temporal expression are known in the art.
[0138] All publications and patent applications mentioned in this specification are indicative of the level of skill of those skilled in the art to which this invention pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.
[0139] Although the foregoing invention has been described in some detail by way of illustration and example, for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims.
[0140] Embodiments of the present invention include the following. 1. A method for modifying a nucleotide sequence at a target site in the genome of a eukaryotic cell, comprising: (i) a DNA-targeting RNA, or a DNA polynucleotide encoding the DNA-targeting RNA, wherein the DNA-targeting RNA comprises: (a) a first segment comprising a nucleotide sequence complementary to a sequence of the target DNA; and (b) a second segment that interacts with a Cms1 polypeptide; and (ii) a Cms1 polypeptide, or a polynucleotide encoding a Cms1 polypeptide, wherein the Cms1 polypeptide comprises: (a) an RNA-binding portion that interacts with a DNA-targeting RNA; and (b) an active portion that exhibits site-specific enzymatic activity. This includes introducing
[0141] 2. A method for modifying a nucleotide sequence at a target site in the genome of a prokaryotic cell, comprising: (i) a DNA-targeting RNA, or a DNA polynucleotide encoding the DNA-targeting RNA, wherein the DNA-targeting RNA comprises: (a) a first segment comprising a nucleotide sequence complementary to a sequence of the target DNA; and (b) a second segment that interacts with a Cms1 polypeptide; and (ii) a Cms1 polypeptide, or a polynucleotide encoding a Cms1 polypeptide, wherein the Cms1 polypeptide comprises: (a) an RNA-binding portion that interacts with a DNA-targeting RNA; and (b) an active portion that exhibits site-specific enzymatic activity, wherein the prokaryotic cell is not a natural host for the gene encoding the Cms1 polypeptide. This includes introducing
[0142] 3. A method for modifying a nucleotide sequence at a target site in the genome of a plant cell, comprising: (i) a DNA-targeting RNA, or a DNA polynucleotide encoding the DNA-targeting RNA, wherein the DNA-targeting RNA comprises: (a) a first segment comprising a nucleotide sequence complementary to a sequence of the target DNA; and (b) a second segment that interacts with a Cms1 polypeptide; and (ii) a Cms1 polypeptide, or a polynucleotide encoding a Cms1 polypeptide, wherein the Cms1 polypeptide comprises: (a) an RNA-binding portion that interacts with a DNA-targeting RNA; and (b) an active portion that exhibits site-specific enzymatic activity. This includes introducing
[0143] The method of embodiment 3, further comprising: culturing the plant under conditions in which the Cms1 polypeptide is expressed and cleaves the nucleotide sequence at the target site to produce a modified nucleotide sequence; and selecting plants containing the modified nucleotide sequence.
[0144] 5. The method of any one of embodiments 1-4, wherein cleavage of the nucleotide sequence at the target site comprises a double-stranded break at or near the sequence to which the DNA-targeting RNA sequence is targeted.
[0145] 6. The method of embodiment 5, wherein the double-stranded break is a staggered double-stranded break.
[0146] 7. The method of embodiment 6, wherein the staggered double-stranded break creates a 5' overhang of 3 to 6 nucleotides.
[0147] 8. The method of any one of embodiments 1 to 7, wherein the DNA-targeting RNA is a guide RNA (gRNA).
[0148] 9. The method of any one of embodiments 1-8, wherein said modified nucleotide sequence comprises an insertion of heterologous DNA into the genome of the cell, a deletion of a nucleotide sequence from the genome of the cell, or a mutation of at least one nucleotide in the genome of the cell.
[0149] 10. The method of any one of embodiments 1 to 9, wherein the Cms1 polypeptide is selected from the group consisting of SEQ ID NOs: 20-23, 30-69, 208-211, and 222-254.
[0150] 11. The method of any one of embodiments 1 to 10, wherein the polynucleotide encoding the Cms1 polypeptide is selected from the group consisting of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287.
[0151] 12. The method of any one of embodiments 1 to 11, wherein the Cms1 polypeptide has at least 80% identity to one or more polypeptide sequences selected from the group consisting of SEQ ID NOs: 20-23, 30-69, 208-211, and 222-254.
[0152] 13. The method of any one of embodiments 1 to 12, wherein the polynucleotide encoding the Cms1 polypeptide has at least 70% identity to one or more nucleic acid sequences selected from the group consisting of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287.
[0153] 14. The method of any one of embodiments 1 to 13, wherein the Cms1 polypeptide forms a homodimer or a heterodimer.
[0154] 15. The method of embodiment 3, wherein the plant cell is derived from a monocotyledonous plant species.
[0155] 16. The method of embodiment 3, wherein the plant cell is derived from a dicotyledonous plant species.
[0156] 17. The method according to any one of embodiments 1 to 16, wherein expression of the Cms1 polypeptide is under the control of an inducible or constitutive promoter.
[0157] 18. The method of any one of embodiments 1 to 17, wherein expression of the Cms1 polypeptide is under the control of a cell type-specific or developmentally preferred promoter.
[0158] 19. The method of any one of embodiments 1-18, wherein the PAM sequence comprises 5'-TTN, where N can be any nucleotide.
[0159] 20. The method of embodiment 3, wherein the nucleotide sequence at the target site in the genome of the plant cell encodes an SBPase, FBPase, FBP aldolase, AGPase large subunit, AGPase small subunit, sucrose phosphate synthase, starch synthase, pyruvate phosphate dikinase encoding a PEP carboxylase, transketolase, Rubisco small subunit, or Rubisco activase protein, or encodes a transcription factor that regulates expression of one or more genes encoding an SBPase, FBPase, FBP aldolase, AGPase large subunit, AGPase small subunit, sucrose phosphate synthase, starch synthase, PEP carboxylase, pyruvate phosphate dikinase, transketolase, Rubisco small subunit, or Rubisco activase protein.
[0160] 21. The method of any one of embodiments 1-20, further comprising contacting the target site with a donor polynucleotide, which is a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a donor polynucleotide, and wherein the copy of the donor polynucleotide is integrated into the target DNA.
[0161] 22. The method of any one of embodiments 1 to 21, wherein the target DNA is modified such that nucleotides in the target DNA are deleted.
[0162] 23. The method of any one of embodiments 1 to 22, wherein the polynucleotide encoding the Cms1 polypeptide is codon optimized for expression in a plant cell.
[0163] 24. The method of any one of embodiments 1 to 23, wherein expression of the nucleotide sequence is increased or decreased.
[0164] 25. The method of any one of embodiments 1 to 24, wherein the polynucleotide encoding the Cms1 polypeptide is operably linked to a promoter that is constitutive, cell-specific, inducible, or activated by alternative splicing of a suicide exon.
[0165] 26. The method of any one of embodiments 1 to 25, wherein the Cms1 polypeptide comprises one or more mutations that reduce or eliminate the nuclease activity of the Cms1 polypeptide.
[0166] 27. The method described in embodiment 26, wherein the mutant Cms1 polypeptide contains a mutation at a position corresponding to position 701 or 922 of SmCms1 (SEQ ID NO: 10), or a position corresponding to position 848 or 1213 of SulfCms1 (SEQ ID NO: 11), when aligned to maximize identity.
[0167] 28. The method of embodiment 27, wherein the mutations at positions corresponding to positions 701 or 922 of SmCms1 (SEQ ID NO: 10) are D701A and E922A, respectively, or the mutations at positions corresponding to positions 848 and 1213 of SulfCms1 (SEQ ID NO: 11) are D848A and D1213A, respectively.
[0168] 29. The method of any one of embodiments 26 to 28, wherein the mutant Cms1 polypeptide is fused to a transcription activation domain.
[0169] 30. The method of embodiment 29, wherein the mutant Cms1 polypeptide is fused directly to the transcription activation domain or fused to the transcription activation domain using a linker.
[0170] 31. The method of any one of embodiments 26 to 28, wherein the mutant Cms1 polypeptide is fused to a transcriptional repressor domain.
[0171] 32. The method of embodiment 31, wherein the mutant Cms1 polypeptide is fused to a transcriptional repressor domain using a linker.
[0172] 33. The method of any one of embodiments 1 to 32, wherein the Cms1 polypeptide further comprises a nuclear localization signal.
[0173] 34. The method of embodiment 33, wherein the nuclear localization signal comprises SEQ ID NO: 1 or is encoded by SEQ ID NO: 2.
[0174] 35. The method of any one of embodiments 1 to 32, wherein the Cms1 polypeptide further comprises a chloroplast signal peptide.
[0175] 36. The method of any one of embodiments 1 to 32, wherein the Cms1 polypeptide further comprises a mitochondrial signal peptide.
[0176] 37. The method of any one of embodiments 1 to 32, wherein the Cms1 polypeptide further comprises a signal peptide that targets the Cms1 polypeptide to multiple intracellular locations.
[0177] 38. A nucleic acid molecule comprising a polynucleotide sequence encoding a Cms1 polypeptide, wherein the polynucleotide sequence is codon-optimized for expression in a plant cell.
[0178] 39. A nucleic acid molecule comprising a polynucleotide sequence encoding a Cms1 polypeptide, wherein said polynucleotide sequence is codon-optimized for expression in a eukaryotic cell.
[0179] 40. A nucleic acid molecule comprising a polynucleotide sequence encoding a Cms1 polypeptide, wherein the polynucleotide sequence is codon-optimized for expression in a prokaryotic cell, and the prokaryotic cell is not a natural host for the Cms1 polypeptide.
[0180] 41. The nucleic acid molecule of any one of embodiments 38-40, wherein the polynucleotide sequence is selected from the group consisting of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, or a fragment or variant thereof; or the polynucleotide sequence encodes a Cms1 polypeptide selected from the group consisting of SEQ ID NOs: 20-23, 30-69, 208-211, and 222-254, wherein the polynucleotide sequence encoding the Cms1 polypeptide is operably linked to a promoter that is heterologous to the polynucleotide sequence encoding the Cms1 polypeptide.
[0181] 42. The nucleic acid molecule of any one of embodiments 38-40, wherein the mutant polynucleotide sequence has at least 70% sequence identity to a polynucleotide sequence selected from the group consisting of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, or the polynucleotide sequence encodes a Cms1 polypeptide having at least 80% sequence identity to a polypeptide selected from the group consisting of SEQ ID NOs: 20-23, 30-69, 208-211, and 222-254, and wherein the polynucleotide sequence encoding the Cms1 polypeptide is operably linked to a promoter heterologous to the polynucleotide sequence encoding the Cms1 polypeptide.
[0182] 43. The nucleic acid molecule of any one of embodiments 38-40, wherein the Cms1 polypeptide comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 20-23, 30-69, 208-211, and 222-254, or a fragment or variant thereof.
[0183] 44. The nucleic acid molecule of embodiment 43, wherein the variant polypeptide sequence has at least 70% sequence identity to a polypeptide sequence selected from the group consisting of SEQ ID NOs: 20-23, 30-69, 208-211, and 222-254.
[0184] 45. The nucleic acid molecule of any one of embodiments 38-44, wherein said polynucleotide sequence encoding a Cms1 polypeptide is operably linked to a promoter that is active in plant cells.
[0185] 46. The nucleic acid molecule according to any one of embodiments 38 to 44, wherein the polynucleotide sequence encoding the Cms1 polypeptide is operably linked to a promoter that is active in eukaryotic cells.
[0186] 47. The nucleic acid molecule according to any one of embodiments 38 to 44, wherein the polynucleotide sequence encoding the Cms1 polypeptide is operably linked to a promoter active in a prokaryotic cell.
[0187] 48. A nucleic acid molecule described in any one of embodiments 38 to 44, wherein the polynucleotide sequence encoding the Cms1 polypeptide is operably linked to a constitutive promoter, an inducible promoter, a cell type-specific promoter, or a developmentally preferred promoter.
[0188] 49. The nucleic acid molecule of any one of embodiments 38 to 44, wherein the nucleic acid molecule encodes a fusion protein comprising the Cms1 polypeptide and an effector domain.
[0189] 50. The nucleic acid molecule of embodiment 49, wherein the effector domain is selected from the group consisting of a transcriptional activator, a transcriptional repressor, a nuclear localization signal, and a cell penetration signal.
[0190] 51. The nucleic acid molecule of embodiment 50, wherein the Cms1 polypeptide is mutated to reduce or eliminate nuclease activity.
[0191] 52. A nucleic acid molecule described in embodiment 51, wherein the mutant Cms1 polypeptide contains a mutation at a position corresponding to positions 701 or 922 of SmCms1 (sequence number 10) or positions 848 and 1213 of SulfCms1 (sequence number 11) when aligned for maximum identity.
[0192] 53. The nucleic acid molecule of any one of embodiments 49-52, wherein said Cms1 polypeptide is fused to said effector domain using a linker.
[0193] 54. The nucleic acid molecule of any one of embodiments 38 to 53, wherein the Cms1 polypeptide forms a dimer.
[0194] 55. A fusion protein encoded by the nucleic acid molecule of any one of embodiments 49 to 54.
[0195] 56. A Cms1 polypeptide encoded by the nucleic acid molecule of any one of embodiments 38 to 44.
[0196] 57. A Cms1 polypeptide mutated to reduce or eliminate nuclease activity.
[0197] 58. A Cms1 polypeptide described in embodiment 57, wherein the mutant Cms1 polypeptide comprises a mutation at a position corresponding to positions 701 or 922 of SmCms1 (sequence number 10), or at positions corresponding to positions 848 and 1213 of SulfCms1 (sequence number 11), when aligned for maximum identity.
[0198] 59. A plant cell, a eukaryotic cell, or a prokaryotic cell comprising a nucleic acid molecule of any one of embodiments 38 to 54.
[0199] 60. A plant, eukaryotic, or prokaryotic cell comprising the fusion protein or polypeptide of any one of embodiments 55 to 58.
[0200] 61. A plant cell produced by the method of any one of embodiments 1 and 3-37.
[0201] 62. A plant comprising the nucleic acid molecule of any one of embodiments 38 to 54.
[0202] 63. A plant comprising the fusion protein or polypeptide of any one of embodiments 55 to 58.
[0203] 64. A plant produced by the method of any one of embodiments 1 and 3-37.
[0204] 65. Seeds of the plant of any one of embodiments 62 to 64.
[0205] 66. The method of any one of embodiments 1 and 3 to 37, wherein the modified nucleotide sequence comprises an insertion of a polynucleotide encoding a protein that confers antibiotic or herbicide resistance to the transformed cell.
[0206] 67. The method of embodiment 66, wherein the polynucleotide encoding a protein that confers antibiotic or herbicide resistance comprises SEQ ID NO: 7 or encodes a protein comprising SEQ ID NO: 8.
[0207] 68. The method of any one of embodiments 3 to 37, wherein the target site in the genome of the plant cell comprises SEQ ID NO: 12 or shares at least 80% identity with a part or fragment of SEQ ID NO: 12.
[0208] 69. The method of any one of embodiments 1 to 37, wherein the DNA polynucleotide encoding the DNA-targeting RNA comprises SEQ ID NO: 15.
[0209] 70. The nucleic acid molecule of any one of embodiments 38 to 54, wherein said polynucleotide sequence encoding a Cms1 polypeptide further comprises a polynucleotide sequence encoding a nuclear localization signal.
[0210] 71. The nucleic acid molecule of embodiment 70, wherein the nuclear localization signal comprises SEQ ID NO: 1 or is encoded by SEQ ID NO: 2.
[0211] 72. The nucleic acid molecule according to any one of embodiments 38 to 54, wherein the polynucleotide sequence encoding the Cms1 polypeptide further comprises a polynucleotide sequence encoding a chloroplast signal peptide.
[0212] 73. The nucleic acid molecule according to any one of embodiments 38 to 54, wherein the polynucleotide sequence encoding the Cms1 polypeptide further comprises a polynucleotide sequence encoding a mitochondrial signal peptide.
[0213] 74. A nucleic acid molecule according to any one of embodiments 38 to 54, wherein the polynucleotide sequence encoding the Cms1 polypeptide further comprises a polynucleotide sequence encoding a signal peptide that targets the Cms1 polypeptide to multiple intracellular locations.
[0214] 75. The fusion protein of embodiment 55, wherein the fusion protein further comprises a nuclear localization signal, a chloroplast signal peptide, a mitochondrial signal peptide, or a signal peptide that targets the Cms1 polypeptide to multiple intracellular locations.
[0215] 76. The Cms1 polypeptide of any one of embodiments 56 to 58, wherein the Cms1 polypeptide further comprises a nuclear localization signal, a chloroplast signal peptide, a mitochondrial signal peptide, or a signal peptide that targets the Cms1 polypeptide to multiple subcellular locations.
[0216] 77. The method of any one of embodiments 1 to 37, wherein the Cms1 polypeptide comprises one or more sequence motifs selected from the group consisting of SEQ ID NOs: 177 to 186.
[0217] 78. The method of any one of embodiments 1 to 37, wherein the Cms1 polypeptide comprises one or more sequence motifs selected from the group consisting of SEQ ID NOs: 288-289 and 187-201.
[0218] 79. The method of any one of embodiments 1 to 37, wherein the Cms1 polypeptide comprises one or more sequence motifs selected from the group consisting of SEQ ID NOs: 290 to 296.
[0219] 80. The nucleic acid molecule of any one of embodiments 38-54, wherein the Cms1 polypeptide comprises one or more sequence motifs selected from the group consisting of SEQ ID NOs: 177-186.
[0220] 81. The nucleic acid molecule of any one of embodiments 38-54, wherein the Cms1 polypeptide comprises one or more sequence motifs selected from the group consisting of SEQ ID NOs: 288-289 and 187-201.
[0221] 82. The nucleic acid molecule of any one of embodiments 38 to 54, wherein the Cms1 polypeptide comprises one or more sequence motifs selected from the group consisting of SEQ ID NOs: 290 to 296.
[0222] The following examples are offered by way of illustration and not by way of limitation. [Example]
[0223] experiment Example 1 - Cloning of plant transformation constructs Constructs containing Cms1 are summarized in Table 1. Briefly, Cms1 genes were plant codon-optimized, de novo synthesized with GenScript (Piscataway, NJ), and amplified by PCR to include an N-terminal SV40 nuclear localization tag (SEQ ID NO: 2) in frame with the Cms1 coding sequence of interest and restriction enzyme sites for cloning. Using appropriate restriction enzyme sites, each Cms1 gene was cloned downstream of the 2x35s promoter (SEQ ID NO: 3). Note that SEQ ID NO: 16, which encodes the ADurb.160Cms1 protein (SEQ ID NO: 20), is derived from an organism that appears to encode glycine using the TGA codon rather than a stop codon, as used by most organisms in the universal genetic code. Thus, the native gene encoding the ADurb.160Cms1 protein (SEQ ID NO: 24) contains what appear to be multiple premature stop codons. However, analysis of this gene using the TGA codon, which encodes glycine, reveals a full-length open reading frame. Similarly, SEQ ID NOs: 82, 91, 92, 100, 105, 213, 255, 259, 266, 267, 268, 270, 271, 272, 273, 275, 276, 277, 279, 280, 284, 285, and 286 use the non-universal genetic code with the TGA codon encoding glycine.
[0224] A plasmid encoding a guide RNA targeting a region of the rice (Oryza sativa cv. Kitaake) gene, CAO1 (SEQ ID NO: 12), was synthesized with the guide RNA flanked by the rice U6 (OsU6) promoter (SEQ ID NO: 5) at its 5' end and the OsU6 terminator (SEQ ID NO: 6) at its 3' end. The guide RNA had the sequence of SEQ ID NO: 15. The guide RNA plasmids are summarized in Table 2.
[0225] Plasmid 131632, containing the repair donor cassette (SEQ ID NO:13), was designed with approximately 1,000 base pairs of homology upstream and downstream of the target site within the OsCAO1 gene. The repair donor cassette contained a maize ubiquitin promoter (SEQ ID NO:9) operably linked to a hygromycin resistance gene (SEQ ID NO:7, encoding SEQ ID NO:8) flanked at its 3' end by a cauliflower mosaic virus 35S polyA sequence (SEQ ID NO:4). Plasmid 131592 was designed similarly to plasmid 131632, but lacked homology arms upstream or downstream of the hygromycin cassette. Thus, plasmid 131592 contains nucleotides 1,001 to 4,302 of SEQ ID NO:13, containing a maize ubiquitin promoter (SEQ ID NO:9) operably linked to a hygromycin resistance gene (SEQ ID NO:7, encoding SEQ ID NO:8) flanked at its 3' end by a cauliflower mosaic virus 35S polyA sequence (SEQ ID NO:4).
[0226] [Table 1A]
[0227] [Table 1B]
[0228] [Table 1C] 1 Each Cms1 gene was fused in frame at its 5' end with the SV40 nuclear localization signal (SEQ ID NO: 2, encoding the amino acid sequence of SEQ ID NO: 1).
[0229] [Table 2]
[0230] Example 2 - Rice transformation Particle bombardment was used to introduce the Cms1 cassette, gRNA-containing plasmid, and repair donor cassette into rice cells. For bombardment, 2 mg of 0.6 μm gold particles were weighed and transferred to a sterile 1.5 mL tube. 500 mL of 100% ethanol was added, and the tube was sonicated for 10–15 seconds. After centrifugation, the ethanol was removed. Next, 1 mL of sterile double-distilled water was added to the tube containing the gold beads. The bead pellet was briefly vortexed and reconstituted by centrifugation, after which the water was removed from the tube. DNA was coated onto the beads in a sterile laminar flow hood. Table 3 shows the amount of DNA added to the beads. The Cms1 cassette-containing plasmid, gRNA-containing plasmid, and repair donor cassette were added to the beads, and sterile double-distilled water was added to bring the total volume to 50 μL. To this, 20 μL of spermidine (1 M) followed by 50 μL of CaCl2 (2.5 M) was added. The gold particles were allowed to pellet by gravity for several minutes and then centrifuged. The supernatant was removed, and 800 μL of 100% ethanol was added. Following brief sonication, the gold particles were allowed to pellet by gravity for 3–5 minutes, and the tube was then centrifuged to form a pellet. The supernatant was removed, and 30 μL of 100% ethanol was added to the tube. The DNA-coated gold particles were resuspended in this ethanol by vortexing, and 10 μL of the resuspended gold particles was added to each of three macrocarriers (Bio-Rad, Hercules, CA). The macrocarriers were allowed to air-dry in a laminar flow hood for 5–10 minutes to evaporate the ethanol.
[0231] [Table 3]
[0232] Rice callus tissue was used for bombardment. Rice callus was maintained in callus induction medium (CIM; 3.99 g / L N6 salts and vitamins, 0.3 g / L casein hydrolysate, 30 g / L sucrose, 2.8 g / L L-proline, 2 mg / L 2,4-D, 8 g / L agar, adjusted to pH 5.8) at 28 °C in the dark for 4–7 days before bombardment. Approximately 80–100 callus pieces (each 0.2–0.3 cm in size, weighing 1–1.5 g in total) were placed in the center of a Petri dish containing osmotic solid medium (CIM supplemented with 0.4 M sorbitol and 0.4 M mannitol) for 4 hours of infiltration pretreatment before particle bombardment. For bombardment, macrocarriers containing DNA-coated gold particles were assembled into a macrocarrier holder. The rupture disk (1,100 psi), stopping screen, and macrocarrier holder were assembled according to the manufacturer's instructions. Plates containing the rice callus to be bombarded were placed 6 cm below the stopping screen, and the callus pieces were bombarded after the vacuum chamber reached 25–28 in. Hg. After bombardment, the callus was placed on the infiltration medium for 16–20 hours, after which the callus pieces were transferred to selective medium (CIM supplemented with 50 mg / L hygromycin and 100 mg / L thymetentin). The plates were transferred to an incubator and kept in the dark at 28°C to initiate the recovery of transformed cells. Every two weeks, the callus was subcultured onto fresh selective medium. Hygromycin-resistant callus pieces began to appear after approximately 5–6 weeks on selective medium. Individual hygromycin-resistant callus pieces were transferred to new selective plates, and cells were allowed to divide and grow to generate enough tissue to sample for molecular analysis. Table 4 summarizes the DNA vector combinations used in these rice bombardment experiments.
[0233] [Table 4A]
[0234] [Table 4B]
[0235] Example 3 - Molecular analysis of rice Individual hygromycin-resistant callus pieces from each transformation experiment were transferred to new plates and allowed to grow to a size sufficient for sampling. A small amount of tissue was harvested from each piece of hygromycin-resistant rice callus, and DNA was extracted from these tissue samples for PCR, DNA sequencing, and T7 endonuclease (T7EI) analysis. PCR analysis was designed using primers that did not generate amplicons from either wild-type rice DNA or the repair donor plasmid alone, but instead had one primer binding site in the rice genome outside the homology arms and another in the insertion cassette, thus indicating an insertion event at the rice CAO1 locus.
[0236] Sanger sequencing and / or next-generation sequencing of the PCR amplicons generated from the above PCR analyses was performed to confirm that the PCR amplicons indeed represented insertions at the intended genomic loci and not actual experimental artifacts. Table 5 summarizes the results of these sequence analyses.
[0237] [Table 5A]
[0238] [Table 5B]
[0239] In addition to PCR and DNA sequence analysis, T7EI analysis was performed to detect the presence of small insertions and / or deletions at the CAO1 locus. T7EI analysis was performed as previously described (Begemann et al. (2017) Sci Reports 7:11606). For callus samples in which T7EI analysis indicated a possible insertion or deletion, DNA sequence analysis was performed to detect the presence of insertions and / or deletions at the CAO1 locus.
[0240] Example 4 - Regeneration of rice plants by genetic modification at the CAO1 locus The transformed rice callus is cultured in tissue culture medium to produce shoots. These shoots are then transferred to rooting medium, and the rooted plants are transferred to soil for greenhouse cultivation. DNA is extracted from the rooted plants and subjected to PCR and DNA sequence analysis. The T0 generation plants are grown to maturity and self-pollinated to produce T1 generation seeds. These T1 generation seeds are planted, and the resulting T1 generation plants are genotyped to identify homozygous, hemizygous, and null segregants. The plants are phenotyped to detect the yellow leaf phenotype associated with a homozygous knockout of the CAO1 gene (Lee et al. (2005) Plant Mol Biol 57:805-818).
[0241] Example 5 - Editing of selected genomic loci in maize (Zea mays) One or more gRNAs are designed to anneal to a desired site in the maize genome and enable interaction with one or more Cms1 proteins. These gRNAs are cloned into a vector so that they are operably linked to a promoter ("gRNA cassette") that is operable in plant cells. One or more genes encoding Cms1 proteins are cloned into a vector so that they are operably linked to a promoter ("Cms1 cassette") that functions in plant cells. The gRNA cassette and Cms1 cassette are cloned into a single vector, or alternatively, into two separate vectors suitable for plant transformation, and this or these vectors are then transformed into Agrobacterium cells. These cells are contacted with maize tissue suitable for transformation. Following incubation with the Agrobacterium cells, the maize cells are cultured in tissue culture medium suitable for regenerating intact plants. Maize plants are regenerated from cells contacted with Agrobacterium cells harboring vectors containing the Cms1 cassette and the gRNA cassette. Following regeneration of the corn plants, plant tissues are harvested and DNA is extracted from the tissues. Optionally, T7EI, PCR, and / or sequencing assays are performed to determine whether a DNA sequence change has occurred at the genomic location of interest.
[0242] Alternatively, the Cms1 cassette and gRNA cassette are introduced into corn cells using biolistic bombardment. A single vector containing the Cms1 cassette and gRNA cassette, or separate vectors containing the Cms1 cassette and gRNA cassette, respectively, are coated onto gold or titanium beads, which are then used to bombard corn tissue suitable for regeneration. Following bombardment, the corn tissue is transferred to tissue culture medium for corn plant regeneration. Following corn plant regeneration, the plant tissue is harvested and DNA is extracted from the tissue. If necessary, T7EI assays, PCR assays, and / or sequencing assays are performed to determine whether DNA sequence changes have occurred at the desired genomic location.
[0243] Example 6 - Computational analysis of Cms1 nuclease and other type V nucleases CRISPR nucleases are often classified by type; for example, Cas9 nuclease is classified as a type II nuclease, and Cpf1 nuclease is classified as type V (Koonin et al. (2017) Curr Opin Microbiol 37:67-78). Examination of the Cms1 nuclease protein sequence suggests that these nucleases should be grouped as type V nucleases, based in part on the presence of a RuvC domain and the absence of an HNH domain. Several groups of nucleases have been described in the scientific literature, including Cpf1 (also known as type VA), C2c1 (also known as type VB), C2c3 (also known as type VC), CasY (also known as type VD), and CasX (also known as type VE).
[0244] Muscle alignments of type V amino acid sequences typically fail to correctly align the catalytic residues of the RuvCI, RuvCII, and RuvCIII domains of these proteins. Given the central importance of these domains in protein function, correct placement of these residues is essential. For the amino acid sequences of Cms1 nuclease disclosed herein and in U.S. Patent No. 9,896,696 (SEQ ID NOS: 10, 11, 20-23, 30-69, and 154-156), the RuvCI, RuvCII, and RuvCIII catalytic residues have been identified. Three Cpf1 nucleases (SEQ ID NOs: 147-149), C2c1 nuclease (SEQ ID NOs: 150 and 157-164), C2c3 nuclease (SEQ ID NOs: 152 and 166-168) (Shmakov et al. (2016) Mol Cell 60:385-397), CasX nuclease (SEQ ID NOs: 151 and 165), and CasY nuclease (SEQ ID NOs: 153 and 169-173) (Burstein et al. (2017) Nature 542:237-241). Table 6 shows the catalytic residues of each of these domains, as well as the three amino acids immediately preceding and following the catalytic residue.
[0245] Table 6A
[0246] Table 6B
[0247] Table 6C
[0248] Sequence alignments and other computational analyses did not reveal any clear RuvCIII catalytic residues in CasY.5 or CasY.6. The putative catalytic residues for Unk64 and Unk69 are lysine and asparagine, respectively, while all other residues have an invariant aspartate residue at this position. For the remaining type V nucleases, the RuvC catalytic residues summarized in Table 6 were used to generate RuvC-anchored sequence alignments, with the catalytic residues acting as fixed anchors, using a previously described method (Begemann et al. (2017) BioRxiv doi:10.1101 / 192799). The resulting RuvC-anchored amino acid alignments were used to construct a phylogenetic tree (see Figure 1). As this diagram shows, Cms1 nuclease is in a separate clade from the other type V nucleases. Furthermore, there are at least three distinct groups of Cms1 nucleases that cluster together in this analysis (Table 6, these groups consist of MicroCms1 to Unk78Cms1, SulfCms1 to Unk71Cms1, and Unk40Cms1 to Unk76Cms1, respectively). This suggests the existence of at least three groups of Cms1 nucleases within this larger group. These three groups are labeled "Sm-type," "Sulf-type," and "Unk40-type," respectively, for the group of nucleases containing SmCms1 (SEQ ID NO: 10), SulfCms1 (SEQ ID NO: 11), and Unk40Cms1 (SEQ ID NO: 68), respectively.
[0249] Amino acid sequence alignments of Cms1 nucleases were examined to identify motifs within the protein sequences that are well conserved among these nucleases. It was observed that Cms1 nucleases were found in three well-separated clades in the phylogenetic tree shown in Figure 1. One of these clades contains SmCms1 (SEQ ID NO: 10), another contains SulfCms1 (SEQ ID NO: 11), and another contains Unk40Cms1 (SEQ ID NO: 68). Therefore, members of each of these clades were aligned separately to identify amino acid motifs that are partially and / or completely conserved among these nucleases. For the alignment of SmCms1-like nucleases, SEQ ID NOs: 10, 20, 23, 30, 32-34, 37-39, 41, 43, 44, 46-60, 67, 154-156, 208-211, 222, 223, 225, 228, 229, 232, 234, 236, 237, 241, 243, 245, 248, 250, 251, 253, and 254 were aligned. For the alignment of SulfCms1-like nucleases, SEQ ID NOs: 11, 21, 22, 31, 35, 36, 40, 42, 45, 61-66, 69, 227, 230, 231, 235, 239, 240, 242, 244, and 247 were aligned. For the alignment of Unk40-like nucleases, SEQ ID NOS: 68, 224, 226, 233, 238, 246, 249, and 252 were aligned. These alignments were performed using MUSCLE, and the resulting alignments were manually inspected to identify regions of conservation among all aligned proteins. The amino acid motifs shown in SEQ ID NOS: 177-186 were identified from the alignment of SmCms1-like nucleases. The amino acid motifs shown in SEQ ID NOS: 288-289 and 187-201 were identified from the alignment of SulfCms1-like nucleases. The amino acid motifs shown in SEQ ID NOS: 290-296 were identified from the alignment of Unk40Cms1-like nucleases. Weblogs were generated using sequence alignments and are illustrated in Figures 2–4 (SmCms1-like, SulfCms1-like, and Unk40Cms1-like sequence motifs, respectively, weblogo.berkeley.edu).Also shown is a schematic diagram showing the location of these conserved motifs on the SmCms1, SulfCms1, and Unk40Cms1 protein sequences.
[0250] Plant genome editing by Cms1 nuclease described herein suggests that TTTN or TTN PAM sites are accessible to many, if not all, Cms1 nucleases, consistent with several other descriptions of type V nucleases. Computational analysis was performed to identify BLAST hits corresponding to CRISPR spacers present in contigs encoding Cms1 nucleases. CRISPR spacers were identified using CRISPRfinder online (crispr.i2bc.paris-saclay.fr / Server / ); these spacers were used as seeds for BLAST searches against the metagenome. BLAST hits were identified against CRISPR spacers from contigs encoding AuxCms1, Unk15Cms1, Unk19Cms1, and Unk40Cms1 (SEQ ID NOS: 297-300, respectively). These BLAST hits are shown in SEQ ID NOS: 301-307 and are summarized in Table 7, along with the nucleotides surrounding the BLAST hits.
[0251] [Table 7]
[0252] In Table 7, underlined bases represent CRISPR spacer BLAST hits. Notably, the base immediately 5' of all BLAST hits shows either TTA or TTC, and seven of the 11 BLAST hits in this table show either TTTA or TTTC. These data, combined with the plant genome editing data described above, strongly suggest that at least these Cms1 nucleases (and likely most or all Cms1 nucleases) can access target sites downstream from at least the TTM PAM site, with a preference for TTTM PAM sites. Notably, these types of computationally identified PAM sites take into account not only nuclease PAM requirements but also CRISPR spacer acquisition machinery requirements, so it is possible that nucleases can access a wider range of PAM sites than those computationally identified.
Claims
1. 1. A method for modifying a nucleotide sequence at a target site in the genome of an animal, fungal, or prokaryotic cell, comprising administering to said cell (i) a guide RNA (gRNA) or a DNA polynucleotide encoding the gRNA, wherein the gRNA comprises: (a) a first segment comprising a nucleotide sequence complementary to a nucleotide sequence at a target site; and (b) a second segment that interacts with a Cmsl polypeptide; and (ii) A Cms1 polypeptide or a polynucleotide encoding a Cms1 polypeptide, wherein the Cms1 polypeptide has an amino acid sequence that has at least 90% sequence identity with any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254, and has endonuclease activity. The method as described above, comprising introducing:
2. culturing the cells under conditions in which the Cmsl polypeptide is expressed and cleaves the nucleotide sequence at the target site to produce a modified nucleotide sequence; and Selecting cells containing the modified nucleotide sequence The method of claim 1, further comprising:
3. 3. The method of claim 2, wherein the modified nucleotide sequence comprises an insertion of heterologous DNA into the genome of the cell, a deletion of a nucleotide sequence from the genome of the cell, or a mutation of at least one nucleotide in the genome of the cell.
4. 3. The method of claim 2, wherein the modified nucleotide sequence comprises the insertion of a polynucleotide encoding a protein that confers antibiotic resistance to the transformed cell.
5. The method according to any one of claims 1 to 4, wherein the genome of the cell is a nuclear genome or a mitochondrial genome.
6. The method of any one of claims 1 to 5, wherein the polynucleotide encoding the Cms1 polypeptide is codon-optimized for expression in a cell.
7. The method of any one of claims 1 to 6, wherein the gRNA is a DNA-targeting RNA.
8. The method according to any one of claims 1 to 7, wherein the cell is an animal cell, and the animal cell is a mammalian cell.
9. 1. A nucleic acid molecule comprising a polynucleotide sequence, wherein the polynucleotide sequence (i) encodes a Cms1 polypeptide having at least 90% identity to any one of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, and wherein the Cms1 polypeptide has endonuclease activity; or (ii) encodes a Cms1 polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254, and wherein the Cms1 polypeptide has endonuclease activity, wherein the polynucleotide sequence is heterologous to the polynucleotide sequence and is operably linked to a promoter active in an animal cell, a fungal cell, or a prokaryotic cell.
10. The nucleic acid molecule of claim 9, wherein the promoter is active in an animal cell, and the animal cell is a mammalian cell.
11. 11. The nucleic acid molecule of claim 9 or 10, wherein the polynucleotide sequence is set forth as any one of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, or wherein the polynucleotide sequence encodes a Cms1 polypeptide comprising the amino acid sequence set forth as any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254.
12. 11. The nucleic acid molecule of claim 9 or 10, wherein the Cms1 polypeptide is mutated to reduce or eliminate nuclease activity.
13. The nucleic acid molecule of claim 12, wherein the mutant Cms1 polypeptide contains a mutation at a position corresponding to positions 701 or 922 of SmCms1 (SEQ ID NO: 10) or positions 848 and 1213 of SulfCms1 (SEQ ID NO: 11) when aligned for maximum identity.
14. The nucleic acid molecule of any one of claims 9 to 13, wherein the polynucleotide sequence encoding the Cms1 polypeptide is codon-optimized for expression in a cell.
15. (a) a Cms1 polypeptide, wherein the Cms1 polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254, and has endonuclease activity, or the Cms1 polypeptide is encoded by a polynucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, and has endonuclease activity; or (b) a nucleic acid molecule comprising a polynucleotide sequence, wherein the polynucleotide sequence is (i) encoding a Cms1 polypeptide having at least 90% identity to any one of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, and having an endonuclease; or (ii) A nucleic acid molecule comprising a polynucleotide sequence that encodes a Cms1 polypeptide comprising an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254, and that has endonuclease activity. An animal cell, a fungal cell, or a prokaryotic cell comprising:
16. 16. The cell of claim 15, wherein the Cms1 polypeptide of (a) comprises the amino acid sequence set forth as any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254, or the polynucleotide sequence of (a) is set forth as any one of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287.
17. The cell of claim 15, wherein the polynucleotide sequence of (b) is set forth as any one of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, or wherein the polynucleotide sequence of (b) encodes a Cms1 polypeptide comprising an amino acid sequence set forth as any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254.
18. The cell of claim 15, wherein the Cms1 polypeptide is mutated to reduce or eliminate nuclease activity.
19. The cell described in claim 18, wherein the mutant Cms1 polypeptide contains a mutation at a position corresponding to positions 701 or 922 of SmCms1 (SEQ ID NO: 10) or positions 848 and 1213 of SulfCms1 (SEQ ID NO: 11) when aligned for maximum identity.
20. (i) a Cms1 polypeptide, wherein the Cms1 polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254, and has endonuclease activity, or the Cms1 polypeptide is encoded by a polynucleotide sequence having at least 90% sequence identity to any one of SEQ ID NOs: 16-19, 24-27, 70-146, 174-176, 212-215, and 255-287, and has endonuclease activity; or (ii) a heterologous polypeptide An animal cell, a fungal cell, or a prokaryotic cell comprising a fusion protein comprising:
21. 21. The cell of claim 20, wherein the heterologous polypeptide comprises an effector domain selected from the group consisting of a cleavage domain, an epigenetic modification domain, a transcriptional activation domain, and a transcriptional repressor domain.
22. 22. The cell of claim 20 or 21, wherein the fusion protein further comprises a nuclear localization signal, a plastid signal peptide, a mitochondrial signal peptide, a signal peptide capable of transporting the protein to multiple intracellular locations, a cell penetration domain, or a marker domain.
23. (i) a Cms1 polypeptide, wherein the Cms1 polypeptide is (a) an RNA-binding moiety that interacts with a guide RNA (gRNA); and (b) an active moiety exhibiting site-specific enzymatic activity; wherein the Cms1 polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254, and has endonuclease activity; and (ii) a gRNA, wherein the gRNA is (a) a first segment comprising a nucleotide sequence that is complementary to a target sequence in the genome of an animal cell, a fungal cell, or a prokaryotic cell, located near 3′ of a protospacer adjacent motif (PAM) site; and (b) a second segment that interacts with the Cmsl polypeptide. gRNA comprising A ribonucleoprotein complex containing
24. 24. The ribonucleoprotein complex of claim 23, wherein the Cms1 polypeptide comprises an amino acid sequence set forth as any one of SEQ ID NOs: 10, 11, 20-23, 30-69, 154-156, and 222-254.
25. 25. The ribonucleoprotein complex of claim 23 or 24, wherein the gRNA is a DNA-targeting RNA.
Citation Information
Patent Citations
Compositions and methods for modifying genomes
WO2017141173A2