MÉTODOS PARA MODIFICAR GENOMAS, MOLÉCULA DE ÁCIDO NUCLEICO, POLIPEPTÍDEO CSM1 e PROTEÍNA DE FUSÃO
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2017-02-15
- Publication Date
- 2026-08-04
Smart Images

Figure 00000177_0000 
Figure 00000178_0000 
Figure 00000180_0000
Description
1 / 172 Descriptive Report of the Invention Patent for: "METHODS FOR MODIFYING GENOMES, NUCLEIC ACID MOLECULE, CSM1 POLYPEPTIDE and FUSION PROTEIN." Divided from the Invention Patent application BR 11 2018 016408 9, filed on February 15, 2017. FIELD OF THE INVENTION
[001] The present invention relates to compositions and methods for editing genomic sequences at pre-selected locations and for modulating gene expression. Reference to a sequence listing submitted as a text file via the EFS network.
[002] The official copy of the sequence listing is submitted simultaneously with the description as a text file via the EFS Network, in accordance with the American Standard Code for Information Interchange (ASCII), with a filename B88552_1060WO_0057_1_Seq_List.txt, a creation date of February 14, 2017, and a size of 1.62 MB. The sequence listing deposited via the EFS Network forms part of the description and is incorporated herein in its entirety by reference herein. FUNDAMENTALS OF THE INVENTION
[003] Genomic DNA modification is of immense importance for basic and applied research. Genomic modifications have the potential to elucidate and, in some cases, cure the causes of disease and provide traits. Petition 870240102758, dated 02 / 12 / 2024, page 24 / 196 2 / 172 desirable in cells and / or individuals comprising such modifications. Genomic modification can include, for example, plant, animal, fungal and / or prokaryotic genomic modification. One area in which genomic modification is practiced is in the modification of plant genomic DNA.
[004] Modification of plant genomic DNA is of immense importance for basic and applied plant research. Transgenic plants with stably modified genomic DNA can have new traits, such as herbicide tolerance, insect resistance, and / or the accumulation of valuable proteins, including pharmaceutical proteins and industrial enzymes transmitted to them. The expression of native plant genes can be positively and negatively regulated or otherwise altered (e.g., by altering the tissue(s) in which native plant genes are expressed), their expression can be completely abolished, DNA sequences can be altered (e.g., through mutations, insertions, or point deletions), or new non-native genes can be inserted into a plant's genome to confer new traits to the plant.
[005] The most common methods for modifying plant genomic DNA tend to modify DNA at random sites within the genome. Such methods include, for example, Agrobacterium-mediated plant transformation. Petition 870240102758, dated 02 / 12 / 2024, page 25 / 196 3 / 172 and biolistic transformation, also referred to as particle bombardment. In many cases, however, it is desirable to modify genomic DNA at a predetermined target site in the genome of a plant of interest, for example, to avoid breaking native plant genes or to insert a transgenic cassette at a genomic location known to provide robust gene expression. Only recently have technologies for targeted modification of plant genomic DNA become available. Such technologies rely on the creation of a double-strand break (DSB) at the desired site. This DSB triggers the recruitment of the plant's native DNA repair machinery to the DSB. The DNA repair machinery can be leveraged to insert heterologous DNA at a predetermined site, to eliminate native plant genomic DNA, or to produce point mutations, insertions, or deletions at a desired site. SUMMARY OF THE INVENTION
[006] Compositions and methods for modifying genomic DNA sequences are provided. As used herein, genomic DNA refers to linear and / or chromosomal DNA and / or plasmid or other extrachromosomal DNA sequences present in the cell or cells of interest. The methods produce double-strand breaks (DSBs) at predetermined target sites in a genomic DNA sequence, resulting Petition 870240102758, dated 02 / 12 / 2024, page 26 / 196 4 / 172 in mutation, insertion and / or deletion of DNA sequences at target site(s) in a genome. The compositions comprise DNA constructs comprising nucleotide sequences encoding a Cpf1 or Csm1 protein operationally linked to a promoter that is operable in the cells of interest. The DNA constructs can be used to target genomic DNA modification at predetermined genomic sites. Methods for using these DNA constructs to modify genomic DNA sequences are described herein. Modified plants, plant cells, plant parts and seeds are also included. Compositions and methods for modulating gene expression are also provided. The methods target protein(s) at predetermined sites in a genome to effect positive or negative regulation of a gene or genes whose expression is regulated by the target site in the genome.The compositions comprise DNA constructs comprising nucleotide sequences encoding a modified Cpf1 or Csm1 protein with decreased or abolished nuclease activity, optionally fused with a transcription activation or repression domain. The methods for using these DNA constructs to modify gene expression are described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[007] Figure 1 shows a schematic representation of the insertion of a resistance gene cassette. Petition 870240102758, dated 02 / 12 / 2024, page 27 / 196 5 / 172 hygromycin at the CAO1 genomic site of rice. The star indicates the site of the intended Cpf1-mediated double-strand break in wild-type DNA. Dashed lines indicate homology between the repair donor cassette and wild-type DNA. Small arrows indicate the primer binding sites for the PCR reactions used to verify insertion at the intended genomic site. 35S Term., Term. CaMV 35S; hph, hygromycin resistance gene; ZmUbi, maize ubiquitin promoter.
[008] Figure 2 shows sequence data obtained from rice calluses generated during Experiment 1. Figure 2A shows the results of an hph cassette insertion at the CAO1 site. The PAM sequence is divided into boxes, and the guide RNA-directed sequence is underlined. The ellipse indicates that a large insertion occurred, but the complete sequence data is not shown here. Figures 2B, 2C, and 2D show data obtained from rice calluses in which an FnCpf1-mediated elimination event occurred in Experiment 1 (Table 7). In Figures 2B and 2C, the bands represent callus parts #1-16, from left to right, followed by a molecular weight marker line. Figure 2B shows the PCR amplification of the FnCpf1 gene cassette, indicating the insertion of this cassette into the rice genome in callus fragments 1, 2, 4, 6, 7, and 15. Figure 2C shows the results of a T7EI assay with Petition 870240102758, dated 02 / 12 / 2024, p. 28 / 196 6 / 172 DNA extracted from these same callus fragments, with the double-band pattern for callus #15 indicating a possible insertion or deletion. Similar results from the T7EI assay were obtained for additional calluses in a repetition of Experiment 01, which resulted in the production of callus fragments 01-20, 01-21, 01-30, and 01-31. Figure 2D shows an alignment of the sequence data obtained from callus #15 (01-15), along with the sequence data from callus fragments 01-20, 01-21, 01-30, and 01-31. The PAM sequence is divided into boxes, and the guide RNA-directed sequence is underlined.
[009] Figure 3 shows sequence data from Experiments 31, 46, 80, 81, 91 and 93, verifying Cpf1-mediated and Csm1-mediated indels at the rice genomic locus CAO1. Figure 3A shows an alignment of the CAO1 site of wild-type rice with sequence data from callus fragment #21 of Experiment 31 (31-21), callus fragment #33 of Experiment 80 (80-33), callus fragments 9, 30, and 46 of Experiment 81 (81-09, 81-30, and 81-46, respectively), callus fragment #47 of Experiment 93 (93-47), callus fragment #4 of Experiment 91 (91-04), callus fragments #112 and 141 of Experiment 97 (97-112 and 97-141), and callus fragments #4 and 11 of Experiment 119 (119-04 and 119-11). Figure 3B shows the sequence data for callus fragments 46-38, 46-77, 46-86, 46-88, and 46-90 from experiment 46. In both 4A and 4B, the website Petition 870240102758, dated 02 / 12 / 2024, p. 29 / 196 7 / 172 of PAM is divided into boxes, and the region targeted by the guide RNA is underlined.
[0010] Figure 4 shows an overview of the unexpected recombination events recovered from Experiments 70 and 75. Figure 4A shows a schematic overview of a portion of plasmid 131633 including the homologous regions of the 35S terminator and the downstream branch that led to the recombination events recovered from Experiment 70. Homology regions that appear to have mediated the unintended HDR events are underlined. Figure 4B shows the sequencing data from callus chunk 70-15. WT, wild-type sequence; GE70, sequence of callus chunks 70-15; 131633_upstream, 35S terminator sequence and upstream branch, 131633_downstream, downstream branch sequence of plasmid 131633. Figure 4C shows a schematic overview of a portion of plasmid 131633 including the homologous regions of the 35S terminator and the downstream branch that led to the recombination events recovered from experiment 75.Regions of homology that appear to have mediated the unintended HDR events are underlined. Figure 4D shows the sequencing data for callus fragment 75-46. WT, wild-type sequence; GE75, callus fragment 75-46; 131633_up, upstream branch and Term 35S sequence of plasmid 131633; 131633_downstream, downstream branch sequence of plasmid 131633. Term 35S,. Petition 870240102758, dated 02 / 12 / 2024, p. 30 / 196 8 / 172 terminator 35S CaMV; hph, coding region for hygromycin phosphotransferase; pZmUbi, maize ubiquitin promoter. In Figures 4B and 4D, the PAM site is divided into boxes.
[0011] Figure 5 shows the sequence of the upstream region of callus fragment #46-161 from experiment 46 (Table 7). The PAM site is divided into boxes, showing the expected mutation of this site in the transformed rice callus, and the sequence data indicate the successful insertion of vector 131633 into the CAO1 genomic locus of rice. DETAILED DESCRIPTION OF THE INVENTION
[0012] The methods and compositions provided herein for controlling gene expression involving sequence targeting, such as genome perturbation or gene editing, which relate to the CRISPR-Cpf or CRISPR-Csm system and components thereof. In certain embodiments, the CRISPR enzyme is a Cpf enzyme, for example, a Cpf1 ortholog. In certain embodiments, the CRISPR enzyme is a Csm enzyme, for example, a Csm1 ortholog. The methods and compositions include nucleic acids for ligating target DNA sequences. This is advantageous because nucleic acids are much easier and less expensive to produce than, for example, peptides, and the specificity can be varied according to the length of the extension where homology is sought. Positioning Petition 870240102758, dated 02 / 12 / 2024, p. 31 / 196 9 / 172 of the 3D multi-finger complex, for example, is not necessary.
[0013] Nucleic acids encoding the Cpf1 and Csm1 polypeptides are also provided, as well as methods for using Cpf1 and Csm1 polypeptides to modify chromosomal (i.e., genomic) or organelle DNA sequences of host cells, including plant cells. The Cpf1 polypeptides interact with specific guide RNAs (gRNAs), which direct the Cpf1 or Csm1 endonuclease to a specific target site, at which site the Cpf1 or Csm1 endonuclease introduces a double-strand break that can be repaired by a DNA repair process whereby the DNA sequence is modified. Because specificity is provided by the guide RNA, the Cpf1 or Csm1 polypeptide is universal and can be used with different guide RNAs to target different genomic sequences. Cpf1 and Csm1 endonucleases have certain advantages over Cas nucleases (e.g., Cas9) traditionally used with CRISPR arrays.For example, CRISPR arrays associated with Cpf1 are processed into mature crRNAs without the requirement of an additional transactivated crRNA (tracrRNA). Furthermore, Cpfl-crRNA complexes can cleave target DNA preceded by a motif adjacent to the short protospacer (PAM) that is often T-rich, in contrast to the G-rich PAM following the target DNA for many Cas9 systems. Additionally, Cppl can introduce a... Petition 870240102758, dated 02 / 12 / 2024, p. 32 / 196 10 / 172 double-strand DNA break with a 5' overhang of 4 or 5 nucleotides (nt). Without being limited by theory, it is likely that Csm1 proteins similarly process their CRISPR arrays into mature crRNAs without the need for an additional transactivated crRNA (tracrRNA) and produce blind cuts instead of staggered cuts. The methods described here can be used to target and modify specific chromosomal sequences and / or introduce exogenous sequences at target locations in the genome of plant cells or plant embryos. The methods can also be used to introduce sequences or modify regions within organelles (e.g., chloroplasts and / or mitochondria). Furthermore, the targeting is specific with limited target effects. I. Cpf1 and Csm1 Endonucleases
[0014] Cpf1 and Csm1 endonucleases, and fragments and variants thereof, are provided herein for use in genome modification, including plant genomes. As used herein, the term Cpf1 endonucleases or Cpf1 polypeptides refers to homologs and orthologs of the Cpf1 polypeptides described in Zetsche et al. (2015) Cell 163: 759-771 and the Cpf1 polypeptides described in US Patent Application 2016 / 0208243, and fragments and variants thereof. Examples of Cpf1 polypeptides are presented in SEQ ID Nos: 3, 6, 9, 12, 15, 18, 20, 23, 106-133, 135-146, 148-158, 161-173 and Petition 870240102758, dated 02 / 12 / 2024, page 33 / 196 11 / 172 231-236. As used herein, the term Csml endonucleases or Csm1 polypeptides refers to homologs and orthologs of SEQ ID NOs: 134, 147, 159, 160, and 230. Typically, Cpf1 and Csm1 endonucleases can act without the use of tracrRNAs and can introduce a stepped double-strand break in DNA. In general, Cpf1 and Csm1 polypeptides comprise at least one RNA recognition and / or RNA binding domain. The RNA recognition and / or RNA binding domains interact with guide RNAs. The Cpf1 and Csm1 polypeptides may also comprise nuclease domains (i.e., DNase or RNase domains), DNA-binding domains, helicase domains, RNase domains, protein-protein interaction domains, dimerization domains, as well as other domains.In specific embodiments, a Cpf1 or Csm1 polypeptide, or a polynucleotide encoding a Cpf1 or Csm1 polypeptide, comprises: an RNA-binding moiety that interacts with RNA directed to DNA and an activity moiety that exhibits site-directed enzymatic activity, such as a RuvC endonuclease domain.
[0015] Cpf1 or Csm1 polypeptides may be wild-type Cpf1 or Csm1 polypeptides, modified Cpf1 or Csm1 polypeptides, or a fragment of a wild-type or modified Cpf1 or Csm1 polypeptide. The Cpf1 or Csm1 polypeptide may be modified to increase affinity. Petition 870240102758, dated 02 / 12 / 2024, page 34 / 196 12 / 172 and / or nucleic acid binding specificity, alter an enzymatic activity and / or change another protein property. For example, the nuclease domains (i.e., DNase, RNase) of the Cppl or Csml polypeptide can be modified, eliminated, or inactivated. Alternatively, the Cpf1 or Csm1 polypeptide can be truncated to remove domains that are not essential for protein function. In specific embodiments, the Cpf1 or Csm1 polypeptide forms a homodimer or a heterodimer.
[0016] In some embodiments, the Cpf1 or Csm1 polypeptide may be derived from a wild-type Cpf1 or Csm1 polypeptide or a fragment thereof. In other embodiments, the Cpf1 or Csm1 polypeptide may be derived from a modified Cpf1 or Csm1 polypeptide. For example, the amino acid sequence of the Cpf1 or Csm1 polypeptide may be modified to alter one or more properties (e.g., nuclease activity, affinity, stability, etc.) of the protein. Alternatively, domains of the Cpf1 or Csm1 polypeptide not involved in RNA-guided cleavage may be eliminated from the protein, so that the modified Cpf1 or Csm1 polypeptide is smaller than the wild-type Cpf1 or Csm1 polypeptide.
[0017] In general, a Cpf1 or Csm1 polypeptide comprises at least one nuclease (i.e., DNAase) domain, but need not contain an HNH domain such as Petition 870240102758, dated 02 / 12 / 2024, page 35 / 196 13 / 172 that found in Cas9 proteins. For example, a Cpf1 or Csm1 polypeptide may comprise a RuvC-type nuclease domain. In some embodiments, the Cpf1 or Csm1 polypeptide may be modified to inactivate the nuclease domain so that it is no longer functional. In some embodiments where one of the nuclease domains is inactive, the Cpf1 or Csm1 polypeptide does not cleave double-stranded DNA. In specific embodiments, the mutated Cpf1 or Csm1 polypeptide comprises a mutation at a position corresponding to positions 917 or 1006 of FnCpf1 (SEQ ID NO: 3) or positions 701 or 922 of SmCsm1 (SEQ ID NO: 160) when aligned to maximum identity that reduces or eliminates nuclease activity.For example, a conversion of aspartate to alanine (D917A) and glutamate to alanine (E1006A) in a RuvC-like domain completely inactivated the DNA cleavage activity of FnCpf1 (SEQ ID NO: 3), while aspartate to alanine (D1255A) significantly reduced cleavage activity (Zetsche et al. (2015) Cell 163: 759-771). The nuclease domain can be modified using well-known methods such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, as well as other methods known in the art. Cpf1 or Csm1 proteins with inactivated nuclease domains (dCpf1 or dCsm1 proteins) can be used to modulate gene expression without modifying DNA sequences. In certain embodiments, a. Petition 870240102758, dated 02 / 12 / 2024, page 36 / 196 14 / 172 The dCpf1 or dCsm1 protein can be targeted to particular regions of a genome, such as promoters for a gene or genes of interest, through the use of appropriate gRNAs. The dCpf1 or dCsm1 protein can bind to the desired DNA region and interfere with RNA polymerase binding to this DNA region and / or with the binding of transcription factors to this DNA region. This technique can be used to positively or negatively regulate the expression of one or more genes of interest. In certain other embodiments, the dCpf1 or dCsm1 protein can be fused to a repressor domain to further negatively regulate the expression of a gene or genes whose expression is regulated by interactions of RNA polymerase, transcription factors, or other transcriptional regulators with the chromosomal DNA region targeted by the gRNA.In certain other embodiments, the dCpf1 or dCsm1 protein can be fused to an activation domain to effect positive regulation of a gene or genes whose expression is regulated by interactions of RNA polymerase, transcription factors, or other transcriptional regulators with the region of chromosomal DNA targeted by the gRNA.
[0018] The Cpf1 and Csm1 polypeptides described herein may also comprise at least one nuclear localization signal (NLS). In general, an NLS comprises an extension of basic amino acids. Nuclear localization signals are Petition 870240102758, dated 02 / 12 / 2024, page 37 / 196 15 / 172 known in the art (see, for example, Lange et al., J. Biol. Chem. (2007) 282: 5101-5105). The NLS may be located at the N-terminus, the C-terminus, or an internal location of the Cpf1 or Csm1 polypeptide. In some embodiments, the Cpf1 or Csm1 polypeptide may further comprise at least one cell-penetrating domain. The cell-penetrating domain may be located at the N-terminus, the C-terminus, or an internal location of the protein.
[0019] The Cpf1 or Csm1 polypeptide described herein may further comprise at least one plastid-directed signal peptide, at least one mitochondrial-directed signal peptide, or a Cpf1 or Csm1 polypeptide signal peptide directed to both plastids and mitochondria. The localization signals of the dual-targeting signal peptide, to plastids, to mitochondria are known in the art (see, for example, Nassoury and Morse (2005) Biochim Biophys Acta 1743: 5-19; Kunze and Berger (2015) Front Physiol dx.doi.org / 10.3389 / fphys.2015.00259; Herrmann and Neupert (2003) IUBMB Life 55:219-225; Soll Opin (2002) Curr Opin Plant Biol 5:529-535; Carrie and Small (2013) Biochim Biophys Acta 1833:253-259; Curr Opin Plant Biol 6: 589-595; Peeters and Small (2001) Biochim Biophys Acta 1541: 54-63; Murcha et al. Petition 870240102758, dated 02 / 12 / 2024, page 38 / 196 16 / 172 38:311-338). The dual-targeting signal peptide, to plastids, to mitochondria, may be located at the N-terminus, at the C-terminus, or in an internal location of the Cpf1 or Csm1 polypeptide.
[0020] In other embodiments, the Cpf1 or Csm1 polypeptide may also comprise at least one marker domain. Non-limiting examples of marker domains include fluorescent proteins, purification tags, and epitope tags. In certain embodiments, the marker domain may be a fluorescent protein.Non-limiting examples of suitable fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, GFP tag, turboGFP, EGFP, Emerald, Azami Green, Azami Green Monomeric, CopGFP, AceGFP, Zsgreen1), yellow fluorescent proteins (e.g., YFP, EYFP, Citrine, Venus, YPet, PhiYFP, ZsYellow1), blue fluorescent proteins (e.g., EBFP, EBFP2, Azurite, mKalama1, GFPuv, Sapphire, T-Sapphire), cyan fluorescent proteins (e.g., ECFP, Cerulean, CyPet, AmCyan1, Midoriishi Cyan), red fluorescent proteins (mKate, mKate2, mPlum, DsRed Monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRed1, AsRed2, eqFP611, mRasberry, mStrawberry, Jred) and orange fluorescent proteins (mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, tdTomato) or any other protein. Petition 870240102758, dated 02 / 12 / 2024, page 39 / 196 17 / 172 suitable fluorescent. In other embodiments, the marker domain may be a purification label and / or an epitope label. Exemplary labels include, but are not limited to, glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly (NANP), tandem affinity purification label (TAP), myc, AcV5, AU1, AU5, E, ECS, E2, FLAG, HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, biotin carboxyl-binding protein (BCCP), and calmodulin.
[0021] In certain embodiments, the Cpf1 or Csm1 polypeptide may be part of a protein-RNA complex comprising a guide RNA. The guide RNA interacts with the Cpf1 or Csm1 polypeptide to direct the Cpf1 or Csm1 polypeptide to a specific target site, where the 5' end of the guide RNA may base-pair with a specific protospacer sequence of the nucleotide sequence of interest in the plant genome, whether part of the nuclear, plastid, and / or mitochondrial genome. As used herein, the term “DNA-directed RNA” refers to a guide RNA that interacts with the Cpf1 or Csm1 polypeptide and the target site of the nucleotide sequence of interest in the genome of a plant cell. A DNA-directed RNA or a DNA polynucleotide encoding a DNA-directed RNA Petition 870240102758, dated 02 / 12 / 2024, p. 40 / 196 18 / 172 may comprise: a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA, and a second segment that interacts with a Cpf1 or Csm1 polypeptide.
[0022] The polynucleotides encoding Cpf1 and Csm1 polypeptides described herein can be used to isolate corresponding sequences from other prokaryotic or eukaryotic organisms. In this way, methods such as PCR, hybridization, and the like can be used to identify such sequences based on their homology or sequence identity with the sequences presented herein. Isolated sequences based on their sequence identity with the complete Cpf1 or Csm1 sequences presented herein, or with variants and fragments thereof, are covered by the present invention. Such sequences include sequences that are orthologous to the described Cpf1 and Csm1 sequences. Orthologous is intended to mean genes derived from a common ancestral gene and that are found in different species as a result of speciation.Genes found in different species are considered orthologous when their nucleotide sequences and / or their encoded protein sequences share at least approximately 75%, approximately 80%, approximately 85%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, approximately. Petition 870240102758, dated 02 / 12 / 2024, p. 41 / 196 19 / 172 of 99% or more sequence identity. Ortholog functions are often highly conserved across species. Thus, isolated polynucleotides encoding polypeptides having Cppl or Csm1 endonuclease activity and sharing at least about 75% or more sequence identity with the sequences described herein are covered by the present invention. As used herein, Cppl or Csm1 endonuclease activity refers to CRISPR endonuclease activity, wherein a guide RNA (gRNA) associated with a Cpf1 or Csm1 polypeptide causes the Cpf1-gRNA or Csm1-gRNA complex to bind to a predetermined nucleotide sequence that is complementary to the gRNA; and wherein the Cpf1 or Csm1 activity can introduce a double-strand break at or near the site targeted by the gRNA. In certain modalities, this double-strand break can be a stepped double-strand break of DNA.As used herein, a “stepped DNA double-strand break” can result in a double-strand break with about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotide overhangs at the 3' or 5' ends after cleavage. In specific embodiments, the Cpfl or Csml polypeptide introduces a stepped DNA double-strand break with a 5-nt 4' or 5' overhang. Petition 870240102758, dated 02 / 12 / 2024, p. 42 / 196 20 / 172 the RNA sequence directed to DNA (e.g., guide RNA) is directed.
[0023] Fragments and variants of the Cpf1 and Csm1 polynucleotides and the Cpf1 and Csm1 amino acid sequences encoded in this way are covered herein. By “fragment” is meant a portion of the polynucleotide or a portion of the amino acid sequence. “Variant” is meant substantially similar sequences. For polynucleotides, a variant comprises a polynucleotide having deletions (i.e., truncations) at the 5' and / or 3' end; deletion and / or addition of one or more nucleotides at one or more internal sites in the native polynucleotide; and / or substitution of one or more nucleotides at one or more sites in the native polynucleotide. As used herein, a “native” polynucleotide or polypeptide comprises a naturally occurring nucleotide sequence or an amino acid sequence, respectively.Generally, variants of a particular polynucleotide of the invention will have at least about 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity with that particular polynucleotide, as determined by the sequence alignment programs and parameters as described elsewhere herein.
[0024] The amino acid or protein “Variant” is intended to mean an amino acid or a protein Petition 870240102758, dated 02 / 12 / 2024, p. 43 / 196 21 / 172 derived from the native amino acid or protein by elimination (so-called truncation) of one or more amino acids at the N-terminal and / or C-terminal end of the native protein; elimination and / or addition of one or more amino acids at one or more internal sites of the native protein; or substitution of one or more amino acids at one or more sites in the native protein. The variant proteins covered by the present invention are biologically active, that is, they continue to possess the desired biological activity of the native protein. Biologically active variants of a native polypeptide will have at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identity with the amino acid sequence to the native sequence, as determined by the sequence alignment programs and parameters described herein.A biologically active variant of a protein of the invention may differ from that protein by only 1-15 amino acid residues, only 1-10, such as 6-10, only 5, such as 4, 3, 2 or even 1 amino acid residue.
[0025] Variant sequences can also be identified by analyzing existing databases of sequenced genomes. In this way, the corresponding sequences can be identified and used in the methods of the invention.
[0026] Sequence alignment methods for Petition 870240102758, dated 02 / 12 / 2024, p. 44 / 196 22 / 172 comparisons are well known in the art. Thus, determining the percentage of sequence identity between any two sequences can be performed using a mathematical algorithm. Non-limiting examples of such mathematical algorithms are the Myers and Miller algorithm (1988)CABIOS 4:11-17; the local alignment algorithm of Smith et al. (1981) Adv. Appl. Math. 2:482; the global alignment algorithm of Needleman and Wunsch (1970) J. Mol. Biol. 48:443-453; the local search alignment method of Pearson and Lipman (1988) Proc. Natl. Acad. Sci. 85:2444-2448; the algorithm of Karlin and Altschul (1990) Proc. Natl. Acad. Sci. EUA 87:2264-2268, modified as in Karlin and Altschul (1993) Proc. Natl. Acad. Sci. EUA 90:5873-5877.
[0027] Computational implementations of these mathematical algorithms can be used for sequence comparison to determine sequence identity. Such implementations include, but are not limited to: CLUSTAL in the PC / Gene program (available from Intelligenetics, Mountain View, California); the ALIGN program (Version 2.0) and GAP, BESTFIT, BLAST, FASTA, and TFASTA in the Wisconsin Genetics Software Package (GCG), Version 10 (available from Accelrys Inc., 9685 Scranton Road, San Diego, California, USA). Alignments using these programs can be performed using standard parameters. The CLUSTAL program is well described by Higgins et al. (1988) Gene 73:237-244; Higgins Petition 870240102758, dated 02 / 12 / 2024, page 45 / 196 23 / 172 et al. (1989) CABIOS 5:151-153; Corpet et al. (1988) Nucleic Acids Res. 16:10881-90; Huang et al. (1992) CABIOS 8:155-65; and Pearson et al. (1994) Meth. Mol. Biol. 24:307-331. The ALIGN program is based on the algorithm of Myers and Miller (1988) supra. A PAM120 weight residue table, a gap length penalty of 12, and a gap penalty of 4 can be used with the ALIGN program when comparing amino acid sequences. The BLAST programs of Altschul et al (1990) J. Mol. Biol. 215: 403 are based on the algorithm of Karlin and Altschul (1990) supra. BLAST nucleotide searches can be performed using the BLASTN program, score = 100, word length = 12, to obtain nucleotide sequences homologous to a nucleotide sequence encoding a protein of the invention.BLAST protein searches can be performed using the BLASTX program, score = 50, word length = 3, to obtain amino acid sequences homologous to a protein or polypeptide of the invention. To obtain gap alignments for comparison purposes, gap BLAST (in BLAST 2.0) can be used as described in Altschul et al. (1997) Nucleic Acids Res. 25: 3389. Alternatively, PSI-BLAST (in BLAST 2.0) can be used to perform an iterated search that detects distant relationships between molecules. See Altschul et al. (1997) above. When using BLAST, gap BLAST, PSI-BLAST, the standard parameters can be used. Petition 870240102758, dated 02 / 12 / 2024, p. 46 / 196 24 / 172 respective programs (e.g., BLASTN for nucleotide sequences, BLASTX for proteins). See the website at www.ncbi.nlm.nih.gov. Alignment can also be performed manually by inspection.
[0028] Nucleic acid molecules encoding Cpf1 and Csm1 polypeptides, or fragments or variants thereof, can be codon-optimized for expression in a plant of interest or other cell or organism of interest. A “codon-optimized gene” is a gene having its codon usage frequency configured to mimic the preferred codon usage frequency of the host cell. Nucleic acid molecules can be codon-optimized, wholly or partially. Since any amino acid (except methionine and tryptophan) is encoded by multiple codons, the sequence of the nucleic acid molecule can be changed without changing the encoded amino acid. Codon optimization occurs when one or more codons is / are altered at the nucleic acid level so that the amino acids are not altered, but expression in a particular host organism is increased.Those skilled in the art will recognize that codon tables and other references providing preference information for a wide range of organisms are available in the art (see, for example, Zhang et al. (1991) Gene 105: 61-72; Murray et al. (1989) Nucl. Acids Res. 17: 477-508). The methodology for optimizing... Petition 870240102758, dated 02 / 12 / 2024, p. 47 / 196 25 / 172 A nucleotide sequence for expression in a plant is provided, for example, in Pat. No. 6,015,891, and the references cited herein. Examples of polynucleotides optimized with codons for expression in a plant are presented in: SEQ ID Nos: 5, 8, 11, 14, 17, 19, 22, 25 and 174-206.
[0029] II. Fusion proteins
[0030] Fusion proteins are hereby provided comprising a Cpf1 or Csm1 polypeptide, or a fragment or variant thereof, and an effector domain. The Cpf1 or Csm1 polypeptide may be directed to a target site by a guide RNA, at which site the effector domain may modify or effect the targeted nucleic acid sequence. The effector domain may be a cleavage domain, an epigenetic modification domain, a transcription activation domain, or a transcription repressor domain. The fusion protein may further comprise at least one additional domain chosen from a nuclear localization signal, plastid signal peptide, mitochondrial signal peptide, signal peptide capable of trafficking protein to multiple subcellular locations, a cell penetration domain, or a marker domain, any of which may be located at the N-terminus, the C-terminus, or an internal location of the fusion protein.The polypeptide Cpf1 or Csm1 may be located in Petition 870240102758, dated 02 / 12 / 2024, page 48 / 196 26 / 172 N-terminal, C-terminal, or an internal location of the fusion protein. The Cpf1 or Csm1 polypeptide can be directly fused to the effector domain, or it can be fused with a linker. In specific embodiments, the linker sequence fusing the Cpf1 or Csm1 polypeptide with the effector domain can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, or 50 amino acids in length. For example, the linker can range from 1-5, 1-10, 1-20, 1-50, 23, 3-10, 3-20, 5-20, or 10-50 amino acids in length.
[0031] In some embodiments, the Cpf1 or Csm1 polypeptide of the fusion protein may be derived from a wild-type Cpf1 or Csm1 protein. The Cpf1-derived or Csm1-derived protein may be a variant or a modified fragment. In some embodiments, the Cpf1 or Csm1 polypeptide may be modified to contain a nuclease domain (e.g., a RuvC-type domain) with reduced or eliminated nuclease activity. For example, the Cpf1-derived or Csm1-derived polypeptide may be modified so that the nuclease domain is eliminated or mutated so that it is no longer functional (i.e., nuclease activity is absent). Specifically, a Cpf1 or Csm1 polypeptide may have a mutation at a position corresponding to positions 917 or 1006 of FnCpf1 (SEQ ID NO: 3) or positions 701 or 922 of SmCsm1 (SEQ ID NO: 160) when aligned for maximum identity. For example, a conversion Petition 870240102758, dated 02 / 12 / 2024, p. 49 / 196 Aspartate to alanine (D917A) and glutamate to alanine (E1006A) mutations in a RuvC-like domain completely inactivated the DNA cleavage activity of FnCpf1, while aspartate to alanine (D1255A) significantly reduced cleavage activity (Zetsche et al. 2015; Cell 163: 759-771). Examples of Cpf1 polypeptides having mutations in the RuvC domain are presented in SEQ ID NOs: 26-41 and 63-70. The nuclease domain can be inactivated by one or more deletion mutations, insertion mutations, and / or substitution mutations using known methods such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, as well as other methods known in the art. In one exemplary embodiment, the Cpf1 or Csm1 polypeptide of the fusion protein is modified by mutation of the RuvC-like domain such that the Cpf1 or Csm1 polypeptide lacks nuclease activity.
[0032] The fusion protein also comprises an effector domain located at the N-terminus, the C-terminus, or an internal location within the fusion protein. In some embodiments, the effector domain is a cleavage domain. As used herein, a “cleavage domain” refers to a domain that cleaves DNA. The cleavage domain can be obtained from any endonuclease or exonuclease. Non-limiting examples of endonucleases from which a cleavage domain can be derived include, but are not limited to, Petition 870240102758, dated 02 / 12 / 2024, page 50 / 196 28 / 172 Restriction endonucleases and endonucleases of origin. See, for example, the New England Biolabs Catalog or Belfort et al. (1997) Nucleic Acids Res. 25: 3379-3388. Additional enzymes that cleave DNA are known (e.g., S1 nuclease, mung bean nuclease, pancreatic DNase I, micrococcal nuclease, yeast HO endonuclease). See also Linn et al. (eds.) Nucleases, Cold Spring Harbor Laboratory Press, 1993. One or more of these enzymes (or functional fragments thereof) can be used as a source of cleavage domains.
[0033] In some embodiments, the cleavage domain may be derived from a type II-S endonuclease. Type II-S endonucleases cleave DNA at sites that are typically several base pairs away from the recognition site and, as such, possess separable recognition and cleavage domains. These enzymes are generally monomers that transiently associate to form dimers to cleave each DNA strand at staggered locations. Non-limiting examples of suitable type II-S endonucleases include BfiI, BpmI, BsaI, BsgI, BsmBI, BsmI, BspMI, FokI, MbolI, and SapI.
[0034] In certain embodiments, type IIS cleavage can be modified to facilitate dimerization of two different cleavage domains (each of which is linked to a Cpf1 or Csm1 polypeptide or fragment thereof). Petition 870240102758, dated 02 / 12 / 2024, page 51 / 196 29 / 172 In embodiments where the effector domain is a cleavage domain, the Cpf1 or Csm1 polypeptide can be modified as discussed herein, so that its endonuclease activity is eliminated. For example, the Cpf1 or Csm1 polypeptide can be modified by mutation of the RuvC-like domain so that the polypeptide no longer exhibits endonuclease activity.
[0035] In other embodiments, the effector domain of the fusion protein may be an epigenetic modification domain. In general, epigenetic modification domains alter histone structure and / or chromosome structure without altering the DNA sequence. Changes in histone and / or chromatin structure can lead to changes in gene expression. Examples of epigenetic modification include, but are not limited to, acetylation or methylation of lysine residues in histone proteins, and methylation of cytosine residues in DNA. Non-limiting examples of suitable epigenetic modification domains include histone acetyltransferase domains, histone deacetylase domains, histone methyltransferase domains, histone demethylase domains, DNA methyltransferase domains, and DNA demethylase domains.
[0036] In embodiments where the effector domain is a histone acetyltransferase (HAT) domain, the HAT domain may be derived from EP300 (i.e., E1A-binding protein). Petition 870240102758, dated 02 / 12 / 2024, page 52 / 196 30 / 172 p300), CREBBP (i.e., CREB-binding protein), CDY1, CDY2, CDYL1, CLOCK, ELP3, ESA1, GCN5 (KAT2A), HAT1, KAT2B, KAT5, MYST1, MYST2, MYST3, MYST4, NCOA1, NCOA2, NCOA3, NCOAT, P / CAF, Tip60, TAFII250, or TF3C4. In embodiments where the effector domain is an epigenetic modification domain, the Cpf1 or Csm1 polypeptide can be modified as discussed herein so that its endonuclease activity is eliminated. For example, the Cpf1 or Csm1 polypeptide can be modified by mutation of the RuvC-like domain so that the polypeptide no longer possesses nuclease activity.
[0037] In some embodiments, the effector domain of the fusion protein may be a transcription activation domain. In general, a transcription activation domain interacts with transcription control elements and / or transcription regulatory proteins (i.e., transcription factors, RNA polymerases, etc.) to increase and / or activate the transcription of one or more genes. In some embodiments, the transcription activation domain may be, without limitation, a VP16 activation domain of the herpes simplex virus, VP64 (which is a tetrameric derivative of VP16), an NFkB p65 activation domain, p53 1 and 2 activation domains, a CREB (cAMP response element-binding protein), an E2A activation domain, and an NFAT (activated T-cell nuclear factor) activation domain. In other modalities, the transcription activation domain may be Petition 870240102758, dated 02 / 12 / 2024, page 53 / 196 31 / 172 Gal4, Gcn4, MLL, Rtg3, Gln3, Oaf1, Pip2, Pdr1, Pdr3, Pho4, and Leu3. The transcription activation domain can be wild-type or a modified version of the original transcription activation domain. In some embodiments, the effector domain of the fusion protein is a VP16 or VP64 transcription activation domain. In embodiments where the effector domain is a transcription activation domain, the Cpf1 or Csm1 polypeptide can be modified as discussed herein so that its endonuclease activity is eliminated. For example, the Cpf1 or Csm1 polypeptide can be modified by mutation of the RuvC-like domain so that the polypeptide no longer possesses nuclease activity.
[0038] In other embodiments, the effector domain of the fusion protein may be a transcriptional repressor domain. In general, a transcriptional repressor domain interacts with transcriptional control elements and / or transcriptional regulatory proteins (i.e., transcription factors, RNA polymerases, etc.) to decrease and / or terminate the transcription of one or more genes. Non-limiting examples of suitable transcriptional repressor domains include inducible cAMP early repressor domains (ICERs), Kruppel-associated A-box repressor domains (KRAB-A), YY1 glycine-rich repressor domains, Sp1-type repressors, E(spl) repressors, I.kappa.B repressor, and MeCP2. In embodiments where the effector domain is Petition 870240102758, dated 02 / 12 / 2024, page 54 / 196 32 / 172 a transcriptional repressor domain, the Cpfl or Csml polypeptide can be modified as discussed here, so that its endonuclease activity is eliminated. For example, the Cpf1 or Csm1 polypeptide can be modified by mutation of the RuvC-like domain so that the polypeptide no longer possesses nuclease activity.
[0039] In some embodiments, the fusion protein further comprises at least one additional domain. Non-limiting examples of suitable additional domains include nuclear localization signals, cell penetration or translocation domains, and marker domains.
[0040] When the effector domain of the fusion protein is a cleavage domain, a dimer comprising at least one fusion protein may be formed. The dimer may be a homodimer or a heterodimer. In some embodiments, the heterodimer comprises two different fusion proteins. In other embodiments, the heterodimer comprises one fusion protein and an additional protein.
[0041] The dimer may be a homodimer in which the two fusion protein monomers are identical with respect to the primary amino acid sequence. In an embodiment in which the dimer is a homodimer, the Cpf1 or Csm1 polypeptide may be modified so that endonuclease activity is eliminated. In certain embodiments in which the Cppl or Csm1 polypeptide is modified so that endonuclease activity Petition 870240102758, dated 02 / 12 / 2024, page 55 / 196 33 / 172 is eliminated, each fusion protein monomer may comprise an identical Cpcl or Csm1 polypeptide and an identical cleavage domain. The cleavage domain may be any cleavage domain, such as any of the exemplary cleavage domains provided herein. In such embodiments, specific guide RNAs would direct the fusion protein monomers to different but adjacently close sites, so that, upon dimer formation, the nuclease domains of the two monomers would create a double-strand break in the target DNA.
[0042] The dimer can also be a heterodimer of two different fusion proteins. For example, the Cpf1 or Csm1 polypeptide of each fusion protein may be derived from a different Cpf1 or Csm1 polypeptide or from an orthologous Cpf1 or Csm1 polypeptide from a different bacterial species. For example, each fusion protein may comprise a Cpf1 or Csm1 polypeptide derived from a different bacterial species. In these embodiments, each fusion protein would recognize a different target site (i.e., specified by the protospacer and / or PAM sequence). For example, guide RNAs may position the heterodimer at different but closely adjacent sites so that their nuclease domains produce an effective double-strand break in the target DNA.
[0043] Alternatively, two fusion proteins of Petition 870240102758, dated 02 / 12 / 2024, page 56 / 196 34 / 172 A heterodimer can have different effector domains. In embodiments where the effector domain is a cleavage domain, each fusion protein can contain a different modified cleavage domain. In these embodiments, the Cpf1 or Csm1 polypeptide can be modified so that its endonuclease activities are eliminated. The two fusion proteins that form a heterodimer can differ in the Cpf1 or Csm1 polypeptide domain and the effector domain.
[0044] In any of the embodiments described above, the homodimer or heterodimer may comprise at least one additional domain chosen from nuclear localization signals (NLS), plastid signal peptides, mitochondrial signal peptides, signal peptides capable of trafficking proteins to multiple subcellular locations, translocation or cell penetration domains, and marker domains, as detailed above. In any of the embodiments described above, one or both of the Cpf1 or Csm1 polypeptides may be modified such that the endonuclease activity of the polypeptide is eliminated or modified.
[0045] The heterodimer may also comprise a fusion protein and an additional protein. For example, the additional protein may be a nuclease. In one embodiment, the nuclease is a zinc finger nuclease. A zinc finger nuclease comprises a DNA-binding domain of Petition 870240102758, dated 02 / 12 / 2024, page 57 / 19635 / 172 zinc finger and a cleavage domain. A zinc finger recognizes and binds three (3) nucleotides. A zinc finger DNA binding domain may comprise from about three zinc fingers to about seven zinc fingers. The zinc finger DNA binding domain may be derived from a naturally occurring protein or may be manipulated. See, for example, Beerli et al. (2002) Nat. Biotechnol. 20:135-141; Pabo et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan et al. (2001) Nat. Biotechnol. 19:656-660; Segal et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo et al. (2000) Curr. Opin. Struct. Biol. 10:411-416; Zhang et al. (2000) J. Biol. Chem. 275(43):33850-33860; Doyon et al. (2008) Nat. Biotechnol. 26:702-708; and Santiago et al. (2008) Proc. Natl. Acad. Sci. EUA 105:5809-5814. The cleavage domain of zinc finger nuclease can be any cleavage domain detailed here.In some embodiments, the zinc finger nuclease may comprise at least one additional domain chosen from nuclear localization signals, plastid signal peptides, mitochondrial signal peptides, signal peptides capable of trafficking proteins to multiple subcellular locations, translocation or cell penetration domains, which are detailed here.
[0046] In certain embodiments, any of the fusion proteins detailed above or a dimer comprising at least one fusion protein can do Petition 870240102758, dated 02 / 12 / 2024, page 58 / 196 36 / 172 part of a protein-RNA complex comprising at least one guide RNA. A guide RNA interacts with the Cpf1 or Csm1 polypeptide of the fusion protein to direct the fusion protein to a specific target site, where the 5' end of the guide RNA base pairs with a specific protospacer sequence. III. Nucleic acids encoding Cpf1 or Csm1 polypeptides or fusion proteins
[0047] Nucleic acids encoding either of the Cpf1 and Csm1 polypeptides or fusion proteins described herein are provided. The nucleic acid may be RNA or DNA. Examples of polynucleotides encoding Cpf1 polypeptides are presented in SEQ ID NOs: 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19, 21, 22, 24, 25, and 174-184, 187-192, 194-201, and 203-206. Examples of polynucleotides encoding Csm1 polypeptides are presented in SEQ ID NOs: 185, 186, 193, and 202. In one embodiment, the nucleic acid encoding the Cpf1 or Csm1 polypeptide or fusion protein is mRNA. The mRNA can be capped at 5' and / or polyadenylated at 3'. In another embodiment, the nucleic acid encoding the Cpf1 or Csm1 polypeptide or fusion protein is DNA. The DNA may be present in a vector.
[0048] Nucleic acids encoding the Cpf1 or Csm1 polypeptide or fusion proteins can be codon-optimized for efficient translation into proteins in Petition 870240102758, dated 02 / 12 / 2024, page 59 / 196 37 / 172 plant cell of interest. Programs for codon optimization are available in the technique (e.g., OPTIMIZER at genomes.urv.es / OPTIMIZER; OptimumGene.TM from GenScript at www.genscript.com / codon-opt.html).
[0049] In certain embodiments, the DNA encoding the Cpf1 or Csm1 polypeptide or fusion protein can be operationally linked to at least one promoter sequence. The DNA coding sequence can be operationally linked to a promoter control sequence for expression in a host cell of interest. In some embodiments, the host cell is a plant cell. “Operationally linked” is intended to mean a functional link between two or more elements. For example, an operational link between a promoter and a coding region of interest (e.g., encoding the region of a Cpf1 or Csm1 polypeptide or guide RNA) is a functional link that allows expression of the coding region of interest. Operationally linked elements can be contiguous or non-contiguous.When used to refer to the junction of two protein coding regions, operationally linked terms mean that the coding regions are in the same reading frame.
[0050] The promoter sequence can be constitutive, regulated, or growth stage specific. Petition 870240102758, dated 02 / 12 / 2024, page 60 / 196 38 / 172 or tissue-specific. It is recognized that different applications can be enhanced by the use of different promoters on nucleic acid molecules to modulate the timing, location, and / or level of expression of the Cpcl or Csml polypeptide and / or guide RNA. Such nucleic acid molecules may also contain, if desired, a promoter regulatory region (e.g., one that confers expression, is constitutively inducible, environmentally or developmentally regulated, or selectively / specific to tissue or cell), a transcription initiation site, a ribosome binding site, an RNA processing signal, a transcription termination site, and / or a polyadenylation signal.
[0051] In some embodiments, the nucleic acid molecules provided herein may be combined with constitutive, tissue-preferred, developmentally preferred, or other promoters for expression in plants. Examples of functional constitutive promoters in plant cells include the 35S transcription initiation region of cauliflower mosaic virus (CaMV), the 1' or 2' T-DNA derived promoter of Agrobacterium tumefaciens, the ubiquitin 1 promoter, the Smas promoter, the cinnamyl alcohol dehydrogenase promoter (US Patent No. 5,683,439), the Nos promoter, the pEmu promoter, the rubisco promoter, the GRP1-8 promoter, and other initiation regions of Petition 870240102758, dated 02 / 12 / 2024, page 61 / 196 39 / 172 transcription of several known plant genes from the enabled. If low-level expression is desired, weak promoter(s) may be used. Weak constitutive promoters include, for example, the core promoter of the Rsyn7 promoter (WO 99 / 43838 and US Patent No. 6,072,050), the core promoter of CaMV 35S and the like. Other constitutive promoters include, for example, US Pat. Nos. 5,608,149; 5,608,144; 5,604,121; 5,569,597; 5,466,785; 5,399,680; 5,268,463; and 5,608,142. See also U.S. Pat. No. 6,177,611, incorporated herein by reference.
[0053] Examples of inducible promoters are the Adh1 promoter, which is inducible by hypoxia or cold stress; the Hsp70 promoter, which is inducible by heat stress; the PPDK promoter and the pepcarboxylase promoter, which are both inducible by light. Also useful are chemically inducible promoters, such as the In22 promoter, which is induced by protector (US Patent No. 5,364,780); the ERE promoter, which is induced by estrogen; and the Axig1 promoter, which is induced by auxin and specific to the tapetum, but also active in callus (PCT US01 / 22169).
[0054] Examples of promoters under developmental control in plants include promoters that preferentially initiate transcription in certain tissues, such as leaves, roots, fruits, seeds, or flowers. A promoter Petition 870240102758, dated 02 / 12 / 2024, page 62 / 196 40 / 172 'Tissue-specific' is a promoter that initiates transcription only in certain tissues. Unlike constitutive gene expression, tissue-specific expression is the result of
[0055] various levels of gene regulatory interaction. As such, promoters from homologous or closely related plant species may be preferable to use in order to achieve efficient and reliable expression of transgenes in particular tissues. In some embodiments, expression comprises a tissue-preferred promoter. A “tissue-preferred” promoter is a promoter that initiates transcription preferentially, but not necessarily entirely or only in certain tissues.
[0056] In some embodiments, nucleic acid molecules encoding a Cpfl or Csml polypeptide and / or guide RNA comprise a cell-type-specific promoter. A “cell-type-specific” promoter is a promoter that primarily triggers expression in certain cell types in one or more organs. Some examples of plant cells in which functional plant cell-type-specific promoters may be primarily active include, for example, BETL cells, vascular cells in roots, leaves, stem cells, and stem cells. Nucleic acid molecules may also include cell-type-preferred promoters. A “cell-type” promoter Petition 870240102758, dated 02 / 12 / 2024, page 63 / 196 41 / 172 preferred promoter is a promoter that primarily drives expression mainly, but not necessarily entirely, or only in certain cell types in one or more organs. Some examples of plant cells in which functional cell-type preferred promoters in plants may be preferentially active include, for example, BETL cells, vascular cells in roots, leaves, stem cells, and stem cells. The nucleic acid molecules described herein may also comprise seed preferred promoters. In some embodiments, seed preferred promoters are expressed in the embryo sac, early embryo, early endosperm, aleurone, and / or basal endosperm transfer cell layer (BETL).
[0057] Examples of preferred seed promoters include, but are not limited to, 27 kD gamma zein promoter and fatty promoter, Boronat, A. et al. (1986) Plant Sci. 47:95-102; Reina, M. et al. Nucl. Acids Res. 18(21):6426; and Kloesgen, RB et al. (1986) Mol. O Gen. Genet. 203:237-244. Promoters expressed in the embryo, pericarp, and endosperm are described in US Pat. No. 6,225,529 and PCT publication WO 00 / 12733. The descriptions for each of these are incorporated herein by reference in their entirety.
[0058] Promoters that can direct gene expression in a preferred manner in seeds of Petition 870240102758, dated 02 / 12 / 2024, p. 64 / 196 42 / 172 plants with expression in the embryo sac, early embryo, early endosperm, aleurone and / or basal endosperm transfer cell layer (BETL) can be used in the compositions and methods described herein.These promoters include, but are not limited to, promoters that are naturally linked to the Zea mays early endosperm gene 5, Zea mays early endosperm gene 1, Zea mays early endosperm gene 2, GRMZM2G124663, GRMZM2G006585, GRMZM2G120008, GRMZM2G157806, GRMZM2G176390, GRMZM2G472234, GRMZM2G138727, Zea mays CLAVATA1, Zea mays MRP1, Oryza sativa PR602, Oryza sativa PR9a, Zea mays BET1, Zea mays BETL-2, Zea mays BETL-3, Zea mays BETL-4, Zea mays BETL-9, BETL-10 from Zea mays, MEG1 from Zea mays, TCCR1 from Zea mays, ASP1 from Zea mays, ASP1 from Oryza sativa, PR60 from Triticum durum, PR91 from Triticum durum, GL7 from Triticum durum, AT3G10590, AT4G18870, AT4G21080, AT5G23650, AT3G05860, AT5G42910, AT2G26320, AT3G03260, AT5G26630, AtIPT4, AtIPT8, AtLEC2, LFAH12. Other promoters of this type are described in US Patents Nos. 7803990, 8049000, 7745697, 7119251, 7964770, 7847160, 7700836, US Patent Application Publication.Nos20100313301, 20090049571, 20090089897, 20100281569, 20100281570, 20120066795, 20040003427; US PCT Publications NosWO / 1999 / 050427, WO / 2010 / 129999, WO / 2009 / 094704, WO / 2010 / 019996 and WO / 2010 / 147825, each of which is incorporated herein by. Petition 870240102758, dated 02 / 12 / 2024, page 65 / 196 43 / 172 reference in its entirety for all purposes. Functional variants or functional fragments of the promoters described herein may also be operationally linked to the nucleic acids described herein.
[0059] Chemically regulated promoters can be used to modulate gene expression through the application of an exogenous chemical regulator. Depending on the objective, the promoter can be a chemically inducible promoter, where the application of the chemical induces gene expression, or a chemically repressible promoter, where the application of the chemical represses gene expression. Chemically inducible promoters are known in the art and include, but are not limited to, the maize In2-2 promoter, which is activated by benzenesulfonamide herbicide protectants, the maize GST promoter, which is activated by hydrophobic electrophilic compounds used as pre-emergent herbicides, and the tobacco PR-1a promoter, which is activated by salicylic acid. Other chemically regulated promoters of interest include steroid-responsive promoters (see, for example, the glucocorticoid-inducible promoter in Schena et al.(1991) Proc. Natl. Acad. Sci. USA 88: 1042110425 and McNellis et al. (1998) Plant J. 14(2):247-257) and tetracycline-repressible or tetracycline-inducible promoters (see, for example, Gatz et al. (1991) Mol. Gen. Petition 870240102758, dated 02 / 12 / 2024, p. 66 / 196 44 / 172 Genet. 227: 229-237, and U.S. Pat. Nos. 5,814,618 and 5,789,156), incorporated herein by reference.
[0060] Preferred tissue promoters can be used to target enhanced expression of an expression construct within a particular tissue. In certain modalities, preferred tissue promoters may be active in plant tissue. Preferred tissue promoters are known in the art. See, for example, Yamamoto et al. (1997) Planta J. 12(2):255-265; Kawamata et al. (1997) Plant Cell Physiol. 38(7):792-803; Hansen et al. (1997) Mol. Gen Genet. 254(3):337-343; Russell et al. (1997) Transgenic Res. 6(2):157-168; Rinehart et al. (1996) Plant Physiol. 112(3):1331-1341; Van Camp et al. (1996) Plant Physiol. 112(2):525-535; Canevascini et al. (1996) Plant Physiol. 112(2):513-524; Yamamoto et al. (1994) Plant Cell Physiol. 35(5):773-778; Lam (1994) Results Probl. Cell Difference. 20:181-196; Orozco et al. (1993) Plant Mol Biol. 23(6):11291138; Matsuoka et al. (1993) Proc Natl. Academic. Sci. USA 90(20):9586-9590; and Guevara-Garcia et al. (1993) Plant J. 4(3):495-505. Such promoters can be modified, if necessary, for weak expression.
[0061] Preferred leaf promoters are known in the art. See, for example, Yamamoto et al. (1997) Plant J. 12(2):255-265; Kwon et al. (1994) Plant Physiol. 105:357-67; Yamamoto et al. (1994) Plant Cell Petition 870240102758, dated 02 / 12 / 2024, page 67 / 196 45 / 172 Physiol. 35(5):773-778; Gotor et al. (1993) Plant J. 3:50918; Orozco et al. (1993) Plant Mol. Biol. 23(6):112 9-1138; and Matsuoka et al. (1993) Proc. Natl. Acad. Sci. USA 90(20):9586-9590. In addition, cab and rubisco promoters can also be used. See, for example, Simpson et al. (1958) EMBO J 4:2723-2729 and Timko et al. (1988) Nature 318:57-58.
[0062] Preferred root promoters are known and can be selected from the many available in the literature or isolated de novo from various compatible species. See, for example, Hire et al. (1992) Plant Mol. Biol. 20 (2): 207-218 (soybean root-specific glutamine synthetase gene); Keller and Baumgartner (1991) Plant Cell 3 (10): 1051-1061 (root-specific control element in the GRP 1.8 gene of French bean); Sanger et al. (1990) Plant Mol. Biol. 14(3):433-443 (root-specific promoter of the mannopin synthase (MAS) gene of Agrobacterium tumefaciens); and Miao et al. (1991) Plant Cell 3(1):11-22 (full-length cDNA clone encoding cytosolic glutamine synthetase (GS), which is expressed in soybean roots and root nodules). See also Bogusz et al.(1990) Plant Cell 2(7):633-641, where two root-specific promoters isolated from nitrogen-fixing hemoglobin genes of Parasponia andersonii and Trema tomentosa are described. Petition 870240102758, dated 02 / 12 / 2024, page 68 / 196 46 / 172 without nitrogen fixation. The promoters of these genes were linked to a β-glucuronidase reporter gene and introduced into the non-leguminous Nicotiana tabacum and the leguminous Lotus corniculatus, and in both cases the root-specific promoter activity was preserved. Leach and Aoyagi (1991) describe their analysis of the promoters of the highly expressed root-inducing genes roIC and roID from Agrobacterium rhizogenes (see Plant Science (Limerick) 79(1):69-76). They concluded that the enhancer and tissue-preferred DNA determinants are dissociated in these promoters. Teeri et al. (1989) used gene fusion for lacZ to show that the Agrobacterium T-DNA gene encoding octopine synthase is especially active in the root tip epidermis and that the TR2' gene is root-specific in the intact plant and stimulated by injury to the leaf tissue, a particularly desirable combination of features for use with an insecticidal or larvicidal gene (see EMBO J.8(2):343-350). The TR1' gene, fused with nptII (neomycin phosphotransferase II), showed similar characteristics. Additional preferred root promoters include the VfENOD-GRP3 gene promoter (Kuster et al. (1995) Plant Mol. Biol. 29(4):759-772); and the roIB promoter (Capana et al. (1994) Plant Mol. Biol. 25(4):681-691. See also US Pat. Nos. 5,837,876; 5,750,386; 5,633,363; 5,459,252; 5,401,836; 5,110,732; and 5,023,179. Petition 870240102758, dated 02 / 12 / 2024, p. 69 / 196 47 / 172 et al. (1983) Science 23:476-482 and Sengopta-Gopalen et al. (1988) PNAS 82:3320-3324. The promoter sequence can be wild-type or it can be modified for more efficient or effective expression.
[0063] Nucleic acid sequences encoding the Cpf1 or Csm1 polypeptide or fusion protein can be operationally linked to a promoter sequence that is recognized by a phage RNA polymerase for in vitro mRNA synthesis. In such embodiments, the RNA transcribed in vitro can be purified for use in the genome modification methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence or a variation of a T7, T3, or SP6 promoter sequence. In some embodiments, the sequence encoding the Cpf1 or Csm1 polypeptide or fusion protein can be operationally linked to a promoter sequence for in vitro expression of the Cpf1 or Csm1 polypeptide or fusion protein in plant cells. In such embodiments, the expressed protein can be purified for use in the genome modification methods described herein.
[0064] In certain embodiments, the DNA encoding the Cpf1 or Csm1 polypeptide or fusion protein may also be linked to a polyadenylation signal (e.g., SV40 polyA signal and other functional signals in plants) and / or at least one transcription termination sequence. Petition 870240102758, dated 02 / 12 / 2024, page 70 / 196 48 / 172 Additionally, the sequence encoding the Cpf1 or Csm1 polypeptide or fusion protein may also be linked to the sequence encoding at least one nuclear localization signal, at least one plastid signal peptide, at least one mitochondrial signal peptide, at least one signal peptide capable of trafficking proteins to multiple subcellular locations, at least one cell penetration domain, and / or at least one marker domain, described elsewhere herein.
[0065] The DNA encoding the Cpf1 or Csm1 polypeptide or fusion protein may be present in a vector. Suitable vectors include plasmid vectors, phagomids, cosmids, minichromosomes / artificial chromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated viral vectors, etc.). In one embodiment, the DNA encoding the Cpf1 or Csm1 polypeptide or fusion protein is present in a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, pCAMBIA, and variants thereof. The vector may comprise additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication Petition 870240102758, dated 02 / 12 / 2024, page 71 / 196 49 / 172 and similar. Additional information can be found in “Current Protocols in Molecular Biology”, Ausubel et al., John Wiley & Sons, New York, 2003 or “Molecular Cloning: A Laboratory Manual” Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.
[0066] In some embodiments, the expression vector comprising the sequence encoding the Cpf1 or Csm1 polypeptide or fusion protein may further comprise a sequence encoding a guide RNA. The sequence encoding the guide RNA may be operationally linked to at least one transcriptional control sequence for expression of the guide RNA in the plant or plant cell of interest. For example, the DNA encoding the guide RNA may be operationally linked to a promoter sequence that is recognized by RNA polymerase III (Pol III). Examples of suitable Pol III promoters include, but are not limited to, mammalian U6, U3, H1, and 7SL RNA promoters and rice U6 and U3 promoters. IV. Methods for Modifying a Nucleotide Sequence in a Plant Genome
[0067] The methods are provided here for modifying a nucleotide sequence of a plant cell, plant organelle, or plant embryo. The methods comprise introducing into a plant cell, organelle, or embryo, a DNA-directed RNA or a DNA polynucleotide that Petition 870240102758, dated 02 / 12 / 2024, page 72 / 196 50 / 172 encodes a DNA-directed RNA, wherein the DNA-directed RNA comprises: (a) a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA; and (b) a second segment that interacts with a Cpf1 or Csm1 polypeptide and also introduces into the plant cell a Cpf1 or Csm1 polypeptide, or a polynucleotide encoding a Cpf1 or Csm1 polypeptide, wherein the Cpf1 or Csm1 polypeptide comprises: (a) an RNA-binding portion that interacts with the DNA-directed RNA; and (b) an activity portion that exhibits site-directed enzymatic activity. The plant cell or plant embryo can then be cultured under conditions in which the Cpf1 or Csm1 polypeptide is expressed and cleaves the nucleotide sequence. Note that the system described herein does not require the addition of exogenous Mg+2 or any other ions. Finally, a plant cell or organelle comprising the modified nucleotide sequence can be selected.
[0068] In some embodiments, the method may comprise the introduction of a Cpf1 or Csm1 polypeptide (or nucleic acid-coding peptide) and a guide RNA (or DNA-coding peptide) into a plant cell, organelle, or embryo, wherein the Cpf1 or Csm1 polypeptide introduces a double-strand break in the target nucleotide sequence of the plant chromosomal DNA. In embodiments where an optional donor polynucleotide is not present, the double-strand break in Petition 870240102758, dated 02 / 12 / 2024, page 73 / 196 51 / 172 nucleotide sequences can be repaired by a non-homologous end-join (NHEJ) repair process. Because NHEJ is error-prone, deletions of at least one nucleotide, insertions of at least one nucleotide, substitutions of at least one nucleotide, or combinations thereof, can occur during break repair. Consequently, the targeted nucleotide sequence may be modified or inactivated. For example, a single nucleotide alteration (SNP) may give rise to an altered protein product, or a frameshift in a coding sequence may inactivate or knock out the sequence so that no protein product is produced. In embodiments where the optional donor polynucleotide is present, the donor sequence in the donor polynucleotide may be exchanged with, or integrated into, the nucleotide sequence at the targeted site during double-strand break repair.For example, in embodiments where the donor sequence is flanked by upstream and downstream sequences having substantial sequence identity with upstream and downstream sequences, respectively, of the targeted site in the plant nucleotide sequence, the donor sequence can be exchanged with, or integrated into, the nucleotide sequence at the target site during homology-directed repair. Alternatively, in embodiments... Petition 870240102758, dated 02 / 12 / 2024, p. 74 / 196 52 / 172 where the donor sequence is flanked by compatible overhangs (or the compatible overhangs are generated in situ by the Cppl or Csm1 polypeptide), the donor sequence can be directly linked to the cleaved nucleotide sequence by a non-homologous repair process during double-strand break repair. The exchange or integration of the donor sequence into the nucleotide sequence modifies the targeted plant nucleotide sequence or introduces an exogenous sequence into the nucleotide sequence of the plant cell, plant organelle, or plant embryo.
[0069] The methods described herein may also comprise the introduction of two Cpf1 or Csm1 polypeptides (or nucleic acids) and two guide RNAs (or coding DNAs) into a plant cell, organelle, or plant embryo, wherein the Cpf1 or Csm1 polypeptides introduce two double-strand breaks in the nucleotide sequence of the nuclear and / or organellar chromosomal DNA. The two breaks may be within several base pairs, within tens of base pairs, or may be separated by many thousands of base pairs. In embodiments where an optional donor polynucleotide is not present, the resulting double-strand breaks may be repaired by a non-homologous repair process such that the sequence between the two cleavage sites is lost and / or deletions of at least one nucleotide, insertions of at least Petition 870240102758, dated 02 / 12 / 2024, page 75 / 196 53 / 172 minus one nucleotide, substitutions of at least one nucleotide, or combinations thereof, may occur during the repair of the break(s). In embodiments where an optional donor polynucleotide is present, the donor sequence in the donor polynucleotide may be exchanged or integrated into the plant nucleotide sequence during the repair of double-strand breaks by a homology-based repair process (e.g., in embodiments where the donor sequence is flanked by upstream and downstream sequences having substantial sequence identity with upstream and downstream sequences, respectively, of the targeted sites in the nucleotide sequence) or a non-homologous repair process (e.g., in embodiments where the donor sequence is flanked by compatible overhangs).
[0070] By “altering” or “modulating” the expression level of a gene, the intention is that the gene expression be positively regulated or negatively regulated. It is recognized that, in some cases, plant growth and yield are increased by decreasing the expression levels of one or more genes encoding proteins involved in photosynthesis, i.e., negatively regulated expression. Thus, the invention encompasses the positive or negative regulation of one or more genes encoding proteins involved in photosynthesis, using the Petition 870240102758, dated 02 / 12 / 2024, page 76 / 196 54 / 172 Cpfl or Csm1 polypeptides described herein. Furthermore, the methods include the upregulation of at least one gene encoding a protein involved in photosynthesis and the downregulation of at least one gene encoding a protein involved in photosynthesis in a plant of interest. By modulating the concentration and / or activity of at least one of the genes encoding a protein involved in photosynthesis in a transgenic plant, it is intended that the concentration and / or activity will be increased or decreased by at least about 1%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, or about 90% or more, relative to a native control plant, plant part, or cell that has not had the sequence of the invention introduced.
[0071] Plant cells possess nuclear, plastid, and mitochondrial genomes. The compositions and methods of the present invention can be used to modify the sequence of the nuclear, plastid, and / or mitochondrial genomes, or can be used to modulate the expression of a gene or genes encoded by the nuclear, plastid, and / or mitochondrial genome. Consequently, by "chromosome" or "chromosomal" is meant the genomic DNA of the nucleus, plastid, or mitochondria. The "genome," as applied to plant cells, encompasses not only the chromosomal DNA found in the nucleus, but also the organelle DNA found Petition 870240102758, dated 02 / 12 / 2024, page 77 / 196 55 / 172 within the subcellular components (e.g., mitochondria or plastids) of the cell. Any nucleotide sequence of interest in a plant cell, organelle, or embryo can be modified using the methods described herein. In specific embodiments, the methods described herein are used to modify a nucleotide sequence encoding an agronomically important trait, such as a plant hormone, plant defense protein, nutrient transport protein, biotic association protein, a desirable input trait, a desirable output trait, a stress resistance gene, a disease / pathogen resistance gene, a male sterility gene, a developmental gene, a regulatory gene, a gene involved in photosynthesis, a DNA repair gene, a transcriptional regulatory gene, or any other polynucleotide and / or polypeptide of interest.Agronomically important traits, such as oil, starch, and protein content, can also be modified. Modifications include increasing the content of oleic acid, saturated and unsaturated oils, increasing lysine and sulfur levels, providing essential amino acids, and also modifying the starch. Modifications of the hordothionin protein are described in US Patents Nos. 5,703,049, 5,885,801, 5,885,802, and 5,990,389, which are incorporated herein by reference. Another example. Petition 870240102758, dated 02 / 12 / 2024, page 78 / 196 56 / 172 is the sulfur- and / or lysine-rich seed protein encoded by soybean 2S albumin described in U.S. Patent No. 5,850,106, and the barley chymotrypsin inhibitor described in Williamson et al. (1987) Eur. J. Biochem. 165:99106, descriptions of which are incorporated herein by reference.
[0072] Derivatives of coding sequences can be made using the methods described herein to increase the level of pre-selected amino acids in the encoded polypeptide. For example, the gene encoding the barley high-lysine polypeptide (BHL) is derived from the barley chymotrypsin inhibitor, US Serial No. 08 / 740682, filed November 1, 1996, and WO 98 / 20133, the descriptions of which are incorporated herein by reference. Other proteins include methionine-rich plant proteins, such as sunflower seeds (Lilley et al. (1989) Proceedings of the World Congress on Vegetable Protein Utilization in Human Foods and Animal Feedstuffs, ed. Applewhite (American Oil Chemists Society, Champaign, Illinois), pp. 497-502; incorporated herein by reference); maize (Pedersen et al. (1986) J. Biol. Chem. 261:6279; Kirihara et al. (1988) Gene 71:359; both incorporated herein by reference); and rice (Musumura et al. (1989) Plant Mol. Biol.12:123, incorporated here by reference). Other agronomically important genes encode latex, starch 2, growth factors, and seed storage factors. Petition 870240102758, dated 02 / 12 / 2024, page 79 / 196 57 / 172 and transcription factors.
[0073] The methods described herein can be used to modify herbicide resistance traits including genes encoding resistance to herbicides that act to inhibit the action of acetolactate synthase (ALS), in particular sulfonylurea-type herbicides (e.g., the acetolactate synthase (ALS) gene containing mutations that lead to this resistance, in particular the S4 and / or Hra mutations), genes encoding resistance to herbicides that act to inhibit the action of glutamine synthase, such as phosphinothricin or basta (e.g., the bar gene); glyphosate (e.g., the EPSPS gene and the GAT gene; see, for example, US Publication No. 20040082770 and WO 03 / 092360); or other genes known in the art. The bar gene encodes resistance to the herbicide basta, the nptII gene encodes resistance to the antibiotics kanamycin and geneticin, and mutants of the ALS gene encode resistance to the herbicide chlorsulfuron.Additional herbicide resistance traits are described, for example, in US Patent Application 2016 / 0208243, which is incorporated herein by reference.
[0074] Sterility genes can also be modified and provide an alternative to physical dehairing. Examples of genes used in this method include genes preferred for male tissues and genes with male sterility phenotypes, such as QM, described in Patent Petition 870240102758, dated 02 / 12 / 2024, page 80 / 196 58 / 172 US No. 5,583,210. Other genes include kinases and those encoding compounds toxic to male or female gametophyte development. Additional sterility traits are described, for example, in US Patent Application 2016 / 0208243, incorporated herein by reference.
[0075] Grain quality can be altered by modifying genes that encode traits such as levels and types of saturated and unsaturated oils, quality and quantity of essential amino acids, and cellulose levels. In maize, modified hordothionin proteins are described in U.S. Patents Nos. 5,703,049, 5,885,801, 5,885,802, and 5,990,389.
[0076] Commercial traits can also be altered by modifying a gene that can, for example, increase starch for ethanol production, or provide protein expression. Another important commercial use of modified plants is the production of polymers and bioplastics, as described in US Patent No. 5,602,321. Genes such as β-ketothiolase, PHBase (polyhydroxybutyrate synthase), and acetoacetyl-CoA reductase (see Schubert et al. (1988) J. Bacteriol. 170: 5837-5847) facilitate the expression of polyhydroxyalkanoates (PHA).
[0077] Exogenous products include enzymes and plant products, as well as those from other sources, including Petition 870240102758, dated 02 / 12 / 2024, page 81 / 196 59 / 172 prokaryotes and other eukaryotes. Such products include enzymes, cofactors, hormones, and the like. The level of proteins, particularly modified proteins having an improved amino acid distribution to enhance the plant's nutritional value, can be increased. This is achieved by expressing such proteins with an enhanced amino acid content.
[0078] The methods described herein may also be used for the insertion of heterologous genes and / or modification of gene expression in native plants to achieve desirable plant traits. Such traits include, for example, disease resistance, herbicide tolerance, drought tolerance, salt tolerance, insect resistance, resistance to parasitic weeds, improved plant nutritional value, improved forage digestibility, higher grain yield, cytoplasmic male sterility, altered fruit ripening, increased storage life of plants or plant parts, reduced allergen production, and increased or decreased lignin content. The genes capable of conferring these desirable traits are described in US Patent Application 2016 / 0208243, incorporated herein by reference. (a) Cpf1 or Csm1 polypeptide
[0079] The methods described herein comprise the introduction into a plant cell, plant organelle or Petition 870240102758, dated 02 / 12 / 2024, p. 82 / 196 60 / 172 plant embryo of at least one Cpcl or Csml polypeptide or a nucleic acid encoding at least one Cpf1 or Csm1 polypeptide, as described herein. In some embodiments, the Cpf1 or Csm1 polypeptide may be introduced into the plant cell, organelle, or plant embryo as an isolated protein. In such embodiments, the Cpf1 or Csm1 polypeptide may further comprise at least one cell-penetrating domain, which facilitates cellular uptake of the protein. In some embodiments, the Cpf1 or Csm1 polypeptide may be introduced into the plant cell, organelle, or plant embryo as a ribonucleoprotein in complex with a guide RNA. In other embodiments, the Cpf1 or Csm1 polypeptide may be introduced into the plant cell, organelle, or plant embryo as an mRNA molecule. In other modes, the Cpf1 or Csm1 polypeptide can be introduced into the plant cell, organelle, or plant embryo as a DNA molecule.In general, the DNA sequences encoding the Cpf1 or Csm1 polypeptide or fusion protein described herein are operationally linked to a promoter sequence that will function in the plant cell, organelle, or plant embryo of interest. The DNA sequence may be linear or the DNA sequence may be part of a vector. In other embodiments, the Cpf1 or Csm1 polypeptide or fusion protein may be introduced into the cell, organelle, or plant embryo as a protein complex. Petition 870240102758, dated 02 / 12 / 2024, page 83 / 196 61 / 172 of RNA comprising guide RNA or a fusion protein and guide RNA.
[0080] In certain embodiments, mRNA encoding the Cppl or Csml polypeptide can be targeted to an organelle (e.g., plastid or mitochondrion). In certain embodiments, mRNA encoding one or more guide RNAs can be targeted to an organelle (e.g., plastid or mitochondrion). In certain embodiments, mRNA encoding the Cpf1 or Csm1 polypeptide and one or more guide RNAs can be targeted to an organelle (e.g., plastid or mitochondrion). Methods for targeting mRNAs to organelles are known in the art (see, for example, U.S. Patent Application 2011 / 0296551; U.S. Patent Application 2011 / 0321187; Gómez and Pallás (2010) PLoS One 5: e12269), and are incorporated herein by reference.
[0081] In certain embodiments, the DNA encoding the Cpf1 or Csm1 polypeptide may further comprise a sequence encoding a guide RNA. In general, each of the sequences encoding the Cpf1 or Csm1 polypeptide and the guide RNA is operationally linked to one or more appropriate promoter control sequences that allow the expression of the Cpf1 or Csm1 polypeptide and the guide RNA, respectively, in the plant cell or plant embryo. The DNA sequence encoding the Cpf1 or Csm1 polypeptide and the guide RNA may further comprise additional control sequences of Petition 870240102758, dated 02 / 12 / 2024, p. 84 / 196 62 / 172 expression, regulatory and / or processing. The DNA sequence encoding the Cpf1 or Csm1 polypeptide and the guide RNA can be linear or can be part of a vector. (b) guide RNA
[0082] The methods described herein may also comprise the introduction, into a plant cell, organelle, or plant embryo, of at least one guide RNA or DNA encoding at least one guide RNA. A guide RNA interacts with the Cpf1 or Csm1 polypeptide to direct the Cpf1 or Csm1 polypeptide to a specific target site, at which site the 5' end of the guide RNA base pairs with a specific protospacer sequence in the plant nucleotide sequence. Guide RNAs may comprise three regions: a first region that is complementary to the target site in the targeted chromosomal sequence, a second region that forms a stem-folding structure, and a third region that remains essentially single-stranded. The first region of each guide RNA is different, so that each guide RNA guides a Cpf1 or Csm1 polypeptide to a specific target site. The second and third regions of each guide RNA may be the same in all guide RNAs.
[0083] A guide RNA region is complementary to a sequence (i.e., protospacer sequence) at the target site in the plant genome, including the nuclear chromosome sequence, as well as plastid or mitochondrial sequences, Petition 870240102758, dated 02 / 12 / 2024, p. 85 / 196 63 / 172 so that the first region of the guide RNA can be base-paired with the target site. In various embodiments, the first region of the guide RNA can comprise from about 8 nucleotides to more than about 30 nucleotides. For example, the base-pairing region between the first region of the guide RNA and the target site in the nucleotide sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 22, about 23, about 24, about 25, about 27, about 30 or more than 30 nucleotides in length. In one exemplary embodiment, the first region of the guide RNA is approximately 23, 24, or 25 nucleotides in length. The guide RNA may also comprise a second region that forms a secondary structure. In some embodiments, the secondary structure comprises a rod or hairpin. The length of the rod may vary.For example, the stem can vary from about 6, to about 10, to about 15, to about 20, to about 25 base pairs in length. The stem may comprise one or more protrusions of 1 to about 10 nucleotides. Thus, the total length of the second region can vary from about 16 to about 25 nucleotides in length. In certain embodiments, the fold is about 5 nucleotides long and the stem comprises about 10 base pairs. Petition 870240102758, dated 02 / 12 / 2024, p. 86 / 196 64 / 172
[0084] The guide RNA may also comprise a third region that remains essentially single-stranded. Thus, the third region has no complementarity with any nucleotide sequence in the cell of interest and has no complementarity with the rest of the guide RNA. The length of the third region may vary. In general, the third region is more than about 4 nucleotides long. For example, the length of the third region may vary from about 5 to about 60 nucleotides long. The combined length of the second and third regions (also called the universal or attachment region) of the guide RNA may vary from about 30 to about 120 nucleotides long. In one aspect, the combined length of the second and third regions of the guide RNA varies from about 40 to about 45 nucleotides long.
[0085] In some embodiments, the guide RNA comprises a single molecule comprising all three regions. In other embodiments, the guide RNA may comprise two separate molecules. The first RNA molecule may comprise the first region of the guide RNA and half of the “stem” of the second region of the guide RNA. The second RNA molecule may comprise the other half of the “stem” of the second region of the guide RNA and the third region of the guide RNA. Thus, in this embodiment, the first and second RNA molecules each contain a sequence of nucleotides that are Petition 870240102758, dated 02 / 12 / 2024, p. 87 / 196 65 / 172 complementary to each other. For example, in one embodiment, each of the first and second RNA molecules comprises a sequence (of about 6 to about 25 nucleotides) that is base-paired to the other sequence to form a functional guide RNA. In specific embodiments, the guide RNA is a single molecule (i.e., crRNA) that interacts with the target site on the chromosome and the Cpf1 polypeptide without the need for a second guide RNA (i.e., a tracrRNA).
[0086] In certain embodiments, the guide RNA can be introduced into the plant cell, organelle, or plant embryo as an RNA molecule. The RNA molecule can be transcribed in vitro. Alternatively, the RNA molecule can be chemically synthesized. In other embodiments, the guide RNA can be introduced into the plant cell, organelle, or embryo as a DNA molecule. In such cases, the DNA encoding the guide RNA can be operationally linked to a promoter control sequence for expression of the guide RNA in the plant cell, organelle, or plant embryo of interest. For example, the RNA coding sequence can be operationally linked to a promoter sequence that is recognized by RNA polymerase III (Pol III). In exemplary embodiments, the RNA coding sequence is linked to a plant-specific promoter.
[0087] The DNA molecule that encodes guide RNA can Petition 870240102758, dated 02 / 12 / 2024, p. 88 / 196 66 / 172 can be linear or circular. In some embodiments, the DNA sequence encoding the guide RNA may be part of a vector. Suitable vectors include plasmid vectors, phagomid vectors, cosmid vectors, minichromosome / artificial chromosome vectors, transposon vectors, and viral vectors. In an exemplary embodiment, the DNA encoding the Cpf1 or Csm1 polypeptide is present in a plasmid vector. Non-limiting examples of suitable plasmid vectors include pUC, pBR322, pET, pBluescript, pCAMBIA, and variants thereof. The vector may comprise additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, and the like.
[0088] In embodiments where the Cpf1 or Csm1 polypeptide and guide RNA are introduced into the cell, organelle, or plant embryo as DNA molecules, each may be part of a separate molecule (e.g., a vector containing the coding sequence for the Cpf1 or Csm1 polypeptide or fusion protein and a second vector containing the coding sequence for the guide RNA) or both may be part of the same molecule (e.g., a vector containing the coding (and regulatory) sequence for the Cpf1 or Csm1 polypeptide or fusion protein and the guide RNA). Petition 870240102758, dated 02 / 12 / 2024, p. 89 / 196 67 / 172 (c) Target site
[0089] A Cpfl or Csm1 polypeptide together with a guide RNA is directed to a target site in a plant, including the chromosomal sequence of a plant, plant cell, plant organelle (e.g., plastid or mitochondrion), or plant embryo, where the Cpfl or Csm1 polypeptide introduces a double-strand break in the chromosomal sequence. The target site has no sequence limitation, except that the sequence is immediately preceded (upstream) by a consensus sequence. This consensus sequence is also known as a proto-adjacent spacer (PAM) motif. Examples of PAM sequences include, but are not limited to, TTN, CTN, TCN, CCN, TTTN, TCTN, TTCN, CTTN, ATTN, TCCN, TTGN, GTTN, CCCN, CCTN, TTAN, TCGN, CTCN, ACTN, GCTN, TCAN, GCCN, and CCGN (where N is defined as any nucleotide). It is well known in the art that the specificity of the PAM sequence for a given nuclease enzyme is affected by the enzyme concentration (Karvelis et al.).(2015) Genome Biol 16: 253). Thus, modulating the concentrations of Cpf1 or Csm1 protein distributed to the cell or in vitro system of interest represents a way to alter the PAM site or sites associated with that Cpf1 or Csm1 enzyme. The modulation of Cpf1 or Csm1 protein concentration in the system of interest can be achieved, for example, by altering the promoter used to express the gene encoding Cpf1 or Csm1. Petition 870240102758, dated 02 / 12 / 2024, pp. 90 / 196 68 / 172, which encodes Csm1, alters the concentration of ribonucleoprotein dispensed to the cell or system in vitro, or adds or removes introns that may play a role in modulating gene expression levels. As detailed here, the first region of the guide RNA is complementary to the protospacer of the target sequence. Typically, the first region of the guide RNA is about 19 to 21 nucleotides long.
[0090] The target site may be in the coding region of a gene, in an intron of a gene, in a control region of a gene, in a non-coding region between genes, etc. The gene may be a protein-coding gene or an RNA-coding gene. The gene may be any gene of interest as described herein. d) Donor Polynucleotide
[0091] In some embodiments, the methods described herein further comprise the introduction of at least one donor polynucleotide into a plant cell, organelle, or plant embryo. A donor polynucleotide comprises at least one donor sequence. In some respects, a donor sequence of the donor polynucleotide corresponds to an endogenous or native plant genomic sequence found in the cell nucleus or in an organelle of interest (e.g., plastid or mitochondrion). For example, the donor sequence may essentially be Petition 870240102758, dated 02 / 12 / 2024, page 91 / 196 69 / 172 identical to a portion of the chromosomal sequence at or near the target site, but comprising at least one nucleotide change. Thus, the donor sequence may comprise a modified version of the wild-type sequence at the target site, such that, upon integration or exchange with the native sequence, the sequence at the targeted location comprises at least one nucleotide change. For example, the change may be an insertion of one or more nucleotides, a deletion of one or more nucleotides, a substitution of one or more nucleotides, or combinations thereof. As a consequence of the integration of the modified sequence, the plant, plant cell, or plant embryo may produce a modified gene product from the targeted chromosomal sequence.
[0092] The donor sequence of the donor polynucleotide may alternatively correspond to an exogenous sequence. As used herein, an exogenous sequence refers to a sequence that is not native to the plant cell, organelle, or embryo, or a sequence whose native location in the genome of the cell, organelle, or embryo is in a different location. For example, the exogenous sequence may comprise a protein-coding sequence, which may be operationally linked to an exogenous promoter control sequence so that, upon integration into the genome, the plant cell or organelle is able to express the Petition 870240102758, dated 02 / 12 / 2024, page 92 / 196 70 / 172 protein encoded by the integrated sequence. For example, the donor sequence can be any gene of interest, such as those encoding agronomically important traits, as described elsewhere here. Alternatively, the exogenous sequence can be integrated into the chromosomal sequence of the nucleus, plastid, and / or mitochondria so that its expression is regulated by an endogenous promoter control sequence. In other iterations, the exogenous sequence can be a transcriptional control sequence, another expression control sequence, or an RNA coding sequence. The integration of an exogenous sequence into a chromosomal sequence is termed “knock-in.” The donor sequence can vary in length from several nucleotides to hundreds of nucleotides to hundreds of thousands of nucleotides.
[0093] In some embodiments, the donor sequence in the donor polynucleotide is flanked by an upstream sequence and a downstream sequence, which have substantial sequence identity with the sequences located upstream and downstream, respectively, of the target site in the plant nucleus, plastid, and / or mitochondrial genomic sequence. Due to these sequence similarities, the upstream and downstream sequences of the donor polynucleotide allow homologous recombination between the donor polynucleotide and the target sequence, such that the donor sequence Petition 870240102758, dated 02 / 12 / 2024, p. 93 / 196 71 / 172 can be integrated (or swapped with) with the targeted plant sequence.
[0094] The upstream sequence, as used herein, refers to a nucleic acid sequence that shares substantial sequence identity with a chromosomal sequence upstream of the targeted site. Similarly, the downstream sequence refers to a nucleic acid sequence that shares substantial sequence identity with a chromosomal sequence downstream of the target site. As used herein, the phrase “substantial sequence identity” refers to sequences having at least about 75% sequence identity. Thus, the upstream and downstream sequences in the donor polynucleotide may have approximately 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the upstream or downstream sequence to the target site.In one exemplary embodiment, the upstream and downstream sequences in the donor polynucleotide may have approximately 95% or 100% sequence identity with nucleotide sequences upstream or downstream of the target site. In another embodiment, the upstream sequence shares substantial sequence identity with a nucleotide sequence located immediately upstream of the target site (i.e., adjacent to the target site). In others... Petition 870240102758, dated 02 / 12 / 2024, p. 94 / 196 In 72 / 172 embodiments, the upstream sequence shares substantial sequence identity with a nucleotide sequence that is located within about one hundred (100) nucleotides upstream of the targeted site. Thus, for example, the upstream sequence may share substantial sequence identity with a nucleotide sequence that is located about 1 to about 20, about 21 to about 40, about 41 to about 60, about 61 to about 80, or about 81 to about nucleotides upstream of the targeted site. In one embodiment, the downstream sequence shares substantial sequence identity with a nucleotide sequence located immediately downstream of the targeted site (i.e., adjacent to the targeted site). In other embodiments, the downstream sequence shares substantial sequence identity with a nucleotide sequence that is located within about one hundred (100) nucleotides downstream of the targeted site.Thus, for example, the downstream sequence may share substantial sequence identity with a nucleotide sequence that is located approximately 1 to approximately 20, approximately 21 to approximately 40, approximately 41 to approximately 60, approximately 61 to approximately 80, or approximately 81 to approximately 80 nucleotides downstream from the targeted site.
[0095] Each upstream or downstream sequence can vary in length from about 20 nucleotides to about Petition 870240102758, dated 02 / 12 / 2024, pp. 95 / 196 73 / 172 5000 nucleotides. In some forms, the upstream and downstream sequences may comprise approximately 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2800, 3000, 3200, 3400, 3600, 3800, 4000, 4200, 4400, 4600, 4800, or 5000. nucleotides. In exemplary embodiments, the upstream and downstream sequences can vary in length from about 50 to about 1500 nucleotides.
[0096] Donor polynucleotides comprising upstream and downstream sequences with sequence similarity to the target nucleotide sequence may be linear or circular. In embodiments where the donor polynucleotide is circular, it may be part of a vector. For example, the vector may be a plasmid vector.
[0097] In certain embodiments, the donor polynucleotide may additionally comprise at least one targeted cleavage site that is recognized by the Cpf1 or Csm1 polypeptide. The targeted cleavage site added to the donor polynucleotide may be placed upstream or downstream or both upstream and downstream of the donor sequence. For example, the donor sequence may be flanked by targeted cleavage sites such that, upon cleavage by the Cpf1 or Csm1 polypeptide, the donor sequence is flanked by overhangs that are compatible with those in the sequence. Petition 870240102758, dated 02 / 12 / 2024, pp. 96 / 196 74 / 172 nucleotides generated in the cleavage by the Cpfl or Csml polypeptide. Therefore, the donor sequence can be linked with the nucleotide sequence cleaved during double-strand break repair by a non-homologous repair process. Generally, the donor polynucleotides comprising the targeted cleavage site(s) are circular (e.g., they may be part of a plasmid vector).
[0098] The donor polynucleotide may be a linear molecule comprising a short donor sequence with optional short overhangs that are compatible with the overhangs generated by the Cpf1 or Csm1 polypeptide. In such embodiments, the donor sequence may be directly ligated to the chromosomal sequence cleaved during double-strand break repair. In some cases, the donor sequence may be less than about 1,000, less than about 500, less than about 250, or less than about 100 nucleotides. In certain cases, the donor polynucleotide may be a linear molecule comprising a short donor sequence with blunt ends. In other iterations, the donor polynucleotide may be a linear molecule comprising a short donor sequence with 5' and / or 3' overhangs. The overhangs may comprise 1, 2, 3, 4, or 5 nucleotides.
[0099] In some forms, the polynucleotide Petition 870240102758, dated 02 / 12 / 2024, pp. 97 / 196 75 / 172 The donor will be DNA. The DNA may be single-stranded or double-stranded and / or linear or circular. The donor polynucleotide may be a DNA plasmid, a bacterial artificial chromosome (BAC), a yeast artificial chromosome (YAC), a viral vector, a linear piece of DNA, a PCR fragment, a naked nucleic acid, or a nucleic acid complexed with a dispensing vehicle, such as a liposome or poloxamer. In certain embodiments, the donor polynucleotide comprising the donor sequence may be part of a plasmid vector. In any of these situations, the donor polynucleotide comprising the donor sequence may further comprise at least one additional sequence. (e) Introduction into the plant cell
[00100] The Cpf1 or Csm1 polypeptide (or the coding nucleic acid), guide RNA(s), or optional donor polynucleotide(s) may be introduced into a plant cell, organelle, or plant embryo by a variety of means, including transformation. Transformation protocols, as well as protocols for introducing polypeptide or polynucleotide sequences into plants, may vary depending on the type of plant or plant cell, i.e., monocotyledons or dicotyledons, targeted for transformation. Suitable methods for introducing polypeptides and polynucleotides into plant cells Petition 870240102758, dated 02 / 12 / 2024, pp. 98 / 196 76 / 172 include microinjection (Crossway et al. (1986) Biotechniques 4:320-334), electroporation (Riggs et al. (1986) Proc. Natl. Acad. Sci. USA 83:5602-5606), Agrobacterium-mediated transformation (US Patent No. 5,563,055 and US Patent No. 5,981,840), direct gene transfer (Paszkowski et al. (1984) EMBO J. 3:2717-2722), and ballistic particle acceleration (see, for example, US Patents Nos. 4,945,050; US Patent No. 5,879,918; US Patents Nos. 5,886,244; and 5,932,782; Tomes et al. (1995) in Plant Cell, Tissue, and Organ Culture: Fundamental Methods, ed. Gamborg and Phillips (Springer-Verlag, Berlin); McCabe et al. (1988) Biotechnology 6:923-926); 22:421-477; Sanford et al. (1987) Particulate Science and Technology 5:27-37 (onion); Christou et al. (1988) Plant Physiol. Vitro Cell Dev.27P:175-182 (soy); Singh et al. (1998) Theor. Appl. Genet. 96:319-324 (soybean); Datta et al. (1990) Biotechnology 8:736-740 (rice); Klein et al. (1988) Proc. Natl. Acad. Sci. USA 85:4305-4309 (corn); Klein et al. (1988) Biotechnology 6:559-563 (corn); US Patent Nos5,240,855; 5,322,783; e, 5,324,646; Klein et al. (1988) Plant Physiol. 91:440–444 (corn); Fromm et al. (1990). Petition 870240102758, of 02 / 12 / 2024, p. 99 / 196 77 / 172 Biotechnology 8: 833-839 (corn); Hooykaas-Van Slogteren et al. (1984) Nature (London) 311:763-764; US Patent No5. 736,369 (cereals); Bytebier et al. (1987) Proc. Natl. Academic. Sci. USA 84:5345-5349 (Liliaceae); De Wet et al. (1985) in The Experimental Manipulation of Ovule Tissues, ed. Chapman et al. (Longman, New York), pp. 197-209 (pollen); Kaeppler et al. (1990) Plant Cell Reports 9:415-418 and Kaeppler et al. (1992) Theor. Appl. Genet. 84:560-566 (vibrissae-mediated transformation); D'Halluin et al. (1992) Plant Cell 4:1495-1505 (electroporation); Li et al. (1993) Plant Cell Reports 12:250-255 and Christou and Ford (1995) Annals of Botany 75:407-413 (rice); Osjoda et al. (1996) Nature Biotechnology 14:745-750 (maize via Agrobacterium tumefaciens); all of which are incorporated herein by reference.Site-specific genome editing of plant cells by biolistic introduction of a ribonucleoprotein comprising a nuclease and a suitable guide RNA has been demonstrated (Svitashev et al (2016) Nat Commun doi: 10.1038 / ncomms13274); these methods are incorporated here by reference. “Stable transformation” is intended to mean that the nucleotide construct introduced into a plant integrates into the plant genome and is capable of being inherited by its offspring. The nucleotide construct can be integrated into the plant's nuclear, plastid, or mitochondrial genome. The methods for this... Petition 870240102758, dated 02 / 12 / 2024, pages 100 / 196 78 / 172 plastid transformation methods are known in the art (see, for example, Chloroplast Biotechnology: Methods and Protocols (2014) Pal Maliga, ed. and US Patent Application 2011 / 0321187), and methods for plant mitochondrial transformation have been described in the art (see, for example, US Patent Application 2011 / 0296551), which is incorporated herein by reference.
[00101] The cells that have been transformed can be grown into plants (i.e., cultivated) according to conventional methods. See, for example, McCormick et al. (1986) Plant Cell Reports 5:81-84. In this way, the present invention provides transformed seed (also referred to as “transgenic seed”) having a nucleic acid modification stably incorporated into its genome.
[00102] “Introduced in the context of inserting a nucleic acid fragment (e.g., a recombinant DNA construct) into a cell, it means “transfection” or “transformation” or “transduction” and includes reference to the incorporation of a nucleic acid fragment into a plant cell where the nucleic acid fragment may be incorporated into the cell's genome (e.g., nuclear chromosome, plasmid, plastid chromosome, or mitochondrial chromosome), converted into a stand-alone replicon, or transiently expressed (e.g., transfected mRNA). Petition 870240102758, dated 02 / 12 / 2024, page 101 / 196 79 / 172
[00103] The present invention can be used for the transformation of any plant species, including, but not limited to, monocotyledons and dicotyledons (i.e., monocotyledons and dicotyledons, respectively). Examples of plant species of interest include, but are not limited to, maize (Zea mays), Brassica sp. (e.g., B. napus, B. rapa, B. juncea), particularly those Brassica species useful as sources of seed oil, alfalfa (Medicago sativa), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), millet (e.g., pearl millet (Pennisetum glaucum)), sorghum (Panicum miliaceum), pearl millet (Setaria italica), goosegrass (Eleusine coracana), sunflower (Helianthus annuus), safflower (Carthamus tinctorius), durum wheat (Triticum aestivum), soybean (Glycine max), tobacco (Nicotiana tabacum), potato (Solanum tuberosum), peanut (Arachis hypogaea), cotton (Gossypium barbadense, Gossypium hirsutum), sweet potato (Ipomoea batatus), cassava (Manihot esculenta), coffee (Coffea spp.), coconut (Cocos nucifera), pineapple (Ananas comosus), citrus trees (Citrus spp.)), cocoa (Theobroma cacao), tea tree (Camellia sinensis), banana (Musa spp.), avocado (Persea americana), fig tree (Ficus casica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia nut (Macadamia). Petition 870240102758, dated 02 / 12 / 2024, page 102 / 196 80 / 172 integrifolia), almond (Prunus amygdalus), sugar beet (Beta vulgaris), sugar cane (Saccharum spp.), oil palm (Elaeis guineensis), poplar (Populus spp.), eucalyptus (Eucalyptus spp.), oat (Avena sativa), barley (Hordeum vulgare), vegetables, ornamental plants and conifers.
[00104] Cpf1 or Csm1 polypeptides (or those encoding nucleic acid), guide RNA(s) (or DNAs encoding guide RNA), and optional donor polynucleotide(s) may be introduced into the plant cell, organelle, or plant embryo simultaneously or sequentially. The ratio of Cpf1 polypeptides (or encoding nucleic acid) to guide RNA(s) (or encoding DNA) will generally be approximately stoichiometric so that the two components can form an RNA-protein complex with the target DNA. In one embodiment, the DNA encoding a Cpf1 or Csm1 polypeptide and the DNA encoding a guide RNA are dispensed together within the plasmid vector.
[00105] The compositions and methods described herein can be used to alter the expression of genes of interest in a plant, such as genes involved in photosynthesis. Therefore, the expression of a gene encoding a protein involved in photosynthesis can be modulated compared to a control plant. A “target plant or Petition 870240102758, dated 02 / 12 / 2024, page 103 / 196 81 / 172 “plant cell” is one in which the genetic alteration, such as a mutation, has been effected in relation to a gene of interest, or is a plant or plant cell that is descended from such an altered plant or cell and that comprises the alteration. A “control” or “control plant” or “control plant cell” provides a reference point for measuring changes in the phenotype of the plant or plant cell in question. Thus, the expression levels are higher or lower than those of the control plant, depending on the methods of the invention.
[00106] A control plant or plant cell may comprise, for example: (a) a wild-type plant or cell, that is, of the same genotype as the starting material for the genetic alteration that resulted in the plant or cell in question; (b) a plant or plant cell of the same genotype as the starting material, but which has been transformed with a null construct (that is, with a construct that has no known effect on the trait of interest, such as a construct comprising a marker gene); (c) a plant or plant cell that is an untransformed segregant among offspring of a plant or plant cell in question; (d) a plant or plant cell genetically identical to the plant or plant cell in question, but which is not exposed to conditions or stimuli that would induce expression of the gene of interest; or (e) a Petition 870240102758, dated 02 / 12 / 2024, page 104 / 196 82 / 172 plant or the plant cell itself in question, under conditions in which the gene of interest is not expressed.
[00107] Although the invention is described in terms of transformed plants, it is recognized that the transformed organisms of the invention also include plant cells, plant protoplastids, plant cell tissue cultures from which plants can be regenerated, plant calluses, plant clusters, and plant cells that are intact in plants or plant parts, such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, grains, ears, cobs, husks, culms, roots, root tips, anthers, and the like. Grain is intended to mean mature seed produced by commercial growers for purposes other than the growth or reproduction of the species. Offspring, variants, and mutants of regenerated plants are also included within the scope of the invention, provided that these parts comprise the introduced polynucleotides.
[00108] (f) Method for Using a Fusion Protein to Modify a Plant Sequence or Regulate the Expression of a Plant Sequence
[00109] The methods described herein also cover the modification of a nucleotide sequence or the regulation of the expression of a nucleotide sequence in a plant cell, plant organelle, or plant embryo. The methods Petition 870240102758, dated 02 / 12 / 2024, pp. 105 / 196 83 / 172 may comprise the introduction into a plant cell or plant embryo of at least one fusion protein or nucleic acid encoding at least one fusion protein, wherein the fusion protein comprises a Cpf1 or Csm1 polypeptide or a fragment or variant thereof and an effector domain, and (b) at least one guide RNA or DNA encoding the guide RNA, wherein the guide RNA guides the Cpf1 or Csm1 polypeptide of the fusion protein to a target site in the chromosome sequence and the effector domain of the fusion protein modifies the chromosome sequence or regulates the expression of the chromosome sequence.
[00110] Fusion proteins comprising a Cpf1 or Csm1 polypeptide or a fragment thereof or variant thereof and an effector domain are described herein. In general, the fusion proteins described herein may further comprise at least one nuclear localization signal, plastid signal peptide, mitochondrial signal peptide, or signal peptide capable of trafficking proteins to multiple subcellular locations. Nucleic acids encoding fusion proteins are described herein. In some embodiments, the fusion protein may be introduced into the cell or embryo as an isolated protein (which may further comprise a cell penetration domain). In addition, the isolated fusion protein may be part of a protein-RNA complex comprising guide RNA. In other Petition 870240102758, dated 02 / 12 / 2024, pp. 106 / 196 In 84 / 172 embodiments, the fusion protein can be introduced into the cell or embryo as an RNA molecule (which may be capped and / or polyadenylated). In still other embodiments, the fusion protein can be introduced into the cell or embryo as a DNA molecule. For example, the fusion protein and guide RNA can be introduced into the cell or embryo as discrete DNA molecules or as part of the same DNA molecule. Such DNA molecules may be plasmid vectors.
[00111] In some embodiments, the method further comprises the introduction into the cell, organelle or embryo of at least one donor polynucleotide as described elsewhere herein. Means for introducing molecules into plant cells, organelles or plant embryos, as well as means for culturing cells (including cells comprising organelles) or embryos are described herein.
[00112] In certain embodiments, where the effector domain of the fusion protein is a cleavage domain, the method may comprise introducing into a plant cell, organelle, or plant embryo a fusion protein (or nucleic acid encoding a fusion protein) and two guide RNAs (or DNA encoding two guide RNAs). The two guide RNAs direct the fusion protein to two different target sites in the chromosome sequence, where the fusion protein dimerizes (i.e., forms a homodimer) so that the two Petition 870240102758, dated 02 / 12 / 2024, page 107 / 196 85 / 172 cleavage domains can introduce a double-strand break in the chromosomal sequence. In embodiments where the optional donor polynucleotide is not present, the double-strand break in the chromosomal sequence can be repaired by a non-homologous end-join (NHEJ) repair process. Because NHEJ is error-prone, deletions of at least one nucleotide, insertions of at least one nucleotide, substitutions of at least one nucleotide, or combinations thereof, can occur during break repair. Consequently, the targeted chromosomal sequence can be modified or inactivated. For example, a single nucleotide change (SNP) can give rise to an altered protein product, or a frameshift in a coding sequence can inactivate or knock out the sequence so that no protein product is produced.In embodiments where the optional donor polynucleotide is present, the donor sequence in the donor polynucleotide can be exchanged or integrated into the chromosomal sequence at the targeted site during double-strand break repair. For example, in embodiments where the donor sequence is flanked by upstream and downstream sequences having substantial sequence identity with upstream and downstream sequences, respectively, of the targeted site in the chromosomal sequence, the donor sequence can be exchanged or integrated. Petition 870240102758, dated 02 / 12 / 2024, p. 108 / 196 86 / 172 integrated into the chromosome sequence at the targeted site during repair mediated by the homology-directed repair process. Alternatively, in embodiments where the donor sequence is flanked by compatible overhangs (or the compatible overhangs are generated in situ by the Cpf1 or Csm1 polypeptide), the donor sequence can be directly linked to the cleaved chromosome sequence by a non-homologous repair process during double-strand break repair. The exchange or integration of the donor sequence into the chromosome sequence modifies the targeted chromosome sequence or introduces an exogenous sequence into the chromosome sequence of the cell, organelle, or plant embryo.
[00113] In other embodiments where the effector domain of the fusion protein is a cleavage domain, the method may comprise the introduction into the plant cell, organelle, or plant embryo of two different fusion proteins (or nucleic acid encoding two different fusion proteins) and two guide RNAs (or DNA encoding two guide RNAs). The fusion proteins may differ as detailed elsewhere here. Each guide RNA directs a fusion protein to a specific target site in the chromosome sequence, where the fusion proteins may dimerize (e.g., form a heterodimer) so that the two cleavage domains can introduce a double-strand break in the chromosome sequence. In embodiments where the Petition 870240102758, dated 02 / 12 / 2024, pp. 109 / 196 87 / 172 optional donor polynucleotide is not present, the resulting double-strand breaks can be repaired by a non-homologous repair process so that deletions of at least one nucleotide, insertions of at least one nucleotide, substitutions of at least one nucleotide, or combinations thereof can occur during break repair.In embodiments where the optional donor polynucleotide is present, the donor sequence in the donor polynucleotide can be exchanged or integrated into the chromosomal sequence during double-strand break repair by a homology-based repair process (e.g., in embodiments where the donor sequence is flanked by upstream and downstream sequences having substantial sequence identity with upstream and downstream sequences, respectively, of the targeted sites in the chromosomal sequence) or a non-homologous repair process (e.g., in embodiments where the donor sequence is flanked by compatible overhangs).
[00114] In certain embodiments where the effector domain of the fusion protein is a transcription activation domain or a transcription repressor domain, the method may comprise introducing into the protein of a plant cell, organelle, or plant embryo a fusion protein (or nucleic acid encoding a fusion protein) and a guide RNA. Petition 870240102758, dated 02 / 12 / 2024, page 110 / 196 88 / 172 (or DNA encoding a guide RNA). The guide RNA directs the fusion protein to a specific chromosomal sequence, where the transcription activation domain or a transcription repressor domain activates or represses the expression, respectively, of a gene or genes located near the target chromosomal sequence. That is, transcription can be affected by genes near the target chromosomal sequence or it can be affected by genes located at a greater distance from the target chromosomal sequence. It is well known in the art that gene transcription can be regulated by distantly located sequences that may be located thousands of bases from the transcription start site or even on a separate chromosome (Harmston and Lenhard (2013) Nucleic Acids Res 41:7185-7199).
[00115] In alternative embodiments in which the effector domain of the fusion protein is an epigenetic modification domain, the method may comprise introducing into a plant cell, organelle, or plant embryo a fusion protein (or nucleic acid encoding a fusion protein) and a guide RNA (or DNA encoding a guide RNA). The guide RNA directs the fusion protein to a specific chromosome sequence, where the epigenetic modification domain modifies the structure of the targeted chromosome sequence. Epigenetic modifications include acetylation, histone protein methylation, and / or methylation. Petition 870240102758, dated 02 / 12 / 2024, page 111 / 196 89 / 172 of nucleotides. In some cases, structural modification of the chromosome sequence leads to changes in the expression of the chromosome sequence. V. Plants and Plant Cells Understanding Genetic Modification
[00116] Plants, plant cells, plant organelles and plant embryos comprising at least one nucleotide sequence that has been modified using a protein-mediated process mediated by Cppl or Csm1 polypeptide or by a fusion protein as described herein are provided herein. Plant cells, organelles and plant embryos comprising at least one DNA or RNA molecule encoding Cppl or Csm1 polypeptide or a fusion protein targeting a chromosomal sequence of interest or a fusion protein, at least one guide RNA and optionally one or more donor polynucleotide(s) are also provided. The genetically modified plants described herein may be heterozygous for the modified nucleotide sequence or homozygous for the modified nucleotide sequence. Plant cells comprising one or more genetic modifications in organellar DNA may be heteroplasmic or homoplasmic.
[00117] The modified chromosomal sequence of a plant, plant organelle, or plant cell can be modified so that it is inactivated, has expression Petition 870240102758, dated 02 / 12 / 2024, page 112 / 196 90 / 172 positively regulated or negatively regulated, or produces an altered protein product or comprises an integrated sequence. The modified chromosome sequence may be inactivated so that the sequence is not transcribed and / or a functional protein product is not produced. Thus, a genetically modified plant comprising an inactivated chromosome sequence may be termed a “knockout” or a “conditional knockout.” The inactivated chromosome sequence may include a deletion mutation (i.e., deletion of one or more nucleotides), an insertion mutation (i.e., insertion of one or more nucleotides), or a nonsense mutation (i.e., substitution of a single nucleotide for another nucleotide so that a stop codon is introduced). As a consequence of the mutation, the targeted chromosome sequence is inactivated and a functional protein is not produced. The inactivated chromosome sequence does not comprise an exogenously introduced sequence.Also included here are genetically modified plants in which two, three, four, five, six, seven, eight, nine, or ten or more chromosomal sequences are inactivated.
[00118] The modified chromosome sequence can also be altered so that it codes for a variant protein product. For example, a genetically modified plant comprising a modified chromosome sequence can Petition 870240102758, dated 02 / 12 / 2024, page 113 / 196 91 / 172 comprise a targeted point mutation(s) or other modification such that a changed protein product is produced. In one embodiment, the chromosome sequence may be modified so that at least one nucleotide is changed and the expressed protein comprises a changed amino acid residue (missense mutation). In another embodiment, the chromosome sequence may be modified to comprise more than one missense mutation, so that more than one amino acid is altered. Additionally, the chromosome sequence may be modified to have a deletion or insertion of three nucleotides so that the expressed protein comprises a deletion or insertion of a single amino acid. The altered protein or variant may have altered properties or activities compared to the wild-type protein, such as altered substrate specificity, altered enzymatic activity, altered kinetic rates, etc.
[00119] In some embodiments, the genetically modified plant may comprise at least one chromosomally integrated nucleotide sequence. A genetically modified plant comprising an integrated sequence may be termed a “knock-in” or “conditional knock-in.” The nucleotide sequence that is the integrated sequence may, for example, encode an orthologous protein, an endogenous protein, or combinations of both. In Petition 870240102758, dated 02 / 12 / 2024, page 114 / 196 92 / 172 In one embodiment, a sequence encoding an orthologous protein or an endogenous protein may be integrated into a nuclear or organelle chromosomal sequence encoding a protein such that the chromosomal sequence is inactivated, but the exogenous sequence is expressed. In this case, the sequence encoding the orthologous protein or endogenous protein may be operationally linked to a promoter control sequence. Alternatively, a sequence encoding an orthologous protein or an endogenous protein may be integrated into a nuclear or organelle chromosomal sequence without affecting the expression of a chromosomal sequence. For example, a sequence encoding a protein may be integrated into a “safe haven” location. The present description also covers genetically modified plants in which two, three, four, five, six, seven, eight, nine, or ten or more sequences, including sequences encoding protein(s), are integrated into the genome.Any gene of interest as described herein can be introduced integrated into the chromosomal sequence of the plant nucleus or organelle. In particular embodiments, genes that increase plant growth or yield are integrated into the chromosome.
[00120] The chromosomally integrated sequence encoding a protein may encode the wild-type form of a protein of interest or it may encode a Petition 870240102758, dated 02 / 12 / 2024, pp. 115 / 196 93 / 172 protein comprising at least one modification such that an altered version of the protein is produced. For example, a chromosomally integrated sequence encoding a protein related to a disease or disorder may comprise at least one modification such that the altered version of the protein produced causes or enhances the associated disorder. Alternatively, the chromosomally integrated sequence encoding a protein related to a disease or disorder may comprise at least one modification such that the altered version of the protein protects the plant against the development of the associated disease or disorder.
[00121] In certain embodiments, the genetically modified plant may comprise at least one modified chromosomal sequence encoding a protein, such that the protein expression pattern is altered. For example, regulatory regions that control protein expression, such as a promoter or a transcription factor binding site, may be altered so that the protein is overexpressed or the tissue-specific or temporal expression of the protein is altered, or a combination thereof. Alternatively, the protein expression pattern may be altered using a conditional knockout system. A non-limiting example of a conditional knockout system includes a Cre-lox recombination system. Petition 870240102758, dated 02 / 12 / 2024, pages 116 / 196 The 94 / 172 Cre-lox recombination system comprises a Cre recombinase enzyme, a site-specific DNA recombinase that can catalyze the recombination of a nucleic acid sequence between specific sites (lox sites) in a nucleic acid molecule. Methods of using this system to produce temporal and tissue-specific expression are known in the art. VI. Methods for Modifying a Nucleotide Sequence in a Non-Plant Eukaryotic Genome and in Non-Plant Eukaryotic Cells Comprising a Genetic Modification
[00122] Methods are provided herein for modifying a nucleotide sequence of a non-plant eukaryotic cell or non-plant eukaryotic organelle. The methods comprise introducing into a targeted cell or organelle a DNA-targeted RNA or a DNA polynucleotide encoding a DNA-targeted RNA, wherein the DNA-targeted RNA comprises: (a) a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA; and (b) a second segment that interacts with a Cpf1 or Csm1 polypeptide, and also introducing into the target cell or organelle a Cpf1 or Csm1 polypeptide, or a polynucleotide encoding a Cpf1 or Csm1 polypeptide, wherein the Cpf1 or Csm1 polypeptide comprises: (a) an RNA-binding portion that Petition 870240102758, dated 02 / 12 / 2024, p. 117 / 196 95 / 172 interacts with RNA directed to DNA; and (b) an activity moiety that exhibits site-directed enzymatic activity. The target cell or organelle can then be cultured under conditions in which the chimeric nuclease polypeptide is expressed and cleaves the nucleotide sequence. Note that the system described here does not require the addition of exogenous Mg+2 or any other ions. Finally, a non-plant eukaryotic cell or organelle comprising the modified nucleotide sequence can be selected.
[00123] In some embodiments, the method may comprise the introduction of a Cpf1 or Csm1 polypeptide (or coding nucleic acid) and a guide RNA (or coding DNA) into a non-plant eukaryotic cell or organelle wherein the Cpf1 or Csm1 polypeptide introduces a double-strand break in the target nucleotide sequence of the nuclear or organelle chromosomal DNA. In some embodiments, the method may comprise the introduction of a Cpf1 or Csm1 polypeptide (or coding nucleic acid) and at least one guide RNA (or coding DNA) into a non-plant eukaryotic cell or organelle wherein the Cpf1 or Csm1 polypeptide introduces more than one double-strand break (i.e., two, three, or more than three double-strand breaks) in the target nucleotide sequence of the nuclear or organelle chromosomal DNA. In embodiments where an optional donor polynucleotide is not present, the break Petition 870240102758, dated 02 / 12 / 2024, pp. 118 / 196 96 / 172 double-strand breaks in the nucleotide sequence can be repaired by a non-homologous end-join (NHEJ) repair process. Because NHEJ is error-prone, deletions of at least one nucleotide, insertions of at least one nucleotide, substitutions of at least one nucleotide, or combinations thereof, can occur during break repair. Consequently, the targeted nucleotide sequence may be modified or inactivated. For example, a single nucleotide change (SNP) may give rise to an altered protein product, or a frameshift in a coding sequence may inactivate or knock out the sequence so that no protein product is produced. In embodiments where an optional donor polynucleotide is present, the donor sequence in the donor polynucleotide may be exchanged or integrated into the nucleotide sequence at the target site during double-strand break repair.For example, in modes where the donor sequence is flanked by upstream and downstream sequences having substantial sequence identity with upstream and downstream sequences, respectively, of the target site in the nucleotide sequence of the non-eukaryotic cell or organelle, the donor sequence can be exchanged or integrated into the nucleotide sequence at the targeted site during repair mediated by the targeted repair process. Petition 870240102758, dated 02 / 12 / 2024, p. 119 / 196 97 / 172 by homology. Alternatively, in embodiments where the donor sequence is flanked by compatible overhangs (or the compatible overhangs are generated in situ by the Cppl or Csm1 polypeptide), the donor sequence can be directly linked to the cleaved nucleotide sequence by a non-homologous repair process during double-strand break repair. The exchange or integration of the donor sequence into the nucleotide sequence modifies the target nucleotide sequence or introduces an exogenous sequence into the target nucleotide sequence of the non-plant eukaryotic cell or organelle.
[00124] In some embodiments, double-strand breaks caused by the action of nucleases or Cppl or Csm1 nucleases are repaired so that DNA is deleted from the chromosome of the non-plant eukaryotic cell or organelle. In some embodiments, one base, a few bases (i.e., 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases), or a large section of DNA (i.e., more than 10, more than 50, more than 100, or more than 500 bases) is deleted from the chromosome of the non-plant eukaryotic cell or organelle.
[00125] In some embodiments, the expression of non-plant eukaryotic genes can be modulated as a result of double-strand breaks caused by the nuclease or nucleases Cpf1 or Csm1. In some embodiments, the expression of non-plant eukaryotic genes Petition 870240102758, dated 02 / 12 / 2024, p. 120 / 196 98 / 172 can be modulated by variant Cpfl or Csm1 enzymes comprising a mutation that renders the Cpfl or Csm1 nuclease incapable of producing a double-strand break. In some preferred embodiments, the variant Cpfl or Csm1 nuclease comprising a mutation that renders the Cpfl or Csm1 nuclease incapable of producing a double-strand break can be fused to a transcription activation or transcription repression domain.
[00126] In some embodiments, a eukaryotic cell comprising mutations in its nuclear and / or organelle chromosomal DNA caused by the action of a Cpf1 or Csm1 nuclease or nucleases is cultured to produce a eukaryotic organism. In some embodiments, a eukaryotic cell in which gene expression is modulated as a result of one or more Cpf1 or Csm1 nucleases, or one or more variant Cpf1 or Csm1 nucleases, is cultured to produce a eukaryotic organism. Methods for culturing non-plant eukaryotic cells to produce eukaryotic organisms are known in the art, for example, in U.S. Patent Applications 2016 / 0208243 and 2016 / 0138008, which are incorporated herein by reference.
[00127] The present invention can be used for the transformation of any eukaryotic species, including, but not limited to, animals (including but not limited to mammals, insects, fish, birds and reptiles), fungi, Petition 870240102758, dated 02 / 12 / 2024, pp. 121 / 196 99 / 172 amoebas and yeasts.
[00128] Methods for introducing nuclease proteins, DNA or RNA molecules encoding nuclease proteins, DNA or guide RNA molecules encoding guide RNAs, and DNA molecules of optional donor sequences into non-plant eukaryotic cells or organelles are known in the art, for example, in U.S. Patent Application 2016 / 0208243, incorporated herein by reference. Exemplary genetic modifications for non-plant eukaryotic cells or organelles that may have particular value for industrial applications are also known in the art, for example, in U.S. Patent Application 2016 / 0208243, incorporated herein by reference. VII. Methods for Modifying a Nucleotide Sequence in a Prokaryotic Genome and Prokaryotic Cells Comprising a Genetic Modification
[00129] Methods are provided here for modifying a nucleotide sequence of a prokaryotic cell (e.g., bacterial or archaeal). The methods comprise introducing into a target cell a DNA-directed RNA or a DNA polynucleotide encoding a DNA-directed RNA, wherein the DNA-directed RNA comprises: (a) a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA; and (b) a second segment that interacts with a polypeptide Petition 870240102758, dated 02 / 12 / 2024, page 122 / 196 100 / 172 Cpf1 or Csm1, and also introduces into the target cell a Cpf1 or Csm1 polypeptide, or a polynucleotide encoding a Cpf1 or Csm1 polypeptide, wherein the Cpf1 or Csm1 polypeptide comprises: (a) an RNA-binding portion that interacts with RNA directed to DNA; and (b) an activity portion that exhibits site-directed enzymatic activity. The target cell can then be cultured under conditions in which the Cpf1 or Csm1 polypeptide is expressed and cleaves the nucleotide sequence. Note that the system described herein does not require the addition of exogenous Mg+2 or any other ions. Finally, prokaryotic cells comprising the modified nucleotide sequence can be selected.It is further noted that the prokaryotic cells comprising the modified nucleotide sequence(s) are not the natural host cells of the polynucleotides encoding the Cpf1 or Csm1 polypeptide of interest, and that a non-naturally occurring guide RNA is used to effect the desired changes in the prokaryotic nucleotide sequence(s). It is further noted that the targeted DNA may be present as part of the prokaryotic chromosome(s) or may be present in one or more plasmids or other non-chromosomal DNA molecules in the prokaryotic cell.
[00130] In some embodiments, the method may involve the introduction of a Cpf1 or Csm1 polypeptide (or Petition 870240102758, dated 02 / 12 / 2024, pp. 123 / 196 101 / 172 coding nucleic acid) and a guide RNA (or coding DNA) into a prokaryotic cell wherein the Cpf1 or Csm1 polypeptide introduces a double-strand break in the target nucleotide sequence of the prokaryotic cellular DNA. In some embodiments, the method may comprise the introduction of a Cpf1 or Csm1 polypeptide (or coding nucleic acid) and at least one guide RNA (or coding DNA) into a prokaryotic cell wherein the Cpf1 or Csm1 polypeptide introduces more than one double-strand break (i.e., two, three, or more than three double-strand breaks) in the target nucleotide sequence of the prokaryotic cellular DNA. In embodiments where an optional donor polynucleotide is not present, the double-strand break in the nucleotide sequence may be repaired by a non-homologous end-join (NHEJ) repair process.Because NHEJ is error-prone, deletions of at least one nucleotide, insertions of at least one nucleotide, substitutions of at least one nucleotide, or combinations thereof, can occur during break repair. Consequently, the targeted nucleotide sequence may be modified or inactivated. For example, a single nucleotide change (SNP) can give rise to an altered protein product, or a frameshift in a coding sequence can inactivate or “knock out” the sequence so that no protein product is produced. Petition 870240102758, dated 02 / 12 / 2024, page 124 / 196 102 / 172 In embodiments where the optional donor polynucleotide is present, the donor sequence in the donor polynucleotide can be exchanged with, or integrated into, the nucleotide sequence at the targeted site during double-strand break repair. For example, in embodiments where the donor sequence is flanked by upstream and downstream sequences having substantial sequence identity with upstream and downstream sequences, respectively, of the target site in the prokaryotic cell nucleotide sequence, the donor sequence can be exchanged with or integrated into the nucleotide sequence at the targeted site during homology-directed repair-mediated repair.Alternatively, in embodiments where the donor sequence is flanked by compatible overhangs (or the compatible overhangs are generated in situ by the Cppl or Csm1 polypeptide), the donor sequence can be directly linked to the cleaved nucleotide sequence by a non-homologous repair process during double-strand break repair. The exchange or integration of the donor sequence into the nucleotide sequence modifies the target nucleotide sequence or introduces an exogenous sequence into the targeted nucleotide sequence of prokaryotic cellular DNA.
[00131] In some embodiments, the double-strand breaks caused by the action of the nuclease or nucleases Cpf1 or Csm1 are repaired so that the DNA is removed from the DNA Petition 870240102758, dated 02 / 12 / 2024, pp. 125 / 196 103 / 172 prokaryotic cell. In some embodiments, one base, a few bases (i.e., 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases), or a large section of DNA (i.e., more than 10, more than 50, more than 100, or more than 500 bases) is / are deleted from the prokaryotic cellular DNA.
[00132] In some embodiments, the expression of prokaryotic genes can be modulated as a result of double-strand breaks caused by the Cpf1 or Csm1 nuclease or nucleases. In some embodiments, prokaryotic gene expression can be modulated by Cpf1 or Csm1 variant nucleases comprising a mutation that renders the Cpf1 or Csm1 nuclease incapable of producing a double-strand break. In some preferred embodiments, the Cpf1 or Csm1 variant nuclease comprising a mutation that renders the Cpf1 or Csm1 nuclease incapable of producing a double-strand break can be fused to a transcription activation or transcription repression domain.
[00133] The present invention can be used for the transformation of any prokaryotic species, including, but not limited to, cyanobacteria, Corynebacterium sp., Bifidobacterium sp., Mycobacterium sp., Streptomyces sp., Thermobifida sp., Chlamydia sp., Prochlorococcus sp., Synechococcus sp., Thermosynechococcus sp., Thermus sp., Bacillus sp., Clostridium sp., Geobacillus sp., Lactobacillus sp., Listeria sp., Staphylococcus sp., Petition 870240102758, dated 02 / 12 / 2024, p. 126 / 196 104 / 172 Streptococcus sp., Fusobacterium sp., Agrobacterium sp., Bradyrhizobium sp., Ehrlichia sp., Mesorhizobium sp., Nitrobacter sp., Rickettsia sp., Wolbachia sp., Zymomonas sp., Burkholderia sp., Neisseria sp., Ralstonia sp., Acinetobacter sp., Erwinia sp., Escherichia sp., Haemophilus sp., Legionella sp., Pasteurella sp., Pseudomonas sp., Psychrobacter sp., Salmonella sp., Shewanella sp., Shigella sp., Vibrio sp., Xanthomonas sp., Xylella sp., Yersinia sp., Campylobacter sp., Desulfovibrio sp., Helicobacter sp., Geobacter sp., Leptospira sp., Treponema sp., Mycoplasma sp., e Thermotoga sp.
[00134] Methods for introducing nuclease proteins, DNA or RNA molecules encoding nuclease proteins, guide DNA or RNA molecules encoding guide RNA, and DNA molecules of optional donor sequences into prokaryotic cells or organelles are known in the art, for example, in U.S. Patent Application 2016 / 0208243, incorporated herein by reference. Exemplary genomic modifications for prokaryotic cells that may have particular value for industrial applications are also known in the art, for example, in U.S. Patent Application 2016 / 0208243, incorporated herein by reference.
[00135] All publications and patent applications mentioned in the description are indicative of the level of technical skill of those qualified to apply this invention. Petition 870240102758, dated 02 / 12 / 2024, page 127 / 196 105 / 172 belongs. All publications and patent applications are incorporated herein by reference to the same extent as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference.
[00136] Although the prior invention has been described in some detail by way of illustration and example for the sake of clarity, it will be obvious that certain alterations and modifications may be made within the scope of the appended claims. Embodiments of the invention include: 1. A method for modifying a nucleotide sequence at a target site in the genome of a eukaryotic cell comprising: introduce into said eukaryotic cell (i) a DNA-directed RNA or a DNA polynucleotide encoding a DNA-directed RNA, wherein the DNA-directed RNA comprises: (a) a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA; and (b) a second segment interacting with a Cpf1 or Csm1 polypeptide; and (ii) a Cpf1 or Csm1 polypeptide, or a polynucleotide encoding a Cpf1 or Csm1 polypeptide, wherein the Cpf1 or Csm1 polypeptide comprises: (a) an RNA-binding portion that interacts with DNA-directed RNA; and Petition 870240102758, dated 02 / 12 / 2024, p. 128 / 196 106 / 172 (b) a portion of activity that exhibits site-directed enzymatic activity.
[00137] 2. A method for modifying a nucleotide sequence at a target site in the genome of a prokaryotic cell comprising: introduce into said prokaryotic cell (i) a DNA-directed RNA or a DNA polynucleotide encoding a DNA-directed RNA, wherein the DNA-directed RNA comprises: (a) a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA; and (b) a second segment interacting with a Cpf1 or Csm1 polypeptide; and (ii) a Cpf1 or Csm1 polypeptide, or a polynucleotide encoding a Cpf1 or Csm1 polypeptide, wherein the Cpf1 or Csm1 polypeptide comprises: (a) an RNA-binding portion that interacts with DNA-directed RNA; and (b) an activity portion that exhibits site-directed enzymatic activity, wherein said prokaryotic cell is not the native host of a gene encoding said Cpf1 or Csm1 polypeptide.
[00138] 3. A method for modifying a nucleotide sequence at a target site in the genome of a plant cell comprising: introduce into the aforementioned plant cell Petition 870240102758, dated 02 / 12 / 2024, p. 129 / 196 107 / 172 (i) a DNA-directed RNA or a DNA polynucleotide encoding a DNA-directed RNA, wherein the DNA-directed RNA comprises: (a) a first segment comprising a nucleotide sequence that is complementary to a sequence in the target DNA; and (b) a second segment interacting with a Cpf1 or Csm1 polypeptide; and (ii) a Cpf1 or Csm1 polypeptide, or a polynucleotide encoding a Cpf1 or Csm1 polypeptide, wherein the Cpf1 or Csm1 polypeptide comprises: (a) an RNA-binding portion that interacts with DNA-directed RNA; and (b) an activity portion that exhibits site-directed enzymatic activity.
[00139] 4. The method of any of the modalities 1-3, also including: cultivate the plant under conditions in which the Cpf1 or Csm1 polypeptide is expressed and cleaves the nucleotide sequence at the target site to produce a modified nucleotide sequence; and select a plant comprising said modified nucleotide sequence.
[00140] 5. The method of any of the embodiments 1-4, in which the cleavage of the nucleotide sequence at the target site comprises a double-strand break at or near the sequence to which the DNA-targeted RNA sequence is directed. Petition 870240102758, dated 02 / 12 / 2024, p. 130 / 196 108 / 172
[00141] 6. The method of embodiment 5, in which the aforementioned double-strand break is a staggered double-strand break.
[00142] 7. The method of embodiment 6, in which the aforementioned staggered double-strand break creates a 5' overhang of 3-6 nucleotides.
[00143] 8. The method of any of the embodiments 1-7, in which the said RNA directed to DNA is a guide RNA (gRNA).
[00144] 9. The method of any of the embodiments 1-8, in which the said modified nucleotide sequence comprises the insertion of heterologous DNA into the cell genome, the deletion of a nucleotide sequence from the cell genome or the mutation of at least one nucleotide in the cell genome.
[00145] 10. The method of any of the embodiments 1-9, wherein the said polypeptide Cppl or Csm1 is selected from the group consisting of: SEQ ID NOs: 3, 6, 9, 12, 15, 18, 20, 23, 106-173, and 230-236.
[00146] 11. The method of any of the embodiments 1-10, wherein the said polynucleotide encoding a Cpf1 or Csm1 polypeptide is selected from the group of SEQ ID NOs: 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19, 21, 22, 24, 25, and 174-206.
[00147] 12. The method of any of the modalities Petition 870240102758, dated 02 / 12 / 2024, p. 131 / 196 109 / 172 1-11, wherein the said polypeptide Cpf1 or Csml has at least 80% identity with one or more selected polypeptide sequence(s) from the group of SEQ ID Nos: 3, 6, 9, 12, 15, 18, 20, 23, 106-173 and 230-236.
[00148] 13. The method of any of the embodiments 1-12, in which the said polynucleotide encoding a Cpf1 or Csm1 polypeptide has at least 70% identity with one or more nucleic acid sequences selected from the group of SEQ ID NOs: 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19, 21, 22, 24, 25, and 174-206.
[00149] 14. The method of any of the embodiments 1-13, in which the polypeptide Cppl or Csm1 forms a homodimer or heterodimer.
[00150] 15. The method of any of the embodiments 1-14, in which the said plant cell is of a monocotyledonous species.
[00151] 16. The method of any of the modalities 1-14, in which the said plant cell is of a dicotyledonous species.
[00152] 17. The method of any of the embodiments 1-16, in which the expression of the Cppl or Csm1 polypeptide is under the control of an inducible or constitutive promoter.
[00153] 18. The method of any of the modalities 1-17, in which the expression of the Cpf1 or Csm1 polypeptide is under the control of a cell type-specific promoter. Petition 870240102758, dated 02 / 12 / 2024, p. 132 / 196 110 / 172 or developmentally preferred.
[00154] 19. The method of any of the embodiments 1-18, wherein the PAM sequence comprises 5'-TTN, wherein N can be any nucleotide.
[00155] 20. The method of any of the embodiments 1-19, wherein said nucleotide sequence at a target site in the genome of a cell encodes an SBPase, FBPase, FBP aldolase, large subunit of AGPase, small subunit of AGPase, sucrose phosphate synthase, amido synthase, PEP carboxylase, pyruvate phosphate dikinase, transketolase, small subunit of rubisco, or rubisco protein activase, or encodes a transcription factor that regulates the expression of one or more genes encoding an SBPase, FBPase, FBP aldolase, large subunit of AGPase, small subunit of AGPase, sucrose phosphate synthase, amido synthase, PEP carboxylase, pyruvate phosphate dikinase, transketolase, small subunit of rubisco, or rubisco protein activase.
[00156] 21. The method of any of the embodiments 1-20, further comprising the method of contacting the target site with a donor polynucleotide, wherein the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide or a portion of a copy of the donor polynucleotide integrates into the target DNA.
[00157] 22. The method of any of the modalities Petition 870240102758, dated 02 / 12 / 2024, p. 133 / 196 111 / 172 1-21, in which the target DNA is modified so that nucleotides within the target DNA are eliminated.
[00158] 23. The method of any of the embodiments 1-22, in which the said polynucleotide encoding a Cpf1 or Csm1 polypeptide is codon-optimized for expression in a plant cell.
[00159] 24. The method of any of the embodiments 1-23, in which the expression of the said nucleotide sequence is increased or decreased.
[00160] 25. The method of any of the embodiments 1-24, in which the polynucleotide encoding a Cpf1 or Csm1 polypeptide is operationally linked to a promoter that is constitutive, cell-specific, inducible, or activated by alternative splicing of a suicide exon.
[00161] 26. The method of any of the embodiments 1-25, wherein said Cpf1 or Csm1 polypeptide comprises one or more mutations that reduce or eliminate the nuclease activity of said Cpf1 or Csm1 polypeptide.
[00162] 27. Method of embodiment 26, wherein said mutated Cppl or Csm1 polypeptide comprises a mutation at a position corresponding to positions 917 or 1006 of FnCpfl (SEQ ID NO: 3) or to positions 701 or 922 of SmCsm1 (SEQ ID NO: 160) when aligned for maximum identity, or wherein said mutated Cpf1 or Csm1 polypeptide comprises a mutation at positions 917 and 1006 of FnCpf1 (SEQ ID NO: 3) or Petition 870240102758, dated 02 / 12 / 2024, p. 134 / 196 112 / 172 positions 701 and 922 of SmCsml (SEQ ID NO: 160) when aligned for maximum identity.
[00163] 28. The method of embodiment 27, wherein the said mutations in positions corresponding to positions 917 or 1006 of FnCpf1 (SEQ ID NO: 3) are D917A and E1006A, respectively, or wherein the said mutations in positions corresponding to positions 701 or 922 of SmCsm1 (SEQ ID NO: 160) are D701A and E922A, respectively.
[00164] 29. The method of any of the embodiments 26-28, in which the said mutated Cppl or Csm1 polypeptide comprises the amino acid sequence shown in the SEQ ID group NOs: 26-41 and 63-70.
[00165] 30. The method of any of the embodiments 26-29, in which the mutated Cppl or Csm1 polypeptide is fused with a transcription activation domain.
[00166] 31. The method of modality 30, in which the mutated Cpcl or Csm1 polypeptide is directly fused with a transcription activation domain or fused with a transcription activation domain with a ligand.
[00167] 32. The method of any of the modalities 26-29, in which the mutated Cppl or Csm1 polypeptide is fused with a transcription repressor domain.
[00168] 33. The method of modality 32, in which the mutated Cppl or Csm1 polypeptide is fused with a transcription repressor domain with a ligand. Petition 870240102758, dated 02 / 12 / 2024, pp. 135 / 196 113 / 172
[00169] 34. The method of any of the embodiments 1-33, in which the said Cpf1 or Csm1 polypeptide further comprises a nuclear localization signal.
[00170] 35. The method of modality 34, in which the said nuclear locator signal comprises SEQ ID NO: 1 or is encoded by SEQ ID NO: 2.
[00171] 36. The method of any of the embodiments 1-33, in which the said polypeptide Cpf1 or Csm1 further comprises a chloroplast signal peptide.
[00172] 37. The method of any of the embodiments 1-33, in which the said Cppl or Csm1 polypeptide further comprises a mitochondrial signal peptide.
[00173] 38. The method of any of the embodiments 1-33, wherein said Cpf1 or Csm1 polypeptide further comprises a signal peptide that is directed to said Cpf1 or Csm1 polypeptide at multiple subcellular locations.
[00174] 39. Nucleic acid molecule comprising a polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide, wherein said polynucleotide sequence has been codon-optimized for expression in a plant cell.
[00175] 40. Nucleic acid molecule comprising a polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide, wherein said sequence of Petition 870240102758, dated 02 / 12 / 2024, p. 136 / 196 114 / 172 polynucleotides were optimized per codon for expression in a eukaryotic cell.
[00176] 41. A nucleic acid molecule comprising a polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide, wherein said polynucleotide sequence has been codon-optimized for expression in a prokaryotic cell, wherein said prokaryotic cell is not the natural host of said Cpf1 or Csm1 polypeptide.
[00177] 42. The nucleic acid molecule of any of the embodiments 39-41, wherein said polynucleotide sequence is selected from the group consisting of: SEQ ID NOS: 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19, 21, 22, 24, 25, and 174-206 or a fragment or variant thereof, or wherein said polynucleotide sequence encodes a Cpf1 or Csm1 polypeptide selected from the group consisting of SEQ ID NOS: 3, 6, 9, 12, 15, 18, 20, 23, 106-173 and 230-236, and wherein said polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide is operationally linked to a promoter that It is heterologous to the polynucleotide sequence that codes for a Cpf1 or Csm1 polypeptide.
[00178] 43. The nucleic acid molecule of any of the embodiments 39-41, in which the variant polynucleotide sequence has at least 70% sequence identity with a sequence of Petition 870240102758, dated 02 / 12 / 2024, p. 137 / 196 115 / 172 polynucleotides selected from the group consisting of: SEQ ID NOs: 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19, 21, 22, 24, 25 and 174-206, or wherein said polynucleotide sequence encodes a Cpf1 or Csm1 polypeptide that has at least 80% sequence identity with a polypeptide selected from the group consisting of SEQ ID NOs: 3, 6, 9, 12, 15, 18, 20, 23, 106-173 and 230-236, and wherein said polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide is operationally linked to a promoter that is heterologous to the polynucleotide sequence encoding a polypeptide Cpf1 or Csm1.
[00179] 44. The nucleic acid molecule of any of the embodiments 39-41, wherein said Cpf1 or Csm1 polypeptide comprises an amino acid sequence selected from the group consisting of: SEQ ID NOs: 3, 6, 9, 12, 15, 18, 20, 23, 106-173 and 230-236, or a fragment or variant thereof.
[00180] 45. The nucleic acid molecule of embodiment 44, in which the variant polypeptide sequence has at least 70% sequence identity with a polypeptide sequence selected from the group consisting of: SEQ ID NOs: 3, 6, 9, 12, 15, 18, 20, 23, 106-173 and 230-236.
[00181] 46. The nucleic acid molecule of any of the embodiments 39-45, in which the aforementioned sequence of Petition 870240102758, dated 02 / 12 / 2024, p. 138 / 196 116 / 172 polynucleotides that encode a Cpfl or Csml polypeptide are operationally linked to a promoter that is active in a plant cell.
[00182] 47. The nucleic acid molecule of any of the embodiments 39-45, in which the said polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide is operationally linked to a promoter that is active in a eukaryotic cell.
[00183] 48. The nucleic acid molecule of either embodiment 39-45, in which the said polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide is operationally linked to a promoter that is active in a prokaryotic cell.
[00184] 49. The nucleic acid molecule of any of the embodiments 39-45, in which the said polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide is operationally linked to a constitutive promoter, inducible promoter, cell-type-specific promoter, or developmentally preferred promoter.
[00185] 50. The nucleic acid molecule of any of the embodiments 39-45, wherein said nucleic acid molecule encodes a fusion protein comprising said Cppl or Csm1 polypeptide and an effector domain.
[00186] 51. The nucleic acid molecule of embodiment 50, in which the aforementioned effector domain is selected from the group Petition 870240102758, dated 02 / 12 / 2024, p. 139 / 196 117 / 172 which consists of: transcription activator, transcription repressor, nuclear localization signal and cell penetration signal.
[00187] 52. The nucleic acid molecule of embodiment 51, in which said polypeptide Cppl or Csm1 is mutated to reduce or eliminate nuclease activity.
[00188] 53. The nucleic acid molecule of embodiment 52, wherein said mutated Cppl or Csm1 polypeptide comprises a mutation at a position corresponding to positions 917 or 1006 of FnCpfl (SEQ ID NO: 3) or to positions 701 or 922 of SmCsm1 (SEQ ID NO: 160) when aligned for maximum identity, or wherein said mutated Cpf1 or Csm1 polypeptide comprises a mutation at positions corresponding to positions 917 and 1006 of FnCpf1 (SEQ ID NO: 3) or to positions 701 and 922 of SmCsm1 (SEQ ID NO: 160) when aligned for maximum identity.
[00189] 54. The nucleic acid molecule of any of the embodiments 50-53, in which said Cppl or Csm1 polypeptide is fused to said effector domain with a ligand.
[00190] 55. The nucleic acid molecule of any of the embodiments 39-54, in which said polypeptide Cppl or Csm1 forms a dimer.
[00191] 56. A fusion protein encoded by the nucleic acid molecule of any of the following forms Petition 870240102758, dated 02 / 12 / 2024, pp. 140 / 196 118 / 172 50-55.
[00192] 57. A Cpfl or Csm1 polypeptide encoded by the nucleic acid molecule of any of the embodiments 39-45.
[00193] 58. A Cpf1 or Csm1 polypeptide mutated to reduce or eliminate nuclease activity.
[00194] 59. The Cpf1 or Csm1 polypeptide of embodiment 58, wherein said mutated Cpf1 or Csm1 polypeptide comprises a mutation at a position corresponding to positions 917 or 1006 of FnCpf1 (SEQ ID NO: 3) or to positions 701 or 922 of SmCsm1 (SEQ ID NO: 160) when aligned for maximum identity or wherein said mutated Cpf1 or Csm1 polypeptide comprises mutations at positions corresponding to positions 917 and 1006 of FnCpf1 (SEQ ID NO: 3) or positions 701 and 922 of SmCsm1 (SEQ ID NO: 160) when aligned for maximum identity.
[00195] 60. A plant cell, eukaryotic cell or prokaryotic cell comprising the nucleic acid molecule of any of the embodiments 39-55.
[00196] 61. A plant cell, eukaryotic cell or prokaryotic cell comprising the fusion protein or polypeptide of any of the embodiments 56-59.
[00197] 62. A plant cell produced by the method of either of the embodiments 1 and 3-38.
[00198] 63. A plant comprising the molecule of Petition 870240102758, dated 02 / 12 / 2024, pp. 141 / 196 119 / 172 nucleic acid of any of the modalities 39-55.
[00199] 64. A plant comprising the fusion protein or polypeptide of any of the embodiments 5659.
[00200] 65. A plant produced by the method of either of the modalities 1 and 3-38.
[00201] 66. The seed of the plant of any of the modalities 63-65.
[00202] 67. The method of either embodiment 1 and 3-38, in which the said modified nucleotide sequence comprises the insertion of a polynucleotide encoding a protein that confers antibiotic or herbicide tolerance to transformed cells.
[00203] 68. The method of embodiment 67, in which the polynucleotide encoding a protein that confers antibiotic or herbicide tolerance comprises SEQ ID NO: 76, or encodes a protein comprising SEQ ID NO: 77.
[00204] 69. The method of any of the embodiments 3-38 in which the said target site in the genome of a plant cell comprises SEQ ID NO: 71, or shares at least 80% identity with a portion or fragment of SEQ ID NO: 71.
[00205] 70. The method of any of the modalities 1-38, in which the said DNA polynucleotide that codes Petition 870240102758, dated 02 / 12 / 2024, p. 142 / 196 120 / 172 an RNA directed to DNA comprises SEQ ID NO:73, SEQ ID NO:91, SEQ ID NO:92, SEQ ID NO:93, or SEQ ID NO:95.
[00206] 71. The nucleic acid molecule of either embodiment 39-55, in which the polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide further comprises a polynucleotide sequence encoding a nuclear localization signal.
[00207] 72. The nucleic acid molecule of embodiment 71, wherein the said nuclear localization signal comprises SEQ ID NO: 1 or is encoded by SEQ ID NO: 2.
[00208] 73. The nucleic acid molecule of either embodiment 39-55, in which the polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide further comprises a polynucleotide sequence encoding a chloroplast signal peptide.
[00209] 74. The nucleic acid molecule of either embodiment 39-55, in which the polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide further comprises a polynucleotide sequence encoding a mitochondrial signal peptide.
[00210] 75. The nucleic acid molecule of any of the embodiments 39-55 in which the polynucleotide sequence encoding a Cpf1 or Csm1 polypeptide further comprises a polynucleotide sequence encoding a signal peptide that directs the aforementioned Petition 870240102758, dated 02 / 12 / 2024, pp. 143 / 196 121 / 172 Cpfl or Csml polypeptides for multiple subcellular locations.
[00211] 76. Fusion protein of embodiment 56, wherein said fusion protein further comprises a nuclear localization signal, chloroplast signal peptide, mitochondrial signal peptide or signal peptide that is directed to said Cppl or Csm1 polypeptide to multiple subcellular locations.
[00212] 77. The Cpf1 or Csm1 polypeptide of any of the embodiments 57-59 in which said Cpf1 or Csm1 polypeptide further comprises a nuclear localization signal, chloroplast signal peptide, mitochondrial signal peptide or signal peptide that is directed to said Cpf1 or Csm1 polypeptide to multiple subcellular localizations.
[00213] The following examples are offered for illustrative purposes only and not as a limitation. EXPERIMENTAL Example 1 - Cloning of cpf1 Constructs
[00214] The constructs containing Cpf1 (construct numbers 131306-131311 and 131313) are summarized in Table 1. Briefly, the cpf1 genes were synthesized de novo by GenScript (Piscataway, NJ) and amplified by PCR to add an SV40 N-terminal nuclear localization tag (SEQ ID NO: 2) in the frame with the sequence of Petition 870240102758, dated 02 / 12 / 2024, p. 144 / 196 122 / 172 coding and cpfl of interest as well as restriction enzyme sites for cloning. Using the appropriate restriction enzyme sites, each individual cpf1 gene was cloned downstream of the 2x35s promoter (SEQ ID NO: 43).
[00215] Guide RNAs targeting a DNA region spanning the junction between the promoter and the 5' end of the GFP coding region were synthesized by Integrated DNA Technologies (Coralville, IA) as complete cassettes. Each cassette included a U3 rice promoter (SEQ ID NO: 42) operationally linked to the appropriate gRNA (SEQ ID NOs: 47-53) which was operationally linked to the U3 rice terminator (SEQ ID NO: 44). While each gRNA was targeted to the same region of the GFP gene and promoter, each gRNA was designed to ensure it included the appropriate platform to correctly interact with its respective Cpf1 enzyme.
[00216] The constructs were assembled and cloned into a modified pSB11 vector skeleton containing the hptII gene which can confer resistance to hygromycin b in plants (SEQ ID NO: 45). The hptII gene was located downstream of the maize ubiquitin and 5'UTR promoter (pZmUbi; SEQ ID NO: 46). Table 1: Cpf1 Vectors Promoter Number Gene Cpfl1 Terminator Promoter Sequence Terminator Construct Cpfl Cpfl gRNA gRNA gRNA Petition 870240102758, dated 02 / 12 / 2024, pp. 145 / 196 123 / 172 131306 2X 35S (SEQ ID NO: 43) Francisella tularensis subsp. novicida U112 (SEQ ID NO: 5) 35S poly A (SEQ ID NO: 54) Rice U3 (SEQ ID NO: 42) Francisella GFP GRNA (SEQ ID NO: 47) Rice U3 (SEQ ID NO: 44) 131307 2X 35S (SEQ ID NO: 43) Acidaminococc us sp. BV3L6 (SEQ ID NO: 8) 35S poly A (SEQ ID NO: 54) Rice U3 (SEQ ID NO: 42) Acidaminococcus GFP GRNA (SEQ ID NO: 48) Rice U3 (SEQ ID NO: 44) 131308 2X 35S (SEQ ID NO: 43) Lachnospirace ae bacterium MA2020 (SEQ ID NO: 11) 35S poly A (SEQ ID NO: 54) Rice U3 (SEQ ID NO: 42) Lachnos GFP GRNAMA2020 (SEQ ID NO: 49) Rice U3 (SEQ ID NO: 44) 131309 2X 35S (SEQ ID NO: 43) Candidatus Methanoplasma termitum (SEQ ID NO: 14) 35S poly A (SEQ ID NO: 54) Rice U3 (SEQ ID NO: 42) Candidatus GFP GRNA (SEQ ID NO: 50) Rice U3 (SEQ ID NO: 44) 131310 2X 35S (SEQ ID NO: 43) Moraxella bovoculi 237 (SEQ ID NO: 17) 35S poly A (SEQ ID NO: 54) Rice U3 (SEQ ID NO: 42) Moraxella GFP GRNA (SEQ ID NO: 51) U3 ofrice (SEQ ID NO: 44) 131311 2X 35S (SEQ ID NO: 43) Lachnospirace ae bacterium ND2006 (SEQ ID NO: 19) 35S poly A (SEQ ID NO: 54) U3 from rice (SEQ ID NO: 42) GFP GRNA from LanchnosND2006 (SEQ ID NO: 52) U3 from rice (SEQ ID NO: 44) 131313 2X 35S (SEQ ID NO: 43) Prevotella disiens (SEQ ID NO: 25) 35S poly A (SEQ ID NO: 54) Rice U3 (SEQ ID NO: 42) Prevo GFP GRNA (SEQ ID NO: 53) Rice U3 (SEQ ID NO: 44) Each cpfl gene was fused into a frame with the SV40 nuclear localization signal (SEQ ID NO: 2, which encodes the amino acid sequence of SEQ ID NO: 1) at its 5' end. Example 2 - Rice Transformation mediated by Agrobacterium
[00217] Rice calluses (Oryza sativa cv. Kitaake) were infected with Agrobacterium cells harboring a superbinary plasmid containing a gene encoding the Petition 870240102758, dated 02 / 12 / 2024, pp. 146 / 196 124 / 172 green fluorescent protein (GFP; SEQ ID NO: 55 encoding SEQ ID NO: 56) operationally linked to a constitutive promoter. Three infected calluses showing high levels of GFP-derived fluorescence based on visual inspection were selected and divided into several sections. These sections were allowed to propagate in selection medium. After the callus pieces were allowed to recover and grow, these calluses were reinfected with Agrobacterium cells harboring genes encoding Cpf1 enzymes and their respective guide RNAs (gRNAs). After infection with the cpf1-containing vectors, the calluses were propagated in selection medium containing hygromycin b. Callus pieces that presumably expressed functional Cpf1 proteins were visually selected by inspecting the callus pieces for regions that were no longer visibly fluorescent.This loss of fluorescence was likely the result of a successful Cpf1-mediated edit of the GFP coding sequence, resulting in a non-functional GFP gene. For example, rice callus transformed first with a GFP construct and then with the 131307 construct, containing a gene encoding the Cpf1 protein from Acidaminococcus sp. BV3L6 (SEQ ID NO: 8, which encodes SEQ ID NO: 6), resulted in parts of the callus showing an apparent loss of GFP-derived fluorescence. Those parts of rice callus that contained clusters of cells that did not... Petition 870240102758, dated 02 / 12 / 2024, pp. 147 / 196 Samples 125 / 172 exhibiting GFP-derived fluorescence were prioritized for further molecular characterization. Example 3 - T7EI Test
[00218] The T7 I endonuclease assay (T7EI) is used to identify samples with insertions and / or deletions at the desired site and to evaluate the effectiveness of targeting genome editing enzymes. The assay protocol is modified from Shan et al (2014) Nature Protocols 9: 23952410. The basis of the assay is that T7EI recognizes and cleaves non-perfectly matched DNA. Briefly, a PCR reaction is performed to amplify a region of DNA containing the DNA sequence targeted by the gRNA. As both edited and unedited DNAs are expected to be included in the sample, a mixture of PCR products is obtained. The PCR products are fused and then allowed to reanneal. When an unedited PCR product reanneals with an edited PCR product, a DNA mismatch results. These DNA mismatches are digested by T7EI and can be identified by gel-based assays.DNA is extracted from rice callus that appears to exhibit a loss of fluorescence derived from GFP. PCR is performed with this DNA as a template using primers configured to amplify a region of DNA spanning the junction between the promoter and the open reading frame of GFP. The PCR products are fused and reannealed, then digested with T7EI. Petition 870240102758, dated 02 / 12 / 2024, pages 148 / 196 126 / 172 (New England Biolabs, Ipswich, MA) according to the manufacturer's protocol. The resulting DNA is subjected to electrophoresis on a 2% agarose gel. In samples where Cpf1 produced an insertion or deletion at the desired site, the initial band is digested to produce two smaller bands. Example 4 - DNA sequencing of rice callus
[00219] DNA extracted from rice callus that has appeared, based on visual inspection for loss of fluorescence and / or based on the results of T7EI assays, to understand genomically edited DNA as a result of the accumulation of functional Cpf1 enzyme is selected for sequence-based analysis. DNA is extracted from appropriate rice callus fragments and primers are used to PCR amplify the GFP-coding sequence from this DNA. The resulting PCR products are cloned into plasmids which are subsequently transformed into E. coli cells. These plasmids are recovered and Sanger sequencing is used to analyze the DNA to identify insertions, deletions, and / or point mutations in the DNA encoding GFP. Example 5 - Use of deactivated Cpf1 proteins to modulate gene expression.
[00220] The RuvC-like domain of Cpf1 has been shown to mediate DNA cleavage (Zetsche et al (2015) Cell 163: 759-771), with specific residues identified in the Cpf1 enzyme of Francisella tularensis subsp. novicida U112 (i.e., Petition 870240102758, dated 02 / 12 / 2024, pp. 149 / 196 127 / 172 D917 and E1006) which completely inactivated DNA cleavage activity when mutated from the native amino acid to alanine. Amino acid-based alignments using Multiple Clustal Alignment (Thompson et al (1994) Nucleic Acid Research 22: 4673-4680) of the eight Cpf1 enzymes investigated here were performed to identify the corresponding amino acid residues in the other enzymes. Table 2 lists these amino acid residues. The amino acid sequences of deactivated Cpf1 proteins corresponding to point mutations in each of the amino acid residues listed in Table 2 are found in SEQ ID NOs: 26-41. The amino acid sequences of double mutant deactivated Cpf1 proteins comprising mutations in the two residues listed for each cpf1 protein in Table 2 are found in SEQ ID NOs: 63-70. Table 2: Amino acid residues mutated to generate deactivated Cpf1 enzymes Protein First aa Second aa FnCpfl (SEQ ID NO: 3) D917 E1006 AsCpfl (SEQ ID NO: 6) D908 E993 Lb2Cpf1 (SEQ ID NO: 9) D815 E906 CMtCpf1 (SEQ ID NO: 12) D859 E944 MbCpf1 (SEQ ID NO: 15) D986 E1080 LbCpf1 (SEQ ID NO: 18) D832 E925 PcCpf1 (SEQ ID NO: 20) D878 E963 Petition 870240102758, dated 02 / 12 / 2024, pp. 150 / 196 128 / 172 PdCpfl (SEQ ID NO: 23) D943 E1032
[00221] The appropriate primers are designed so that the Quikchange PCR (Agilent Technologies, Santa Clara, CA) can be run to produce genes encoding the deactivated Cpf1 sequences listed in SEQ ID NOs: 26-41 and to produce genes encoding the deactivated Cpf1 sequences listed in SEQ ID NOs: 6370. The PCR is performed to produce genes encoding a fusion protein containing a deactivated Cpf1 protein fused to a gene expression activation or repression domain such as the EDL or TAL activation domains or the SRDX repressor domain, with the SV40 nuclear localization signal (SEQ ID NO: 2, which encodes SEQ ID NO: 1) fused into the frame at the 5' end of the gene. Guide RNAs (gRNAs) are designed to allow the gRNA to interact with the deactivated Cpf1 protein and to guide the deactivated Cpf1 protein to a desired location in a plant genome.Cassettes containing the gRNA(s) of interest, operationally linked to operable promoter(s) in plant cells and containing the gene(s) encoding Cpf1 fusion protein(s) fused to the activation and / or repression domain, are cloned into a vector suitable for plant transformation. This vector is transformed into a plant cell, resulting in the production of the gRNA(s) and Cpf1 fusion protein(s) in the plant cell. The fusion protein containing... Petition 870240102758, dated 02 / 12 / 2024, pp. 151 / 196 129 / 172 the deactivated Cpfl protein and the activator or repressor domain affect the modulation of the expression of nearby genes in the plant genome. Example 6 - Editing predetermined genomic locations in maize (Zea mays)
[00222] One or more gRNAs are designed to anneal to a desired site in the maize genome and to allow interaction with one or more Cpf1 or Csm1 proteins. These gRNAs are cloned into a vector so that they are operationally linked to a promoter that is operable in a plant cell (the “gRNA cassette”). One or more genes encoding a Cpf1 or Csm1 protein are cloned into a vector so that they are operationally linked to a promoter that is operable in a plant cell (the “cpf1 cassette” or “csm1 cassette”). Each of the gRNA cassette and the cpf1 cassette or csm1 cassette is cloned into a vector that is suitable for plant transformation, and this vector is subsequently transformed into Agrobacterium cells. These cells are placed in contact with corn tissue that is suitable for transformation.Following this incubation with Agrobacterium cells, the corn cells are grown in a tissue culture medium suitable for the regeneration of intact plants. Corn plants are regenerated from the cells that were placed in contact with the Agrobacterium cells. Petition 870240102758, dated 02 / 12 / 2024, pages 152 / 196 130 / 172 contain the vector that held the cpfl or csml cassette and the gRNA cassette. After regeneration of the maize plants, plant tissue is harvested and DNA is extracted from the tissue. T7EI assays and / or sequencing assays are performed, as appropriate, to determine if a change in the DNA sequence has occurred at the desired genomic location.
[00223] Alternatively, particle bombardment is used to introduce the cpf1 or csml cassette and the gRNA cassette into maize cells. Vectors containing a cpf1 or csml cassette and a gRNA cassette are coated in gold or titanium beads which are then used to bombard maize tissue that is suitable for regeneration. After bombardment, the maize tissue is transferred to tissue culture medium for maize plant regeneration. After maize plant regeneration, the plant tissue is harvested and DNA is extracted from the tissue. T7EI assays and / or sequencing assays are performed, as appropriate, to determine if a change in the DNA sequence has occurred at the desired genomic location. Example 7 - Editing predetermined genomic locations in Setaria viridis
[00224] One or more gRNAs is / are designed to anneal to a desired site in the Setaria viridis genome and to allow interaction with one or more Cpf1 or Csm1 proteins. These gRNAs are cloned into a vector so that Petition 870240102758, dated 02 / 12 / 2024, pp. 153 / 196 131 / 172 they are operationally linked to a promoter that is operable in a plant cell (the gRNA cassette). One or more genes encoding a Cpfl or Csml protein are cloned into a vector so that they are operationally linked to a promoter that is operable in a plant cell (the cpfl cassette or csml cassette). Each of the gRNA cassette and the cpfl1 cassette or csml cassette is cloned into a vector that is suitable for plant transformation, and this vector is subsequently transformed into Agrobacterium cells. These cells are placed in contact with Setaria viridis tissue that is suitable for transformation. After this incubation with Agrobacterium cells, the Setaria viridis cells are grown in a tissue culture medium that is suitable for the regeneration of intact plants.Setaria viridis plants are regenerated from cells that have been placed in contact with Agrobacterium cells harboring the vector containing the cpfl cassette or the csml cassette and the gRNA cassette. After regeneration of the Setaria viridis plants, plant tissue is harvested and DNA is extracted from the tissue. T7EI assays and / or sequencing assays are performed, as appropriate, to determine if a change in the DNA sequence has occurred at the desired genomic location.
[00225] Alternatively, particle bombardment is used to introduce the cpf1 cassette or the Petition 870240102758, dated 02 / 12 / 2024, pages 154 / 196 132 / 172 csml cassette and gRNA cassette in S. viridis cells. Vectors containing a cpf1 cassette or csm1 cassette and a gRNA cassette are coated in gold or titanium beads which are then used to bombard S. viridis tissue suitable for regeneration. After bombardment, the S. viridis tissue is transferred to tissue culture medium for regeneration of intact plants. After plant regeneration, the plant tissue is harvested and DNA is extracted from this tissue. T7EI assays and / or sequencing assays are performed, as appropriate, to determine if a change in the DNA sequence has occurred at the desired genomic location. Example 8 - Eliminate DNA from a predetermined genomic location.
[00226] A first gRNA is configured to anneal to a first desired site in the genome of a plant of interest and to allow interaction with one or more Cpf1 or Csm1 proteins. A second gRNA is designed to anneal to a second desired site in the genome of a plant of interest and to allow interaction with one or more Cpf1 or Csm1 proteins. Each of these gRNAs is operationally ligated to a promoter that is operable in a plant cell and is subsequently cloned into a vector that is suitable for plant transformation. One or more genes encoding a Cpf1 or Csm1 protein is / are Petition 870240102758, dated 02 / 12 / 2024, pages 155 / 196 133 / 172 cloned into a vector so that they are operationally linked to a promoter that is operable in a plant cell (the “cpf1 cassette” or “csm1 cassette”). The cpf1 cassette or csm1 cassette and the gRNA cassettes are cloned into a single plant transformation vector which is subsequently transformed into Agrobacterium cells. These cells are placed in contact with plant tissue that is suitable for transformation. After this incubation with Agrobacterium cells, the plant cells are grown in a tissue culture medium that is suitable for the regeneration of intact plants. Alternatively, the vector containing the cpf1 cassette or csm1 cassette and the gRNA cassettes is coated onto gold or titanium beads suitable for bombarding plant cells. The cells are bombarded and then transferred to the tissue culture medium that is suitable for the regeneration of intact plants.The gRNA-Cpf1 or gRNA-Csm1 complexes effect double-strand breaks at the desired genomic locations and, in some cases, the DNA repair machinery causes the DNA to be repaired so that the native DNA sequence located between the two targeted genomic locations is deleted. Plants are regenerated from cells that are placed in contact with Agrobacterium cells containing the vector containing the cpf1 cassette or the csm1 cassette and the gRNA cassettes. Petition 870240102758, dated 02 / 12 / 2024, pages 156 / 196 134 / 172 bombarded with pellets coated with this vector. After plant regeneration, plant tissue is harvested and DNA is extracted from the tissue. T7EI assays and / or sequencing assays are performed, as appropriate, to determine if DNA has been removed from the desired genomic location(s). Example 9 - Insertion of DNA into a predetermined genomic location.
[00227] A gRNA is designed to anneal to a desired site in the genome of a plant of interest and to allow interaction with one or more Cpf1 or Csm1 proteins. The gRNA is operationally linked to a promoter that is operable in a plant cell and is subsequently cloned into a vector that is suitable for plant transformation. One or more gene(s) encoding a Cpf1 or Csm1 protein is / are cloned into a vector so that it is / are operationally linked to a promoter that is operable in a plant cell (the “cpf1 cassette” or “csm1 cassette”). The cpf1 cassette or csm1 cassette and the gRNA cassette are both cloned into a single plant transformation vector that is subsequently transformed into Agrobacterium cells. These cells are placed in contact with plant tissue that is suitable for transformation. Simultaneously, the donor DNA is introduced into these same plant cells.The donor's DNA includes a DNA molecule that must... Petition 870240102758, dated 02 / 12 / 2024, pp. 157 / 196 135 / 172 to be inserted into the desired site in the plant genome, flanked by upstream and downstream flanking regions. The upstream flanking region is homologous to the genomic DNA region upstream of the genomic site targeted by the gRNA, and the downstream flanking region is homologous to the genomic DNA region downstream of the genomic site targeted by the gRNA. The upstream and downstream flanking regions mediate the insertion of DNA into the desired site of the plant genome. After this incubation with Agrobacterium cells and introduction of donor DNA, the plant cells are cultured in a tissue culture medium suitable for the regeneration of intact plants. Plants are regenerated from cells that have been placed in contact with Agrobacterium cells harboring the vector containing the cpf1 cassette or csml cassette or gRNA cassettes. After plant regeneration, plant tissue is harvested and DNA is extracted from the tissue.T7EI assays and / or sequencing assays are performed, as appropriate, to determine whether the DNA was inserted at the desired genomic location(s). Example 10 - Biolytic insertion of DNA into the CAO1 genomic site of rice
[00228] For biolistic DNA insertion into a predetermined genomic location, vectors with cpf1 cassettes or csm1 cassettes were designed. These vectors contained Petition 870240102758, dated 02 / 12 / 2024, pages 158 / 196 136 / 172 a 2X35S promoter (SEQ ID NO: 43) upstream of the cpf1 or csm1 ORF and a 35S poliA terminator sequence (SEQ ID NO: 54) downstream of the cpf1 or csm1 ORF. Table 3 summarizes these cpf1 and csm1 vectors. Table 3: Summary of cpf1 and csm1 vectors used for biological experiments Número do Promotor Fonte de ORF Cpf1 ou Csm1 Terminador Vetor 131272 (SEQ NO:81) ID 2X35S NO:43) (SEQ ID Francisella tularensis NO:5) (SEQ ID 35S poliA NO:54) (SEQ ID 131273 (SEQ NO:82) ID 2X35S NO:43) (SEQ ID Acidaminococcus sp. (SEQ ID NO:8) 35S poliA NO:54) (SEQ ID 131274 (SEQ NO:83) ID 2X35S NO:43) (SEQ ID Lachnospiraceae bacterium (SEQ ID NO:11) MA2020 35S poliA NO:54) (SEQ ID 131275 (SEQ NO:84) ID 2X35S NO:43) (SEQ ID Candidatus Methanoplasma (SEQ ID NO:14) termitum 35S poliA NO:54) (SEQ ID 131276 (SEQ NO:85) ID 2X35S NO:43) (SEQ ID Moraxella bovoculi 237 NO:17) (SEQ ID 35S poliA NO:54) (SEQ ID 131277 (SEQ NO:86) ID 2X35S NO:43) (SEQ ID Lachnospiraceae bacterium (SEQ ID NO:19) ND2006 35S poliA NO:54) (SEQ ID 131278 (SEQ NO:87) ID 2X35S NO:43) (SEQ ID Porphyromonas creviorican ID NO:22) is (SEQ 35S poliA NO:54) (SEQ ID 131279 (SEQ NO:88) ID 2X35S NO:43) (SEQ ID Prevotella disiens (SEQ ID NO:25) 35S poliA NO:54) (SEQ ID 132058 2X35S NO:43) (SEQ ID Anaerovibrio sp.RM50 NO:176) (SEQ ID 35S polyA NO:54) (SEQ ID 132059 2X35S NO:43) (SEQ ID Lachnospiraceae bacterium (SEQ ID NO:174) MC2017 35S polyA NO:54) (SEQ ID. Petition 870240102758, dated 02 / 12 / 2024, pp. 159 / 196 137 / 172 132065 2X35S NO:43) (SEQ ID Moraxella caprae DSM 19149 (SEQ ID NO:175) 35S poliA NO:54) (SEQ ID 132066 2X35S NO:43) (SEQ ID Succinivibrio dextrinosolvens H5 (SEQ ID NO:177) 35S poliA NO:54) (SEQ ID 132067 2X35S NO:43) (SEQ ID Prevotella bryantii B14 (SEQ ID NO:179) 35S poliA NO:54) (SEQ ID 132068 2X35S NO:43) (SEQ ID Flavobacterium branchiophilum FL15 (SEQ ID NO:178) 35S poliA NO:54) (SEQ ID 132075 2X35S NO:43) (SEQ ID Lachnospiraceae bacterium NC2008 (SEQ ID NO:180) 35S poliA NO:54) (SEQ ID 132082 2X35S NO:43) (SEQ ID Pseudobutyrivibrio ruminis (SEQ ID NO:181) 35S poliA NO:54) (SEQ ID 132083 2X35S NO:43) (SEQ ID Helcococcus kunzii ATCC 51366 (SEQ ID NO:183) 35S poliA NO:54) (SEQ ID 132084 2X35S NO:43) (SEQ ID Smithella sp.SCADC (SEQ ID NO:185) 35S poliA NO:54) (SEQ ID 132095 2X35S NO:43) (SEQ ID Uncultured bacterium (gcode 4) ACD 3C00058 (SEQ ID NO:187) 35S poliA NO:54) (SEQ ID 132096 2X35S NO:43) (SEQ ID Proteocatella sphenisci (SEQ ID NO:191) 35S poliA NO:54) (SEQ ID 132098 2X35S NO:43) (SEQ ID Bactéria WS6 de divisão candidata GW2011_GWA2_37_6 US52_C0007 (SEQ ID NO:182) 35S poliA NO:54) (SEQ ID 132099 2X35S NO:43) (SEQ ID Butyrivibrio sp. NC3005 (SEQ ID NO:190) 35S poliA NO:54) (SEQ ID 132105 2X35S NO:43) (SEQ ID Flavobacterium sp.316 (SEQ ID NO:196) 35S polyA NO:54) (SEQ ID 132100 2X35S NO:43) (SEQ ID Butyrivibrio fibrisolvens (SEQ ID NO:192) 35S polyA NO:54) (SEQ ID 132094 2X35S NO:43) (SEQ ID Bacteroidetes oral taxon 274 strain F0058 (SEQ ID NO:188) 35S polyA NO:54) (SEQ ID 132093 2X35S NO:43) (SEQ ID Lachnospiraceae bacterium COE1 (SEQ ID NO:189) 35S polyA NO:54) (SEQ ID 132111 2X35S NO:43) (SEQ ID Parcubacteria bacterium GW2011 (SEQ ID NO:197) 35S polyA NO:54) (SEQ ID 132092 2X35S NO:43) (SEQ ID Sulfuricurvum sp. PC08-66 (SEQ ID NO:186) 35S polyA NO:54) (SEQ ID 132097 2X35S NO:43) (SEQ ID Candidatus Methanomethylophilus alvus Mx1201 (SEQ ID NO:184) 35S polyA NO:54) (SEQ ID 132106 2X35S NO:43) (SEQ ID Eubacterium sp. (SEQ ID NO:200) 35S polyA NO:54) (SEQ ID Petition 870240102758, dated 02 / 12 / 2024, p. 160 / 196 138 / 172 132107 2X35S NO:43) (SEQ ID Microgenomates (Roizmanbacteria) bacterium GW2011_GWA2_37_7 (SEQ ID NO:201) 35S poliA NO:54) (SEQ ID 132102 2X35S NO:43) (SEQ ID Microgenomates (Roizmanbacteria) bacterium GW2011_GWA2_37_7 (SEQ ID NO:193) 35S poliA NO:54) (SEQ ID 132104 2X35S NO:43) (SEQ ID Prevotella brevis ATCC 19188 (SEQ ID NO:199) 35S poliA NO:54) (SEQ ID 132109 2X35S NO:43) (SEQ ID Smithella sp. SCADC (SEQ ID NO:203) 35S poliA NO:54) (SEQ ID 132101 2X35S NO:43) (SEQ ID Oribacterium sp. NK2B42 (SEQ ID NO:194) 35S poliA NO:54) (SEQ ID 132103 2X35S NO:43) (SEQ ID Synergistes jonesii cepa 78-1 (SEQ ID NO:195) 35S poliA NO:54) (SEQ ID 132108 2X35S NO:43) (SEQ ID Smithella sp.SC_K08D17 (SEQ ID NO:202) 35S polyA NO:54) (SEQ ID 132110 2X35S NO:43) (SEQ ID Prevotella albensis (SEQ ID NO:198) 35S polyA NO:54) (SEQ ID 132143 2X35S NO:43) (SEQ ID Moraxella lacunata (SEQ ID NO:206) 35S polyA NO:54) (SEQ ID 132144 2X35S NO:43) (SEQ ID Eubacterium coprostanoligenes (SEQ ID NO:205) 35S polyA NO:54) (SEQ ID 132145 2X35S NO:43) (SEQ ID Succiniclasticum ruminis (SEQ ID NO:204) 35S polyA NO:54) (SEQ ID.
[00229] In addition to the cpfl and csml vectors described in Table 3 shows that vectors with gRNA cassettes were designed so that the gRNA would anneal to a region of the CAO1 gene site in the rice (Oryza sativa) genome (SEQ ID NO: 71) and also allow interaction with the appropriate Cpf1 or Csm1 protein. In these vectors, the gRNA was operationally ligated to the rice U6 promoter (SEQ ID NO: 72) and terminator (SEQ ID NO: 74). Table 4 summarizes these gRNA vectors. Table 4: Summary of gRNA vectors used for biolistic experiments at the CAO1 genomic locus of rice. Petition 870240102758, dated 02 / 12 / 2024, pp. 161 / 196 139 / 172 Vetor Promotor gRNA Terminador 131608 OsU6 (SEQ ID NO:72) SEQ ID NO:73 Terminador OsU6 (SEQ ID NO:74) 131609 OsU6 (SEQ ID NO:72) SEQ ID NO:91 Terminador OsU6 (SEQ ID NO:74) 131610 OsU6 (SEQ ID NO:72) SEQ ID NO:92 Terminador OsU6 (SEQ ID NO:74) 131611 OsU6 (SEQ ID NO:72) SEQ ID NO:93 Terminador OsU6 (SEQ ID NO:74) 131612 OsU6 (SEQ ID NO:72) SEQ ID NO:94 Terminador OsU6 (SEQ ID NO:74) 131613 OsU6 (SEQ ID NO:72) SEQ ID NO:95 Terminador OsU6 (SEQ ID NO:74) 131912 OsU6 (SEQ ID NO:72) SEQ ID NO:207 Terminador OsU6 (SEQ ID NO:74) 131913 OsU6 (SEQ ID NO:72) SEQ ID NO:208 Terminador OsU6 (SEQ ID NO:74) 131914 OsU6 (SEQ ID NO:72) SEQ ID NO:209 Terminador OsU6 (SEQ ID NO:74) 131980 OsU6 (SEQ ID NO:72) SEQ ID NO:210 Terminador OsU6 (SEQ ID NO:74) 131981 OsU6 (SEQ ID NO:72) SEQ ID NO:211 Terminador OsU6 (SEQ ID NO:74) 131982 OsU6 (SEQ ID NO:72) SEQ ID NO:212 Terminador OsU6 (SEQ ID NO:74) 131983 OsU6 (SEQ ID NO:72) SEQ ID NO:213 Terminador OsU6 (SEQ ID NO:74) 131984 OsU6 (SEQ ID NO:72) SEQ IDNO:214 Terminador OsU6 (SEQ ID NO:74) 131985 OsU6 (SEQ ID NO:72) SEQ ID NO:215 Terminador OsU6 (SEQ ID NO:74) 131986 OsU6 (SEQ ID NO:72) SEQ ID NO:216 Terminador OsU6 (SEQ ID NO:74) 132033 OsU6 (SEQ ID NO:72) SEQ ID NO:228 Terminador OsU6 (SEQ ID NO:74) 132051 OsU6 (SEQ ID NO:72) SEQ ID NO:217 Terminador OsU6 (SEQ ID NO:74) 132052 OsU6 (SEQ ID NO:72) SEQ ID NO:218 Terminador OsU6 (SEQ ID NO:74) 132053 OsU6 (SEQ ID NO:72) SEQ ID NO:219 Terminador OsU6 (SEQ ID NO:74) 132054 OsU6 (SEQ ID NO:72) SEQ ID NO:220 Terminador OsU6 (SEQ ID NO:74) 132164 OsU6 (SEQ ID NO:72) SEQ ID NO:229 Terminador OsU6 (SEQ ID NO:74)
[00230] To facilitate the insertion of a hygromycin gene cassette into the CAO1 genomic site of rice, repair donor cassettes with approximately 1,000 base pairs of homology upstream and downstream of the double-strand break site to be caused by the action of the Cpf1 or Csm1 enzyme coupled to the gRNA targeting the site were configured. Figure 1 provides a schematic view of the CAO1 genomic site and the homology branches that were used. Petition 870240102758, dated 02 / 12 / 2024, pages 162 / 196 140 / 172 to guide homologous recombination and insertion of the hygromycin gene cassette into the CAO1 genomic site. The hygromycin gene cassette that was inserted into the rice CAO1 genomic site included the maize ubiquitin promoter (SEQ ID NO: 46) triggering the expression of the hygromycin resistance gene (SEQ ID NO: 76, which encodes SEQ ID NO: 77), which was flanked at its 3' end by the cauliflower mosaic virus 35S polyA sequence (SEQ ID NO: 54). Table 5 summarizes the repair donor cassette vectors that were constructed for hygromycin insertion into the rice CAO1 genomic site. Table 5: Rice CAO1 repair donor cassettes for insertion of hygromycin resistance genes. Repair Donor Cassette Vector 131760 SEQ ID NO:75 131632 SEQ ID NO:89 131633 SEQ ID NO:90 131987 SEQ ID NO:221 131988 SEQ ID NO:222 131990 SEQ ID NO:223 131991 SEQ ID NO:224 131992 SEQ ID NO:225 131993 SEQ ID NO:226 131994 SEQ ID NO:227
[00231] For the introduction of the cpf1 cassette or csml cassette, plasmid containing gRNA, and donor cassette Petition 870240102758, dated 02 / 12 / 2024, pages 163 / 196 For repairing rice cells (141 / 172), particle bombardment was used. For bombardment, 2 mg of 0.6 µm gold particles were weighed and transferred to sterile 1.5 mL tubes. 500 mL of 100% ethanol were added, and the tubes were sonicated for 10–15 seconds. After centrifugation, the ethanol was removed. One milliliter of sterile, double-distilled water was then added to the tube containing the gold beads. The bead pellet was briefly vortexed and then reformed by centrifugation, after which the water was removed from the tube. Under a sterile laminar flow hood, the DNA was coated onto the beads. Table 6 shows the amounts of DNA added to the beads. The plasmid containing the Cpf1 cassette or Csm1 cassette, the plasmid containing gRNA, and the repair donor cassette were added to the beads, and sterile double-distilled water was added to bring the total volume to 50 mL.For this, 20 mL of spermidine (1 M) were added, followed by 50 mL of CaCl2 (2.5 M). The gold particles were allowed to settle by gravity for several minutes and were then sedimented by centrifugation. The supernatant liquid was removed, and 800 µL of 100% ethanol were added. After a brief sonication, the gold particles were allowed to settle by gravity for 3-5 minutes, then the tube was centrifuged to form a sediment. Petition 870240102758, dated 02 / 12 / 2024, pages 164 / 196 142 / 172 supernatant was removed and 30 pL of 100% ethanol was added to the tube. The DNA-coated gold particles were resuspended in this ethanol by vortexing, and 10 resuspended gold particles were added to each of the microcarriers (Bio-Rad, Hercules, CA). The macrocarriers were allowed to air dry for 5-10 minutes under a laminar flow hood to allow the ethanol to evaporate. Table 6: Quantities of DNA used for particle bombardment experiments (all quantities are per 2 mg of gold particles) Cpfl or Csml plasmid 1.5 pg Plasmid containing gRNA- 1.5 pg Plasmid donor cassette Repair 3-15 pg Sterile double distilled water Add to bring the total volume to 50 pL
[00232] Rice callus tissue was used for bombardment. The rice callus was maintained in callus induction medium (MIC; 3.99 g / L N6 salts and vitamins, 0.3 g / L casein hydrolysates, 30 g / L sucrose, 2.8 g / L L-proline, 2 mg / L 2,4-D, 8 g / L agar, adjusted to pH 5.8) for 4-7 days at 28°C in the dark before bombardment. Approximately 80-100 callus pieces, each 0.2-0.3 cm in size and totaling 1-1.5 g in weight, were arranged in the center of a Petri dish containing Petition 870240102758, dated 02 / 12 / 2024, pages 165 / 196 143 / 172 Solid osmotic medium (MIC supplemented with 0.4 M sorbitol and 0.4 M mannitol) for a 4-hour osmotic pretreatment before particle bombardment. For bombardment, macrocarriers containing the DNA-coated gold particles were mounted on a macrocarrier. The rupture disc (1,100 psi), stop screen, and macrocarrier were assembled according to the manufacturer's instructions. The plate containing the rice callus to be bombarded was placed 6 cm below the stop screen, and the callus pieces were bombarded after the vacuum chamber reached 25-28 inches of Hg. After bombardment, the callus was left in osmotic medium for 16-20 hours, then the callus pieces were transferred to selection medium (MIC supplemented with 50 mg / L hygromycin and 100 mg / L timentine). The plates were transferred to an incubator and kept at 28°C in the dark to initiate the recovery of the transformed cells.Every two weeks, the callus was subcultured in fresh selection medium. Hygromycin-resistant callus fragments began to appear after approximately five to six weeks in the selection medium. Individual hygromycin-resistant callus fragments were transferred to new selection plates to allow the cells to divide and grow to produce sufficient tissue to be sampled for molecular analysis. Table 7 summarizes the... Petition 870240102758, dated 02 / 12 / 2024, pages 166 / 196 144 / 172 combinations of DNA vectors that were used for these rice bombardment experiments. Table 7: Summary of rice particle bombardment experiments for insertion of the hygromycin resistance gene at the CAO1 site. Cpf1 or Csm1 Plasmid Experiment gRNA Plasmid Repair Donor Plasmid 1 131272 131608 131760 2 131272 131608 131632 3 131273 131610 131632 4 131276 131612 131632 5 131272 131609 131633 6 131273 131611 131633 7 131276 131613 131633 13 131279 131912 131632 14 131274 131914 131632 15 131275 131913 131632 31 131277 132033 131632 32 131272 131985 131987 33 131272 131986 131988 43 131279 132051 131633 44 131274 132052 131633 45 131275 132053 131633 46 131277 132054 131633 53 131272 131982 131992 54 131272 131983 131993 55 131272 131980 131990 56 131272 131981 131991 57 131272 131984 131994 58 132058 131609 131633 59 132059 131609 131633 66 132068 131609 131633 67 132066 131609 131633 Petition 870240102758, dated 02 / 12 / 2024, pages 167 / 196 145 / 172 68 132065 131609 131633 69 132075 131609 131633 70 132067 131609 131633 71 132082 131609 131633 75 132096 131609 131633 76 132098 131609 131633 78 132083 131608 131632 79 132066 131608 131632 80 132065 131608 131632 81 132084 131608 131632 85 132075 131608 131632 86 132095 131608 131632 87 132099 131608 131632 88 132105 131608 131632 89 132100 131608 131632 90 132094 131608 131632 91 132093 131608 131632 92 132111 131608 131632 93 132092 131608 131632 94 132097 131608 131632 95 132106 131608 131632 96 132107 131608 131632 97 132102 131608 131632 98 132104 131608 131632 99 132109 132164 131632 100 132101 132164 131632 101 132103 132164 131632 102 132108 132164 131632 104 132143 132164 131632 105 132145 132164 131632 106 132059 131608 131632 107 132067 131608 131632 108 132096 131608 131632 109 132058 131608 131632 118 132110 131608 131632 119 132144 131608 131632 Petition 870240102758, dated 02 / 12 / 2024, pp. 168 / 196 146 / 172
[00233] After individual hygromycin-resistant callus fragments from each experiment were transferred to new plates, they were grown to a size sufficient for sampling. A small amount of tissue was harvested from each individual hygromycin-resistant rice callus fragment, and DNA was extracted from these tissue samples for PCR and DNA sequencing analysis. For experiments using repair donor plasmids 131760 or 131632, PCR was performed on these DNA extracts using primers with the sequences of SEQ ID NOs: 78 and 79 configured to amplify a DNA region extending from the ZmUbi promoter to a region of the rice genome that falls off the downstream repair donor branch, as schematically represented in Figure 1.For those experiments that used the repair donor plasmid 131633, primers with SEQ ID sequences NOS: 102 and 103 were used to amplify a region of DNA extending from the 35S terminator of CaMV into a region of the rice genome that lies outside the upstream repair donor branch, as shown schematically in Figure 1. These PCR reactions do not produce an amplicon from wild-type rice DNA of the repair donor plasmid and are therefore indicative of an insertion event at the site.
[00234] CAO1 of rice. Table 8 summarizes the number of Petition 870240102758, dated 02 / 12 / 2024, pp. 169 / 196 147 / 172 hygromycin-resistant callus fragments were produced from each experiment described in Table 7, as well as the number of PCR-positive callus fragments in which a putative insertion event occurred. The number of callus fragments used for each bombardment experiment was estimated by weight based on a survey of ten plates, with 159 ± 11.1 callus fragments per plate. Table 8: Summary of rice callus bombardment experiments Experiment #Bombardment of Callus Pieces (approximate) #Hygromycin-resistant callus pieces #PCR Events - Positive for Insertion 01 4134 290 12 02 795 20 0 03 1749 39 1 04 954 24 0 05 2067 46 3 06 1908 57 5 07 3339 49 3 13 3339 90 0 14 3180 81 0 15 4134 138 0 31 2067 55 0 43 1908 68 0 44 2544 117 0 45 1908 143 0 46 1908 192 3 58 1908 192 1 59 1431 192 1 66 1431 192 0 67 477 192 0 68 477 192 0 70 1431 192 1 Petition 870240102758, dated 02 / 12 / 2024, pp. 170 / 196 148 / 172 71 1431 192 0 75 1431 192 1 76 1431 192 0 78 1908 192 0 79 954 192 0 80 954 160 0 81 1431 192 0 85 1113 96 0 86 1113 133 0 87 1113 155 0 88 1272 192 0 89 954 192 0 90 954 192 0 91 954 192 0 92 954 192 0 93 954 192 0 94 954 192 0 95 954 192 0 97 954 192 0 98 1431 192 0 99 1272 192 0 100 1272 192 0 101 1272 192 0 102 1272 192 0 104 1272 192 0 105 1113 192 0 106 1113 192 0 107 1113 192 0 108 1272 96 0 109 1272 96 0 118 954 96 0 119 954 96 0
[00235] For the PCR-positive callus fragments listed in Table 8, an additional PCR analysis was performed to amplify along the junctions between the homology branches and the rice genome. Primers with the sequence SEQ ID NOs: 96 and 97 were used to amplify the Petition 870240102758, dated 02 / 12 / 2024, pp. 171 / 196 149 / 172 upstream region for experiments using repair donor plasmids 131760 or 131632. The location of these primer binding sites is shown schematically in Figure 1.
[00236] Sanger sequencing of PCR amplicons produced using the primer pairs described above to amplify the region downstream of the insertion event showed that the expected sequence was present in the transformed rice callus, confirming the insertion of the hygromycin gene cassette at the expected genomic site mediated by double-strand breaks produced by the Cpf1 or Csm1 enzyme. Sanger sequencing of PCR amplicons produced using the primer pairs described above to amplify the region upstream of the insertion events also showed that the expected sequence was present in the transformed rice callus, further confirming the insertion of the hygromycin gene cassette at the expected genomic site mediated by double-strand breaks produced by the Cpf1 or Csm1 enzyme.It is important to highlight that a five-base pair deletion (GCCTT) from the rice genomic sequence would occur at the upstream insertion site after Cpf1-mediated DSB formation, and this deletion was confirmed from sequencing data, further verifying that the observed insertion events were mediated by Cpf1 action. Figure 2A shows an alignment summarizing the... Petition 870240102758, dated 02 / 12 / 2024, pages 172 / 196 150 / 172 sequencing data confirmed the insertion events at the targeted rice CAO1 site in Experiment 1 (see Table 7).
[00237] Sequencing of the PCR products used to confirm the presence of a targeted insertion at the CAO1 site of rice as directed in Experiments 5 and 7 (see Table 7) was performed. Primers with SEQ ID numbers 104 and 105 were used to amplify the downstream region of these insertion events. These PCR products were sequenced, and the expected sequences were observed for insertion events mediated by DSB production by FnCpf1 (Experiment 5) and MbCpf1 (Experiment 7). The hph cassette was inserted at the CAO1 site at the target site, without base changes in the downstream branch.
[00238] Experiment 70 (Table 7) resulted in an insertion of a portion of the 35S terminator present in plasmid 131633 at the intended insertion site in the CAO1 genomic site of rice instead of an insertion of the entire hph cassette. Sequence analysis showed that the 35S terminator contained an eleven-base-pair region that shared ten bases with the downstream branch (Figure 4A). It appears that this region in the 35S terminator mediated an unintended homologous recombination event with the downstream branch in rice callus chunk #70-15, while the upstream branch in plasmid 131633 mediated Petition 870240102758, dated 02 / 12 / 2024, pages 173 / 196 151 / 172 the intended recombination event between this plasmid and the upstream sequence of the site in the rice CAO1 gene targeted by the guide RNA and the Cpf1 enzyme, resulting in the insertion sequence shown in Figure 4B. The resulting insertion led to a deletion of 179 base pairs and an insertion of 133 base pairs at the rice CAO1 site. While the insertion event discovered in experiment 70 included only a portion of the 35S terminator instead of the complete hph cassette intended for insertion, the recovered event was at the intended site in the CAO1 site targeted by the Cpf1 enzyme from Prevotella bryantii (SEQ ID NO: 138, encoded by SEQ ID NO: 179), indicating that this Cpf1 enzyme was effective in producing the intended DSB at the CAO1 genomic site.
[00239] Experiment 75 (Table 7) resulted in the insertion of a portion of the 35S terminator present in plasmid 131633 into the intended insertion site in the CAO1 genomic location of rice instead of an insertion of the entire hph cassette. Sequence analysis showed that the 35S terminator contained a twelve-base-pair region that shared eight bases with the downstream branch (Figure 4C). It appears that this region in the 35S terminator mediated an unintended homologous recombination event with the downstream branch in rice callus chunk #75-46, while the upstream branch in plasmid 131633 mediated Petition 870240102758, dated 02 / 12 / 2024, pages 174 / 196 152 / 172 the intended recombination event between this plasmid and the upstream sequence of the site in the rice CAO1 gene, directed by guide RNA and the Cpf1 enzyme, resulting in the insertion sequence shown in Figure 4D. The resulting insertion led to a 47 base pair deletion and a 24 base pair insertion at the rice CAO1 site. While the insertion event discovered in experiment 75 included only a portion of the 35S terminator instead of the complete hph cassette intended for insertion, the recovered event was at the intended site in the CAO1 site directed by the Cpf1 enzyme from Proteocatella sphenisci (SEQ ID NO: 142, encoded by SEQ ID NO: 191), indicating that this Cpf1 enzyme was effective in producing the intended DSB at the CAO1 genomic site.
[00240] Experiment 46 (Table 7) resulted in an insertion at the intended insertion site in the CAO1 genomic locus of rice, mediated by the Cpfl enzyme from Lachnospiraceae bacterium ND2006 (SEQ ID NO: 18, encoded by SEQ ID NO: 19). PCR analysis of the intended insertion site region at CAO1 resulted in the amplification of a band that is diagnostic of an insertion in callus fragment #46-161. This genomic region was subjected to sequence analysis to confirm the presence of the intended DNA insertion at the CAO1 locus of rice. Figure 5 shows the results of this sequence analysis, with the insertion Petition 870240102758, dated 02 / 12 / 2024, pages 175 / 196 153 / 172 expected from vector 131633 present in rice DNA at the expected site. The mutated PAM site (TTTC>TAGC) present in vector 131633 was also detected in rice DNA from callus fragment #46-161, further supporting HDR-mediated insertion of vector 131633 into the CAO1 site of rice as mediated by site-specific DSB induction by the Cpf1 enzyme from Lachnospiraceae bacterium ND2006.
[00241] Experiment 58 (Table 7) resulted in an insertion at the intended insertion site in the CAO1 genomic locus of rice, mediated by the Cpf1 enzyme of Anaerovibrio sp. RM50 (SEQ ID NO: 143, encoded by SEQ ID NO: 176). PCR analysis of the intended insertion site region at the CAO1 locus resulted in the amplification of a band that is diagnostic of an insertion in callus fragment #58-169. This genomic region is subjected to sequence analysis to confirm the presence of the intended DNA insertion at the CAO1 locus of rice.
[00242] Example 11 - Cpf1-mediated genomic DNA modification at the CAO1 site in rice
[00243] Rice callus was bombarded as described above with gold beads that were coated with a cpf1 vector and a gRNA vector. The rice callus that was bombarded as described for experiment 01 (Table 7) was left in osmotic medium for 16-20 hours after bombardment, then the callus pieces were transferred to medium of Petition 870240102758, dated 02 / 12 / 2024, pages 176 / 196 154 / 172 selection (MIC supplemented with 50 mg / L hygromycin and 100 mg / L timentine). The plates were transferred to an incubator and maintained at 28°C in the dark to initiate the recovery of transformed cells. Every two weeks, the callus was subcultured in fresh selection medium. Hygromycin-resistant callus fragments began to appear after approximately five to six weeks in the selection medium. Individual hygromycin-resistant callus fragments were transferred to new selection plates to allow the cells to divide and grow to produce sufficient tissue to be sampled for molecular analysis.
[00244] DNA was extracted from sixteen hygromycin-resistant callus fragments produced in Experiment 01 (Table 7), and PCR was performed using primers with SEQ ID sequences 100 and 101 to test for the presence of the cpf1 cassette. This PCR reaction showed that the DNA extracted from callus fragments numbered 1, 2, 4, 6, 7, and 15 produced the expected 853-pair amplicon consistent with the insertion of the cpf1 cassette into the rice genome (Figure 2B). PCR was also performed with DNA extracted from these hygromycin-resistant rice callus fragments using primers with SEQ ID sequences 98 and 99 to amplify a region of the rice CAO1 genomic site that was targeted by gRNA in vector 131608. The PCR reaction produced an amplicon Petition 870240102758, dated 02 / 12 / 2024, pp. 177 / 196 155 / 172 of 595 base pairs when wild-type DNA was used as a template. After PCR reaction with SEQ ID NOs: 98 and 99 as primers, a T7 endonuclease assay was performed with the resulting PCR product to test for small insertions and / or deletions at this site. Callus piece number DNA 15 showed a band pattern consistent with a small insertion or deletion (Figure 2C). PCR products produced from the reaction using primers with SEQ ID NO: 98 and 99 were cloned into E. coli cells using the pGEM® system (Promega, Madison, WI) according to the manufacturer's instructions. DNA was extracted from eight E. coli colonies for sequencing. Five of the eight colonies showed the same seven base pair deletion at the predicted Cpf1-mediated double-strand break site at the CAO1 site (Figure 2D).Without being limited by theory, a likely explanation for this deletion is that the DNA repair machinery of rice cells produced the deletion after repairing the double-strand break caused by FnCpf1 at the CAO1 site.
[00245] Experiment 01 (Table 7) was repeated with additional pieces of rice callus to confirm the reproducibility of the results obtained initially. The repetition of Experiment 01 resulted in the identification of four additional pieces of callus that appeared to be positive for indel production based on the assay results. Petition 870240102758, dated 02 / 12 / 2024, pages 178 / 196 156 / 172 of T7EI. DNA was extracted from these callus fragments for sequence analysis. PCR was performed to amplify the rice genome region around the targeted site in the CAO1 gene, and Sanger sequencing was performed. The sequencing results confirmed the results of the T7EI assay. Figure 2D shows the resulting sequence data. These four callus fragments exhibited variable deletion sizes ranging from a three-base pair to a seventy-five-base pair deletion, all located at the expected targeted site for FnCpf1 (SEQ ID NO: 3, encoded by SEQ ID NO: 5).
[00246] Experiments 31 and 46 (Table 7) tested the ability of LbCpf1 (SEQ ID NO: 18, encoded by SEQ ID NO: 19) to perform DSBs at two different locations in the CAO1 site of rice. Experiment 31 used plasmid 132033 as a gRNA source, while experiment 46 used plasmid 132054 as a gRNA source. After bombarding rice callus with the plasmids used for these experiments, DNA was extracted from hygromycin-resistant rice callus fragments and subjected to T7EI assays. After PCR amplification of the CAO1 genomic site of rice, T7EI assays identified one callus fragment from experiment 31 and five callus fragments from experiment 46 that appeared to contain indel at the expected site. The PCR products Petition 870240102758, dated 02 / 12 / 2024, pp. 179 / 196 157 / 172 of these rice callus fragments were analyzed by Sanger sequencing to identify the sequence(s) present at the CAO1 site in these callus fragments. Figure 3 shows the results of the Sanger sequencing analyses, confirming the presence of indels at the expected locations in the rice CAO1 site. Figure 3A shows the results of experiment 31 and Figure 3B shows the results of experiment 46. As Figure 3A shows, callus fragment 31-21 showed a deletion of fifty-six base pairs along with an insertion of ten base pairs. The calluses from experiment 46 (data presented in Figure 3B) showed deletions with sizes ranging from three to fifteen base pairs. It should be noted that callus fragments 46-38 and 46-77 showed two different indels, indicating that multiple indel production events occurred in independent cells within these callus fragments.All indels from these experiments were located at the predicted site in the CAO1 location directed by the respective guide RNA, indicating faithful production of DSBs at this location by the LbCpf1 enzyme.
[00247] Experiment 80 (Table 7) tested the ability of the Cpf1 enzyme from Moraxella caprae (SEQ ID NO: 133, encoded by SEQ ID NO: 175) to perform DSBs at the CAO1 site of rice. After bombarding the rice callus with the plasmids used for this experiment, the DNA was Petition 870240102758, dated 02 / 12 / 2024, pp. 180 / 196 158 / 172 extracted from hygromycin-resistant rice callus fragments and subjected to T7EI assays. After PCR amplification of the CAO1 genomic locus of rice, T7EI assays identified a callus fragment from the experiment containing an indel at the expected site. A PCR product of this rice callus fragment was analyzed by Sanger sequencing to identify the sequence present at the CAO1 site in this callus fragment. Figure 3A shows the results of these sequencing assays, with an elimination of eight base pairs present in callus fragment #80-33 at the predicted site in the directed CAO1 site of the respective guide RNA, indicating faithful DSB production at this site by the Cpf1 enzyme from Moraxella caprae.
[00248] Experiment 91 (Table 7) tested the ability of the Lachnospiraceae bacterium COE1 enzyme Cpf1 (SEQ ID NO: 125, encoded by SEQ ID NO: 189) to perform DSBs at the CAO1 site in rice. After bombarding rice callus with the plasmids used for this experiment, DNA was extracted from hygromycin-resistant rice callus fragments and subjected to T7EI assays. After PCR amplification of the rice CAO1 genomic site, T7EI assays identified a callus fragment from the experiment containing an indel at the expected site. A PCR product from this rice callus fragment was analyzed by Sanger sequencing to identify the sequence. Petition 870240102758, dated 02 / 12 / 2024, pages 181 / 196 159 / 172 present at the CAO1 site in this callus fragment. Figure 3A shows the results of these sequencing assays, with a nine-base-pair elimination present in callus fragment #91-4 at the predicted site in the CAO1 location directed by the respective guide RNA, indicating faithful DSB production at this site by the Cpf1 enzyme from Lachnospiraceae bacterium COE1.
[00249] Experiment 119 (Table 7) tested the ability of the Cpf1 enzyme from Eubacterium coprostanoligenes (SEQ ID NO: 173, encoded by SEQ ID NO: 205) to perform DSBs at the CAO1 site of rice. After bombarding rice callus with the plasmids used for this experiment, DNA was extracted from hygromycin-resistant rice callus fragments and subjected to T7EI assays. After PCR amplification of the CAO1 genomic site of rice, T7EI assays identified two callus fragments from the experiment that contained an indel at the expected site. A PCR product of these rice callus fragments was analyzed by Sanger sequencing to identify the sequence present at the CAO1 site in these calluses.Figure 3A shows the results of these sequencing assays, with an identical deletion of eight base pairs present in both callus fragments #119-4 and #119-11 at the predicted site in the CAO1 location directed by the respective guide RNA, indicating faithful DSB production at this site by the Cpf1 enzyme from Eubacterium coprostanoligenes. Petition 870240102758, dated 02 / 12 / 2024, pp. 182 / 196 160 / 172
[00250] Example 12 - Rice Plant Regeneration with an In-Place Insertion CAO1
[00251] Rice callus transformed with an hph cassette targeting the CAO1 site by an FnCpf1-mediated DSB in Experiment 1 (see Tables 7 and 8) was grown in tissue culture medium to produce shoots. These shoots were subsequently transferred to rooting medium, and the rooted plants were transferred to soil for cultivation in a greenhouse. The rooted plants appeared to be phenotypically normal in the soil. DNA was extracted from the rooted plants for PCR analysis. PCR amplification of upstream and downstream branches confirmed that the hph cassette was present at the CAO1 genomic site of rice.
[00252] T0 generation rice plants generated in Experiment 1 with the insertion of an hph cassette at the CAO1 site were grown and self-pollinated to produce T1 generation seeds. This seed was planted and the resulting T1 generation plants were genotyped to identify homozygous, hemizygous, and null plants. The T1 plants segregated as expected, with approximately 25% of the T1 plants being hemizygous for the hph insertion, 25% being null segregants, and 50% being heterozygous. Homozygous plants were observed phenotypically, with the expected yellow leaf phenotype associated with knockout. Petition 870240102758, dated 02 / 12 / 2024, pages 183 / 196 161 / 172 of the CAO1 gene (Lee et al. (2005) Plant Mol Biol 57:805818).
[00253] Generation T0 plants were regenerated from callus GE0046 numbers 33, 40, 62, and 90, which showed positive results for indels through T7EI assays and sequence verification (for callus piece #90) (Figure 3B). Regenerated plants derived from callus pieces 46-33, 40, 62, and 90 were positive for the presence of an indel at the CAO1 site based on T7EI assays using DNA extracted from regenerated plant tissue. Plants were also regenerated from callus pieces GE0046 46-96 and 46-161, which had previously shown an insertion of the hygromycin marker at the CAO1 site. Plants derived from callus pieces 46-96 and 46-161 were all positive for the insertion as detected by a PCR screening.Sequence data obtained from DNA extracted from two plants regenerated from callus fragment #46-90 showed the same eight base pair deletion detected in the callus (Figure 3B), indicating that this deletion was stable throughout the regeneration process. Sequence data obtained from DNA extracted from plants derived from callus fragments #46-40 and #46-62 showed deletions of 8, 9, 10, and 11 base pairs (data not shown). Example 13 - Identifying a new putative class Petition 870240102758, dated 02 / 12 / 2024, pages 184 / 196 162 / 172 of Cpfl-type proteins
[00254] Examination of phylogenetic trees of putative Cpf1 proteins (Zetsche et al. (2015) Cell 163: 759-771 and data not shown), together with sequence analyses of Cpf1 proteins and Cpf1-like proteins identified through BLAST searches, revealed a small group of proteins that appeared to be related to Cpf1 proteins, but with significantly altered sequences relative to known Cpf1 proteins. As two of these proteins are found in Smithella sp. SCADC and in Microgenomates, this new putative class of proteins was named Csm1 (CRISPR-associated proteins of Smithella and Microgenomates). Like Cpf1 proteins, these Csm1 proteins comprise the RuvCl, RuvCII, and RuvCIII domains, but importantly, the amino acid sequences of these domains are often quite divergent compared to those found in the amino acid sequences of the Cpf1 protein, particularly for the RuvCIII domain.Additionally, the spacing between RuvCI-RuvCII and RuvCIIRuvCIII is significantly altered in Csm1 proteins compared to Cpf1 proteins.
[00255] Alignment of the Csm1 protein from Smithella sp. SCADC (SmCsm1; SEQ ID NO: 160) with known Cpf1 proteins using the standard parameters of the BLASTP algorithm blast.ncbi.nlm.nih.gov / Blast.cgi) showed very little Petition 870240102758, dated 02 / 12 / 2024, pages 185 / 196 163 / 172 apparent sequence identity between these proteins. It was particularly apparent that while the RuvCI domain in the SmCsm1 protein appeared to be present and well aligned with the corresponding sequences in Cpf1 proteins, the RuvCII and RuvCIII regions, well conserved in Cpf1 proteins (Shmakov et al. (2016) Mol Cell 60:385-397), initially did not appear to be present in the putative Csm1 protein. Further analyses using HHPred (toolkit.tuebingen.mpg.de / hhpred; Soding et al. (2006) Nucleic Acids Res 34:W374-W378) revealed putative RuvCII and RuvCIII domains in this SmCsm1 protein. Table 9 shows the putative RuvCII domains in several putative Cppl and Csml proteins, and a representative C2c1 protein, along with the amino acid residue numbers in each sequence listing corresponding to the listed RuvCII sequence. The putative active residue is underlined for each listed protein. Table 9: RuvCII sequences of Cpf1 and Csm1 proteins Protein Sequence of RuvCII Number of Amino Acid Residues AsCpf1 (SEQ ID NO:6) Q—AVVVLENLNFGF 987-999 LbCpf1 (SEQ ID NO:18) D—AVIALEDLNSGF 919-931 SsCsm1 (SEQ ID NO:147) K—AYISLEDLSRAY 1057-1069 SmCsm1 (SEQ ID NO:160) R—GIISIEDLKQTK 920-932 ObCsm (SEQ ID NO:230) FPETIVALENLAKGT 931-945 Sm2Csm1 (SEQ ID NO:159) R—GIISIEDLKQTK 920-932 MbCsm1 (SEQ ID NO:134) Q—GVIALENLDTVR 916-928 AacC2c1 (SEQ ID NO:237) PPCQLILLEELS-EY 840-853 Petition 870240102758, dated 02 / 12 / 2024, pp. 186 / 196 164 / 172
[00256] Table 10 shows the putative RuvCIII domains in several putative Cpf1 and Csm1 proteins along with a representative C2c1 protein, along with the amino acid residue numbers in each sequence listing corresponding to the listed RuvCIII sequence. The putative active residue is underlined for each listed protein. Table 10: RuvCIII sequences of Cpf1 and Csm1 proteins Protein Sequence of RuvCIII Amino acid numbers of AsCpf1 NO:6) (SEQ ID WPM---------- -DADANGAYHIALK 1258-1273 LbCpf1 NO:18) (SEQ ID LPK---------- -NADANGAYNIARK 1175-1190 SsCsm1 NO:147) (SEQ ID RENNIHYIH------- -NGDDNGAYHIALK 1202-1223 SmCsm1 NO:160) (SEQ ID FDTRNDLKGFEGLNDPDKVAAFNIAKR 1029-1055 ObCsm1 NO:230) (SEQ ID SLN---------- -SPDTVAAYNVARK 1048-1063 Sm2Csm1 NO:159) (SEQ ID SLD---------- -SNDKVAAFNIAKR 1061-1076 MbCsm1 NO:134) (SEQ ID NLH---------- -NSDDVAAFNIAKR 1035-1050 AacC2c1 NO:237) (SEQ ID 972-987
[00257] As Tables 9 and 10 show, the domains RuvCII and RuvCIII identified by HHPred for putative Csm1 proteins (SEQ ID NOs: 134, 147, 159, 160, and 230) are significantly divergent from those found in Cpf1 proteins (representative sequences SEQ ID NOs: 6 and 18 shown above). Of particular interest, the ANGAY motif Petition 870240102758, dated 02 / 12 / 2024, pp. 187 / 196 The 165 / 172 residue after the active residue in the RuvCIII domain is extremely well conserved among Cpf1 proteins (Shmakov et al. (2016) Mol Cell 60: 385-397 and data not shown), but is absent or altered in most of these Csm1 proteins. Analysis of the RuvCII and RuvCIII domains in Csm1, Cpf1, and C2c1 proteins (Shmakov et al. (2016) Mol Cell 60:385-397) suggests that Csm1 proteins appear to be intermediate between Cpf1 and C2c1 proteins, as the RuvCII sequences of Csm1 are similar to those found in Cpf1 proteins, while the RuvCIII sequences of Csm1 are similar to those found in C2c1 proteins. The RuvCIII domains of Csm1 proteins primarily contain a DXXAA motif that is conserved in the C2c1 protein sequence.
[00258] While Csm1 proteins share some sequence similarity with C2c1 proteins, their genomic context suggests that Csm1 proteins function in several ways, like Cpf1 proteins. Specifically, C2c1 proteins require a crRNA and a tracrRNA, with the tracrRNA being partially complementary to the crRNA sequence. The genomic locus comprising the ORF encoding Csm1 from SCADC of Smithella sp. (SEQ ID NO: 238) includes a CRISPR array with Cpf1-like direct repeats, preceded by a Csm1 ORF, Cas4 ORF, Cas1 ORF, and Cas2 ORF. This is consistent with the genomic organization found in Cpf1-coding genomes (Shmakov et al. (2017) Nat Petition 870240102758, dated 02 / 12 / 2024, pages 188 / 196 166 / 172 Rev Microbiol doi:10.1038 / nrmicro.2016.184). In contrast, the C2c1 genomic organization tends to contain a fused Cas1 / Cas4 ORF. Furthermore, genomic sites containing C2c1 tend to encode a crRNA array such as a tracrRNA with partial complementarity to the direct crRNA repeat. Examination of the Smithella sp. SCADC genomic site containing the Csm1-coding ORF and associated crRNA sequences did not reveal any tracrRNA-like sequences, strongly suggesting that Csm1 does not require a tracrRNA to produce double-strand breaks.
[00259] Recently, a new class of nucleases termed CasX proteins was described (Burstein et al. (2016) Nature http: / / dx.doi.org10.1038 / nature20159). The Deltaproteobacterial CasX protein (SEQ ID NO: 239) was described as a ~980 amino acid protein found in a genomic region that included the Cas1, Cas4, and Cas2 protein coding regions as well as a CRISPR repeat region and a tracrRNA. The report describing CasX conclusively showed that this tracrRNA was required for endonuclease function, in contrast to Csm1 proteins which do not require a tracrRNA. BLASTP alignments of SmCsm1 (SEQ ID NO: 160) and Deltaproteobacterial CasX (SEQ ID NO: 239) showed very weak alignment (data not shown). HHPred analysis of this CasX protein was used to identify the Petition 870240102758, dated 02 / 12 / 2024, pages 189 / 196 167 / 172 putative domains RuvCl, RuvCII and RuvCIII and their respective residues in the active site.
[00260] In addition to the altered amino acid sequences of the putative RuvCII and RuvCIII domains in Csm1 proteins relative to Cpf1 proteins, the protein organization is significantly altered such that the spacing between these domains is significantly different between Csm1 and Cpf1 proteins. Table 11 shows a comparison of the spacing between active residues in RuvC subdomains in known Cpf1 proteins (AsCpf1 and LbCpf1; SEQ ID NOs: 6 and 18) compared with the spacing in these putative Csm1 proteins (SEQ ID NOs: 134, 147, 159, 160 and 230), the CasX Deltaproteobacterial protein (SEQ ID NO: 239) and a representative C2c1 protein (SEQ ID NO: 237). The data in Table 11 clearly show that the Cpf1, CasX, C2c1, and Csm1 proteins have a characteristic RuvC domain spacing, with the RuvCI-RuvCII spacing of CasX similar to Cpf1 and the RuvCII-RuvCIII spacing similar to the Csm1 / C2c1 spacing.The spacing of the RuvCI, RuvCII, and RuvCIII domains in the C2c1 and Csm1 proteins is similar, but the divergent sequences of RuvCIII and the absence of a tracrRNA in the Csm1 systems support the classification of Csm1 nucleases as separate from C2c1 nucleases. Table 11: Comparison of RuvC subdomain spacing Petition 870240102758, dated 02 / 12 / 2024, pp. 190 / 196 168 / 172 Protein Spacing of RuvCI-RuvCII RuvCII-RuvCIII (#amino acids) (#amino acids) AsCpf1 (SEQ ID NO:6) 84 269 LbCpf1 (SEQ ID NO:18) 92 254 CasX (SEQ ID NO:239) 97 166 SmCsm1 (SEQ ID NO:160) 220 122 ObCsm1 (SEQ ID NO:230) 211 113 Sm2Csm1 (SEQ ID NO:159) 224 139 MbCsm1 (SEQ ID NO:134) 225 117 SsCsm1 (SEQ ID NO:147) 214 149 AacC2c1 (SEQ ID NO:237) 278 129
[00261] Along with the divergent amino acid sequences RuvCII and RuvCIII and the altered spacing of these domains in Csm1 proteins compared to Cpf1 proteins, it should be noted that in many cases, HHPred analyses did not find any Csm1 sequence corresponding to the amino acid residues corresponding to D1225 in FnCpf1 (SEQ ID NO: 3) (D1234 in AsCpf1 (SEQ ID NO: 6) and D1148 in LbCpf1 (SEQ ID NO: 18)). Analysis of the mutation of the D1225 residue of FnCpf1 showed that the mutation of this residue significantly reduces the catalytic activity of this nuclease (Zetsche et al. (2015) Cell 163: 759-771), suggesting that the enzymatic function of this residue is very important for Cpf1 enzymes.
[00262] In addition to the altered amino acid sequence of Petition 870240102758, dated 02 / 12 / 2024, pp. 191 / 196 169 / 172 putative RuvC domains in Csml proteins in relation to Cpf1 proteins, HHPred analyses with Csm1 proteins show no matches with Cpf1 proteins at their N-terminus, in contrast to HHPred analyses based on known Cpf1 proteins. An HHPred analysis with the FnCpf1 amino acid sequence (SEQ ID NO: 3) resulted in only two matches, for AsCpf1 (SEQ ID NO: 6) and LbCpf1 (SEQ ID NO: 18) with 100% probability and covering the entire FnCpf1 amino acid sequence. In contrast, an HHPred analysis with SmCsm1 (SEQ ID NO: 160) finds only matches with Cpf1 proteins covering the amino acid regions 391-1017 and 1003-1064 in SmCsm1. Amino acids 1003-1030 correspond to a variety of proteins, including a lysine biosynthesis protein, an amino acid transporter protein, a transcription initiation factor, 50S and 30S ribosomal proteins, and DNA-directed RNA polymerases.No matches for the first 390 amino acids of Csm1 were found in this HHPred analysis. Similar HHPred analyses with additional Csm1 proteins (SEQ ID NOs: 134, 147, 159, 160 and 230) also found no matches for the N-terminal portions of these Csm1 proteins, further supporting the conclusion that these proteins share some similarity with Cpf1 proteins, but are not themselves Cpf1 proteins. Petition 870240102758, dated 02 / 12 / 2024, pages 192 / 196 170 / 172 Example 14 - Functional characterization Csml
[00263] Given the divergent nature of Csm1 proteins compared to Cpf1 proteins, we sought to confirm that these proteins were capable of producing DSBs in vivo. Although the amino acid sequences of Csm1 proteins are quite divergent from those of Cpf1 proteins, genomic analyses of the organisms that are the source of these Csm1 proteins revealed CRISPR arrays (data not shown), suggesting that these proteins could in fact be functional.
[00264] Experiment 81 (Table 7) tested the ability of a Csm1 enzyme from SCADC of Smithella sp. (SEQ ID NO: 160, encoded by SEQ ID NO: 185) to perform DSBs at the CAO1 site of rice. After bombarding the rice callus with the plasmids used for this experiment, DNA was extracted from hygromycin-resistant rice callus fragments and subjected to T7EI assays. After PCR amplification of the rice CAO1 genomic site, T7EI assays identified three callus fragments from the experiment that contained an indel at the expected site. The PCR products of these rice callus fragments were analyzed by Sanger sequencing to identify the sequence present at the CAO1 site in this callus. Figure 3A shows the results of these sequencing assays, with an elimination of eight base pairs present in callus fragment #81-46, a Petition 870240102758, dated 02 / 12 / 2024, pages 193 / 196 171 / 172 identical deletion of eight base pairs present in callus fragment #81-30, and a deletion of twelve base pairs present in callus fragment #81-9 at the predicted site in the CAO1 location directed by the respective guide RNA, indicating faithful DSB production at this site by the Csm1 enzyme of SCADC Smithella sp.
[00265] Experiment 93 (Table 7) tested the ability of a Csm1 enzyme from Sulfuricurvum sp. (SEQ ID NO: 147, encoded by SEQ ID NO: 186) to perform DSBs at the CAO1 site of rice. After bombarding the rice callus with the plasmids used for this experiment, DNA was extracted from hygromycin-resistant rice callus fragments and subjected to T7EI assays. After PCR amplification of the rice CAO1 genomic site, T7EI assays identified a callus fragment from the experiment containing an indel at the expected site. A PCR product of rice callus fragment #93-47 was analyzed by Sanger sequencing to identify the sequence present at the CAO1 site in this callus fragment.Figure 3A shows the results of these sequencing assays, with a deletion of forty-two base pairs present in the callus fragment #93-47 at the predicted site in the directed CAO1 site of the respective guide RNA, indicating faithful DSB production at this site by the Csm1 enzyme from Sulfuricurvum sp. Petition 870240102758, dated 02 / 12 / 2024, pp. 194 / 196 172 / 172
[00266] Experiment 97 (Table 7) tested the ability of a Csm1 enzyme from Microgenomates (Roizmanbactería) bacterium (SEQ ID NO: 134, encoded by SEQ ID NO: 193) to perform DSBs at the CAO1 site of rice. After bombarding rice callus with the plasmids used for this experiment, DNA was extracted from hygromycin-resistant rice callus fragments and subjected to T7EI assays. After PCR amplification of the CAO1 genomic locus of rice, T7EI assays identified three callus fragments from the experiment that contained an indel at the expected site. Callus fragments #97-112, 97-130, and 97-141 showed a banding pattern in the T7EI analysis consistent with faithful DSB production at this site by the Csm1 enzyme from Microgenomates (Roizmanbactería) bacterium. The DNA extracted from callus fragments #97-112 and #97-141 was subjected to sequence analysis (Fig. 3A).This sequence analysis showed a deletion of eight identical base pairs present in both calluses, indicating the faithful production of DSB at this site by the Csm1 enzyme from Microgenomates (Roizmanbacteria) bacterium. Petition 870240102758, dated 02 / 12 / 2024, pp. 195 / 196
Claims
1 / 4 CLAIMS 1. A method for modifying a nucleotide sequence at a target site in the genome of a plant cell, characterized in that it comprises: introducing into said plant cell (i) a guide RNA (gRNA) or a DNA polynucleotide encoding a gRNA, wherein the gRNA comprises: (a) a first segment comprising a nucleotide sequence that is complementary to the nucleotide sequence at the target site; and (b) a second segment that interacts with a Csm1 polypeptide; and (ii) a Csm1 polypeptide, or a polynucleotide encoding a Csm1 polypeptide, wherein the polynucleotide is defined by the sequence of SEQ ID NO: 186 and encodes the Csm1 polypeptide which has endonuclease activity, and the Csm1 polypeptide is defined by the amino acid sequence of SEQ ID NO: 147 and has endonuclease activity.
2. Method according to claim 1, characterized in that it further comprises: cultivating the plant cell under conditions in which the Csm1 polypeptide is expressed and cleaves the nucleotide sequence at the target site to produce a modified nucleotide sequence; and selecting a plant comprising said modified nucleotide sequence. Petition 870260060588, dated 06 / 22 / 2026, p. 11 / 19 2 / 4 3. A method according to claim 2, characterized in that said modified nucleotide sequence comprises the insertion of heterologous DNA into the cell genome, the deletion of a nucleotide sequence from the cell genome, or the mutation of at least one nucleotide in the cell genome.
4. Method according to claim 2, characterized in that said modified nucleotide sequence comprises the insertion of a polynucleotide encoding a protein that confers antibiotic or herbicide tolerance to the transformed cells.
5. Method, according to claim 4, characterized in that said polynucleotide encoding a protein that confers antibiotic or herbicide tolerance defined by the nucleotide sequence of SEQ ID NO: 76, or encodes a protein defined by the amino acid sequence of SEQ ID NO:
77.
6. A method, according to any one of claims 1 to 5, characterized in that said plant cell genome is a nuclear, plastid or mitochondrial genome.
7. Method, according to any one of claims 1 to 6, characterized in that said gRNA is an RNA directed to DNA. Petition 870260060588, dated 06 / 22 / 2026, page 12 / 19 3 / 4 8. Nucleic acid molecule characterized in that it comprises a polynucleotide sequence, wherein said polynucleotide sequence is defined by SEQ ID NO: 186, and encodes a Csm1 polypeptide that has endonuclease activity, wherein said polynucleotide sequence is operationally linked to a promoter that is heterologous to said polynucleotide sequence.
9. Nucleic acid molecule, according to claim 8, characterized in that said polynucleotide sequence encodes a Csm1 polypeptide comprising one or more mutations at one or more positions corresponding to positions 701 or 922 of SmCsm1 (SEQ ID NO: 160) when aligned for maximum identity, wherein said mutations at positions 701 or 922 are D701A and E922A, respectively.
10. Fusion protein characterized in that it comprises: (i) a Csm1 polypeptide defined by an amino acid sequence SEQ ID NO: 147 and has endonuclease activity; and (ii) a heterologous polypeptide.
11. Fusion protein, according to claim 10, characterized in that the heterologous polypeptide comprises an effector domain selected from the group consisting of: a cleavage domain, an epigenetic modification domain, a transcriptional activation domain, and a transcriptional repressor domain.
12. Fusion protein according to any one of claims 10 or 11, characterized in that the fusion protein further comprises a nuclear localization signal, a plastid signal peptide, a mitochondrial signal peptide, a signal peptide capable of trafficking proteins to multiple subcellular locations, a cell penetration domain, or a marker domain. Petition 870260060588, dated 06 / 22 / 2026, p. 14 / 19