RNA transfection into plant cells using modified RNA
Modified RNA molecules with pseudouridine and N1-methyl-pseudouridine in plant cells enhance protein expression and genome editing by optimizing 5'-UTR and 3'-UTR structures, addressing the inefficiencies of current transient expression methods.
Patent Information
- Application Number
- JP2025522138
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-20
- Publication Date
- 2025-10-21
AI Technical Summary
Current methods for transient gene expression in plant cells, such as using mRNA, face challenges with rapid degradation and inefficient protein production, particularly when using pseudouridine-modified mRNAs, leading to low levels of genome editing and protein production that decline over time.
The use of modified RNA molecules with pseudouridine and N1-methyl-pseudouridine modifications in the 5'-UTR, 3'-UTR, and poly(A) tail, optimized for plant cells, enhances protein expression by incorporating a 5'-UTR from a positive-strand RNA virus or plant RNA transcript, and introducing these molecules via methods like PEG transformation or biolistics.
This approach results in significantly increased and sustained protein expression, with levels up to 10-fold higher than unmodified mRNAs, facilitating effective transient genetic modification and targeted genome editing in plant cells.
Smart Images

Figure 2025534904000004 
Figure 2025534904000005 
Figure 2025534904000006
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of molecular plant biology. More particularly, the present invention relates to transient gene expression in plant cells. The present invention relates to the introduction of modified RNA molecules into plant cells to achieve long-term transient expression, for example, for the purpose of efficient targeted genetic modification of plant cells. [Background technology]
[0002] Crop plant improvement is fundamental to modern agriculture. The classical way to introduce new traits is by the introgression of foreign genes into fertile plant species. Such genes can, for example, result in crops with increased resistance to (a-)biotic stresses. Beyond the requirement for fertility, introgression is a time-consuming and labor-intensive process, requiring many breeding and selection steps. Another way to introduce (trans)genes is through genetic modification, which opens up the possibility of gene transfer even across kingdom boundaries. For example, bacterial genes, such as cellular programmed endonucleases of the CRISPR / CAS system, can be expressed in plant cells.
[0003] The typical method for expressing proteins in plant cells involves stably integrating a transgene into the plant genome using Agrobacterium or another transformation method. When a transgene is driven by a constitutive plant promoter, the protein is expressed throughout the plant's life cycle, whereas inducible or cell-specific promoters can be used to express the transgene at specific times and / or in specific cell types. For example, to be effective, the expression of a gene conferring herbicide resistance must be stable and produced throughout the plant's life cycle. However, the expression of most (trans)genes only needs to be present for a more or less narrow window of time. This is true for genes involved in processes such as development and regeneration (e.g., WOX5 and PLT1), which are tightly regulated and, when expressed for a short period of time, can be used only to promote regeneration. The same is true for transgenes involved in site-directed mutagenesis, such as proteins in the CRISPR / Cas complex, which only need to be present to induce modification and may be unnecessary thereafter. Therefore, there is a need for transient expression methods in plant cells.
[0004] Controlled gene expression can be more easily achieved using mRNA. However, while this is a more common technique in mammalian cells, there is little literature on the introduction of mRNA into plant cells (usually protoplasts) for protein production. After messenger RNA is produced in vitro, it can be introduced into cells and tissues using several different methods. Once introduced into the cell, the mRNA is immediately translated without the need for a promoter sequence. This is particularly beneficial for plant cells, as promoter sequences need to be optimized for different plant species.
[0005] Zhang et al. (Nat Commun;7:12617, 2016) reported the use of mRNA for transient (transgene-free) Cas9 delivery into maize using biolistic bombardment. In this case, the delivered mRNA contained the 5'- and 3'-untranslated repeats of the ZmUbi1 gene. However, because the mRNA is non-replicative and is degraded by cellular RNases over time, relatively low levels of genome editing were observed. Thus, a typical transient expression profile was obtained upon mRNA transfection, with high levels of protein production achieved immediately after the mRNA was introduced into the cells, followed by a rapid decline over time as the mRNA was degraded and became less dominant.
[0006] In mammalian cells, modified mRNAs containing pseudouridine (Ψ) have been studied for transient protein expression. Kariko et al. (Mol Ther. 16(11):1833-1840, 2008) reported the use of pseudouridine-containing mRNAs for translation enhancement in mammalian cells. However, Ψ-mediated translation enhancement was not predictable; in wheat extracts, protein production from Ψ-containing mRNAs was reduced by approximately 50%, and in bacterial cell lysates, mRNAs with the Ψ modification were not translated at all.
[0007] To meet the unmet needs mentioned above, optimized (transgene-free) transient expression systems in plant cells are needed. Summary of the Invention
[0008] The present invention is outlined below: Embodiment 1. A method for producing a plant cell comprising a modified RNA molecule, comprising: i) providing a plant cell; and ii) introducing the modified RNA molecule into a plant cell Including, A method wherein the modified RNA molecule comprises a modified uridine, the modified uridine being at least one of pseudouridine and N1-methyl-pseudouridine.
[0009] Embodiment 2. The modified RNA molecule is a modified messenger (m)RNA molecule comprising a 5'-UTR, a coding sequence, and a 3'-UTR; the 5'-UTR comprises a 5'-UTR of a positive-strand RNA virus or a 5'-UTR of a plant RNA transcript; The 3'-UTR contains a poly(A) tail. 2. The method of embodiment 1, wherein the modified mRNA molecule has increased expression from the coding sequence compared to an identical unmodified mRNA molecule.
[0010] Embodiment 3. The method of embodiment 2, wherein the 5'-UTR of the modified mRNA molecule is a 5'-UTR of a potyvirus, preferably, the 5'-UTR is a 5'-UTR of Tobacco Etch Virus (TEV).
[0011] Embodiment 4. The method of embodiment 2 or 3, wherein the 5'-UTR of the modified mRNA molecule has at least 80% sequence identity to any one of SEQ ID NOs: 1-46.
[0012] Embodiment 5. The method of any one of the preceding embodiments, wherein the modified RNA molecule comprises a 5'-cap.
[0013] Embodiment 6. The method of any one of the preceding embodiments, wherein not all uridines in the modified RNA molecule are modified uridines, preferably the percentage of uridines that are modified uridines is less than 95%, preferably about 10%-90% of the total uridines are modified uridines.
[0014] Embodiment 7. The method of any one of the preceding embodiments, wherein the only modification of the modified RNA molecule is modification of uridine to pseudouridine and / or N1-methyl-pseudouridine.
[0015] Embodiment 8. The method of any one of embodiments 2 to 7, wherein the coding sequence of the modified mRNA molecule encodes at least one of a site-specific nuclease and a plant morphogenetic polypeptide.
[0016] Embodiment 9. i) the site-specific nuclease is selected from the group consisting of TAL effector nucleases (TALENS), zinc finger nucleases (ZFNs), and CRISPR Cas proteins; and ii) the morphogenic polypeptide is a PLETHORA (PLT) polypeptide or a WUS / WOX homeobox polypeptide; and preferably at least one of - the PLT polypeptide is selected from the group consisting of PLT1, PLT2, PLT3, PLT4, PLT5 and PLT7, preferably PLT1; - the method of embodiment 8, wherein the WUS / WOX homeobox polypeptide is selected from the group consisting of WUS1, WUS2, WUS3, WOX2A, WOX4, WOX5, or WOX9, preferably WOX5.
[0017] Embodiment 10. The method of embodiment 9, wherein the coding sequence of the modified RNA molecule encodes a TALEN, preferably a TALEN having at least 80% sequence identity to any one of SEQ ID NOs: 47-50.
[0018] Embodiment 11. A method according to any one of embodiments 2 to 10, wherein in step ii), a first and a second modified RNA molecule are introduced, wherein the coding sequence of the first modified RNA molecule encodes a first portion of a TALEN and the coding sequence of the second modified RNA molecule encodes a second portion of the TALEN, and the first and second portions of the TALEN form a functional TALEN when expressed in a plant cell.
[0019] Embodiment 12. The method of any one of embodiments 8 to 11, wherein the plant cells produced comprise a targeted genome modification.
[0020] Embodiment 13. The method of embodiment 8, wherein the coding sequence of the modified mRNA molecule encodes a site-specific nuclease, preferably a CRISPR nuclease.
[0021] Embodiment 14. The method of embodiment 13, wherein in step ii) the modified mRNA molecule is introduced in combination with a guide RNA.
[0022] Embodiment 15. The method of any one of the preceding embodiments, wherein in step ii), the modified RNA molecule is introduced by at least one of PEG transformation, cell-penetrating peptide (CPP), and biolistics.
[0023] Embodiment 16. The method of any one of the preceding embodiments, wherein the plant cells provided in step i) are part of a multicellular tissue.
[0024] Embodiment 17. The method according to claim 1, further comprising the step iii) of selecting the plant cells produced or their progeny, wherein the selected plant cells are: - modified RNA molecules; - a protein expressed from the modified RNA molecule; and - genome sequence containing targeted genome modifications 10. The method of any one of the preceding embodiments, comprising at least one of:
[0025] Embodiment 18. The method includes a step iii) of selecting the plant cells produced, or their progeny, wherein the plant cells are: - modified RNA molecules; - a protein expressed from the modified RNA molecule; and - genome sequence containing targeted genome modifications 10. The method of any one of the preceding embodiments, selected to include at least one of:
[0026] Embodiment 19 The method of any one of the preceding embodiments, further comprising regenerating a plant from the plant cell, optionally the selected plant cell.
[0027] Embodiment 20. A modified RNA molecule according to any one of embodiments 1 to 11.
[0028] Embodiment 21. A modified mRNA molecule comprising a modified uridine, the modified uridine is at least one of pseudouridine and N1-methyl-pseudouridine; The mRNA molecule comprises a 5'-UTR, a coding sequence, and a 3'-UTR, the 5'-UTR comprises a 5'-UTR of a positive-strand RNA virus or a 5'-UTR of a plant RNA transcript; The 3'-UTR contains a poly(A) tail. Modified mRNA molecules.
[0029] Embodiment 22. The modified mRNA molecule of embodiment 21, wherein the coding sequence encodes a site-specific nuclease, preferably a CRISPR nuclease.
[0030] Embodiment 23. A plant cell comprising a modified RNA molecule according to any one of embodiments 1 to 11, preferably wherein the plant cell is a protoplast.
[0031] Embodiment 24. A plant cell comprising a modified RNA molecule according to embodiment 21 or 22, preferably wherein the plant cell is a protoplast.
[0032] Embodiment 25. A plant cell according to embodiment 23 or 24, wherein the plant cell is a tomato (Solanum Lycopersicon) plant cell.
[0033] Embodiment 26. Use of a modified RNA molecule according to any one of embodiments 1 to 11 for the transient expression of a gene product in a plant cell.
[0034] Embodiment 27. Use of a modified RNA molecule according to embodiment 21 or 22 for the transient expression of a gene product in a plant cell.
[0035] Embodiment 28. Use of a modified RNA molecule according to embodiment 22 for the targeted modification of a plant cell.
[0036] definition Various terms relating to the methods, compositions, uses, and other aspects of the present invention are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art to which the invention pertains, unless otherwise indicated. Other specifically defined terms are to be interpreted consistent with the definition set forth herein.
[0037] It will be apparent to one skilled in the art that any methods and materials similar or equivalent to those described herein can be used to practice the present invention.
[0038] Those skilled in the art will understand how to carry out the conventional techniques used in the methods of the present invention. The practice of conventional techniques in molecular biology, biochemistry, computational chemistry, cell culture, recombinant DNA, bioinformatics, genomics, sequencing, and related fields is well known to those skilled in the art and is discussed, for example, in the following references: Sambrook et al., Molecular Cloning. A Laboratory Manual, 2nd Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989; Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1987 and periodic updates; and series Methods in Enzymology, Academic Press, San Diego.
[0039] The singular terms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a combination of two or more cells, etc. Thus, the indefinite article "a" or "an" typically means "at least one."
[0040] The term "and / or" refers to a situation in which one or more of the stated cases may occur alone or together with at least one of the stated cases (up to all of the stated cases).
[0041] As used herein, the term "about" is used to describe and account for slight variations. For example, this term can refer to ±(+ or -) 10% or less, e.g., ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1% or less, or ±0.05% or less. Furthermore, quantities, ratios, and other numerical values may be expressed in range format herein. It is understood that such range formats are used for convenience and clarity and should be interpreted flexibly to include not only the numerical values explicitly stated as the limits of the range, but also all individual numerical values or subranges subsumed within the range, as if each numerical value and subrange were explicitly stated. For example, a ratio in the range of about 1 to about 200 should be understood to include not only the explicitly stated limits of about 1 to about 200, but also individual ratios such as about 2, about 3, and about 4, and subranges such as about 10 to about 50, about 20 to about 100, etc.
[0042] The term "comprises" should be interpreted as inclusive and open-ended, and not exclusive. Specifically, this term and its variations mean that the specified features, steps, or components are included. These terms should not be interpreted to exclude the presence of other features, steps, or components.
[0043] The terms "protein" or "polypeptide" are used interchangeably herein to refer to a molecule consisting of a chain of amino acids, regardless of a specific mode of action, size, three-dimensional structure, or origin. Thus, a "fragment" or "portion" of a protein may also be referred to as a "protein." An "isolated protein" is used to refer to a protein that is no longer in its natural environment, for example, in vitro or in a recombinant bacterial or plant host cell.
[0044] "Plant" refers to a whole plant or a portion of a tissue or organ obtained from a plant, such as pollen, seeds, gametes, roots, leaves, flowers, flower buds, anthers, fruit, etc., and any derivatives thereof, and progeny obtained from such a plant by self-pollination or cross-pollination. Non-limiting examples of plants include crop plants and cultivated plants, such as African eggplant, allium, artichoke, asparagus, barley, beet, bell pepper, bitter melon, ground cherry, bottle gourd, cabbage, canola, carrot, cassava, cauliflower, celery, chicory, kidney bean, corn salad, cotton, cucumber, eggplant, endive, fennel, gherkin, grape, chili pepper, lemongrass, and the like. Lettuce, corn, melon, rapeseed, okra, parsley, parsnip, pepino, pepper, potato, pumpkin, radish, rice, loofah, rocket, rye, snake gourd, sorghum, spinach, loofah, pumpkin, sugar beet, sugarcane, sunflower, tomatillo, tomato, tomato rootstock, Brassica vegetables, watermelon, wax gourd, wheat, and zucchini.
[0045] "Plant cells" include protoplasts, gametes, suspension cultures, microspores, pollen grains, etc., either isolated or of tissue, organ or organism derived from a plant source. A plant cell can be part of a multicellular structure such as, for example, a callus, a meristem, a plant organ, or an explant.
[0046] "Similar conditions" for culturing plants / plant cells means, inter alia, the use of similar temperature, humidity, nutrient and light conditions, as well as similar irrigation and day / night rhythms.
[0047] As used herein, the terms "homology," "sequence identity," and the like are used interchangeably. Sequence identity is defined herein as the relationship between two or more amino acid (polypeptide or protein) sequences or two or more nucleotide (polynucleotide) sequences, as determined by comparing the sequences. In the art, "identity" refers to the degree of sequence relatedness between amino acid sequences or nucleic acid sequences, as the case may be, as determined by the match between stretches of such sequences. "Similarity" between two amino acid sequences is determined by comparing the amino acid sequence of one polypeptide and its conserved amino acid substitutes with the sequence of a second polypeptide. "Identity" and "similarity" can be readily calculated by known methods. The percentage of sequence identity / similarity can be determined over the entire length of the sequences.
[0048] As used herein, "sequence identity" refers to the degree to which two optimally aligned polynucleotide or peptide sequences remain constant throughout the window of alignment of components, e.g., nucleotides or amino acids. The "fractional identity" of an aligned segment of a test sequence and a reference sequence is the number of identical components shared by the two aligned sequences divided by the total number of components in the reference sequence segment, i.e., the entire reference sequence or a smaller, defined portion of the reference sequence. "Percent identity" is the fractional identity multiplied by 100.
[0049] "Sequence identity" and "sequence similarity" can be determined by aligning two peptide sequences or two nucleotide sequences using a global or local alignment algorithm, depending on the length of the two sequences. Sequences of similar length are preferably aligned using a global alignment algorithm (e.g., Needleman Wunsch) that optimally aligns the sequences over their entire length, while sequences of substantially different lengths are preferably aligned using a local alignment algorithm (e.g., Smith Waterman). Sequences can be referred to as "substantially identical" or "essentially similar" if they share at least a certain minimum percentage of sequence identity (as defined herein) (e.g., when optimally aligned using the programs GAP or BESTFIT with default parameters). The percentage of sequence identity is preferably determined using "BESTFIT" or "GAP" from the Sequence Analysis Software Package™ (Version 10; Genetics Computer Group, Inc., Madison, Wis.). GAP uses the Needleman-Wunsch global alignment algorithm (Needleman and Wunsch, Journal of Molecular Biology 48:443-453, 1970) to align two sequences over their full length, maximizing the number of matches and minimizing the number of gaps. Global alignment is suitable for use in determining sequence identity when two sequences are similar in length. Generally, the GAP default parameters are used, with a gap creation penalty of 50 (nucleotides) / 8 (proteins) and a gap extension penalty of 3 (nucleotides) / 2 (proteins). For nucleotides, the default scoring matrix used is nwsgapdna, and for proteins, the default scoring matrix is Blosum62 (Henikoff & Henikoff, 1992, PNAS 89, 915-919).Sequence alignments and percent sequence identity scores can be determined using computer programs such as the GCG Wisconsin Package; Version 10.3, available from Accelrys Inc., 9685 Scranton Road, San Diego, CA 92121-3752 USA, or using the programs "needle" (using the global Needleman-Wunsch algorithm) or "water" (using the local Smith-Waterman algorithm) in EmbossWIN version 2.10.0, using the same parameters as for GAP described above, or using default settings (for both "needle" and "water," and for both protein and DNA alignments, the default gap opening penalty is 10.0, the default gap extension penalty is 0.5; the default scoring matrix is Blossum62 for proteins and DNAFull for DNA). The gap extension penalty is 0.5; the default scoring matrix is Blossum62 for proteins and DNAFull for DNA). "BESTFIT" uses the local homology algorithm of Smith and Waterman (Smith and Waterman, Advances in Applied Mathematics, 2:482-489, 1981; Smith et al., Nucleic Acids Research 11:2205-2220, 1983) to optimally align the most similar segments between two sequences and insert gaps to maximize the number of matches. When the total length of the sequences differs substantially, local alignments such as the Smith-Waterman algorithm are preferred.
[0050] Useful methods for determining sequence identity are also disclosed in Guide to Huge Computers, Martin J. Bishop, ed., Academic Press, San Diego, 1994, and Carillo, H., and Lipton, D., Applied Math (1988) 48:1073. Among other computer programs suitable for determining sequence identity are the Basic Local Alignment Search Tool (BLAST) programs, publicly available from the National Center for Biotechnology Information (NCBI), National Library of Medicine, National Institutes of Health, Bethesda, Md. 20894; see BLAST Manual, Altschul et al., NCBI, NLM, NIH; Altschul et al., J. Mol. Biol. 215:403-410 (1990); BLAST program versions 2.0 and later allow gaps (deletions and insertions) to be introduced into the alignment; for peptide sequences, BLASTX can be used to determine sequence identity; for polynucleotide sequences, BLASTN can be used to determine sequence identity.
[0051] Alternatively, percent similarity or identity may be determined by searching public databases using algorithms such as FASTA, BLAST, etc. Thus, the nucleic acid and protein sequences of the present invention can also be used as "query sequences" to search public databases, for example, to identify other family members or related sequences. Such searches can be performed using the BLASTn and BLASTx programs (version 2.0) of Altschul, et al. (1990) J. Mol. Biol. 215:403-10. BLAST nucleotide searches can be performed using the NBLAST program, score = 100, word length = 12, to obtain nucleotide sequences homologous to the nucleic acid molecules of the present invention. BLAST protein searches can be performed with the BLASTx program, score = 50, word length = 3, to obtain amino acid sequences homologous to the protein molecules of the present invention. To obtain gapped alignments for comparison purposes, Gapped BLAST can be used as described in Altschulet et al. (1997) Nucleic Acids Res. 25(17):3389-3402. When using BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., BLASTx and BLASTn) can be used. See the National Center for Biotechnology Information homepage at http: / / www.ncbi.nlm.nih.gov / .
[0052] "Similar to," with respect to a domain, sequence, or position of a protein relative to a designated domain, sequence, or position of a reference protein, is understood herein as a domain, sequence, or position that aligns with the designated domain, sequence, or position of the reference protein when the protein is aligned with the reference protein using an alignment algorithm such as Needleman Wunsch, as described herein.
[0053] "Similar to," with respect to a domain, sequence, or position of a nucleic acid relative to a designated domain, sequence, or position of a reference nucleic acid, is understood herein as a domain, sequence, or position that aligns with the designated domain, sequence, or position of the reference nucleic acid when the nucleic acid is aligned with the reference nucleic acid using an alignment algorithm such as Needleman Wunsch, as described herein.
[0054] A "nucleic acid" or "polynucleotide" according to the present invention can comprise any polymer or oligomer of pyrimidine and purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, at 793-800 (Worth Pub. 1982), incorporated herein by reference in its entirety for all purposes). The present invention contemplates any deoxyribonucleotide, ribonucleotide, or nucleic acid component, and any chemical variant thereof, e.g., methylated, hydroxymethylated, or glycosylated forms of the above bases. Particularly preferred variants are pseudouridine and N1-methyl-pseudouridine. The polymer or oligomer can be heterogeneous or homogeneous in composition and can be isolated from naturally occurring sources or produced artificially or synthetically. Furthermore, the nucleic acid may be DNA (optionally cDNA) or RNA, or a mixture thereof, and may exist permanently or transiently in single- or double-stranded form, including homoduplexes, heteroduplexes, and hybrid states. "Isolated nucleic acid" is used herein to refer to a nucleic acid that is no longer in its natural environment, for example, in vitro or in a recombinant bacterial or plant cell. The nucleic acids and / or proteins of the present invention may be at least one of recombinant, synthetic, or artificial nucleic acids and / or proteins.
[0055] The terms "nucleic acid construct," "nucleic acid vector," "vector," and "expression construct" are used interchangeably herein and are defined herein as artificial nucleic acid molecules resulting from the use of recombinant DNA technology. Thus, the terms "nucleic acid construct" and "nucleic acid vector" do not include naturally occurring nucleic acid molecules, although a nucleic acid construct may include (parts of) naturally occurring nucleic acid molecules.
[0056] The vector backbone is known in the art and described elsewhere herein, and may be, for example, a binary or superbinary vector (see, e.g., U.S. Pat. No. 5,591,616, U.S. Patent Application No. 2002138879, and WO 95 / 06722), an integrative vector, or a T-DNA vector, into which the chimeric gene is incorporated, or, if appropriate transcriptional regulatory sequences are already present, into which only the desired nucleic acid sequence (e.g., coding sequence, antisense sequence, or inverted repeat sequence) is incorporated downstream of the transcriptional regulatory sequence. Vectors may contain additional genetic elements, such as selectable markers, multiple cloning sites, etc., to facilitate their use in molecular cloning.
[0057] The term "gene" refers to a nucleic acid fragment comprising a region (transcribed region) that is transcribed into an RNA molecule (e.g., mRNA) in a cell and operably linked to a suitable regulatory region (e.g., a promoter). A gene may comprise several operably linked fragments, such as a promoter sequence, a 5' untranslated region, a coding region, and a 3' untranslated sequence containing a polyadenylation site. The promoter sequence of a gene may bind transcription factors that recruit RNA polymerase and help the RNA polymerase initiate transcription. Apart from the promoter sequence, regulatory sequences may further comprise sites that act as enhancers and / or silencers of transcription, for example, by binding to specific enhancer or inhibitory elements and / or by influencing chromatin structure. The transcribed region of a gene may be annotated herein as an open reading frame (ORF), beginning with a three-letter code designated as a start codon and ending with one of three possible stop codons. An ORF may comprise exons and one or more introns. The transcribed region is preceded by a 5'-UTR and followed by a 3'-UTR, which constitute the boundaries of the transcribed RNA. In the case of pre-mRNA, during maturation into mRNA, introns are spliced out, and a polyadenylation tail (abbreviated as polyA tail) and, optionally, a 5'-cap are added to the 3' and 5' ends of the RNA, respectively, to form the mature mRNA. Thus, the mature mRNA contains at least the following nucleotide sequence elements: a 5'-UTR, exons, a 3'-UTR, and a polyA-tail. Optionally, the mRNA also contains the following nucleotide sequence elements: a 5'-cap, a 5'-UTR, exons, a 3'-UTR, and a polyA-tail. The composite exons of an ORF, i.e., the sequence that is translated into a protein, are referred to herein as a coding sequence or CDS. The mature mRNA is then translated into a protein in a process called translation.
[0058] "Expression" with respect to (m)RNA refers to the process by which said (m)RNA is translated into a biologically active protein. This process is also referred to as "translation" or "protein expression." "Expression" with respect to a gene refers to the process by which a nucleic acid region operably linked to an appropriate regulatory region, particularly a promoter, is transcribed into a biologically active RNA, e.g., a RNA that can be translated into a biologically active protein, or, e.g., a regulatory non-coding RNA (this process is sometimes referred to herein as "transcription"). Optionally, "expression" with respect to a gene can encompass both the processes of transcription and translation.
[0059] The term "operably linked" refers to a linkage of polynucleotide elements in a functional relationship. A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. For example, a promoter or transcriptional regulatory sequence is operably linked to a coding sequence if it affects the transcription of the coding sequence. Operably linked can mean that the linked nucleic acid sequences are contiguous.
[0060] "Promoter" refers to a nucleic acid fragment that functions to control the transcription of one or more nucleic acids. A promoter fragment is preferably located upstream (5') of a gene's transcription initiation site in the direction of transcription, and is structurally distinguished by the presence of an RNA polymerase binding site and a transcription initiation site, and may further include any other nucleic acid sequences, such as, but not limited to, transcription factor binding sites, repressor and activator protein binding sites, and other sequences of nucleotides known to those skilled in the art to act directly or indirectly to regulate the amount of transcription from the promoter.
[0061] A "constitutive" promoter is a promoter that is active in most tissues and / or under most physiological and developmental conditions. An "inducible" promoter is a promoter that is physiologically (e.g., by the external application of certain compounds) or developmentally regulated. A "tissue-specific" promoter is active only in specific tissue or cell types.
[0062] The term "cDNA" refers to complementary DNA. Complementary DNA is produced by reverse transcribing RNA into a complementary DNA sequence. Thus, a cDNA sequence corresponds to the RNA sequence expressed from a gene. Because RNA sequences expressed from the genome may undergo splicing before being translated into proteins in the cytoplasm, i.e., introns are spliced out from pre-mRNA and exons are joined together, the sequence of a cDNA is understood to correspond to the sequence of an mRNA. Thus, in the case of a protein, a cDNA may encode only a complete open reading frame consisting of joined exons, whereas a genomic DNA sequence may contain exon sequences flanked by intron sequences, and therefore a cDNA sequence may not be identical to its corresponding genomic DNA sequence. Genetic modification of a protein-encoding gene may involve not only modification of the protein-encoding sequence but also mutations in the intron sequences and / or other gene regulatory sequences of the genomic DNA.
[0063] The term "regeneration" is defined herein as the formation of new tissues and / or new organs from a single plant cell, callus, explant, tissue, or organ. The regeneration pathway can be somatic embryogenesis or organogenesis. Somatic embryogenesis is understood herein as the formation of a somatic embryo that can be developed to regenerate a whole plant. Organogenesis is understood herein as the formation of a new organ from an (undifferentiated) cell. Preferably, the regeneration is at least one of ectopic apical meristem formation, shoot regeneration, and root regeneration. Regeneration as defined herein can preferably involve at least de novo shoot formation. For example, regeneration can be the regeneration of an (n) (elongated) hypocotyl explant into an (n) (inflorescence) shoot. Regeneration also includes the formation of a new plant from a single plant cell, or from a callus, explant, tissue, or organ. The regeneration process can occur directly from the parent tissue or indirectly, for example, via callus formation.
[0064] As used herein, "conditions that allow regeneration" are understood to mean an environment in which plant cells or tissues can regenerate. Such conditions include at least suitable temperature (i.e., 1°C to 60°C), nutrients, day-night rhythm, irrigation, and one or more plant hormones and / or plant hormone-like compounds. Furthermore, "optimal conditions that allow regeneration" are environmental conditions that allow maximum regeneration of plant cells.
[0065] The term "wild-type" as used in conjunction with a protein or nucleic acid in the context of the present invention means that said protein or nucleic acid consists of an amino acid sequence or nucleotide sequence, respectively, that occurs in its entirety in nature and that can be isolated intact from a natural organism, and has not been obtained by modification techniques such as, for example, targeted or random mutagenesis. A wild-type protein is expressed, for example, as it occurs in nature, under particular environmental conditions and at least at a particular developmental stage.
[0066] The term "endogenous" as used in connection with a protein or nucleic acid in the context of the present invention means that the protein or nucleic acid is still contained within the plant, i.e., is present in its natural environment. In many cases, an endogenous gene is present in its normal genetic background in the plant.
[0067] "Targeted mutagenesis" refers to mutagenesis that can be designed to modify specific nucleotides or nucleic acid sequences, such as, but not limited to, oligo-directed mutagenesis, RNA-guided endonucleases (e.g., CRISPR technology), TALEN, or zinc finger technology.
[0068] The term "sequence of interest" includes, but is not limited to, genetic sequences that are preferably present in a cell, such as genes, portions of genes, and non-coding sequences within or adjacent to genes. Sequences of interest can be present in, for example, chromosomes, episomes, organelle genomes such as mitochondrial or chloroplast genomes, or genetic material, but they can also exist independently of the body of genetic material, such as infectious viral genomes, plasmids, episomes, and transposons. A sequence of interest can be present within the coding sequence of a gene or within a transcribed non-coding sequence, such as a leader sequence, trailer sequence, or intron. The sequence of interest can be present in a double-stranded or single-stranded nucleic acid molecule. Preferably, the nucleic acid sequence is present in a double-stranded nucleic acid molecule. A sequence of interest can be any sequence within a nucleic acid, such as a gene, gene complex, locus, pseudogene, regulatory region, highly repetitive region, polymorphic region, or portion thereof. A sequence of interest can also be a region containing a gene or epigenetic mutation that is indicative of a phenotype or disease. Preferably, the sequence of interest is a short or long contiguous stretch of nucleotides (i.e., a polynucleotide) of double-stranded DNA, said double-stranded DNA further comprising a sequence complementary to the target sequence in the complementary strand of said double-stranded DNA.
[0069] As referred to herein, a "control plant" is a plant of the same species as the plant of the present invention, and preferably having the same genetic background, i.e., the plant that has been subjected to the methods taught herein. Preferably, the control plant differs from the putative test plant only in that the control plant does not contain a targeted modification as detailed herein. Preferably, the control plant is grown under the same conditions as the plant that is subjected to the methods of the present invention.
[0070] "Guide RNA" or "gRNA" is herein understood as an RNA molecule comprising a guide sequence for targeting a gRNA-CAS complex, preferably to a protospacer sequence near, at, or within a sequence of interest within a nucleic acid molecule, and may be an sgRNA, or a combination of crRNA and tracrRNA (e.g., in the case of Cas9), or crRNA alone (e.g., in the case of Cpfl). Optionally, two or more guide RNAs may be used in the same experiment, e.g., targeting two or more different sequences of interest, or the same sequence of interest.
[0071] As used herein, a "guide sequence" is understood to be a portion of a guide RNA that recognizes, binds to, and / or hybridizes to a specific site in an RNA or DNA molecule. Preferably, the guide sequence is the portion of the sgRNA or crRNA required to target the gRNA-CAS complex to a specific site in the double-stranded DNA. DETAILED DESCRIPTION OF THE INVENTION
[0072] In plants, transient expression of a transgene is often more desirable than persistent expression, for example, due to regulatory challenges or when long-term expression inhibits plant development. However, transient expression is often too short to achieve the desired effect. Therefore, transient expression ideally lasts at least long enough to achieve the effect. The present inventors have developed a method to achieve such desirable long-term transient expression.
[0073] Thus, in a first aspect, there is provided a method of producing a plant cell, wherein the produced plant cell comprises a modified RNA molecule. Preferably, the method comprises: i) providing a plant cell; and ii) introducing the modified RNA molecule into a plant cell Includes.
[0074] The modified RNA molecule as defined herein comprises a modified uridine, wherein the modified uridine is at least one of pseudouridine and N1-methyl-pseudouridine. Accordingly, in this specification, a modified RNA molecule is understood as an RNA molecule comprising a modified uridine, wherein the modified uridine is preferably at least one of pseudouridine and N1-methyl-pseudouridine.
[0075] Optionally, the modified RNA molecule is a non-coding RNA. Optionally, the non-coding RNA plays a role in regulating gene expression and / or guiding RNA-guided endonucleases to specific nucleotide sequences. Non-coding RNA may be involved in gene translation (e.g., rRNA or tRNA), splicing (e.g., snRNA), ribosomal RNA modification (e.g., snoRNA), regulating gene expression at the post-transcriptional level (e.g., microRNA, siRNA, or piRNA), or may be involved in chromatin remodeling, transcriptional regulation, and / or post-transcriptional processing (e.g., lnoRNA) (Kukurba and Montgomery, Cold Spring Harb Protoc 2015;11:951-969). Thus, the modified RNA molecule may be selected from the group consisting of guide RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), microRNA (miRNA), small interfering RNA (siRNA), trans-acting siRNA, tasiRNA, piwi-interacting RNA (piRNA) and long non-coding RNA (lnoRNA).
[0076] Alternatively, the modified RNA molecule may be a coding RNA, i.e., a messenger RNA (mRNA) that encodes a protein. Preferably, the modified RNA molecule comprises at least one of a 5'-UTR, a coding sequence, and a 3'-UTR. Preferably, the modified RNA molecule has increased protein expression from its coding sequence, preferably a coding sequence as defined herein, compared to an identical unmodified RNA molecule.
[0077] An identical unmodified RNA molecule is herein understood to mean a molecule that contains the same nucleotide sequence as a modified RNA molecule, but does not contain any pseudouridine or N1-methylpseudouridine, and preferably does not contain any nucleic acid modifications. Thus, an identical unmodified RNA molecule is preferably completely identical to a modified RNA molecule, except that the unmodified RNA molecule does not contain the uridine modifications described, i.e., does not contain substitutions of uridine with pseudouridine and / or N1-methyl-pseudouridine. If a modified RNA molecule contains a 5' cap, preferably the control (unmodified) RNA molecule also contains a 5' cap.
[0078] Increased protein expression is seen in: i) the increase in total amount (i.e., the sum of the amount of protein expressed over a defined period after the transfection event); ii) an increase in protein levels (i.e., the level of the protein at a particular time point after transfection, preferably the peak level of the protein after transfection); and iii) Increased protein levels over time (i.e., the longer the modified mRNA is present, the longer the protein can be detected in the cell). It can be at least one of:
[0079] The level of translated protein and / or the total amount of translated protein can be increased compared to the same unmodified RNA molecule. The total amount and / or protein level can preferably be increased by at least 1.2, 1.4, 1.6, 1.8, 2, 2.5, 3, 3.5, 4.5, 5, 6, 7, 8, 9, or at least about 10-fold when determined about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more days after introduction of the RNA molecule. Optionally, the total amount of protein is increased when determined over a period of about 2, 3, 4, 5, 6, 7, 8, 9, 10, or more days, the first day of which period preferably beginning 1 day after introduction of the RNA molecule.
[0080] An increase in the total amount of protein and / or increased levels of protein expressed from modified RNA compared to the same unmodified RNA molecule, preferably under otherwise identical experimental conditions (e.g., the same transfection method, the same or similar host cells, and the same or similar culture conditions), is believed to be due to at least one of increased translation efficiency and increased RNA retention in the cell.
[0081] Alternatively, or in addition, the level of a protein encoded by a modified RNA molecule may be increased in plant cells over a longer period of time compared to the level of a protein encoded by the same unmodified RNA molecule. The increased level over a longer period of time may be due to the presence of the modified RNA molecule for a longer period of time compared to the same unmodified RNA molecule. Preferably, the protein expressed from the modified RNA molecule can be detected for a longer period of time compared to the protein expressed from the same, but unmodified, RNA molecule under otherwise identical experimental conditions. Preferably, the protein encoded by the modified RNA molecule can be detected in plant cells for at least another 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more days when the protein encoded by the same unmodified RNA molecule cannot be detected any more. Preferably, the modified RNA molecule can still be detected in plant cells for at least another 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more days when the protein encoded by the same unmodified RNA molecule cannot be detected any more.
[0082] Optionally, increased expression can be determined by transfecting plant cells with a modified RNA molecule and then transfecting the plant cells with the same unmodified RNA molecule, where the separate transfections are preferably performed using the same experimental conditions. Because the modified RNA can be present for an extended period of time after transfection, the number of cells expressing a protein encoded by the modified RNA molecule can be increased compared to the number of cells expressing a protein encoded by the same unmodified RNA molecule under otherwise identical circumstances. The number of cells expressing a protein from a modified RNA molecule is preferably increased by at least 1.2, 1.4, 1.6, 1.8, 2, 2.5, 3, 3.5, 4.5, 5, 6, 7, 8, 9, or at least about 10-fold compared to the number of cells expressing said protein from an unmodified RNA molecule, when determined at about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more days after introduction of the RNA molecule. Preferably, the number of cells expressing the protein is increased by at least about two-fold after transfection with the modified RNA molecule compared to the number of cells expressing the protein after transfection with the same unmodified RNA molecule, preferably determined after about two days.
[0083] The detection and, optionally, quantification of proteins can be carried out using any common method known to those skilled in the art, such as, but not limited to, Western blotting, FACS analysis, microscopy, etc. Similarly, the detection and, optionally, quantification of nucleic acids can be carried out using any common method known to those skilled in the art, such as, but not limited to, (quantitative) PCR, deep sequencing, Northern blotting, etc.
[0084] Introduction of the modified RNA molecule preferably results in the translation of a protein encoded by the RNA molecule. The translated protein can subsequently have a specific effect on the plant cell, preferably a measurable and / or phenotypic effect. As a non-limiting example, the expressed protein can target a specific region in the plant genome, resulting in a (detectable) site-directed mutagenesis event. As another non-limiting example, the expressed protein can induce and / or enhance plant cell regeneration.
[0085] The methods detailed herein may also be considered, for example, as follows: - A method for increasing transient gene expression in plant cells; - a method for producing plant cells containing targeted genome modifications; - a method for producing plant cells expressing a transgene; - Methods for site-directed mutagenesis in plant cells; - a method for producing de novo shoots; and - Methods of de novo shoot regeneration.
[0086] Those skilled in the art will readily appreciate that additional methods involving the introduction of modified RNA molecules into plant cells are also part of the present invention.
[0087] plant cells Step i) of the provided method is providing a plant cell. The plant cell may be an isolated cell or part of a multicellular structure such as a callus, meristem, plant part, plant organ, or explant. Those skilled in the art will readily understand that the method of the present invention is not limited to a particular plant cell type. In particular, the method of the present invention disclosed herein can be applied to dividing cells as well as non-dividing cells. The cell may be transgenic or non-transgenic. The plant cell can be obtained, for example, from plant cell tissue cultures capable of regenerating plants, plant callus, plant mass, and intact plant cells in plants or plant parts such as embryos, pollen, ovaries, seeds, leaves, flowers, branches, fruits, kernels, panicles, cobs, stems, roots, root tips, anthers, and grains. A preferred plant cell is a protoplast.
[0088] The plant cell may be a cell from any plant, such as a cultivated plant or a wild-type plant. The plant cell may be a cell from a crop plant or a grain plant. The plant cell may preferably be obtained from a crop plant, such as a monocotyledonous or dicotyledonous plant, or a crop or grain plant, such as cassava, maize, sorghum, soybean, wheat, oat, or rice. Crop plants are plant species cultivated and bred by humans. Crop plants may be grown for food and / or feed purposes (e.g., field crops), or for ornamental purposes (e.g., production of cut flowers, lawn grass, etc.). Crop plants, as defined herein, also include plants from which non-food products are harvested, such as fuel oil, plastic polymers, pharmaceuticals, cork, etc.
[0089] The plant cell may be derived from algae, trees or productive plants, fruit or vegetables (e.g., citrus trees, e.g., orange, grapefruit or lemon trees; peach or nectarine trees; apple or pear trees; nut trees such as almond or walnut or pistachio trees; nightshade plants; plants of the genus Brassica; plants of the genus Lactuca; plants of the genus Spinacia; plants of the genus Capsicum; plants of the genus Solanum, preferably tomato (Solanum lycopersicum).
[0090] The plant cell may be, derived from, or obtainable from a plant belonging to the Brassicaceae, Cucurbitaceae, Fabaceae, Gramineae, Solanaceae, Asteraceae (Compositae), Rosaceae, or Poaceae families.
[0091] Preferably, the plant cell is selected from the group consisting of maize / corn (Zea spp.), wheat (Triticum spp.), barley (e.g., Hordeum vulgare), oats (e.g., Avena sativa), sorghum (Sorghum bicolor), rye (Secale cereale), soybean (Glycine spp., e.g., Glycine max), cotton (Gossypium spp., e.g., G. hirsutum, G. barbadense), Brassica spp. spp.) (e.g., rapeseed (B. napus), mustard (B. juncea), kale (B. oleracea), Brassica rapa, etc.), sunflower (Helianthus annuus), safflower, yam, cassava, alfalfa (Medicago sativa), rice (Oryza species, e.g., the indica or japonica cultivars), forage grasses, pearl millet (Pennisetumspp.), e.g. pearl millet (P. glaucum), tree species (pine, poplar, fir, plantain, etc.), tea, coffee, oil palm, coconut, vegetable species, e.g. pea, zucchini, beans (e.g. Phaseolus spp.), pepper, cucumber, artichoke, asparagus, eggplant, broccoli, garlic, leek, lettuce, onion, radish, turnip, tomato, potato, Brussels sprouts, carrot, cauliflower, chicory, celery, spinach, endive, fennel, beet, fleshy fruiting plants (grape, or derived from a plant selected from the group consisting of: peach, plum, strawberry, mango, apple, plum, cherry, apricot, banana, blackberry, blueberry, citrus fruit, kiwi, fig, lemon, lime, nectarine, raspberry, watermelon, orange, grapefruit, etc.), ornamental species (e.g., rose, petunia, chrysanthemum, lily, Gerbera species), herbs (mint, parsley, basil, thyme, etc.), woody plants (e.g., Populus, Salix, Quercus, Eucalyptus species), fiber species such as flax (Linum usitatissimum) and hemp (Cannabis sativa, etc.). Plant tissue may also be derived from trees or producing plants, fruit or vegetables (e.g., citrus trees, e.g., orange, grapefruit or lemon trees; peach or nectarine trees; apple or pear trees; nut trees such as almond or walnut or pistachio trees; nightshade plants; Brassica plants; Lactuca plants; Spinacia plants; Capsicum plants; Solanum plants, preferably tomato (SolanumThe plant may be derived from Solanum lycopersicum. Preferably, the plant is a sorghum (Solanum) plant. Preferred plants for use in the methods provided herein are, are derived from, or are obtained from tomato (Solanum lycopersicum) or pepper (Capsicum annuum) plants.
[0092] In another preferred embodiment, the plant tissue is derived from a plant selected from the following: asparagus, barley, blackberry, blueberry, broccoli, cabbage, canola, carrot, cassava, cauliflower, chicory, cocoa, coffee, cotton, cucumber, eggplant, grape, chili pepper, lettuce, corn, melon, rapeseed, pepper, potato, pumpkin, raspberry, rice, rye, sorghum, spinach, squash, strawberry, sugarcane, sugar beet, sunflower, bell pepper, tobacco, tomato, watermelon, wheat, and zucchini.
[0093] The plant cell may be an isolated plant cell, preferably a protoplast, or the plant cell may be contained in a multicellular structure. The plant cell may be part of a specific plant structure, such as, but not limited to, pollen, a seed, a gamete, a root, a leaf, a flower, a flower bud, an anther, and / or a fruit. The modified RNA molecule may be introduced into cells of any part of the provided plant. The modified RNA molecule may be introduced into cells of the shoot system and / or cells of the root system. The modified RNA molecule may be introduced into cells of the root, stem, fruit, leaf, internode, and / or flower. The modified RNA molecule may be introduced into cells of a seedling. The modified RNA molecule may also be introduced into seeds. As a non-limiting example, the modified RNA molecule may be introduced into cells of the true leaves, epicotyl, cotyledon, hypocotyl, and / or radicle. Optionally, the modified RNA molecule is introduced into cells of the cotyledon. The modified RNA molecule can be introduced into cells of a provided plant, preferably into cells of the cotyledons of said plant, where the plant is a young seedling, which is composed of a radicle (embryonic root), hypocotyl (embryonic shoot), and cotyledons. Optionally, the modified RNA molecule is introduced into cells of a provided plant, preferably into cells of the cotyledons of said plant, about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 days, preferably about 7-10 days, after sowing.
[0094] modified RNA molecules Preferably, the introduced modified RNA molecule comprises a 5'-UTR, a coding sequence, and a 3'-UTR. Compared to the same unmodified RNA molecule, the modified RNA molecule defined herein results in increased expression of the encoded protein when introduced into a plant cell. The 5'-UTR of the modified RNA molecule preferably comprises the 5'-UTR of a positive-strand RNA virus or the 5'-UTR of a plant RNA transcript. Preferably, the 3'-UTR of the modified RNA molecule comprises a poly(A) tail.
[0095] Modified Uridine The RNA molecule used in the methods provided herein contains a modified uridine. The modified uridine is preferably at least one of pseudouridine (Ψ) and N1-methyl-pseudouridine (m1Ψ). The modified RNA molecule may contain uridine, pseudouridine, and N1-methyl-pseudouridine. Alternatively, the RNA molecule may contain only one type of uridine modification, i.e., pseudouridine or N1-methyl-pseudouridine. The RNA molecule may contain uridine and modified uridine. That is, the RNA molecule may contain uridine and pseudouridine and / or N1-methyl-pseudouridine. Alternatively, all uridines in the RNA molecule are substituted with modified uridines, i.e., pseudouridine and / or N1-methyl-pseudouridine.
[0096] Preferably, the modified RNA molecule comprises uridine and modified uridine, and the modified uridine is pseudouridine and / or N1-methyl-pseudouridine. Thus, preferably, not all uridines are substituted with modified uridines. Preferably, at most about 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 78%, 76%, 74%, 72%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, or at most about 20% of the total uridines in the RNA molecule are substituted with modified uridines. Preferably, at least about 8%, 10%, 12%, 14%, 16%, 18%, 20%, 22%, 24%, 26%, 28%, or at least about 30% of the uridines in the RNA molecule are substituted with modified uridines. Preferably, about 10% to 90%, 20% to 85%, 30% to 80%, 40% to 70%, or about 45% to 65% of all uridines in the RNA molecule are substituted with modified uridines. Preferably, about 10% to 90%, about 25% to 85%, or about 50% to 80% of all uridines in the RNA molecule are substituted with modified uridines. Uridines not substituted with modified uridines are unmodified uridines.
[0097] In one embodiment, at most about 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 78%, 76%, 74%, 72%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, or at most about 20% of the total uridines in the RNA molecule are substituted with pseudouridine. Preferably, at least about 8%, 10%, 12%, 14%, 16%, 18%, 20%, 22%, 24%, 26%, 28%, or at least about 30% of the uridines in the RNA molecule are substituted with pseudouridine. Preferably, about 10% to 90%, 20% to 85%, 30% to 80%, 40% to 70%, or about 45% to 60% of all uridines in the RNA molecule are substituted with pseudouridine. Preferably, about 10% to 90%, about 25% to 85%, or about 50% to 80% of all uridines in the RNA molecule are substituted with pseudouridine. Uridines not substituted with pseudouridine are unmodified uridines.
[0098] In another embodiment, at most about 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 78%, 76%, 74%, 72%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, or at most about 20% of the total uridines in the RNA molecule are substituted with N1-methyl-pseudouridine. Preferably, at least about 8%, 10%, 12%, 14%, 16%, 18%, 20%, 22%, 24%, 26%, 28%, or at least about 30% of the uridines in the RNA molecule are substituted with N1-methyl-pseudouridine. Preferably, about 10% to 90%, 20% to 85%, 30% to 80%, 40% to 70%, or about 45% to 60% of all uridines in the RNA molecule are substituted with N1-methyl-pseudouridine. Preferably, about 10% to 90%, about 25% to 85%, or about 50% to 80% of all uridines in the RNA molecule are substituted with N1-methyl-pseudouridine. Uridines not substituted with N1-methyl-pseudouridine are unmodified uridines.
[0099] Uridine or "unmodified uridine" is a naturally occurring glycosylated pyrimidine analog containing uracil attached to a ribose ring and one of the five standard nucleosides that make up (e.g., naturally occurring) nucleic acids. Pseudouridine (5-(β-D-ribofuranosyl)pyrimidine-2,4(1H,3H)-dione), "Ψ," or "5-ribosyluracil," is an isomer of the nucleoside uridine in which uracil is attached via a carbon-carbon linkage instead of a nitrogen-carbon glycosidic bond. Pseudouridine is known to be one of the most abundant RNA modifications in cellular RNA. N1-methyl-pseudouridine (5-[(2S,3R,4S,5R)-3,4-dihydroxy-5-(hydroxymethyl)oxolan-2-yl]-1-methylpyrimidine-2,4-dione) or "m1Ψ" is a naturally occurring archaeal tRNA component and a known synthetic pyrimidine nucleoside. It is a methylated derivative of pseudouridine and is used as a component of the SARS-CoV-2 mRNA vaccines Tojinamelan and Elasomeran.
[0100] Preferably, the only nucleic acid modification of the modified RNA molecule is a modified uridine as described herein. Preferably, the modified RNA molecule does not contain cytidine, guanosine, and / or adenosine modifications. The modified RNA molecule may include a 5'-cap.
[0101] Elements of an RNA molecule The modified RNA molecules used in the methods described herein preferably contain the following three elements: i) 5'-UTR; ii) a coding sequence; and iii) 3'-UTR Includes.
[0102] 5'-UTR The modified RNA molecule comprises a 5'-UTR. Preferably, the 5'-UTR induces or enhances translation of the coding sequence. Optionally, the 5'-UTR enhances translation even in the absence of a 5'm7G cap. Preferably, the 5'-UTR is from or derived from a 5'-UTR of a positive-strand RNA virus or a 5'-UTR of a plant transcript.
[0103] Preferably, the 5'-UTR is from or derived from the 5'-UTR of a positive-strand RNA virus. The genome of a positive-strand RNA virus can act as a messenger RNA and therefore can use the 5'-UTR to promote translation of the coding sequence. The 5'-UTR can be derived from a positive-strand RNA virus of the phylum Kitrinoviricota, Lenarviricota, and Pisviricota (particularly the classes Pisoniviricetes and Stelpaviricetes). Preferably, the 5'-UTR is obtained from or derived from a positive-strand RNA virus of the phylum Pisviricota, preferably the class Stelpaviricetes. Preferably, the sequence of the 5'-UTR is or is derived from a 5'-UTR of the order Patatavirales, preferably of the family Potyviridae. Preferably, the sequence of the 5'-UTR of the modified RNA molecule is or is derived from a 5'-UTR of the genus Potyvirus.
[0104] The 5'-UTR of the modified RNA molecule is preferably a 5'-UTR of or derived from a Potyvirus selected from the group consisting of: Bean common mosaic virus (BCMV), Bean common mosaic necrosis virus (BCMCV), Bean yellow mosaic virus (BYMV), Beet mosaic virus (BtMV), Chili vein mottle virus (ChiVMV), Clover yellow vein virus (ClYVV), Cocksfoot streak virus (CSV), Cowpea aphid-borne mosaic virus (CABMV), Daphne virus Y, Y) (DVY), Dasheen mosaic virus (DMV), East Asian Passiflora virus (EAPV), Fritillary virus Y (FVY), Japanese yam mosaic virus (JYMV), Johnsongrass mosaic virus (JGMV), Konjak mosaic virus (KoMV), Leek yellow stripe virus (LYSV), Lettuce mosaic virus (LMV), Lily mottle virus (LMoV), Maize dwarf mosaic virus (MDMV), Narcissus yellow stripe virus (NYSV), Onion dwarf virusOyster yellow dwarf virus (OYDV), Papaya leaf distortion mosaic virus (PLDMV), Papaya ringspot virus (PRSV), Pea seed-borne mosaic virus (PSbMV), Peanut mottle virus (PeMV), Peanut stripe virus (PStV), Pennisetum mosaic virus (PenMV), Pepper mottle virus (PepMoV), Peru tomato mosaic virus (PTV), Plum pox virus (PPV), Potato virus A (PVA), Potato virus V (PVV), Potato virus Y (PVY), Scallion mosaic virus (ScaMV), Shallot yellow stripe virus (SYSV), Soybean mosaic virus (SMV), Sugarcane mosaic virus (SCMV), Sweet potato feathery mottle virus (SPFMV), Thunberg fritillary mosaic virus (TFMV), Tobacco etch virus (TEV), Tobacco vein mottling virus (TVMV), Turnip mosaic virus (TuMV), Watermelon mosaic virus (WMV), Wild potato mosaic virus (WMV),virus (WPMV), Wisteria vein mosaic virus (WVMV), Yam mosaic virus (YMV), Zucchini yellow mosaic virus (ZYMV), and Ryegrass mosaic virus (RMV). A preferred 5'-UTR of the modified RNA molecule is from or derived from Tobacco etch virus (TEV). The 5'-UTR of the modified RNA molecule may have a sequence similar to the 5'-UTR of Tobacco etch virus (TEV).
[0105] Preferably, the 5'-UTR of the modified RNA molecule comprises a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or about 100% sequence identity to any one of SEQ ID NOs: 1-45. Preferably, the 5'-UTR of the modified RNA molecule comprises a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or about 100% sequence identity to SEQ ID NO: 37, the 5'-UTR of Tobacco etch virus (TEV).
[0106] Alternatively, the 5'-UTR of the modified RNA molecule may comprise a 5'-UTR sequence from or derived from a plant RNA transcript. A preferred 5'-UTR is a 5'-UTR from a highly expressed gene. A preferred 5'-UTR is from or derived from a plant ubiquitin gene, preferably the maize ubiquitin gene (ZmUbi). The 5'-UTR of ZmUbi has previously been used in the art to control Cas9 expression in maize (Zhang et al. (2016) supra). Thus, a preferred 5'-UTR of the modified RNA molecule may comprise a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO: 46 (ZmUbi).
[0107] The RNA molecule comprising a 5'-UTR as defined herein is preferably a modified RNA molecule as defined herein. However, the present specification also provides a method for producing a plant cell comprising an RNA molecule, the method comprising: i) providing a plant cell; and ii) introducing a modified RNA molecule into a plant cell, wherein the RNA molecule comprises a 5'-UTR, a coding sequence, and a 3'-UTR; The 5'-UTR comprises a 5'-UTR of a positive-strand RNA virus or a 5'-UTR of a plant RNA transcript; and The 3'-UTR comprises a poly(A) tail. Also provided is a method comprising: wherein the RNA molecule has increased protein expression compared to an identical RNA molecule (control) that does not contain the 5'-UTR. The increased protein expression can be determined by: i) the total increase (i.e., the sum of the amount of protein expressed after a transfection event over a defined period of time); ii) an increase in protein levels (i.e., the level of the protein at a particular time point after transfection, preferably the peak level of the protein after transfection); and iii) Increased protein levels over time (i.e., the longer the mRNA is present, the longer the protein can be detected in the cell). Preferably, in this embodiment, the 5'-UTR of an identical (control) RNA molecule comprises the same length and number of nucleotides, but the nucleotide sequence is scrambled.
[0108] 3'-UTR The modified RNA molecule comprises a 3'-UTR. Regulatory regions within the 3'-UTR are known to affect mRNA polyadenylation, translation efficiency, localization, and stability. The modified RNA molecule is preferably designed to have optimal translation efficiency and stability. As a non-limiting example, it is known that the 3'-UTR may contain binding sites for regulatory proteins, as well as microRNAs (miRNAs). Preferably, the 3'-UTR of the modified RNA molecule lacks a binding site for a plant microRNA, and preferably, the modified RNA molecule lacks a binding site for a plant microRNA known to be expressed in the provided plant cell. Thus, preferably, the modified RNA molecule lacks a plant microRNA response element (MRE).
[0109] Alternatively or additionally, the 3'-UTR of the modified RNA molecule does not contain a silencer sequence. A silencer RNA sequence is a sequence that can bind to a repressor that inhibits protein translation. The 3'-UTR of the modified RNA molecule preferably does not contain a silencer sequence that can bind to a repressor known to be expressed in the provided cell.
[0110] Alternatively, or in addition, the 3'-UTR of the modified RNA molecule does not contain an AU-rich element (ARE), which may affect the stability of the modified RNA molecule. The modified RNA molecule preferably does not contain at least one of Class I, Class II, Class III, Group 1, Group 2, Group 3, Group 4, and Group 5 AREs.
[0111] Alternatively, or in addition, the 3'-UTR of the modified RNA molecule may comprise a 3'-UTR sequence from or derived from a plant RNA transcript. Preferred 3'-UTRs are 3'-UTRs from highly expressed genes. Preferred 3'-UTRs are from or derived from plant ubiquitin genes, preferably the maize ubiquitin gene (ZmUbi).
[0112] The length of the 3'-UTR can be short (1-500 bp), medium (501-2,000 bp), or long (>2,000 bp) (Srivastava AK et al., Trends Plant Sci. 2018, 23(3):248-259). Generally, shorter 3'-UTRs are more stable than longer 3'-UTRs. Therefore, preferably, the modified RNA molecule has a short 3'-UTR. The 3'-UTR can have a length of about 1-5 nt, about 1-10 nt, about 1-50 nt, about 1-100 nt, about 1-200 nt, or about 1-300 nt. Preferably, the 3'-UTR comprises or consists of the nucleotide sequence ACCCAGCTT.
[0113] The 3'-UTR may be designed to contain secondary structures, such as stem-loop structures, that further aid in the stability of the modified RNA molecule.
[0114] The 3'-UTR preferably comprises a poly(A) tail. The poly(A) tail can protect the modified RNA molecule from degradation in the cytoplasm. The poly(A) tail comprises a stretch of adenine bases. Poly(A) tails are well known to those skilled in the art, and those skilled in the art will readily understand that such a stretch can have various lengths. Optionally, the poly(A) tail is herein referred to as A. nwhere n is about 10 to about 1000, preferably about 20 to about 800, about 40 to about 700, about 60 to about 600, about 80 to about 500, or about 90 to about 400, or n is about 100 to about 300. Preferably, n is about 50, 75, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, or n is about 500. Preferably, n is about 200.
[0115] 5'-cap Optionally, modified RNA molecules for use in the methods described herein include a 5'-cap. The 5'-cap may further contribute to the stability of the modified RNA molecule. 5'-caps are well known to those skilled in the art. They comprise a guanine nucleotide linked to RNA via a rare 5'-5' triphosphate linkage. The guanosine is methylated at the 7th position. This 5'-cap is also referred to as a 7-methylguanylate cap, or m7G. Optionally, the 5'-cap may include additional modifications, such as a methylated 2'-hydroxy group on the first ribose sugar or methylated 2'-hydroxy groups on the first two ribose sugars. Optionally, the modified RNA molecule may include a 5'-trimethylguanosine cap, a 5'-monomethylphosphate cap, or the modified RNA molecule is optionally capped with NAD+, NADH, or 3'-dephosphorylated coenzyme A. Optionally, the modified RNA molecule does not include a 5'-end (cap) modification.
[0116] Code Sequence The modified RNA molecule comprises a sequence encoding a protein of interest. Preferably, the sequence encoding the protein of interest is codon-optimized for expression in a plant cell. Optionally, the protein of interest is a protein encoded by a transgene or an endogenous gene for the plant cell of the methods provided herein. Optionally, the protein of interest is derived from a protein encoded by an endogenous gene having one or more mutations, which may result in increased activity and / or gain of function. The modification of the RNA molecule provides increased expression of the protein of interest compared to a control. The control is preferably an identical RNA molecule but containing an unmodified uridine. Preferably, the control RNA molecule is introduced into the plant cell using the same experimental conditions as those for the introduction of the modified RNA molecule. Preferably, the control RNA molecule and the modified RNA molecule can be introduced into the same plant cell under the same conditions.
[0117] The present invention is not limited to a particular protein of interest. The protein of interest can be an enzyme, a reporter protein, a hormone, a morphogenic polypeptide, or a site-specific nuclease. The protein of interest can be an enzyme, for example, selected from the group consisting of carbohydrases (including cellulases, amylases, pectinases, and lactases), proteases, lipases, phytases, laccases, polymerases, and nucleases. The protein of interest can also be a reporter protein, for example, but not limited to, green FP ((e)GFP), blue FP (BFP), cyan FP (CFP), yellow FP (YFP), orange FP (OFP), and red FP (RFP). The protein of interest can also be a hormone, preferably a plant hormone, for example, but not limited to, an auxin or a cytokinin. The protein of interest can also be a morphogenic polypeptide, preferably a morphogenic polypeptide described herein. Preferably, the protein of interest is a site-specific nuclease.
[0118] Preferably, the site-specific nuclease is a programmable nuclease or "site-specific nuclease," such as, but not limited to, transcription activator-like endonuclease (TALEN), zinc finger nuclease (ZFN), meganuclease, clustered regularly interspaced short palindromic repeats (CRISPR) nuclease, and Argonaute. The site-specific nuclease may preferably have endonuclease activity capable of introducing double-strand breaks into double-stranded DNA, or may be modified to exhibit reduced endonuclease activity, for example, to produce a nickase that can preferably introduce single-strand breaks into double-stranded DNA, or to eliminate nuclease activity and produce an inactive nuclease. Preferably, the coding sequence encodes a TALEN, CRISPR-nuclease, WOX5, or PLT1 protein.
[0119] The coding sequence of the modified RNA molecule defined herein can encode a protein of interest, wherein the protein of interest is a site-specific nuclease. A site-specific (endo)nuclease is herein understood as a protein (optionally when complexed with a nucleic acid) that modifies, for example, cleaves, a (double-stranded) nucleic acid molecule at a target sequence.
[0120] The protein of interest, preferably a site-specific nuclease, may contain a nuclear localization signal (NLS) to direct the expressed protein to the nucleus of the plant cell. The NLS can be located at the C-terminus and / or N-terminus of the protein of interest. Any known nuclear localization signal is suitable for use in the present invention. Preferred nuclear localization signals include, but are not limited to, the SV40 large T antigen NLS: MEDPTMAPKKKRKV (SEQ ID NO: 81), the monopartite NLS: PKKKRKV (SEQ ID NO: 82), and the nucleoplasmin NLS: KRPAATKKAGQAKKKK (SEQ ID NO: 83). The protein of interest may contain two or more nuclear localization signals, for example, one or more at the N-terminus and one or more at the C-terminus. The protein of interest may contain two or more nuclear localization signals, for example, the C-terminal NLS may be different from the N-terminal NLS.
[0121] Preferably, targeted modification of the plant genome is achieved by introducing expression of a site-specific nuclease in a plant cell, optionally in combination with the presence of a guide, preferably a guide RNA. The targeted modification is preferably a mutation of one or more nucleotides, such as an insertion or deletion (indel). Thus, the produced plant cell containing the modified RNA molecule may contain the targeted genome modification. The targeted genome modification can be transferred to one or more cells, for example, by cell division. Thus, the present specification encompasses a plant cell or its progeny that contains the targeted genome modification but does not contain the modified RNA molecule or the optional guide. The cell containing the targeted genome modification can develop into a multicellular tissue containing the genome modification. Thus, the method of the present invention may further comprise a step of generating or producing a multicellular tissue, preferably a plant, having the targeted genome modification, wherein preferably the multicellular tissue does not contain the modified RNA molecule.
[0122] Developing a plant from a plant cell containing a targeted genome modification can be accomplished using any conventional method known in the art. For example, and not by way of limitation, the cell containing the targeted genome modification can be regenerated and developed into pollen and / or egg cells, which can then be self-pollinated or, preferably, pollinated with a second plant containing cells with the same or a different targeted genome modification.
[0123] Alternatively, the meristematic cells having the targeted genome modification can be developed into plants, for example, by regeneration after top removal. As a non-limiting example, one or more plant cells, preferably one or more cells of the cotyledons, can be transfected with a modified RNA molecule, followed by preferably top removal of the shoot apical meristem. Thereafter, newly formed meristematic cells may contain the targeted genome modification, and these newly formed meristematic cells can be regenerated into plants containing cells having the targeted genome modification. Alternatively, the plant can first be top removed, then transfected, and then the newly formed meristematic cells can be genome-modified. The newly formed meristematic cells can be regenerated into plants containing cells having the targeted genome modification.
[0124] TALEN The protein of interest is preferably a transcription activator-like effector nuclease (TALEN). Thus, the modified RNA molecule preferably comprises a coding sequence, which encodes a TALEN.
[0125] TALENs are well known to those skilled in the art and are constructed by fusing a TAL effector DNA-binding domain (TALE) to an effector domain, preferably a (non-specific) DNA-cleavage domain such as the FokI cleavage domain. TALENs are known in the art to be effective in plant cells. As used herein, the term "Transcriptional Activator-Like Effector," "TALE," or "TAL effector DNA-binding domain" refers to a protein containing a DNA-binding domain, which comprises a highly conserved 33-34 amino acid sequence containing a highly variable two-amino acid motif (Repeat Variable Diresidue, RVD). RVD motifs are known to determine binding specificity for nucleic acid sequences and can be engineered to specifically bind to desired DNA sequences according to methods well known to those skilled in the art (see, e.g., WO 2010 / 079430, WO 2011 / 072246, and WO 2015027134, the entire contents of each of which are incorporated herein by reference). The simple relationship between amino acid sequence and DNA recognition has made it possible to engineer specific DNA-binding domains by selecting combinations of repeat segments containing appropriate RVDs.
[0126] A preferred conserved 34 amino acid sequence of a TALE has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with any of SEQ ID NOs: 47-50. Positions 12 and 13 constitute a two-amino acid motif (RVD), which can be modified to achieve sequence-specific DNA binding. Preferred RVDs are selected from the group consisting of two amino acid combinations: HD, NG, NI, NN, NS, N-, HG, H-, IG, NK, HA, ND, HI, HN, NA, SN, and YG, as described, for example, in WO2011072246, which is incorporated herein by reference. Preferably, the RVD is "NI" to target adenine, "NG" to target thymine, "NN" to target guanine, and / or "HD" to target cytosine. Preferably, when targeting adenine, the TALE has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 47. Preferably, when targeting thymine, the TALE has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 48. Preferably, when targeting guanine, the TALE has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 49. Preferably, when targeting cytosine, the TALE has at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 50.
[0127] The coding sequence of the modified RNA molecule can optionally encode several TALEs linked to a nuclease domain, thereby forming a TALEN. Individual TALE domains of the TALEN preferably have at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 51, whereby positions 34-39 (comprising the RVD) can optionally be modified to achieve sequence-specific DNA binding. Preferably, nucleotide positions 34-39 of SEQ ID NO: 51 can encode the amino acid residues "NI" to target adenine, "NG" to target thymine, "NN" to target guanine, or "HD" to target cytosine.
[0128] The coding sequence of the modified RNA molecule preferably encodes a TALEN, and the TALEN comprises several TALEs, where each TALE can target a single nucleotide.Therefore, the number of TALEs depends on the length of the targeting nucleotide sequence.As a non-limiting example, when targeting a 17 or 18 bp nucleotide sequence, the coded TALEN preferably comprises 17 or 18 TALEs, respectively.
[0129] As used herein, the term "transcription activator-like element nuclease" or "TALEN" refers to an engineered nuclease that contains a transcription activator-like effector DNA-binding domain (TALE) linked to a DNA cleavage domain, e.g., a Fokl domain. Several modular assembly schemes for generating engineered TALE constructs have been previously reported and are well known to those skilled in the art.
[0130] Transcription activator-like effector nucleases (TALENs) are fusions of a restriction endonuclease cleavage domain, preferably a FokI domain, with a DNA-binding transcription activator-like effector (TALE) repeat sequence. Other useful endonuclease domains include, for example, Hhal, Hindlll, Notl, BbvCl, EcoRl, Bgl II, and AlwI. Preferably, the cleavage domain is a FokI domain.
[0131] TALENs can be engineered to reduce off-target cleavage activity and thereby specifically bind to target DNA sequences, and can be used to cleave target DNA sequences in, for example, a (plant) genome in vitro or in vivo. Such engineered TALENs can be used to edit genomes in vivo or in vitro, for example, for the purpose of creating indels, gene knockout or knock-in by inducing DNA breaks at target genomic sites, targeted gene knockout by non-homologous end joining (NHEJ), or targeted genomic sequence replacement by homology-directed repair (HDR) using an exogenous DNA template.
[0132] TALENs can be designed to actually cleave any desired target DNA sequence, including naturally occurring and synthetic sequences. A preferred TALEN is a fusion of a Fokl restriction endonuclease cleavage domain with a DNA-binding TALE repeat array. These arrays contain multiple 34-amino acid TALE repeats, each of which recognizes a single nucleotide using a repeat variable dinucleotide (RVD), preferably amino acids 12 and 13. Examples of RVDs that allow recognition of each of the four DNA base pairs are known, allowing the construction of arrays of TALE repeats that can bind to virtually any DNA sequence.
[0133] A TALE may be linked to the catalytic domain of Fokl, also annotated herein as a "Fokl domain." The Fokl domain acts as a dimer and is therefore only active upon dimerization (e.g., forming a homodimer or heterodimer). Optionally, TALENs can be engineered to be only active as heterodimers by the use of forced heterodimeric Fokl mutants (e.g., Cade, L. et al. 2012, Highly efficient generation of heritable zebrafish gene mutations using homo- and heterodimeric TALENs. Nucleic Acids Res 40, 8001-8010). In this configuration, two different TALEN monomers are designed to each bind to one target half-site and cleave within the DNA spacer sequence between the two half-sites.
[0134] Therefore, preferably, the method includes step ii) of introducing a first and a second modified RNA molecule as defined herein into a plant cell, i.e., introducing a combination of modified RNA molecules. The first modified RNA molecule comprises a first coding sequence encoding a first portion of a TALEN, and the second modified RNA molecule comprises a second coding sequence encoding a second portion of the TALEN, wherein the first and second portions of the TALEN form a functional TALEN, i.e., a TALEN capable of introducing a double-stranded break in a target sequence. The first portion of the TALEN preferably comprises a first TALE linked to a FokI monomer, and the second portion of the TALEN preferably comprises a second TALE linked to a FokI monomer. Two FokI monomers can form a homodimer or heterodimer capable of cleaving double-stranded DNA. Thus, the first and second TALEs can hybridize to or near a sequence of interest, or to or near the complement of the sequence of interest, preferably a sequence of interest as defined herein. Preferably, the first and second TALEs hybridize to complementary DNA strands in an orientation and at a spacing that allows the FokI domains to dimerize and cleave the DNA.
[0135] In cells, such as plant cells, the double-strand breaks induced by TALENs result in site-specific mutagenesis events, such as the generation of indels, targeted gene knockout by non-homologous end joining (NHEJ), or targeted genome sequence replacement by homology-directed repair (HDR) using foreign DNA templates. To date, TALENs have been successfully used to manipulate the genomes of various organisms, including plant cells (e.g., Tzfira, T. et al. (2012) Genome modifications in plant cells by custom-made restriction enzymes. Plant Biotechnol. J. 10, 373-389; Curtin, SJ et al. 2012) Genome engineering of crops with designer nucleases. Plant Genome 5, 42-50, as reviewed).
[0136] CRISPR nuclease The modified RNA molecule may comprise a coding sequence encoding a CRISPR nuclease, preferably a CRISPR nuclease as defined herein.Preferably, the nuclease is a type II CRISPR nuclease, such as Cas9 (for example, the protein of SEQ ID NO:74 encoded by SEQ ID NO:75, or the protein of SEQ ID NO:76), or a type V CRISPR nuclease, such as Cpf1 (for example, the protein of SEQ ID NO:77 encoded by SEQ ID NO:78) or Mad7 (for example, the protein of SEQ ID NO:79 or 80), or a protein derived therefrom, preferably having at least about 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with said protein over its entire length.Preferably, the site-specific nuclease is a type II CRISPR nuclease, preferably a Cas9 nuclease.
[0137] Those skilled in the art know how to prepare modified RNA molecules as defined herein, including sequences encoding site-specific nucleases, for example, sequences encoding CRISPR-nucleases. Numerous reports on the design and use of CRISPR-nucleases are available in the prior art. For example, see the review by Haeussler et al. on the design of guide RNAs and their use in combination with CAS-proteins (originally obtained from Streptococcus pyogenes) (J Genet Genomics. (2016) 43 (5): 239-50. Doi: 10.1016 / j.jgg.2016.04.008.), or the review by Lee et al. (Plant Biotechnology Journal (2016) 14 (2) 448-462). Optionally, the site-specific nuclease is a CRISPR-nuclease, which is either a nickase or an (endo)nuclease.
[0138] The site-specific nuclease expressed by the modified RNA comprises or consists of the entire Type II or Type V CRISPR nuclease, or a variant or functional fragment thereof. Optionally, such a fragment binds to the guide RNA but may lack, for example, one or more residues required for nuclease activity. Preferably, the site-specific nuclease is a Cas9 protein.The Cas9 protein may be derived from the following bacteria: Streptococcus pyogenes (SpCas9; NCBI Reference Sequence NC_017053.1; UniProtKB - Q99ZW2), Geobacillus thermodenitrificans (UniProtKB - A0A178TEJ9), Corynebacterium ulcerous (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Ref: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Ref:NC_021284.1); Prevotella intermedia (NCBI Ref:NC_017861.1); Spiroplasma taiwanense (NCBI Ref:NC_021846.1); Streptococcus iniae (NCBI Ref:NC_021314.1); Belliella baltica (NCBI Ref:NC_018010.1); Psychroflexus torquisl (NCBI Ref:NC_018721.1); Streptococcus thermophilus (NCBI Ref:YP_820832.1); Listeria innocua innocua (NCBI Ref: NP_472073.1); Campylobacter jejuni (NCBI Ref: YP_002344900.1); or Neisseria meningitidis (NCBI Ref: YP_002342100.1).Cas9 variants derived from these that have inactive HNH or RuvC domains homologous to SpCas9, such as SpCas9_D10A or SpCas9_H840A, or Cas9s with equivalent substitutions at positions corresponding to D10 or H840 in the SpCas9 protein that result in a nickase, are included.
[0139] The site-specific nuclease may be or be derived from Cpf1, such as Cpf1 from Acidaminococcus sp. UniProtKB-U2UMQ6. The mutant may be a Cpf1-nickase with an inactivated RuvC or NUC domain, where the RuvC or NUC domain no longer has nuclease activity. Those skilled in the art are well aware of techniques available in the art, such as site-directed mutagenesis, PCR-mediated mutagenesis, and total gene synthesis, that enable inactivated nucleases, such as inactivated RuvC or NUC domains. An example of a Cpf1 nickase with an inactive NUC domain is Cpf1 R1226A (Gao et al. Cell Research (2016) 26:901-913, Yamano et al. Cell (2016) 165(4):949-962). In this mutant, there is an arginine to alanine conversion (R1226A) in the NUC domain, which inactivates the NUC domain.
[0140] The site-specific nuclease may be or be derived from CRISPR-CasΦ, a nuclease approximately half the size of Cas9. CRISPR-CasΦ uses a single crRNA to target and cleave nucleic acids, as described, for example, in Pausch et al. (CRISPR-CasΦ from huge phages is a hypercompact genome editor, Science (2020); 369(6501):333-337).
[0141] Active, partially inactive or inactive site-specific nuclease, preferably active, partially inactive or inactive CRISPR-nuclease complex, can guide fused functional domain to specific site in DNA, as determined by guide RNA.Therefore, site-specific nuclease can be fused with functional domain.Optionally, this functional domain is endonuclease domain or domain for epigenetic modification, for example, histone modification domain.
[0142] In one embodiment, the modified mRNA molecule disclosed herein preferably comprises a coding sequence encoding an inactive CAS protein (e.g., dCas9, dCpf1) fused to a restriction enzyme such as, but not limited to, Fok1 or Clo51, as described in WO 2014 / 144288, WO 2016 / 205554, Tsai et al. Nat Biotechnol. 2014 Jun;32(6):569-576, or Cheng et al. Biotechnol J. 2022 Jul;17(7):e2100571, all of which are incorporated herein by reference. Preferably, the fusion protein has the sequence of any one of SEQ ID NOs: 84-86, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 805, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 84-86.
[0143] In one embodiment, the modified mRNA molecule disclosed herein comprises a coding sequence encoding an inactive CAS protein (e.g., dCas9, dCpf1) fused with a domain for epigenetic modification. The domain for epigenetic modification may be selected from the group consisting of deaminase, methyltransferase, demethylase, deacetylase, methylase, deacetylase, deoxygenase, glycosylase, and acetylase (Cano-Rodriguez et al., Curr Genet Med Rep (2016) 4:170-179). The methyltransferase may be selected from the group consisting of G9a, Suv39h1, DNMT3, PRDM9, and Dot1L. The demethylase may be LSD1. The deacetylase may be SIRT6 or SIRT3.
[0144] Optionally, the functional domain is a deaminase, or a functional fragment thereof, selected from the group consisting of apolipoprotein B mRNA editing complex (APOBEC) family deaminases, activation-induced cytosine deaminase (AID), ACF1 / ASE deaminase, adenine deaminase, and ADAT family deaminases. Alternatively or additionally, the deaminase or functional fragment thereof may be ADAR1 or ADAR2, or a variant thereof. The apolipoprotein B mRNA editing complex (APOBEC) family of cytosine deaminase enzymes includes 11 proteins that play a role in initiating mutagenesis in a regulated and beneficial manner. Preferably, the APOBEC deaminase is selected from the group consisting of APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminase. Preferably, the cytosine deaminase of the APOBEC family is activation-induced cytosine (or cytidine) deaminase (AID) or apolipoprotein B editing complex 3 (APOBEC3). Preferably, the deaminase domain fused to the CRISPR-nuclease is a deaminase of the APOBEC1 family.
[0145] Another exemplary suitable type of deaminase domain that can be fused to a site-specific nuclease, preferably a CRISPR nuclease, is an adenine or adenosine deaminase, such as an adenine deaminase of the ADAT family. Furthermore, the adenine deaminase may preferably be TadA or a variant thereof, as described in Gaudelli et al., 2017 (Gaudelli et al. 2017 Nature 551:464-471). Furthermore, the nuclease, preferably a CRISPR nuclease, may be fused to an adenine deaminase domain, such as an adenine deaminase domain derived from ADAR1 or ADAR2. The deaminase domain of the present invention may comprise or consist of the entire catalytically active deaminase protein or a fragment thereof. Preferably, the deaminase domain has deaminase activity. Optionally, the nuclease element, preferably a CRISPR-nuclease, is further fused to a UDG inhibitor (UGI) domain.
[0146] The site-specific nuclease, preferably a CRISPR nuclease, may be fused to a reverse transcriptase, for example, as described in Anzalone AV, et al. (Nature (2019), 576, pages 149-157), for use in prime editing.
[0147] The site-specific nuclease encoded by the modified RNA molecule may be an Argonaute protein. Argonaute (Ago) proteins bind to small RNA or DNA guides and confer base-pairing specificity for recognizing and cleaving complementary nucleic acid targets (Kaya et al. PNAS 2016 Apr 12;113(15):4057-4062). Argonaute proteins can cleave DNA in a process known as DNA interference, as described, for example, in Kuzmenko et al. (Nature (2020), 587, 632-637). Optionally, the Argonaute protein is complexed or fused with one or more functional domains or proteins, preferably at least one of a helicase and a topoisomerase domain, preferably for unwinding genomic DNA.
[0148] Preferably, in step ii) of the method of the present invention, the modified RNA molecule can be introduced in combination with a guide, preferably a guide RNA. The guide can be introduced simultaneously with, before, or after the introduction of the modified RNA molecule. Preferably, the guide is introduced after the introduction of the modified RNA molecule. Preferably, the site-specific nuclease is expressed from the modified RNA molecule before the introduction of the guide. Preferably, step ii) comprises introducing the modified RNA molecule and the guide into the plant cell, wherein the guide is introduced simultaneously with, or about 1, 2, 3, 4, or about 5 days after, the introduction of the modified RNA molecule.
[0149] The guide directs the complex to a defined target site in the (double-stranded) nucleic acid molecule, also called a protospacer sequence. The guide contains a sequence for targeting the site-specific nuclease complex to the protospacer sequence, which is preferably near, at, or within the sequence of interest in the genome of the plant cell.
[0150] When the site-specific nuclease forms a CRISPR-endonuclease complex, the guide may be a guide RNA that is a single guide (sg) RNA molecule, or a combination of crRNA and tracrRNA (e.g., in the case of Cas9) as separate molecules, or only a crRNA molecule (e.g., in the case of Cpfl and CasΦ). Optionally, the guide RNA is a single guide (sg) RNA (e.g., in the case of Cas9) or only a crRNA (e.g., in the case of Cpfl and CasΦ).
[0151] The guide, preferably the guide RNA, used in the methods of the present invention may comprise a sequence that can hybridize to or near a sequence of interest, preferably a sequence of interest as defined herein. The guide, preferably the guide RNA, preferably comprises a nucleotide sequence that is perfectly complementary to a sequence in the sequence of interest, i.e., the sequence of interest comprises a protospacer sequence. Alternatively or additionally, the guide, preferably the guide RNA, for use in the present invention may comprise a sequence that can hybridize to or near the complement of the sequence of interest.
[0152] Optionally, the modified RNA molecules defined herein may comprise additional functional elements. Preferably, said elements may be located in at least one of the following: - after the (optional) 5'-cap of the modified RNA molecule; - before the 5'-UTR; - between the 5'-UTR and the coding sequence; - between the coding sequence and the 3'-UTR; - within the 3'-UTR and before the poly(A) tail; and - after the 3'-UTR (and after the poly(A) tail).
[0153] As a non-limiting example, the additional functional element may be a guide, preferably a guide RNA. The guide may or may not contain a modified uridine as defined herein. The modified RNA molecule may include a 5'-UTR, a coding sequence, and a 3'-UTR, preceding or following the guide sequence. Optionally, the modified RNA molecule includes a cleavable spacer sequence, such as, but not limited to, a tRNA or tRNA-like structure. As a non-limiting example, the cleavable sequence may be located between the additional functional element, preferably the guide RNA, and the 5'-UTR, coding sequence, and / or 3'-UTR of the modified RNA molecule.
[0154] Additionally or alternatively, the additional functional element may be a mobile element, which allows intercellular translocation of the modified RNA molecule from one plant cell to another. Preferably, the mobile element is a (plant) tRNA. Optionally, the modified RNA molecule comprises an element and / or an editing RNA as described in WO2022219175, which is incorporated herein by reference.
[0155] Optionally, the modified RNA molecule encodes a site-specific nuclease, wherein the site-specific nuclease is TALEN, ZFP or ZFN, or a meganuclease such as I-SceI, I-CreI or I-DmoI. Optionally, the site-specific nuclease is active as a dimer, and can induce, for example, DSB. The expressed site-specific nuclease can be designed to target a specific location in the genome of a plant cell to achieve targeted gene modification. The target site is preferably located near, at, or within the sequence of interest in the genome of a plant cell.
[0156] Morphogenetic polypeptides The coding sequence of the modified RNA molecule may encode a morphogenic polypeptide. As used herein, the terms "regeneration factor," "morphogenic polypeptide," or "morphogenetic polypeptide" refer to a polypeptide that, when ectopically expressed, stimulates the formation of somatically derived structures capable of producing plants. More precisely, ectopic expression of a morphogenic polypeptide stimulates the de novo formation of organogenic structures, such as somatic embryos or shoot meristems, thereby producing plants. This stimulated de novo formation occurs either in the cell in which the morphogenic polypeptide is expressed or in adjacent cells. A morphogenic polypeptide can be a transcription factor that controls the expression of other genes or a polypeptide that affects hormone levels in plant tissues, either of which can stimulate morphogenetic changes.
[0157] Preferably, the encoded morphogenetic polypeptide is at least one of a WUS / WOX homeobox polypeptide, a PLT (PLETHORA) protein, a polypeptide containing two AP-2 DNA binding domains, and WIND1. Preferably, the morphogenetic polypeptide encoded by the modified RNA molecule is at least one of a WUS / WOX homeobox polypeptide and a PLT protein. Preferably, the morphogenetic polypeptide encoded by the coding sequence of the modified RNA molecule is at least one of a WOX5 homeobox polypeptide and a PLT1 protein.
[0158] The WUS / WOX homeobox polypeptide encoded by the modified RNA molecule is preferably selected from the group consisting of WUS1, WUS2, WUS3, WOX2A, WOX4, WOX5, WOX5A, or WOX9 polypeptides (see, e.g., U.S. Pat. Nos. 7,348,468 and 7,256,322 and U.S. Patent Application Publication Nos. 2017 / 0121722 and 2007 / 0271628, which are incorporated by reference in their entireties, and van der Graaff et al., 2009, Genome Biology 10:248). The functional WUS / WOX homeobox polypeptide encoded by the modified RNA molecule can be obtained from or derived from any plant. Functional WUS / WOX polypeptides comprise a homeobox DNA-binding domain, a WUS box, and an EAR repressor domain, and are described, for example, in Table 1 of WO2020214986, in particular SEQ ID NOs: 246-310 of WO2020214986, which sequences are incorporated herein by reference.
[0159] A preferred WUS / WOX homeobox polypeptide encoded by the modified RNA molecule is WOX5. The amino acid sequence of the WOX5 protein preferably has at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 87. SEQ ID NO: 87 is the Arabidopsis thaliana WOX5 protein. In one embodiment, the WOX5 amino acid sequence is or is derived from AT3G11260, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT3G11260 or a homolog thereof. In one embodiment, the WOX5 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 88. The nucleotide sequence encoding the WOX5 protein may be or be derived from the gene AT3G11260, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with AT3G11260 or a homolog thereof. The identity percentage can be determined over the entire length of the genomic sequence. Alternatively, the identity percentage can be determined over the entire length of the coding sequence of the gene.
[0160] The PLT protein is preferably selected from the group consisting of PLT1, PLT2, PLT3, PLT4, PLT5, and PLT7. Preferably, the PLT protein is at least one of PLT1, PLT4, and PLT5. Preferably, the PLT protein is PLT1. The amino acid sequence of the PLT1 protein may have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 89. SEQ ID NO: 89 is the Arabidopsis thaliana PLT1 protein. In one embodiment, the PLT1 amino acid sequence is or is derived from AT3G20840, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to AT3G20840 or a homolog thereof. In one embodiment, the PLT1 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity to SEQ ID NO:90. The nucleotide sequence encoding the PLT1 protein may be or may be derived from the gene AT3G20840, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with AT3G20840 or a homolog thereof. The percentage identity may be determined over the entire length of the genomic sequence. Alternatively, the percentage identity may be determined over the entire length of the coding sequence of the gene.
[0161] The amino acid sequence of the PLT2 protein can have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 91. SEQ ID NO: 91 is the Arabidopsis thaliana PLT2 protein. In one embodiment, the PLT2 amino acid sequence is or is derived from AT1G51190, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT1G51190 or a homolog thereof. In one embodiment, the PLT2 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 92. The nucleotide sequence encoding the PLT2 protein may be or be derived from the gene AT1G51190, its homologs, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with AT1G51190 or its homologs. The identity percentage can be determined over the entire length of the genomic sequence. Alternatively, the identity percentage can be determined over the entire length of the coding sequence of the gene.
[0162] The amino acid sequence of the PLT3 protein can have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 93. SEQ ID NO: 93 is the Arabidopsis thaliana PLT3 protein. In one embodiment, the PLT3 amino acid sequence is or is derived from AT5G10510, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT5G10510 or a homolog thereof. In one embodiment, the PLT3 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 94. The nucleotide sequence encoding the PLT3 protein may be or be derived from the gene AT5G10510, its homologs, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with AT5G10510 or its homologs. The identity percentage can be determined over the entire length of the genomic sequence. Alternatively, the identity percentage can be determined over the entire length of the coding sequence of the gene.
[0163] The amino acid sequence of the PLT4 protein can have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 95. SEQ ID NO: 95 is the Arabidopsis thaliana PLT4 protein. In one embodiment, the PLT4 amino acid sequence is or is derived from AT5G17430, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT5G17430 or a homolog thereof. In one embodiment, the PLT4 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 96. The nucleotide sequence encoding the PLT4 protein may be or be derived from the gene AT5G17430, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT5G17430 or a homolog thereof. The identity percentage can be determined over the entire length of the genomic sequence. Alternatively, the identity percentage can be determined over the entire length of the coding sequence of the gene.
[0164] The amino acid sequence of the PLT5 protein may have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 97. SEQ ID NO: 97 is the Arabidopsis thaliana PLT5 protein. In one embodiment, the PLT5 amino acid sequence is or is derived from AT5G57390, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT5G57390 or a homolog thereof. In one embodiment, the PLT5 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 98. The nucleotide sequence encoding the PLT5 protein may be or be derived from the gene AT5G57390, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT5G57390 or a homolog thereof. The identity percentage can be determined over the entire length of the genomic sequence. Alternatively, the identity percentage can be determined over the entire length of the coding sequence of the gene.
[0165] The amino acid sequence of the PLT7 protein may have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 99. SEQ ID NO: 99 is the Arabidopsis thaliana PLT7 protein. In one embodiment, the PLT7 amino acid sequence is or is derived from AT5G65510, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT5G65510 or a homolog thereof. In one embodiment, the PLT7 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 100. The nucleotide sequence encoding the PLT7 protein may be or be derived from the gene AT5G65510, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT5G65510 or a homolog thereof. The identity percentage can be determined over the entire length of the genomic sequence. Alternatively, the identity percentage can be determined over the entire length of the coding sequence of the gene.
[0166] The polypeptide comprising two AP-2 DNA-binding domains is preferably a polypeptide selected from the group consisting of ODP2, BBM2, BMN2, or BMN3 polypeptides. The amino acid sequence of the ODP2 protein may have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NO: 101. In one embodiment, the OPD2 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 102. The amino acid sequence of the BBM2 protein can have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 103. In one embodiment, the BBM2 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 104.
[0167] The amino acid sequence of the WOUND INDUCED DEDIFFERENTIATION 1 (WIND1) protein can have at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 105. SEQ ID NO: 105 is the Arabidopsis thaliana WIND1 protein. In one embodiment, the WIND1 amino acid sequence is or is derived from AT1G78080, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity to AT1G78080 or a homolog thereof. In one embodiment, the WIND1 protein is encoded by a nucleotide sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with SEQ ID NO: 106. The nucleotide sequence encoding the WIND1 protein may be or be derived from the gene AT1G78080, a homolog thereof, or a sequence having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity with AT1G78080 or a homolog thereof. The identity percentage can be determined over the entire length of the genomic sequence. Alternatively, the identity percentage can be determined over the entire length of the coding sequence of the gene.
[0168] Optionally, the morphogenetic polypeptide is LEC1 (preferably any one of SEQ ID NOs: 2, 8, 10, 12, 14, 16, 18, 20, or 22 of U.S. Pat. No. 6,825,397, herein incorporated by reference, or a homolog thereof having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity thereto), SHORT a ROOT protein (preferably an SHR having SEQ ID NO: 107 or having a sequence encoded by SEQ ID NO: 108, or a homolog thereof having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity thereto), or a SCARECROW protein (preferably an SCR having SEQ ID NO: 109 or having a sequence encoded by SEQ ID NO: 110, or a homolog thereof having at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or 100% sequence identity thereto).
[0169] Optionally, the modified RNA molecule comprises two or more coding sequences, for example, separated by an internal ribosome entry site. Optionally, the modified RNA molecule comprises two, three, four, five, six, seven, eight, nine, ten, or more coding sequences. Optionally, a first coding sequence may encode a WUS / WOX homeobox polypeptide, preferably WOX5, and a second coding sequence may encode a PLETHORA polypeptide, preferably PLT1. Alternatively, the first coding sequence may encode a first portion of a TALEN as defined herein, and the second coding sequence may encode a second portion of a TALEN as defined herein.
[0170] Expression of one more morphogenic polypeptides, preferably one or more morphogenic polypeptides as defined herein, preferably leads to the de novo formation of a shoot or shoot meristem and optional regeneration into a plant. Thus, the present specification also encompasses a plant part, preferably a shoot, or a plant obtained by the methods described herein.
[0171] Combination of modified RNA molecules Optionally, in step ii) of the method of the present invention, a combination of two or more RNA molecules is introduced into the provided plant cell, wherein at least one of the RNA molecules is a modified RNA molecule. Optionally, a combination of 2, 3, 4, 5, 6, 7, 8, 9, 10 or more RNA molecules is introduced, wherein at least one of the molecules is a modified RNA molecule. Optionally, at least 25%, 50%, 75%, or about 100% of the RNA molecules introduced into the plant cell are modified RNA molecules.
[0172] A preferred combination of RNA molecules is a combination of RNA molecules encoding morphogenic polypeptides, preferably a combination of 2, 3, 4, 5, 6, 7, 8, 9, or 10 RNA molecules, where each RNA molecule encodes a different morphogenic polypeptide, preferably a morphogenic polypeptide as defined herein. A preferred combination of RNA molecules is a combination of a first and a second RNA molecule, where the first RNA molecule encodes a PLETHORA polypeptide and the second RNA molecule encodes a WUS / WOX homeobox polypeptide. Preferably, the PLETHORA polypeptide is PLT1, and / or preferably, the WUS / WOX homeobox polypeptide is WOX5. Preferably, at least one of the first and second RNA molecules is a modified RNA molecule.
[0173] Alternatively, a preferred combination is a combination of RNA molecules, in which one of the RNA molecules encodes a site-specific nuclease and one of the RNA molecules is a guide RNA. Optionally, the combination of RNA molecules includes one RNA molecule encoding a site-specific nuclease and 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 1000 or more RNA molecules encoding guide RNAs. These different guide RNAs can optionally target different genes, different positions of the same gene, or different allelic variants of the same gene. The encoded site-specific nuclease is preferably a CRISPR protein as defined herein. Preferably, at least the RNA molecule encoding the site-specific nuclease is a modified RNA molecule. Optionally, all or part of the guide RNA contains pseudouridine and / or N1-methyl-pseudouridine as defined herein.
[0174] Alternatively, preferred combination of RNA molecules is the combination of RNA molecules encoding TALEN, preferably the combination described herein.Preferred combination is the combination of first and second modified RNA molecules, in which the coding sequence of first modified RNA molecule encodes the first part of TALEN, and the coding sequence of second modified RNA molecule encodes the second part of TALEN, and the first and second parts of TALEN form functional TALEN when expressed in plant cells.Above-mentioned combination can further comprise the first and second parts of TALEN, whereby different functional TALEN can optionally target different genes, different positions of the same gene, or different allelic variants of the same gene.
[0175] Alternatively, preferred combinations are: i) one or more modified RNA molecules encoding a site-specific nuclease, preferably a TALEN; and ii) one or more modified RNA molecules encoding a morphogenic polypeptide, preferably WOX5 and / or PLT1.
[0176] Introduction into plant cells Modified RNA molecules can be introduced into plant cells by any conventional method known in the art.Non-limiting examples of suitable introduction systems include chemical-based transfection (for example, using calcium phosphate, dendrimer, cyclodextrin, polymer, liposome or nanoparticle), non-chemical-based methods (for example, electroporation, cell transformation, sonoporation, optical transfection, protoplast fusion, impalefection, heat shock and hydrodynamic introduction), and particle-based methods (for example, gene gun or magnetic-assisted transfection).
[0177] The modified RNA molecule can be introduced into cells using a carrier suitable for delivering RNA into cells.Preferred carriers are selected from the group consisting of lipoplexes, liposomes, polymersomes, polyplexes, dendrimers, inorganic nanoparticles, virosomes, and cell-penetrating peptides.Optionally, the modified RNA molecule can be introduced using particle bombardment or cell-penetrating peptides (CPPs).Preferred cell-penetrating peptides are described in Miyamoto et al.2021. Preferred CPPs are BP100-(KH)9 or dTat-Sar-EED4, where BP100-(KH)9 preferably comprises the sequence KKLFKKILKYLKHKHKHKHKHKHKHKHKH (3810 Da, SEQ ID NO: 52), and / or dTat-Sar-EED4 preferably comprises the sequence RRRQRRKKR (SEQ ID NO: 53)-(Sar)6-GWWG (SEQ ID NO: 54) (2253 Da, Sar = sarcosine linker). The molar ratio of modified RNA molecule to CPP is preferably at least about 1:1, preferably about 1:5, about 1:10, about 1:16, or about 1:20, preferably about 1:16. Optionally, modified RNA molecules can be introduced into plant pollen using particle bombardment.
[0178] The modified RNA molecule can be introduced into plant cells using an aqueous solvent, which contains PEG. Any appropriate medium can be used, and the pH of the medium is preferably 5 to 8, preferably 6 to 7.5. In addition to the modified RNA molecule, the medium may also contain polyethylene glycol. Polyethylene glycol (PEG) is a polyether compound with many applications, from industrial manufacturing to pharmaceuticals. PEG is also known as polyethylene oxide (PEO) or polyoxyethylene (POE). The structure of PEG is generally represented as H—(O—CH—CH)—OH. Preferably, the PEG used is an oligomer and / or polymer, or a mixture thereof, and preferably has a molecular weight of less than 20,000 g / mol.
[0179] The aqueous medium preferably contains 100-400 mg / ml of PEG, e.g., 150-300 mg / ml, e.g., 180-250 mg / ml. A preferred PEG is PEG 4000 Sigma-Aldrich no. 81240 (i.e., having an average Mn of 4000; Mn is the average molecular weight). Preferably, the PEG used has an Mn of about 1000-10000, e.g., between 2000-6000). Optionally, the aqueous medium containing PEG does not contain more than about 0.001%, 0.01%, 0.05%, 0.1%, 1%, 2%, 5%, 10%, or 20% (v / v) of glycerol. Preferably, the medium contains less than about 0.001%, 0.01%, 0.05%, 0.1%, 1%, 2%, 5%, 10%, or 20% (v / v) of glycerol. Preferably, the aqueous medium contains less than about 0.1%, e.g., 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.009%, 0.008%, 0.007%, 0.006%, 0.005%, 0.004%, 0.003%, 0.002%, 0.001%, 0.0009%, 0.0008%, 0.0007%, 0.0006%, 0.0005%, 0.0004%, 0.0003%, 0.0002%, or 0.0001% (v / v) of glycerol. Optionally, the aqueous medium containing the modified RNA molecules is completely free of glycerol.
[0180] Optionally, the cell cycle of the plant cells is synchronized upon introduction of the modified RNA molecule, preferably the plant cells are synchronized in the S, M, G1 and / or G2 phase of the cell cycle.
[0181] Also provided herein are aqueous solvents as defined herein that contain modified RNA molecules, as well as compositions that contain plant cells, modified RNA molecules, and aqueous solvents as defined herein.
[0182] Selection and regeneration of produced plant cells The method of the present invention may further comprise a step of selecting the produced plant cells or their progeny. Preferably, the selected plant cells contain at least one of the modified RNA molecule, the protein expressed from the modified RNA molecule, and, preferably, if the modified RNA molecule encodes a programmed endonuclease, a genomic sequence comprising the targeted genome modification. Thus, plant cells can be selected by detecting the expressed protein, detecting the modified RNA molecule, and / or detecting the targeted modification.
[0183] Such detection methods are well known to those skilled in the art.Depending on the expressed protein, modified RNA molecule or targeted genome modification, those skilled in the art can select appropriate experimental design.Such exemplary methods include but are not limited to Northern blotting, Western blotting, (quantitative) PCR, deep sequence analysis, FACS analysis, phenotype analysis, etc.
[0184] The methods provided herein may optionally include the step of regenerating the plant cell or its progeny. Optionally, the regenerated plant cell comprises a targeted modification. Optionally, the step of introducing expression of a morphogenic polypeptide as defined herein is essential to enable the plant cell to regenerate.
[0185] Further aspects In one aspect, a modified RNA molecule as described herein is provided. Preferably, the modified RNA molecule is for use in the methods described herein. The modified RNA molecule may have any length, preferably suitable for delivery to a plant cell and / or expression of an encoded protein in a plant cell. Preferably, the length of the modified RNA molecule is about 50 to 50,000 nucleotides (nt), 100 to 10,000 nt, 200 to 8,000 nt, 300 to 7,000 nt, 400 to 6,000 nt, 500 to 5,000 nt, 600 to 4,000 nt, 700 to 3,000 nt, 800 to 2,000 nt, or about 900 to 1,000 nt.
[0186] In one aspect, there is provided a combination of modified RNA molecules as described herein, preferably for use in a method as described herein.
[0187] In one aspect, a plant, plant part, or plant cell, or progeny thereof, obtained by the methods described herein is provided. Optionally, the plant, plant part, or plant cell comprises a modified RNA molecule described herein. Optionally, the plant cell is a protoplast. Optionally, the plant part is plant pollen. Optionally, at least a portion of a plant obtained by the methods of the present invention comprises a targeted genome modification. Optionally, the plant is not necessarily solely obtainable by an essentially biological process. Optionally, the plant is a transgenic plant. Optionally, the plant cell, plant part, or plant is a tomato (Solanum lycopersicon). The plant cell, plant part, or plant obtained by the methods provided herein may subsequently be propagated, for example, to obtain cell cultures, (parts of) plants, or progeny thereof.
[0188] The present invention also relates to the progeny, or descendants, of plant cells, plant parts, or plants obtained by the methods of the present invention. Optionally, the progeny or descendants do not comprise modified RNA molecules as defined herein. Optionally, the progeny or descendants comprise a genomic modification previously introduced by a method described herein.
[0189] Further provided herein are plant products obtainable from plant cells, plant parts or plants as defined herein, e.g., selected from the group consisting of fruits, leaves, plant organs, plant fats, plant oils, plant starches, and plant protein fractions, whether crushed, ground, whole, mixed with other materials, dried, frozen, etc. These products may be non-reproductive. Optionally, said plant products comprise at least one or at least a portion of a (part of) genome comprising a modified RNA molecule and / or a targeted genome modification.
[0190] In one aspect, a kit of parts is provided, preferably for use in the methods described herein. Preferably, the kit of parts comprises a modified RNA molecule as defined herein and a solution for dissolving or diluting the modified RNA molecule. The solution may be, for example, a physiological buffer, a plant or plant cell growth medium, or an aqueous solvent as described herein.
[0191] Further provided is the use of the modified RNA molecules described herein for the transient expression of a gene product in a plant cell.
[0192] Additionally, there is provided the use of the modified RNA molecules described herein for the targeted genome modification of plant cells.
[0193] Additionally, there is provided the use of the modified RNA molecules described herein for the regeneration of plant cells.
[0194] Having generally described the invention, the same will be more readily understood by reference to the following examples, which are provided by way of illustration and are not intended to be limiting of the invention. [Brief explanation of the drawings]
[0195] [Figure 1] The percentage of fluorescent cells after transfection of tomato protoplasts with either a control GFP plasmid (pKG7460), unmodified GFP mRNA (GFP), pseudouridine-modified GFP mRNA (GFP-Pu), or N1-methyl-pseudouridine-modified GFP mRNA (GFP-mPu) is shown. The y-axis shows the percentage of fluorescent cells, and the x-axis shows time. [Figure 2] The percentage of fluorescent cells 48 h after transfection of tomato protoplasts with either a control GFP plasmid (20 μg of plasmid DNA carrying a 35S::GFP cassette) or GFP mRNA carrying various uridine / N1-methyluridine ratios is shown. For controls, no mRNA was added to the transfection. The y-axis shows the percentage of fluorescent cells, and the x-axis shows the percentage of mPu modification. [Figure 3] Figure 1 shows the activity of unmodified and modified SpCas9 mRNA in tomato protoplasts. The various sample times are shown on the x-axis (hours) and the percentage of indels detected on the y-axis. [Figure 4] The GFP-positive rate of tomato protoplasts at each time point after transfection with a GFP control plasmid (7460), no mRNA (control), TEV 5′UTR-GFP mRNA (TEV-GFP), PsBMV 5′UTR-GFP mRNA (PSbMV-GFP), or ZmUbi 5′UTR-GFP-ZmUbi 3′UTR mRNA (ZmUTR-GFP) is shown. [Figure 5]Determination of the optimal mRNA:CPP molar ratio. Either TEV-WOX5 or TEV-PLT1 mRNA was incubated with increasing molar ratios of the CPP BP100-(KH)9. mRNA:CPP molar ratio: Lane 1, no CPP; Lane 2, 5-fold molar excess CPP; Lane 3, 8-fold molar excess CPP; Lane 4, 16-fold molar excess CPP; Lane 5, 40-fold molar excess CPP. [Example]
[0196] Example 1. Increased expression of mRNA containing the modified nucleotides pseudouridine or N1-methyl-pseudouridine in tomato protoplasts To investigate whether the modified RNA nucleotides pseudouridine (Pu) or N1-methyl-pseudouridine (mPu) enhance mRNA expression in plant cells, we conducted experiments using GFP and tomato protoplasts. Using in vitro transcription, we generated GFP mRNA with the 5'UTR sequence derived from Tobacco Etch Virus (TEV) (SEQ ID NO: 64, nts 1-18 are the T7 promoter, nts 19-162 are the TEV 5'-UTR, and nts 163-924 are the GFP ORF) with or without the modified nucleotides pseudouridine or N1-methyl-pseudouridine. These were then transfected into tomato protoplasts, and GFP protein levels were quantified using fluorescence flow cytometry. GFP mRNA containing either of the modified nucleotides significantly enhanced GFP expression compared to unmodified GFP mRNA (Figure 1).
[0197] construct A GFP ORF with an NLS tag fused to the 3' end was synthesized with the Tobacco Etch Virus (TEV) 5'-UTR. A T7 promoter was added for in vitro transcription, and a unique restriction site (MfeI) was added for vector linearization. A similar plasmid design was used for in vitro transcription of SpCas9 (KG11683; SEQ ID NO: 65, nt 1-18 is the T7 promoter, nt 19-162 is the TEV 5'-UTR, and SpCas9 ORF is nt 163-4302). The plasmid was linearized with MfeI and then purified using a Qiagen PCR purification kit.
[0198] mRNA synthesis One microgram of linearized plasmid vector was used per reaction for synthesis using the High Scribe T7 ARCA mRNA Kit (New England Biolabs, E2060S) according to the manufacturer's instructions. Partially modified mRNA (including UTP-triphosphate in the reaction) was synthesized using the same procedure, but with the addition of 25 nmol of pseudouridine or N1-methyl-pseudouridine (Trilink Biotechnologies N-1019 and N-1105). Fully modified mRNA (all UTP replaced with either Pu or mPu) was synthesized using the High Yield T7 ARCA mRNA Synthesis Kit (Jena Biosciences). The mRNA was then purified (Qiagen RNAeasy kit).
[0199] Guide RNA synthesis The guide RNA for gene Solyc07g043010 was identified using primer 19_03079. [ka] : italics, T7 promoter; underline, seed sequence) and was generated using the New England Biolabs EnGen sgRNA Synthesis Kit (E3322) to yield an sgRNA with SEQ ID NO: 66. A modified sgRNA with improved endonuclease resistance and the same sequence as SEQ ID NO: 66 was synthesized by Synthego (Synthego.com).
[0200] Protoplast transfection In vitro shoot cultures of tomato (Solanum lycopersicon var. Moneyberg) were maintained in high-plastic jars on MS20 medium containing 0.8% agar at 25°C and 60-70% RH under a 16 / 8-h photoperiod (2000 lux). Young leaves (1 g) were carefully sliced vertically up to the central vein to facilitate penetration with the enzyme mixture. The sliced leaves were transferred to an enzyme mixture (2% Cellulase Onozuka RS, 0.4% Macerozyme Onozuka R10 in CPW9M), and cell wall digestion proceeded overnight at 25°C in the dark. The resulting protoplasts were filtered through a 50 μm nylon sieve and harvested by centrifugation at 800 rpm for 5 min. Protoplasts were resuspended in CPW9M (Frearson, 1973) medium, and 3 mL of CPW18S (Frearson, 1973) was added to the bottom of each tube using a long-neck glass Pasteur pipette. Viable protoplasts were collected as the cell fraction at the interface between the sucrose and CPW9M medium by centrifugation at 800 rpm for 10 min. Protoplasts were counted and resuspended in MaMg (Negrutiu, 1987) medium to a final density of 10 6 The cells were resuspended at 1000 cells / mL.
[0201] 40 μg of mRNA was mixed with 250 μL (250,000 protoplasts) of the protoplast suspension, and 250 μL of PEG solution (400 g / L poly(ethylene glycol) 4000, Sigma-Aldrich #81240; 0.1 M Ca(NO3)2) was added. Transfection was carried out at room temperature for 20 minutes. As a control, a plasmid carrying a 35S::GFP cassette was also transfected. Next, 10 mL of 0.275 M Ca(NO3)2 solution was added and mixed thoroughly but gently. Protoplasts were collected by centrifugation at 800 rpm for 5 minutes and resuspended in 9 M medium at a concentration of 0.5 × 10 6 The cells were resuspended at a density of 1 / ml and transferred to 4-cm diameter Petri dishes. The number and expression level of GFP-expressing cells were quantified using an Accurri fluorescence flow cytometer by sampling (62,500 cells per measurement) at four different time points post-transfection.
[0202] result Partial incorporation of pseudouridine or N1-methyl-pseudouridine enhances mRNA expression in tomato protoplasts Tomato protoplasts were transfected with equimolar amounts of unmodified GFP mRNA or GFP mRNA incorporating pseudouridine (Pu) or N1-methyl-pseudouridine (mPu) nucleotides. The number of GFP-fluorescent cells was then measured at different time points. For unmodified GFP mRNA, the percentage of GFP-positive cells declined linearly over the course of the experiment, whereas for modified GFP mRNA, the decrease in GFP signal remained constant or showed a diminishing decrease over time. GFP fluorescence may be related to the lifespan and / or translation level of GFP mRNA in the cells.
[0203] As a control, a plasmid carrying a 35S::GFP cassette was used. Because this construct must be transcribed first, the signal is initially lower than that of the mRNA but then increases. The intensity of the GFP signal when using the control GFP plasmid is also four-fold higher than when using either of the mRNAs.
[0204] The Tobacco Etch Virus (TEV) 5' UTR (Carrington et al. 1990) stimulates GFP expression, as GFP mRNA lacking the TEV 5' UTR showed no expression in protoplasts (data not shown). Additionally, we scrambled the TEV 5' UTR sequence and fused it to GFP, producing mRNA from this construct. Transfection of this mRNA into tomato protoplasts did not result in a GFP signal (data not shown). Thus, the TEV 5' UTR contains a specific sequence that promotes translational activity. The TEV 5' UTR has previously been used to enhance mRNA expression in both plant and animal cells (Nicolaisen et al. FEBS Lett. 303(2-3):169-72, 1992; Kariko et al. 2008, supra). Here, we demonstrate that the TEV 5' UTR also promotes the translation of modified RNA molecules in plant cells.
[0205] The results shown in Figure 1 were obtained using partially modified GFP mRNA, which was generated by adding both uridine and either Pu or mPu to the synthesis reaction. According to the manufacturer's protocol, this results in up to 50% replacement of UTP with the modified nucleotide. To examine the expression of fully modified mRNA in plant cells, an alternative synthesis kit (Jena Biosciences) was used to generate GFP mRNA in which all uridines were replaced with Pu or mPu. Fully modified GFP mRNA was not expressed in tomato protoplasts.
[0206] Therefore, we investigated the effect of varying amounts of the modified nucleotide N1-methyl-pseudouridine on mRNA translation in tomato protoplasts. To this end, GFP mRNA was synthesized in vitro using both uridine (UTP) and the modified nucleotide N1-methyl-uridine (mPu) in the reaction. When equimolar amounts of UTP and mPu were added to the synthesis reaction, it can be reasonably assumed that half (50%) of the U in the mRNA is actually mPu. By adding different ratios of UTP and mPu to the reaction, mRNA modified to different degrees can be produced. mPu-modified GFP mRNA was transfected into tomato protoplasts, and the percentage of fluorescent cells was measured after 48 h. As shown in Figure 2, fully modified mRNA (100% mPu) is inactive at this time point (and at all other time points tested (2 h, 6 h, 24 h), data not shown).
[0207] High activity of SpCas9 mRNA incorporating Pu and mPu nucleotides Initial experiments demonstrated that Pu- or mPu-modified GFP mRNA showed higher expression in tomato protoplasts than unmodified mRNA. To confirm that this effect was not specific to the GFP ORF, we replaced the GFP ORF with SpCas9.
[0208] Vector KG11683 contains the TEV 5'UTR fused to the SpCas9 ORF (SEQ ID NO: 65). This vector was used to generate unmodified and partially modified SpCas9 mRNA, which was then transfected into tomato protoplasts along with a guide RNA targeting the Solyc07g043010 gene. This guide RNA has been shown to be highly efficient at generating indels in experiments using SpCas9 ribonucleoprotein (RNP). SpCas9 mRNA and guide RNA were transfected into tomato protoplasts at a 1:1 molar ratio, and the cells were sampled at multiple time points. Genomic DNA was isolated from each sample and used as a template to generate amplicons of the Solyc07g04310 target region using 19_03122 + 19_03123. Nested PCR products were generated using 19_03152 + 19_03153, and a final PCR round was performed to add barcodes to the products. The primers are shown in Table 1. These amplicons were sequenced and the percentage of reads containing indel mutations was quantified (Figure 3).
[0209] [Table 1]
[0210] The partially modified SpCas9 mRNA generated significantly more indels, especially at later time points. Introduction of Pu or mPu nucleotides increased the percentage of indels by approximately twofold after 48 hours. There were no significant differences in the types and proportions of small indels generated by the different mRNAs, and all were similar to the indels generated after RNP transfection into protoplasts.
[0211] Activity of TALEN mRNA incorporating Pu and mPu nucleotides To evaluate the efficacy of modified RNA molecules encoding TALENs, we targeted the tomato gene Solyc04g045660 using TALEN mRNA. For this purpose, we used two sequences encoding TALENs (Talen1, SEQ ID NO: 71 (nt 1-18 is the T7 promoter, nt 19-162 is the TEV 5'-UTR, nt 163-4086 is the TALEN1 ORF) and Talen2, SEQ ID NO: 72 (nt 1-18 is the T7 promoter, nt 19-162 is the TEV 5'-UTR, nt 163-4188 is the TALEN2 ORF) fused to the tobacco etch virus 5' UTR. The ORFs (which are the target sequences of the TALENs) were designed and synthesized on vectors. These vectors were used to generate unmodified and partially modified mRNAs as described above. The mRNAs for both TALENs were transfected into tomato protoplasts, and the cells were analyzed for indels in the sequence of interest using amplicon sequencing and gene-specific primers. The sequences targeted by TALEN1 and TALEN2 are shown in SEQ ID NO: 73, where nt 1-17 is targeted by TALEN1 and nt 43-60 is targeted by TALEN2.
[0212] Example 2. Potyvirus 5'UTRs promote high levels of translation Experiments were performed to compare the levels of translation conferred by different potyvirus 5'UTRs as well as the UTR of a highly expressed maize gene (ZmUbi) that has been shown to drive SpCas9 mRNA translation in wheat seedlings (Zhang et al., 2016). Plasmid constructs carrying TEV 5'UTR-GFP (SEQ ID NO: 64), PSbMV 5'UTR-GFP (SEQ ID NO: 69, where nt 1-18 are the T7 promoter, nt 19-165 are the PSbMV 5'UTR, and the GFP ORF is nt 166-927), and ZmUbi 5'UTR-GFP-ZmUbi 3'UTR (SEQ ID NO: 70, where nt 1-18 are the T7 promoter, nt 19-323 are the ZmUbi 5'UTR, the GFP ORF is nt 324-1085, and the ZmUbi 3'UTR is nt 1086-1297) were used for mRNA production as described, in this case without modified nucleotides. These mRNAs were then transfected into tomato protoplasts, and GFP expression was measured over time (Figure 4).
[0213] Both the TEV 5'UTR and PSbMV 5'UTR promoted expression in large numbers of protoplasts at early time points, with the TEV 5'UTR being slightly better at later time points. The ZmUTR showed some activity but was significantly less active at all time points. These results indicate that potyvirus 5'UTRs are optimal for driving mRNA expression in plant cells and are superior to the maize UTR from the highly constitutively expressed ubiquitin gene.
[0214] Example 3. Cell-penetrating peptide-mediated delivery of regeneration-promoting mRNA into tomato and pepper When expressed together, the WOX5 and PLT1 genes can induce plant regeneration upon transient, induced expression (WO 2019 / 211296). While such transient expression of WOX5 and PLT1 is essential for proper regeneration and is effective, this approach currently requires the generation of transgenic lines and a substantial investment in tissue culture processes. We hypothesized that in vitro synthesized modified WOX5 and PLT1 mRNAs could also be introduced into plant tissues via CPPs to stimulate regeneration. The transient mRNA expression profile would be sufficient to switch on the plant regeneration pathway. The modified mRNAs could be introduced into leaves of young seedlings grown under normal greenhouse conditions, thereby avoiding tissue culture.
[0215] Construct and mRNA synthesis Plasmids were synthesized in which the WOX5 and PLT1 ORFs were fused to the TEV 5'UTR (WOX5 (KG11681; SEQ ID NO: 67, nt 1-18 are the T7 promoter, nt 19-162 are the TEV 5'UTR, and the WOX5 ORF is nt 163-711); PLT1 (KG11685; SEQ ID NO: 68, nt 1-18 are the T7 promoter, nt 19-162 are the TEV 5'UTR, and the PLT1 ORF is nt 163-1887). These plasmids were digested with MfeI and purified using a Qiagen PCR purification kit. mRNA synthesis was performed as described above, and modified RNA nucleotides (pseudouridine or N1-methyl-pseudouridine) were also incorporated.
[0216] Cell-penetrating peptides We synthesized the cell-penetrating peptides BP100-(KH)9 and dTat-Sar-EED4 (https: / / activotec.com), which are described in Miyamoto et al. 2021. The CPPs have the following sequences: BP100-(KH)9: KKLFKKILKYLKHKHKHKHKHKHKHKHKH (3810 Da, SEQ ID NO: 52); dTat-Sar-EED4: RRRQRRKKR (SEQ ID NO: 53)-(Sar)6-GWWG (SEQ ID NO: 54) (2253 Da, Sar = sarcosine linker).
[0217] Preparation of mRNA / CPP complexes To determine the optimal mRNA:CPP molar ratio for correct complex formation (N / P = 0.5; Watanabe et al. 2021), a gel shift assay was performed. 0.5 pmol of mRNA was mixed with a 5x, 8x, 16x, or 40x molar excess of the CPP BP100-(KH)9 in 10 μl of solution. Complex formation was achieved by incubation at 25°C for 15 min. 2 μl of 6x loading buffer (without SDS) was added, and the samples were run on a 1% agarose gel. For plant treatments, an optimal mRNA:CPP molar ratio of 1:16 was used. A total of 3 μg of mRNA (857 ng WOX5 mRNA (3.3 pmol) and 2120 ng PLT1 mRNA (3.3 pmol)) was used and mixed with 106 pmol of BP100-(KH)9. Complex formation was achieved at 25°C for 15 min. Next, 10 nmol of CPP dTat-Sar-EED4 was added. The volume was then increased to 100 μl with MS10 medium and used for leaf infiltration. Similarly, when SpCas9 mRNA was also added to the mixture, a total of 3.05 μg of RNA was used (337 ng WOX5 mRNA (1.3 pmol); 842 ng PLT1 mRNA (1.3 pmol); 1820 ng SpCas9 mRNA (1.3 pmol); 50 ng sgRNA (1.3 pmol)).
[0218] Infiltration experiment Seeds of tomato (Tomato (Moneyberg) TMV+) or pepper (Capsicum annuum) cvMaor; Israel) were surface-sterilized and germinated on MS20 medium in a growth chamber. After 14 days, the leaves were infiltrated with mRNA:CPP complexes. The apical meristem was then removed, and the seedlings were maintained on hormone-free MS20 medium until regeneration occurred.
[0219] result Determining the optimal mRNA:CPP binding ratio Prior to infiltration, we determined the optimal mRNA:CPP molar ratio at which the most active complexes formed. This was defined as an N / P ratio of 0.5, which can be visualized in an agarose gel shift assay as a point at which the nucleic acid exhibits a clear mobility shift but is able to enter the gel and not remain in the loading well. To measure this for unmodified and modified WOX5 and PLT1 mRNAs, we mixed them with different molar ratios of the CPP BP100-(KH)9 and measured the resulting gel shift on an agarose gel (Figure 5).
[0220] Binding studies revealed that an mRNA:CPP molar ratio of 1:16 was closest to the optimal N / P = 0.5 ratio for subsequent experiments. Similar results were obtained with modified mRNAs, which were found to have no effect on CPP binding.
[0221] Induced regeneration of tomato and pepper seedlings Leaves of tomato or pepper seedlings are infiltrated with the (modified) mRNA:CPP complex and maintained until regeneration occurs.
[0222] Combining regeneration and genome editing The modified WOX5 and PLT1 mRNAs induce regeneration of plant tissues, and these can be combined with genome editing reagents: modified SpCas9 mRNA and guide RNA. CPPs can deliver all of these mRNAs to the same cell, where mutations can occur, leading to the regeneration of mutant plants.
[0223] WOX5, PLT1, SpCas9, and guide RNA are mixed in an equimolar ratio and complexed with CPPs. These are infiltrated into seedlings, and regenerated plants are generated. These are then genotyped for the presence of indel mutations at the target site.
[0224] 4. Bombardment of Pollen with SpCas9 mRNA and Guide RNA to Generate Plant Mutants Without a Tissue Culture Step We demonstrated that modified SpCas9 mRNA (in combination with guide RNA) can generate indel mutations at high frequencies when introduced into protoplasts. Therefore, modified Cas9 mRNA can also be used to generate mutations in other plant tissues. We tested the ability of modified SpCas9 RNA molecules in combination with guide RNA to generate mutations in pollen, which can then be used to fertilize plants. These pollen mutations can then be passed on to the next generation. This avoids the need for transgenic lines or tissue culture infrastructure. It may also be applicable to any plant species. Biolistics (bombardment) is a suitable method for introducing RNA molecules into pollen.
[0225] As a non-limiting example, tobacco pollen can be bombarded with both modified SpCas9 mRNA and (optionally modified) guide RNA to introduce indel mutations into the PDS1 gene, the frequency of indel mutations can then be measured, and the pollen can be used to fertilize tobacco flowers to introduce indel mutations into the next generation.
[0226] Guide RNA and sequencing Guide RNAs targeting the tobacco PDS1 gene were generated using the EnGen sgRNA synthesis kit (neb.com) with the following primers: 5'-TTCTAATACGACTCACTATAGGCTGCATGGAAAGATGATGAGTTTTAGAGCTAGA-3' (SEQ ID NO: 111). After pollen bombardment, genomic DNA was isolated and used to prepare a sequencing library. Primers NtPDS1 F and NtPDS1 R were used to generate an amplicon containing the PDS1 target site. These PCR products were then subjected to nested PCR (using NtPDS1 nF and NtPDS1 nR) to generate smaller amplicons, which were then used for library construction in an additional PCR using P5 and barcoded P7 Illumina primers.
[0227] [Table 2]
[0228] Pollen Bombardment Three micrograms of SpCas9 mRNA (2.1 pmol; unmodified, pseudouridine-modified, or N1-methyl-pseudouridine-modified) and PDS1 sgRNA (2.1 pmol) were conjugated to gold particles using established protocols. Tobacco (Nicotiana tabacum) SR1 pollen was germinated for 1 hour under high humidity and bombarded with coated gold particles using a BioRad instrument. The pollen was then incubated in pollen germination medium for 24 hours before being used for genomic DNA isolation.
Claims
1. 1. A method for producing a plant cell comprising a modified RNA molecule, comprising: i) providing a plant cell; and ii) introducing said modified messenger (m)RNA molecule into said plant cell. Including, the modified RNA molecule comprises a modified uridine, the modified uridine being at least one of pseudouridine and N1-methyl-pseudouridine; the modified RNA molecule comprises a 5'-UTR, a coding sequence, and a 3'-UTR; the 5'-UTR comprises a 5'-UTR of a positive-strand RNA virus or a 5'-UTR of a plant RNA transcript; the 3'-UTR comprises a poly(A) tail; The method, wherein the modified mRNA molecule has increased expression from the coding sequence compared to an identical unmodified mRNA molecule.
2. 2. The method of claim 1, wherein the 5'-UTR of the modified mRNA molecule is a potyvirus 5'-UTR, preferably the 5'-UTR is a Tobacco Etch Virus (TEV) 5'-UTR.
3. The method of claim 1 or 2, wherein the 5'-UTR of the modified mRNA molecule has at least 80% sequence identity with any one of SEQ ID NOs: 1 to 46.
4. The method of any one of claims 1 to 3, wherein the modified RNA molecule comprises a 5'-cap.
5. 5. The method of any one of claims 1 to 4, wherein not all uridines in the modified RNA molecule are modified uridines, preferably the percentage of uridines that are modified uridines is less than 95%, preferably about 10% to 90% of the total uridines are modified uridines.
6. The method of any one of claims 1 to 5, wherein the only modification of the modified RNA molecule is modification of uridine to pseudouridine and / or N1-methyl-pseudouridine.
7. The method of any one of claims 1 to 6, wherein the coding sequence of said modified mRNA molecule encodes a plant morphogenetic polypeptide.
8. The method of any one of claims 1 to 6, wherein the coding sequence of the modified mRNA molecule encodes a site-specific nuclease, preferably a CRISPR nuclease or a TALEN.
9. 9. The method of claim 8, wherein in step ii) the modified mRNA molecule is introduced in combination with a guide RNA.
10. 10. The method of any one of claims 1 to 9, wherein in step ii), the modified RNA molecule is introduced by at least one of PEG transformation, cell penetrating peptide (CPP), and biolistics.
11. and iii) selecting the produced plant cells or their progeny, wherein the plant cells are: - said modified RNA molecule; - a protein expressed from said modified RNA molecule; and - the genome sequence containing the targeted genome modification The method of any one of claims 1 to 10, wherein the method is selected to include at least one of the following:
12. 12. The method of any one of claims 1 to 11, further comprising the step of regenerating a plant from said plant cell, optionally from said selected plant cell.
13. A modified mRNA molecule comprising a modified uridine, wherein the modified uridine is at least one of pseudouridine and N1-methyl-pseudouridine; the mRNA molecule comprises a 5'-UTR, a coding sequence, and a 3'-UTR; the 5'-UTR comprises a 5'-UTR of a positive-strand RNA virus or a 5'-UTR of a plant RNA transcript; the 3'-UTR comprises a poly(A) tail; Modified mRNA molecules.
14. A plant cell comprising a modified RNA molecule according to any one of claims 1 to 8, wherein preferably said plant cell is a protoplast.
15. Use of the modified RNA molecule of claim 13 for the transient expression of a gene product in a plant cell.