Methods of selecting and producing eucalyptus plants resistant to physiological disturbance
Patent Information
- Application Number
- EP2023852123
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-11
- Filing Date
- 2023-08-09
- Publication Date
- 2025-06-18
AI Technical Summary
Current methods for selecting Eucalyptus plants resistant to physiological disturbance are time-consuming and costly, requiring extensive field trials to identify resistant genotypes, and there is a lack of genetic markers for early prediction of resistance before adult age.
The method involves determining the expression levels of specific genes in Eucalyptus plants to identify resistant or susceptible phenotypes, using gene upregulation or downregulation through transgenesis, DNA or RNA editing, and breeding to produce plants with increased resistance to physiological disturbance.
This approach allows for early identification and production of Eucalyptus plants resistant to physiological disturbance, reducing the time and cost associated with traditional selection methods and enabling the development of resistant plant varieties before visible symptoms appear.
Smart Images

Figure 00000072_0000 
Figure 000073 
Figure IMGF000066_0001
Abstract
Description
[0001] METHODS OF SELECTING AND PRODUCING EUCALYPTUS PLANTS RESISTANT TO PHYSIOLOGICAL DISTURBANCE
[0002] RELATED APPLICATIONS:
[0003] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 397,000 filed August 11, 2022, which is hereby incorporated by reference in its entirety.
[0004] SEQUENCE LISTING STATEMENT
[0005] The XML file, entitled 97275 Sequence Listing.xml, created on August 9, 2023, comprising 192,174 bytes, submitted concurrently with the filing of this application is incorporated herein by reference.
[0006] FIELD AND BACKGROUND OF THE INVENTION
[0007] The present invention, in some embodiments thereof, relates to methods of selecting and producing Eucalyptus plants resistant to physiological disturbance.
[0008] ‘Physiological disturbance’ in Eucalyptus is a significant abiotic disease in broad regions, e.g., low altitudes, of important Eucalyptus producing states in Brazil including Bahia, Espirito Santo and Maranhao as well as possibly other regions of the world where Eucalyptus is grown commercially. This phenomenon, first reported in Brazil in 2005, dramatically reduces the productivity of susceptible genotypes. Plant die-back, with progressive death often starting at the plant tip of apical shoots and leaves, cankers and epicormic sprouting along the main stem are major symptoms of the physiological disturbance. Entire stands of highly susceptible clones can collapse and die in specific sites. All plants of a susceptible genotype in large plantation stands express the symptoms equally in terms of temporal and spatial distribution. However, the actual trigger of this important abiotic disease is still unknown. A solution that has been effective for physiological disturbance phenomena is the selection of resistant genotypes, as there is large genetic variability among commercial Eucalyptus varieties such as E. urophylla x E. grandis hybrids. Phenotypes of trees growing in the field can vary from highly resistant to highly susceptible clones, with a full range of susceptibility levels among those two extremes. This suggests that resistance to physiological disturbance has a strong genetic base indicating that a potential genetic marker can be identified. Though, planting of resistant hybrid genotypes has become a solution to the problem, the identification and selection of resistant clones is an expensive and time-consuming process requiring the establishment of a large network of field trials in different sites and different seasons. Extensive testing and evaluations must be done 3 years after planting or more, when the correct Eucalyptus resistance level to the physiological disturbance can be characterized. Until now, the scientific community believed that the genetic control of Eucalyptus resistance to physiological disturbance was complexed, with several genetic loci associated to it. Thus, there is no previous genetic data or genomic markers which can indicate or predict the selection, at an early age, of resistant genotypes, before adult age (3 years +), when the negative phenotype is apparent.
[0009] SUMMARY OF THE INVENTION
[0010] According to an aspect of the invention there is provided a method of identifying a physiological disturbance phenotype in Eucalyptus, the method comprising determining in a cell or tissue of a Eucalyptus plant a level of expression of at least one gene of Table 1 and / or Table 2, wherein: a statistically significant upregulation in expression of a gene of Table 1 relative to a resistant control is indicative of a susceptible physiological disturbance phenotype; and / or a statistically significant upregulation in expression of a gene of Table 2 relative to a susceptible control is indicative of a resistant physiological disturbance phenotype.
[0011] According to an aspect of the invention there is provided a method of producing Eucalyptus plants exhibiting resistance to a physiological disturbance in Eucalyptus, the method comprising upregulating expression and / or activity in a Eucalyptus of at least one gene of Table 2 and / or downregulating expression and / or activity in a Eucalyptus of at least one gene of Table 1, thereby increasing resistance to a physiological disturbance phenotype in Eucalyptus.
[0012] According to an aspect of the invention there is provided a method of producing Eucalyptus, the method comprising: identifying a physiological disturbance as described herein; growing a plant exhibiting resistance to the physiological disturbance.
[0013] According to an aspect of the invention there is provided a method of producing a Eucalyptus processed product, the method comprising: producing Eucalyptus plants as described herein; processing the Eucalyptus plants to produce a processed product selected form the group consisting of a paper product, nanocellulose, microfibrilated cellulose, dissolving pulp, fluff pulp, timber, wood-boards, oil, lignin-derived product, plywood, particleboard, wood-based composite, Medium-density fibreboard (MDF), oil, dye and charcoal.
[0014] According to some embodiments of the invention, the determining is at the RNA level.
[0015] According to some embodiments of the invention, the determining is at the protein level. According to some embodiments of the invention, the upregulating expression is by transgenesis, DNA or RNA editing and / or breeding.
[0016] According to some embodiments of the invention, the downnregulating expression is by transgenesis, DNA or RNA editing, RNA silencing and / or breeding.
[0017] According to an aspect of the invention there is provided a kit for identifying a physiological disturbance phenotype in Eucalyptus, the kit comprising at least one reagent for identifying expression of at least one gene of Table 1 and / or Table 2, wherein the kit does not detect more than 100, 80, 50 genes.
[0018] According to some embodiments of the invention, the at least one reagent comprises a primer pair or a probe.
[0019] According to an aspect of the invention there is provided a nucleic acid construct comprising a nucleic acid sequence encoding the gene of Table 2 or a functional homolog of same and a heterologous cis-acting regulatory element for driving expression of the gene or functional homolog, such as described herein.
[0020] According to an aspect of the invention there is provided a nucleic acid agent which down- regulates expression of at least one gene of Table 1.
[0021] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the invention pertains. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting.
[0022] BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWING(S)
[0023] Some embodiments of the invention are herein described, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is stressed that the particulars shown are by way of example and for purposes of illustrative discussion of embodiments of the invention. In this regard, the description taken with the drawings makes apparent to those skilled in the art how embodiments of the invention may be practiced.
[0024] In the drawings:
[0025] FIG. 1 is a Volcano plot describing differential expression and RNA levels of genes between a set of 6 physiological disorder resistant clones (left side) compared to a set of 6 susceptible clones (right side). Differentially expressed genes in the resistant clones are depicted to the left and differentially expressed genes in the susceptible clones are depicted to the right. The selected genes are listed in Table 1 and Table 2 below.
[0026] DESCRIPTION OF SPECIFIC EMBODIMENTS OF THE INVENTION
[0027] The present invention, in some embodiments thereof, relates to kits and methods for selecting Eucalyptus resistant to physiological disturbance.
[0028] Before explaining at least one embodiment of the invention in detail, it is to be understood that the invention is not necessarily limited in its application to the details set forth in the following description or exemplified by the Examples. The invention is capable of other embodiments or of being practiced or carried out in various ways.
[0029] Physiological disturbance in Eucalyptus is an important abiotic disease considered to be the result of environmentally-induced factors in broad regions of important eucalyptus producing states such as Bahia and Espirito Santo, in Brazil.
[0030] Whilst conceiving embodiments of the invention and reducing them to practice, the present invention have tested differential gene expression (DGE) between 6 resistant and 6 susceptible phenotypes to physiological disturbance in Eucalyptus. The present inventors quantified RNA levels and detected genes that express uniquely in either the resistant or in the susceptible clones but not in both. Expression level comparisons were analyzed between the set of resistant clones and the set of susceptible clones, whereby a volcano plot was generated (Figure 1). The genes in Table 1 show significantly higher expression levels in susceptible (sensitive) clones grown in physiological disorder affected plantations. Without being bound by theory, it is suggested that the level of expression is increased due to the plant response to the abiotic stress and can be even detected before visual symptoms appear.
[0031] Conversely, the genes in Table 2 show significantly higher expression levels in resistant plants grown in physiological disorder affected plantations. Without being bound by theory, it is suggested that the level of expression is increased due to the plant’s response to the disturbance stress and might be detected before visual observations of resistant phenotype is verified.
[0032] Hence the embodiments of the invention relate to the use of these genes as molecular markers for the identification of resistant and susceptible phenotypes to the abiotic disease.
[0033] Thus, embodiments of the invention relate to early markers of the plant’s phenotypic response to the physiological disturbance that can help in identifying potential resistant plants versus susceptible plants at an early stage before visible disturbance symptoms appear. Further provided herein are methods of genetically manipulating Eucalyptus to achieve a resistance to physiological disturbance. Thus, according to an aspect of the invention there is provided a method of identifying a physiological disturbance phenotype in Eucalyptus, the method comprising determining in a cell or tissue of a Eucalyptus plant a level of expression of at least one gene of Table 1 and / or Table 2, wherein: a statistically significant upregulation in expression of a gene of Table 1 relative to a resistant control is indicative of a susceptible physiological disturbance phenotype; and / or a statistically significant upregulation in expression of a gene of Table 2 relative to a susceptible control is indicative of a resistant physiological disturbance phenotype.
[0034] As used herein “physiological disturbance” refers to an abiotic physiological disease which is considered to be a result of a combination of environmental factors such as precipitation, temperature and atmospheric pressure. The phenomenon is endemic in some regions in Brazil, such as Bahia and Espirito Santo, though other areas where Eucalyptus grows commercially may also be affected.
[0035] Physiological disturbance is manifested by a plurality of symptoms including necrotic lesions on young and mature leaves and branches, loss of apical dominance for example as can be seen by increased branching in the canopy, cankers and / or epicormic sprouting along the main stem.
[0036] Physiological disturbance is progressively characterized by levels of severity, from completely healthy plants with no symptoms in the canopy and in the stem, to severe symptoms on the tree canopy, with apical shoot curling, necrosis in the plant tip and dieback, with loss of the apical dominance. Severe symptoms are also observed in the main stem of susceptible plants, with bark sunk, cankers and epicormic sprouting along the main stem.
[0037] According to a specific embodiment, the scoring system is based on visual symptoms on the main stem and general healthy aspect of the tree canopy. Score 0: no symptoms on the main stem and healthy canopy; Score 1: Sparse and mild cankers along the main stem, with healthy canopy; Score 2: Visible cankers along the main stem and die-back and necrosis symptoms in the canopy; Score 3: Severe cankers covering the whole main stem and severe die-back symptoms in the canopy. Scores 0 and 1 are ranked as resistant, while Scores 2 and 3 are considered susceptible
[0038] As used herein, "Resistance” and "improved resistance" are used interchangeably herein and refer to any type of increase in the frequency of resistant individuals within a given breeding population or resistance to physiological disturbance or a symptom thereof.
[0039] It will be appreciated that resistance may be interchanged with tolerance.
[0040] Likewise, sensitivity may replace susceptibility. A "resistant plant" or "tolerant plant variety" possess absolute or complete resistance, with no visual symptoms. A "susceptible plant," "susceptible plant variety," or a plant or plant variety with "higher level of susceptibility" will display symptoms of dieback, cankers along the main stem and cankers in branches. In between those two extremes there are levels of resistance or susceptibility, with plants displaying mild symptoms of disturbance, considered as moderately resistant and some plants displaying symptoms of intermediate severity, being classified as moderately susceptible.
[0041] As used herein “Eucalyptus” refers to a genus of over eight hundred species of flowering trees, shrubs or mallees in the myrtle family, Myrtaceae. Along with several other genera in the tribe Eucalypteae, including Corymbia, they are commonly known as “Eucalypts” . Plants in the genus Eucalyptus have bark that is either smooth, fibrous, hard or stringy, leaves with oil glands, and sepals and petals that are fused to form a "cap" or operculum over the stamens. The fruit is a woody capsule commonly referred to as a "gumnut".
[0042] Examples of species that can be used in the teachings of the present invention include, but are not limited to:
[0043] E. grandis
[0044] E. urophylla
[0045] E. pellita
[0046] E. robusta
[0047] E. cloeziana
[0048] E. saligna
[0049] E. tereticornis
[0050] E. camaldulensis
[0051] The plant can be a plant line or clone.
[0052] According to another embodiment, the plant is hybrid plant.
[0053] The hybrid can be intraspecific or interspecific. Hybrids among the same species, as for example different genotypes of E. grandis or other pure species are known as intraspecific hybrid. Hybrids between genotypes of different species, as for example obtained from a cross between a clone of E. grandis with a clone of E. urophylla are known as interspecific hybrids.
[0054] According to a specific embodiment, the Eucalypt is of a species selected for commercial use.
[0055] According to a specific embodiment, the Eucalyptus parent plant and / or second parent Eucalyptus plant are selected from the group consisting of E. grandis, E. urophylla, E. pellita, E. cloeziana, E. camaldulensis, E. globulus, E. robusta, E. saligna, , C. torelliana or other Eucalyptus species. As used herein, the term "plant" refers to an entire plant, its organs (i.e., leaves, stems, roots, flowers etc.), seeds, plant cells, and progeny of the same. The term "plant cell" includes without limitation cells within seeds, suspension cultures, embryos, meristematic regions, callus tissue, leaves, shoots, gametophytes, sporophytes, pollen, and microspores. According to a specific embodiment, the plant is a breeding line or clone.
[0056] According to a specific embodiment the plant line is an elite line. The phrase "plant part" refers to a part of a plant, including single cells and cell tissues such as plant cells that are intact in plants, cell clumps, and tissue cultures from which plants can be regenerated. Examples of plant parts include, but are not limited to, single cells and tissues from pollen, ovules, leaves, embryos, roots, root tips, anthers, flowers, fruits, stems, shoots, and seeds; as well as scions, rootstocks, protoplasts, calli, and the like. According to a specific embodiment, the plant part comprises the nucleic acid variation as described below. According to a specific embodiment, the plant part is a seed.
[0057] As used herein, the phrases "progeny plant" refers to any plant resulting as progeny from sexual reproduction from one or more parent plants or descendants thereof. An individual genotype within a progeny can be propagated by cuttings or any in vitro (= vegetative propagation) method, thus becoming a clone.
[0058] As used herein “upregulation” or an “increase” refers to at least about 2 %, at least about 3 %, at least about 4 %, at least about 5 %, at least about 10 %, at least about 15 %, at least about 20 %, at least about 30 %, at least about 40 %, at least about 50 %, at least about 60 %, at least about 70 %, at least about 80 %, at least about 2 fold, at least about 3 fold, at least about 5 fold, at least about 10 fold increase in the level of expression of the gene of Table 1 and / or 2 in a plant or part thereof as compared to a control plant of the invention of the same species which is grown under the same (e.g., identical) growth conditions.
[0059] As used herein “a statistically significant” refers to a statistical test known in the art such as student’s t-test.
[0060] Also provided herein is a kit for identifying a physiological disturbance phenotype in Eucalyptus, the kit comprising at least one reagent for identifying expression of at least one gene of Table 1 and / or Table 2, wherein said kit does not detect more than 100, 80, 50, 20 or 10 genes.
[0061] As used herein “reagent” refers to a substance or a molecule which is used in the art of cell biology, protein biochemistry or molecular biology which is used in the specific detection of gene expression. Thus the term may refer to primers, probes, primary antibodies but not to buffers, secondary antibodies and the like, which are not specific reagents.
[0062] According to a specific embodiment, the at least one reagent comprises a primer pair or a probe or a plurality of same which can be used in a sequential or simultaneous manner (e.g., multiplex).. For any of the aspects disclosed herein, the term “measuring” or “measurement,” or alternatively “detecting” or “detection,” means assessing the presence, absence, quantity or amount gene product (mRNA or protein), including the derivation of qualitative or quantitative concentration levels of such products.
[0063] According to some embodiments, said determining is at the DNA level.
[0064] According to some embodiments, said determining is at the RNA level.
[0065] According to some embodiments, said determining is at the protein level.
[0066] Methods of measuring the level of proteins are well known in the art and include, e.g., immunoassays based on antibodies to proteins, aptamers or molecular imprints.
[0067] Proteins can be detected in any suitable manner, but are typically detected by contacting a sample from the plant with an antibody, which binds the protein and then detecting the presence or absence of a reaction product. The antibody may be monoclonal, polyclonal, chimeric, or a fragment of the foregoing, and the step of detecting the reaction product may be carried out with any suitable immunoassay.
[0068] In one embodiment, the antibody which specifically binds the protein is attached (either directly or indirectly) to a signal producing label, including but not limited to a radioactive label, an enzymatic label, a hapten, a reporter dye or a fluorescent label.
[0069] Immunoassays carried out in accordance with some embodiments of the present invention may be homogeneous assays or heterogeneous assays. In a homogeneous assay the immunological reaction usually involves the specific antibody (e.g., proteins or Table 1 or 2), a labeled analyte, and the sample of interest. The signal arising from the label is modified, directly or indirectly, upon the binding of the antibody to the labeled analyte. Both the immunological reaction and detection of the extent thereof can be carried out in a homogeneous solution. Immunochemical labels, which may be employed, include free radicals, radioisotopes, fluorescent dyes, enzymes, bacteriophages, or coenzymes.
[0070] In a heterogeneous assay approach, the reagents are usually the sample, the antibody, and means for producing a detectable signal. Samples as described above may be used. The antibody can be immobilized on a support, such as a bead (such as protein A and protein G agarose beads), plate, pipette tip or slide, and contacted with the specimen suspected of containing the antigen in a liquid phase.
[0071] The support is then separated from the liquid phase and either the support phase or the liquid phase is examined for a detectable signal employing means for producing such signal. The signal is related to the presence of the analyte in the sample. Means for producing a detectable signal include the use of radioactive labels, fluorescent labels, or enzyme labels. For example, if the antigen to be detected contains a second binding site, an antibody which binds to that site can be conjugated to a detectable group and added to the liquid phase reaction solution before the separation step. The presence of the detectable group on the solid support indicates the presence of the antigen in the test sample. Examples of suitable immunoassays are oligonucleotides, immunoblotting, immunofluorescence methods, immunoprecipitation, chemiluminescence methods, electrochemiluminescence (ECL) or enzyme-linked immunoassays.
[0072] Those skilled in the art will be familiar with numerous specific immunoassay formats and variations thereof which may be useful for carrying out the method disclosed herein. See generally E. Maggio, Enzyme-Immunoassay, (1980) (CRC Press, Inc., Boca Raton, Fla.); see also U.S. Pat. No. 4,727,022 to Skold et al., titled “Methods for Modulating Ligand-Receptor Interactions and their Application,” U.S. Pat. No. 4,659,678 to Forrest et al., titled “Immunoassay of Antigens,” U.S. Pat. No. 4,376,110 to David et al., titled “Immunometric Assays Using Monoclonal Antibodies,” U.S. Pat. No. 4,275,149 to Litman et al., titled “Macromolecular Environment Control in Specific Receptor Assays,” U.S. Pat. No. 4,233,402 to Maggio et al., titled “Reagents and Method Employing Channeling,” and U.S. Pat. No. 4,230,767 to Boguslaski et al., titled “Heterogenous Specific Binding Assay Employing a Coenzyme as Label.” The protein can also be detected with antibodies using flow cytometry. Those skilled in the art will be familiar with flow cytometric techniques which may be useful in carrying out the methods disclosed herein (Shapiro 2005). These include, without limitation, Cytokine Bead Array (Becton Dickinson) and Luminex technology.
[0073] Antibodies can be conjugated to a solid support suitable for a detection assay (e.g., beads such as magnetic beads, protein A or protein G agarose, microspheres, plates, slides, pipette tip or wells formed from materials such as latex or polystyrene) in accordance with known techniques, such as passive binding. Antibodies as described herein may likewise be conjugated to detectable labels or groups such as radiolabels (e.g.,35S,125I,131I), enzyme labels (e.g., horseradish peroxidase, alkaline phosphatase), and fluorescent labels (e.g., fluorescein, Alexa, green fluorescent protein, rhodamine) in accordance with known techniques.
[0074] In particular embodiments, the antibodies of the present invention are monoclonal antibodies.
[0075] The presence of a label can be detected by inspection, or a detector which monitors a particular probe or probe combination is used to detect the detection reagent label. Typical detectors include spectrophotometers, phototubes and photodiodes, microscopes, scintillation counters, cameras, film and the like, as well as combinations thereof. Those skilled in the art will be familiar with numerous suitable detectors that widely available from a variety of commercial sources and may be useful for carrying out the method disclosed herein. Commonly, an optical image of a substrate comprising bound labeling moieties is digitized for subsequent computer analysis. See generally The Immunoassay Handbook [The Immunoassay Handbook. Third Edition. 2005].
[0076] RNA Analysis:
[0077] Isolation, extraction or derivation of RNA may be carried out by any suitable method. Isolating RNA from a biological sample generally includes treating a biological sample in such a manner that the RNA present in the sample is extracted and made available for analysis. Any isolation method that results in extracted RNA may be used in the practice of the present invention. It will be understood that the particular method used to extract RNA will depend on the nature of the source.
[0078] Methods of RNA extraction are well-known in the art and further described herein under.
[0079] Phenol based extraction methods: These single-step RNA isolation methods based on Guanidine isothiocyanate (GITC) / phenol / chloroform extraction require much less time than traditional methods (e.g. CsCh ultracentrifugation). Many commercial reagents (e.g. Trizol, RNAzol, RNAWIZ) are based on this principle. The entire procedure can be completed within an hour to produce high yields of total RNA.
[0080] Silica gel - based purification methods: RNeasy is a purification kit marketed by Qiagen. It uses a silica gel-based membrane in a spin-column to selectively bind RNA larger than 200 bases. The method is quick and does not involve the use of phenol.
[0081] Oligo-dT based affinity purification of mRNA: Due to the low abundance of mRNA in the total pool of cellular RNA, reducing the amount of rRNA and tRNA in a total RNA preparation greatly increases the relative amount of mRNA. The use of oligo-dT affinity chromatography to selectively enrich poly (A)+ RNA has been practiced for over 20 years. The result of the preparation is an enriched mRNA population that has minimal rRNA or other small RNA contamination. mRNA enrichment is essential for construction of cDNA libraries and other applications where intact mRNA is highly desirable. The original method utilized oligo-dT conjugated resin column chromatography and can be time consuming. Recently more convenient formats such as spin-column and magnetic bead based reagent kits have become available.
[0082] The sample may also be processed prior to carrying out the detection methods of the present invention. Processing of the sample may involve one or more of: filtration, distillation, centrifugation, extraction, concentration, dilution, purification, inactivation of interfering components, addition of reagents, and the like.
[0083] After obtaining the RNA sample, cDNA may be generated therefrom. For synthesis of cDNA, template mRNA may be obtained directly from lysed cells or may be purified from a total RNA or mRNA sample. The total RNA sample may be subjected to a force to encourage shearing of the RNA molecules such that the average size of each of the RNA molecules is between 100-300 nucleotides, e.g. about 200 nucleotides. To separate the heterogeneous population of mRNA from the majority of the RNA found in the cell, various technologies may be used which are based on the use of oligo(dT) oligonucleotides attached to a solid support. Examples of such oligo(dT) oligonucleotides include: oligo(dT) cellulose / spin columns, oligo(dT) / magnetic beads, and oligo(dT) oligonucleotide coated plates.
[0084] Generation of single stranded DNA from RNA requires synthesis of an intermediate RNA- DNA hybrid. For this, a primer is required that hybridizes to the 3’ end of the RNA. Annealing temperature and timing are determined both by the efficiency with which the primer is expected to anneal to a template and the degree of mismatch that is to be tolerated.
[0085] The annealing temperature is usually chosen to provide optimal efficiency and specificity, and generally ranges from about 50 °C to about 80°C, usually from about 55 °C to about 70 °C, and more usually from about 60 °C to about 68 °C. Annealing conditions are generally maintained for a period of time ranging from about 15 seconds to about 30 minutes, usually from about 30 seconds to about 5 minutes.
[0086] According to a specific embodiment, the primer comprises a polydT oligonucleotide sequence.
[0087] Preferably the polydT sequence comprises at least 5 nucleotides. According to another is between about 5 to 50 nucleotides, more preferably between about 5-25 nucleotides, and even more preferably between about 12 to 14 nucleotides.
[0088] Following annealing of the primer (e.g. polydT primer) to the RNA sample, an RNA-DNA hybrid is synthesized by reverse transcription using an RNA-dependent DNA polymerase. Suitable RNA-dependent DNA polymerases for use in the methods and compositions of the invention include reverse transcriptases (RTs). Examples of RTs include, but are not limited to, Moloney murine leukemia virus (M-MEV) reverse transcriptase, human immunodeficiency virus (HIV) reverse transcriptase, rous sarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, rous associated virus (RAV) reverse transcriptase, and myeloblastosis associated virus (MAV) reverse transcriptase or other avian sarcoma-leukosis virus (ASEV) reverse transcriptases, and modified RTs derived therefrom. See e.g. U.S. Patent No. 7,056,716. Many reverse transcriptases, such as those from avian myeloblastosis virus (AMV-RT), and Moloney murine leukemia virus (MMEV-RT) comprise more than one activity (for example, polymerase activity and ribonuclease activity) and can function in the formation of the double stranded cDNA molecules. Additional components required in a reverse transcription reaction include dNTPS (dATP, dCTP, dGTP and dTTP) and optionally a reducing agent such as Dithiothreitol (DTT) and MnCh.
[0089] Methods of analyzing the amount of RNA are known in the art and are summarized infra:
[0090] Northern Blot analysis: This method involves the detection of a particular RNA in a mixture of RNAs. An RNA sample is denatured by treatment with an agent (e.g., formaldehyde) that prevents hydrogen bonding between base pairs, ensuring that all the RNA molecules have an unfolded, linear conformation. The individual RNA molecules are then separated according to size by gel electrophoresis and transferred to a nitrocellulose or a nylon-based membrane to which the denatured RNAs adhere. The membrane is then exposed to labeled DNA probes. Probes may be labeled using radio-isotopes or enzyme linked nucleotides. Detection may be using autoradiography, colorimetric reaction or chemiluminescence. This method allows both quantitation of an amount of particular RNA molecules and determination of its identity by a relative position on the membrane which is indicative of a migration distance in the gel during electrophoresis.
[0091] RT-PCR analysis: This method uses PCR amplification of relatively rare RNAs molecules. First, RNA molecules are purified from the cells and converted into complementary DNA (cDNA) using a reverse transcriptase enzyme (such as an MMLV-RT) and primers such as, oligo dT, random hexamers or gene specific primers. Then by applying gene specific primers and Taq DNA polymerase, a PCR amplification reaction is carried out in a PCR machine. Those of skills in the art are capable of selecting the length and sequence of the gene specific primers and the PCR conditions (z.e., annealing temperatures, number of cycles and the like) which are suitable for detecting specific RNA molecules. It will be appreciated that a semi-quantitative RT-PCR reaction can be employed by adjusting the number of PCR cycles and comparing the amplification product to known controls. Isothermal amplification is also contemplated.
[0092] RNA in situ hybridization stain: In this method DNA or RNA probes are attached to the RNA molecules present in the cells. Generally, the cells are first fixed to microscopic slides to preserve the cellular structure and to prevent the RNA molecules from being degraded and then are subjected to hybridization buffer containing the labeled probe. The hybridization buffer includes reagents such as formamide and salts (e.g., sodium chloride and sodium citrate) which enable specific hybridization of the DNA or RNA probes with their target mRNA molecules in situ while avoiding non-specific binding of probe. Those of skills in the art are capable of adjusting the hybridization conditions (z.e., temperature, concentration of salts and formamide and the like) to specific probes and types of cells. Following hybridization, any unbound probe is washed off and the bound probe is detected using known methods. For example, if a radio-labeled probe is used, then the slide is subjected to a photographic emulsion which reveals signals generated using radio-labeled probes; if the probe was labeled with an enzyme then the enzyme-specific substrate is added for the formation of a colorimetric reaction; if the probe is labeled using a fluorescent label, then the bound probe is revealed using a fluorescent microscope; if the probe is labeled using a tag (e.g., digoxigenin, biotin, and the like) then the bound probe can be detected following interaction with a tag-specific antibody which can be detected using known methods.
[0093] In situ RT-PCR stain: This method is described in Nuovo GJ, et al. [Intracellular localization of polymerase chain reaction (PCR)-amplified hepatitis C cDNA. Am J Surg Pathol. 1993, 17: 683- 90] and Komminoth P, et al. [Evaluation of methods for hepatitis C virus detection in archival liver biopsies. Comparison of histology, immunohistochemistry, in situ hybridization, reverse transcriptase polymerase chain reaction (RT-PCR) and in situ RT-PCR. Pathol Res Pract. 1994, 190: 1017-25]. Briefly, the RT-PCR reaction is performed on fixed cells by incorporating labeled nucleotides to the PCR reaction. The reaction is carried on using a specific in situ RT-PCR apparatus such as the lasercapture microdissection PixCell I LCM system available from Arcturus Engineering (Mountainview, CA).
[0094] DNA microarray s / DN A chips:
[0095] The expression of thousands of genes may be analyzed simultaneously using DNA microarrays, allowing analysis of the complete transcriptional program of an organism during specific developmental processes or physiological responses. DNA microarrays consist of thousands of individual gene sequences attached to closely packed areas on the surface of a support such as a glass microscope slide. Various methods have been developed for preparing DNA microarrays. In one method, an approximately 1 kilobase segment of the coding region of each gene for analysis is individually PCR amplified. A robotic apparatus is employed to apply each amplified DNA sample to closely spaced zones on the surface of a glass microscope slide, which is subsequently processed by thermal and chemical treatment to bind the DNA sequences to the surface of the support and denature them. Typically, such arrays are about 2 x 2 cm and contain about individual nucleic acids 6000 spots. In a variant of the technique, multiple DNA oligonucleotides, usually 20 nucleotides in length, are synthesized from an initial nucleotide that is covalently bound to the surface of a support, such that tens of thousands of identical oligonucleotides are synthesized in a small square zone on the surface of the support. Multiple oligonucleotide sequences from a single gene are synthesized in neighboring regions of the slide for analysis of expression of that gene. Hence, thousands of genes can be represented on one glass slide. Such arrays of synthetic oligonucleotides may be referred to in the art as “DNA chips”, as opposed to “DNA microarrays”, as described above [Lodish et al. (eds.). Chapter 7.8: DNA Microarrays: Analyzing Genome-Wide Expression. In: Molecular Cell Biology, 4th ed., W. H. Freeman, New York. (2000)]. Oligonucleotide microarray - In this method oligonucleotide probes capable of specifically hybridizing with the polynucleotides of some embodiments of the invention are attached to a solid surface (e.g., a glass wafer). Each oligonucleotide probe is of approximately 20-25 nucleic acids in length. To detect the expression pattern of the polynucleotides of some embodiments of the invention in a specific cell sample (e.g., leaf cells), RNA is extracted from the cell sample using methods known in the art (using e.g., a TRIZOL solution, Gibco BRL, USA). Hybridization can take place using either labeled oligonucleotide probes (e.g., 5 '-biotinylated probes) or labeled fragments of complementary DNA (cDNA) or RNA (cRNA). Briefly, double stranded cDNA is prepared from the RNA using reverse transcriptase (RT) (e.g., Superscript II RT), DNA ligase and DNA polymerase I, all according to manufacturer’s instructions (Invitrogen Life Technologies, Frederick, MD, USA). To prepare labeled cRNA, the double stranded cDNA is subjected to an in vitro transcription reaction in the presence of biotinylated nucleotides using e.g., the BioArray High Yield RNA Transcript Labeling Kit (Enzo, Diagnostics, Affymetix Santa Clara CA). For efficient hybridization the labeled cRNA can be fragmented by incubating the RNA in 40 mM Tris Acetate (pH 8.1), 100 mM potassium acetate and 30 mM magnesium acetate for 35 minutes at 94 °C. Following hybridization, the microarray is washed and the hybridization signal is scanned using a confocal laser fluorescence scanner which measures fluorescence intensity emitted by the labeled cRNA bound to the probe arrays.
[0096] For example, in the Affymetrix microarray (Affymetrix®, Santa Clara, CA) each gene on the array is represented by a series of different oligonucleotide probes, of which, each probe pair consists of a perfect match oligonucleotide and a mismatch oligonucleotide. While the perfect match probe has a sequence exactly complimentary to the particular gene, thus enabling the measurement of the level of expression of the particular gene, the mismatch probe differs from the perfect match probe by a single base substitution at the center base position. The hybridization signal is scanned using the Agilent scanner, and the Microarray Suite software subtracts the non-specific signal resulting from the mismatch probe from the signal resulting from the perfect match probe.
[0097] RNA sequencing: Methods for RNA sequence determination are generally known to the person skilled in the art. Preferred sequencing methods are next generation sequencing methods or parallel high throughput sequencing methods. An example of an envisaged sequence method is pyrosequencing, in particular 454 pyrosequencing, e.g. based on the Roche 454 Genome Sequencer. This method amplifies DNA inside water droplets in an oil solution with each droplet containing a single DNA template attached to a single primer-coated bead that then forms a clonal colony. Pyrosequencing uses luciferase to generate light for detection of the individual nucleotides added to the nascent DNA, and the combined data are used to generate sequence read-outs. Yet another envisaged example is Illumina or Solexa sequencing, e.g. by using the Illumina Genome Analyzer technology, which is based on reversible dye-terminators. DNA molecules are typically attached to primers on a slide and amplified so that local clonal colonies are formed. Subsequently one type of nucleotide at a time may be added, and non-incorporated nucleotides are washed away. Subsequently, images of the fluorescently labeled nucleotides may be taken and the dye is chemically removed from the DNA, allowing a next cycle. Yet another example is the use of Applied Biosystems' SOLiD technology, which employs sequencing by ligation. This method is based on the use of a pool of all possible oligonucleotides of a fixed length, which are labeled according to the sequenced position. Such oligonucleotides are annealed and ligated. Subsequently, the preferential ligation by DNA ligase for matching sequences typically results in a signal informative of the nucleotide at that position. Since the DNA is typically amplified by emulsion PCR, the resulting bead, each containing only copies of the same DNA molecule, can be deposited on a glass slide resulting in sequences of quantities and lengths comparable to Illumina sequencing. A further method is based on Helicos' Heliscope technology, wherein fragments are captured by polyT oligomers tethered to an array. At each sequencing cycle, polymerase and single fluorescently labeled nucleotides are added and the array is imaged. The fluorescent tag is subsequently removed and the cycle is repeated. Further examples of sequencing techniques encompassed within the methods of the present invention are sequencing by hybridization, sequencing by use of nanopores, microscopy-based sequencing techniques, microfluidic Sanger sequencing, or microchip-based sequencing methods. The present invention also envisages further developments of these techniques, e.g. further improvements of the accuracy of the sequence determination, or the time needed for the determination of the genomic sequence of an organism etc.
[0098] According to one embodiment, the sequencing method comprises deep sequencing.
[0099] As used herein, the term “deep sequencing” refers to a sequencing method wherein the target sequence is read multiple times in the single test. A single deep sequencing run is composed of a multitude of sequencing reactions run on the same target sequence and each, generating independent sequence readout.
[0100] It will be appreciated that in order to analyze the amount of an RNA marker, oligonucleotides may be used that are capable of hybridizing thereto or to cDNA generated therefrom. According to one embodiment a single oligonucleotide is used to determine the presence of a particular RNA marker, at least two oligonucleotides are used to determine the presence of a particular RNA marker, at least three oligonucleotides are used to determine the presence of a particular RNA marker, at least four oligonucleotides are used to determine the presence of a particular RNA marker, at least five or more oligonucleotides are used to determine the presence of a particular RNA marker. When more than one oligonucleotide is used, the sequence of the oligonucleotides may be selected such that they hybridize to the same exon of the RNA marker or different exons of the RNA marker. In one embodiment, at least one of the oligonucleotides hybridizes to the 3’ exon of the RNA markder. In another embodiment, at least one of the oligonucleotides hybridizes to the 5’ exon of the RNA marker.
[0101] In one embodiment, the method of this aspect of the present invention is carried out using an isolated oligonucleotide which hybridizes to the RNA or cDNA of any of the RNA markers disclosed herein by complementary base-pairing in a sequence specific manner, and discriminates the gene sequence from other nucleic acid sequence in the sample. Oligonucleotides (e.g. DNA or RNA oligonucleotides) typically comprises a region of complementary nucleotide sequence that hybridizes under stringent conditions to at least about 8, 10, 13, 16, 18, 20, 22, 25, 30, 40, 50, 55, 60, 65, 70, 80, 90, 100, 120 (or any other number in-between) or more consecutive nucleotides in a target nucleic acid molecule. Depending on the particular assay, the consecutive nucleotides include the nucleic acid sequence.
[0102] The term "isolated", as used herein in reference to an oligonucleotide, means an oligonucleotide, which by virtue of its origin or manipulation, is separated from at least some of the components with which it is naturally associated or with which it is associated when initially obtained. By "isolated", it is alternatively or additionally meant that the oligonucleotide of interest is produced or synthesized by the hand of man.
[0103] In order to identify an oligonucleotide specific for any of the RNA markers disclosed herein, the gene / transcript of interest is typically examined using a computer algorithm which starts at the 5' or at the 3' end of the nucleotide sequence. Typical algorithms will then identify oligonucleotides of defined length that are unique to the gene, have a GC content within a range suitable for hybridization, lack predicted secondary structure that may interfere with hybridization, and / or possess other desired characteristics or that lack other undesired characteristics.
[0104] Following identification of the oligonucleotide it may be tested for specificity towards the gene (of Table 1 or 2 or mRNA product thereof) under wet or dry conditions. Thus, for example, in the case where the oligonucleotide is a primer, the primer may be tested for its ability to amplify a sequence of the gene (of Table 1 or 2 or mRNA product thereof) using PCR to generate a detectable product and for its non ability to amplify other mRNA products in the sample. The products of the PCR reaction may be analyzed on a gel and verified according to presence and / or size.
[0105] Additionally, or alternatively, the sequence of the oligonucleotide may be analyzed by computer analysis to see if it is homologous (or is capable of hybridizing to) other known sequences. A BLAST 2.2.10 (Basic Local Alignment Search Tool) analysis may be performed on the chosen oligonucleotide (www(dot)ncbi(dot)nlm(dot)nih(dot)gov / blast / ). The BLAST program finds regions of local similarity between sequences. It compares nucleotide or protein sequences to sequence databases and calculates the statistical significance of matches thereby providing valuable information about the possible identity and integrity of the ‘query’ sequences.
[0106] According to one embodiment, the oligonucleotide is a probe. As used herein, the term "probe" refers to an oligonucleotide which hybridizes to the specific nucleic acid sequence of the gene (of Table 1 or 2 or mRNA product thereof) to provide a detectable signal under experimental conditions and which does not hybridize to additional sequences to provide a detectable signal under identical experimental conditions.
[0107] The probes of this embodiment of this aspect of the present invention may be, for example, affixed to a solid support (e.g., arrays or beads).
[0108] Solid supports are solid-state substrates or supports onto which the nucleic acid molecules of the present invention may be associated. The nucleic acids may be associated directly or indirectly. Solid-state substrates for use in solid supports can include any solid material with which components can be associated, directly or indirectly. This includes materials such as acrylamide, agarose, cellulose, nitrocellulose, glass, gold, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, poly silicates, polycarbonates, teflon, fluorocarbons, nylon, silicon rubber, polyanhydrides, polyglycolic acid, polylactic acid, polyorthoesters, functionalized silane, polypropylfumerate, collagen, glycosaminoglycans, and polyamino acids. Solid-state substrates can have any useful form including thin film, membrane, bottles, dishes, fibers, woven fibers, shaped polymers, particles, beads, microparticles, or a combination. Solid-state substrates and solid supports can be porous or non-porous. A chip is a rectangular or square small piece of material. Preferred forms for solid-state substrates are thin films, beads, or chips. A useful form for a solid-state substrate is a microtiter dish. In some embodiments, a multiwell glass slide can be employed.
[0109] In one embodiment, the solid support is an array which comprises a plurality of nucleic acids which hybridize to RNA markers of the present invention immobilized at identified or predefined locations on the solid support. Each predefined location on the solid support generally has one type of component (that is, all the components at that location are the same). Alternatively, multiple types of components can be immobilized in the same predefined location on a solid support. Each location will have multiple copies of the given components. The spatial separation of different components on the solid support allows separate detection and identification. According to particular embodiments, the array does not comprise nucleic acids that specifically bind to more than 50 RNA markers, more than 40 RNA markers, 30 RNA markers, 20 RNA markers, 15 RNA markers, 10 RNA markers, 5 RNA markers or even 3 RNA markers.
[0110] Methods for immobilization of oligonucleotides to solid-state substrates are well established. Oligonucleotides, including address probes and detection probes, can be coupled to substrates using established coupling methods. For example, suitable attachment methods are described by Pease et al., Proc. Natl. Acad. Sci. USA 91(l l):5022-5026 (1994), and Khrapko et al., Mol Biol (Mosk) (USSR) 25:718-730 (1991). A method for immobilization of 3'-amine oligonucleotides on casein- coated slides is described by Stimpson et al., Proc. Natl. Acad. Sci. USA 92:6379-6383 (1995). A useful method of attaching oligonucleotides to solid-state substrates is described by Guo et al., Nucleic Acids Res. 22:5456-5465 (1994).
[0111] According to another embodiment, the oligonucleotide is a primer of a primer pair. As used herein, the term "primer" refers to an oligonucleotide which acts as a point of initiation of a template- directed synthesis using methods such as PCR (polymerase chain reaction) or LCR (ligase chain reaction) under appropriate conditions (e.g., in the presence of four different nucleotide triphosphates and a polymerization agent, such as DNA polymerase, RNA polymerase or reverse-transcriptase, DNA ligase, etc, in an appropriate buffer solution containing any necessary co-factors and at suitable temperature(s)). Such a template directed synthesis is also called "primer extension". For example, a primer pair may be designed to amplify a region of DNA using PCR. Such a pair will include a "forward primer" and a "reverse primer" that hybridize to complementary strands of a DNA molecule and that delimit a region to be synthesized / amplified. A primer of this aspect of the present invention is capable of amplifying, together with its pair (e.g. by PCR) a specific nucleic acid sequence (of Table 1 or 2) to provide a detectable signal under experimental conditions and which does not amplify other nucleic acid sequences to provide a detectable signal under identical experimental conditions.
[0112] According to additional embodiments, the oligonucleotide is about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides in length. While the maximal length of a probe can be as long as the target sequence to be detected, depending on the type of assay in which it is employed, it is typically less than about 50, 60, 65, or 70 nucleotides in length. In the case of a primer, it is typically less than about 30 nucleotides in length. In a specific preferred embodiment of the invention, a primer or a probe is within the length of about 18 and about 28 nucleotides. It will be appreciated that when attached to a solid support, the probe may be of about 30-70, 75, 80, 90, 100, or more nucleotides in length.
[0113] The oligonucleotide of this aspect of the present invention need not reflect the exact sequence of the RNA marker nucleic acid sequence (i.e. need not be fully complementary), but must be sufficiently complementary to hybridize with the nucleic acid sequence (of Table 1 or 2) under the particular experimental conditions. Accordingly, the sequence of the oligonucleotide typically has at least 70 % homology, preferably at least 80 %, 90 %, 95 %, 97 %, 99 % or 100 % homology, for example over a region of at least 13 or more contiguous nucleotides with the target nucleic acid sequence. The conditions are selected such that hybridization of the oligonucleotide to the nucleic acid sequence (of Table 1 or 2) is favored and hybridization to other nucleic acid sequences is minimized.
[0114] By way of example, hybridization of short nucleic acids (below 200 bp in length, e.g. 13-50 bp in length) can be effected by the following hybridization protocols depending on the desired stringency; (i) hybridization solution of 6 x SSC and 1 % SDS or 3 M TMAC1, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS, 100 qg / ml denatured salmon sperm DNA and 0.1 % nonfat dried milk, hybridization temperature of 1 - 1.5 °C below the Tm, final wash solution of 3 M TMAC1, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS at 1 - 1.5 °C below the Tm (stringent hybridization conditions) (ii) hybridization solution of 6 x SSC and 0.1 % SDS or 3 M TMACI, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS, 100 p.g / ml denatured salmon sperm DNA and 0.1 % nonfat dried milk, hybridization temperature of 2 - 2.5 °C below the Tm, final wash solution of 3 M TMACI, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS at 1 - 1.5 °C below the Tm, final wash solution of 6 x SSC, and final wash at 22 °C (stringent to moderate hybridization conditions); and (iii) hybridization solution of 6 x SSC and 1 % SDS or 3 M TMACI, 0.01 M sodium phosphate (pH 6.8), 1 mM EDTA (pH 7.6), 0.5 % SDS, 100 qg / ml denatured salmon sperm DNA and 0.1 % nonfat dried milk, hybridization temperature at 2.5-3 °C below the Tm and final wash solution of 6 x SSC at 22 °C (moderate hybridization solution).
[0115] Oligonucleotides of the invention may be prepared by any of a variety of methods (see, for example, J. Sambrook et al., "Molecular Cloning: A Laboratory Manual", 1989, 2. sup. nd Ed., Cold Spring Harbour Laboratory Press: New York, N.Y.; "PCR Protocols: A Guide to Methods and Applications", 1990, M. A. Innis (Ed.), Academic Press: New York, N.Y.; P. Tijssen "Hybridization with Nucleic Acid Probes— Laboratory Techniques in Biochemistry and Molecular Biology (Parts I and II)", 1993, Elsevier Science; "PCR Strategies", 1995, M. A. Innis (Ed.), Academic Press: New York, N.Y.; and "Short Protocols in Molecular Biology", 2002, F. M. Ausubel (Ed.), 5. sup. th Ed., John Wiley & Sons: Secaucus, N.J.). For example, oligonucleotides may be prepared using any of a variety of chemical techniques well-known in the art, including, for example, chemical synthesis and polymerization based on a template as described, for example, in S. A. Narang et al., Meth. Enzymol. 1979, 68: 90-98; E. L. Brown et al., Meth. Enzymol. 1979, 68: 109-151; E. S. Belousov et al., Nucleic Acids Res. 1997, 25: 3440-3444; D. Guschin et al., Anal. Biochem. 1997, 250: 203-211; M. J. Blommers et al., Biochemistry, 1994, 33: 7886-7896; and K. Frenkel et al., Free Radic. Biol. Med. 1995, 19: 373-380; and U.S. Pat. No. 4,458,066.
[0116] For example, oligonucleotides may be prepared using an automated, solid-phase procedure based on the phosphoramidite approach. In such a method, each nucleotide is individually added to the 5'-end of the growing oligonucleotide chain, which is attached at the 3 '-end to a solid support. The added nucleotides are in the form of trivalent 3'-phosphoramidites that are protected from polymerization by a dimethoxytriyl (or DMT) group at the 5'-position. After base-induced phosphoramidite coupling, mild oxidation to give a pentavalent phosphotriester intermediate and DMT removal provides a new site for oligonucleotide elongation. The oligonucleotides are then cleaved off the solid support, and the phosphodiester and exocyclic amino groups are deprotected with ammonium hydroxide. These syntheses may be performed on oligo synthesizers such as those commercially available from Perkin Elmer / Applied Biosystems, Inc. (Foster City, Calif.), DuPont (Wilmington, Del.) or Milligen (Bedford, Mass.). Alternatively, oligonucleotides can be custom made and ordered from a variety of commercial sources well-known in the art, including, for example, the Midland Certified Reagent Company (Midland, Tex.), ExpressGen, Inc. (Chicago, Ill.), Operon Technologies, Inc. (Huntsville, Ala.), and many others.
[0117] Purification of the oligonucleotides of the invention, where necessary or desirable, may be carried out by any of a variety of methods well-known in the art. Purification of oligonucleotides is typically performed either by native acrylamide gel electrophoresis, by anion-exchange HPLC as described, for example, by J. D. Pearson and F. E. Regnier (J. Chrom., 1983, 255: 137-149) or by reverse phase HPLC (G. D. McFarland and P. N. Borer, Nucleic Acids Res., 1979, 7: 1067-1080).
[0118] The sequence of oligonucleotides can be verified using any suitable sequencing method including, but not limited to, chemical degradation (A. M. Maxam and W. Gilbert, Methods of Enzymology, 1980, 65: 499-560), matrix-assisted laser desorption ionization time-of-flight (MALDL TOF) mass spectrometry (U. Pieles et al., Nucleic Acids Res., 1993, 21: 3191-3196), mass spectrometry following a combination of alkaline phosphatase and exonuclease digestions (H. Wu and H. Aboleneen, Anal. Biochem., 2001, 290: 347-352), and the like.
[0119] As already mentioned above, modified oligonucleotides may be prepared using any of several means known in the art. Non-limiting examples of such modifications include methylation, "caps", substitution of one or more of the naturally occurring nucleotides with an analog, and internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoroamidates, carbamates, etc), or charged linkages (e.g., phosphorothioates, phosphorodithioates, etc). Oligonucleotides may contain one or more additional covalently linked moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L- lysine, etc), intercalators (e.g., acridine, psoralen, etc), chelators (e.g., metals, radioactive metals, iron, oxidative metals, etc), and alkylators. The oligonucleotide may also be derivatized by formation of a methyl or ethyl phosphotriester or an alkyl phosphoramidate linkage. Furthermore, the oligonucleotide sequences of the present invention may also be modified with a label.
[0120] In certain embodiments, the detection probes or amplification primers or both probes and primers are labeled with a detectable agent or moiety before being used in amplification / detection assays. In certain embodiments, the detection probes are labeled with a detectable agent. Preferably, a detectable agent is selected such that it generates a signal which can be measured and whose intensity is related (e.g., proportional) to the amount of amplification products in the sample being analyzed.
[0121] The association between the oligonucleotide and detectable agent can be covalent or non- covalent. Labeled detection probes can be prepared by incorporation of or conjugation to a detectable moiety. Labels can be attached directly to the nucleic acid sequence or indirectly (e.g., through a linker). Linkers or spacer arms of various lengths are known in the art and are commercially available, and can be selected to reduce steric hindrance, or to confer other useful or desired properties to the resulting labeled molecules (see, for example, E. S. Mansfield et al., Mol. Cell. Probes, 1995, 9: 145- 156).
[0122] Methods for labeling nucleic acid molecules are well-known in the art. For a review of labeling protocols, label detection techniques, and recent developments in the field, see, for example, L. J. Kricka, Ann. Clin. Biochem. 2002, 39: 114-129; R. P. van Gijlswijk et al., Expert Rev. Mol. Diagn. 2001, 1: 81-91; and S. Joos et al., J. Biotechnol. 1994, 35: 135-153. Standard nucleic acid labeling methods include: incorporation of radioactive agents, direct attachments of fluorescent dyes (L. M. Smith et al., Nucl. Acids Res., 1985, 13: 2399-2412) or of enzymes (B. A. Connoly and O. Rider, Nucl. Acids. Res., 1985, 13: 4485-4502); chemical modifications of nucleic acid molecules making them detectable immunochemically or by other affinity reactions (T. R. Broker et al., Nucl. Acids Res. 1978, 5: 363-384; E. A. Bayer et al., Methods of Biochem. Analysis, 1980, 26: 1-45; R. Langer et al., Proc. Natl. Acad. Sci. USA, 1981, 78: 6633-6637; R. W. Richardson et al., Nucl. Acids Res. 1983, 11: 6167-6184; D. J. Brigati et al., Virol. 1983, 126: 32-50; P. Tchen et al., Proc. Natl. Acad. Sci. USA, 1984, 81: 3466-3470; J. E. Landegent et al., Exp. Cell Res. 1984, 15: 61-72; and A. H. Hopman et al., Exp. Cell Res. 1987, 169: 357-368); and enzyme-mediated labeling methods, such as random priming, nick translation, PCR and tailing with terminal transferase (for a review on enzymatic labeling, see, for example, J. Temsamani and S. Agrawal, Mol. Biotechnol. 1996, 5: 223- 232). More recently developed nucleic acid labeling systems include, but are not limited to: ULS (Universal Linkage System), which is based on the reaction of mono-reactive cisplatin derivatives with the N7 position of guanine moieties in DNA (R. J. Heetebrij et al., Cytogenet. Cell. Genet. 1999, 87: 47-52), psoralen-biotin, which intercalates into nucleic acids and upon UV irradiation becomes covalently bonded to the nucleotide bases (C. Levenson et al., Methods Enzymol. 1990, 184: 577- 583; and C. Pfannschmidt et al., Nucleic Acids Res. 1996, 24: 1702-1709), photoreactive azido derivatives (C. Neves et al., Bioconjugate Chem. 2000, 11: 51-55), and DNA alkylating agents (M. G. Sebestyen et al., Nat. Biotechnol. 1998, 16: 568-576).
[0123] Any of a wide variety of detectable agents can be used in the practice of the present invention. Suitable detectable agents include, but are not limited to, various ligands, radionuclides (such as, for example,32P,35S,3H,14C,1251,131I, and the like); fluorescent dyes (for specific exemplary fluorescent dyes, see below); chemiluminescent agents (such as, for example, acridinium esters, stabilized dioxetanes, and the like); spectrally resolvable inorganic fluorescent semiconductor nanocrystals (i.e., quantum dots), metal nanoparticles (e.g., gold, silver, copper and platinum) or nanoclusters; enzymes (such as, for example, those used in an ELISA, i.e., horseradish peroxidase, beta-galactosidase, luciferase, alkaline phosphatase); colorimetric labels (such as, for example, dyes, colloidal gold, and the like); magnetic labels (such as, for example, Dynabeads™); and biotin, dioxigenin or other haptens and proteins for which antisera or monoclonal antibodies are available.
[0124] In certain embodiments, the detection probes are fluorescently labeled. Numerous known fluorescent labeling moieties of a wide variety of chemical structures and physical characteristics are suitable for use in the practice of this invention. Suitable fluorescent dyes include, but are not limited to, fluorescein and fluorescein dyes (e.g., fluorescein isothiocyanine or FITC, naphthofluorescein, 4',5'-dichloro-2',7'-dimethoxy-fluorescein, 6 carboxyfluorescein or FAM), carbocyanine, merocyanine, styryl dyes, oxonol dyes, phycoerythrin, erythrosin, eosin, rhodamine dyes (e.g., carboxytetramethylrhodamine or TAMRA, carboxyrhodamine 6G, carboxy-X-rhodamine (ROX), lissamine rhodamine B, rhodamine 6G, rhodamine Green, rhodamine Red, tetramethylrhodamine or TMR), coumarin and coumarin dyes (e.g., methoxycoumarin, dialkylaminocoumarin, hydroxycoumarin and aminomethylcoumarin or AMCA), Oregon Green Dyes (e.g., Oregon Green 488, Oregon Green 500, Oregon Green 514), Texas Red, Texas Red-X, Spectrum Red.TM., Spectrum Green.TM., cyanine dyes (e.g., Cy-3™, Cy-5™, Cy-3.5™, Cy-5.5™), Alexa Fluor dyes (e.g., Alexa Fluor 350, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 633, Alexa Fluor 660 and Alexa Fluor 680), BODIPY dyes (e.g., BODIPY FL, BODIPY R6G, BODIPY TMR, BODIPY TR, BODIPY 530 / 550, BODIPY 558 / 568, BODIPY 564 / 570, BODIPY 576 / 589, BODIPY 581 / 591, BODIPY 630 / 650, BODIPY 650 / 665), IRDyes (e.g., IRD40, IRD 700, IRD 800), and the like. For more examples of suitable fluorescent dyes and methods for linking or incorporating fluorescent dyes to nucleic acid molecules see, for example, "The Handbook of Fluorescent Probes and Research Products", 9th Ed., Molecular Probes, Inc., Eugene, Oreg. Fluorescent dyes as well as labeling kits are commercially available from, for example, Amersham Biosciences, Inc. (Piscataway, N.J.), Molecular Probes Inc. (Eugene, Oreg.), and New England Biolabs Inc. (Berverly, Mass.).
[0125] As mentioned, identification of the RNA marker may be carried out using an amplification reaction.
[0126] As used herein, the term "amplification" refers to a process that increases the representation of a population of specific nucleic acid sequences in a sample by producing multiple (i.e., at least 2) copies of the desired sequences. Methods for nucleic acid amplification are known in the art and include, but are not limited to, polymerase chain reaction (PCR) and ligase chain reaction (LCR). In a typical PCR amplification reaction, a nucleic acid sequence of interest is often amplified at least fifty thousand fold in amount over its amount in the starting sample. A "copy" or "amplicon" does not necessarily mean perfect sequence complementarity or identity to the template sequence. For example, copies can include nucleotide analogs such as deoxyinosine, intentional sequence alterations (such as sequence alterations introduced through a primer comprising a sequence that is hybridizable but not complementary to the template), and / or sequence errors that occur during amplification.
[0127] A typical amplification reaction is carried out by contacting a forward and reverse primer (a primer pair) to the sample DNA together with any additional amplification reaction reagents under conditions which allow amplification of the target sequence.
[0128] The terms "forward primer" and "forward amplification primer" are used herein interchangeably, and refer to a primer that hybridizes (or anneals) to the target (template strand). The terms "reverse primer" and "reverse amplification primer" are used herein interchangeably, and refer to a primer that hybridizes (or anneals) to the complementary target strand. The forward primer hybridizes with the target sequence 5' with respect to the reverse primer.
[0129] The term "amplification conditions", as used herein, refers to conditions that promote annealing and / or extension of primer sequences. Such conditions are well-known in the art and depend on the amplification method selected. Thus, for example, in a PCR reaction, amplification conditions generally comprise thermal cycling, i.e., cycling of the reaction mixture between two or more temperatures. In isothermal amplification reactions, amplification occurs without thermal cycling although an initial temperature increase may be required to initiate the reaction. Amplification conditions encompass all reaction conditions including, but not limited to, temperature and temperature cycling, buffer, salt, ionic strength, and pH, and the like. As used herein, the term "amplification reaction reagents", refers to reagents used in nucleic acid amplification reactions and may include, but are not limited to, buffers, reagents, enzymes having reverse transcriptase and / or polymerase activity or exonuclease activity, enzyme cofactors such as magnesium or manganese, salts, nicotinamide adenine dinuclease (NAD) and deoxynucleoside triphosphates (dNTPs), such as deoxy adenosine triphospate, deoxyguanosine triphosphate, deoxycytidine triphosphate and thymidine triphosphate. Amplification reaction reagents may readily be selected by one skilled in the art depending on the amplification method used.
[0130] According to this aspect of the present invention, the amplifying may be effected using techniques such as polymerase chain reaction (PCR), which includes, but is not limited to Allelespecific PCR, Assembly PCR or Polymerase Cycling Assembly (PCA), Asymmetric PCR, Helicasedependent amplification, Hot-start PCR, Intersequence- specific PCR (ISSR), Inverse PCR, Ligation- mediated PCR, Methylation-specific PCR (MSP), Miniprimer PCR, Multiplex Ligation-dependent Probe Amplification, Multiplex-PCR, Nested PCR, Overlap-extension PCR, Quantitative PCR (Q- PCR), Reverse Transcription PCR (RT-PCR), Solid Phase PCR: encompasses multiple meanings, including Polony Amplification (where PCR colonies are derived in a gel matrix, for example), Bridge PCR (primers are covalently linked to a solid-support surface), conventional Solid Phase PCR (where Asymmetric PCR is applied in the presence of solid support bearing primer with sequence matching one of the aqueous primers) and Enhanced Solid Phase PCR (where conventional Solid Phase PCR can be improved by employing high Tm and nested solid support primer with optional application of a thermal 'step' to favour solid support priming), Thermal asymmetric interlaced PCR (TAIL-PCR), Touchdown PCR (Step-down PCR), PAN-AC and Universal Fast Walking.
[0131] The PCR (or polymerase chain reaction) technique is well-known in the art and has been disclosed, for example, in K. B. Mullis and F. A. Faloona, Methods Enzymol., 1987, 155: 350-355 and U.S. Pat. Nos. 4,683,202; 4,683,195; and 4,800,159 (each of which is incorporated herein by reference in its entirety). In its simplest form, PCR is an in vitro method for the enzymatic synthesis of specific DNA sequences, using two oligonucleotide primers that hybridize to opposite strands and flank the region of interest in the target DNA. A plurality of reaction cycles, each cycle comprising: a denaturation step, an annealing step, and a polymerization step, results in the exponential accumulation of a specific DNA fragment ("PCR Protocols: A Guide to Methods and Applications", M. A. Innis (Ed.), 1990, Academic Press: New York; "PCR Strategies", M. A. Innis (Ed.), 1995, Academic Press: New York; "Polymerase chain reaction: basic principles and automation in PCR: A Practical Approach", McPherson et al. (Eds.), 1991, IRL Press: Oxford; R. K. Saiki et al., Nature, 1986, 324: 163-166). The termini of the amplified fragments are defined as the 5' ends of the primers. Examples of DNA polymerases capable of producing amplification products in PCR reactions include, but are not limited to: E. coli DNA polymerase I, Klenow fragment of DNA polymerase I, T4 DNA polymerase, thermostable DNA polymerases isolated from Thermus aquaticus (Taq), available from a variety of sources (for example, Perkin Elmer), Thermus thermophilus (United States Biochemicals), Bacillus stereothermophilus (Bio-Rad), or Thermococcus litoralis ("Vent" polymerase, New England Biolabs). RNA target sequences may be amplified by reverse transcribing the mRNA into cDNA, and then performing PCR (RT-PCR), as described above. Alternatively, a single enzyme may be used for both steps as described in U.S. Pat. No. 5,322,770.
[0132] The duration and temperature of each step of a PCR cycle, as well as the number of cycles, are generally adjusted according to the stringency requirements in effect. Annealing temperature and timing are determined both by the efficiency with which a primer is expected to anneal to a template and the degree of mismatch that is to be tolerated. The ability to optimize the reaction cycle conditions is well within the knowledge of one of ordinary skill in the art. Although the number of reaction cycles may vary depending on the detection analysis being performed, it usually is at least 15, more usually at least 20, and may be as high as 60 or higher. However, in many situations, the number of reaction cycles typically ranges from about 20 to about 40.
[0133] The denaturation step of a PCR cycle generally comprises heating the reaction mixture to an elevated temperature and maintaining the mixture at the elevated temperature for a period of time sufficient for any double-stranded or hybridized nucleic acid present in the reaction mixture to dissociate. For denaturation, the temperature of the reaction mixture is usually raised to, and maintained at, a temperature ranging from about 85 °C to about 100 °C, usually from about 90 °C to about 98 °C, and more usually from about 93 °C to about 96 °C, for a period of time ranging from about 3 to about 120 seconds, usually from about 5 to about 30 seconds.
[0134] Following denaturation, the reaction mixture is subjected to conditions sufficient for primer annealing to template DNA present in the mixture. The temperature to which the reaction mixture is lowered to achieve these conditions is usually chosen to provide optimal efficiency and specificity, and generally ranges from about 50 °C to about °C, usually from about 55 °C, to about 70 °C, and more usually from about 60 °C to about 68 °C. Annealing conditions are generally maintained for a period of time ranging from about 15 seconds to about 30 minutes, usually from about 30 seconds to about 5 minutes.
[0135] Following annealing of primer to template DNA or during annealing of primer to template DNA, the reaction mixture is subjected to conditions sufficient to provide for polymerization of nucleotides to the primer's end in a such manner that the primer is extended in a 5' to 3' direction using the DNA to which it is hybridized as a template, (i.e., conditions sufficient for enzymatic production of primer extension product). To achieve primer extension conditions, the temperature of the reaction mixture is typically raised to a temperature ranging from about 65 °C to about 75 °C, usually from about 67 °C. to about 73 °C, and maintained at that temperature for a period of time ranging from about 15 seconds to about 20 minutes, usually from about 30 seconds to about 5 minutes.
[0136] The above cycles of denaturation, annealing, and polymerization may be performed using an automated device typically known as a thermal cycler or thermocycler. Thermal cyclers that may be employed are described in U.S. Pat. Nos. 5,612,473; 5,602,756; 5,538,871; and 5,475,610 (each of which is incorporated herein by reference in its entirety). Thermal cyclers are commercially available, for example, from Perkin Elmer-Applied Biosystems (Norwalk, Conn.), BioRad (Hercules, Calif.), Roche Applied Science (Indianapolis, Ind.), and Stratagene (La Jolla, Calif.).
[0137] Amplification products obtained using primers of the present invention may be detected using agarose gel electrophoresis and visualization by ethidium bromide staining and exposure to ultraviolet (UV) light or by sequence analysis of the amplification product.
[0138] According to one embodiment, the amplification and quantification of the amplification product may be effected in real-time (qRT-PCR). Typically, QRT-PCR methods use double stranded DNA detecting molecules to measure the amount of amplified product in real time.
[0139] As used herein the phrase "double stranded DNA detecting molecule" refers to a double stranded DNA interacting molecule that produces a quantifiable signal (e.g., fluorescent signal). For example such a double stranded DNA detecting molecule can be a fluorescent dye that (1) interacts with a fragment of DNA or an amplicon and (2) emits at a different wavelength in the presence of an amplicon in duplex formation than in the presence of the amplicon in separation. A double stranded DNA detecting molecule can be a double stranded DNA intercalating detecting molecule or a primerbased double stranded DNA detecting molecule.
[0140] A double stranded DNA intercalating detecting molecule is not covalently linked to a primer, an amplicon or a nucleic acid template. The detecting molecule increases its emission in the presence of double stranded DNA and decreases its emission when duplex DNA unwinds. Examples include, but are not limited to, ethidium bromide, YO-PRO-1, Hoechst 33258, SYBR Gold, and SYBR Green I. Ethidium bromide is a fluorescent chemical that intercalates between base pairs in a double stranded DNA fragment and is commonly used to detect DNA following gel electrophoresis. When excited by ultraviolet light between 254 nm and 366 nm, it emits fluorescent light at 590 nm. The DNA-ethidium bromide complex produces about 50 times more fluorescence than ethidium bromide in the presence of single stranded DNA. SYBR Green I is excited at 497 nm and emits at 520 nm. The fluorescence intensity of SYBR Green I increases over 100 fold upon binding to double stranded DNA against single stranded DNA. An alternative to SYBR Green I is SYBR Gold introduced by Molecular Probes Inc. Similar to SYBR Green I, the fluorescence emission of SYBR Gold enhances in the presence of DNA in duplex and decreases when double stranded DNA unwinds. However, SYBR Gold's excitation peak is at 495 nm and the emission peak is at 537 nm. SYBR Gold reportedly appears more stable than SYBR Green I. Hoechst 33258 is a known bisbenzimide double stranded DNA detecting molecule that binds to the AT rich regions of DNA in duplex. Hoechst 33258 excites at 350 nm and emits at 450 nm. YO-PRO-1, exciting at 450 nm and emitting at 550 nm, has been reported to be a double stranded DNA specific detecting molecule. In a particular embodiment of the present invention, the double stranded DNA detecting molecule is SYBR Green I.
[0141] A primer-based double stranded DNA detecting molecule is covalently linked to a primer and either increases or decreases fluorescence emission when amplicons form a duplex structure. Increased fluorescence emission is observed when a primer-based double stranded DNA detecting molecule is attached close to the 3' end of a primer and the primer terminal base is either dG or dC. The detecting molecule is quenched in the proximity of terminal dC-dG and dG-dC base pairs and dequenched as a result of duplex formation of the amplicon when the detecting molecule is located internally at least 6 nucleotides away from the ends of the primer. The dequenching results in a substantial increase in fluorescence emission. Examples of these type of detecting molecules include but are not limited to fluorescein (exciting at 488 nm and emitting at 530 nm), FAM (exciting at 494 nm and emitting at 518 nm), JOE (exciting at 527 and emitting at 548), HEX (exciting at 535 nm and emitting at 556 nm), TET (exciting at 521 nm and emitting at 536 nm), Alexa Fluor 594 (exciting at 590 nm and emitting at 615 nm), ROX (exciting at 575 nm and emitting at 602 nm), and TAMRA (exciting at 555 nm and emitting at 580 nm). In contrast, some primer-based double stranded DNA detecting molecules decrease their emission in the presence of double stranded DNA against single stranded DNA. Examples include, but are not limited to, rhodamine, and BODIPY-FI (exciting at 504 nm and emitting at 513 nm). These detecting molecules are usually covalently conjugated to a primer at the 5' terminal dC or dG and emit less fluorescence when amplicons are in duplex. It is believed that the decrease of fluorescence upon the formation of duplex is due to the quenching of guanosine in the complementary strand in close proximity to the detecting molecule or the quenching of the terminal dC-dG base pairs.
[0142] According to one embodiment, the primer-based double stranded DNA detecting molecule is a 5' nuclease probe. Such probes incorporate a fluorescent reporter molecule at either the 5' or 3' end of an oligonucleotide and a quencher at the opposite end. The first step of the amplification process involves heating to denature the double stranded DNA target molecule into a single stranded DNA. During the second step, a forward primer anneals to the target strand of the DNA and is extended by Taq polymerase. A reverse primer and a 5' nuclease probe then anneal to this newly replicated strand. In this embodiment, at least one of the primer pairs or 5' nuclease probe should hybridize with a unique sequence (of Table 1 or 2). The polymerase extends and cleaves the probe from the target strand. Upon cleavage, the reporter is no longer quenched by its proximity to the quencher and fluorescence is released. Each replication will result in the cleavage of a probe. As a result, the fluorescent signal will increase proportionally to the amount of amplification product.
[0143] DNA, RNA or proteins may be extracted from any plant tissue or cells sample, for example, from leaves, phloem, bark, roots, seeds, and flowers.
[0144] According to a specific embodiment, the sample is a leaf sample (see Examples section which follows).
[0145] As mentioned, according to other embodiments, the genes are targets for intervention to select resistant plants.
[0146] Thus, according to an aspect of the invention there is provided a method of producing Eucalyptus plants exhibiting resistance to a physiological disturbance in Eucalyptus, the method comprising upregulating expression and / or activity in a Eucalyptus of at least one gene of Table 2 and / or downregulating expression and / or activity in a Eucalyptus of at least one gene of Table 1, thereby increasing resistance to a physiological disturbance phenotype in Eucalyptus.
[0147] Methods of modifying gene expression are well known in the art, and are provided infra in more details yet in a non-limiting manner.
[0148] As used herein the term "increasing" refers to at least about 2 %, at least about 3 %, at least about 4 %, at least about 5 %, at least about 10 %, at least about 15 %, at least about 20 %, at least about 30 %, at least about 40 %, at least about 50 %, at least about 60 %, at least about 70 %, at least about 80 %, increase in the trait [e.g., resistance] of a plant as compared to a control plant (a plant which is not modified according to the present teachings and exhibits physiological disturbance susceptibility by nature), such as a native plant, a wild type plant, a non-transformed plant or a non-genomic edited plant of the same species which is grown under the same (e.g., identical) growth conditions.
[0149] The phrase “over-expressing a polypeptide” as used herein refers to increasing the level of the polypeptide within the plant as compared to a control plant of the same species under the same growth conditions.
[0150] According to some embodiments of the invention the increased level of the polypeptide is in a specific cell type or organ of the plant.
[0151] According to some embodiments of the invention, the increased level of the polypeptide is in a temporal time point of the plant. According to some embodiments of the invention, the increased level of the polypeptide is during the whole life cycle of the plant.
[0152] For example, over-expression of a polypeptide can be achieved by elevating the expression level of a native gene of a plant as compared to a control plant. This can be done for example, by means of genome editing which are further described hereinunder, e.g., by introducing mutation(s) in regulatory element(s) (e.g., an enhancer, a promoter, an untranslated region, an intronic region) which result in upregulation of the native gene, and / or by Homology Directed Repair (HDR), e.g., for introducing a “repair template” encoding the polypeptide-of- interest.
[0153] Additionally, and / or alternatively, over-expression of a polypeptide can be achieved by increasing a level of a polypeptide-of-interest due to expression of a heterologous polynucleotide by means of recombinant DNA technology, e.g., using a nucleic acid construct comprising a polynucleotide encoding the polypeptide-of-interest.
[0154] In some embodiments, in case the plant-of-interest (e.g., a plant for which over-expression of a polypeptide is desired) has no detectable expression level of the polypeptide-of-interest prior to employing the method of some embodiments of the invention, qualifying an “over-expression” of the polypeptide in the plant is performed by determination of a positive detectable expression level of the polypeptide-of-interest in a plant cell and / or a plant.
[0155] Additionally and / or alternatively in case the plant-of-interest (e.g., a plant for which overexpression of a polypeptide is desired) has some degree of detectable expression level of the polypeptide-of-interest prior to employing the method of some embodiments of the invention, qualifying an “over-expression” of the polypeptide in the plant is performed by determination of an increased level of expression of the polypeptide-of-interest in a plant cell and / or a plant as compared to a control plant cell and / or plant, respectively, of the same species which is grown under the same (e.g., identical) growth conditions.
[0156] Methods of detecting presence or absence of a polypeptide in a plant cell and / or in a plant, as well as quantification of protein expression levels are well known in the art (e.g., protein detection methods), and are further described hereinunder.
[0157] As used herein the phrase "expressing an exogenous polynucleotide encoding a polypeptide" refers to expression at the mRNA level.
[0158] As used herein the phrase "expressing an exogenous polynucleotide encoding a polypeptide" refers to expression at the mRNA level.
[0159] As used herein, the phrase "exogenous polynucleotide" refers to a heterologous nucleic acid sequence which may not be naturally expressed within the plant (e.g., a nucleic acid sequence from a different species) or which overexpression in the plant is desired. The exogenous polynucleotide may be introduced into the plant in a stable or transient manner, so as to produce a ribonucleic acid (RNA) molecule and / or a polypeptide molecule. It should be noted that the exogenous polynucleotide may comprise a nucleic acid sequence which is identical or partially homologous to an endogenous nucleic acid sequence of the plant.
[0160] The term “endogenous” as used herein refers to any polynucleotide or polypeptide which is present and / or naturally expressed within a plant or a cell thereof.
[0161] According to some embodiments of the invention, the exogenous polynucleotide of the invention comprises a nucleic acid sequence encoding a polypeptide having an amino acid sequence at least about 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, or more say 100 % homologous (e.g., identical) to the amino acid sequence selected from the group consisting of SEQ ID NOs: 60-118.
[0162] Homologous sequences include both orthologous and paralogous sequences. The term “paralogous” relates to gene-duplications within the genome of a species leading to paralogous genes. The term “orthologous” relates to homologous genes in different organisms due to ancestral relationship. Thus, orthologs are evolutionary counterparts derived from a single ancestral gene in the last common ancestor of given two species (Koonin EV and Galperin MY (Sequence - Evolution - Function: Computational Approaches in Comparative Genomics. Boston: Kluwer Academic; 2003. Chapter 2, Evolutionary Concept in Genetics and Genomics. Available from: ncbi(dot)nlm(dot)nih(dot)gov / books / NBK20255) and therefore have great likelihood of having the same function.
[0163] One option to identify orthologues in monocot plant species is by performing a reciprocal blast search. This may be done by a first blast involving blasting the sequence-of-interest against any sequence database, such as the publicly available NCBI database which may be found at: ncbi(dot)nlm(dot)nih(dot)gov. If orthologues in rice were sought, the sequence-of-interest would be blasted against, for example, the 28,469 full-length cDNA clones from Oryza sativa Nipponbare available at NCBI. The blast results may be filtered. The full-length sequences of either the filtered results or the non-filtered results are then blasted back (second blast) against the sequences of the organism from which the sequence-of-interest is derived. The results of the first and second blasts are then compared. An orthologue is identified when the sequence resulting in the highest score (best hit) in the first blast identifies in the second blast the query sequence (the original sequence- of-interest) as the best hit. Using the same rational a paralogue (homolog to a gene in the same organism) is found. In case of large sequence families, the ClustalW program may be used [ebi(dot)ac(dot)uk / Tools / clustalw2 / index(dot)html], followed by a neighbor-joining tree (wikipedia(dot)org / wiki / Neighbor-joining) which helps visualizing the clustering.
[0164] Homology (e.g., percent homology, sequence identity + sequence similarity) can be determined using any homology comparison software computing a pairwise sequence alignment.
[0165] As used herein, "sequence identity" or "identity" in the context of two nucleic acid or polypeptide sequences includes reference to the residues in the two sequences which are the same when aligned. When percentage of sequence identity is used in reference to proteins it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule. Where sequences differ in conservative substitutions, the percent sequence identity may be adjusted upwards to correct for the conservative nature of the substitution. Sequences which differ by such conservative substitutions are considered to have "sequence similarity" or "similarity". Means for making this adjustment are well-known to those of skill in the art. Typically this involves scoring a conservative substitution as a partial rather than a full mismatch, thereby increasing the percentage sequence identity. Thus, for example, where an identical amino acid is given a score of 1 and a non-conservative substitution is given a score of zero, a conservative substitution is given a score between zero and 1. The scoring of conservative substitutions is calculated, e.g., according to the algorithm of Henikoff S and Henikoff JG. [Amino acid substitution matrices from protein blocks. Proc. Natl. Acad. Sci. U.S.A. 1992, 89(22): 10915- 9].
[0166] Identity (e.g., percent homology) can be determined using any homology comparison software, including for example, the BlastN software of the National Center of Biotechnology Information (NCBI) such as by using default parameters.
[0167] According to some embodiments of the invention, the identity is a global identity, z.e., an identity over the entire amino acid or nucleic acid sequences of the invention and not over portions thereof.
[0168] According to some embodiments of the invention, the term “homology” or “homologous” refers to identity of two or more nucleic acid sequences; or identity of two or more amino acid sequences; or the identity of an amino acid sequence to one or more nucleic acid sequence.
[0169] According to some embodiments of the invention, the homology is a global homology, z.e., a homology over the entire amino acid or nucleic acid sequences of the invention and not over portions thereof. The degree of homology or identity between two or more sequences can be determined using various known sequence comparison tools. Following is a non-limiting description of such tools which can be used along with some embodiments of the invention.
[0170] Pairwise global alignment was defined by S. B. Needleman and C. D. Wunsch, "A general method applicable to the search of similarities in the amino acid sequence of two proteins" Journal of Molecular Biology, 1970, pages 443-53, volume 48).
[0171] For example, when starting from a polypeptide sequence and comparing to other polypeptide sequences, the EMBOSS-6.0.1 Needleman-Wunsch algorithm (available from emboss(dot)sourceforge(dot)net / apps / cvs / emboss / apps / needle(dot)html) can be used to find the optimum alignment (including gaps) of two sequences along their entire length - a “Global alignment”. Default parameters for Needleman-Wunsch algorithm (EMBOSS-6.0.1) include: gapopen=10; gapextend=0.5; datafile= EBLOSUM62; brief=YES.
[0172] According to some embodiments of the invention, the parameters used with the EMBOSS- 6.0.1 tool (for protein-protein comparison) include: gapopen=8; gapextend=2; datafile= EBLOSUM62; brief=YES.
[0173] According to some embodiments of the invention, the threshold used to determine homology using the EMBOSS-6.0.1 Needleman-Wunsch algorithm is 80%, 81%, 82 %, 83 %, 84 %, 85 %, 86 %, 87 %, 88 %, 89 %, 90 %, 91 %, 92 %, 93 %, 94 %, 95 %, 96 %, 97 %, 98 %, 99 %, or 100 %.
[0174] When starting from a polypeptide sequence and comparing to polynucleotide sequences, the OneModel FramePlus algorithm [Halperin, E., Faigler, S. and Gill-More, R. (1999) - FramePlus: aligning DNA to protein sequences. Bioinformatics, 15, 867-873) (available from biocceleration(dot)com / Products(dot)html] can be used with following default parameters: model=frame+_p2n.model mode=local.
[0175] According to some embodiments of the invention, the parameters used with the OneModel FramePlus algorithm are model=frame+_p2n.model, mode=qglobal.
[0176] According to some embodiments of the invention, the threshold used to determine homology using the OneModel FramePlus algorithm is 80%, 81%, 82 %, 83 %, 84 %, 85 %, 86 %, 87 %, 88 %, 89 %, 90 %, 91 %, 92 %, 93 %, 94 %, 95 %, 96 %, 97 %, 98 %, 99 %, or 100 %.
[0177] When starting with a polynucleotide sequence and comparing to other polynucleotide sequences the EMBOSS-6.0.1 Needleman-Wunsch algorithm (available from emboss(dot)sourceforge(dot)net / apps / cvs / emboss / apps / needle(dot)html) can be used with the following default parameters: (EMBOSS-6.0.1) gapopen=10; gapextend=0.5; datafile= EDNAFULL; brief=YES. According to some embodiments of the invention, the parameters used with the EMBOSS- 6.0.1 Needleman-Wunsch algorithm are gapopen=10; gapextend=0.2; datafile= EDNAFULL; brief=YES.
[0178] According to some embodiments of the invention, the threshold used to determine homology using the EMBOSS-6.0.1 Needleman-Wunsch algorithm for comparison of polynucleotides with polynucleotides is 80%, 81%, 82 %, 83 %, 84 %, 85 %, 86 %, 87 %, 88 %, 89 %, 90 %, 91 %, 92 %, 93 %, 94 %, 95 %, 96 %, 97 %, 98 %, 99 %, or 100 %.
[0179] According to some embodiment, determination of the degree of homology further requires employing the Smith- Waterman algorithm (for protein-protein comparison or nucleotidenucleotide comparison).
[0180] Default parameters for GenCore 6.0 Smith- Waterman algorithm include: model =sw.model.
[0181] According to some embodiments of the invention, the threshold used to determine homology using the Smith- Waterman algorithm is 80%, 81%, 82 %, 83 %, 84 %, 85 %, 86 %, 87 %, 88 %, 89 %, 90 %, 91 %, 92 %, 93 %, 94 %, 95 %, 96 %, 97 %, 98 %, 99 %, or 100 %.
[0182] According to some embodiments of the invention, the global homology is performed on sequences which are pre-selected by local homology to the polypeptide or polynucleotide of interest (e.g., 60% identity over 60% of the sequence length), prior to performing the global homology to the polypeptide or polynucleotide of interest (e.g., 80% global homology on the entire sequence). For example, homologous sequences are selected using the BLAST software with the Blastp and tBlastn algorithms as filters for the first stage, and the needle (EMBOSS package) or Frarne+ algorithm alignment for the second stage. Local identity (Blast alignments) is defined with a very permissive cutoff - 60% Identity on a span of 60% of the sequences lengths because it is used only as a filter for the global alignment stage. In this specific embodiment (when the local identity is used), the default filtering of the Blast package is not utilized (by setting the parameter “-F F”).
[0183] In the second stage, homologs are defined based on a global identity of at least 80% to the core gene polypeptide sequence.
[0184] According to some embodiments the homology is a local homology or a local identity.
[0185] Local alignments tools include, but are not limited to the BlastP, BlastN, BlastX or TBLASTN software of the National Center of Biotechnology Information (NCBI), FASTA, and the Smith- Waterman algorithm.
[0186] A tblastn search allows the comparison between a protein sequence to the six-frame translations of a nucleotide database. It can be a very productive way of finding homologous protein coding regions in unannotated nucleotide sequences such as expressed sequence tags (ESTs) and draft genome records (HTG), located in the BLAST databases est and htgs, respectively.
[0187] Default parameters for blastp include: Max target sequences: 100; Expected threshold: e“5; Word size: 3; Max matches in a query range: 0; Scoring parameters: Matrix - BLOSUM62; filters and masking: Filter - low complexity regions.
[0188] Local alignments tools, which can be used include, but are not limited to, the tBLASTX algorithm, which compares the six-frame conceptual translation products of a nucleotide query sequence (both strands) against a protein sequence database. Default parameters include: Max target sequences: 100; Expected threshold: 10; Word size: 3; Max matches in a query range: 0; Scoring parameters: Matrix - BLOSUM62; filters and masking: Filter - low complexity regions.
[0189] According to some embodiments of the invention, the exogenous polynucleotide of the invention encodes a polypeptide having an amino acid sequence at least about 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, or more say 100 % identical to the amino acid sequence selected from the group consisting of SEQ ID NOs:60- 118.
[0190] According to some embodiments of the invention, the exogenous polynucleotide of the invention encodes a polypeptide having the amino acid sequence selected from the group consisting of SEQ ID NOs: 60-118.
[0191] According to some embodiments of the invention, the method of increasing resistance of a plant to physiological disturbance, is effected by expressing within the plant an exogenous polynucleotide comprising a nucleic acid sequence encoding a polypeptide at least at least about 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, or more say 100 % identical to the amino acid sequence selected from the group consisting of SEQ ID NOs:60-118, thereby increasing the resistance of the plant.
[0192] According to some embodiments of the invention, the exogenous polynucleotide encodes a polypeptide consisting of the amino acid sequence set forth by SEQ ID NO: 60-118. According to an aspect of some embodiments of the invention, there is provided a method increasing resistance of a plant to physiological disturbance comprising expressing within the plant an exogenous polynucleotide comprising a nucleic acid sequence at least about 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, e.g., 100 % identical to the nucleic acid sequence selected from the group consisting of SEQ ID NOs: 1-59, thereby increasing resistance of a plant to physiological disturbance.
[0193] According to some embodiments of the invention the exogenous polynucleotide is at least about 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, e.g., 100 % identical to the polynucleotide selected from the group consisting of SEQ ID NOs: 1-59.
[0194] According to some embodiments of the invention the exogenous polynucleotide is set forth by SEQ ID NO: 1-59.
[0195] As used herein the term “polynucleotide” refers to a single or double stranded nucleic acid sequence which is isolated and provided in the form of an RNA sequence, a complementary polynucleotide sequence (cDNA), a genomic polynucleotide sequence and / or a composite polynucleotide sequences (e.g., a combination of the above).
[0196] The term “isolated” refers to at least partially separated from the natural environment e.g., from a plant cell.
[0197] As used herein the phrase "complementary polynucleotide sequence" refers to a sequence, which results from reverse transcription of messenger RNA using a reverse transcriptase or any other RNA dependent DNA polymerase. Such a sequence can be subsequently amplified in vivo or in vitro using a DNA dependent DNA polymerase.
[0198] As used herein the phrase "genomic polynucleotide sequence" refers to a sequence derived (isolated) from a chromosome and thus it represents a contiguous portion of a chromosome.
[0199] As used herein the phrase "composite polynucleotide sequence" refers to a sequence, which is at least partially complementary and at least partially genomic. A composite sequence can include some exonal sequences required to encode the polypeptide of the present invention, as well as some intronic sequences interposing therebetween. The intronic sequences can be of any source, including of other genes, and typically will include conserved splicing signal sequences. Such intronic sequences may further include cis acting expression regulatory elements.
[0200] According to some embodiments of the invention, the exogenous polynucleotide encodes a polypeptide consisting of the amino acid sequence set forth by SEQ ID NO: 60-118.
[0201] Nucleic acid sequences encoding the polypeptides of the present invention may be optimized for expression. Examples of such sequence modifications include, but are not limited to, an altered G / C content to more closely approach that typically found in the plant species of interest, and the removal of codons atypically found in the plant species commonly referred to as codon optimization.
[0202] The phrase "codon optimization" refers to the selection of appropriate DNA nucleotides for use within a structural gene or fragment thereof that approaches codon usage within the plant of interest. Therefore, an optimized gene or nucleic acid sequence refers to a gene in which the nucleotide sequence of a native or naturally occurring gene has been modified in order to utilize statistically-preferred or statistically-favored codons within the plant. The nucleotide sequence typically is examined at the DNA level and the coding region optimized for expression in the plant species determined using any suitable procedure, for example as described in Sardana et al. (1996, Plant Cell Reports 15:677-681). In this method, the standard deviation of codon usage, a measure of codon usage bias, may be calculated by first finding the squared proportional deviation of usage of each codon of the native gene relative to that of highly expressed plant genes, followed by a calculation of the average squared deviation. The formula used is: 1 SDCU = n = 1 N [ ( Xn - Yn ) / Yn ] 2 / N, where Xn refers to the frequency of usage of codon n in highly expressed plant genes, where Yn to the frequency of usage of codon n in the gene of interest and N refers to the total number of codons in the gene of interest. A Table of codon usage from highly expressed genes of dicotyledonous plants is compiled using the data of Murray et al. (1989, Nuc Acids Res. 17:477-498).
[0203] One method of optimizing the nucleic acid sequence in accordance with the preferred codon usage for a particular plant cell type is based on the direct use, without performing any extra statistical calculations, of codon optimization Tables such as those provided on-line at the Codon Usage Database through the NIAS (National Institute of Agrobiological Sciences) DNA bank in Japan (kazusa(dot)or(dot)jp / codon / ). The Codon Usage Database contains codon usage tables for a number of different species, with each codon usage Table having been statistically determined based on the data present in Genbank.
[0204] By using the above Tables to determine the most preferred or most favored codons for each amino acid in a particular species (for example, rice), a naturally-occurring nucleotide sequence encoding a protein of interest can be codon optimized for that particular plant species. This is effected by replacing codons that may have a low statistical incidence in the particular species genome with corresponding codons, in regard to an amino acid, that are statistically more favored. However, one or more less-favored codons may be selected to delete existing restriction sites, to create new ones at potentially useful junctions (5' and 3' ends to add signal peptide or termination cassettes, internal sites that might be used to cut and splice segments together to produce a correct full-length sequence), or to eliminate nucleotide sequences that may negatively effect mRNA stability or expression.
[0205] The naturally-occurring encoding nucleotide sequence may already, in advance of any modification, contain a number of codons that correspond to a statistically-favored codon in a particular plant species. Therefore, codon optimization of the native nucleotide sequence may comprise determining which codons, within the native nucleotide sequence, are not statistically- favored with regards to a particular plant, and modifying these codons in accordance with a codon usage table of the particular plant to produce a codon optimized derivative. A modified nucleotide sequence may be fully or partially optimized for plant codon usage provided that the protein encoded by the modified nucleotide sequence is produced at a level higher than the protein encoded by the corresponding naturally occurring or native gene. Construction of synthetic genes by altering the codon usage is described in for example PCT Patent Application 93 / 07278.
[0206] Thus, the invention encompasses nucleic acid sequences described hereinabove; fragments thereof, sequences hybridizable therewith, sequences homologous thereto, sequences encoding similar polypeptides with different codon usage, altered sequences characterized by mutations, such as deletion, insertion or substitution of one or more nucleotides, either naturally occurring or man induced, either randomly or in a targeted fashion.
[0207] According to some embodiments of the invention, the exogenous polynucleotide encodes a polypeptide comprising an amino acid sequence at least 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, e.g., 100 % identical to the amino acid sequence of a naturally occurring plant orthologue of the polypeptide selected from the group consisting of SEQ ID NOs: 60-118.
[0208] According to some embodiments of the invention, the polypeptide comprising an amino acid sequence at least 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, e.g., 100 % identical to the amino acid sequence of a naturally occurring plant orthologue of the polypeptide selected from the group consisting of SEQ ID NOs: 60-118.
[0209] The invention provides an isolated polynucleotide comprising a nucleic acid sequence at least about 80 %, at least about 81 %, at least about 82 %, at least about 83 %, at least about 84 %, at least about 85 %, at least about 86 %, at least about 87 %, at least about 88 %, at least about 89 %, at least about 90 %, at least about 91 %, at least about 92 %, at least about 93 %, at least about 93 %, at least about 94 %, at least about 95 %, at least about 96 %, at least about 97 %, at least about 98 %, at least about 99 %, e.g., 100 % identical to the polynucleotide selected from the group consisting of SEQ ID NOs: 1-59.
[0210] Also provided is a nucleic acid construct comprising a nucleic acid sequence encoding the gene of Table 2 or a functional homolog of same and a heterologous cis-acting regulatory element for driving expression of said gene or functional homolog, such as described herein.
[0211] According to some embodiments of the invention, there is provided a plant cell exogenously expressing the polynucleotide of some embodiments of the invention, the nucleic acid construct of some embodiments of the invention and / or the polypeptide of some embodiments of the invention.
[0212] According to some embodiments of the invention, expressing the exogenous polynucleotide of the invention within the plant is effected by transforming one or more cells of the plant with the exogenous polynucleotide, followed by generating a mature plant from the transformed cells and cultivating the mature plant under conditions suitable for expressing the exogenous polynucleotide within the mature plant.
[0213] According to some embodiments of the invention, the transformation is effected by introducing to the plant cell a nucleic acid construct which includes the exogenous polynucleotide of some embodiments of the invention and at least one promoter for directing transcription of the exogenous polynucleotide in a host cell (a plant cell). Further details of suitable transformation approaches are provided hereinbelow.
[0214] As mentioned, the nucleic acid construct according to some embodiments of the invention comprises a promoter sequence and the isolated polynucleotide of some embodiments of the invention.
[0215] According to some embodiments of the invention, the isolated polynucleotide is operably linked to the promoter sequence. A coding nucleic acid sequence is “operably linked” to a regulatory sequence (e.g., promoter) if the regulatory sequence is capable of exerting a regulatory effect on the coding sequence linked thereto.
[0216] As used herein, the term “promoter” refers to a region of DNA which lies upstream of the transcriptional initiation site of a gene to which RNA polymerase binds to initiate transcription of RNA. The promoter controls where (e.g., which portion of a plant) and / or when (e.g., at which stage or condition in the lifetime of an organism) the gene is expressed.
[0217] According to some embodiments of the invention, the promoter is heterologous to the isolated polynucleotide and / or to the host cell.
[0218] As used herein the phrase “heterologous promoter” refers to a promoter from a different species with respect to the spcies from which the polynucleotide is isolated, or to a promoter from the same species but from a different gene locus within the plant’ s genome with respect to the gene locus from which the polynucleotide sequence is isolated.
[0219] According to some embodiments of the invention, the isolated polynucleotide is heterologous to the plant cell (e.g., the polynucleotide is derived from a different plant species when compared to the plant cell, thus the isolated polynucleotide and the plant cell are not from the same plant species).
[0220] Any suitable promoter sequence can be used by the nucleic acid construct of the present invention. Preferably the promoter is a constitutive promoter, a tissue-specific, or an abiotic stressinducible promoter.
[0221] According to some embodiments of the invention, the promoter is a plant promoter, which is suitable for expression of the exogenous polynucleotide in a plant cell.
[0222] Suitable constitutive promoters include, for example, CaMV 35S promoter CaMV 35S (pQXNc) Promoter); PJJ 35S from Brachypodium; CaMV 35S (OLD) Promoter, Odell et al., Nature 313:810-812, 1985, Arabidopsis At6669 promoter see PCT Publication No. W004081173A2 or the new At6669 promoter; maize Ubl Promoter [cultivar Nongda 105; GenBank: DQ141598.1; Taylor et al., Plant Cell Rep 1993 12: 491-495, which is fully incorporated herein by reference; and cultivar B73; Christensen, AH, et al. Plant Mol. Biol. 18 (4), 675-689 (1992), which is fully incorporated herein by reference] ; rice actin 1 (McElroy et al., Plant Cell 2:163-171, 1990); pEMU (Last et al., Theor. Appl. Genet. 81:581-588, 1991); CaMV 19S (Nilsson et al., Physiol. Plant 100:456-462, 1997); rice GOS2 [(rice GOS2 longer Promoter and GOS2 Promoter), de Pater et al, Plant J Nov;2(6):837-44, 1992] ; RBCS promoter; Rice cyclophilin (Bucholz et al, Plant Mol Biol. 25(5):837-43, 1994); Maize H3 histone (Lepetit et al, Mol. Gen. Genet. 231: 276-285, 1992); Actin 2 (An et al, Plant J. 10( 1); 107- 121 , 1996) and Synthetic Super MAS (Ni et al., The Plant Journal 7: 661-76, 1995). Other constitutive promoters include those in U.S. Pat. Nos. 5,659,026, 5,608,149; 5.608,144; 5,604,121; 5.569,597: 5.466,785; 5,399,680; 5,268,463; and 5,608,142.
[0223] Suitable tissue-specific promoters include, but not limited to, leaf-specific promoters [e.g., AT5G06690 (Thioredoxin) (high expression), AT5G61520 (AtSTP3) (low expression) described in Buttner et al 2000 Plant, Cell and Environment 23, 175-184, or the promoters described in Yamamoto et al., Plant J. 12:255-265, 1997; Kwon et al., Plant Physiol. 105:357-67, 1994; Yamamoto et al., Plant Cell Physiol. 35:773-778, 1994; Gotor et al., Plant J. 3:509-18, 1993; Orozco et al., Plant Mol. Biol. 23:1129-1138, 1993; and Matsuoka et al., Proc. Natl. Acad. Sci. USA 90:9586-9590, 1993; as well as Arabidopsis STP3 (AT5G61520) promoter (Buttner et al., Plant, Cell and Environment 23:175-184, 2000)], seed-preferred promoters [e.g., Napin (originated from Brassica napus which is characterized by a seed specific promoter activity; Stuitje A. R. et. al. Plant Biotechnology Journal 1 (4): 301-309; (Brassica napus NAPIN Promoter) from seed specific genes (Simon, et al., Plant Mol. Biol. 5. 191, 1985; Scofield, et al., J. Biol. Chem. 262: 12202, 1987; Baszczynski, et al., Plant Mol. Biol. 14: 633, 1990), rice PG5a (US 7,700,835), early seed development Arabidopsis BAN (AT1G61720) (US 2009 / 0031450 Al), late seed development Arabidopsis ABI3 (AT3G24650) (Arabidopsis ABI3 (AT3G24650) longer Promoter) or Arabidopsis AB 13 (AT3G24650) Promoter) (Ng et al., Plant Molecular Biology 54: 25-38, 2004), Brazil Nut albumin (Pearson' et al., Plant Mol. Biol. 18: 235- 245, 1992), legumin (Ellis, et al. Plant Mol. Biol. 10: 203-214, 1988), Glutelin (rice) (Takaiwa, et al., Mol. Gen. Genet. 208: 15-22, 1986; Takaiwa, et al., FEBS Letts. 221: 43-47, 1987), Zein (Matzke et al Plant Mol Biol, 143).323-32 1990), napA (Stalberg, et al, Planta 199: 515-519, 1996), Wheat SPA ( Albanietal, Plant Cell, 9: 171- 184, 1997), sunflower oleosin (Cummins, et al., Plant Mol. Biol. 19: 873- 876, 1992)], endosperm specific promoters [Thomas and Flavell, The Plant Cell 2:1171- 1180, 1990; Mol Gen Genet 216:81-90, 1989; NAR 17:461-2), wheat alpha, beta and gamma gliadins (wheat alpha gliadin (B genome) promoter); wheat gamma gliadin promoter; EMBO 3:1409-15, 1984), Barley Itrl promoter, barley B l, C, D hordein (Theor Appl Gen 98:1253-62, 1999; Plant J 4:343-55, 1993; Mol Gen Genet 250:750- 60, 1996), Barley DOF (Mena et al, The Plant Journal, 116(1): 53- 62, 1998), Biz2 (EP99106056.7), Barley SS2 (Barley SS2 Promoter); Guerin and Carbonero Plant Physiology 114: 1 55-62, 1997), wheat Tarp60 (Kovalchuk et al., Plant Mol Biol 71:81-98, 2009), barley D-hordein (D-Hor) and B-hordein (B-Hor) (Agnelo Furtado, Robert J. Henry and Alessandro Pellegrineschi (2009)], Synthetic promoter (Vicente - Carbajosa et al., Plant J. 13: 629-640, 1998), rice prolamin NRP33, rice -globulin Glb-1 (Wu et al, Plant Cell Physiology 39(8) 885- 889, 1998), rice alpha-globulin REB / OHP-1 (Nakase et al. Plant Mol. Biol. 33: 513-S22, 1997), rice ADP-glucose PP (Trans Res 6:157-68, 1997), maize ESR gene family (Plant J 12:235-46, 1997), sorgum gamma- kafirin (PMB 32:1029-35, 1996)], embryo specific promoters [e.g., rice OSHI (Sato et al, Proc. Natl. Acad. Sci. USA, 93: 8117-8122), KNOX (Postma-Haarsma et al, Plant Mol. Biol. 39:257-71, 1999), rice oleosin (Wu et at, J. Biochem., 123:386, 1998)], and flower- specific promoters [e.g., AtPRP4, chalene synthase (chsA) (Van der Meer, et al., Plant Mol. Biol. 15, 95-109, 1990), LAT52 (Twell et al Mol. Gen Genet. 217:240-245; 1989).
[0224] The nucleic acid construct of some embodiments of the invention can further include an appropriate selectable marker and / or an origin of replication. According to some embodiments of the invention, the nucleic acid construct utilized is a shuttle vector, which can propagate both in E. coli (wherein the construct comprises an appropriate selectable marker and origin of replication) and be compatible with propagation in cells. The construct according to the present invention can be, for example, a plasmid, a bacmid, a phagemid, a cosmid, a phage, a virus or an artificial chromosome.
[0225] The nucleic acid construct of some embodiments of the invention can be utilized to stably or transiently transform plant cells. In stable transformation, the exogenous polynucleotide is integrated into the plant genome and as such it represents a stable and inherited trait. In transient transformation, the exogenous polynucleotide is expressed by the cell transformed but it is not integrated into the genome and as such it represents a transient trait.
[0226] There are various methods of introducing foreign genes into both monocotyledonous and dicotyledonous plants (Potrykus, I., Annu. Rev. Plant. Physiol., Plant. Mol. Biol. (1991) 42:205- 225; Shimamoto et al., Nature (1989) 338:274-276).
[0227] The principle methods of causing stable integration of exogenous DNA into plant genomic DNA include two main approaches:
[0228] (i) Agrobacterium-mediated gene transfer: Klee et al. (1987) Annu. Rev. Plant Physiol. 38:467-486; Klee and Rogers in Cell Culture and Somatic Cell Genetics of Plants, Vol. 6, Molecular Biology of Plant Nuclear Genes, eds. Schell, J., and Vasil, L. K., Academic Publishers, San Diego, Calif. (1989) p. 2-25; Gatenby, in Plant Biotechnology, eds. Kung, S. and Amtzen, C. J., Butterworth Publishers, Boston, Mass. (1989) p. 93-112.
[0229] (ii) Direct DNA uptake: Paszkowski et al., in Cell Culture and Somatic Cell Genetics of Plants, Vol. 6, Molecular Biology of Plant Nuclear Genes eds. Schell, J., and Vasil, L. K., Academic Publishers, San Diego, Calif. (1989) p. 52-68; including methods for direct uptake of DNA into protoplasts, Toriyama, K. et al. (1988) Bio / Technology 6:1072-1074. DNA uptake induced by brief electric shock of plant cells: Zhang et al. Plant Cell Rep. (1988) 7:379-384. Fromm et al. Nature (1986) 319:791-793. DNA injection into plant cells or tissues by particle bombardment, Klein et al. Bio / Technology (1988) 6:559-563; McCabe et al. Bio / Technology (1988) 6:923-926; Sanford, Physiol. Plant. (1990) 79:206-209; by the use of micropipette systems: Neuhaus et al., Theor. Appl. Genet. (1987) 75:30-36; Neuhaus and Spangenberg, Physiol. Plant. (1990) 79:213-217; glass fibers or silicon carbide whisker transformation of cell cultures, embryos or callus tissue, U.S. Pat. No. 5,464,765 or by the direct incubation of DNA with germinating pollen, DeWet et al. in Experimental Manipulation of Ovule Tissue, eds. Chapman, G. P. and Mantell, S. H. and Daniels, W. Longman, London, (1985) p. 197-209; and Ohta, Proc. Natl. Acad. Sci. USA (1986) 83:715-719.
[0230] The Agrobacterium system includes the use of plasmid vectors that contain defined DNA segments that integrate into the plant genomic DNA. Methods of inoculation of the plant tissue vary depending upon the plant species and the Agrobacterium delivery system. A widely used approach is the leaf disc procedure which can be performed with any tissue explant that provides a good source for initiation of whole plant differentiation. See, e.g., Horsch et al. in Plant Molecular Biology Manual A5, Kluwer Academic Publishers, Dordrecht (1988) p. 1-9. A supplementary approach employs the Agrobacterium delivery system in combination with vacuum infiltration. The Agrobacterium system is especially viable in the creation of transgenic dicotyledonous plants.
[0231] There are various methods of direct DNA transfer into plant cells. In electroporation, the protoplasts are briefly exposed to a strong electric field. In microinjection, the DNA is mechanically injected directly into the cells using very small micropipettes. In microparticle bombardment, the DNA is adsorbed on microprojectiles such as magnesium sulfate crystals or tungsten particles, and the microprojectiles are physically accelerated into cells or plant tissues.
[0232] Following stable transformation plant propagation is exercised. The most common method of plant propagation is by seed. Regeneration by seed propagation, however, has the deficiency that due to heterozygosity there is a lack of uniformity in the crop, since seeds are produced by plants according to the genetic variances governed by Mendelian rules. Basically, each seed is genetically different and each will grow with its own specific traits. Therefore, it is preferred that the transformed plant be produced such that the regenerated plant has the identical traits and characteristics of the parent transgenic plant. Therefore, it is preferred that the transformed plant be regenerated by micropropagation which provides a rapid, consistent reproduction of the transformed plants.
[0233] Micropropagation is a process of growing new generation plants from a single piece of tissue that has been excised from a selected parent plant or cultivar. This process permits the mass reproduction of plants having the preferred tissue expressing the fusion protein. The new generation plants which are produced are genetically identical to, and have all of the characteristics of, the original plant. Micropropagation allows mass production of quality plant material in a short period of time and offers a rapid multiplication of selected cultivars in the preservation of the characteristics of the original transgenic or transformed plant. The advantages of cloning plants are the speed of plant multiplication and the quality and uniformity of plants produced.
[0234] Micropropagation is a multi-stage procedure that requires alteration of culture medium or growth conditions between stages. Thus, the micropropagation process involves four basic stages: Stage one, initial tissue culturing; stage two, tissue culture multiplication; stage three, differentiation and plant formation; and stage four, greenhouse culturing and hardening. During stage one, initial tissue culturing, the tissue culture is established and certified contaminant- free. During stage two, the initial tissue culture is multiplied until a sufficient number of tissue samples are produced from the seedlings to meet production goals. During stage three, the tissue samples grown in stage two are divided and grown into individual plantlets. At stage four, the transformed plantlets are transferred to a greenhouse for hardening where the plants' tolerance to light is gradually increased so that it can be grown in the natural environment.
[0235] According to some embodiments of the invention, the transgenic plants are generated by transient transformation of leaf cells, meristematic cells or the whole plant.
[0236] Transient transformation can be effected by any of the direct DNA transfer methods described above or by viral infection using modified plant viruses.
[0237] Viruses that have been shown to be useful for the transformation of plant hosts include CaMV, Tobacco mosaic virus (TMV), brome mosaic virus (BMV) and Bean Common Mosaic Virus (BV or BCMV). Transformation of plants using plant viruses is described in U.S. Pat. No. 4,855,237 (bean golden mosaic virus; BGV), EP-A 67,553 (TMV), Japanese Published Application No. 63-14693 (TMV), EPA 194,809 (BV), EPA 278,667 (BV); and Gluzman, Y. et al., Communications in Molecular Biology: Viral Vectors, Cold Spring Harbor Laboratory, New York, pp. 172-189 (1988). Pseudovirus particles for use in expressing foreign DNA in many hosts, including plants are described in WO 87 / 06261.
[0238] According to some embodiments of the invention, the virus used for transient transformations is avirulent and thus is incapable of causing severe symptoms such as reduced growth rate, mosaic, ring spots, leaf roll, yellowing, streaking, pox formation, tumor formation and pitting. A suitable avirulent virus may be a naturally occurring avirulent virus or an artificially attenuated virus. Virus attenuation may be effected by using methods well known in the art including, but not limited to, sub-lethal heating, chemical treatment or by directed mutagenesis techniques such as described, for example, by Kurihara and Watanabe (Molecular Plant Pathology 4:259-269, 2003), Gal-on et al. (1992), Atreya et al. (1992) and Huet et al. (1994).
[0239] Suitable virus strains can be obtained from available sources such as, for example, the American Type culture Collection (ATCC) or by isolation from infected plants. Isolation of viruses from infected plant tissues can be effected by techniques well known in the art such as described, for example by Foster and Taylor, Eds. “Plant Virology Protocols: From Virus Isolation to Transgenic Resistance (Methods in Molecular Biology (Humana Pr), Vol 81)”, Humana Press, 1998. Briefly, tissues of an infected plant believed to contain a high concentration of a suitable virus, preferably young leaves and flower petals, are ground in a buffer solution (e.g., phosphate buffer solution) to produce a virus infected sap which can be used in subsequent inoculations.
[0240] Construction of plant RNA viruses for the introduction and expression of non-viral exogenous polynucleotide sequences in plants is demonstrated by the above references as well as by Dawson, W. O. et al., Virology (1989) 172:285-292; Takamatsu et al. EMBO J. (1987) 6:307- 311; French et al. Science (1986) 231:1294-1297; Takamatsu et al. FEBS Fetters (1990) 269:73- 76; and U.S. Pat. No. 5,316,931.
[0241] When the virus is a DNA virus, suitable modifications can be made to the virus itself. Alternatively, the virus can first be cloned into a bacterial plasmid for ease of constructing the desired viral vector with the foreign DNA. The virus can then be excised from the plasmid. If the virus is a DNA virus, a bacterial origin of replication can be attached to the viral DNA, which is then replicated by the bacteria. Transcription and translation of this DNA will produce the coat protein which will encapsidate the viral DNA. If the virus is an RNA virus, the virus is generally cloned as a cDNA and inserted into a plasmid. The plasmid is then used to make all of the constructions. The RNA virus is then produced by transcribing the viral sequence of the plasmid and translation of the viral genes to produce the coat protein(s) which encapsidate the viral RNA.
[0242] In one embodiment, a plant viral polynucleotide is provided in which the native coat protein coding sequence has been deleted from a viral polynucleotide, a non-native plant viral coat protein coding sequence and a non-native promoter, preferably the subgenomic promoter of the non-native coat protein coding sequence, capable of expression in the plant host, packaging of the recombinant plant viral polynucleotide, and ensuring a systemic infection of the host by the recombinant plant viral polynucleotide, has been inserted. Alternatively, the coat protein gene may be inactivated by insertion of the non-native polynucleotide sequence within it, such that a protein is produced. The recombinant plant viral polynucleotide may contain one or more additional non-native subgenomic promoters. Each non-native subgenomic promoter is capable of transcribing or expressing adjacent genes or polynucleotide sequences in the plant host and incapable of recombination with each other and with native subgenomic promoters. Non-native (foreign) polynucleotide sequences may be inserted adjacent the native plant viral subgenomic promoter or the native and a non-native plant viral subgenomic promoters if more than one polynucleotide sequence is included. The non-native polynucleotide sequences are transcribed or expressed in the host plant under control of the subgenomic promoter to produce the desired products.
[0243] In a second embodiment, a recombinant plant viral polynucleotide is provided as in the first embodiment except that the native coat protein coding sequence is placed adjacent one of the non-native coat protein subgenomic promoters instead of a non-native coat protein coding sequence.
[0244] In a third embodiment, a recombinant plant viral polynucleotide is provided in which the native coat protein gene is adjacent its subgenomic promoter and one or more non-native subgenomic promoters have been inserted into the viral polynucleotide. The inserted non-native subgenomic promoters are capable of transcribing or expressing adjacent genes in a plant host and are incapable of recombination with each other and with native subgenomic promoters. Non- native polynucleotide sequences may be inserted adjacent the non-native subgenomic plant viral promoters such that the sequences are transcribed or expressed in the host plant under control of the subgenomic promoters to produce the desired product.
[0245] In a fourth embodiment, a recombinant plant viral polynucleotide is provided as in the third embodiment except that the native coat protein coding sequence is replaced by a non-native coat protein coding sequence.
[0246] The viral vectors are encapsidated by the coat proteins encoded by the recombinant plant viral polynucleotide to produce a recombinant plant virus. The recombinant plant viral polynucleotide or recombinant plant virus is used to infect appropriate host plants. The recombinant plant viral polynucleotide is capable of replication in the host, systemic spread in the host, and transcription or expression of foreign gene(s) (exogenous polynucleotide) in the host to produce the desired protein.
[0247] Techniques for inoculation of viruses to plants may be found in Foster and Taylor, eds. “Plant Virology Protocols: From Virus Isolation to Transgenic Resistance (Methods in Molecular Biology (Humana Pr), Vol 81)”, Humana Press, 1998; Maramorosh and Koprowski, eds. “Methods in Virology” 7 vols, Academic Press, New York 1967-1984; Hill, S.A. “Methods in Plant Virology”, Blackwell, Oxford, 1984; Walkey, D.G.A. “Applied Plant Virology”, Wiley, New York, 1985; and Kado and Agrawa, eds. “Principles and Techniques in Plant Virology”, Van Nostrand-Reinhold, New York. In addition to the above, the polynucleotide of the present invention can also be introduced into a chloroplast genome thereby enabling chloroplast expression.
[0248] A technique for introducing exogenous polynucleotide sequences to the genome of the chloroplasts is known. This technique involves the following procedures. First, plant cells are chemically treated so as to reduce the number of chloroplasts per cell to about one. Then, the exogenous polynucleotide is introduced via particle bombardment into the cells with the aim of introducing at least one exogenous polynucleotide molecule into the chloroplasts. The exogenous polynucleotides selected such that it is integratable into the chloroplast's genome via homologous recombination which is readily effected by enzymes inherent to the chloroplast. To this end, the exogenous polynucleotide includes, in addition to a gene of interest, at least one polynucleotide stretch which is derived from the chloroplast's genome. In addition, the exogenous polynucleotide includes a selectable marker, which serves by sequential selection procedures to ascertain that all or substantially all of the copies of the chloroplast genomes following such selection will include the exogenous polynucleotide. Further details relating to this technique are found in U.S. Pat. Nos. 4,945,050; and 5,693,507 which are incorporated herein by reference. A polypeptide can thus be produced by the protein expression system of the chloroplast and become integrated into the chloroplast's inner membrane.
[0249] As used herein “downregulation” or a “decrease” refers to at least about 2 %, at least about 3 %, at least about 4 %, at least about 5 %, at least about 10 %, at least about 15 %, at least about 20 %, at least about 30 %, at least about 40 %, at least about 50 %, at least about 60 %, at least about 70 %, at least about 80 %, at least about 2 fold, at least about 3 fold, at least about 5 fold, at least about 10 fold decrease in the level of expression of the gene of Table 1 and / or 2 in a plant or part thereof as compared to a control plant of the invention of the same species which is grown under the same (e.g., identical) growth conditions.
[0250] Downregulation (gene silencing) of the transcription or translation product of an endogenous gene can be achieved by methods which are well known in the art such as genome editing, homologous recombination and co-suppression, antisense suppression, RNA interference and ribozyme molecules.
[0251] Alternatively, there is provided a nucleic acid agent as detailed here below for example which down-regulates expression of at least one gene of Table 1.
[0252] Note that methods of genome editing and manipulation of gene expression at the genomic level are provided herein in the document.
[0253] Co-suppression (sense suppression) - Inhibition of the endogenous gene can be achieved by co- suppression, using an RNA molecule (or an expression vector encoding same) which is in the sense orientation with respect to the transcription direction of the endogenous gene. The polynucleotide used for co-suppression may correspond to all or part of the sequence encoding the endogenous polypeptide and / or to all or part of the 5' and / or 3' untranslated region of the endogenous transcript; it may also be an unpolyadenylated RNA; an RNA which lacks a 5' cap structure; or an RNA which contains an unsplicable intron. In some embodiments, the polynucleotide used for co-suppression is designed to eliminate the start codon of the endogenous polynucleotide so that no protein product will be translated. Methods of co-suppression using a full-length cDNA sequence as well as a partial cDNA sequence are known in the art (see, for example, U.S. Pat. No. 5,231,020).
[0254] According to some embodiments of the invention, downregulation of the endogenous gene is performed using an amplicon expression vector which comprises a plant virus-derived sequence that contains all or part of the target gene but generally not all of the genes of the native virus. The viral sequences present in the transcription product of the expression vector allow the transcription product to direct its own replication. The transcripts produced by the amplicon may be either sense or antisense relative to the target sequence [see for example, Angell and Baulcombe, (1997) EMBO J. 16:3675-3684; Angell and Baulcombe, (1999) Plant J. 20:357-362, and U.S. Pat. No. 6,646,805, each of which is herein incorporated by reference] .
[0255] Antisense suppression - Antisense suppression can be performed using an antisense polynucleotide or an expression vector which is designed to express an RNA molecule complementary to all or part of the messenger RNA (mRNA) encoding the endogenous polypeptide and / or to all or part of the 5' and / or 3' untranslated region of the endogenous gene. Over expression of the antisense RNA molecule can result in reduced expression of the native (endogenous) gene. The antisense polynucleotide may be fully complementary to the target sequence (z.e., 100 % identical to the complement of the target sequence) or partially complementary to the target sequence (z.e., less than 100 % identical, e.g., less than 90 %, less than 80 % identical to the complement of the target sequence). Antisense suppression may be used to inhibit the expression of multiple proteins in the same plant (see e.g., U.S. Pat. No. 5,942,657). In addition, portions of the antisense nucleotides may be used to disrupt the expression of the target gene. Generally, sequences of at least about 50 nucleotides, at least about 100 nucleotides, at least about 200 nucleotides, at least about 300, at least about 400, at least about 450, at least about 500, at least about 550, or greater may be used. Methods of using antisense suppression to inhibit the expression of endogenous genes in plants are described, for example, in Liu, et al., (2002) Plant Physiol. 129:1732-1743 and U.S. Pat. Nos. 5,759,829 and 5,942,657, each of which is herein incorporated by reference. Efficiency of antisense suppression may be increased by including a poly-dT region in the expression cassette at a position 3' to the antisense sequence and 5' of the polyadenylation signal [See, U.S. Patent Publication No. 20020048814, herein incorporated by reference] .
[0256] RNA interference - RNA interference can be achieved using a polynucleotide, which can anneal to itself and form a double stranded RNA having a stem-loop structure (also called hairpin structure), or using two polynucleotides, which form a double stranded RNA.
[0257] For hairpin RNA (hpRNA) interference, the expression vector is designed to express an RNA molecule that hybridizes to itself to form a hairpin structure that comprises a single-stranded loop region and a base-paired stem.
[0258] In some embodiments of the invention, the base-paired stem region of the hpRNA molecule determines the specificity of the RNA interference. In this configuration, the sense sequence of the base-paired stem region may correspond to all or part of the endogenous mRNA to be downregulated, or to a portion of a promoter sequence controlling expression of the endogenous gene to be inhibited; and the antisense sequence of the base-paired stem region is fully or partially complementary to the sense sequence. Such hpRNA molecules are highly efficient at inhibiting the expression of endogenous genes, in a manner which is inherited by subsequent generations of plants [See, e.g., Chuang and Meyerowitz, (2000) Proc. Natl. Acad. Sci. USA 97:4985-4990; Stoutjesdijk, et al., (2002) Plant Physiol. 129:1723-1731; and Waterhouse and Helliwell, (2003) Nat. Rev. Genet. 4:29-38; Chuang and Meyerowitz, (2000) Proc. Natl. Acad. Sci. USA 97:4985-4990; Pandolfini et al., BMC Biotechnology 3:7; Panstruga, et al., (2003) Mol. Biol. Rep. 30:135-140; and U.S. Patent Publication No. 2003 / 0175965; each of which is incorporated by reference] .
[0259] According to some embodiments of the invention, the sense sequence of the base-paired stem is from about 10 nucleotides to about 2,500 nucleotides in length, e.g., from about 10 nucleotides to about 500 nucleotides, e.g., from about 15 nucleotides to about 300 nucleotides, e.g., from about 20 nucleotides to about 100 nucleotides, e.g., or from about 25 nucleotides to about 100 nucleotides.
[0260] According to some embodiments of the invention, the antisense sequence of the basepaired stem may have a length that is shorter, the same as, or longer than the length of the corresponding sense sequence.
[0261] According to some embodiments of the invention, the loop portion of the hpRNA can be from about 10 nucleotides to about 500 nucleotides in length, for example from about 15 nucleotides to about 100 nucleotides, from about 20 nucleotides to about 300 nucleotides or from about 25 nucleotides to about 400 nucleotides in length. According to some embodiments of the invention, the loop portion of the hpRNA can include an intron (ihpRNA), which is capable of being spliced in the host cell. The use of an intron minimizes the size of the loop in the hairpin RNA molecule following splicing and thus increases efficiency of the interference [See, for example, Smith, et al., (2000) Nature 407:319-320; Wesley, et al., (2001) Plant J. 27:581-590; Wang and Waterhouse, (2001) Curr. Opin. Plant Biol. 5:146- 150; Helliwell and Waterhouse, (2003) Methods 30:289-295; Brummell, et al. (2003) Plant J. 33:793-800; and U.S. Patent Publication No. 2003 / 0180945; WO 98 / 53083; WO 99 / 32619; WO 98 / 36083; WO 99 / 53050; US 20040214330; US 20030180945; U.S. Pat. No. 5,034,323; U.S. Pat. No. 6,452,067; U.S. Pat. No. 6,777,588; U.S. Pat. No. 6,573,099 and U.S. Pat. No. 6,326,527; each of which is herein incorporated by reference] .
[0262] In some embodiments of the invention, the loop region of the hairpin RNA determines the specificity of the RNA interference to its target endogenous RNA. In this configuration, the loop sequence corresponds to all or part of the endogenous messenger RNA of the target gene. See, for example, WO 02 / 00904; Mette, et al., (2000) EMBO J 19:5194-5201; Matzke, et al., (2001) Curr. Opin. Genet. Devel. 11:221-227; Scheid, et al., (2002) Proc. Natl. Acad. Sci., USA 99:13659- 13662; Aufsaftz, et al., (2002) Proc. Nat'l. Acad. Sci. 99(4): 16499- 16506; Sijen, et al., Curr. Biol. (2001) 11:436-440), each of which is incorporated herein by reference.
[0263] For double-stranded RNA (dsRNA) interference, the sense and antisense RNA molecules can be expressed in the same cell from a single expression vector (which comprises sequences of both strands) or from two expression vectors (each comprising the sequence of one of the strands). Methods for using dsRNA interference to inhibit the expression of endogenous plant genes are described in Waterhouse, et al., (1998) Proc. Natl. Acad. Sci. USA 95:13959-13964; and WO 99 / 49029, WO 99 / 53050, WO 99 / 61631, and WO 00 / 49035; each of which is herein incorporated by reference.
[0264] According to some embodiments of the invention, RNA interference is effected using an expression vector designed to express an RNA molecule that is modeled on an endogenous micro RNAs (miRNA) gene. Micro RNAs (miRNAs) are regulatory agents consisting of about 22 ribonucleotides and highly efficient at inhibiting the expression of endogenous genes [Javier, et al., (2003) Nature 425:257-263]. The miRNA gene encodes an RNA that forms a hairpin structure containing a 22-nucleotide sequence that is complementary to the endogenous target gene.
[0265] Ribozyme - Catalytic RNA molecules, ribozymes, are designed to cleave particular mRNA transcripts, thus preventing expression of their encoded polypeptides. Ribozymes cleave mRNA at site-specific recognition sequences. For example, “hammerhead ribozymes” (see, for example, U.S. Pat. No. 5,254,678) cleave mRNAs at locations dictated by flanking regions that form complementary base pairs with the target mRNA. The sole requirement is that the target RNA contains a 5'-UG-3' nucleotide sequence. Hammerhead ribozyme sequences can be embedded in a stable RNA such as a transfer RNA (tRNA) to increase cleavage efficiency in vivo [Perriman et al. (1995) Proc. Natl. Acad. Sci. USA, 92(13):6175-6179; de Feyter and Gaudron Methods in Molecular Biology, Vol. 74, Chapter 43, "Expressing Ribozymes in Plants", Edited by Turner, P. C, Humana Press Inc., Totowa, N.J.; U.S. Pat. No. 6,423,885]. RNA endoribonucleases such as that found in Tetrahymena thermophila are also useful ribozymes (U.S. Pat. No. 4,987,071).
[0266] Plant lines transformed with any of the downregulating molecules described hereinabove are screened to identify those that show the greatest inhibition of the endogenous polypeptide-of- interest, and thereby the increase of the desired plant trait (e.g., yield, WUE, NUE, FUE and / or ABST).
[0267] For example, according to some embodiment, upregulating expression is by transgenesis, DNA or RNA editing and / or breeding.
[0268] According to some embodiment, downregulating expression is by transgenesis, DNA or RNA editing, RNA silencing and / or breeding.
[0269] The plant cell transformed with the construct including a plurality of different exogenous polynucleotides, can be regenerated into a mature plant, using the methods described hereinabove.
[0270] Alternatively, expressing a plurality of exogenous polynucleotides in a single host plant can be effected by introducing different nucleic acid constructs, including different exogenous polynucleotides, into a plurality of plants. The regenerated transformed plants can then be crossbred and resultant progeny selected for superior traits, using conventional plant breeding techniques.
[0271] According to some embodiments of the invention, over-expression of the polypeptide of the invention is achieved by means of genome editing.
[0272] Genome editing is a powerful mean to impact target traits by modifications of the target plant genome sequence. Such modifications can result in new or modified alleles or regulatory elements. Thus, genome editing employs reverse genetics by artificially engineered nucleases to cut and create specific double- stranded breaks at a desired location(s) in the genome, which are then repaired by cellular endogenous processes such as, homology directed repair (HDR) and non- homologous end-joining (NHEJ). NHEJ directly joins the DNA ends in a double-stranded break, while HDR utilizes a homologous sequence as a template for regenerating the missing DNA sequence at the break point. In order to introduce specific nucleotide modifications to the genomic DNA, a DNA repair template containing the desired sequence must be present during HDR. Genome editing cannot be performed using traditional restriction endonucleases since most restriction enzymes recognize a few base pairs on the DNA as their target and the probability is very high that the recognized base pair combination will be found in many locations across the genome resulting in multiple cuts not limited to a desired location. To overcome this challenge and create site-specific single- or double- stranded breaks, several distinct classes of nucleases have been discovered and bioengineered to date. These include the meganucleases, Zinc finger nucleases (ZFNs), transcription-activator like effector nucleases (TALENs) and CRISPR / Cas system.
[0273] Since most genome-editing techniques can leave behind minimal traces of DNA alterations evident in a small number of nucleotides as compared to transgenic plants, crops created through gene editing could avoid the stringent regulation procedures commonly associated with genetically modified (GM) crop development. On the other hand, the traces of genome-edited techniques can be used for marker assisted selection (MAS) as is further described hereinunder. Target plants for the mutagenesis / genome editing methods according to the invention are any plants of interest including monocot or dicot plants.
[0274] Over expression of a polypeptide by genome editing can be achieved by: (i) replacing an endogenous sequence encoding the polypeptide of interest or a regulatory sequence under the control which it is placed, and / or (ii) inserting a new gene encoding the polypeptide of interest in a targeted region of the genome, and / or (iii) introducing point mutations which result in upregulation of the gene encoding the polypeptide of interest (e.g., by altering the regulatory sequences such as promoter, enhancers, 5'-UTR and / or 3'-UTR, or mutations in the coding sequence).
[0275] Homology Directed Repair (HDR)
[0276] Homology Directed Repair (HDR) can be used to generate specific nucleotide changes (also known as gene “edits”) ranging from a single nucleotide change to large insertions. In order to utilize HDR for gene editing, a DNA “repair template” containing the desired sequence must be delivered into the cell type of interest with the guide RNA [gRNA(s)] and Cas9 or Cas9 nickase. The repair template must contain the desired edit as well as additional homologous sequence immediately upstream and downstream of the target (termed left and right homology arms). The length and binding position of each homology arm is dependent on the size of the change being introduced. The repair template can be a single stranded oligonucleotide, double-stranded oligonucleotide, or double-stranded DNA plasmid depending on the specific application. It is worth noting that the repair template must lack the Protospacer Adjacent Motif (PAM) sequence that is present in the genomic DNA, otherwise the repair template becomes a suitable target for Cas9 cleavage. For example, the PAM could be mutated such that it is no longer present, but the coding region of the gene is not affected (i.e., a silent mutation).
[0277] The efficiency of HDR is generally low (<10% of modified alleles) even in cells that express Cas9, gRNA and an exogenous repair template. For this reason, many laboratories are attempting to artificially enhance HDR by synchronizing the cells within the cell cycle stage when HDR is most active, or by chemically or genetically inhibiting genes involved in Non-Homologous End Joining (NHEJ). The low efficiency of HDR has several important practical implications. First, since the efficiency of Cas9 cleavage is relatively high and the efficiency of HDR is relatively low, a portion of the Cas9~induced double strand breaks (DSBs) will be repaired via NHEJ. In other words, the resulting population of cells will contain some combination of wildtype alleles, NHEJ-repaired alleles, and / or the desired HDR-edited allele. Therefore, it is important to confinn the presence of the desired edit experimentally, and if necessary, isolate clones containing the desired edit.
[0278] The HDR method was successfully used for targeting a specific modification in a coding sequence of a gene in plants (Budhagatapalli Nagaveni et al. 2015. “Targeted Modification of Gene Function Exploiting Homology-Directed Repair of T ADEN -Mediated Double-Strand Breaks in Barley”. G3 (Bethesda). 2015 Sep; 5(9): 1857-1863). Thus, the g / p-specific transcription activator dike effector nucleases were used along with a repair template that, via HDR, facilitates conversion of gfp into yfp, which is associated with a single amino acid exchange in the gene product. The resulting yellow-fluorescent protein accumulation along with sequencing confirmed the success of the genomic editing.
[0279] Similarly, Zhao Yongping et al. 2016 (An alternative strategy for targeted gene replacement in plants using a dual-sgRNA / Cas9 design. Scientific Reports 6, Article number: 23890 (2016)) describe co-transformation of Arabidopsis plants with a combinatory dual-sgRNA / Cas9 vector that successfully deleted miRNA gene regions (MIR169a and MIR827a) and second construct that contains sites homologous to Arabidopsis TERMINAL FLOWER 1 (TFL1) for homology-directed repair (HDR) with regions corresponding to the two sgRNAs on the modified construct to provide both targeted deletion and donor repair for targeted gene replacement by HDR.
[0280] One example of such approach includes editing a selected genomic region as to express the polypeptide of interest.
[0281] Activation of Target Genes Using CRISPR / Cas9
[0282] Many bacteria and archea contain endogenous RNA-based adaptive immune systems that can degrade nucleic acids of invading phages and plasmids. These systems consist of clustered regularly interspaced short palindromic repeat (CRISPR) genes that produce RNA components and CRISPR associated (Cas) genes that encode protein components. The CRISPR RNAs (crRNAs) contain short stretches of homology to specific viruses and plasmids and act as guides to direct Cas nucleases to degrade the complementary nucleic acids of the corresponding pathogen. Studies of the type II CRISPR / Cas system of Streptococcus pyogenes have shown that three components form an RNA / protein complex and together are sufficient for sequence-specific nuclease activity: the Cas9 nuclease, a crRNA containing 20 base pairs of homology to the target sequence, and a trans-activating crRNA (tracrRNA) (Jinek et al. Science (2012) 337: 816-821.). It was further demonstrated that a synthetic chimeric guide RNA (gRNA) composed of a fusion between crRNA and tracrRNA could direct Cas9 to cleave DNA targets that are complementary to the crRNA in vitro. It was also demonstrated that transient expression of CRISPR- associated endonuclease (Cas9) in conjunction with synthetic gRNAs can be used to produce targeted doublestranded brakes in a variety of different species.
[0283] The CRISPR / Cas9 system is a remarkably flexible tool for genome manipulation. A unique feature of Cas9 is its ability to bind target DN A independently of its ability to cleave target DNA. Specifically, both RuvC- and HNH- nuclease domains can be rendered inactive by point mutations (D10A and H840A in SpCas9), resulting in a nuclease dead Cas9 (dCas9) molecule that cannot cleave target DNA. The dCas9 molecule retains the ability to bind to target DNA based on the gRNA targeting sequence. The dCas9 can be tagged with transcriptional activators, and targeting these dCas9 fusion proteins to the promoter region results in robust transcription activation of downstream target genes. The simplest dCas9-based activators consist of dCas9 fused directly to a single transcriptional activator. Importantly, unlike the genome modifications induced by Cas9 or Cas9 nickase, dCas9-mediated gene activation is reversible, since it does not permanently modify the genomic DNA.
[0284] Indeed, genome editing was successfully used to over-express a protein of interest in a plant by, for example, mutating a regulatory sequence, such as a promoter to overexpress the endogenous polynucleotide operably linked to the regulatory sequence. For example, U.S. Patent Application Publication No. 20160102316 to Rubio Munoz, Vicente et al. which is fully incorporated herein by reference, describes plants with increased expression of an endogenous DDA1 plant nucleic acid sequence wherein the endogenous DDA1 promoter carries a mutation introduced by mutagenesis or genome editing which results in increased expression of the DDA1 gene, using for example, CRISPR. The method involves targeting of Cas9 to the specific genomic locus, in this case DDA1, via a 20-nucleotide guide sequence of the single-guide RNA. An online CRISPR Design Tool can identify suitable target sites (www(dot)tools.genome- engineering(dot)org. Ran et al. Genome engineering using the CRISPR-Cas9 system nature protocols, V0L.8 NO.l l, 2281-2308, 2013).
[0285] The CRISPR-Cas system was used for altering gene expression in plants as described in U.S. Patent Application publication No. 20150067922 to Yang; Yinong et al., which is fully incorporated herein by reference. Thus, the engineered, non-naturally occurring gene editing system comprises two regulatory elements, wherein the first regulatory element (a) operable in a plant cell operably linked to at least one nucleotide sequence encoding a CRISPR-Cas system guide RNA (gRNA) that hybridizes with the target sequence in the plant, and a second regulatory element (b) operable in a plant cell operably linked to a nucleotide sequence encoding a Type-II CRISPR-associated nuclease, wherein components (a) and (b) are located on same or different vectors of the system, whereby the guide RNA targets the target sequence and the CRISPR- associated nuclease cleaves the DNA molecule, thus altering the expression of a gene product in a plant. It should be noted that the CRISPR-associated nuclease and the guide RNA do not naturally occur together.
[0286] In addition, as described above, point mutations which activate a gene-of-interest and / or which result in over-expression of a polypeptide-of-interest can be also introduced into plants by means of genome editing. Such mutation can be for example, deletions of repressor sequences which result in activation of the gene-of-interest; and / or mutations which insert nucleotides and result in activation of regulatory sequences such as promoters and / or enhancers.
[0287] Meganucleases - Meganucleases are commonly grouped into four families: the LAGLIDADG family, the GIY-YIG family, the His-Cys box family and the HNH family. These families are characterized by structural motifs, which affect catalytic activity and recognition sequence. For instance, members of the LAGLIDADG family are characterized by having either one or two copies of the conserved LAGLIDADG motif. The four families of meganucleases are widely separated from one another with respect to conserved structural elements and, consequently, DNA recognition sequence specificity and catalytic activity. Meganucleases are found commonly in microbial species and have the unique property of having very long recognition sequences (>14bp) thus making them naturally very specific for cutting at a desired location. This can be exploited to make site-specific double-stranded breaks in genome editing. One of skill in the art can use these naturally occurring meganucleases, however the number of such naturally occurring meganucleases is limited. To overcome this challenge, mutagenesis and high throughput screening methods have been used to create meganuclease variants that recognize unique sequences. For example, various meganucleases have been fused to create hybrid enzymes that recognize a new sequence. Alternatively, DNA interacting amino acids of the meganuclease can be altered to design sequence specific meganucleases (see e.g., US Patent 8,021,867). Meganucleases can be designed using the methods described in e.g., Certo, MT et al. Nature Methods (2012) 9:073-975; U.S. Patent Nos. 8,304,222; 8,021,867; 8, 119,381; 8, 124,369; 8, 129,134; 8,133,697; 8,143,015; 8,143,016; 8, 148,098; or 8, 163,514, the contents of each are incorporated herein by reference in their entirety. Alternatively, meganucleases with site specific cutting characteristics can be obtained using commercially available technologies e.g., Precision Biosciences' Directed Nuclease Editor™ genome editing technology.
[0288] ZFNs and TALENs - Two distinct classes of engineered nucleases, zinc-finger nucleases (ZFNs) and transcription activator- like effector nucleases (TALENs), have both proven to be effective at producing targeted double- stranded breaks (Christian et al., 2010; Kim et al., 1996; Li et al., 2011; Mahfouz et al., 2011; Miller et al., 2010).
[0289] Basically, ZFNs and TALENs restriction endonuclease technology utilizes a non-specific DNA cutting enzyme which is linked to a specific DNA binding domain (either a series of zinc finger domains or TALE repeats, respectively). Typically, a restriction enzyme whose DNA recognition site and cleaving site are separate from each other is selected. The cleaving portion is separated and then linked to a DNA binding domain, thereby yielding an endonuclease with very high specificity for a desired sequence. An exemplary restriction enzyme with such properties is Fokl. Additionally, Fokl has the advantage of requiring dimerization to have nuclease activity and this means the specificity increases dramatically as each nuclease partner recognizes a unique DNA sequence. To enhance this effect, Fokl nucleases have been engineered that can only function as heterodimers and have increased catalytic activity. The heterodimer functioning nucleases avoid the possibility of unwanted homodimer activity and thus increase specificity of the doublestranded break.
[0290] Thus, for example to target a specific site, ZFNs and TALENs are constructed as nuclease pairs, with each member of the pair designed to bind adjacent sequences at the targeted site. Upon transient expression in cells, the nucleases bind to their target sites and the Fokl domains heterodimerize to create a double-stranded break. Repair of these double- stranded breaks through the nonhomologous end-joining (NHEJ) pathway most often results in small deletions or small sequence insertions. Since each repair made by NHEJ is unique, the use of a single nuclease pair can produce an allelic series with a range of different deletions at the target site. The deletions typically range anywhere from a few base pairs to a few hundred base pairs in length, but larger deletions have successfully been generated in cell culture by using two pairs of nucleases simultaneously (Carlson et al., 2012; Lee et al., 2010). In addition, when a fragment of DNA with homology to the targeted region is introduced in conjunction with the nuclease pair, the double- stranded break can be repaired via homology directed repair to generate specific modifications (Li et al., 2011; Miller et al., 2010; Umov et al., 2005).
[0291] Although the nuclease portions of both ZFNs and TALENs have similar properties, the difference between these engineered nucleases is in their DNA recognition peptide. ZFNs rely on Cys2- His2 zinc fingers and TALENs on TALEs. Both of these DNA recognizing peptide domains have the characteristic that they are naturally found in combinations in their proteins. Cys2-His2 Zinc fingers typically found in repeats that are 3 bp apart and are found in diverse combinations in a variety of nucleic acid interacting proteins. TALEs on the other hand are found in repeats with a one-to-one recognition ratio between the amino acids and the recognized nucleotide pairs. Because both zinc fingers and TALEs happen in repeated patterns, different combinations can be tried to create a wide variety of sequence specificities. Approaches for making site-specific zinc finger endonucleases include, e.g., modular assembly (where Zinc fingers correlated with a triplet sequence are attached in a row to cover the required sequence), OPEN (low- stringency selection of peptide domains vs. triplet nucleotides followed by high- stringency selections of peptide combination vs. the final target in bacterial systems), and bacterial one-hybrid screening of zinc finger libraries, among others. ZFNs can also be designed and obtained commercially from e.g., Sangamo Biosciences™ (Richmond, CA).
[0292] Method for designing and obtaining TALENs are described in e.g., Reyon et al. Nature Biotechnology 2012 May;30(5):460-5; Miller et al. Nat Biotechnol. (2011) 29: 143-148; Cermak et al. Nucleic Acids Research (2011) 39 (12): e82 and Zhang et al. Nature Biotechnology (2011) 29 (2): 149-53. A recently developed web-based program named Mojo Hand was introduced by Mayo Clinic for designing TAL and TALEN constructs for genome editing applications (can be accessed through www(dot)talendesign(dot)org). TALEN can also be designed and obtained commercially from e.g., Sangamo Biosciences™ (Richmond, CA).
[0293] The CRIPSR / Cas system for genome editing contains two distinct components: a gRNA and an endonuclease e.g., Cas9.
[0294] The gRNA is typically a 20-nucleotide sequence encoding a combination of the target homologous sequence (crRNA) and the endogenous bacterial RNA that links the crRNA to the Cas9 nuclease (tracrRNA) in a single chimeric transcript. The gRNA / Cas9 complex is recruited to the target sequence by the base-pairing between the gRNA sequence and the complement genomic DNA. For successful binding of Cas9, the genomic target sequence must also contain the correct Protospacer Adjacent Motif (PAM) sequence immediately following the target sequence. The binding of the gRNA / Cas9 complex localizes the Cas9 to the genomic target sequence so that the Cas9 can cut both strands of the DNA causing a double-strand break. Just as with ZFNs and TALENs, the double-stranded brakes produced by CRISPR / Cas can undergo homologous recombination or NHEJ.
[0295] The Cas9 nuclease has two functional domains: RuvC and HNH, each cutting a different DNA strand. When both of these domains are active, the Cas9 causes double strand breaks in the genomic DNA.
[0296] A significant advantage of CRISPR / Cas is that the high efficiency of this system coupled with the ability to easily create synthetic gRNAs enables multiple genes to be targeted simultaneously. In addition, the majority of cells carrying the mutation present biallelic mutations in the targeted genes.
[0297] However, apparent flexibility in the base-pairing interactions between the gRNA sequence and the genomic DNA target sequence allows imperfect matches to the target sequence to be cut by Cas9.
[0298] Modified versions of the Cas9 enzyme containing a single inactive catalytic domain, either RuvC- or HNH-, are called ‘nickases’. With only one active nuclease domain, the Cas9 nickase cuts only one strand of the target DNA, creating a single-strand break or 'nick'. A single-strand break, or nick, is normally quickly repaired through the HDR pathway, using the intact complementary DNA strand as the template. However, two proximal, opposite strand nicks introduced by a Cas9 nickase are treated as a double-strand break, in what is often referred to as a 'double nick' CRISPR system. A double-nick can be repaired by either NHEJ or HDR depending on the desired effect on the gene target. Thus, if specificity and reduced off-target effects are crucial, using the Cas9 nickase to create a double-nick by designing two gRNAs with target sequences in close proximity and on opposite strands of the genomic DNA would decrease off- target effect as either gRNA alone will result in nicks that will not change the genomic DNA.
[0299] Modified versions of the Cas9 enzyme containing two inactive catalytic domains (dead Cas9, or dCas9) have no nuclease activity while still able to bind to DNA based on gRNA specificity. The dCas9 can be utilized as a platform for DNA transcriptional regulators to activate or repress gene expression by fusing the inactive enzyme to known regulatory domains. For example, the binding of dCas9 alone to a target sequence in genomic DNA can interfere with gene transcription.
[0300] There are a number of publically available tools available to help choose and / or design target sequences as well as lists of bioinformatically determined unique gRNAs for different genes in different species such as the Feng Zhang lab's Target Finder, the Michael Boutros lab's Target Finder (E-CRISP), the RGEN Tools: Cas-OFFinder, the CasFinder: Flexible algorithm for identifying specific Cas9 targets in genomes and the CRISPR Optimal Target Finder. In order to use the CRISPR system, both gRNA and Cas9 should be expressed in a target cell. The insertion vector can contain both cassettes on a single plasmid or the cassettes are expressed from two separate plasmids. CRISPR plasmids are commercially available such as the px33O plasmid from Addgene.
[0301] “Hit and run” or “in-out” - involves a two-step recombination procedure. In the first step, an insertion-type vector containing a dual positive / negative selectable marker cassette is used to introduce the desired sequence alteration. The insertion vector contains a single continuous region of homology to the targeted locus and is modified to carry the mutation of interest. This targeting construct is linearized with a restriction enzyme at a one site within the region of homology, electroporated into the cells, and positive selection is performed to isolate homologous recombinants. These homologous recombinants contain a local duplication that is separated by intervening vector sequence, including the selection cassette. In the second step, targeted clones are subjected to negative selection to identify cells that have lost the selection cassette via intrachromosomal recombination between the duplicated sequences. The local recombination event removes the duplication and, depending on the site of recombination, the allele either retains the introduced mutation or reverts to wild type. The end result is the introduction of the desired modification without the retention of any exogenous sequences.
[0302] The “double-replacement” or “tag and exchange” strategy - involves a two-step selection procedure similar to the hit and run approach, but requires the use of two different targeting constructs. In the first step, a standard targeting vector with 3' and 5' homology arms is used to insert a dual positive / negative selectable cassette near the location where the mutation is to be introduced. After electroporation and positive selection, homologously targeted clones are identified. Next, a second targeting vector that contains a region of homology with the desired mutation is electroporated into targeted clones, and negative selection is applied to remove the selection cassette and introduce the mutation. The final allele contains the desired mutation while eliminating unwanted exogenous sequences.
[0303] Site-Specific Recombinases - The Cre recombinase derived from the Pl bacteriophage and Flp recombinase derived from the yeast Saccharomyces cerevisiae are site-specific DNA recombinases each recognizing a unique 34 base pair DNA sequence (termed “Lox” and “FRT”, respectively) and sequences that are flanked with either Lox sites or FRT sites can be readily removed via site-specific recombination upon expression of Cre or Flp recombinase, respectively. For example, the Lox sequence is composed of an asymmetric eight base pair spacer region flanked by 13 base pair inverted repeats. Cre recombines the 34 base pair lox DNA sequence by binding to the 13 base pair inverted repeats and catalyzing strand cleavage and religation within the spacer region. The staggered DNA cuts made by Cre in the spacer region are separated by 6 base pairs to give an overlap region that acts as a homology sensor to ensure that only recombination sites having the same overlap region recombine.
[0304] Basically, the site-specific recombinase system offers means for the removal of selection cassettes after homologous recombination. This system also allows for the generation of conditional altered alleles that can be inactivated or activated in a temporal or tissue-specific manner. Of note, the Cre and Flp recombinases leave behind a Lox or FRT “scar” of 34 base pairs. The Lox or FRT sites that remain are typically left behind in an intron or 3' UTR of the modified locus, and current evidence suggests that these sites usually do not interfere significantly with gene function.
[0305] Thus, Cre / Lox and Flp / FRT recombination involves introduction of a targeting vector with 3' and 5' homology arms containing the mutation of interest, two Lox or FRT sequences and typically a selectable cassette placed between the two Lox or FRT sequences. Positive selection is applied and homologous recombinants that contain targeted mutation are identified. Transient expression of Cre or Flp in conjunction with negative selection results in the excision of the selection cassette and selects for cells where the cassette has been lost. The final targeted allele contains the Lox or FRT scar of exogenous sequences.
[0306] Transposases - As used herein, the term “transposase” refers to an enzyme that binds to the ends of a transposon and catalyzes the movement of the transposon to another part of the genome.
[0307] As used herein the term “transposon” refers to a mobile genetic element comprising a nucleotide sequence which can move around to different positions within the genome of a single cell. In the process the transposon can cause mutations and / or change the amount of a DNA in the genome of the cell.
[0308] A number of transposon systems that are able to also transpose in cells e.g. vertebrates have been isolated or designed, such as Sleeping Beauty [Izsvak and Ivies Molecular Therapy (2004) 9, 147-156] , piggyBac [Wilson et al. Molecular Therapy (2007) 15, 139-145], Tol2 [Kawakami et al. PNAS (2000) 97 (21): 11403-11408] or Frog Prince [Miskey et al. Nucleic Acids Res. Dec 1, (2003) 31(23): 6873-6881]. Generally, DNA transposons translocate from one DNA site to another in a simple, cut-and-paste manner. Each of these elements has their own advantages, for example, Sleeping Beauty is particularly useful in region- specific mutagenesis, whereas Tol2 has the highest tendency to integrate into expressed genes. Hyperactive systems are available for Sleeping Beauty and piggyBac. Most importantly, these transposons have distinct target site preferences, and can therefore introduce sequence alterations in overlapping, but distinct sets of genes. Therefore, to achieve the best possible coverage of genes, the use of more than one element is particularly preferred. The basic mechanism is shared between the different transposases, therefore we will describe piggyBac (PB) as an example.
[0309] PB is a 2.5 kb insect transposon originally isolated from the cabbage looper moth, Trichoplusia ni. The PB transposon consists of asymmetric terminal repeat sequences that flank a transposase, PBase. PBase recognizes the terminal repeats and induces transposition via a “cut- and-paste” based mechanism, and preferentially transposes into the host genome at the tetranucleotide sequence TTAA. Upon insertion, the TTAA target site is duplicated such that the PB transposon is flanked by this tetranucleotide sequence. When mobilized, PB typically excises itself precisely to reestablish a single TTAA site, thereby restoring the host sequence to its pretransposon state. After excision, PB can transpose into a new location or be permanently lost from the genome.
[0310] Typically, the transposase system offers an alternative means for the removal of selection cassettes after homologous recombination quit similar to the use Cre / Lox or Flp / FRT. Thus, for example, the PB transposase system involves introduction of a targeting vector with 3' and 5' homology arms containing the mutation of interest, two PB terminal repeat sequences at the site of an endogenous TTAA sequence and a selection cassette placed between PB terminal repeat sequences. Positive selection is applied and homologous recombinants that contain targeted mutation are identified. Transient expression of PBase removes in conjunction with negative selection results in the excision of the selection cassette and selects for cells where the cassette has been lost. The final targeted allele contains the introduced mutation with no exogenous sequences.
[0311] For PB to be useful for the introduction of sequence alterations, there must be a native TTAA site in relatively close proximity to the location where a particular mutation is to be inserted.
[0312] Genome editing using recombinant adeno-associated virus (rAAV) platform - this genome-editing platform is based on rAAV vectors which enable insertion, deletion or substitution of DNA sequences in the genomes of live mammalian cells. The rAAV genome is a singlestranded deoxyribonucleic acid (ssDNA) molecule, either positive- or negative- sensed, which is about 4.7 kb long. These single-stranded DNA viral vectors have high transduction rates and have a unique property of stimulating endogenous homologous recombination in the absence of doublestrand DNA breaks in the genome. One of skill in the art can design a rAAV vector to target a desired genomic locus and perform both gross and / or subtle endogenous gene alterations in a cell. rAAV genome editing has the advantage in that it targets a single allele and does not result in any off-target genomic alterations. rAAV genome editing technology is commercially available, for example, the rAAV GENESIS™ system from Horizon™ (Cambridge, UK). Methods for qualifying efficacy and detecting sequence alteration are well known in the art and include, but not limited to, DNA sequencing, electrophoresis, an enzyme-based mismatch detection assay and a hybridization assay such as PCR, RT-PCR, RNase protection, in-situ hybridization, primer extension, Southern blot, Northern Blot and dot blot analysis.
[0313] Sequence alterations in a specific gene can also be determined at the protein level using e.g. chromatography, electrophoretic methods, immunodetection assays such as ELISA and western blot analysis and immunohistochemistry.
[0314] In addition, one ordinarily skilled in the art can readily design a knock-in / knock-out construct including positive and / or negative selection markers for efficiently selecting transformed cells that underwent a homologous recombination event with the construct. Positive selection provides a means to enrich the population of clones that have taken up foreign DNA. Non-limiting examples of such positive markers include glutamine synthetase, dihydrofolate reductase (DHFR), markers that confer antibiotic resistance, such as neomycin, hygromycin, puromycin, and blasticidin S resistance cassettes. Negative selection markers are necessary to select against random integrations and / or elimination of a marker sequence (e.g. positive marker). Non-limiting examples of such negative markers include the herpes simplex-thymidine kinase (HSV-TK) which converts ganciclovir (GCV) into a cytotoxic nucleoside analog, hypoxanthine phosphoribosyltransferase (HPRT) and adenine phosphoribosytransferase (ARPT).
[0315] Once expressed, suppressed, silenced or knocked out within the plant cell or the entire plant, the level of the polypeptide can be determined by methods well known in the art such as, activity assays, Western blots using antibodies capable of specifically binding the polypeptide, Enzyme-Linked Immuno Sorbent Assay (ELISA), radio-immuno-assays (RIA), immunohistochemistry, immunocytochemistry, immunofluorescence and the like.
[0316] Methods of determining the level in the plant of the RNA transcribed from an exogenous or endogebous polynucleotide are well known in the art and include, for example, Northern blot analysis, reverse transcription polymerase chain reaction (RT-PCR) analysis (including quantitative, semi-quantitative or real-time RT-PCR) and RNA-zn situ hybridization.
[0317] The sequence information and annotations uncovered by the present teachings can be harnessed in favor of classical breeding. Thus, sub-sequence data of those polynucleotides described above, can be used as markers for marker assisted selection (MAS), in which a marker is used for indirect selection of a genetic determinant or determinants of a trait of interest (e.g., physiological disturbance). Nucleic acid data of the present teachings (DNA or RNA sequence) may contain or be linked to polymorphic sites or genetic markers on the genome such as restriction fragment length polymorphism (RFLP), microsatellites and single nucleotide polymorphism (SNP), DNA fingerprinting (DFP), amplified fragment length polymorphism (AFLP), expression level polymorphism, polymorphism of the encoded polypeptide and any other polymorphism at the DNA or RNA sequence.
[0318] It will be appreciated that upregulating or downregulating expression can also be achieved by way of breeding, i.e., crossing with a donor plant having the desired trait.
[0319] Thus according to an aspect of the invention there is provided a method of producing Eucalyptus, the method comprising: identifying a physiological disturbance as described herein; and growing a plant exhibiting resistance to the physiological disturbance.
[0320] Such plants can be the subject for further breeding and can be used as donor plants for introducing the trait.
[0321] Plants produced according to the present teachings can be used in commercial settings to produce eucalyptus-derived products such as pulp and paper products, nanocellulose, microfibrilated cellulose, dissolving pulp, fluff pulp, timber, wood-boards, oil, lignin-derived products, Plywood, particleboard, wood-based composites, and charcoal.
[0322] Accordingly, there is provided a method of producing a processed product of a Eucalyptus plant, the method comprising: growing Eucalyptus plants as described herein; and producing pulp, nanocellulose, timber, oil, lignin-derived products, dyes and / or charcoal from the Eucalyptus plants.
[0323] According to a specific embodiment, the processed product comprises DNA of the plant in which the marker(s) can be detected.
[0324] As used herein the term “about” refers to ± 10 %.
[0325] The terms "comprises", "comprising", "includes", "including", “having” and their conjugates mean "including but not limited to".
[0326] The term “consisting of’ means “including and limited to”.
[0327] The term "consisting essentially of" means that the composition, method or structure may include additional ingredients, steps and / or parts, but only if the additional ingredients, steps and / or parts do not materially alter the basic and novel characteristics of the claimed composition, method or structure.
[0328] As used herein, the singular form "a", "an" and "the" include plural references unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" may include a plurality of compounds, including mixtures thereof. Throughout this application, various embodiments of this invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0329] Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integral) within the indicated range. The phrases “ranging / ranges between” a first indicate number and a second indicate number and “ranging / ranges from” a first indicate number “to” a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numerals between.
[0330] As used herein the term "method" refers to manners, means, techniques and procedures for accomplishing a given task including, but not limited to, those manners, means, techniques and procedures either known to, or readily developed from known manners, means, techniques and procedures by practitioners of the chemical, pharmacological, biological, biochemical and medical arts.
[0331] When reference is made to particular sequence listings, such reference is to be understood to also encompass sequences that substantially correspond to its complementary haplotype sequence as including minor sequence variations, resulting from, e.g., sequencing errors, or other alterations resulting in base substitution, base deletion or base addition.
[0332] It is understood that a polynucleotide provided in the sequence listing as a single strand refers to the sense direction which is equivalent to the mRNA transcribed from the polynucleotide.
[0333] It is understood that any Sequence Identification Number (SEQ ID NO) disclosed in the instant application can refer to either a DNA sequence or a RNA sequence, depending on the context where that SEQ ID NO is mentioned, even if that SEQ ID NO is expressed only in a DNA sequence format or a RNA sequence format.
[0334] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment of the invention. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.
[0335] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section below find experimental support in the following examples.
[0336] EXAMPLES
[0337] Reference is now made to the following examples, which together with the above descriptions illustrate some embodiments of the invention in a non-limiting fashion.
[0338] MATERIAL AND METHODS
[0339] RNA purification and quantification:
[0340] Young leaves from one physiological disturbance resistant Eucalyptus parent and one physiological disturbance susceptible Eucalyptus parent, 10 of their progenies, 5 resistant and 5 susceptible, were sampled in a physiological disturbance hot spot region in Bahia, Brazil in a plantation where susceptible clones were displaying severe disturbance symptoms. The samples were collected from 4.5-year-old trees during the onset of the related physiological disturbance symptoms (in this case the disorder is the disturbance). Total RNA was extracted in triplicates from each tested tree, 50 mg leaf tissue per test was subjected to RNA extraction using Plant / Fungi Total RNA Purification kit (Norgen BioTek corp. Cat#258OO) and the isolated RNA was subjected to RNA sequencing. Read data (PE 150, 6 Gb clean data per sample) of next generation sequencing of all 36 samples was quantified and analyzed using Geneious Prime software 2022.2 (www(dot)geneious(dot)com). The reads of each sample were mapped to Eucalyptus grandis V2.0 CDS sequences (Phytozome 13: www(dot)phytozome-next(dot)jgi(dot)doe(dot)gov / ) and quantified using Geneious Prime “Calculate expression Level” module. Transcript per million (TPM) results of each gene was normalized by total transcript count (Wagner et al 2012). The expression levels were compared between resistant and susceptible samples using DESeq2 plugin and detected genes that significantly expressed in either the set of the resistant samples or the set of the susceptible samples, but not in both.
[0341] Wagner, Gunter P., Koryu Kin, and Vincent J. Lynch. "Measurement of mRNA abundance using RNA-seq data: RPKM measure is inconsistent among samples." Theory in biosciences 131.4 (2012): 281-285. EXAMPLE 1
[0342] RNA expression level of 6 resistant and 6 susceptible Eucalyptus trees was compared using read count of mRNA reads mapped to all Eucalyptus CDS sequences. Expression level comparisons were analyzed between the set of resistant clones and the set of susceptible clones, whereby a volcano plot was generated (p-value=0.01) (Figure 1). Genes that showed more than 5- fold difference between the sample sets with a p-value smaller than 0.01 are listed in Table 1.
[0343] Table 1: Significantly differentiated genes in susceptible phenotypes
[0344]
[0345] Table 2: Significantly differentiated genes in resistant phenotypes
[0346]
[0347] EXAMPLE 2
[0348] Phenotypic Markers
[0349] The genes in Table 1 show significantly higher expression levels (more than x5) in susceptible (sensitive) clones grown in physiological disorder affected plantations. Without being bound by theory, it is suggested that the level of expression is increased due to the plant response to the abiotic stress and can be even detected before visual symptoms appear. Therefore, the present inventors have tested their potential as early and / or late indicator / marker for disturbance sensitivity and to detect susceptible plants in the field as early as 3-6 months old and to understand the biological trigger of the disturbance phenomena for identifying the causal variant and genetic manipulations thereof.
[0350] The genes in Table 2 show significantly higher expression levels (more than x5) in resistant plants grown in physiological disorder affected plantations. Without being bound by theory, it is suggested that the level of expression is increased due to the plant’s response to the disturbance stress and might be detected before visual observations of resistant phenotype is verified. Therefore, as above, the present inventors have tested their potential as an early and / or late indicator / marker for disturbance resistance and to detect resistant plants in the field as early as 3- 6 months old.
[0351] Any RNA detection method can be used like:
[0352] 1. Reverse transcription and or realtime PCR / Taqman / Kasps
[0353] 2. Transcriptome next generation sequencing.
[0354] 3. Northern blot 4. Any differential display / expression method.
[0355] EXAMPLE 3
[0356] Transgenic Plants
[0357] The genes in Table 1 show significantly higher expression levels (more than x5) in susceptible trees grown in disturbance affected plantations. These genes can affect the trees and cause, promote or respond to the disease and / or enhance or mitigate the symptoms. Since these genes are not expressed or expressed in relatively low levels in resistant trees, these genes are silenced and / or knocked out to thereby reduce the disturbance symptoms and increase resistance level up to a total immunity.
[0358] Biotechnology transgenic methods that are used include:
[0359] 1. dsRNA
[0360] 2. antisense RNA
[0361] 3. Genome editing (like CRISPR / ZNF / TALENS / Meganucleases)
[0362] 4. EMS / Radiation and select a null mutation.
[0363] 5. Selection from knockout libraries.
[0364] The genes in Table 2 show significantly higher expression levels (more than x5) in resistant (tolerant) trees grown in disturbance affected plantations. These genes affect the trees and cause or promote resistance and / or reduce the symptoms. Since these genes are not expressed or expressed in relatively low levels in susceptible trees, contemplated herein is increasing expression transgenically in susceptible trees to convert disturbance susceptible trees to disturbance resistant trees.
[0365] Biotechnology transgenic promoters that are used are:
[0366] 1. High / medium / low general constitutive promoter
[0367] 2. Xylem specific promoter
[0368] 3. Vessel specific promoter
[0369] 4. Root specific promoter
[0370] 5. Native expression cassette from a resistant tree
[0371] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.
[0372] It is the intent of the Applicant(s) that all publications, patents and patent applications referred to in this specification are to be incorporated in their entirety by reference into the specification, as if each individual publication, patent or patent application was specifically and individually noted when referenced that it is to be incorporated herein by reference. In addition, citation or identification of any reference in this application shall not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not be construed as necessarily limiting. In addition, any priority document(s) of this application is / are hereby incorporated herein by reference in its / their entirety.
Claims
WHAT IS CLAIMED IS:
1. A method of identifying a physiological disturbance phenotype in Eucalyptus, the method comprising determining in a cell or tissue of a Eucalyptus plant a level of expression of at least one gene of Table 1 and / or Table 2, wherein: a statistically significant upregulation in expression of a gene of Table 1 relative to a resistant control is indicative of a susceptible physiological disturbance phenotype; and / or a statistically significant upregulation in expression of a gene of Table 2 relative to a susceptible control is indicative of a resistant physiological disturbance phenotype.
2. A method of producing Eucalyptus plants exhibiting resistance to a physiological disturbance in Eucalyptus, the method comprising upregulating expression and / or activity in a Eucalyptus of at least one gene of Table 2 and / or downregulating expression and / or activity in a Eucalyptus of at least one gene of Table 1, thereby increasing resistance to a physiological disturbance phenotype in Eucalyptus.
3. A method of producing Eucalyptus, the method comprising: identifying a physiological disturbance according to claim 1 ; growing a plant exhibiting resistance to the physiological disturbance.
4. A method of producing a Eucalyptus processed product, the method comprising: producing Eucalyptus plants according to claim 2 or 3; processing the Eucalyptus plants to produce a processed product selected form the group consisting of a paper product, nanocellulose, microfibrilated cellulose, dissolving pulp, fluff pulp, timber, wood-boards, oil, lignin-derived product, plywood, particleboard, wood-based composite, Medium-density fibreboard (MDF), oil, dye and charcoal.
5. The method of any one of claims 1 and 3, wherein said determining is at the RNA level.
6. The method of any one of claims 1 and 3, wherein said determining is at the protein level.
7. The method of claim 2, wherein said upregulating expression is by transgenesis, DNA or RNA editing and / or breeding.
8. The method of claim 2, wherein said downregulating expression is by transgenesis, DNA or RNA editing, RNA silencing and / or breeding.
9. A kit for identifying a physiological disturbance phenotype in Eucalyptus, the kit comprising at least one reagent for identifying expression of at least one gene of Table 1 and / or Table 2, wherein said kit does not detect more than 100, 80, 50 genes.
10. The kit of claim 9, wherein said at least one reagent comprises a primer pair or a probe.
11. A nucleic acid construct comprising a nucleic acid sequence encoding the gene of Table 2 or a functional homolog of same and a heterologous cis-acting regulatory element for driving expression of said gene or functional homolog, such as described herein.
12. A nucleic acid agent which down-regulates expression of at least one gene of Table