Method for selecting and producing eucalyptus plants against physiological disorders

By measuring the expression level of specific genes in Eucalyptus plants, identifying and selecting plants that are resistant to physiological disorders, and improving resistance through gene expression regulation, the problem of identifying and selecting resistant Eucalyptus plants in the prior art has been solved, and efficient and economical improvement of resistance has been achieved.

CN120019154APending Publication Date: 2025-05-16FUTURAGENE ISRAEL LTD +1
View PDF 82 Cites 0 Cited by

Patent Information

Application Number
CN202380071445.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-11
Filing Date
2023-08-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and select Eucalyptus plants that are resistant to physiological disorders, and the identification and selection of resistance clones is an expensive and time-consuming process.

Method used

The susceptible or resistant physiological disorder phenotype is identified by measuring the expression levels of the genes of Table 1 and/or Table 2 in Eucalyptus plants cells or tissues, and the resistance of plants to physiological disorders is increased by upregulating the expression of the genes of Table 2 or downregulating the expression of the genes of Table 1.

Benefits of technology

Early recognition of resistant or susceptible plants is achieved, reducing the time and cost of resistance cloning identification, and improving the level of resistance to physiological disorders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005346472360000681
    Figure BDA0005346472360000681
  • Figure BDA0005346472360000691
    Figure BDA0005346472360000691
  • Figure BDA0005346472360000692
    Figure BDA0005346472360000692
Patent Text Reader

Abstract

A method of identifying the phenotype of a physiological disorder in Eucalyptus comprises determining the expression level of at least one gene in a cell or tissue of an Eucalyptus plant. Compositions and methods for treating physiological disorders are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications:

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 397,000, filed on August 11, 2022, the entire contents of which are incorporated herein by reference.

[0003] Sequence Listing Statement

[0004] An XML file named 97275SequenceListing.xml, submitted concurrently with the filing of this application, was created on August 9, 2023, comprises 192,174 bytes, and is incorporated herein by reference.

[0005] Technical Field and Background Art

[0006] The present invention, in some embodiments thereof, relates to methods of selecting and producing Eucalyptus plants that are resistant to physiological disturbances.

[0007] Eucalyptus "biodystrophy" is a serious abiotic disease occurring in a wide range of areas (e.g., low altitudes) in important eucalyptus producing states in Brazil (including Bahia, Espírito Santo, and Maranhão) and possibly in other parts of the world. The phenomenon, first reported in Brazil in 2005, significantly reduces the productivity of susceptible genotypes when Eucalyptus is grown commercially. Plant wilt, usually progressive death starting at the plant tips of apical buds and leaves, cankers along the main stem, and epicormic sprouting are the main symptoms of biodystrophy. Entire stands of highly susceptible clones may collapse and die at a particular location. Plants of all susceptible genotypes in large plantation stands show the same symptoms in both temporal and spatial distribution. However, the actual trigger of this important abiotic disease remains unknown. An effective solution to the biodystrophy phenomenon is to select for resistant genotypes, as there is great genetic variability between commercial eucalyptus varieties (such as hybrids of E. urophylla and E. grandis). The phenotype of trees grown in the field can vary from highly resistant clones to highly susceptible clones, with a full range of susceptibility levels between these two extremes. This suggests that resistance to physiological disorders has a strong genetic basis, indicating that potential genetic markers can be identified. Although planting resistant hybrid genotypes has become a solution to this problem, the identification and selection of resistant clones is an expensive and time-consuming process that requires the establishment of a large network of field trials in different locations and different seasons. Extensive testing and evaluation must be carried out 3 years or more after planting before the correct level of resistance of Eucalyptus to physiological disorders can be characterized. Until now, the scientific community believes that the genetic control of resistance to physiological disorders in Eucalyptus is complex, with several loci associated with it. Therefore, before adulthood (3 years +), when negative phenotypes are obvious, there are no previous genetic data or genomic markers that can indicate or predict the selection of early resistant genotypes. Summary of the invention

[0008] According to one aspect of the present invention, there is provided a method for identifying a physiological disorder phenotype in Eucalyptus, the method comprising determining the expression level of at least one gene of Table 1 and / or Table 2 in a cell or tissue of a Eucalyptus plant, wherein:

[0009] Statistically significant upregulation of expression of a gene of Table 1 relative to a resistant control indicates a susceptibility to a physiological disorder phenotype; and / or

[0010] Statistically significant upregulation of expression of the genes of Table 2 relative to susceptible controls indicates a resistant physiological disorder phenotype.

[0011] According to one aspect of the present invention, there is provided a method for producing a Eucalyptus plant exhibiting resistance to a physiological disorder in Eucalyptus, the method comprising upregulating the expression and / or activity of at least one gene of Table 2 in Eucalyptus, and / or downregulating the expression and / or activity of at least one gene of Table 1 in Eucalyptus, thereby increasing resistance to the phenotype of a physiological disorder in Eucalyptus.

[0012] According to one aspect of the present invention, there is provided a method for producing Eucalyptus, the method comprising:

[0013] identifying a physiological disorder as described herein;

[0014] Plants are grown that exhibit resistance to the physiological disorder.

[0015] According to one aspect of the present invention, there is provided a method for producing a processed eucalyptus product, the method comprising:

[0016] Producing eucalyptus plants as described herein;

[0017] The eucalyptus plant is processed to produce a processed product selected from the group consisting of: paper products, nanocellulose, microfibrillated cellulose, dissolving pulp, fluff pulp, wood, wood boards, oils, lignin derived products, plywood, particleboard, wood based composites, medium density fiberboard (MDF), oils, dyes and charcoal.

[0018] According to some embodiments of the invention, the determining is at the RNA level.

[0019] According to some embodiments of the invention, the determining is at the protein level.

[0020] According to some embodiments of the invention, up-regulating expression is performed by transgenics, DNA or RNA editing and / or breeding.

[0021] According to some embodiments of the invention, down-regulating expression is performed by transgenics, DNA or RNA editing, RNA silencing and / or breeding.

[0022] According to one aspect of the present invention, there is provided a kit for identifying a physiological disorder phenotype in Eucalyptus, the kit comprising at least one reagent for identifying the expression of at least one gene of Table 1 and / or Table 2, wherein the kit detects no more than 100, 80, 50 genes.

[0023] According to some embodiments of the invention, the at least one reagent comprises a primer pair or a probe.

[0024] According to one aspect of the present invention, a nucleic acid construct is provided, comprising a nucleic acid sequence encoding a gene of Table 2 or a functional homolog thereof and a heterologous cis-acting regulatory element for driving expression of the gene or functional homolog, such as described herein.

[0025] According to one aspect of the present invention, a nucleic acid agent is provided, which down-regulates the expression of at least one gene in Table 1.

[0026] Unless otherwise defined, all technical and / or scientific terms used herein have the same meanings as those of ordinary skill in the art to which the invention belongs. Although methods and materials similar or equivalent to the methods and materials described herein can be used in the practice or testing of embodiments of the present invention, exemplary methods and / or materials are described below. In the event of a conflict, the patent specification (including definitions) shall prevail. In addition, materials, methods and examples are illustrative only and are not necessarily restrictive. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Some embodiments of the present invention are described herein with reference to the accompanying drawings by way of example only. With specific reference now to the detailed drawings, it is emphasized that the details shown are only examples for the purpose of illustrative discussion of the embodiments of the present invention. In this regard, the description in conjunction with the drawings will enable those skilled in the art to clearly understand how the embodiments of the present invention may be implemented.

[0028] In the attached picture:

[0029] Figure 1 It is a volcano plot depicting differential gene expression and RNA levels for a set of 6 physiological disease resistant clones (left side) compared to a set of 6 susceptible clones (right side). The left side depicts differentially expressed genes in resistant clones, and the right side depicts differentially expressed genes in susceptible clones. The selected genes are listed in Tables 1 and 2 below. DETAILED DESCRIPTION

[0030] The present invention, in some embodiments thereof, relates to kits and methods for selecting Eucalyptus species that are resistant to physiological disorders.

[0031] Before explaining at least one embodiment of the present invention in detail, it should be understood that the application of the present invention is not necessarily limited to the details set forth in the following description or exemplified in the examples. The present invention is capable of other embodiments or can be practiced or carried out in various ways.

[0032] Physiological disorders of eucalyptus are important abiotic diseases believed to be the result of environmental triggers in large areas of important eucalyptus producing states in Brazil, such as Bahia and Espírito Santo.

[0033] While conceiving and putting into practice embodiments of the present invention, the inventors have tested differential gene expression (DGE) between six resistant phenotypes and six susceptible phenotypes to physiological disorders in Eucalyptus. The inventors quantified RNA levels and detected genes that were uniquely expressed in either resistant clones or susceptible clones but not in both. Comparisons of expression levels between groups of resistant and susceptible clones were analyzed to generate a volcano plot ( Figure 1 ). The genes in Table 1 show significantly higher expression levels in susceptible (sensitive) clones grown in plantations affected by physiological diseases. Without being bound by theory, it is suggested that expression levels increase due to the plant's response to abiotic stress and can even be detected before visual symptoms appear.

[0034] In contrast, the genes in Table 2 showed significantly higher expression levels in resistant plants grown in plantations affected by physiological disorders. Without being bound by theory, it is suggested that the expression levels increase due to the plant's response to the stress of the disorder and can be detected before the visual observation of the resistance phenotype is verified.

[0035] Thus, embodiments of the present invention relate to the use of these genes as molecular markers to identify resistance and susceptibility phenotypes to abiotic diseases.

[0036] Thus, embodiments of the present invention relate to early markers of plant phenotypic response to physiological disorders, which can help identify potentially resistant plants from susceptible plants at an early stage before visible disorder symptoms appear. Further provided herein are methods for genetically manipulating Eucalyptus to confer resistance to physiological disorders.

[0037] Therefore, according to one aspect of the present invention, there is provided a method for identifying a physiological disorder phenotype in Eucalyptus, the method comprising determining the expression level of at least one gene of Table 1 and / or Table 2 in a cell or tissue of a Eucalyptus plant, wherein:

[0038] Statistically significant upregulation of expression of a gene of Table 1 relative to a resistant control indicates a susceptibility to a physiological disorder phenotype; and / or

[0039] Statistically significant upregulation of expression of the genes of Table 2 relative to susceptible controls indicates a resistant physiological disorder phenotype.

[0040] As used herein, "physiological disorder" refers to an abiotic physiological disease that is believed to be the result of a combination of environmental factors such as precipitation, temperature and atmospheric pressure. This phenomenon is common in some areas of Brazil, such as Bahia and Espirito Santo, but other areas where Eucalyptus is commercially cultivated may also be affected.

[0041] Physiological disturbances manifest themselves in a variety of symptoms, including necrotic lesions on young and mature leaves and branches, loss of apical dominance, as evidenced, for example, by increased branching in the crown, cankers along the main stem, and / or latent bud initiation.

[0042] The severity of the physiological disturbance is characterized gradually, from completely healthy plants without any symptoms in the crown and stem, to severe symptoms in the crown, with curling of apical buds, necrosis of plant tips and dieback, and loss of apical dominance. Severe symptoms are also observed in the main stem of susceptible plants, with depression of the bark, ulcers along the main stem, and the initiation of latent buds.

[0043] According to a specific embodiment, the scoring system is based on visual symptoms on the main stem and general health of the crown. 0 points: no symptoms on the main stem and a healthy crown; 1 point: sparse and mild cankers along the main stem, a healthy crown; 2 points: visible cankers along the main stem, and dieback and necrosis symptoms in the crown; 3 points: severe cankers covering the entire main stem and severe dieback symptoms in the crown. 0 and 1 points are rated as resistant, while 2 and 3 points are considered susceptible.

[0044] As used herein, "resistance" and "improved resistance" are used interchangeably herein and refer to an increase in the frequency of resistant individuals within a given breeding population or any type of resistance to a physiological disorder or symptoms thereof.

[0045] It should be understood that resistance is interchangeable with tolerance.

[0046] Likewise, sensitivity can be substituted for susceptibility. "Resistant plants" or "tolerant plant varieties" have absolute or complete resistance, with no visual symptoms. "Susceptible plants," "susceptible plant varieties," or plants or plant varieties with a "higher level of susceptibility" will show symptoms of dieback, cankers along the main stem, and cankers in branches. Between these two extremes there are varying levels of resistance or susceptibility, with plants showing mild symptoms of the disorder, considered moderately resistant, and some plants showing symptoms of moderate severity, classified as moderately susceptible.

[0047] As used herein, "Eucalyptus" refers to a genus of more than eight hundred species of flowering trees, shrubs, or small eucalypts in the myrtle family (Myrtaceae). Together with several other genera in the Eucalyptus family (including Corymbia), they are commonly referred to as "Eucalypts". The bark of Eucalyptus plants is either smooth, fibrous, hard, or fibrous, the leaves have oil glands, and the sepals and petals are fused to the stamens to form a "cap" or capsule. The fruit is a woody capsule, commonly called a "gumnut".

[0048] Examples of species that can be used in the teachings of the present invention include, but are not limited to:

[0049] Eucalyptus grandis

[0050] Eucalyptus urophylla

[0051] E.pellita

[0052] Eucalyptus robusta

[0053] Eucalyptus cloeziana

[0054] Eucalyptus saligna

[0055] Eucalyptus tereticornis

[0056] Camelidum (Eucalyptus camaldulensis)

[0057] The plants may be plant lines or clones.

[0058] According to another embodiment, the plant is a hybrid plant.

[0059] Hybrids can be intraspecific or interspecific. Hybrids between the same species, such as hybrids between different genotypes of Eucalyptus grandis or other pure species, are called intraspecific hybrids. Hybrids between genotypes of different species, such as hybrids obtained from crossing between Eucalyptus grandis clones and Eucalyptus urophylla clones, are called interspecific hybrids.

[0060] According to a particular embodiment, eucalyptus is the species selected for commercial use.

[0061] According to a specific embodiment, the eucalyptus parent plant and / or the second parent eucalyptus plant are selected from the group consisting of: Eucalyptus grandis, Eucalyptus urophylla, Eucalyptus crassa, Eucalyptus grandis, Eucalyptus camaldulensis, Eucalyptus globulus, Eucalyptus robusta, Eucalyptus willow-leaf, Eucalyptus torelliana or other eucalyptus species. As used herein, the term "plant" refers to whole plant, its organ (i.e. leaf, stem, root, flower etc.), seed, plant cell and offspring thereof. The term "plant cell" includes but is not limited to the cell in seed, suspension culture, embryo, meristem zone, callus, leaf, bud, gametophyte, sporophyte, pollen and microspore. According to a specific embodiment, the plant is a breeding line or clone.

[0062] According to a specific embodiment, the plant strain is a refined strain. The phrase "plant part" refers to a part of a plant, including single cells and cell tissues, such as complete plant cells, cell clumps, and tissue cultures that can regenerate plants in plants. Examples of plant parts include, but are not limited to, single cells and tissues from pollen, ovules, leaves, embryos, roots, root tips, anthers, flowers, fruits, stems, buds, and seeds; as well as scions, rootstocks, protoplasts, callus, etc. According to a specific embodiment, the plant part includes nucleic acid variations as described below. According to a specific embodiment, the plant part is a seed.

[0063] As used herein, the phrase "progeny plant" refers to any plant produced as an offspring by sexual reproduction of one or more parent plants or their offspring. Individual genotypes in the offspring can be propagated by cuttings or any in vitro (= vegetative propagation) method, thereby becoming clones.

[0064] As used herein, "up-regulate" or "increase" refers to an increase in the expression level of a gene of Table 1 and / or 2 in a plant or part thereof by at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 2-fold, at least about 3-fold, at least about 5-fold, at least about 10-fold, compared to a control plant of the same species of the present invention grown under the same (e.g., identical) growth conditions.

[0065] As used herein, "statistically significant" refers to statistical tests known in the art, such as Student's t-test.

[0066] Also provided herein is a kit for identifying a physiological disorder phenotype in Eucalyptus, the kit comprising at least one reagent for identifying the expression of at least one gene of Table 1 and / or Table 2, wherein the kit detects no more than 100, 80, 50, 20 or 10 genes.

[0067] As used herein, "reagent" refers to a substance or molecule used in the field of cell biology, protein biochemistry or molecular biology for specific detection of gene expression. Therefore, the term can refer to primers, probes, primary antibodies, but not buffers, secondary antibodies, etc., which are not specific reagents.

[0068] According to a specific embodiment, at least one reagent comprises a primer pair or a probe or a plurality of primer pairs or probes that can be used in a sequential or simultaneous manner (eg, multiplex).

[0069] For any aspect disclosed herein, the term "measuring" or "measurement", or alternatively "detecting" or "detection" means assessing the presence, absence, amount or quantity of a gene product (mRNA or protein), including deducing the qualitative or quantitative concentration level of such product.

[0070] According to some embodiments, the determining is at the DNA level.

[0071] According to some embodiments, the determining is at the RNA level.

[0072] According to some embodiments, the determining is at the protein level.

[0073] Methods of measuring protein levels are well known in the art and include, for example, immunoassays based on protein antibodies, aptamers, or molecular imprinting.

[0074] The protein may be detected in any suitable manner, but is typically detected by contacting a sample from the plant with an antibody that binds to the protein and then detecting the presence or absence of the reaction product. The antibody may be a monoclonal antibody, a polyclonal antibody, a chimeric antibody, or a fragment of the above, and the step of detecting the reaction product may be performed using any suitable immunoassay.

[0075] In one embodiment, an antibody that specifically binds to a protein is attached (directly or indirectly) to a signal-generating label, including but not limited to a radiolabel, an enzyme label, a hapten, a reporter dye, or a fluorescent label.

[0076] The immunoassays performed according to some embodiments of the present invention can be homogeneous assays or heterogeneous assays. In homogeneous assays, immune responses are generally directed to specific antibodies (e.g., proteins or Table 1 or Table 2), labeled analytes, and samples of interest. When antibodies are combined with labeled analytes, the signal produced by the labeling is changed directly or indirectly. Both immune responses and the detection of their degree can be carried out in homogeneous solutions. Applicable immunochemical labels include free radicals, radioisotopes, fluorescent dyes, enzymes, bacteriophages, or coenzymes.

[0077] In heterogeneous assays, the reagents are typically a sample, an antibody, and a means for producing a detectable signal. The sample may be used as described above. The antibody may be immobilized on a support such as beads (such as protein A and protein G agarose beads), a plate, a pipette tip, or a slide and contacted with a sample suspected of containing a liquid phase antigen.

[0078] The support is then separated from the liquid phase, and the means for producing such a signal is applied to check the detectable signal of the support phase or the liquid phase. This signal is related to the presence of the analyte in the sample. The means for producing a detectable signal include the use of radioactive labels, fluorescent labels or enzyme labels. For example, if the antigen to be detected comprises a second binding site, the antibody bound to the site can be conjugated to a detectable group and added to the liquid phase reaction solution before the separation step. The presence of a detectable group on the solid support indicates the presence of an antigen in the test sample. Examples of suitable immunoassays are oligonucleotides, immunoblotting, immunofluorescence methods, immunoprecipitation, chemiluminescence methods, electrochemiluminescence (ECL) or enzyme-linked immunosorbent assays.

[0079] Those skilled in the art will be familiar with the many specific immunoassay formats and variations thereof that can be used to practice the methods disclosed herein. See generally E. Maggio, Enzyme-Immunoassay, (1980) (CRC Press, Inc., Boca Raton, Fla.); see also U.S. Pat. No. 4,727,022 to Skold et al., entitled “Methods for Modulating Ligand-Receptor Interactions and their Application,” U.S. Pat. No. 4,659,678 to Forrest et al., entitled “Immunoassay of Antigens,” U.S. Pat. No. 4,376,110 to David et al., entitled “Immunometric Assays Using Monoclonal Antibodies,” U.S. Pat. No. 4,275,149 to Litman et al., entitled “Macromolecular Environment Control in Specific Receptor Assays,” U.S. Pat. No. 4,233,402 to Maggio et al., entitled “Reagents and Method Employing Channeling,” and U.S. Pat. No. 4,233,402 to Boguslaski et al., entitled “Heterogenous Specific Binding Assay Employing a Coenzyme as Label" in U.S. Patent No. 4,230,767. The protein can also be detected with antibodies using flow cytometry. Those skilled in the art will be familiar with flow cytometry techniques, which can be used to implement the methods disclosed herein (Shapiro 2005). These include, but are not limited to, cytokine bead arrays (Becton Dickinson) and Luminex technology.

[0080] The antibodies can be conjugated to a solid support suitable for a detection assay (e.g., beads such as magnetic beads, protein A or protein G agarose, microspheres, plates, slides, pipette tips, or wells formed from materials such as latex or polystyrene) according to known techniques, such as passive binding. The antibodies described herein can also be conjugated to a detectable label or group, such as a radiolabel (e.g., 35 S. 125 I. 131 I), enzyme labels (e.g., horseradish peroxidase, alkaline phosphatase), and fluorescent labels (e.g., fluorescein, Alexa, green fluorescent protein, rhodamine).

[0081] In certain embodiments, the antibodies of the invention are monoclonal antibodies.

[0082] The presence of the label can be detected by inspection, or the detection reagent label is detected using a detector that monitors a specific probe or probe combination. Typical detectors include spectrophotometers, phototubes and photodiodes, microscopes, scintillation counters, cameras, films, etc., and combinations thereof. Those skilled in the art will be familiar with a variety of suitable detectors that are widely available from a variety of commercial sources, and can be used to implement the method disclosed herein. Typically, the optical image of the substrate including the combined label portion is digitized for subsequent computer analysis. See generally "Immunoassay Handbook" ["Immunoassay Handbook", third edition, 2005].

[0083] RNA analysis:

[0084] The separation, extraction or derivatization of RNA can be carried out by any suitable method. Isolation of RNA from a biological sample generally includes processing the biological sample in a manner that extracts the RNA present in the sample and makes it available for analysis. Any separation method that produces the extracted RNA can be used in the practice of the present invention. It should be understood that the particular method used to extract RNA will depend on the nature of the source.

[0085] RNA extraction methods are well known in the art and are described further below.

[0086] Phenol-based extraction methods: These single-step RNA isolation methods based on guanidine isothiocyanate (GITC) / phenol / chloroform extraction require much less time than traditional methods (e.g., CsCl2 ultracentrifugation). Many commercial reagents (e.g., Trizol, RNAzol, RNAWIZ) are based on this principle. The entire process can be completed in less than an hour to produce a high yield of total RNA.

[0087] Silica-based purification method: RNeasy is a purification kit sold by Qiagen. It uses a silica-based membrane in a spin column to selectively bind RNA larger than 200 bases. This method is rapid and does not involve the use of phenol.

[0088] Affinity purification of mRNA based on oligo-dT: Due to the low abundance of mRNA in the total cellular RNA pool, reducing the amount of rRNA and tRNA in the total RNA preparation will greatly increase the relative amount of mRNA. The use of oligo-dT (oligo-dT) affinity chromatography to selectively enrich poly (A) + RNA has been practiced for more than 20 years. The result of the preparation is that the enriched mRNA population has minimal rRNA or other small RNA contamination. mRNA enrichment is essential for the construction of cDNA libraries and other applications where complete mRNA is very much needed. The initial method uses oligo-dT conjugated resin column chromatography and can be time-consuming. Recently, more convenient formats have emerged, such as kits based on spin columns and magnetic beads.

[0089] The sample may also be processed before the detection method of the present invention is performed. The sample processing may involve one or more of the following: filtration, distillation, centrifugation, extraction, concentration, dilution, purification, inactivation of interfering components, addition of reagents, etc.

[0090] After obtaining RNA sample, cDNA can be produced therefrom.For the synthesis of cDNA, template mRNA can be directly obtained from the cell of cracking, or can be purified from total RNA or mRNA sample.Total RNA sample can be subjected to force to promote the shearing of RNA molecule, so that the average size of each of RNA molecule is between 100-300 nucleotides, for example, about 200 nucleotides.In order to separate most of RNA found in heterologous population of mRNA from cell, various techniques based on oligo (dT) oligonucleotides attached to solid support can be used.The example of such oligo (dT) oligonucleotides includes: oligo (dT) cellulose / spin column, oligo (dT) / magnetic beads and oligo (dT) oligonucleotide coated plates.

[0091] Generating single-stranded DNA from RNA requires the synthesis of an intermediate RNA-DNA hybrid. To this end, a primer that hybridizes to the 3' end of the RNA is required. The annealing temperature and time are determined by the efficiency of the expected primer annealing to the template and the tolerable mismatch degree.

[0092] The annealing temperature is typically selected to provide optimal efficiency and specificity, and typically ranges from about 50° C. to about 80° C., typically from about 55° C. to about 70° C., and more typically from about 60° C. to about 68° C. Annealing conditions are typically maintained for a period of time ranging from about 15 seconds to about 30 minutes, typically from about 30 seconds to about 5 minutes.

[0093] According to a specific embodiment, the primer comprises a poly dT oligonucleotide sequence.

[0094] Preferably, the poly dT sequence comprises at least 5 nucleotides. According to another, it is between about 5 and 50 nucleotides, more preferably between about 5-25 nucleotides, and even more preferably between about 12 and 14 nucleotides.

[0095] After the primer (e.g., poly dT primer) is annealed to the RNA sample, an RNA-DNA hybrid is synthesized by reverse transcription using an RNA-dependent DNA polymerase. RNA-dependent DNA polymerases suitable for use in the methods and compositions of the present invention include reverse transcriptases (RTs). Examples of RTs include, but are not limited to, Moloney murine leukemia virus (M-MLV) reverse transcriptase, human immunodeficiency virus (HIV) reverse transcriptase, Rous sarcoma virus (RSV) reverse transcriptase, avian myeloblastosis virus (AMV) reverse transcriptase, Rous-associated virus (RAV) reverse transcriptase, and myeloblastosis-associated virus (MAV) reverse transcriptase or other avian sarcoma-leukemia virus (ASLV) reverse transcriptases, and modified RTs derived therefrom. See, e.g., U.S. Patent No. 7,056,716. Many reverse transcriptases, such as those from avian myeloblastic leukemia virus (AMV-RT) and Moloney murine leukemia virus (MMLV-RT), include more than one activity (e.g., polymerase activity and ribonuclease activity) and can function in the formation of double-stranded cDNA molecules.

[0096] Additional components required in the reverse transcription reaction include dNTPS (dATP, dCTP, dGTP and dTTP) and optionally a reducing agent such as dithiothreitol (DTT) and MnCl2.

[0097] Methods for analyzing RNA quantities are known in the art and are summarized below:

[0098] Northern Blot Analysis: This method involves the detection of a specific RNA in a mixture of RNA. The RNA sample is denatured by treatment with an agent that prevents the formation of hydrogen bonds between base pairs, such as formaldehyde, thereby ensuring that all RNA molecules have an unfolded linear conformation. Individual RNA molecules are then separated by size by gel electrophoresis and transferred to a nitrocellulose or nylon-based membrane to which the denatured RNA adheres. The membrane is then exposed to a labeled DNA probe. The probe can be labeled using radioactive isotopes or enzyme-linked nucleotides. Detection can use autoradiography, colorimetric reactions, or chemiluminescence. This method allows both the quantification of the amount of a specific RNA molecule and its identity to be determined by its relative position on the membrane, which indicates the migration distance in the gel during electrophoresis.

[0099] RT-PCR analysis: This method uses PCR amplification of relatively rare RNA molecules. First, RNA molecules are purified from cells and converted into complementary DNA (cDNA) using reverse transcriptase (such as MMLV-RT) and primers (such as oligo dT, random hexamer or gene-specific primers). Then, the PCR amplification reaction is performed in a PCR instrument by using gene-specific primers and Taq DNA polymerase. Those skilled in the art can select the length and sequence of gene-specific primers and PCR conditions (i.e., annealing temperature, number of cycles, etc.) that are suitable for detecting specific RNA molecules. It should be understood that semi-quantitative RT-PCR reactions can be applied by adjusting the number of PCR cycles and comparing the amplified product with a known control. Isothermal amplification is also considered.

[0100] RNA in situ hybridization staining: In this method, DNA or RNA probes are attached to RNA molecules present in cells. Generally speaking, cells are first fixed on microscope slides to retain cell structure and prevent RNA molecules from being degraded, and then placed in a hybridization buffer containing labeled probes. Hybridization buffer includes reagents such as formamide and salts (e.g., sodium chloride and sodium citrate) that allow DNA or RNA probes to hybridize specifically with their target mRNA molecules in situ while avoiding non-specific binding of the probes. Those skilled in the art can adjust hybridization conditions (i.e., temperature, concentration of salt and formamide, etc.) for specific probes and cell types. After hybridization, any unbound probes are washed off and bound probes are detected using known methods. For example, if a radioactively labeled probe is used, the slide is placed in a photographic emulsion, which displays the signal generated using the radioactively labeled probe; if the probe is labeled with an enzyme, an enzyme-specific substrate is added to form a colorimetric reaction; if a fluorescent marker is used to label the probe, a fluorescent microscope is used to display the bound probe; if the probe is labeled with a tag (e.g., digoxigenin, biotin, etc.), the bound probe can be detected after interacting with a tag-specific antibody, which can be detected using known methods.

[0101] In situ RT-PCR staining: This method is described in Nuovo GJ et al. [Intracellular localization of polymerase chain reaction (PCR)-amplified hepatitis C cDNA. Am J Surg Pathol. 1993, 17: 683-90] and Komminoth P et al. [Evaluation of methods for hepatitis C virus detection in archival liver biopsies. Comparison of histology, immunohistochemistry, in situ hybridization, reverse transcriptase polymerase chain reaction (RT-PCR) and in situ RT-PCR. Pathol Res Pract. 1994, 190: 1017-25]. In brief, fixed cells are subjected to RT-PCR reaction by incorporating labeled nucleotides into the PCR reaction. The reaction is performed using a specific in situ RT-PCR device, such as the laser capture microdissection PixCell I LCM system available from Arcturus Engineering (Mountainview, CA).

[0102] DNA microarray / DNA chip:

[0103] DNA microarrays can be used to analyze the expression of thousands of genes simultaneously, thereby analyzing the complete transcriptional program of an organism during a specific developmental process or physiological response. DNA microarrays consist of thousands of individual gene sequences attached to closely spaced areas on the surface of a support such as a glass microscope slide. A variety of methods have been developed to prepare DNA microarrays. In one method, approximately 1 kilobase fragments of the coding region of each gene for analysis are individually PCR amplified. A robotic device is applied to each amplified DNA sample to closely spaced areas on the surface of a glass microscope slide, which is then subjected to heat and chemical treatment to bind the DNA sequences to the surface of the support and denature them. Typically, such an array is about 2x 2cm and contains about 6000 individual nucleic acid spots. In a variant of this technology, multiple DNA oligonucleotides, typically 20 nucleotides in length, are synthesized from initial nucleotides covalently bound to the surface of the support, thereby synthesizing tens of thousands of identical oligonucleotides in small square areas on the surface of the support. Multiple oligonucleotide sequences from a single gene are synthesized in adjacent areas of the slide to analyze the expression of the gene. Therefore, thousands of genes can be presented on one slide. Such arrays of synthetic oligonucleotides may be referred to in the art as "DNA chips" rather than "DNA microarrays" as described above [Lodish et al. (eds.). Chapter 7.8: DNA Microarrays: Analyzing Genome-Wide Expression. In: Molecular Cell Biology, 4th ed., WH Freeman, New York. (2000)].

[0104] Oligonucleotide microarray - In this method, oligonucleotide probes that can specifically hybridize with the polynucleotides of some embodiments of the present invention are attached to a solid surface (e.g., a glass wafer). Each oligonucleotide probe is about 20-25 nucleic acids in length. In order to detect the expression pattern of the polynucleotides of some embodiments of the present invention in a specific cell sample (e.g., leaf cells), RNA is extracted from the cell sample using methods known in the art (using, for example, TRIZOL solution, Gibco BRL, USA). Hybridization can be performed using labeled oligonucleotide probes (e.g., 5'-biotinylated probes) or labeled fragments of complementary DNA (cDNA) or RNA (cRNA). In short, reverse transcriptase (RT) (e.g., Superscript II RT), DNA ligase, and DNA polymerase I are used to prepare double-stranded cDNA from RNA, and all operations are performed according to the manufacturer's instructions (Invitrogen Life Technologies, Frederick, MD, USA). In order to prepare labeled cRNA, double-stranded cDNA is subjected to in vitro transcription reaction in the presence of biotinylated nucleotides using, for example, BioArray high-yield RNA transcription labeling kit (Enzo, Diagnostics, Affymetix Santa Clara CA). For efficient hybridization, the labeled cRNA can be fragmented by incubating the RNA in 40 mM Tris acetate (pH 8.1), 100 mM potassium acetate, and 30 mM magnesium acetate at 94° C. for 35 minutes. After hybridization, the microarray is washed and the hybridization signal is scanned using a confocal laser fluorescence scanner, which measures the fluorescence intensity emitted by the labeled cRNA bound to the probe array.

[0105] For example, in Affymetrix microarrays ( Santa Clara, CA), each gene on the array is represented by a series of different oligonucleotide probes, wherein each probe pair is composed of a perfect match oligonucleotide and a mismatch oligonucleotide. Although the perfect match probe has a sequence that is completely complementary to a specific gene, it is possible to measure the expression level of a specific gene, but the mismatch probe is different from the perfect match probe in that the single base displacement at the central base position. The hybridization signal is scanned using an Agilent scanner, and the Microarray Suite software subtracts the non-specific signal generated by the mismatch probe from the signal generated by the perfect match probe.

[0106] RNA sequencing: Methods for RNA sequence determination are generally known to those skilled in the art. Preferred sequencing methods are next generation sequencing methods or parallel high throughput sequencing methods. An example of a contemplated sequencing method is pyrophosphate sequencing, in particular 454 pyrophosphate sequencing, for example based on the Roche 454 genome sequencer. This method amplifies DNA in water droplets in an oil solution, each droplet containing a single DNA template attached to a bead coated with a single primer, and then forms clonal colonies. Pyrophosphate sequencing uses luciferase to generate light to detect individual nucleotides added to the nascent DNA, and uses the merged data to generate sequence reads. Yet another contemplated example is Illumina or Solexa sequencing, for example by using Illumina genome analyzer technology based on reversible dye terminators. DNA molecules are typically attached to primers on a slide and amplified to form local clonal colonies. Subsequently, one type of nucleotide can be added at a time, and unincorporated nucleotides can be washed away. Subsequently, an image of the fluorescently labeled nucleotides can be taken, and the dye can be chemically removed from the DNA to proceed to the next cycle. Yet another example is the use of Applied Biosystems' SOLiD technology, which employs ligation sequencing. The method is based on the use of all possible fixed-length oligonucleotide libraries, which are labeled according to the sequencing position. Such oligonucleotides are annealed and connected. Subsequently, the preferential connection of the matching sequence by DNA ligase usually results in the signal information of the nucleotide at this position. Since DNA is usually amplified by emulsion PCR, the resulting beads (each bead only contains a copy of the same DNA molecule) can be deposited on a slide, thereby producing a sequence of quantity and length comparable to Illumina sequencing. Another method is the Heliscope technology based on Helicos, in which the fragment is captured by the poly-T oligomer connected to the array. In each sequencing cycle, polymerase and a single fluorescently labeled nucleotide are added and the array is imaged. The fluorescent label is then removed and the cycle is repeated. Other examples of sequencing techniques encompassed in the method of the present invention are hybrid sequencing, sequencing using nanopores, sequencing techniques based on microscopy, microfluidic Sanger sequencing or sequencing methods based on microchips. The present invention also contemplates further developing these technologies, such as further improving the accuracy of sequence determination or the time required for determining the genome sequence of an organism, etc.

[0107] According to one embodiment, the sequencing method comprises deep sequencing.

[0108] As used herein, the term "deep sequencing" refers to a sequencing method that reads a target sequence multiple times in a single test. A single deep sequencing run consists of multiple sequencing reactions run on the same target sequence, and each reaction generates an independent sequence read.

[0109] It should be understood that, in order to analyze the amount of RNA markers, oligonucleotides capable of hybridizing therewith or with the cDNA generated thereby can be used. According to one embodiment, a single oligonucleotide is used to determine the presence of a specific RNA marker, at least two oligonucleotides are used to determine the presence of a specific RNA marker, at least three oligonucleotides are used to determine the presence of a specific RNA marker, at least four oligonucleotides are used to determine the presence of a specific RNA marker, at least five or more oligonucleotides are used to determine the presence of a specific RNA marker.

[0110] When more than one oligonucleotide is used, the sequences of the oligonucleotides can be selected so that they hybridize to the same exon of the RNA marker or to different exons of the RNA marker. In one embodiment, at least one oligonucleotide hybridizes to the 3' exon of the RNA marker. In another embodiment, at least one oligonucleotide hybridizes to the 5' exon of the RNA marker.

[0111] In one embodiment, the method of this aspect of the invention is carried out using an isolated oligonucleotide that hybridizes with the RNA or cDNA of any RNA marker disclosed herein in a sequence-specific manner by complementary base pairing, and distinguishes the gene sequence from other nucleic acid sequences in the sample. Oligonucleotides (e.g., DNA or RNA oligonucleotides) typically include complementary nucleotide sequence regions that hybridize with at least about 8, 10, 13, 16, 18, 20, 22, 25, 30, 40, 50, 55, 60, 65, 70, 80, 90, 100, 120 (or any other number between the two) or more continuous nucleotides in the target nucleic acid molecule under stringent conditions. Depending on the specific determination, the continuous nucleotides include a nucleic acid sequence.

[0112] As used herein, the term "isolated" refers to an oligonucleotide, meaning an oligonucleotide that, due to its origin or manipulation, is separated from at least some of the components with which it is naturally associated or originally obtained. "Isolated" alternatively or additionally means that the oligonucleotide of interest is produced or synthesized by human hand.

[0113] To identify oligonucleotides specific for any of the RNA markers disclosed herein, a computer algorithm is typically used starting from the 5' or 3' end of the nucleotide sequence to examine the gene / transcript of interest. Typical algorithms will then identify oligonucleotides of a specific length that are unique to the gene, have a GC content within a range suitable for hybridization, lack predicted secondary structure that might interfere with hybridization, and / or have other desired characteristics or lack other undesirable characteristics.

[0114] After the oligonucleotide is identified, it can be tested for specificity to the gene (of Table 1 or 2 or its mRNA product) under wet or dry conditions. Thus, for example, where the oligonucleotide is a primer, PCR can be used to test the ability of the primer to amplify the gene sequence (of Table 1 or Table 2 or its mRNA product) to generate a detectable product and to test its ability not to amplify other mRNA products in the sample. The products of the PCR reaction can be analyzed on a gel and verified based on their presence and / or size.

[0115] Additionally or alternatively, the sequence of the oligonucleotide can be analyzed by computer analysis to see if it is homologous to (or capable of hybridizing to) other known sequences. The selected oligonucleotide can be subjected to BLAST 2.2.10 (Basic Local Alignment Search Tool) analysis (www(.)ncbi(.)nlm(.)nih(.)gov / blast / ). The BLAST program finds regions of local similarity between sequences. It compares the nucleotide or protein sequence to a sequence database and calculates the statistical significance of the match, thereby providing valuable information about the possible identity and integrity of the "query" sequence.

[0116] According to one embodiment, the oligonucleotide is a probe. As used herein, the term "probe" refers to an oligonucleotide that hybridizes with a specific nucleic acid sequence of a gene (of Table 1 or 2 or its mRNA product) to provide a detectable signal under experimental conditions and does not hybridize with additional sequences to provide a detectable signal under the same experimental conditions.

[0117] The probes of this embodiment of this aspect of the invention may, for example, be immobilized to a solid support (eg, an array or beads).

[0118] Solid support is a solid substrate or support to which nucleic acid molecules of the present invention can be attached. Nucleic acids can be attached directly or indirectly. The solid substrate for solid support can include any solid material that can be attached directly or indirectly to a component. This includes materials such as acrylamide, agarose, cellulose, nitrocellulose, glass, gold, polystyrene, polyethylene vinyl acetate, polypropylene, polymethacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, Teflon, fluorocarbon, nylon, silicone rubber, polyanhydride, polyglycolic acid, polylactic acid, polyorthoester, functionalized silane, polypropyl fumarate, collagen, glycosaminoglycan and polyamino acid. The solid substrate can have any useful form, including film, membrane, bottle, dish, fiber, braided fiber, molded polymer, particle, bead, microparticle or combination. Solid substrate and solid support can be porous or non-porous. Chip is a small piece of material in a rectangular or square shape. The preferred form of solid substrate is film, bead or chip. The useful form of solid substrate is micro titration dish. In some embodiments, a porous glass slide may be employed.

[0119] In one embodiment, the solid support is an array comprising a plurality of nucleic acids, which are hybridized with the RNA marker of the present invention at an identification or predetermined position fixed on the solid support. Each predefined position on the solid support generally has a type of component (that is, all components at the position are the same). Alternatively, various types of components can be fixed at the same predetermined position on the solid support. Each position has multiple copies of a given component. The spatial separation of different components on the solid support allows for separate detection and identification.

[0120] According to specific embodiments, the array does not include nucleic acids that specifically bind to more than 50 RNA markers, more than 40 RNA markers, 30 RNA markers, 20 RNA markers, 15 RNA markers, 10 RNA markers, 5 RNA markers, or even 3 RNA markers.

[0121] Methods for fixing oligonucleotides to solid substrates are well established. Oligonucleotides, including address probes and detection probes, can be coupled to substrates using established coupling methods. For example, suitable attachment methods are described by Pease et al., Proc. Natl. Acad. Sci. USA 91(11):5022-5026 (1994) and Khrapko et al., Mol Biol (Mosk) (USSR) 25:718-730 (1991). Stimpson et al., Proc. Natl. Acad. Sci. USA 92:6379-6383 (1995) describes methods for fixing 3'-amine oligonucleotides to casein-coated slides. Guo et al., Nucleic Acids Res. 22:5456-5465 (1994) describes a useful method for attaching oligonucleotides to solid substrates.

[0122] According to another embodiment, oligonucleotide is the primer in primer pair.As used herein, term " primer " refers to under appropriate conditions (for example, in the presence of four different nucleotide triphosphates and polymerizing agents such as DNA polymerase, RNA polymerase or reverse transcriptase, DNA ligase, etc., in a suitable buffer solution containing any necessary cofactors, and at one or more temperatures suitable), using methods such as PCR (polymerase chain reaction) or LCR (ligase chain reaction) as the oligonucleotide of the starting point of template-guided synthesis.This template-guided synthesis is also referred to as "primer extension".For example, primer pairs can be designed to use PCR to amplify DNA regions.Such a pair will include "forward primer" and "reverse primer", which hybridize with the complementary strand of DNA molecule and define the region to be synthesized / amplified.The primer of this aspect of the present invention can amplify (for example, by PCR) a specific nucleic acid sequence (of table 1 or 2) to provide a detectable signal under experimental conditions together with its pairing, and it does not amplify other nucleic acid sequences to provide a detectable signal under the same experimental conditions.

[0123] According to additional embodiments, the length of the oligonucleotide is about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides. Although the maximum length of the probe can be as long as the target sequence to be detected, depending on the type of assay adopted, its length is generally less than about 50, 60, 65 or 70 nucleotides. In the case of primers, its length is generally less than about 30 nucleotides. In a specific preferred embodiment of the present invention, the length of the primer or probe is within about 18 to about 28 nucleotides. It should be understood that when attached to a solid support, the length of the probe can be about 30-70, 75, 80, 90, 100 or more nucleotides.

[0124] The oligonucleotides of this aspect of the invention do not need to reflect the exact sequence of the RNA marker nucleic acid sequence (i.e., they do not need to be completely complementary), but must be sufficiently complementary to hybridize with the nucleic acid sequence (of Table 1 or 2) under specific experimental conditions. Therefore, the oligonucleotide sequence is generally at least 70% homologous to the target nucleic acid sequence, for example, over a region of at least 13 or more consecutive nucleotides, preferably at least 80%, 90%, 95%, 97%, 99% or 100% homologous. Conditions are selected so that hybridization of the oligonucleotide with the nucleic acid sequence (of Table 1 or 2) is favorable and hybridization with other nucleic acid sequences is minimized.

[0125] For example, hybridization of short nucleic acids (less than 200 bp in length, e.g., 13-50 bp in length) can be achieved by the following hybridization schemes depending on the desired stringency; (i) a hybridization solution of 6x SSC and 1% SDS or 3M TMACl, 0.01M sodium phosphate (pH 6.8), 1mM EDTA (pH 7.6), 0.5% SDS, 100 μg / ml denatured salmon sperm DNA, and 0.1% skim milk powder, the hybridization temperature is 1-1.5°C below Tm, and the final wash solution is 3M TMACl, 0.01M sodium phosphate (pH 6.8), 1mM EDTA (pH 7.6), 0.5% SDS, 1-1.5°C below Tm (stringent hybridization conditions) (ii) 6x SSC and 0.1% SDS or 3M TMACl, 0.01M sodium phosphate (pH 6.8), 1mM EDTA (pH 7.6), 0.5% Hybridization solution of SDS, 100 μg / ml denatured salmon sperm DNA and 0.1% skim milk powder, the hybridization temperature is below Tm2-2.5°C, the final wash solution is 3M TMACl, 0.01M sodium phosphate (pH6.8), 1mM EDTA (pH7.6), 0.5% SDS, below Tm 1-1.5°C, the final wash solution is 6x SSC, and the final wash is at 22°C (stringent to moderate hybridization conditions); and (iii) a hybridization solution of 6x SSC and 1% SDS or 3M TMACI, 0.01M sodium phosphate (pH6.8), 1mM EDTA (pH7.6), 0.5% SDS, 100 μg / ml denatured salmon sperm DNA and 0.1% skim milk powder, the hybridization temperature is below Tm2.5-3°C, and the final wash solution is 6x SSC at 22°C (moderate hybridization solution).

[0126] Oligonucleotides of the present invention can be prepared by any of a variety of methods (see, e.g., J. Sambrook et al., "Molecular Cloning: A Laboratory Manual", 1989, 2. sup. nd Ed., Cold Spring Harbour Laboratory Press: New York, NY; "PCR Protocols: A Guide to Methods and Applications", 1990, MA Innis (Ed.), Academic Press: New York, NY; P. Tijssen "Hybridization with Nucleic Acid Probes--Laboratory Techniques in Biochemistry and Molecular Biology (Parts I and II)", 1993, Elsevier Science; "PCR Strategies", 1995, MA Innis (Ed.), Academic Press: New York, NY; and "Short Protocols in Molecular Biology", 2002, FM Ausubel (Ed.), 5. sup. th Ed., John Wiley & Sons: Secaucus, NJ). For example, oligonucleotides can be prepared using any of a variety of chemical techniques well known in the art, including, for example, template-based chemical synthesis and polymerization, as described, for example, in: SA Narang et al., Meth. Enzymol. 1979, 68:90-98; EL Brown et al., Meth. Enzymol. 1979, 68:109-151; ES Belousov et al., Nucleic Acids Res. 1997, 25:3440-3444; D. Guschin et al., Anal. Biochem. 1997, 250:203-211; MJ Blommers et al., Biochemistry, 1994, 33:7886-7896; and K. Frenkel et al., Free Radic. Biol. Med. 1995, 19:373-380; and U.S. Pat. No. 4,458,066.

[0127] For example, the automated solid phase procedure based on the phosphoramidite method can be used to prepare oligonucleotides. In this method, each nucleotide is added to the 5' end of the growing oligonucleotide chain separately, and the oligonucleotide chain is attached to the solid support at the 3' end. The added nucleotide is in the form of trivalent 3'-phosphoramidites, which are protected from polymerization by the dimethoxytrityl (or DMT) group at the 5'-position. After the alkali-induced phosphoramidite coupling, gentle oxidation produces a pentavalent phosphotriester intermediate and removes DMT, providing new sites for oligonucleotide extension. The oligonucleotide is then cut off from the solid support, and the phosphodiester and the exocyclic amino group are deprotected with ammonium hydroxide. These synthesis can be carried out on an oligonucleotide synthesizer, such as those commercially available from Perkin Elmer / Applied Biosystems, Inc. (Foster City, Calif.), DuPont (Wilmington, Del.) or Milligen (Bedford, Mass.). Alternatively, oligonucleotides can be custom made and ordered from a variety of commercial sources well known in the art, including, for example, Midland Certified Reagent Company (Midland, Tex.), ExpressGen, Inc. (Chicago, Ill.), Operon Technologies, Inc. (Huntsville, Ala.), and the like.

[0128] When necessary or desired, the purification of the oligonucleotide of the present invention can be carried out by any of the various methods well known in the art. The purification of the oligonucleotide is usually carried out by native acrylamide gel electrophoresis, by anion exchange HPLC (such as, for example, by JD Pearson and F E Regnier (J. Chrom., 1983, 255: 137-149)) or by reversed phase HPLC (G D M C Farland and P N Borer, Nucleic Acids Res., 1979, 7: 1067-1080).

[0129] Any suitable sequencing method can be used to verify the sequence of the oligonucleotide, including but not limited to chemical degradation (AM Maxam and W. Gilbert, Methods of Enzymology, 1980, 65: 499-560), matrix-assisted laser desorption ionization time of flight (MALDI-TOF) mass spectrometry (U. Pieles et al., Nucleic Acids Res., 1993, 21: 3191-3196), combined alkaline phosphatase and exonuclease digestion followed by mass spectrometry (H. Wu and H. Aboleneen, Anal. Biochem., 2001, 290: 347-352), etc.

[0130] As mentioned above, any one of several methods known in the art can be used to prepare the modified oligonucleotide. The non-limiting examples of such modifications include methylation, "capping", replacement of one or more naturally occurring nucleotides with analogs, and internucleotide modification, such as, for example, modification with uncharged connection (e.g., methylphosphonate, phosphotriester, phosphoramidate, carbamate, etc.) or charged connection (e.g., phosphorothioate, phosphorodithioate, etc.). Oligonucleotide can contain one or more additional covalently linked parts, such as, for example, protein (e.g., nuclease, toxin, antibody, signal peptide, poly-L-lysine, etc.), intercalator (e.g., acridine, psoralen, etc.), chelator (e.g., metal, radioactive metal, iron, oxidized metal, etc.) and alkylating agent. Oligonucleotide can also be connected to derivatize by forming methyl or ethyl phosphotriester or alkylphosphoramidate. In addition, oligonucleotide sequence of the present invention can also be modified with labeling.

[0131] In some embodiments, the detection probe or amplification primer or both the probe and the primer are labeled with a detectable agent or a portion thereof before being used for amplification / detection assays. In some embodiments, the detection probe is labeled with a detectable agent. Preferably, the detectable agent is selected so that it generates a measurable signal, and its intensity is related to (e.g., proportional to) the amount of the amplification product in the sample being analyzed.

[0132] The binding between oligonucleotide and detectable agent can be covalent or non-covalent. The detection probe of labeling can be prepared by incorporating or conjugating detectable part. Labeling can be directly or indirectly (for example, through joint) attached to nucleic acid sequence. Joints or spacer arms of various lengths are known in the art and are commercially available, and can be selected to reduce steric hindrance, or give other useful or desired characteristics of obtained labeling molecule (see, for example, ES Mansfield et al., Mol. Cell. Probes, 1995, 9: 145-156).

[0133] Methods for labeling nucleic acid molecules are well known in the art. For a review of labeling protocols, label detection techniques, and the latest developments in this field, see, for example, LJ Kricka, Ann. Clin. Biochem. 2002, 39: 114-129; RP van Gijlswijk et al., Expert Rev. Mol. Diagn. 2001, 1: 81-91; and S. Joos et al., J. Biotechnol. 1994, 35: 135-153. Standard nucleic acid labeling methods include: incorporation of radioactive agents, direct attachment of fluorescent dyes (L. M. Smith et al., Nucl. Acids Res., 1985, 13: 2399-2412) or direct attachment of enzymes (B. A. Connoly and O. Rider, Nucl. Acids Res., 1985, 13: 4485-4502); chemical modification of nucleic acid molecules so that they can be detected by immunochemistry or by other affinity reactions (T. R. Broker et al., Nucl. Acids Res. 1978, 5: 363-384; E. A. Bayer et al., Methods of Biochem. Analysis, 1980, 26: 1-45; R. Langer et al., Proc. Natl. Acad. Sci. USA, 1981, 78: 6633-6637; R. W. Richardson et al., Nucl. Acids Res. 1983, 11: 6167-6184; DJ Brigati et al., Virol. 1983, 126: 32-50; P. Tchen et al., Proc. Natl. Acad. Sci. USA, 1984, 81: 3466-3470; JE Landegent et al., Exp. Cell Res. 1984, 15: 61-72; and AH Hopman et al., Exp. Cell Res. 1987, 169: 357-368); and enzyme-mediated labeling methods, such as random priming, nick translation, PCR, and terminal transferase tailing (for a review of enzyme labeling, see, e.g., J. Temsamani and S. Agrawal, Mol. Biotechnol. 1996, 5: 223-232).Recently developed nucleic acid labeling systems include, but are not limited to, ULS (Universal Ligation System), which is based on the reaction of a monoreactive cisplatin derivative with the N7 position of the guanine moiety in DNA (R. J. Heetebrij et al., Cytogenet. Cell. Genet. 1999, 87:47-52), psoralen-biotin, which intercalates into nucleic acids and covalently binds to nucleotide bases under UV irradiation (C. Levenson et al., Methods Enzymol. 1990, 184:577-583; and C. Pfannschmidt et al., Nucleic Acids Res. 1996, 24:1702-1709), photoreactive azide derivatives (C. Neves et al., Bioconjugate Chem. 2000, 11:51-55), and DNA alkylating agents (MG Sebestyen et al., Bioconjugate Chem. 2001, 11:51-55). et al., Nat. Biotechnol. 1998, 16: 568-576).

[0134] Any of a variety of detectable agents may be used in the practice of the present invention. Suitable detectable agents include, but are not limited to, various ligands, radionuclides (such as, for example, 32 P. 35 S. 3 H. 14 C. 125 I. 131 I, etc.); fluorescent dyes (see below for specific exemplary fluorescent dyes); chemiluminescent agents (such as, for example, acridinium esters, stable dioxetanes, etc.); spectrally resolvable inorganic fluorescent semiconductor nanocrystals (i.e., quantum dots), metal nanoparticles (e.g., gold, silver, copper, and platinum) or nanoclusters; enzymes (such as, for example, those used in ELISA, i.e., horseradish peroxidase, β-galactosidase, luciferase, alkaline phosphatase); colorimetric labels (such as, for example, dyes, colloidal gold, etc.); magnetic labels (such as, for example, Dynabeads TM ); and biotin, digoxin, or other haptens and proteins for which antisera or monoclonal antibodies are available.

[0135] In certain embodiments, the detection probe is fluorescently labeled. Many known fluorescent labeling moieties with various chemical structures and physical characteristics are suitable for practice of the present invention. Suitable fluorescent dyes include, but are not limited to, fluorescein and fluorescein dyes (e.g., fluorescein isothiocyanate or FITC, naphthyl fluorescein, 4', 5'-dichloro-2', 7'-dimethoxyfluorescein, 6-carboxyfluorescein or FAM), carbocyanine, merocyanine, styryl dyes, oxonol dyes, phycoerythrin, erythrosine, eosin, rhodamine dyes (e.g., carboxytetramethylrhodamine or TAMRA, carboxyrhodamine 6G, carboxy-X-rhodamine (ROX), lissamine (lissa Rhodamine B, Rhodamine 6G, Rhodamine Green, Rhodamine Red, Tetramethylrhodamine or TMR), Coumarin and Coumarin dyes (e.g., methoxycoumarins, dialkylaminocoumarins, hydroxycoumarins and aminomethylcoumarins or AMCA), Oregon Green dyes (e.g., Oregon Green 488, Oregon Green 500, Oregon Green 514), Texas Red, Texas Red-X, Spectra Red.TM., Spectra Green.TM., Cyanine dyes (e.g., Cy-3 TM 、Cy-5 TM 、Cy-3.5 TM 、Cy-5.5 TM), Alexa Fluor dyes (e.g., Alexa Fluor 350, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 633, Alexa Fluor 660, and Alexa Fluor 680), BODIPY dyes (e.g., BODIPY FL, BODIPY R6G, BODIPY TMR, BODIPY TR, BODIPY 530 / 550, BODIPY 558 / 568, BODIPY 564 / 570, BODIPY 576 / 589, BODIPY581 / 591, BODIPY 630 / 650, BODIPY 650 / 665), IR dyes (e.g., IRD40, IRD 700, IRD 800), and the like. For more examples of suitable fluorescent dyes and methods for attaching or incorporating fluorescent dyes into nucleic acid molecules, see, e.g., "The Handbook of Fluorescent Probes and Research Products", 9th Ed., Molecular Probes, Inc., Eugene, Oreg. Fluorescent dyes and labeling kits are commercially available from, e.g., Amersham Biosciences, Inc. (Piscataway, NJ), Molecular Probes Inc. (Eugene, Oreg.), and New England Biolabs Inc. (Berverly, Mass.).

[0136] As described above, identification of RNA markers can be performed using an amplification reaction.

[0137] As used herein, the term "amplification" refers to a process of increasing the representativeness of a specific nucleic acid sequence population in a sample by producing multiple (i.e., at least 2) copies of the desired sequence. Nucleic acid amplification methods are known in the art and include, but are not limited to, polymerase chain reaction (PCR) and ligase chain reaction (LCR). In a typical PCR amplification reaction, the amount of amplification of the target nucleic acid sequence is typically at least 50,000 times higher than the amount in the starting sample. "Copy" or "amplicon" does not necessarily mean a perfect sequence that is complementary or identical to the template sequence. For example, a copy can include nucleotide analogs (such as deoxyinosine), intentional sequence changes (such as sequence changes introduced by primers that include sequences that can hybridize but are not complementary to the template) and / or sequence errors that occur during amplification.

[0138] A typical amplification reaction is carried out by contacting forward and reverse primers (primer pair) with sample DNA along with any additional amplification reaction reagents under conditions that allow amplification of the target sequence.

[0139] The terms "forward primer" and "forward amplification primer" are used interchangeably herein and refer to a primer that hybridizes (or anneals) to a target (template strand). The terms "reverse primer" and "reverse amplification primer" are used interchangeably herein and refer to a primer that hybridizes (or anneals) to a complementary target strand. The forward primer hybridizes 5' to the target sequence relative to the reverse primer.

[0140] The term "amplification conditions" used herein refers to conditions that promote annealing and / or extension of primer sequences. Such conditions are well known in the art and depend on the selected amplification method. Therefore, for example, in a PCR reaction, the amplification conditions generally include thermal cycling, i.e., the circulation of the reaction mixture between two or more temperatures. In an isothermal amplification reaction, although it may be necessary to raise the initial temperature to initiate the reaction, amplification can occur without thermal cycling. Amplification conditions encompass all reaction conditions, including but not limited to temperature and temperature cycling, buffer, salt, ionic strength and pH, etc.

[0141] As used herein, the term "amplification reaction reagent" refers to a reagent for a nucleic acid amplification reaction, and may include, but is not limited to, a buffer, a reagent, an enzyme with reverse transcriptase and / or polymerase activity or exonuclease activity, an enzyme cofactor such as magnesium or manganese, a salt, nicotinamide adenine dinuclease (NAD), and deoxynucleoside triphosphates (dNTPs) such as deoxyadenosine triphosphate, deoxyguanosine triphosphate, deoxycytidine triphosphate, and thymidine triphosphate. A person skilled in the art can easily select an amplification reaction reagent according to the amplification method used.

[0142] According to this aspect of the invention, amplification can be achieved using techniques such as polymerase chain reaction (PCR), including but not limited to allele-specific PCR, assembly PCR or polymerase cycle assembly (PCA), asymmetric PCR, helicase-dependent amplification, hot-start PCR, inter-sequence-specific PCR (ISSR), inverse PCR, ligation-mediated PCR, methylation-specific PCR (MSP), mini-primer PCR, multiplex ligation-dependent probe amplification, multiplex PCR, nested PCR, overlap extension PCR, quantitative PCR (Q-PCR), reverse transcription PCR (RT-PCR), solid phase PCR : covers a variety of meanings, including polony amplification (e.g., in which PCR colonies are derived from a gel matrix), bridge PCR (primers are covalently attached to a solid support surface), conventional solid phase PCR (in which asymmetric PCR is applied in the presence of primers with a solid support, the sequence of which matches one of the aqueous primers) and enhanced solid phase PCR (conventional solid phase PCR can be improved by applying high Tm and nested solid support primers, optionally applying a thermal "step" to facilitate solid support priming), thermal asymmetric interlaced PCR (TAIL-PCR), touchdown PCR (step-down PCR), PAN-AC and universal rapid walking.

[0143] PCR (or polymerase chain reaction) technology is well known in the art and has been disclosed, for example, in K.B. Mullis and F.A. Faloona, Methods Enzymol., 1987, 155:350-355 and U.S. Pat. Nos. 4,683,202, 4,683,195 and 4,800,159 (each of which is incorporated herein by reference in its entirety). In its simplest form, PCR is an in vitro enzymatic method for synthesizing a specific DNA sequence using two oligonucleotide primers that hybridize to opposite strands and flank a region of interest in the target DNA. Multiple reaction cycles, each cycle includes: a denaturation step, an annealing step and a polymerization step, resulting in the exponential accumulation of specific DNA fragments ("PCR Protocols: A Guide to Methods and Applications", MAInnis (Ed.), 1990, Academic Press: New York; "PCR Strategies", MAInnis (Ed.), 1995, Academic Press: New York; "Polymerase chain reaction: basic principles and automation in PCR: A Practical Approach", McPherson et al. (Eds.), 1991, IRL Press: Oxford; RKSaiki et al., Nature, 1986, 324: 163-166). The end of the amplified fragment is defined as the 5' end of the primer. Examples of DNA polymerases capable of producing an amplified product in a PCR reaction include, but are not limited to, E. coli DNA polymerase I, Klenow fragment of DNA polymerase I, T4 DNA polymerase, a thermostable DNA polymerase isolated from Thermus aquaticus (Taq) (available from a variety of sources (e.g., Perkin Elmer)), Thermus thermophilus (United States Biochemicals), Bacillus stereothermophilus (Bio-Rad), or Thermococcus litoralis ("Vent" polymerase, New England Biolabs). RNA target sequences can be amplified by reverse transcribing mRNA into cDNA and then performing PCR (RT-PCR), as described above.Alternatively, a single enzyme can be used in both steps as described in US Pat. No. 5,322,770.

[0144] The duration and temperature of each step of the PCR cycle and the number of cycles are usually adjusted according to existing stringent requirements. The annealing temperature and time are determined by the efficiency of the expected primer and template annealing and the tolerable mismatch degree. The ability to optimize the reaction cycle conditions is completely within the knowledge of those of ordinary skill in the art. Although the number of reaction cycles can vary according to the detection analysis performed, it is usually at least 15, more usually at least 20, and can be up to 60 or higher. However, in many cases, the number of reaction cycles is usually in the range of about 20 to about 40.

[0145] The denaturation step of the PCR cycle typically involves heating the reaction mixture to an elevated temperature and maintaining the mixture at an elevated temperature for a period of time sufficient to dissociate any double-stranded or hybridized nucleic acids present in the reaction mixture. For denaturation, the temperature of the reaction mixture is typically raised to and maintained in the temperature range of about 85°C to about 100°C, typically about 90°C to about 98°C, and more typically about 93°C to about 96°C, for a period of time ranging from about 3 to about 120 seconds, typically about 5 to about 30 seconds.

[0146] After denaturation, the reaction mixture is placed under conditions sufficient to anneal the primers to the template DNA present in the mixture. The temperature at which the reaction mixture is lowered to achieve these conditions is typically selected to provide optimal efficiency and specificity, and typically ranges from about 50° C. to about 10° C., typically from about 55° C. to about 70° C., and more typically from about 60° C. to about 68° C. Annealing conditions are typically maintained for a period of time ranging from about 15 seconds to about 30 minutes, typically from about 30 seconds to about 5 minutes.

[0147] After or during the annealing of the primer to the template DNA, the reaction mixture is subjected to conditions sufficient to provide for polymerization of nucleotides to the termini of the primers in a manner such that the primers are extended in the 5' to 3' direction using the DNA to which they are hybridized as a template (i.e., conditions sufficient to enzymatically produce primer extension products). To achieve primer extension conditions, the temperature of the reaction mixture is typically raised to a temperature in the range of about 65°C to about 75°C, typically about 67°C to about 73°C, and maintained at that temperature for a period of time in the range of about 15 seconds to about 20 minutes, typically about 30 seconds to about 5 minutes.

[0148] The above-mentioned denaturation, annealing and polymerization cycles can be carried out using an automated device commonly referred to as a thermal cycler or a thermal cycler. Applicable thermal cyclers are described in U.S. Patents Nos. 5,612,473, 5,602,756, 5,538,871 and 5,475,610 (each of which is incorporated herein by reference in its entirety). Thermal cyclers are commercially available from, for example, Perkin Elmer-Applied Biosystems (Norwalk, Conn.), BioRad (Hercules, Calif.), Roche Applied Science (Indianapolis, Ind.) and Stratagene (La Jolla, Calif.).

[0149] Amplification products obtained using the primers of the present invention can be detected using agarose gel electrophoresis and visualization by ethidium bromide staining and exposure to ultraviolet (UV) light or by sequence analysis of the amplification products.

[0150] According to one embodiment, the amplification and quantification of the amplification product can be performed in real time (qRT-PCR). Typically, the QRT-PCR method uses a double-stranded DNA detection molecule to measure the amount of the amplification product in real time.

[0151] As used herein, the phrase "double-stranded DNA detection molecule" refers to a double-stranded DNA interactive molecule that produces a quantifiable signal (e.g., a fluorescent signal). For example, such a double-stranded DNA detection molecule can be a fluorescent dye that (1) interacts with a DNA fragment or an amplicon and (2) emits at a different wavelength when a duplex-formed amplicon is present than when a separated amplicon is present. The double-stranded DNA detection molecule can be a double-stranded DNA intercalating detection molecule or a primer-based double-stranded DNA detection molecule.

[0152] Double-stranded DNA intercalating detection molecules are not covalently linked to primers, amplicons, or nucleic acid templates. The detection molecule increases its emission in the presence of double-stranded DNA and reduces its emission when the duplex DNA is unwound. Examples include, but are not limited to, ethidium bromide, YO-PRO-1, Hoechst 33258, SYBR gold, and SYBR green I. Ethidium bromide is a fluorescent chemical that can be intercalated between base pairs of double-stranded DNA fragments and is commonly used to detect DNA after gel electrophoresis. When excited by ultraviolet light between 254nm and 366nm, it emits 590nm fluorescence. In the presence of single-stranded DNA, the fluorescence produced by the DNA-ethidium bromide complex is about 50 times more than that of ethidium bromide. SYBR Green I is excited at 497nm and emitted at 520nm. Compared with single-stranded DNA, the fluorescence intensity of SYBR Green I after binding to double-stranded DNA increases by more than 100 times. An alternative to SYBR Green I is SYBR gold introduced by Molecular Probes Inc. Similar to SYBR Green I, the fluorescence emission of SYBR Gold is enhanced in the presence of duplex DNA and decreases when the double-stranded DNA is unwound. However, the excitation peak of SYBR Gold is at 495nm, and the emission peak is at 537nm. It is reported that SYBR Gold behaves more stably than SYBR Green I. Hoechst 33258 is a known bisbenzimide double-stranded DNA detection molecule that can bind to the AT-rich region of duplex DNA. Hoechst 33258 excites at 350nm and emits at 450nm. It is reported that YO-PRO-1 excites at 450nm and emits at 550nm, and is said to be a double-stranded DNA specific detection molecule. In a specific embodiment of the present invention, the double-stranded DNA detection molecule is SYBR Green I.

[0153] The double-stranded DNA detection molecule based on primers is covalently linked to the primers, and when the amplicon forms a duplex structure, the fluorescence emission can be increased or decreased. When the double-stranded DNA detection molecule based on primers is attached to the 3' end of the primer and the primer terminal base is dG or dC, the fluorescence emission is observed to increase. When the detection molecule is located at least 6 nucleotides inside the primer end, the detection molecule is quenched near the terminal dC-dG and dG-dC base pairs, and is dequenched due to the duplex formation of the amplicon. Dequenching causes a large increase in fluorescence emission. Examples of these types of detection molecules include, but are not limited to, fluorescein (excited at 488nm and emitted at 530nm), FAM (excited at 494nm and emitted at 518nm), JOE (excited at 527 and emitted at 548), HEX (excited at 535nm and emitted at 556nm), TET (excited at 521nm and emitted at 536nm), Alexa Fluor 594 (excited at 590nm and emitted at 615nm), ROX (excited at 575nm and emitted at 602nm), and TAMRA (excited at 555nm and emitted at 580nm). In contrast, some primer-based double-stranded DNA detection molecules reduce their emission relative to single-stranded DNA in the presence of double-stranded DNA. Examples include, but are not limited to, rhodamine and BODIPY-FI (excited at 504nm and emitted at 513nm). These detector molecules are usually covalently conjugated to the primer at the 5' terminal dC or dG and emit less fluorescence when the amplicon is in a duplex. It is believed that the reduction in fluorescence upon duplex formation is due to quenching by guanosine in the complementary strand immediately adjacent to the detector molecule or by quenching by the terminal dC-dG base pair.

[0154] According to one embodiment, the primer-based double-stranded DNA detection molecule is a 5' nuclease probe. Such probes incorporate a fluorescent reporter molecule at the 5' or 3' end of the oligonucleotide and a quencher at the opposite end. The first step of the amplification process involves heating to denature the double-stranded DNA target molecule into single-stranded DNA. During the second step, the forward primer anneals to the target strand of the DNA and is extended by Taq polymerase. The reverse primer and 5' nuclease probe then anneal to this newly replicated strand.

[0155] In this embodiment, at least one of the primer pair or the 5' nuclease probe should hybridize to the unique sequence (of Table 1 or 2). The polymerase extends from the target strand and cleaves the probe. After cleavage, the reporter is no longer quenched by proximity to the quencher and releases fluorescence. Each replication results in cleavage of the probe. As a result, the fluorescent signal will increase in proportion to the amount of amplified product.

[0156] DNA, RNA or protein can be extracted from any plant tissue or cell sample, such as leaves, phloem, bark, roots, seeds and flowers.

[0157] According to a specific embodiment, the sample is a leaf sample (see Examples section below).

[0158] As described above, according to other embodiments, genes are targets for intervention to select for resistant plants.

[0159] Therefore, according to one aspect of the present invention, there is provided a method for producing a Eucalyptus plant exhibiting resistance to a physiological disorder in Eucalyptus, the method comprising up-regulating the expression and / or activity of at least one gene of Table 2 in Eucalyptus, and / or down-regulating the expression and / or activity of at least one gene of Table 1 in Eucalyptus, thereby increasing resistance to the phenotype of a physiological disorder in Eucalyptus.

[0160] Methods of modifying gene expression are well known in the art and are provided in greater detail below in a non-limiting manner.

[0161] As used herein, the term "increase" refers to an increase in a trait [e.g., resistance] of a plant by at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80% compared to a control plant (a plant that has not been modified according to the present teachings and that naturally exhibits susceptibility to a physiological disorder), such as a natural plant, a wild-type plant, a non-transformed plant, or a non-genome-edited plant of the same species grown under the same (e.g., identical) growth conditions.

[0162] As used herein, the phrase "overexpressing a polypeptide" refers to an increase in the level of a polypeptide in a plant compared to control plants of the same species grown under the same conditions.

[0163] According to some embodiments of the invention, the increased polypeptide level is in a specific cell type or organ of the plant.

[0164] According to some embodiments of the invention, the increased level of the polypeptide is at a temporal point in the plant.

[0165] According to some embodiments of the invention, the increased polypeptide level is during the entire life cycle of the plant.

[0166] For example, overexpression of a polypeptide can be achieved by increasing the expression level of a plant's native gene compared to a control plant. This can be accomplished, for example, by means of genome editing, as described further below, for example by introducing one or more mutations in one or more regulatory elements (e.g., enhancers, promoters, untranslated regions, intronic regions) that result in upregulation of the native gene, and / or by homology-directed repair (HDR), for example, for introducing a "repair template" encoding a polypeptide of interest.

[0167] Additionally and / or alternatively, overexpression of a polypeptide may be achieved by increasing the level of the polypeptide of interest due to expression of a heterologous polynucleotide with the aid of recombinant DNA techniques, for example using a nucleic acid construct comprising a polynucleotide encoding the polypeptide of interest.

[0168] In some embodiments, if the plant of interest (e.g., the plant in which the polypeptide is to be overexpressed) does not have a detectable expression level of the polypeptide of interest before the methods of some embodiments of the present invention are used, then "overexpression" of the polypeptide in the plant is identified by determining a positive detectable expression level of the polypeptide of interest in the plant cells and / or plants.

[0169] Additionally and / or alternatively, if the plant of interest (e.g., a plant in which the polypeptide is to be overexpressed) has a detectable expression level of the polypeptide of interest to a certain extent prior to adopting the methods of some embodiments of the present invention, then the "overexpression" of the polypeptide in the plant is identified by measuring the increase in the expression level of the polypeptide of interest in the plant cells and / or plants compared to control plant cells and / or plants of the same species grown under the same (e.g., the same) growth conditions, respectively.

[0170] Methods for detecting the presence or absence of polypeptides in plant cells and / or plants and quantification of protein expression levels are well known in the art (eg, protein detection methods) and are further described below.

[0171] As used herein, the phrase "expressing an exogenous polynucleotide encoding a polypeptide" refers to expression at the mRNA level.

[0172] As used herein, the phrase "expressing an exogenous polynucleotide encoding a polypeptide" refers to expression at the mRNA level.

[0173] As used herein, the phrase "exogenous polynucleotide" refers to a heterologous nucleic acid sequence that may not be naturally expressed in a plant (e.g., a nucleic acid sequence from a different species) or that is desired to be overexpressed in a plant. Exogenous polynucleotides can be introduced into a plant in a stable or transient manner to produce ribonucleic acid (RNA) molecules and / or polypeptide molecules. It should be noted that exogenous polynucleotides can include nucleic acid sequences that are identical or partially homologous to endogenous nucleic acid sequences of a plant.

[0174] The term "endogenous" as used herein refers to any polynucleotide or polypeptide present and / or naturally expressed in a plant or a cell thereof.

[0175] According to some embodiments of the invention, the exogenous polynucleotides of the invention include a nucleic acid sequence encoding a polypeptide having an amino acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more 100% homologous (e.g., identical) to an amino acid sequence selected from the group consisting of SEQ ID NOs: 60-118.

[0176] Homologous sequences include both orthologous sequences and paralogous sequences. The term "paralogous sequence" refers to the gene duplication that causes paralogous genes in species genomes. The term "orthologous sequence" refers to homologous genes in different organisms due to ancestral relationships. Therefore, orthologs are evolutionary counterparts (Koonin EV and Galperin MY (Sequence-Evolution-Function: ComputationalApproaches in Comparative Genomics.Boston: Kluwer Academic) derived from the last common ancestor of a given two species, and therefore are likely to have the same function.

[0177] One option for identifying orthologs in monocot species is to perform mutual blast searches. This can be accomplished by a first blast analysis, which involves blasting the target sequence against any sequence database, such as the publicly available NCBI database found below: ncbi(.)nlm(.)nih(.)gov. If you are looking for orthologs in rice, you can blast the target sequence with 28,469 full-length cDNA clones from Oryza sativa Nipponbare available on NCBI, for example. The blast analysis results may be filtered. The full-length sequence of the filtered results or the unfiltered results is then subjected to a reverse blast analysis (second blast analysis) with the sequence of the organism from which the target sequence is derived. The results of the first and second blast analyses are then compared. When the sequence that produces the highest score (best hit) in the first blast analysis identifies the query sequence (original target sequence) as the best hit in the second blast analysis, the ortholog is identified. Using the same reasoning, paralogs (homologs of a certain gene in the same organism) are found. For large sequence families, the ClustalW program [ebi(.)ac(.)uk / Tools / clustalw2 / index(.)html] can be used, followed by neighbor joining to help visualize the clustered tree (wikipedia(.)org / wiki / Neighbor-joining).

[0178] Homology (eg, percent homology, sequence identity + sequence similarity) can be determined using any homology comparison software that calculates pairwise sequence alignments.

[0179] As used herein, "sequence identity" or "identity" in the context of two nucleic acids or polypeptide sequences includes residues in two sequences that are identical when the two sequences are aligned. When using a percentage of sequence identity with respect to a protein, it is recognized that non-identical residue positions are usually different due to conservative amino acid substitutions, wherein the amino acid residues are replaced by other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity), and therefore the functional properties of the molecule will not be changed. When the sequences are different due to conservative substitutions, the percentage of sequence identity can be adjusted upward to correct the conservative nature of the substitution. Sequences that differ due to such conservative substitutions are considered to have "sequence similarity" or "similarity". The method for making such adjustments is well known to those skilled in the art. Typically, this involves scoring conservative substitutions as partial mismatches rather than complete mismatches, thereby increasing the percentage of sequence identity. Therefore, for example, when the same amino acid score is 1 and the non-conservative substitution score is zero, the conservative substitution score is between zero and 1. The scoring of conservative substitutions is calculated, for example, according to the algorithm of Henikoff S and Henikoff JG. [Amino acid substitution matrices from protein blocks. Proc. Natl. Acad. Sci. USA 1992, 89 (22): 10915-9].

[0180] Identity (eg, percent homology) can be determined using any homology comparison software, including, for example, the BlastN software from the National Center of Biotechnology Information (NCBI), such as by using default parameters.

[0181] According to some embodiments of the invention, the identity is an overall identity, ie, the identity of the entire amino acid or nucleic acid sequence of the invention rather than the identity of a portion thereof.

[0182] According to some embodiments of the invention, the term "homology" or "homologous" refers to the identity of two or more nucleic acid sequences; or the identity of two or more amino acid sequences; or the identity of an amino acid sequence and one or more nucleic acid sequences.

[0183] According to some embodiments of the invention, the homology is overall homology, ie, homology of the entire amino acid or nucleic acid sequence of the invention, rather than homology of a portion thereof.

[0184] The degree of homology or identity between two or more sequences can be determined using various known sequence comparison tools. The following is a non-limiting description of such tools that can be used in conjunction with some embodiments of the present invention.

[0185] Pairwise global alignments are defined by SB Needleman and CD Wunsch, "A general method applicable to the search of similarities in the amino acid sequence of two proteins" Journal of Molecular Biology, 1970, pages 443-53, volume 48).

[0186] For example, when starting with a polypeptide sequence and comparing it to other polypeptide sequences, the EMBOSS-6.0.1 Needleman-Wunsch algorithm (available from emboss(.)sourceforge(.)net / apps / cvs / emboss / apps / needle(.)html) can be used to find the best alignment of two sequences along their entire length (including gaps) - a "global alignment". The default parameters for the Needleman-Wunsch algorithm (EMBOSS-6.0.1) include: gapopen=10; gapextend=0.5; datafile=EBLOSUM62; concise=yes.

[0187] According to some embodiments of the invention, parameters used with the EMBOSS-6.0.1 tool (for protein-protein comparison) include: Gap Opening = 8; Gap Extension = 2; Data File = EBLOSUM62; Brief = Yes.

[0188] According to some embodiments of the invention, the threshold for determining homology using the EMBOSS-6.0.1 Needleman-Wunsch algorithm is 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%.

[0189] When starting with a polypeptide sequence and comparing to a polynucleotide sequence, the OneModelFramePlus algorithm [Halperin, E., Faigler, S. and Gill-More, R. (1999) - FramePlus: aligning DNA to protein sequences. Bioinformatics, 15, 867-873) (available from biocceleration(.)com / Products(.)html] can be used with the following default parameters: Model=Framework+_p2n. ModelMode=Local.

[0190] According to some embodiments of the invention, the parameters used with the single model framework enhancement algorithm are model=framework+_p2n.model, mode=qentire.

[0191] According to some embodiments of the invention, the threshold for determining homology using a single model framework enhancement algorithm is 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%.

[0192] The EMBOSS-6.0.1 Needleman-Wunsch algorithm (available from emboss(.)sourceforge(.)net / apps / cvs / emboss / apps / needle(.)html) can be used with the following default parameters when starting from a polynucleotide sequence and comparing to other polynucleotide sequences: (EMBOSS-6.0.1) gap open = 10; gap extend = 0.5; data file = EDNAFULL; brief = yes.

[0193] According to some embodiments of the invention, the parameters used with the EMBOSS-6.0.1 Needleman-Wunsch algorithm are Gap Open = 10; Gap Extend = 0.2; Data File = EDNAFULL; Brief = Yes.

[0194] According to some embodiments of the invention, when comparing polynucleotides to polynucleotides using the EMBOSS-6.0.1 Needleman-Wunsch algorithm, the threshold for determining homology is 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%.

[0195] According to some embodiments, determining the degree of homology further entails using the Smith-Waterman algorithm (for protein-protein comparisons or nucleotide-nucleotide comparisons).

[0196] The default parameters for the GenCore 6.0 Smith-Waterman algorithm include: model = sw.model.

[0197] According to some embodiments of the invention, the threshold for homology determined using the Smith-Waterman algorithm is 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%.

[0198] According to some embodiments of the present invention, before the overall homology (for example, 80% overall homology of the whole sequence) is carried out to the target polypeptide or polynucleotide, the sequence pre-selected by the local homology (for example, 60% identity of 60% sequence length) with the target polypeptide or polynucleotide is carried out overall homology.For example, BLAST software is used to select homologous sequences, wherein Blastp and tBlastn algorithms are used as filters in the first stage, and needle (EMBOSS software package) or Frame+ algorithm comparison is used as a filter in the second stage.Local identity (Blast comparison) is defined as a very loose cutoff value--with 60% identity in 60% sequence length, because it is only used as a filter in the overall comparison stage.In this specific embodiment (when using local identity), the default filtering (by setting parameter "-FF") of the Blast software package is not used.

[0199] In the second stage, homologs were defined based on having at least 80% overall identity to the core gene polypeptide sequence.

[0200] According to some embodiments, the homology is local homology or local identity.

[0201] Local alignment tools include, but are not limited to, BlastP, BlastN, BlastX or TBLASTN software from the National Center of Biotechnology Information (NCBI), FASTA and the Smith-Waterman algorithm.

[0202] The tblastn search allows the comparison of protein sequences with six-frame translations of nucleotide databases. It can be a very efficient method to find homologous protein coding regions in unlabeled nucleotide sequences, such as expressed sequence tags (ESTs) and draft genome records (HTGs) located in the BLAST databases est and htgs, respectively.

[0203] The default parameters of blastp include: maximum target sequence: 100; expected threshold: e -5 ; Word length: 3; Maximum match value in query range: 0; Scoring parameters: Matrix – BLOSUM62; Filters and masks: Filters – Low complexity regions.

[0204] Local alignment tools that can be used include, but are not limited to, the tBLASTX algorithm, which compares the six-frame conceptual translation products of the nucleotide query sequence (both strands) to a protein sequence database. Default parameters include: maximum target sequence: 100; expected threshold: 10; word length: 3; maximum match value within the query range: 0; scoring parameters: matrix-BLOSUM62; filters and masks: filters-low complexity regions.

[0205] According to some embodiments of the invention, the exogenous polynucleotide of the invention encodes a polypeptide having an amino acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 60-118.

[0206] According to some embodiments of the present invention, the exogenous polynucleotide of the present invention encodes a polypeptide having an amino acid sequence selected from the group consisting of SEQ ID NOs: 60-118.

[0207] According to some embodiments of the invention, a method for increasing resistance of a plant to a physiological disorder is achieved by expressing in a plant an exogenous polynucleotide comprising a nucleic acid sequence encoding a polypeptide, wherein the polypeptide is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or more 100% identical to an amino acid sequence selected from the group consisting of SEQ ID NOs: 60-118, thereby increasing the resistance of the plant.

[0208] According to some embodiments of the present invention, the exogenous polynucleotide encodes a polypeptide consisting of the amino acid sequence shown in SEQ ID NO: 60-118.

[0209] According to an aspect of some embodiments of the present invention, there is provided a method for increasing resistance of a plant to a physiological disorder, the method comprising expressing in a plant an exogenous polynucleotide comprising a nucleic acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, such as 100% identical to a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 1-59, thereby increasing resistance of the plant to the physiological disorder.

[0210] According to some embodiments of the invention, the exogenous polynucleotide is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, such as 100% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 1-59.

[0211] According to some embodiments of the invention, the exogenous polynucleotide is represented by SEQ ID NO: 1-59.

[0212] As used herein, the term "polynucleotide" refers to a single-stranded or double-stranded nucleic acid sequence that is isolated and provided in the form of an RNA sequence, a complementary polynucleotide sequence (cDNA), a genomic polynucleotide sequence, and / or a composite polynucleotide sequence (e.g., a combination of the above).

[0213] The term "isolated" means at least partially separated from the natural environment, for example, separated from plant cells.

[0214] As used herein, the phrase "complementary polynucleotide sequence" refers to a sequence produced by reverse transcription of messenger RNA using reverse transcriptase or any other RNA-dependent DNA polymerase. Such a sequence can then be amplified in vivo or in vitro using a DNA-dependent DNA polymerase.

[0215] As used herein, the phrase "genomic polynucleotide sequence" refers to a sequence that is derived (isolated) from a chromosome, and thus represents a contiguous portion of a chromosome.

[0216] As used herein, the phrase "composite polynucleotide sequence" refers to a sequence that is at least partially complementary and at least partially genomic. A composite sequence may include some exon sequences required to encode a polypeptide of the present invention, as well as some intron sequences inserted therein. Intron sequences may be of any origin, including other genes, and will generally include conserved splicing signal sequences. Such intron sequences may further include cis-acting expression regulatory elements.

[0217] According to some embodiments of the present invention, the exogenous polynucleotide encodes a polypeptide consisting of the amino acid sequence shown in SEQ ID NO: 60-118.

[0218] The nucleic acid sequence encoding the polypeptide of the present invention can be optimized for expression. Examples of such sequence modifications include, but are not limited to, changing the G / C content to be closer to that usually found in the target plant species, and removing the codons found atypically in the plant species, commonly referred to as codon optimization.

[0219] The phrase "codon optimization" refers to the selection of appropriate DNA nucleotides for use in a structural gene or fragment thereof that are close to the codon usage in the target plant. Therefore, an optimized gene or nucleic acid sequence refers to a gene in which the nucleotide sequence of a natural or naturally occurring gene has been modified to utilize statistically preferred or statistically favorable codons in a plant. The nucleotide sequence is usually examined at the DNA level, and any suitable program is used to determine the coding region optimized for expression in a plant species, such as described in Sardana et al. (1996, Plant Cell Reports 15: 677-681). In this method, the standard deviation of codon usage (a measure of codon usage bias) can be calculated by first finding the squared proportional deviation of the usage of each codon of a natural gene relative to the usage of each codon of a highly expressed plant gene, followed by calculating the average squared deviation. The formula used is: 1SDCU=n=1N[(Xn-Yn) / Yn]2 / N, wherein Xn refers to the frequency of use of codon n in a highly expressed plant gene, wherein Yn refers to the frequency of use of codon n in a target gene, and N refers to the total number of codons in a target gene. A codon usage table of highly expressed genes in dicotyledonous plants was compiled using data from Murray et al. (1989, Nuc Acids Res. 17:477-498).

[0220] One method for optimizing nucleic acid sequences according to the preferred codon usage of a particular plant cell type is based on the direct use of codon optimization tables (such as those provided online in the Codon Usage Database by the Japanese NIAS (National Institute of Agrobiological Sciences) DNA library (kazusa (.) or (.) jp / codon / )) without any additional statistical calculations. The Codon Usage Database contains codon usage tables for many different species, each of which is statistically determined based on data present in Genbank.

[0221] By using the above table to determine the most preferred or most favorable codons for each amino acid in a particular species (e.g., rice), the naturally occurring nucleotide sequence encoding the target protein can be codon optimized for that particular plant species. This is achieved by replacing codons that may have low statistical incidences in the genome of a particular species with corresponding codons that are statistically more favorable with respect to amino acids. However, one or more less favorable codons can be selected to delete existing restriction sites, to create new restriction sites at potentially useful junctions (5' and 3' ends to add signal peptides or termination boxes, internal sites that can be used to cut and splice fragments together to produce the correct full-length sequence), or to eliminate nucleotide sequences that may have a negative impact on mRNA stability or expression.

[0222] Prior to any modification, a naturally occurring coding nucleotide sequence may have contained many codons corresponding to statistically favorable codons in a particular plant species. Therefore, codon optimization of a native nucleotide sequence can include determining which codons within the native nucleotide sequence are not statistically favorable for a particular plant, and modifying these codons to produce codon-optimized derivatives according to the codon usage table of the particular plant. The modified nucleotide sequence can be fully or partially optimized for plant codon usage, provided that the protein encoded by the modified nucleotide sequence is produced at a level higher than that of the protein encoded by the corresponding naturally occurring or native gene. Construction of synthetic genes by altering codon usage is described in, for example, PCT patent application 93 / 07278.

[0223] Therefore, the present invention encompasses the nucleic acid sequences described above; fragments thereof, sequences that can hybridize therewith, sequences homologous thereto, sequences encoding similar polypeptides with different codon usage, altered sequences characterized by mutations, such as naturally occurring or artificially induced, random or targeted deletions, insertions or substitutions of one or more nucleotides.

[0224] According to some embodiments of the invention, the exogenous polynucleotide encodes a polypeptide comprising an amino acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, such as 100% identical to the amino acid sequence of a naturally occurring plant ortholog of a polypeptide selected from the group consisting of SEQ ID NOs: 60-118.

[0225] According to some embodiments of the invention, the polypeptide comprises an amino acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, such as 100% identical to the amino acid sequence of a naturally occurring plant ortholog of a polypeptide selected from the group consisting of SEQ ID NOs: 60-118.

[0226] The invention provides an isolated polynucleotide comprising a nucleic acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, e.g., 100% identical to a polynucleotide selected from the group consisting of SEQ ID NOs: 1-59.

[0227] Also provided are nucleic acid constructs comprising a nucleic acid sequence encoding a gene of Table 2 or a functional homolog thereof and a heterologous cis-acting regulatory element for driving expression of the gene or functional homolog, such as described herein.

[0228] According to some embodiments of the present invention, provided are plant cells that exogenously express: a polynucleotide of some embodiments of the present invention, a nucleic acid construct of some embodiments of the present invention, and / or a polypeptide of some embodiments of the present invention.

[0229] According to some embodiments of the present invention, expressing the exogenous polynucleotide of the present invention in a plant is achieved by transforming one or more cells of the plant with the exogenous polynucleotide, then generating a mature plant from the transformed cells and culturing the mature plant under conditions suitable for expressing the exogenous polynucleotide in the mature plant.

[0230] According to some embodiments of the present invention, transformation is achieved by introducing a nucleic acid construct into a plant cell, the nucleic acid construct comprising the exogenous polynucleotides of some embodiments of the present invention and at least one promoter for directing the transcription of the exogenous polynucleotides in a host cell (plant cell). Other details of a suitable transformation method are provided below.

[0231] As mentioned, the nucleic acid construct according to some embodiments of the present invention comprises a promoter sequence and the isolated polynucleotide of some embodiments of the present invention.

[0232] According to some embodiments of the invention, the isolated polynucleotide is operably linked to a promoter sequence.

[0233] A coding nucleic acid sequence is "operably linked" to a regulatory sequence (eg, a promoter) if the regulatory sequence is capable of exerting a regulatory effect on the coding sequence to which it is linked.

[0234] As used herein, the term "promoter" refers to a DNA region located upstream of the gene transcription start site to which RNA polymerase binds to initiate RNA transcription. The promoter controls where (e.g., which part of the plant) and / or when (e.g., which stage or state in the life cycle of an organism) a gene is expressed.

[0235] According to some embodiments of the invention, the promoter is heterologous to the isolated polynucleotide and / or the host cell.

[0236] As used herein, the phrase "heterologous promoter" refers to a promoter from a different species than the species from which the polynucleotide was isolated, or to a promoter from the same species but at a different locus within the plant genome than the locus from which the polynucleotide sequence was isolated.

[0237] According to some embodiments of the invention, the isolated polynucleotide is heterologous to the plant cell (eg, the polynucleotide is derived from a different plant species than the plant cell, and thus the isolated polynucleotide and the plant cell are not from the same plant species).

[0238] The nucleic acid construct of the present invention can use any suitable promoter sequence. Preferably, the promoter is a constitutive promoter, a tissue-specific promoter or an abiotic stress-inducible promoter.

[0239] According to some embodiments of the invention, the promoter is a plant promoter, which is suitable for the expression of exogenous polynucleotides in plant cells.

[0240] Suitable constitutive promoters include, for example, the CaMV 35S promoter (CaMV 35S (pQXNc) promoter); PJJ 35S from Brachypodium; CaMV 35S (OLD) promoter, Odell et al., Nature 313:810-812, 1985, Arabidopsis At6669 promoter see PCT Publication No. WO04081173A2 or the new At6669 promoter; maize Ub1 promoter [variety Nongda 105; GenBank: DQ141598.1; Taylor et al., Plant Cell Rep 199312:491-495, which is incorporated herein by reference in its entirety; and variety B73; Christensen, AH, et al. ... Mol. Biol. 18(4), 675-689 (1992), which is incorporated herein by reference in its entirety]; rice actin 1 (McElroy et al., Plant Cell 2:163-171, 1990); pEMU (Last et al., Theor. Appl. Genet. 81:581-588, 1991); CaMV 19S (Nilsson et al., Physiol. Plant 100:456-462, 1997); rice GOS2 [(rice GOS2 longer promoter and GOS2 promoter), de Pater et al, Plant J Nov; 2(6):837-44, 1992]; RBCS promoter; rice cyclophilin (Bucholz et al, Plant Mol Biol. 25(5):837-43, 1994); maize H3 histone (Lepetit et al, Mol. Gen. Genet. 231:276-285, 1992); actin 2 (An et al, Plant J. 10(1):107-121, 1996) and synthetic super MAS (Ni et al, The Plant Journal 7:661-76, 1995). Other constitutive promoters include those in U.S. Pat. Nos. 5,659,026; 5,608,149; 5.608,144; 5,604,121; 5.569,597; 5.466,785; 5,399,680; 5,268,463; and 5,608,142.

[0241] Suitable tissue-specific promoters include, but are not limited to, leaf-specific promoters [e.g., AT5G06690 (thioredoxin) (high expression), AT5G61520 (AtSTP3) (low expression) described in Buttner et al 2000 Plant, Cell and Environment 23, 175-184; or Yamamoto et al., Plant J. 12: 255-265, 1997; Kwon et al., Plant Physiol. 105: 357-67, 1994; Yamamoto et al., Plant Cell Physiol. 35: 773-778, 1994; Gotor et al., Plant J. 3: 509-18, 1993; Orozco et al., Plant Mol. Biol. 23: 1129-1138, 1993; and Matsuoka ... al., Proc. Natl. Acad. Sci. USA 90:9586-9590, 1993; and Arabidopsis STP3 (AT5G61520) promoter (Buttner et al., Plant, Cell and Environment 23:175-184, 2000)], seed-preferred promoters [e.g. Napin (derived from Brassica napus, which is characterized by seed-specific promoter activity; Stuitje AR et. al. Plant Biotechnology Journal 1(4):301-309; (Brassica napus NAPIN promoter) from seed-specific genes (Simon, et al., Plant Mol. Biol. 5.191, 1985; Scofield, et al. al., J. Biol. Chem. 262:12202, 1987; Baszczynski, et al., Plant Mol. Biol. 14:633, 1990), rice PG5a (US 7,700,835), early seed development Arabidopsis BAN (AT1G61720) (US2009 / 0031450 A1), late seed development Arabidopsis ABI3 (AT3G24650) (Arabidopsis ABI3 (AT3G24650) longer promoter) or Arabidopsis ABI3 (AT3G24650) promoter) (Ng et al., Plant Molecular Biology 54:25-38,2004), Brazil nut albumin (Pearson'et al., Plant Mol. Biol. 18:235-245,1992), legumin (Ellis,et al. Plant Mol. Biol. 10:203-214,1988), glutenin (rice) (Takaiwa,et al., Mol. Gen. Genet. 208:15-22,1986; Takaiwa,et al., FEBS Letts. 221:43-47,1987), zein (Matzke et al Plant Mol Biol, 143.323-32 1990), napA (Stalberg,et al, Planta 199:515-519,1996), wheat SPA (Albanietal, Plant , Cell, 9:171-184, 1997), sunflower oleosin (Cummins, et al., Plant Mol. Biol. 19:873-876, 1992)], endosperm-specific promoter [Thomas and Flavell, The Plant Cell 2:1171-1180, 1990; Mol Gen Genet 216:81-90, 1989; NAR 17:461-2), wheat α, β and γ gliadin (wheat α gliadin (B genome) promoter); wheat γ gliadin promoter; EMBO 3:1409-15, 1984), barley ltrl promoter, barley B1, C, D hordein (Theor Appl Gen 98:1253-62, 1999; Plant J 4:343-55, 1993; Mol Gen Genet 216:81-90, 1989; NAR 17:461-2), wheat α, β and γ gliadin (wheat α gliadin (B genome) promoter); wheat γ gliadin promoter; EMBO 3:1409-15, 1984), barley ltrl promoter, barley B1, C, D hordein (Theor Appl Gen 98:1253-62, 1999; Plant J 4:343-55, 1993; Mol Gen Genet 216:81-90, 1989; NAR 17:461-2 Genet 250:750-60, 1996), barley DOF (Mena et al, The Plant Journal, 116 (1): 53-62, 1998), Biz2 (EP99106056.7), barley SS2 (barley SS2 promoter; Guerin and Carbonero Plant Physiology 114: 1 55-62, 1997), wheat Tarp60 (Kovalchuk et al., Plant Mol Biol 71: 81-98, 2009), barley D-hordein (D-Hor) and B-hordein (B-Hor) (Agnelo Furtado, Robert J.Henry and Alessandro Pellegrineschi (2009)], synthetic promoters (Vicente-Carbajosa et al., Plant J. 13: 629-640, 1998), rice prolamin NRP33, rice globulin Glb-1 (Wu et al., Plant Cell Physiology 39 (8) 885-889, 1998), rice α-globulin REB / OHP-1 (Nakase et al. Plant Mol. Biol. 33: 513-S22, 1997), rice ADP-glucose PP (Trans Res 6: 157-68, 1997), maize ESR gene family (Plant J 12: 235-46, 1997), sorghum γ-kafirin (PMB 32: 1029-35, 1996)], embryo-specific promoters [e.g., rice OSH1 (Sato et al, Proc. Natl. Acad. Sci. USA, 93: 8117-8122), KNOX (Postma-Haarsma et al, Plant Mol. Biol. 39: 257-71, 1999), rice oleosin (Wu et at, J. Biochem., 123: 386, 1998)] and flower-specific promoters [e.g. AtPRP4, chalcone synthase (chsA) (Van der Meer, et al., Plant Mol. Biol. 15, 95-109, 1990), LAT52 (Twell et al Mol. Gen Genet. 217: 240-245; 1989).

[0242] The nucleic acid construct of some embodiments of the present invention may further include a suitable selectable marker and / or replication origin. According to some embodiments of the present invention, the nucleic acid construct utilized is a shuttle vector that can be propagated in E. coli (wherein the construct includes a suitable selectable marker and replication origin) and can be compatible with propagation in cells. The construct according to the present invention can be, for example, a plasmid, a bacmid, a phasmid, a cosmid, a phage, a virus or an artificial chromosome.

[0243] The nucleic acid constructs of some embodiments of the present invention can be used for stable or transient transformation of plant cells. In stable transformation, the exogenous polynucleotide is integrated into the plant genome, so it represents a stable hereditary trait. In transient transformation, the exogenous polynucleotide is expressed by the transformed cell, but it is not integrated into the genome, and so it represents a transient trait.

[0244] There are various methods for introducing foreign genes into both monocotyledonous and dicotyledonous plants (Potrykus, I., Annu. Rev. Plant. Physiol., Plant. Mol. Biol. (1991) 42:205-225; Shimamoto et al., Nature (1989) 338:274-276).

[0245] There are two main approaches to stably integrate foreign DNA into plant genomic DNA:

[0246] (i) Agrobacterium-mediated gene transfer: Klee et al. (1987) Annu. Rev. Plant Physiol. 38: 467-486; Klee and Rogers in Cell Culture and SomaticCell Genetics of Plants, Vol. 6, Molecular Biology of Plant Nuclear Genes, eds. Schell, J., and Vasil, LK, Academic Publishers, San Diego, Calif. (1989) p. 2-25; Gatenby, in Plant Biotechnology, eds. Kung, S. and Arntzen, CJ, Butterworth Publishers, Boston, Mass. (1989) p. 93-112.

[0247] (ii) Direct DNA uptake: Paszkowski et al., in Cell Culture and Somatic Cell Genetics of Plants, Vol. 6, Molecular Biology of Plant Nuclear Genes eds. Schell, J., and Vasil, LK, Academic Publishers, San Diego, Calif. (1989) p. 52-68; including methods for direct uptake of DNA into protoplasts, Toriyama, K. et al. (1988) Bio / Technology 6: 1072-1074. DNA uptake induced by brief electric shock of plant cells: Zhang et al. Plant Cell Rep. (1988) 7: 379-384. Fromm et al. Nature (1986) 319: 791-793. Injection of DNA into plant cells or tissues by particle bombardment, Klein et al. Bio / Technology (1988) 6:559-563; McCabe et al. Bio / Technology (1988) 6:923-926; Sanford, Physiol. Plant. (1990) 79:206-209; by using a micropipette system: Neuhaus et al., Theor. Appl. Genet. (1987) 75:30-36; Neuhaus and Spangenberg, Physiol. Plant. (1990) 79:213-217; glass fiber or silicon carbide whisker transformation of cell culture, embryos or callus tissue, U.S. Pat. No. 5,464,765, or by direct incubation of DNA with germinating pollen, DeWet et al. in Experimental Manipulation of Ovule Tissue, eds. Chapman, GP and Mantell, SH and Daniels, W. Longman, London, (1985) p. 197-209 and Ohta, Proc. Natl. Acad. Sci. USA (1986) 83:715-719.

[0248] Agrobacterium system includes the use of plasmid vectors containing defined DNA fragments integrated into plant genomic DNA. The inoculation method of plant tissue varies according to the plant species and Agrobacterium delivery system. A widely used method is leaf disc surgery, which can be performed with any tissue explant, providing a good source for the initiation of whole plant differentiation. See, for example, Horsch et al. in Plant Molecular Biology Manual A5, Kluwer Academic Publishers, Dordrecht (1988) p. 1-9. Supplementary methods use Agrobacterium delivery system in combination with vacuum infiltration. Agrobacterium system is particularly feasible in the production of transgenic dicotyledons.

[0249] There are several methods for transferring DNA directly into plant cells. In electroporation, protoplasts are briefly exposed to a strong electric field. In microinjection, DNA is mechanically injected directly into cells using a very small micropipette. In microprojectile bombardment, DNA is adsorbed onto microprojectiles such as magnesium sulfate crystals or tungsten particles, and the microprojectiles are physically accelerated into cells or plant tissues.

[0250] Stable transformation is followed by plant propagation. The most common method of plant propagation is through seeds. However, regeneration through seed propagation has the following drawbacks: the crop lacks uniformity due to heterozygosity, since seeds are produced from plants according to genetic variation controlled by Mendel's rules. Basically, each seed is genetically different, and each seed will grow in its own specific shape. Therefore, it is preferred to produce transformed plants so that the regenerated plants have the same traits and characteristics as the parent transgenic plant. Therefore, it is preferred to regenerate transformed plants through micropropagation, which provides rapid and consistent reproduction of transformed plants.

[0251] Micropropagation is the process of growing a new generation of plants from a single piece of tissue cut from a selected parent plant or variety. The process allows for the mass propagation of plants with the preferred tissue expressing the fusion protein. The new generation of plants produced is genetically identical to the original plant and has all of the characteristics of the original plant. Micropropagation allows for the mass production of high-quality plant material in a short period of time and allows for the rapid multiplication of selected varieties while retaining the characteristics of the original transgenic or transformed plant. The advantages of cloning plants are the speed of plant propagation and the quality and consistency of the plants produced.

[0252] Micropropagation is a multi-stage process that requires changing the culture medium or growth conditions between stages. Therefore, the micropropagation process involves four basic stages: the first stage, initial tissue culture; the second stage, tissue culture propagation; the third stage, differentiation and plant formation; and the fourth stage, greenhouse culture and hardening. During the first stage, i.e., the initial tissue culture, the tissue culture is established and proved to be free of contamination. During the second stage, the initial tissue culture will multiply until a sufficient number of tissue samples are produced from the seedlings to meet the production goals. During the third stage, the tissue samples grown in the second stage are divided and grown into individual plantlets. During the fourth stage, the transformed plantlets are transferred to a greenhouse for hardening, wherein the plant's tolerance to light is gradually improved so that it can be grown in a natural environment.

[0253] According to some embodiments of the invention, transgenic plants are generated by transient transformation of leaf cells, meristem cells or the whole plant.

[0254] Transient transformation can be achieved by any of the direct DNA transfer methods described above or by viral infection using modified plant viruses.

[0255] Viruses that have been shown to be useful for transforming plant hosts include CaMV, tobacco mosaic virus (TMV), brome mosaic virus (BMV), and bean common mosaic virus (BV or BCMV). The use of plant viruses to transform plants is described in U.S. Pat. No. 4,855,237 (bean golden mosaic virus; BGV), EP-A 67,553 (TMV), Japanese published application No. 63-14693 (TMV), EPA 194,809 (BV), EPA 278,667 (BV); and Gluzman, Y. et al., Communications in Molecular Biology: Viral Vectors, Cold Spring Harbor Laboratory, New York, pp. 172-189 (1988). Pseudoviral particles for expressing foreign DNA in many hosts including plants are described in WO 87 / 06261.

[0256] According to some embodiments of the present invention, the virus used for transient transformation is avirulent and therefore cannot cause severe symptoms, such as reduced growth rate, mosaic disease, ring spot, leaf curl, yellowing, streaking, pockmark formation, tumor formation and pitting. Suitable avirulent viruses can be naturally occurring avirulent viruses or artificially attenuated viruses. Virus attenuation can be achieved by using methods well known in the art, including but not limited to sublethal heating, chemical treatment or by directed mutagenesis techniques, such as, for example, described by Kurihara and Watanabe (Molecular Plant Pathology 4: 259-269, 2003), Gal-on et al. (1992), Atreya et al. (1992) and Huet et al. (1994).

[0257] Suitable virus strains can be obtained from available sources, such as, for example, the American Type culture Collection (ATCC) or by isolation from infected plants. Isolation of viruses from infected plant tissues can be achieved by techniques well known in the art, such as, for example, described by: Foster and Taylor, Eds. "Plant Virology Protocols: From Virus Isolation to Transgenic Resistance (Methods in Molecular Biology (Humana Pr), Vol 81)", Humana Press, 1998. Briefly, tissues of infected plants believed to contain high concentrations of suitable viruses, preferably young leaves and petals, are ground in a buffer solution (e.g., phosphate buffered solution) to produce a virus-infected sap that can be used for subsequent inoculation.

[0258] The construction of plant RNA viruses for introducing and expressing non-viral exogenous polynucleotide sequences in plants is demonstrated by the above references and by the following: Dawson, WO et al., Virology (1989) 172: 285-292; Takamatsu et al. EMBO J. (1987) 6: 307-311; French et al. Science (1986) 231: 1294-1297; Takamatsu et al. FEBS Letters (1990) 269: 73-76; and U.S. Patent No. 5,316,931.

[0259] When the virus is a DNA virus, the virus itself can be suitably modified. Alternatively, the virus can first be cloned into a bacterial plasmid to facilitate the construction of a desired viral vector with exogenous DNA. The virus can then be excised from the plasmid. If the virus is a DNA virus, the bacterial origin of replication can be attached to the viral DNA and then replicated by bacteria. The transcription and translation of this DNA will produce a coating protein that will encapsulate the viral DNA. If the virus is an RNA virus, the virus is usually cloned as cDNA and inserted into a plasmid. All structures are then constructed using plasmids. RNA viruses are then produced by transcribing the viral sequence of the plasmid and translating the viral gene to produce one or more coating proteins that encapsulate viral RNA.

[0260] In one embodiment, a plant viral polynucleotide is provided, wherein the natural coat protein coding sequence has been deleted from the viral polynucleotide, and the non-natural plant viral coat protein coding sequence and the non-natural promoter, preferably the subgenomic promoter of the non-natural coat protein coding sequence, can express and package the recombinant plant viral polynucleotide in the plant host, and ensure the systemic infection of the inserted recombinant plant viral polynucleotide to the host. Alternatively, the coat protein gene can be inactivated by inserting the non-natural polynucleotide sequence therein, thereby producing the protein. The recombinant plant viral polynucleotide can contain one or more additional non-natural subgenomic promoters. Each non-natural subgenomic promoter can transcribe or express adjacent genes or polynucleotide sequences in the plant host and cannot recombine with each other and cannot recombine with the natural subgenomic promoter. If more than one polynucleotide sequence is included, the non-natural (exogenous) polynucleotide sequence can be inserted adjacent to the natural plant viral subgenomic promoter or the natural and non-natural plant viral subgenomic promoter. The non-natural polynucleotide sequence is transcribed or expressed in the host plant under the control of the subgenomic promoter to produce the desired product.

[0261] In a second embodiment, a recombinant plant viral polynucleotide is provided as in the first embodiment, except that a native coat protein coding sequence is placed adjacent to one of the non-native coat protein subgenomic promoters instead of the non-native coat protein coding sequence.

[0262] In a third embodiment, a recombinant plant viral polynucleotide is provided in which a native coat protein gene is adjacent to its subgenomic promoter and one or more non-native subgenomic promoters have been inserted into the viral polynucleotide. The inserted non-native subgenomic promoters are capable of transcribing or expressing adjacent genes in a plant host and are incapable of recombining with each other and with the native subgenomic promoters. The non-native polynucleotide sequence may be inserted adjacent to the non-native subgenomic plant viral promoter so that the sequence is transcribed or expressed in the host plant under the control of the subgenomic promoter to produce the desired product.

[0263] In a fourth embodiment, a recombinant plant viral polynucleotide is provided as in the third embodiment, except that the native coat protein encoding sequence is replaced with a non-native coat protein encoding sequence.

[0264] The viral vector is coated with a coating protein encoded by a recombinant plant viral polynucleotide to produce a recombinant plant virus. The recombinant plant viral polynucleotide or recombinant plant virus is used to infect an appropriate host plant. The recombinant plant viral polynucleotide is capable of replicating in the host, spreading systemically in the host, and transcribing or expressing one or more exogenous genes (exogenous polynucleotides) in the host to produce a desired protein.

[0265] Techniques for inoculating viruses into plants can be found in: Foster and Taylor, eds. “Plant Virology Protocols: From Virus Isolation to Transgenic Resistance (Methods in Molecular Biology (Humana Pr), Vol 81)”, Humana Press, 1998; Maramorosh and Koprowski, eds. “Methods in Virology” 7 vols, Academic Press, New York 1967-1984; Hill, SA “Methods in Plant Virology”, Blackwell, Oxford, 1984; Walkey, DGA “Applied Plant Virology”, Wiley, New York, 1985; and Kado and Agrawa, eds. “Principles and Techniques in Plant Virology”, Van Nostrand-Reinhold, New York.

[0266] In addition to the above, the polynucleotide of the present invention may also be introduced into the chloroplast genome, thereby enabling chloroplast expression.

[0267] The technology for introducing exogenous polynucleotide sequences into the chloroplast genome is known. The technology involves the following procedures. First, plant cells are chemically treated to reduce the number of chloroplasts per cell to about one. Then, exogenous polynucleotides are introduced into the cells via particle bombardment, with the purpose of introducing at least one exogenous polynucleotide molecule into the chloroplast. The exogenous polynucleotide selected allows it to be integrated into the genome of the chloroplast via homologous recombination, which is easily affected by enzymes inherent to the chloroplast. For this reason, in addition to the target gene, the exogenous polynucleotide also includes at least one polynucleotide fragment derived from the chloroplast genome. In addition, the exogenous polynucleotide includes a selectable marker, which determines by a sequential selection procedure that all or substantially all copies of the chloroplast genome after such selection will include the exogenous polynucleotide. Other details related to the technology can be found in U.S. Patents Nos. 4,945,050 and 5,693,507, which are incorporated herein by reference. Therefore, the polypeptide can be produced by the protein expression system of the chloroplast and integrated into the inner membrane of the chloroplast.

[0268] As used herein, "down-regulate" or "reduced" refers to a decrease in the expression level of a gene of Table 1 and / or 2 in a plant or part thereof by at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 2-fold, at least about 3-fold, at least about 5-fold, at least about 10-fold, compared to a control plant of the same species of the present invention grown under the same (e.g., identical) growth conditions.

[0269] Down-regulation of the transcription or translation products of endogenous genes (gene silencing) can be achieved by methods well known in the art, such as genome editing, homologous recombination and co-suppression, antisense inhibition, RNA interference and ribozyme molecules.

[0270] Alternatively, nucleic acid agents as described in detail below are provided, for example, which downregulate the expression of at least one gene of Table 1.

[0271] Please note that this article provides methods for genome editing and gene expression manipulation at the genome level in the document.

[0272] Cosuppression (sense suppression) - Suppression of endogenous genes can be achieved by cosuppression using RNA molecules (or expression vectors encoding them) that are in a sense orientation relative to the transcription direction of the endogenous gene. The polynucleotide used for cosuppression can correspond to all or part of the sequence encoding the endogenous polypeptide and / or to all or part of the 5' and / or 3' untranslated regions of the endogenous transcript; it can also be an RNA that is not polyadenylated; RNA that lacks a 5' cap structure; or RNA that contains unspliced ​​introns. In some embodiments, the polynucleotide used for cosuppression is designed to eliminate the start codon of the endogenous polynucleotide so that the protein product will not be translated. Cosuppression methods using full-length cDNA sequences as well as partial cDNA sequences are known in the art (see, for example, U.S. Patent No. 5,231,020).

[0273] According to some embodiments of the present invention, downregulation of endogenous genes is performed using an amplicon expression vector comprising a sequence derived from a plant virus, which contains all or part of the target gene, but generally not all the genes of the native virus. The viral sequence present in the transcription product of the expression vector allows the transcription product to direct its own replication. The transcript produced by the amplicon can be sense or antisense relative to the target sequence [see, e.g., Angell and Baulcombe, (1997) EMBO J. 16: 3675-3684; Angell and Baulcombe, (1999) Plant J. 20: 357-362 and U.S. Pat. No. 6,646,805, each of which is incorporated herein by reference].

[0274] Antisense suppression - Antisense suppression can be performed using antisense polynucleotides or expression vectors designed to express RNA molecules that are complementary to all or part of a messenger RNA (mRNA) encoding an endogenous polypeptide and / or to all or part of the 5' and / or 3' untranslated regions of an endogenous gene. Overexpression of antisense RNA molecules can lead to reduced expression of natural (endogenous) genes. Antisense polynucleotides can be completely complementary to the target sequence (i.e., 100% identical to the complement of the target sequence) or partially complementary to the target sequence (i.e., less than 100% identical, e.g., less than 90%, less than 80% identical to the complement of the target sequence). Antisense suppression can be used to inhibit the expression of multiple proteins in the same plant (see, e.g., U.S. Patent No. 5,942,657). In addition, portions of antisense nucleotides can be used to disrupt the expression of target genes. In general, sequences of at least about 50 nucleotides, at least about 100 nucleotides, at least about 200 nucleotides, at least about 300, at least about 400, at least about 450, at least about 500, at least about 550 or more can be used. Methods of using antisense suppression to inhibit endogenous gene expression in plants are described, for example, in Liu, et al., (2002) Plant Physiol. 129: 1732-1743 and U.S. Patents 5,759,829 and 5,942,657, each of which is incorporated herein by reference. The efficiency of antisense suppression can be increased by including a poly dT region at the 3' position of the antisense sequence and at the 5' position of the polyadenylation signal in the expression cassette [see U.S. Patent Publication No. 20020048814, incorporated herein by reference].

[0275] RNA interference - RNA interference can be achieved using a polynucleotide that can anneal to itself and form a double-stranded RNA with a stem-loop structure (also called a hairpin structure), or using two polynucleotides that form a double-stranded RNA.

[0276] For hairpin RNA (hpRNA) interference, an expression vector is designed to express an RNA molecule that hybridizes to itself to form a hairpin structure consisting of a single-stranded loop region and a base-paired stem.

[0277] In some embodiments of the invention, the base-paired stem region of the hpRNA molecule determines the specificity of RNA interference. In this configuration, the sense sequence of the base-paired stem region can correspond to all or part of the endogenous mRNA to be downregulated, or to a portion of the promoter sequence that controls the expression of the endogenous gene to be inhibited; and the antisense sequence of the base-paired stem region is fully or partially complementary to the sense sequence. Such hpRNA molecules are highly effective in inhibiting the expression of endogenous genes in a manner that is inherited by subsequent generations of plants [see, e.g., Chuang and Meyerowitz, (2000) Proc. Natl. Acad. Sci. USA 97:4985-4990; Stoutjesdijk, et al., (2002) Plant Physiol. 129:1723-1731; and Waterhouse and Helliwell, (2003) Nat. Rev. Genet. 4:29-38; Chuang and Meyerowitz, (2000) Proc. Natl. Acad. Sci. USA 97:4985-4990; Pandolfini et al., BMC Biotechnology 3:7; Panstruga, et al., (2003 ... al., (2003) Mol. Biol. Rep. 30: 135-140; and U.S. Patent Publication No. 2003 / 0175965; each of which is incorporated by reference].

[0278] According to some embodiments of the invention, the length of the sense sequence of the base-paired stem is about 10 nucleotides to about 2,500 nucleotides, such as about 10 nucleotides to about 500 nucleotides, such as about 15 nucleotides to about 300 nucleotides, such as about 20 nucleotides to about 100 nucleotides, for example, or about 25 nucleotides to about 100 nucleotides.

[0279] According to some embodiments of the invention, the antisense sequence of the base-paired stem may have a length that is shorter, the same as, or longer than the length of the corresponding sense sequence.

[0280] According to some embodiments of the invention, the loop portion of the hpRNA may have a length of about 10 nucleotides to about 500 nucleotides, such as a length of about 15 nucleotides to about 100 nucleotides, about 20 nucleotides to about 300 nucleotides, or about 25 nucleotides to about 400 nucleotides.

[0281] According to some embodiments of the invention, the loop portion of the hpRNA may include an intron (ihpRNA), which is capable of splicing in the host cell. The use of introns minimizes the size of the loop in the hairpin RNA molecule after splicing and thus increases the efficiency of interference [see, e.g., Smith, et al., (2000) Nature 407:319-320; Wesley, et al., (2001) Plant J. 27:581-590; Wang and Waterhouse, (2001) Curr. Opin. Plant Biol. 5:146-150; Helliwell and Waterhouse, (2003) Methods 30:289-295; Brummell, et al. (2003) Plant J. 33:793-800; and U.S. Patent Publication Nos. 2003 / 0180945; WO 98 / 53083; WO 99 / 32619; WO 98 / 36083; WO 99 / 53050; US20040214330; US20030180945; US Patent No. 5,034,323; US Patent No. 6,452,067; US Patent No. 6,777,588; US Patent No. 6,573,099 and US Patent No. 6,326,527; each of which is incorporated herein by reference].

[0282] In some embodiments of the invention, the loop region of the hairpin RNA determines the specificity of RNA interference to its target endogenous RNA. In this configuration, the loop sequence corresponds to all or part of the endogenous messenger RNA of the target gene. See, e.g., WO02 / 00904; Mette, et al., (2000) EMBO J 19:5194-5201; Matzke, et al., (2001) Curr. Opin. Genet. Devel. 11:221-227; Scheid, et al., (2002) Proc. Natl. Acad. Sci., USA 99:13659-13662; Aufsaftz, et al., (2002) Proc. Nat'l. Acad. Sci. 99(4):16499-16506; Sijen, et al., Curr. Biol. (2001) 11:436-440), each of which is incorporated herein by reference.

[0283] For double-stranded RNA (dsRNA) interference, sense and antisense RNA molecules can be expressed in the same cell from a single expression vector (which includes the sequences of both strands) or from two expression vectors (each including the sequence of one of the strands). Methods for inhibiting endogenous plant gene expression using dsRNA interference are described in Waterhouse, et al., (1998) Proc. Natl. Acad. Sci. USA 95: 13959-13964; and WO 99 / 49029, WO 99 / 53050, WO 99 / 61631, and WO 00 / 49035; each of which is incorporated herein by reference.

[0284] According to some embodiments of the present invention, RNA interference is achieved using an expression vector designed to express an RNA molecule modeled on an endogenous microRNA (miRNA) gene. MicroRNA (miRNA) is a regulator composed of about 22 ribonucleotides and can efficiently inhibit the expression of endogenous genes [Javier, et al., (2003) Nature 425: 257-263]. The miRNA gene encodes an RNA that forms a hairpin structure that contains a 22-nucleotide sequence complementary to an endogenous target gene.

[0285] Ribozymes - catalytic RNA molecules - ribozymes - are designed to cleave specific mRNA transcripts, thereby preventing the expression of the polypeptide encoded therein. Ribozymes cleave mRNA at site-specific recognition sequences. For example, "hammerhead ribozymes" (see, e.g., U.S. Pat. No. 5,254,678) cleave mRNA at positions specified by flanking regions that form complementary base pairs with the target mRNA. The only requirement is that the target RNA contains a 5'-UG-3' nucleotide sequence. Hammerhead ribozyme sequences can be embedded in stable RNAs such as transfer RNAs (tRNAs) to increase in vivo cleavage efficiency [Perriman et al. (1995) Proc. Natl. Acad. Sci. USA, 92(13): 6175-6179; de Feyter and Gaudron Methods in Molecular Biology, Vol. 74, Chapter 43, "Expressing Ribozymes in Plants", Edited by Turner, PC, Humana Press Inc., Totowa, NJ; U.S. Pat. No. 6,423,885]. RNA endoribonucleases such as those found in Tetrahymena thermophila are also useful ribozymes (U.S. Pat. No. 4,987,071).

[0286] Plant lines transformed with any of the above-described down-regulating molecules are screened to identify plant lines that exhibit the greatest suppression of the endogenous polypeptide of interest and thereby enhance the desired plant trait (e.g., yield, WUE, NUE, FUE and / or ABST).

[0287] For example, according to some embodiments, up-regulating expression is performed by transgenics, DNA or RNA editing and / or breeding.

[0288] According to some embodiments, down-regulating expression is performed by transgenics, DNA or RNA editing, RNA silencing and / or breeding.

[0289] Using the methods described above, plant cells transformed with constructs comprising a variety of different exogenous polynucleotides can be regenerated into mature plants.

[0290] Alternatively, multiple exogenous polynucleotides can be expressed in a single host plant by introducing different nucleic acid constructs (comprising different exogenous polynucleotides) into multiple plants. Conventional plant breeding techniques can then be used to cross breed the regenerated transformed plants and select the resulting offspring with superior traits.

[0291] According to some embodiments of the present invention, overexpression of the polypeptide of the present invention is achieved by genome editing.

[0292] Genome editing is a powerful means to affect target traits by modifying target plant genome sequences. Such modifications can produce new or modified alleles or regulatory elements. Therefore, genome editing uses artificially engineered nucleases for reverse genetics to cut and produce specific double-strand breaks at one or more desired locations in the genome, which are then repaired by endogenous cell programs such as homology-directed repair (HDR) and non-homologous end joining (NHEJ). NHEJ directly connects DNA ends at double-strand breaks, while HDR uses homologous sequences as templates to regenerate missing DNA sequences at the breakpoints. In order to introduce specific nucleotide modifications into genomic DNA, a DNA repair template containing the desired sequence must be present during HDR. Genome editing cannot be performed using traditional restriction endonucleases because most restriction enzymes recognize several base pairs on DNA as their targets, and the probability of finding recognized base pair combinations in many locations in the genome is very high, resulting in multiple cuts that are not limited to the desired location. In order to overcome this challenge and produce site-specific single-strand or double-strand breaks, several different classes of nucleases have been discovered and bioengineered to date. These nucleases include meganucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and CRISPR / Cas systems.

[0293] Since most genome editing techniques can leave very small traces of obvious DNA changes in a few nucleotides compared to transgenic plants, crops created by gene editing can avoid the strict regulatory procedures usually associated with the development of genetically modified (GM) crops. On the other hand, traces of genome editing techniques can be used for marker-assisted selection (MAS), as described further below. The target plant of the mutagenesis / genome editing method according to the present invention is any plant of interest, including monocotyledonous or dicotyledonous plants.

[0294] Overexpression of a polypeptide by genome editing can be achieved by: (i) replacing the endogenous sequence encoding the polypeptide of interest or a regulatory sequence under its control, and / or (ii) inserting a new gene encoding the polypeptide of interest in a target region of the genome, and / or (iii) introducing point mutations that result in upregulation of the gene encoding the polypeptide of interest (e.g., by altering regulatory sequences such as promoters, enhancers, 5'-UTRs and / or 3'-UTRs, or mutations in coding sequences).

[0295] Homology-directed repair (HDR)

[0296] Homology-directed repair (HDR) can be used to generate specific nucleotide changes (also known as gene "edits") that range from single nucleotide changes to large insertions. In order to use HDR for gene editing, a DNA "repair template" containing the desired sequence must be delivered to the target cell type with a guide RNA [one or more gRNAs] and Cas9 or Cas9 nickase. The repair template must contain the desired edit as well as additional homologous sequences immediately upstream and downstream of the target (called left and right homology arms). The length and binding position of each homology arm depends on the size of the change being introduced. Depending on the specific application, the repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Notably, the repair template must lack the protospacer adjacent motif (PAM) sequence present in genomic DNA, otherwise the repair template will become a suitable target for Cas9 cleavage. For example, the PAM may mutate so that it no longer exists, but the coding region of the gene is not affected (i.e., a silent mutation).

[0297] Even in cells expressing Cas9, gRNA, and an exogenous repair template, the efficiency of HDR is typically low (<10% modified alleles). Therefore, many laboratories are trying to artificially enhance HDR by synchronizing cells at the cell cycle phase when HDR is most active, or by chemically or genetically inhibiting genes involved in non-homologous end joining (NHEJ). The low efficiency of HDR has several important practical implications. First, because the efficiency of Cas9 cleavage is relatively high and the efficiency of HDR is relatively low, a portion of Cas9-induced double-strand breaks (DSBs) will be repaired via NHEJ. In other words, the resulting cell population will contain some combination of wild-type alleles, NHEJ repair alleles, and / or desired HDR-edited alleles. Therefore, it is important to experimentally confirm the presence of the desired edits and, if necessary, isolate clones containing the desired edits.

[0298] The HDR method has been successfully used for specific modification of coding sequences of targeted plant genes (Budhagatapalli Nagaveni et al. 2015. "Targeted Modification of Gene Function Exploiting Homology-Directed Repair of TALEN-Mediated Double-Strand Breaks in Barley". G3 (Bethesda). 2015 Sep; 5 (9): 1857-1863). Therefore, the gfp-specific transcription activator-like effector nuclease is used together with the repair template to promote the conversion of gfp to yfp via HDR, which is associated with a single amino acid exchange in the gene product. The resulting accumulation of yellow fluorescent protein together with sequencing confirmed the success of genome editing.

[0299] Similarly, Zhao Yongping et al. 2016 (An alternative strategy for targeted gene replacement in plants using a dual-sgRNA / Cas9 design. Scientific Reports 6, Article number: 23890 (2016)) describes the co-transformation of Arabidopsis plants with: a combined dual sgRNA / Cas9 vector that successfully deleted miRNA gene regions (MIR169a and MIR827a); and a second construct containing sites homologous to Arabidopsis TERMINAL FLOWER 1 (TFL1) for homology-directed repair (HDR) with regions corresponding to the two sgRNAs on the modified construct to provide both targeted deletion and donor repair by HDR for targeted gene replacement.

[0300] One example of such an approach involves editing a selected genomic region to express a polypeptide of interest.

[0301] Activating target genes using CRISPR / Cas9

[0302] Many bacteria and archaea contain endogenous RNA-based adaptive immune systems that can degrade nucleic acids of invading phages and plasmids. These systems consist of clustered regularly interspaced short palindromic repeats (CRISPR) genes that produce RNA components and CRISPR-related (Cas) genes that encode protein components. CRISPR RNA (crRNA) contains short fragments homologous to specific viruses and plasmids, and can be used as a guide to instruct Cas nucleases to degrade complementary nucleic acids of corresponding pathogens. Studies on the type II CRISPR / Cas system of Streptococcus pyogenes have shown that three components form RNA / protein complexes, and together are sufficient to achieve sequence-specific nuclease activity: Cas9 nuclease, crRNA containing 20 base pairs homologous to the target sequence, and trans-activating crRNA (tracrRNA) (Jinek et al. Science (2012) 337: 816-821.). It was further demonstrated that a synthetic chimeric guide RNA (gRNA) consisting of a fusion of crRNA and tracrRNA can guide Cas9 to cleave a DNA target complementary to the crRNA in vitro. The study also demonstrated that transient expression of CRISPR-associated endonuclease (Cas9) together with synthetic gRNA can be used to generate targeted double-strand breaks in a variety of different species.

[0303] The CRISPR / Cas9 system is a very flexible genome manipulation tool. A unique feature of Cas9 is its ability to bind to target DNA independent of its ability to cut the target DNA. Specifically, both the RuvC- and HNH-nuclease domains can be inactivated by point mutations (D10A and H840A in SpCas9), resulting in a nuclease-dead Cas9 (dCas9) molecule that is unable to cut the target DNA. The dCas9 molecule retains the ability to bind to the target DNA based on the gRNA targeting sequence. dCas9 can be tagged with a transcriptional activator and these dCas9 fusion proteins can be targeted to promoter regions, resulting in robust transcriptional activation of downstream target genes. The simplest dCas9-based activator consists of dCas9 fused directly to a single transcriptional activator. Importantly, unlike genome modifications induced by Cas9 or Cas9 nickase, dCas9-mediated gene activation is reversible because it does not permanently modify genomic DNA.

[0304] In fact, genome editing has been successfully used to overexpress target proteins in plants, by, for example, mutating regulatory sequences (such as promoters) to overexpress endogenous polynucleotides operably connected to regulatory sequences. For example, Rubio Munoz, Vicente et al. U.S. Patent Application Publication No. 20160102316 (which is incorporated herein by reference in its entirety) describes plants with increased expression of endogenous DDA1 plant nucleic acid sequences, wherein endogenous DDA1 promoters carry mutations introduced by mutagenesis or genome editing, which results in increased DDA1 gene expression, using, for example, CRISPR. The method involves targeting Cas9 to specific genomic loci (in this case, DDA1) via 20 nucleotide guide sequences of a single guide RNA. Online CRISPR design tools can identify suitable target sites (www (.) tools. genome-engineering (.) org. Ran et al. Genome engineering using the CRISPR-Cas9 system nature protocols, VOL. 8 NO. 11, 2281-2308, 2013).

[0305] The CRISPR-Cas system is used to change gene expression in plants, as described in U.S. Patent Application Publication No. 20150067922 of Yang; Yinong et al., which is incorporated herein by reference in its entirety. Thus, the engineered, non-naturally occurring gene editing system includes two regulatory elements, wherein the first regulatory element (a) is operable in plant cells, operably linked to at least one nucleotide sequence encoding a CRISPR-Cas system guide RNA (gRNA) that hybridizes with a target sequence in a plant, and the second regulatory element (b) is operable in plant cells, operably linked to a nucleotide sequence encoding a type II CRISPR-associated nuclease, wherein components (a) and (b) are located on the same or different vectors of the system, so that the guide RNA targets the target sequence and the CRISPR-associated nuclease cuts the DNA molecule, thereby changing the expression of the gene product in the plant. It should be noted that CRISPR-associated nucleases and guide RNAs do not naturally occur together.

[0306] In addition, as described above, point mutations that activate the target gene and / or cause overexpression of the target polypeptide can also be introduced into the plant by means of genome editing. Such mutations can be, for example, deletions of repressor sequences that cause activation of the target gene; and / or insertion of nucleotides and mutations that cause activation of regulatory sequences such as promoters and / or enhancers.

[0307] Meganucleases - Meganucleases are generally divided into four families: LAGLIDADG family, GIY-YIG family, His-Cys box family and HNH family. These families are characterized by structural motifs that affect catalytic activity and recognition sequences. For example, members of the LAGLIDADG family are characterized by having one or two copies of the conserved LAGLIDADG motif. The four meganuclease families are very far apart from each other in terms of conserved structural elements, and therefore also very far apart in terms of DNA recognition sequence specificity and catalytic activity. Meganucleases are common in microbial species and have the unique property of having very long recognition sequences (>14bp), which makes them naturally very specific when cutting at the desired position. This can be used to perform site-specific double-strand breaks in genome editing. Those skilled in the art can use these naturally occurring meganucleases, but the number of such naturally occurring meganucleases is limited. In order to overcome this challenge, mutagenesis and high-throughput screening methods have been used to create meganuclease variants that recognize unique sequences. For example, various meganucleases have been fused to produce hybrid enzymes that recognize new sequences. Alternatively, the DNA-interacting amino acids of the meganuclease can be altered to design sequence-specific meganucleases (see, e.g., U.S. Pat. No. 8,021,867). Meganucleases can be designed using methods described in, e.g., Certo, MT et al. Nature Methods (2012) 9:073-975; U.S. Pat. Nos. 8,304,222; 8,021,867; 8,119,381; 8,124,369; 8,129,134; 8,133,697; 8,143,015; 8,143,016; 8,148,098; or 8,163,514, the contents of each of which are incorporated herein by reference in their entirety. Alternatively, meganucleases with site-specific cleavage characteristics can be obtained using commercially available techniques, such as Precision Biosciences' Directed Nuclease Editor TM Genome editing technology.

[0308] ZFNs and TALENs – Two different classes of engineered nucleases, zinc finger nucleases (ZFNs) and transcription activator-like effector nucleases (TALENs), have been shown to be effective in generating targeted double-strand breaks (Christian et al., 2010; Kim et al., 1996; Li et al., 2011; Mahfouz et al., 2011; Miller et al., 2010).

[0309] Basically, ZFN and TALEN restriction endonuclease technology utilizes non-specific DNA cleavage enzymes connected to specific DNA binding domains (a series of zinc finger domains or TALE repeats, respectively). Typically, a restriction enzyme in which the DNA recognition site and the cleavage site are separated from each other is selected. The cleavage portion is separated and then connected to the DNA binding domain, thereby producing an endonuclease with very high specificity for the desired sequence. An exemplary restriction enzyme with this characteristic is Fok1. In addition, Fok1 has the advantage that dimerization is required to have nuclease activity, and this means that as each nuclease partner recognizes a unique DNA sequence, specificity is significantly increased. In order to enhance this effect, the Fok1 nuclease is engineered so that it can only function as a heterodimer and has an increased catalytic activity. Heterodimer functional nucleases avoid the possibility of unwanted homodimer activity and therefore increase the specificity of double-strand breaks.

[0310] Thus, for example, to target specific sites, ZFNs and TALENs are constructed as nuclease pairs, each member of which is designed to bind to adjacent sequences at the targeted site. After transient expression in cells, the nuclease binds to its target site, and the FokI domain heterodimerizes to produce double-strand breaks. Repair of these double-strand breaks by the non-homologous end joining (NHEJ) pathway typically results in small deletions or small sequence insertions. Since each repair performed by NHEJ is unique, a single nuclease pair can be used to generate an allele series with a series of different deletions at the target site. The length of the deletion typically ranges from a few base pairs to hundreds of base pairs, but larger deletions have been successfully generated in cell culture by using two pairs of nucleases simultaneously (Carlson et al., 2012; Lee et al., 2010). Additionally, when DNA fragments homologous to the target region are introduced together with a nuclease pair, double-strand breaks can be repaired via homology-directed repair, thereby generating specific modifications (Li et al., 2011; Miller et al., 2010; Urnov et al., 2005).

[0311] Although the nuclease parts of both ZFN and TALEN have similar properties, the difference between these engineered nucleases lies in their DNA recognition peptides. ZFN relies on Cys2-His2 zinc fingers, and TALEN relies on TALE. Both DNA recognition peptide domains have the characteristics of being naturally present in their protein combinations. Cys2-His2 zinc fingers are usually present in repetitive sequences 3bp apart, and are present in different combinations in a variety of nucleic acid interaction proteins. On the other hand, TALEs are present in repetitive sequences with a one-to-one recognition ratio between amino acids and recognized nucleotide pairs. Since both zinc fingers and TALEs occur in a repetitive pattern, different combinations can be tried to create various sequence specificities. Methods for making site-specific zinc finger endonucleases include, for example, modular assembly (where zinc fingers associated with triplet sequences are attached in rows to cover the desired sequence), OPEN (low stringency selection of peptide domains with triplet nucleotides, followed by high stringency selection of peptide combinations with the final target in the bacterial system) and bacterial single hybrid screening of zinc finger libraries. ZFNs can also be obtained from, for example, Sangamo Biosciences TM (Richmond, CA) for design and commercial acquisition.

[0312] Methods for designing and obtaining TALEN are described in, for example, Reyon et al. Nature Biotechnology 2012 May; 30(5):460-5; Miller et al. Nat Biotechnol. (2011) 29:143-148; Cermak et al. Nucleic Acids Research (2011) 39(12):e82 and Zhang et al. Nature Biotechnology (2011) 29(2):149-53. Mayo Clinic has introduced a recently developed web-based program called Mojo Hand for designing TAL and TALEN constructs for genome editing applications (accessible via www(.)talendesign(.)org). TALEN can also be obtained from, for example, Sangamo Biosciences. TM (Richmond, CA) for design and commercial acquisition.

[0313] The CRIPSR / Cas system for genome editing consists of two distinct components: guide RNA and an endonuclease, such as Cas9.

[0314] The gRNA is typically a 20-nucleotide sequence encoding a combination of a target homologous sequence (crRNA) and an endogenous bacterial RNA that links the crRNA to the Cas9 nuclease (tracrRNA) in a single chimeric transcript. The gRNA / Cas9 complex is recruited to the target sequence through base pairing between the gRNA sequence and complementary genomic DNA. For successful binding of Cas9, the genomic target sequence must also contain the correct protospacer adjacent motif (PAM) sequence immediately following the target sequence. Binding of the gRNA / Cas9 complex positions Cas9 to the genomic target sequence so that Cas9 can cut both strands of the DNA, resulting in a double-strand break. Just like ZFNs and TALENs, double-strand breaks generated by CRISPR / Cas can undergo homologous recombination or NHEJ.

[0315] Cas9 nuclease has two functional domains: RuvC and HNH, each of which cuts a different DNA strand. When both domains are active, Cas9 causes double-strand breaks in genomic DNA.

[0316] A significant advantage of CRISPR / Cas is that the high efficiency of the system, combined with the ability to easily create synthetic gRNAs, enables the simultaneous targeting of multiple genes. Additionally, most cells carrying mutations have biallelic mutations in the target genes.

[0317] However, the apparent flexibility of the base-pairing interactions between the gRNA sequence and the genomic DNA target sequence allows sequences that do not perfectly match the target sequence to be cleaved by Cas9.

[0318] Modified versions of the Cas9 enzyme that contain a single inactive catalytic domain (RuvC- or HNH-) are called "nickases". The Cas9 nickase has only one active nuclease domain and cuts only one strand of the target DNA, thereby producing a single-strand break or "nick". Single-strand breaks or nicks can usually be quickly repaired by the HDR pathway using an intact complementary DNA strand as a template. However, in what is commonly referred to as a "double-nick" CRISPR system, two adjacent nicks on opposite strands introduced by the Cas9 nickase are treated as double-strand breaks. Double nicks can be repaired by NHEJ or HDR, depending on the desired effect on the gene target. Therefore, if specificity and reduced off-target effects are critical, using the Cas9 nickase to produce double nicks by designing two gRNAs with target sequences that are closely adjacent and located on opposite strands of the genomic DNA will reduce off-target effects, because a single gRNA will only produce nicks that will not alter the genomic DNA.

[0319] A modified version of the Cas9 enzyme that contains two inactive catalytic domains (dead Cas9 or dCas9) has no nuclease activity but can still bind to DNA based on gRNA specificity. dCas9 can be used as a platform for DNA transcription regulators to activate or repress gene expression by fusing the inactive enzyme to known regulatory domains. For example, dCas9 alone binding to a target sequence in genomic DNA interferes with gene transcription.

[0320] There are many publicly available tools that can be used to aid in the selection and / or design of target sequences, as well as bioinformatically determined lists of unique gRNAs for different genes in different species, such as Target Finder from Feng Zhang's lab, Target Finder (E-CRISP) from Michael Boutros' lab, RGEN Tools: Cas-OFFinder, CasFinder: a flexible algorithm for identifying specific Cas9 targets in the genome, and CRISPR Optimal TargetFinder.

[0321] In order to use the CRISPR system, both gRNA and Cas9 should be expressed in the target cell. The insertion vector can contain two cassettes on a single plasmid, or the cassettes can be expressed from two separate plasmids. CRISPR plasmids are commercially available, such as the px330 plasmid from Addgene.

[0322] "Hit and run" or "in-out" - involves a two-step recombination procedure. In the first step, an insertion vector containing a double positive / negative selectable marker box is used to introduce the desired sequence changes. The insertion vector contains a single continuous region homologous to the target locus and is modified to carry the desired mutation. The targeting construct is linearized with a restriction enzyme at a site within the homology region, electroporated into cells, and positively selected to isolate homologous recombinants. These homologous recombinants contain local repeats separated by intervening vector sequences (including selection boxes). In the second step, negative selection is performed on the targeted clones to identify cells that have lost the selection box via intrachromosomal recombination between replicated sequences. Local recombination events eliminate duplications, and depending on the recombination site, the allele either retains the introduced mutation or reverts to the wild type. The end result is the introduction of the desired modification without retaining any exogenous sequence.

[0323] "Double Replacement" or "Tag and Swap" strategy - involves a two-step selection procedure similar to the contact-escape strategy approach, but requires the use of two different targeting constructs. In the first step, a standard targeting vector with 3' and 5' homology arms is used to insert a dual positive / negative selectable cassette near the location where the mutation is to be introduced. After electroporation and positive selection, homology-targeted clones are identified. Next, a second targeting vector containing a region homologous to the desired mutation is electroporated into the targeted clone, and negative selection is applied to remove the selection cassette and introduce the mutation. The final allele contains the desired mutation while eliminating unwanted exogenous sequences.

[0324] Site-specific recombinases - Cre recombinases derived from P1 phage and Flp recombinases derived from yeast Saccharomyces cerevisiae are site-specific DNA recombinases that each recognize a unique 34 base pair DNA sequence (referred to as "Lox" and "FRT", respectively), and sequences flanked by Lox sites or FRT sites can be easily removed via site-specific recombination after expression of Cre or Flp recombinase, respectively. For example, the Lox sequence consists of an asymmetric eight base pair spacer flanked by 13 base pair inverted repeats. Cre recombines 34 base pair lox DNA sequences by binding to 13 base pair inverted repeats and catalyzing chain breaks and reconnections within the spacer. The staggered DNA cuts performed by Cre in the spacer are separated by 6 base pairs to produce overlapping regions that act as homology sensors to ensure that only recombination sites with the same overlapping regions can be recombined.

[0325] Basically, the site-specific recombinase system provides a method for removing the selection cassette after homologous recombination. The system also allows the generation of conditionally altered alleles that can be inactivated or activated in a temporal or tissue-specific manner. It is worth noting that the Cre and Flp recombinases leave a 34-base pair Lox or FRT "scar sequence". The retained Lox or FRT sites are usually left in the introns or 3'UTR of the modified locus, and current evidence suggests that these sites usually do not significantly interfere with gene function.

[0326] Therefore, Cre / Lox and Flp / FRT recombination involves the introduction of a targeting vector with 3' and 5' homology arms containing the desired mutation, two Lox or FRT sequences, and a selectable cassette usually placed between the two Lox or FRT sequences. Positive selection is applied and homologous recombinants containing the targeted mutation are identified. Transient expression of Cre or Flp, together with negative selection, results in excision of the selection cassette and selection of cells in which the cassette has been lost. The final targeted allele contains the Lox or FRT scar sequence of the exogenous sequence.

[0327] Transposase - As used herein, the term "transposase" refers to an enzyme that binds to the ends of a transposon and catalyzes the movement of the transposon to another part of the genome.

[0328] As used herein, the term "transposon" refers to a mobile genetic element comprising a nucleotide sequence that can move around different locations within the genome of a single cell. In the process, the transposon can cause mutations and / or change the amount of DNA in the cell's genome.

[0329] Many transposon systems are also capable of translocation in cells, for example vertebrates have been isolated or engineered, such as Sleeping Beauty [Izsvák and Ivics Molecular Therapy (2004) 9, 147-156], piggyBac [Wilson et al. Molecular Therapy (2007) 15, 139-145], Tol2 [Kawakami et al. PNAS (2000) 97 (21): 11403-11408] or Frog Prince [Miskey et al. Nucleic Acids Res. Dec 1, (2003) 31 (23): 6873-6881]. In general, DNA transposons translocate from one DNA site to another in a simple cut-and-paste manner. Each of these elements has its own advantages, for example, Sleeping Beauty is particularly useful in region-specific mutagenesis, while Tol2 is most likely to integrate into expressed genes. The Hyperactive system can be used for Sleeping Beauty and piggyBac. Most importantly, these transposons have different target site preferences and can therefore introduce sequence changes in overlapping but different sets of genes. Therefore, in order to achieve the best possible gene coverage, it is particularly preferred to use more than one element. The basic mechanism is shared between different transposases, so we will describe it using piggyBac (PB) as an example.

[0330] PB is a 2.5 kb insect transposon originally isolated from the cabbage looper moth (Trichoplusia ni). The PB transposon consists of asymmetric terminal repeats flanking the transposase PBase. PBase recognizes the terminal repeats and induces transposition via a "cut and paste"-based mechanism, and preferentially translocates into the host genome at the tetranucleotide sequence TTAA. After insertion, the TTAA target site is replicated so that the PB transposon is flanked by the tetranucleotide sequence. When activated, PBs typically self-cleave precisely to reestablish a single TTAA site, thereby restoring the host sequence to its pre-transposon state. After cleavage, PBs can be translocated to a new location or permanently lost from the genome.

[0331] In general, the transposase system provides an alternative method for removing the selection box after homologous recombination stops, similar to the use of Cre / Lox or Flp / FRT. Therefore, for example, the PB transposase system involves the introduction of a targeting vector with 3' and 5' homology arms, which contain the target mutation, two PB terminal repeats at the endogenous TTAA sequence site, and a selection box placed between the PB terminal repeats. Positive selection is applied and homologous recombinants containing the targeted mutation are identified. The transient expression of PBase is removed together with negative selection, resulting in the excision of the selection box, and cells in which the box has been lost are selected. The final targeted allele contains the introduced mutation without exogenous sequences.

[0332] In order for a PB to be useful for introducing sequence changes, there must be a native TTAA site relatively close to the position where the specific mutation is to be inserted.

[0333] Genome editing is performed using a recombinant adeno-associated virus (rAAV) platform - the genome editing platform is based on rAAV vectors, which are capable of inserting, deleting, or replacing DNA sequences in the genome of living mammalian cells. The rAAV genome is a single-stranded deoxyribonucleic acid (ssDNA) molecule, either positive or negative sense, and is approximately 4.7 kb in length. These single-stranded DNA viral vectors have high transduction rates and have the unique property of stimulating endogenous homologous recombination in the absence of double-stranded DNA breaks in the genome. One skilled in the art can design rAAV vectors to target desired genomic loci and make gross and / or subtle endogenous gene changes in cells. The advantage of rAAV genome editing is that it targets a single allele and does not result in any off-target genomic changes. rAAV genome editing technology is commercially available, for example from Horizon TM rAAV GENESIS (Cambridge, UK) TM system.

[0334] Methods for identifying efficacy and detecting sequence changes are well known in the art and include, but are not limited to, DNA sequencing, electrophoresis, enzyme-based mismatch detection assays, and hybridization assays such as PCR, RT-PCR, RNase protection, in situ hybridization, primer extension, Southern blotting, Northern blotting, and dot blot analysis.

[0335] Sequence changes in specific genes also can be determined at the protein level, for example using chromatography, electrophoresis, immunodetection assays such as ELISA and Western blot analysis, and immunohistochemistry.

[0336] In addition, those of ordinary skill in the art can easily design knock-in / knock-out constructs including positive and / or negative selection markers for effectively selecting transformed cells that experience homologous recombination events with this construct.Positive selection provides a method for enrichment of clone colonies that have absorbed exogenous DNA.The limiting examples of this type of positive markers include glutamine synthetase, dihydrofolate reductase (DHFR), markers that confer antibiotic resistance, such as neomycin, hygromycin, puromycin and blasticidin S resistance box.Negative selection markers are necessary for selecting random integration and / or eliminating marker sequences (for example positive markers).The limiting examples of this type of negative markers include herpes simplex-thymidine kinase (HSV-TK), hypoxanthine phosphoribosyltransferase (HPRT) and adenine phosphoribosyltransferase (ARPT) that ganciclovir (ganciclovir) (GCV) is converted into cytotoxic nucleoside analogs.

[0337] Once expressed, inhibited, silenced or knocked out in plant cells or whole plants, the level of the polypeptide can be determined by methods well known in the art, such as activity assays, Western blots using antibodies capable of specifically binding to the polypeptide, enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), immunohistochemistry, immunocytochemistry, immunofluorescence, and the like.

[0338] Methods for determining the levels of RNA transcribed from exogenous or endogenous polynucleotides in plants are well known in the art and include, for example, Northern blot analysis, reverse transcription polymerase chain reaction (RT-PCR) analysis (including quantitative, semi-quantitative or real-time RT-PCR), and RNA in situ hybridization.

[0339] The disclosed sequence information and annotation of this teaching can be utilized to support classical breeding.Therefore, the subsequence data of those polynucleotides mentioned above can be used as the marker of marker assisted selection (MAS), and wherein marker is used for indirectly selecting one or more genetic determinants of target traits (for example, physiological disorder).The nucleic acid data (DNA or RNA sequence) of this teaching can comprise following or be connected to following: polymorphic sites or genetic markers on genome, such as restriction fragment length polymorphism (RFLP), microsatellite and single nucleotide polymorphism (SNP), DNA fingerprint analysis (DFP), amplified fragment length polymorphism (AFLP), expression level polymorphism, polymorphism of encoded polypeptide and any other polymorphism at DNA or RNA sequence place.

[0340] It will be appreciated that up- or down-regulation of expression may also be achieved by means of breeding, ie crossing with donor plants having the desired trait.

[0341] Therefore, according to one aspect of the present invention, there is provided a method for producing Eucalyptus, the method comprising:

[0342] identifying a physiological disorder as described herein; and

[0343] Plants are grown that exhibit resistance to the physiological disorder.

[0344] Such plants can be the subject of further breeding and can serve as donor plants for introducing the trait.

[0345] Plants produced according to the present teachings can be used in a commercial setting to produce Eucalyptus-derived products such as pulp and paper products, nanocellulose, microfibrillated cellulose, dissolving pulp, fluff pulp, lumber, wood panels, oils, lignin-derived products, plywood, particleboard, wood-based composites and charcoal.

[0346] Therefore, there is provided a method for producing a processed product of eucalyptus, the method comprising:

[0347] Planting a eucalyptus plant as described herein; and

[0348] Production of pulp, nanocellulose, wood, oil, lignin-derived products, dyes and / or charcoal from Eucalyptus plants.

[0349] According to a specific embodiment, the processed product includes DNA of a plant in which one or more markers can be detected.

[0350] As used herein, the term "about" refers to ± 10%.

[0351] The terms "comprises," "comprising," "includes," "including," "having" and variations thereof mean "including but not limited to."

[0352] The term "consisting of" means "including and limited to."

[0353] The term "consisting essentially of" means that the composition, method or structure may include additional ingredients, steps and / or parts, but only if the additional ingredients, steps and / or parts do not materially change the basic and novel characteristics of the claimed composition, method or structure.

[0354] As used herein, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "compound" or "at least one compound" may include a plurality of compounds, including mixtures thereof.

[0355] Throughout this application, various embodiments of the present invention may be presented in the form of a range. It should be understood that the description in the form of a range is only for convenience and brevity and should not be construed as a strict limitation of the scope of the present invention. Therefore, the description of a range should be deemed to have clearly disclosed all possible sub-ranges and single numerical values ​​within the range. For example, a description of a range such as 1 to 6 should be deemed to clearly disclose sub-ranges, such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., and single numbers within the range, such as 1, 2, 3, 4, 5 and 6. This applies regardless of the width of the range.

[0356] Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integer) within the indicated range. The phrases "ranging / ranges between" a first indicated number and a second indicated number and "ranging / ranges from" a first indicated number "to" a second indicated number are used interchangeably herein and are meant to include the first and second indicated numbers and all decimals and integers therebetween.

[0357] As used herein, the term "method" refers to manners, means, techniques and procedures for accomplishing a given task, including but not limited to those manners, means, techniques and procedures known to practitioners in the fields of chemistry, pharmacology, biology, biochemistry and medicine or readily developable from known manners, means, techniques and procedures.

[0358] When reference is made to a particular sequence listing, such reference should be understood to also encompass sequences that substantially correspond to the complementary haplotype sequence thereof, including minor sequence variations resulting from, for example, sequencing errors or other changes resulting in base substitutions, base deletions or base additions.

[0359] It is to be understood that the polynucleotides provided as single strands in the sequence listing are intended to be in the same sense orientation as the mRNA transcribed from the polynucleotide.

[0360] It should be understood that any sequence identification number (SEQ ID NO) disclosed in this application may refer to a DNA sequence or an RNA sequence, depending on the context in which the SEQ ID NO is mentioned, even if the SEQ ID NO is expressed only in DNA sequence format or RNA sequence format.

[0361] It should be understood that certain features of the invention described in the context of separate embodiments for the sake of clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment for the sake of brevity may also be provided separately or in any suitable subcombination or in any other described embodiment of the invention in an appropriate manner. Certain features described in the context of various embodiments should not be considered essential features of those embodiments unless the embodiment cannot function without those elements.

[0362] Various embodiments and aspects of the present invention as delineated hereinabove and as claimed in the claims section above find experimental support in the following examples.

[0363] Example

[0364] Reference is now made to the following examples, which together with the above descriptions illustrate some embodiments of the invention in a non limiting fashion.

[0365] Materials and methods

[0366] RNA purification and quantification:

[0367] In the physiological disorder hotspot of plantations in Bahia, Brazil, young leaves from a physiological disorder-resistant Eucalyptus parent and a physiological disorder-susceptible Eucalyptus parent and their 10 offspring (5 of which were resistant and 5 were susceptible) were sampled, among which the susceptible clones showed severe disorder symptoms. Samples were collected from 4.5-year-old trees during the onset of symptoms of the relevant physiological disorder (in this case, the disease is the disorder). Total RNA was extracted in triplicate from each tree tested, and RNA was extracted from 50 mg of leaf tissue for each test using the Plant / Fungus Total RNA Purification Kit (Norgen BioTek corp. Cat#25800), and RNA sequencing was performed on the isolated RNA. The read segment data (PE150, 6Gb clean data per sample) of the next generation sequencing of all 36 samples were quantified and analyzed using Geneious Prime software 2022.2 (www(.)geneious(.)com). Reads for each sample were mapped to the Eucalyptus grandis V2.0 CDS sequence (Phytozome 13: www(.)phytozome-next(.)jgi(.)doe(.)gov / ) and quantified using the Geneious Prime "Calculate Expression Levels" module. Transcript per million (TPM) results were normalized by total transcript counts (Wagner et al 2012). The DESeq2 plugin was used to compare expression levels between resistant and susceptible samples and to detect genes that were significantly expressed in either the resistant or susceptible sample set, but not in both.

[0368] Wagner, Günter P., Koryu Kin, and Vincent J. Lynch. "Measurement of mRNAabundance using RNA-seq data: RPKM measure is inconsistent among samples." Theory in biosciences 131.4(2012):281-285.

[0369] Example 1

[0370] Read counts of mRNA reads mapped to all Eucalyptus CDS sequences were used to compare RNA expression levels in six resistant and six susceptible Eucalyptus trees. Expression level comparisons between resistant and susceptible clone groups were analyzed to generate volcano plots (p value = 0.01) ( Figure 1Table 1 lists the genes with a difference greater than 5-fold between sample sets and a p-value less than 0.01.

[0371] Table 1: Genes with significant differentiation in susceptible phenotypes

[0372]

[0373]

[0374] Table 2: Genes with significant differentiation in resistance phenotypes

[0375]

[0376] Example 2

[0377] Phenotypic markers

[0378] The genes in Table 1 show significantly higher expression levels (greater than x5) in susceptible (sensitive) clones grown in plantations affected by physiological disorders. Without being bound by theory, it is suggested that expression levels increase due to the response of the plant to abiotic stress and can even be detected before visual symptoms appear. Therefore, the inventors tested their potential as early and / or late indicators / markers of disorder sensitivity and tested susceptible plants in the field as early as 3-6 months of age and understood the biological triggers of disorder phenomena to identify causal variants and their genetic manipulation.

[0379] The genes in Table 2 show significantly higher expression levels (greater than x5) in resistant plants grown in plantations affected by physiological disorders. Without being bound by theory, it is shown that expression levels increase due to the response of plants to disorder stress and can be detected before visual observation of resistance phenotypes is verified. Therefore, as described above, the inventors tested their potential as early and / or late indicators / markers of disorder resistance and the potential for detecting resistant plants in the field as early as 3-6 months of age.

[0380] Any RNA detection method can be used, such as:

[0381] 1. Reverse transcription and / or real-time PCR / Taqman / Kasps

[0382] 2. Transcriptome Next Generation Sequencing.

[0383] 3. Northern Blot

[0384] 4. Any differential display / expression method.

[0385] Example 3

[0386] Transgenic plants

[0387] The genes in Table 1 show significantly higher expression levels (greater than x5) in susceptible trees grown in plantations affected by the disorder. These genes may affect trees and cause, promote or respond to disease and / or enhance or alleviate symptoms. Since these genes are not expressed or expressed at relatively low levels in resistant trees, these genes are silenced.

[0388] and / or knocked out, thereby reducing disorder symptoms and increasing resistance levels up to complete immunity.

[0389] Biotech transgenic methods used include:

[0390] 1.dsRNA

[0391] 2. Antisense RNA

[0392] 3. Genome editing (such as CRISPR / ZNF / TALENS / meganuclease)

[0393] 4. EMS / irradiation and selection for null mutations.

[0394] 5. Select from the knockout library.

[0395] The genes in Table 2 show significantly higher expression levels (greater than x5) in resistant (tolerant) trees grown in plantations affected by the disorder. These genes affect the trees and cause or promote resistance and / or alleviate symptoms. Since these genes are not expressed or expressed at relatively low levels in susceptible trees, it is contemplated herein to increase transgenic expression in susceptible trees to convert disorder susceptible trees to disorder resistant trees.

[0396] The biotech transgenic promoters used are:

[0397] 1. High / medium / low universal constitutive promoter

[0398] 2. Xylem-specific promoter

[0399] 3. Duct-specific promoter

[0400] 4. Root-specific promoter

[0401] 5. Natural expression cassettes from resistant trees

[0402] Although the invention has been described in conjunction with the specific embodiments of the invention, it is apparent that many substitutions, modifications and variations will be apparent to those skilled in the art. Therefore, it is intended to cover all such substitutions, modifications and variations that fall within the spirit and broad scope of the appended claims.

[0403] It is the intention of one or more applicants that all publications, patents and patent applications mentioned in this specification will be incorporated by reference in their entirety into this specification, just as each individual publication, patent or patent application is specifically and individually indicated when cited that it will be incorporated by reference herein. In addition, the citation or identification of any reference in this application should not be construed as an admission that such reference can be used as prior art of the present invention. In the scope of using section headings, they should not be construed as necessary limitations. In addition, any one or more priority documents of the present application are incorporated by reference herein in their entirety.

Claims

1. A method for identifying a physiological disorder phenotype in Eucalyptus, the method comprising determining the expression level of at least one gene of Table 1 and / or Table 2 in a cell or tissue of a Eucalyptus plant, wherein: Statistically significant upregulation of expression of a gene of Table 1 relative to a resistant control indicates a susceptibility to a physiological disorder phenotype; and / or Statistically significant upregulation of expression of the genes of Table 2 relative to susceptible controls indicates a resistant physiological disorder phenotype.

2. A method for producing a Eucalyptus plant that exhibits resistance to a physiological disorder in Eucalyptus, the method comprising upregulating the expression and / or activity of at least one gene of Table 2 in Eucalyptus, and / or downregulating the expression and / or activity of at least one gene of Table 1 in Eucalyptus, thereby increasing resistance to the phenotype of a physiological disorder in Eucalyptus.

3. A method for producing Eucalyptus, the method comprising: Identifying the physiological disorder of claim 1; Plants are grown that exhibit resistance to the physiological disorder.

4. A method for producing a processed eucalyptus product, the method comprising: Producing eucalyptus plants according to claim 2 or 3; The eucalyptus plant is processed to produce a processed product selected from the group consisting of: paper products, nanocellulose, microfibrillated cellulose, dissolving pulp, fluff pulp, wood, wood boards, oils, lignin derived products, plywood, particleboard, wood based composites, medium density fiberboard (MDF), oils, dyes and charcoal.

5. The method according to any one of claims 1 and 3, wherein: The determination is at the RNA level.

6. The method according to any one of claims 1 and 3, wherein: The determination is at the protein level.

7. The method according to claim 2, wherein: The up-regulated expression is performed by transgenics, DNA or RNA editing and / or breeding.

8. The method according to claim 2, wherein: The down-regulation of expression is performed by transgenics, DNA or RNA editing, RNA silencing and / or breeding.

9. A kit for identifying a physiological disorder phenotype in Eucalyptus, the kit comprising at least one reagent for identifying the expression of at least one gene of Table 1 and / or Table 2, wherein: The kit detects no more than 100, 80, or 50 genes.

10. The kit according to claim 9, wherein The at least one reagent comprises a primer pair or a probe.

11. A nucleic acid construct comprising a nucleic acid sequence encoding a gene of Table 2 or a functional homologue thereof and a heterologous cis-acting regulatory element for driving expression of the gene or functional homologue, such as described herein.

12. A nucleic acid agent that downregulates the expression of at least one gene of Table 1.

Citation Information

Patent Citations

  • Radio paging receiver and method for controlling display auto-reset function

    CN1108068C

  • An RNA plant virus vector or portion thereof, a method of construction thereof, and a method of producing a gene derived product therefrom

    EP0067553A2

  • Plant virus RNA vector

    JP1988014693A

  • Improvement in journal-boxes for cars

    US109159A

  • Improvement in horseshoe-nail machines

    US135140A