Methods for identifying, selecting, and producing crops resistant to southern corn rust

By detecting specific gene alleles on plant chromosome 4 and using recombinant DNA constructs to transform plant cells, the problem of difficulty in identifying southern corn rust resistance in existing technologies was solved, efficient identification and enhanced resistance to SCR were achieved, and the breeding and gene editing of resistant plants were promoted.

CN115335506BActive Publication Date: 2025-09-16PIONEER HI BREED INTERNATIONAL INC +1
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
CN202080077386.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-06
Filing Date
2020-11-05
Publication Date
2025-09-16
Estimated Expiration
2040-11-05

AI Technical Summary

Technical Problem

Existing technologies have not yet been able to effectively identify and characterize the causal genes responsible for southern corn rust resistance, resulting in a lack of targeting in breeding and gene editing, making it difficult to develop plants resistant to southern corn rust.

Method used

By analyzing the gene alleles at specific locations on plant chromosome 4, using PCR and other methods to detect specific nucleotide sequences (such as variations at SEQ ID NO:11-31), screening and hybridization are used to obtain plants with SCR resistance, and resistance is enhanced through transgenic or genome editing methods. Recombinant DNA constructs are used to transform plant cells to change the expression of disease resistance.

Benefits of technology

It has achieved efficient identification and selection of southern corn rust, enhanced plant resistance to SCR, improved the targeting of breeding and gene editing, and promoted the development of resistant plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IDA0003630473270000011
    Figure IDA0003630473270000011
  • Figure IDA0003630473270000021
    Figure IDA0003630473270000021
  • Figure IDA0003630473270000031
    Figure IDA0003630473270000031
Patent Text Reader

Abstract

This field relates to plant breeding and methods for identifying and selecting plants that are resistant to southern corn rust. Methods and uses of identifying novel genes that encode proteins that confer plant resistance to southern corn rust are provided. These disease resistance genes can be used to generate resistant plants through breeding, transgenic modification, or genome editing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to PCT International Application No. PCT / CN2019 / 115925, filed on November 6, 2019, the disclosure of which is expressly incorporated by reference in its entirety. Field of the Invention

[0003] This field relates to plant breeding and methods for identifying and selecting plants that are resistant to southern corn rust. Methods and uses of identifying novel genes that encode proteins that confer plant resistance to southern corn rust are provided. These disease resistance genes can be used to generate resistant plants through breeding, transgenic modification, or genome editing.

[0004] Reference to a sequence listing submitted as a text file via EFS-WEB

[0005] An official copy of the sequence listing is submitted as a text file simultaneously with the description via EFS-Web, which conforms to the American Standard Code for Information Interchange (ASCII) and has the file name RTS21631A_Seq_List.txt, a creation date of October 26, 2020, and a size of 377 Kb. The sequence listing submitted via EFS-Web is part of the description and is hereby incorporated by reference in its entirety.

[0006] background

[0007] Southern corn rust (SCR), a disease caused by Puccinia multifida ( Puccinia polysora Underw), a fungal disease caused by SCR, is a major disease in tropical regions and southern regions of the United States and China. If SCR reaches temperate regions (e.g., the Midwestern United States) at a critical point in the growing season, and if conditions are favorable for rust development, disease intensity can reach epidemic levels very quickly, leading to severe yield losses. Temperate maize germplasm is generally susceptible to SCR. The identification and utilization of resistant lines and QTLs in breeding programs to develop varieties resistant to SCR represents a cost-effective way to control SCR. Alternatively, varieties carrying the genes responsible for SCR resistance can be developed through transgenic or genome editing technologies. The identification of resistance QTLs and genes will accelerate the development of products resistant to SCR.Resistant lines (e.g., Brewbaker, JL, et al. "General resistance in maize to southern rust (Puccinia polysora Underw.)." Crop science 51, no. 4 (2011): 1393-1409) or QTLs (e.g., Jines, MP, et al. "Mapping resistance to Southern rust in a tropical by temperate maize recombinant inbred topcross population." Theoretical and Applied Genetics 114, no. 4 (2007): 659-667. Zhang, Y., et al. "Mapping of southern corn rust-resistant genes in the W2D inbred line of maize (Zea mays L.)." Molecularbreeding 25, no. 3 (2010): 433-439. Zhou CJ, et al. (2007) Characterization and fine mapping of RppQ, a resistance gene to southern corn rust in maize. MolGenet Genomics 278:723–728. Holland, JB, et al. "Inheritance of resistance to southern corn rust in tropical-by-corn-belt maize populations." Theoretical and Applied Genetics 96, no. 2 (1998): 232-241.). However, the causal gene responsible for SCR resistance has not yet been identified and characterized. There is a continuing need for disease-resistant plants and methods for discovering disease-resistance genes.

[0008] Overview

[0009] Provided herein are compositions and methods that can be used to identify and select plant disease resistance genes or "R genes." The compositions and methods can be used to select disease-resistant plants, generate transgenic resistant plants, and / or generate plants with genome-edited resistance. Also provided herein are plants with newly conferred or enhanced resistance to various plant diseases compared to control plants. In some embodiments, the compositions and methods can be used to select southern corn rust (SCR) disease-resistant plants, generate transgenic SCR-resistant plants, and / or generate plants with genome-edited SCR resistance.

[0010] SCR-resistant plants can be crossed with a second plant to obtain progeny plants that possess the resistance gene allele. Disease resistance can be newly conferred or enhanced relative to control plants that do not possess the favorable allele. SCR R gene alleles can be further refined to a chromosomal interval defined by and including defined markers. In some embodiments, methods for identifying and / or selecting plants that are resistant to SCR are presented. In these methods, the DNA of the plant is analyzed for the presence of a resistance gene allele on chromosome 4 that is associated with SCR resistance, wherein the resistance gene allele comprises a "T" at PM01-000058W (position 99 of reference sequence SEQ ID NO: 11), a "G" at SOURST-83_1284720 (position 51 of reference sequence SEQ ID NO: 21), a "C" at SOURST-83_1314662 (position 24 of reference sequence SEQ ID NO: 19), a "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), an "A" at PZE-104005694 (position 25 of reference sequence SEQ ID NO: 23), a "C" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), a "G" at SOURST-83_1936804 (position 33 of reference sequence SEQ ID NO: 24), a "C" at SOURST-83_1936805 (position 47 of reference sequence SEQ ID NO: 25), a "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: NO: 32, position 30), "C" at PZE-104001404 (position 26 of reference sequence SEQ ID NO: 25), "G" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "C" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), "A" at SOURST-83_2036602 (position 24 of reference sequence SEQ ID NO: 20), "T" at SOURST-83_2035716 (position 35 of reference sequence SEQ ID NO: 28), "T" at PZE-104001592 (position 46 of reference sequence SEQ ID NO: 29), "C" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "G" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), In some embodiments, the method for identifying and / or selecting plants resistant to SCR comprises detecting or selecting for a genomic region comprising SEQ ID NO: 9 or 10. SCR resistance can be newly conferred or enhanced relative to a control plant that does not have the favorable allele.In a further embodiment, the SCR resistance region comprises a coding conferring or enhancing pair. SCR In some embodiments, the polypeptide comprises the amino acid sequence shown in SEQ ID NO: 5, 6, 7 or 8.

[0011] In another embodiment, a method for identifying and / or selecting plants with SCR resistance is provided, wherein one or more marker alleles linked to and associated with any of the following: "T" at PM01-000058W (position 99 of reference sequence SEQ ID NO: 11), "G" at SOURST-83_1284720 (position 51 of reference sequence SEQ ID NO: 21), "C" at SOURST-83_1314662 (position 24 of reference sequence SEQ ID NO: 19), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at PZE-104005694 (position 25 of reference sequence SEQ ID NO: 23), "C" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), "G" at SOURST-83_1936804 (position 33 of reference sequence SEQ ID NO: 25), "C" at SOURST-83_1936805 (position 47 of reference sequence SEQ ID NO: 26), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_1 NO: 32, position 30), "C" at PZE-104001404 (position 26 of reference sequence SEQ ID NO: 25), "C" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "G" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), "A" at SOURST-83_2036602 (position 24 of reference sequence SEQ ID NO: 20), "T" at SOURST-83_2035716 (position 35 of reference sequence SEQ ID NO: 28), "T" at PZE-104001592 (position 46 of reference sequence SEQ ID NO: 29), "C" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "G" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), In some embodiments, the present invention provides a method for selecting plants having one or more marker alleles, wherein the one or more marker alleles are linked at 10 cM, 9 cM, 8 cM, 7 cM, 6 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.9 cM, 0.8 cM, 0.7 cM, 0.6 cM, 0.5 cM, 0.4 cM, 0.3 cM, 0.2 cM or 0.1 cM or less on a genetic map based on a single meiotic division.The selected plant can be crossed with a second plant to obtain a progeny plant having one or more marker alleles linked to and associated with any of the following: "T" at PM01-000058W (position 99 of reference sequence SEQ ID NO: 11), "G" at SOURST-83_1284720 (position 51 of reference sequence SEQ ID NO: 21), "C" at SOURST-83_1314662 (position 24 of reference sequence SEQ ID NO: 19), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at PZE-104005694 (position 25 of reference sequence SEQ ID NO: 23), "C" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), "G" at SOURST-83_1936804 (position 33 of reference sequence SEQ ID NO: 24), "C" at SOURST-83_1936805 (position 47 of reference sequence SEQ ID NO: 25), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_1 NO: 32, position 30), "C" at PZE-104001404 (position 26 of reference sequence SEQ ID NO: 25), "C" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "G" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), "A" at SOURST-83_2036602 (position 24 of reference sequence SEQ ID NO: 20), "T" at SOURST-83_2035716 (position 35 of reference sequence SEQ ID NO: 28), "T" at PZE-104001592 (position 46 of reference sequence SEQ ID NO: 29), "C" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "G" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), NO:30, position 51) and "G" at SOURST-83_2679982 (reference sequence SEQ ID NO:31, position 26).

[0012] In another embodiment, methods for introgressing gene alleles associated with SCR resistance are presented herein. In these methods, a population of plants is screened with one or more markers to determine whether any plant has a gene allele associated with SCR resistance, and at least one plant having a gene allele associated with SCR resistance is selected from the population. The gene alleles include: "T" at PM01-000058W (position 99 of reference sequence SEQ ID NO: 11), "G" at SOURST-83_1284720 (position 51 of reference sequence SEQ ID NO: 21), "C" at SOURST-83_1314662 (position 24 of reference sequence SEQ ID NO: 19), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), "C" at SOURST-83_1936804 (position 30 of reference sequence SEQ ID NO: 32), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), "C" at SOURST-83_1936804 (position 30 of reference sequence SEQ ID NO: 32), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: NO: 25), "G" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "C" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), "A" at SOURST-83_2036602 (position 24 of reference sequence SEQ ID NO: 20), "T" at SOURST-83_2035716 (position 35 of reference sequence SEQ ID NO: 28), "T" at PZE-104001592 (position 46 of reference sequence SEQ ID NO: 29), "G" at SOURST-83_2465654 (position 51 of reference sequence SEQ ID NO: 30), and "T" at SOURST-83_2679982 (position 26 of reference sequence SEQ ID NO: 31).

[0013] In some embodiments, introgression of SCR resistance genes from resistant lines into susceptible lines can be achieved through marker-assisted trait introgression, transgenics, or genome editing approaches.

[0014] Embodiments include isolated polynucleotides comprising a coding sequence capable of conferring SCRThe invention also provides a nucleotide sequence of a polypeptide that confers resistance to influenza virus infection, wherein the polypeptide encodes an amino acid sequence that is at least 50%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to any one of SEQ ID NOs: 5-8.

[0015] Additional embodiments of the present disclosure include vectors comprising a polynucleotide of the present disclosure (e.g., any one of SEQ ID NOs: 1-4), or recombinant DNA constructs comprising a polynucleotide disclosed herein operably linked to at least one regulatory sequence. Also included are plant cells and plants each comprising a recombinant DNA construct of an embodiment disclosed herein, and seeds comprising a recombinant DNA construct.

[0016] In some embodiments, the compositions and methods relate to modified plants with enhanced resistance to disease, wherein the allele that causes the enhanced disease resistance comprises a nucleotide sequence encoding an SCR resistance gene, wherein the SCR resistance gene has at least 50%, at least 75%, at least 80%, at least 85%, at least 90% and at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity to the sequence shown in any one of SEQ ID NOs: 5-8.

[0017] The present disclosure encompasses methods relating to methods for transforming host cells, including plant cells, comprising transforming the host cells with a polynucleotide of an embodiment of the present disclosure; methods for producing plants, comprising transforming plant cells with a recombinant DNA construct of an embodiment of the present disclosure, and regenerating plants from the transformed plant cells; and methods for conferring or enhancing disease resistance, comprising transforming plants with a recombinant DNA construct disclosed herein.

[0018] Also included are methods for altering the expression level of a protein capable of conferring disease resistance in a plant or plant cell, comprising (a) transforming a plant cell with a recombinant DNA construct disclosed herein, and (b) growing the transformed plant cell under conditions suitable for expression of the recombinant DNA construct, wherein expression of the recombinant DNA construct results in altered production levels of the protein capable of conferring disease resistance in the transformed host.

[0019] Also provided are plants identified and / or selected using any of the methods presented above.

[0020] Sequence Description

[0021] Table 1. Sequence description

[0022] SEQ ID NO: describe 1 NLR01-3_CDS 2 NLR01-2_CDS 3 NLR01-1_CDS 4 NLR02_CDS 5 CIMBL83_NLR01-3 protein 6 CIMBL83_NLR01-2 protein 7 CIMBL83_NLR01-1 protein 8 NLR02 protein 9 NLR01_genome (including UTRs) 10 NLR02_Genome 11 PM01-000058W 12 PM01-00002MH 13 SOURST-83_1314662 primer 14 SOURST-83_1314662 primer 15 SOURST-83_1314662 probe 16 SOURST-83_2036602 primer 17 SOURST-83_2036602 primer 18 SOURST-83_2036602 probe 19 SOURST-83_1314662 reference sequence 20 SOURST-83_2036602 reference sequence 21 SOURST-83_1284720 reference sequence 22 SOURST-83_1542053 reference sequence 23 PZE-104005694 reference sequence 24 SOURST-83_1II reference sequence 25 PZE-104001404 reference sequence 26 SOURST-83_1926276 reference sequence 27 SOURST-83_1652968 reference sequence 28 SOURST-83_2035716 reference sequence 29 PZE-104001592 reference sequence 30 SOURST-83_2465654 reference sequence 31 SOURST-83_2679982 reference sequence 32 SOURST-83_1936804 reference sequence 33 NLR01-4_CDS 34 NLR01-5_CDS 35 NLR01-6_CDS 36 NLR01-7_CDS 37 NLR01-8_CDS 38 NLR01-9_CDS 39 NLR01-10_CDS 40 NLR01-11_CDS 41 NLR01-12_CDS 42 NLR01-13_CDS 43 NLR01-14_CDS 44 NLR01-15_CDS 45 NLR01-16_CDS 46 NLR01-4_Protein 47 NLR01-5_Protein 48 NLR01-6_Protein 49 NLR01-7_Protein 50 NLR01-8_Protein 51 NLR01-9_Protein 52 NLR01-10_Protein 53 NLR01-11_Protein 54 NLR01-12_Protein 55 NLR01-13_Protein 56 NLR01-14_Protein 57 NLR01-15_Protein 58 NLR01-16_Protein

[0023] Details

[0024] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a plurality of such cells, and reference to "a protein" includes reference to one or more proteins and equivalents thereof, and so forth. All technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs, unless expressly indicated otherwise.

[0025] The NBS-LRR ("NLR") group of R-genes is the largest class of R-genes discovered to date. In Arabidopsis, more than 150 are expected to be present in the genome (Meyers, et al., (2003), Plant Cell, 15:809-834; Monosi, et al., (2004), Theoretical and Applied Genetics, 109:1434-1447), while in rice, approximately 500 NLR genes have been predicted (Monosi, (2004) supra). The NBS-LRR class of R genes consists of two subclasses. Class 1 NLR genes contain a TIR-Toll / interleukin-1-like domain at their N' termini; to date, they have only been found in dicots (Meyers, (2003) supra; Monosi, (2004) supra). The second type of NBS-LRR contains a coiled-coil domain or (nt) domain at its N-terminus (Bai et al. (2002) Genome Research, 12:1871-1884; Monosi, (2004) supra; Pan et al. (2000) Journal of Molecular Evolution, 50:203-213). Type 2 NBS-LRRs are found in both dicotyledonous and monocotyledonous plant species (Bai, (2002) supra; Meyers, (2003) supra; Monosi, (2004) supra; Pan, (2000) supra).

[0026] The NBS domain of this gene appears to play a role in signal transduction of plant defense mechanisms (van der Biezen et al., (1998), Current Biology: CB, 8: R226-R227). The LRR region appears to be a region that interacts with pathogen AVR products (Michelmore, et al., (1998), Genome Res., 8: 1113-1130; Meyers, (2003) ibid). Compared to the NB-ARC (NBS) domain, this LRR region is under much greater selective pressure to achieve diversification (Michelmore, (1998) ibid; Meyers, (2003) ibid; Palomino, et al., (2002), Genome Research, 12: 1305-1315). LRR domains are also found in other contexts; these 20-29 residue motifs are present in tandem arrays in many proteins with various functions (e.g., hormone-receptor interactions, enzyme inhibition, cell adhesion, and cell trafficking). Many recent studies have revealed that LRR proteins are involved in early mammalian development, neural development, cell polarization, regulation of gene expression, and apoptosis signaling.

[0027] An allele is "associated with a trait" when it is part of or linked to a DNA sequence or is an allele that affects the expression of the trait. The presence of the allele is an indicator of how the trait will be expressed.

[0028] As used herein, "disease-resistant" or "having resistance to a disease" refers to a plant that exhibits increased resistance to a disease compared to a control plant. Disease resistance can be manifested as fewer and / or smaller lesions, increased plant health, increased yield, increased root mass, increased plant vigor, less or no discoloration, increased growth, reduced necrotic area, or reduced wilting. In some embodiments, the allele can confer resistance to one or more diseases.

[0029] Diseases affecting corn plants include, but are not limited to, bacterial leaf blight and stalk rot; bacterial leaf spot; bacterial stripe; chocolate spot; Goss's bacterial wilt and blight; brown spot; purple sheath; seed rot-seedling blight; bacterial wilt; corn dwarf; anthracnose leaf blight; anthracnose stalk rot; Aspergillus ear rot; banded leaf and sheath spot; black bunch; black kernel rot; white edge disease (bordeaux). blanco); brown spot; black spot; stalk rot; cephalosporium sclerotium rot; charcoal rot; cortex ear rot; Curvularia leaf spot; Subsporia leaf spot; Chromodiplodia ear and stalk rot; Chromodiplodia ear rot; seed rot; corn seedling blight; Chromodiplodia leaf spot or leaf streak; downy mildew; brown streak downy mildew; crazy top downy mildew; green ear downy mildew; grass downy mildew; Java downy mildew; Philippine downy mildew; sorghum downy mildew; sweetroot downy mildew; sugarcane downy mildew; stem ear rot; ergot; horse's tooth tooth); corn eye spot; Fusarium ear and stem rot; Fusarium wilt; seedling root rot; Gibberellum ear and stem rot; gray ear rot; gray leaf spot; Cercospora leaf spot; Helminthosporium root rot; Cladosporium monosporum ear rot; Cladosporium rot; Penicillium leaf spot; late wilt; northern leaf blight; white blast blast); crown and stalk rot; corn streak; northern leaf spot; Helminthosporium ear rot; Penicillium ear rot; corn blue eye; blue mold; dark spherical stem and root rot; dark spherical leaf spot; cystosporium ear rot; botrytis ear rot; Aspergillus stem and root rot; Pythium root rot; Pythium stem rot; red kernel disease; Rhizoctonia ear rot; sclerotia rot; Rhizoctonia root and stalk rot; rostratum leaf spot; common corn rust; southern corn rust; tropical corn rust; sclerotia ear rot; southern blight; Aspergillus leaf spot; sheath rot; hull rot; silage mildew; common smut; false smut; head smut; southern corn leaf and stalk rot; southern Square leaf spot; tar spot; Trichoderma ear and root rot; white ear rot; root and stem rot; yellow leaf blight; banded leaf spot; American wheat streak (wheat stripe mosaic); barley stripe mosaic; barley yellow dwarf; brome mosaic; cereal chlorotic mottle; lethal necrosis (maize lethal necrosis); cucumber mosaic; Johnsongrass mosaic; corn bushy stunt; corn chlorotic dwarf; corn chlorotic mottle; corn dwarf mosaic; corn leaf spot; corn transparent ring spot; corn red leaf and red stripe; corn red stripe; corn ring spot; corn cob dwarf; corn sterile dwarf; corn mottle; corn stripe; maize tassel abortion; corn vein rise; corn mouse ear; corn white leaf; corn white line mosaic; millet red leaf; and northern grain mosaic.

[0030] Diseases affecting plants include, but are not limited to, bacterial wilt; bacterial leaf streak; root rot; grain rot; brown sheath; wilt; brown spot; crown sheath rot; downy mildew; eyespot; false smut; core smut; leaf smut; leaf scorch; narrow brown leaf spot; root rot; seedling wilt; sheath wilt; sheath rot; sheath spot; Alternaria leaf spot; and stem rot.

[0031] Diseases affecting soybean plants include, but are not limited to, Alternaria leaf spot; anthracnose; black leaf blight; black root rot; brown spot; brown stem rot; charcoal rot; Alternaria leaf blight; downy mildew; umbellifer wilt; frogeye leaf spot; scutellaria leaf spot; mycoleptodiscus root rot; Neotrichos stem rot; Phomopsis seed rot; Phytophthora root and stem rot; Phyllosticta leaf spot; Tuberospora root rot; pod and stem wilt; powdery mildew; purple spot; Echinops leaf spot; Pythium rot; red crown rot; Rhizoctonia leaf spot; Rhizoctonia gas blight; Rhizoctonia root and stem rot disease; rust; scab; sclerotinia stem rot; sclerotinia blight; stem canker; Portobello leaf blight; sudden death syndrome; target spot; yeast spot; lance nematode; lesion nematode; needle nematode; reniform nematode; ring nematode; root knot nematode; sheath nematode; cyst nematode; spiral nematode; thorn nematode; short root nematode; stunting nematode; alfalfa mosaic; bean pod mottle; bean yellow mosaic; Brazilian bud blight; chlorotic mottle; yellow mosaic; peanut mottle; peanut stripe; peanut stunt; chlorotic mottle; wrinkled leaves; stunting; severe stunting; and tobacco ringspot or bud blight.

[0032] Diseases affecting canola plants include, but are not limited to, bacterial black rot; bacterial leaf spot; bacterial pod rot; bacterial soft rot; scab; crown gall; Alternaria black spot; anthracnose; root rot; black mold rot; black root; brown girdling root rot; Cercospora leaf spot; clubroot; downy mildew; Fusarium wilt; gray mold; head rot; leaf spot; light leaf spot; pod rot; powdery mildew; ring spot; root rot; Sclerotinia stem rot; seed rot, damping-off; root gall smut; southern wilt; Verticillium wilt; white blight; white leaf spot; dieback; yellows; crinkle virus; mosaic virus; and yellows virus.

[0033] Diseases affecting sunflower plants include, but are not limited to, root tip chlorosis; bacterial leaf spot; bacterial wilt; crown gall; Erwinia stem and head rot; Iternaria leaf blight, stem spot and head rot; Botrytis head rot; charcoal rot; downy mildew; Fusarium stem rot; Fusarium wilt; Myrocyoma leaf and stem spot; Flask mold yellows; Phoma black stem; Phoma spp. brown stem canker; Tuberospora root rot; Phytophthora stem rot; powdery mildew; Pythium seedling wilt and root rot; Rhizoctonia seedling wilt; Rhizopus head rot; sunflower rust; sclerotium base stem and root rot; Septoria leaf spot; Verticillium wilt; white rust; yellow rust; dagger; pin; lesions; reniform; root knot; and chlorotic mottle.

[0034] Diseases affecting sorghum plants include, but are not limited to, bacterial leaf spot; bacterial leaf spot; bacterial leaf streak; Acremonium wilt; anthracnose; charcoal rot; downy mildew; damping-off and seed rot; ergot; Fusarium head blight, root and stem rot; grain storage mildew; gray leaf spot; trailing leaf spot; leaf blight; milo disease; oval leaf spot; tip rot (pokkah boeng); Pythium root rot; rough leaf spot; rust; seedling blight and seed rot; hard-core smut; head smut; pine-core smut; leaf streak; downy mildew; tar spot; target spot; and banded leaf spot and sheath blight.

[0035] Disease resistant plants can have an increased resistance to disease by 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% compared to control plants. In some embodiments, plants can have an increased plant health in the presence of disease by 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100% compared to control plants.

[0036] As used herein, term " chromosome interval " refers to the genomic DNA scope of continuous linearity on the single chromosome in plant.The genetic element or the gene that are positioned at the single chromosome interval are physically chain.The size of chromosome interval has no concrete restriction.In some respects, the genetic element that is positioned at the single chromosome interval is chain in heredity, and the genetic recombination distance is generally for example less than or equal to 20cM (centimorgan), or alternatively less than or equal to 10cM, 9cM, 8cM, 7cM, 6cM, 5cM, 4cM, 3cM, 2cM or 1 cM.That is, two genetic elements that are positioned at the single chromosome interval are reorganized with being less than or equal to 20% at 20cM, 10% at 10cM or 5% at 5cM frequency.

[0037] In the present application, the phrase "tightly linked" means that the frequency of occurrence of recombination between two linked loci is equal to or less than about 10% (i.e., no more than 10 cM apart on the genetic map). In other words, tightly linked loci co-segregate at least 90% of the time. When marker loci show a significant probability of co-segregation (linkage) with a desired trait (e.g., resistance to southern corn rust), they are particularly useful for the subject matter of the present disclosure. Tightly linked loci such as marker loci and second loci can show 10% or lower, preferably about 9% or lower, still more preferably about 8% or lower, again more preferably about 7% or lower, still more preferably about 6% or lower, again more preferably about 5% or lower, still more preferably about 4% or lower, again more preferably about 3% or lower, and still more preferably about 2% or lower intra-locus recombination frequency. In highly preferred embodiments, the associated loci show a recombination frequency of about 1% or lower, such as about 0.75% or lower, more preferably about 0.5% or lower, or again more preferably about 0.25% or lower recombination frequency. Two loci that are located on the same chromosome and are at a distance therebetween such that the frequency of recombination between the two loci is less than 10% (e.g., about 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or lower) are also referred to as being "close" to each other. In some cases, two different markers can have identical gene map coordinates. In that case, the two markers are so close to each other that the frequency of recombination between them is so low that it cannot be detected.

[0038] The term "hybridized" or "crossing" refers to sexual crossing and involves the fusion of haploid gametes to produce diploid progeny (e.g., cells, seeds, or plants) via pollination. The term includes both pollination of one plant by another and selfing (or self-pollination, e.g., when the pollen and ovules are from the same plant).

[0039] An "elite line" is any line derived from breeding and selection for superior agronomic performance.

[0040] "Exotic varieties," "tropical lines," or "exotic germplasm" are varieties that are derived from plants that are not part of the germplasm of an available elite line or variety. In the case of a hybrid between germplasm from two plants or varieties, the lineage of the exotic germplasm is not closely related to the elite germplasm with which it is hybridized. Most commonly, exotic germplasm is not derived from any known elite line, but rather is selected to introduce new genetic elements (usually new alleles) into a breeding program.

[0041] A "favorable allele" is an allele at a particular locus (marker, QTL, gene, etc.) that confers or contributes to an agronomically desirable phenotype, such as disease resistance, and allows identification of plants having an agronomically desirable phenotype. A favorable allele of a marker is a marker allele that segregates with the favorable phenotype.

[0042] "Genetic marker" is a polymorphic nucleic acid in a colony, and wherein can detect and distinguish their alleles by one or more analytical methods, such as RFLP, AFLP, isozymes, SNP, SSR etc. The term also refers to a nucleic acid sequence complementary to the genomic sequence, such as the nucleic acid used as a probe. The mark corresponding to the genetic polymorphism between the colony members can be detected by the method fully established in this area. These include, for example, sequence-specific amplification methods based on PCR, the detection of restriction fragment length polymorphism (RFLP), the detection of isozyme markers, the detection of polynucleotide polymorphisms (ASH) by allele-specific hybridization, the detection of amplification variable sequences of plant genomes, the detection of self-sustaining sequence replication, the detection of simple sequence repeats (SSR), the detection of single nucleotide polymorphisms (SNPs) or the detection of amplified fragment length polymorphisms (AFLPs). The method fully established is also known for detecting expressed sequence tags (ESTs) and SSR markers and randomly amplified polymorphic DNA (RAPD) derived from EST sequences.

[0043] " germplasm " refers to an individual (for example, plant), a group of individuals (for example, plant strains, varieties or families) or a clone derived from a strain, variety, species or culture, or more specifically, the genetic material of all individuals within a species or several species (for example, corn germplasm collections or Andean germplasm collections), or the genetic material from them. Germplasm can be the part of an organism or cell, or can be separated from an organism or cell. Germplasm usually provides the genetic material with a specific molecular structure that provides the physical basis of some or all of the genetic qualities of an organism or cell culture. As used herein, germplasm includes cells, seeds or tissues from which new plants can be grown, or plant parts such as leaves, stems, pollen or cells that can be cultured into complete plants.

[0044] A "haplotype" is an individual's genotype at multiple loci, i.e., a combination of alleles. The loci described by a haplotype are typically physically and genetically linked, i.e., on the same chromosome segment.

[0045] The term "heterogeneous" is used to indicate that individuals within a group differ in their genotype at one or more specific loci.

[0046] The heterotic response or "heterosis" of a material can be defined by its performance above the average of the parent (or higher parent) when crossed with a different or unrelated population.

[0047] A "heterotic group" comprises a group of genotypes that perform well when crossed with genotypes from a different heterotic group (Hallauer et al. (1998) Corn breeding, pp. 463-564, GF Sprague and JW Dudley (eds.), Corn and corn improvement Inbred lines are classified into heterotic groups and further subdivided into families within heterotic groups based on several criteria such as pedigree, association based on molecular markers, and performance in hybrid combinations (Smith et al. (1990) Theor. Appl. Gen. 80:833-840.) The two most widely used heterotic groups in the United States are called "Iowa Stiff Stalk Synthetic" (also referred to herein as "stiff stalk") and "Lancaster" or "Lancaster Sure Crop" (sometimes referred to as NSS or non-stiff stalk).

[0048] Some heterotic groups have traits that are required to be the female parent, while others have traits that are used for the male parent. For example, in corn, yield results from public inbred lines released from a population called BSSS (Iowa State Hard Stalk Synthesis population) led to these inbred lines and their derivatives becoming the female pool for the central corn belt. BSSS inbred lines have been crossed with other inbred lines (e.g., SD 105 and Maiz Amargo), and this general group of material has been called the Hard Stalk Synthesis (SSS), although not all inbred lines are derived from the original BSSS population (Mikel and Dudley (2006)). Crop Sci :46:1193-1205). By default, all other inbred lines that combined well with the SSS inbred lines have been assigned to the male pool, which, for lack of a better name, is designated NSS, or non-hard stem. This group includes several major heterotic groups, such as Lancaster Surecrop, Iodent, and Leaming Corn.

[0049] The term "homogeneous" means that the members of a group have the same genotype at one or more specific loci.

[0050] The term "hybrid" refers to the progeny obtained from a cross between at least two genetically dissimilar parents.

[0051] The term "inbred line" refers to a line that has been bred to achieve genetic homogeneity.

[0052] The term "indel" refers to an insertion or deletion, wherein one strain may be referred to as having an inserted nucleotide or DNA segment relative to a second strain, or the second strain may be referred to as having a deleted nucleotide or DNA segment relative to the first strain.

[0053] The term "infiltration" refers to the transmission of the desired allele of a locus from a genetic background to another genetic background. For example, the infiltration of the desired allele at a specific locus can be transmitted to at least one filial generation via the sexual hybridization between two parents of the same species, wherein at least one parent has the desired allele in its genome. Alternatively, for example, the transmission of an allele can occur by the recombination between two donor genomes, such as in a fusion protoplast, wherein at least one of the donor protoplasts has the desired allele in its genome. The desired allele can, for example, be detected by a marker associated with the phenotype, at the QTL place, transgenics, etc. The filial generation comprising the desired allele can be repeatedly backcrossed with the strain with the desired genetic background and the desired allele can be selected to produce the allele that becomes fixed in the selection genetic background.

[0054] When the process of "introgression" is repeated two or more times, the process is often called "backcrossing."

[0055] A "line" or "breed" is a group of individuals with the same parents that are generally inbred to some degree and are generally homozygous and homogeneous (isogenic or near isogenic) at most gene loci. A "subline" refers to a subpopulation of inbred lines of progeny that is genetically distinct from other similar subpopulations of inbred lines of the same ancestral origin.

[0056] As used herein, the term "linkage" is used to describe the degree to which a marker locus is associated with another marker locus or some other loci. The linkage relationship between a molecular marker and the locus that affects the phenotype is represented by "probability" or "adjusted probability". Linkage can be expressed as required limit or scope. For example, in some embodiments, when a marker is separated by a single meiotic map (based on a population that has undergone a round of meiosis, such as a genetic map such as F2; ​​The IBM2 map is composed of multiple meiosis) by less than 50, 40, 30, 25, 20, or 15 map units (or cM), any marker is linked (genetically and physically) to any other marker. In some aspects, it is advantageous to limit the linkage range of bracketing, for example, between 10 cM and 20 cM, between 10 cM and 30 cM, or between 10 cM and 40 cM. ​​The tighter the linkage of the marker to the second locus, the better the indication of the marker to the second locus. Thus, "closely linked loci," such as a marker locus and a second locus, exhibit an intralocus recombination frequency of 10% or less, preferably about 9% or less, still more preferably about 8% or less, yet more preferably about 7% or less, still more preferably about 6% or less, yet more preferably about 5% or less, still more preferably about 4% or less, yet more preferably about 3% or less, and still more preferably about 2% or less. In highly preferred embodiments, the associated loci exhibit a recombination frequency of about 1% or less, e.g., about 0.75% or less, more preferably about 0.5% or less, or yet more preferably about 0.25% or less. Two loci that are located on the same chromosome and are at a distance between them such that the frequency of recombination between the two loci is less than 10% (e.g., about 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or less) are also referred to as being "close" to each other. Because 1 cM is the distance between two markers that exhibit a recombination frequency of 1%, any marker is tightly linked (genetically and physically) to any other marker that is in close proximity (e.g., at a distance of 10 cM or less). Two tightly linked markers on the same chromosome can be located at a distance of 9 cM, 8 cM, 7 cM, 6 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.75 cM, 0.5 cM, or 0.25 cM or less from each other.

[0057] The term "linkage disequilibrium" refers to the non-random separation of genetic locus or proterties (or both). In either case, linkage disequilibrium implies that the relevant locus is physically close enough along a certain length of chromosome so that they separate together with a frequency greater than random (i.e., non-random). The marker showing linkage disequilibrium is considered to be linked. The linked locus separates altogether at more than 50% of the time, for example, about 51% to about 100% of the time. In other words, two markers that separate altogether have a recombination frequency less than 50% (and, by definition, separate less than 50cM on the same linkage group). As used herein, linkage can be between two markers or alternatively between a marker and the locus affecting the phenotype. The marker locus can be " associated " (linkage) with proterties. The linkage degree of the marker locus and the locus affecting the phenotypic traits, for example, the statistical probability (such as F statistics or LOD scoring) of molecular marker and phenotype separation.

[0058] Linkage disequilibrium is most commonly measured using the metric r 2 Evaluation, the measure r 2 By Hill, W.G. and Robertson, A. Theor. Appl. Genet. 38: 226-231 (1968). 2 = 1, there is complete LD between the two marker loci, meaning that the markers have not yet separated by recombination and have the same allele frequency. 2 The value will depend on the population used. 2 Values ​​greater than 1 / 3 indicate that LD is strong enough to be used for mapping (Ardlie et al. Nature Reviews Genetics 3: 299-309 (2002). Therefore, when the r between paired marker loci 2 When the value is greater than or equal to 0.33, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0, the alleles are in linkage disequilibrium.

[0059] As used herein, "linkage equilibrium" describes a situation in which two markers segregate independently, i.e., randomly sort in progeny. Markers that exhibit linkage equilibrium are considered unlinked (regardless of whether they are located on the same chromosome).

[0060] A "locus" is a position, eg, on a chromosome, where a nucleotide, gene, sequence, or marker is located.

[0061] " Logarithm of odds (LOD) value " or " LOD score " (Risch, Science 255:803-804 (1992)) is used for genetic interval mapping to describe the degree of linkage between two marker loci. An LOD score of three between two markers indicates that the probability of linkage is 1000 times higher than the probability of non-linkage, while an LOD score of two indicates that the probability of linkage is 100 times higher than the probability of non-linkage. An LOD score greater than or equal to two can be used to detect linkage. The LOD score can also be used to show the strength of association between a marker locus and a quantitative trait in a "quantitative trait locus" mapping. In this case, the size of the LOD score depends on the proximity of the marker locus to the locus that affects the quantitative trait and the size of the quantitative trait effect.

[0062] The term "plant" includes whole plants, plant cells, plant protoplasts, plant cell or tissue cultures from which plants can be regenerated, plant callus, intact plant masses and plant cells in plants or plant parts, such as seeds, flowers, cotyledons, leaves, stems, shoots, roots, root tips, etc. As used herein, "modified plants" means any plant that has a genetic change due to human intervention. The modified plant may have a genetic change introduced by plant transformation, genome editing, or conventional plant breeding.

[0063] " mark " is the means of linkage between the position or mark and the trait locus (the locus that affects proteus) on discovery heredity or physical map.The position that mark detects can be known by detecting polymorphic allele and gene mapping thereof, or is known by hybridizing, sequence matching or amplification of the sequence that has carried out physical mapping.Mark can be the phenotype (such as " waxy " phenotype) of DNA marker (detecting DNA polymorphism), albumen (detecting the variation of coded polypeptide) or simple inheritance.Can be from genomic nucleotide sequence or the nucleotide sequence (such as, the RNA or cDNA of splicing) development DNA marker.According to DNA marking technology, described mark can be made up of the complementary primer of the described locus of side joint and / or the complementary probe of the polymorphic allele hybridization with this locus place.DNA marker, or genetic marker, can also be used to describe gene, dna sequence or Nucleotide (rather than the component for detecting gene or dna sequence) on chromosome itself, and often is used (such as, for example, mark for breast cancer) when this DNA marker is associated with the specific traits in human genetics. The term marker locus is the locus (gene, sequence or nucleotide) at which the marker is detected.

[0064] Mark can be limited by the type of polymorphism detected by it and the labeling technology for detecting the polymorphism.Marker types include, but are not limited to, for example, the detection of restriction fragment length polymorphism (RFLP), the detection of isozyme markers, randomly amplified polymorphic DNA (RAPD), amplified fragment length polymorphism detection (AFLP), the detection of simple sequence repeats (SSR), the detection of the amplification variable sequence of plant genome, the detection of self-maintaining sequence replication or the detection of single nucleotide polymorphism (SNP). SNP can for example be via following detection: DNA sequencing, sequence-specific amplification method based on PCR, the detection of polynucleotide polymorphism by allele-specific hybridization (ASH), dynamic allele-specific hybridization (DASH), molecular beacons, microarray hybridization, oligonucleotide ligase assay, Flap endonuclease, 5 ' endonuclease, primer extension, single-stranded conformational polymorphism (SSCP) or temperature gradient gel electrophoresis (TGGE). DNA sequencing (such as pyrophosphate sequencing technology) has the advantage of being able to detect a series of linked SNP alleles that form haplotype. Haplotypes tend to be more informative (detect higher levels of polymorphism) than SNPs.

[0065] A "marker allele" or "allele of a marker locus" can refer to one of a plurality of polymorphic nucleotide sequences found at a marker locus in a population.

[0066] "Marker Assisted Selection" (MAS) is a method for selecting individual plants based on their marker genotype.

[0067] "Marker-assisted counter-selection" is a method by which marker genotypes are used to identify plants that are not to be selected, allowing them to be removed from a breeding program or planting.

[0068] A "marker haplotype" refers to the combination of alleles at a marker locus.

[0069] "Marker locus" is a specific chromosomal site in the genome of a species where a specific marker can be found. The marker locus can be used to track the presence of a second linked locus, for example a linked locus that affects the expression of a phenotypic trait. For example, the marker locus can be used to monitor the separation of alleles at genetically or physically linked loci.

[0070] The term "molecular marker" can be used to refer to a genetic marker (as defined above) used as a reference point when identifying a linked locus, or its encoded product (e.g., an albumen). The mark can be derived from a genomic nucleotide sequence or from an expressed nucleotide sequence (e.g., RNA, cDNA, etc. derived from splicing), or from a coded polypeptide. The term also refers to a nucleic acid sequence complementary to a marker sequence or a flanking marker sequence, such as a nucleic acid pair as a probe or a primer capable of amplifying a marker sequence. A "molecular marker probe" is a nucleic acid sequence or molecule that can be used to identify the presence of a marker locus, such as a nucleic acid probe complementary to a marker locus sequence. Alternatively, in some aspects, a molecular probe refers to any type of probe that can distinguish (i.e., genotyping) a specific allele present in a marker locus. When nucleic acid specifically hybridizes in a solution, nucleic acid is "complementary." When positioned at an insertion / deletion region, such as a non-collinear region as described herein, some marks as described herein are also referred to as hybridization markers. This is because, by definition, the insertion region is a polymorphism relative to a plant without insertion. Therefore, the mark only needs to indicate whether the insertion / deletion region exists. Any suitable marker detection technology can be used to identify such hybridization markers, for example, SNP technology is used in the examples provided herein.

[0071] An allele is "negatively" associated with a trait when the allele is linked to the trait and when the presence of the allele is an indicator that the desired trait or form of the trait will not occur in a plant containing the allele.

[0072] The term "phenotype", "phenotypic traits" or "trait" can refer to the observable expression of a gene or a series of genes. The phenotype can be observed by the naked eye, or by any other evaluation method, such as weighing, counting, measuring (length, width, angle, etc.), microscopy, biochemical analysis or electromechanical determination. In some cases, phenotype is directly controlled by a single gene or genetic locus, i.e., a "single gene trait" or "simple genetic trait". In the absence of high-level environmental changes, a single gene trait can be separated in a colony to produce a "qualitative" or "discrete" distribution, i.e., the phenotype falls into discrete categories. In other cases, the phenotype is the result of several genes and can be regarded as a "polygenic trait" or "complex trait". Polygenic traits are separated in a colony to produce a "quantitative" or "continuous" distribution, i.e., the phenotype cannot be divided into discrete categories. Both single gene and polygenic traits can be affected by the expression environment at their location, but polygenic traits tend to have larger environmental components.

[0073] A "physical map" of a genome is a map showing the linear order of identifiable landmarks (including genes, markers, etc.) on chromosomal DNA. However, in contrast to a genetic map, the distances between landmarks are absolute (e.g., measured in base pairs, either separated or overlapping contiguous gene segments) and are not based on genetic recombination (which can vary in different populations).

[0074] A "polymorphism" is a change in the DNA between two or more individuals within a population. A polymorphism preferably has a frequency of at least 1% in a population. Useful polymorphisms can include single nucleotide polymorphisms (SNPs), simple sequence repeats (SSRs), or insertion / deletion polymorphisms, also referred to herein as "indels."

[0075] A "production marker" or "production SNP marker" is a marker that has been developed for high throughput purposes. Production SNP markers are developed to detect specific polymorphisms and are designed for use with a variety of chemistries and platforms.

[0076] The term "quantitative trait locus" or "QTL" refers to a region of DNA that is associated with the differential expression of a quantitative phenotypic trait in at least one genetic background (e.g., in at least one breeding population). The region of a QTL encompasses or is closely linked to one or more genes that affect the trait in question.

[0077] A "reference sequence" or "consensus sequence" is a defined sequence used as the basis for sequence comparison. A labeled reference sequence is obtained by sequencing multiple lines at the locus, aligning the nucleotide sequences in a sequence alignment program (e.g., Sequencher), and then obtaining the most common nucleotide sequence of the alignment. Polymorphisms found between the individual sequences are annotated within the consensus sequence. A reference sequence is typically not an exact copy of any single DNA sequence, but rather represents a mixture of available sequences and can be used to design primers and probes for polymorphisms within the sequence.

[0078] An "unfavorable allele" of a marker is a marker allele that segregates with an unfavorable plant phenotype, thus providing the benefit of identifying plants that can be removed from a breeding program or planting.

[0079] The term "yield" refers to the productivity per unit area of ​​a specific plant product with commercial value. Yield is affected by both heredity and environmental factors." agronomy," "agronomic traits" and "agronomic performance" refer to the proterties (and the genetic elements of basis) of a given plant variety, which facilitate the output of the process through the growing season. Single agronomic traits include vigor of emergence, nutrient vitality, stress tolerance, disease resistance or tolerance, herbicide resistance, branching, flowering, fruiting rate, seed size, seed density, lodging resistance, threshing rate, etc. Therefore, yield is the final result of all agronomic traits.

[0080] Provided herein are marker loci that demonstrate statistically significant cosegregation with disease resistance traits that confer broad resistance to a specific disease or diseases. Detection of these loci or additional linked loci and resistance genes can be used as part of a breeding program in marker-assisted selection to produce plants that confer resistance to one or more diseases.

[0081] Gene mapping

[0082] For a long time, it has been recognized that specific loci associated with specific phenotypes (e.g., disease resistance) can be mapped in the genome of an organism. Advantageously, plant breeders can use molecular markers to identify desired individuals by detecting marker alleles that show a statistically significant probability of co-segregating with the desired phenotype, manifesting as linkage disequilibrium. By identifying molecular markers or clusters of molecular markers that co-segregate with the desired trait, breeders can rapidly select for the desired phenotype by selecting the appropriate molecular marker alleles (a process known as marker-assisted selection or MAS).

[0083] Various methods can be used for detecting molecular markers or molecular marker clusters that are co-segregated with the purpose proterties (for example disease resistance proterties).The basic idea of ​​these methods is to detect that its alternative genotype (or allele) has the mark of significantly different average phenotype.Therefore, the amplitude of the difference between the alternative genotype (or allele) between the comparative marker locus or the significance level of this difference.Infer that the proterties gene is positioned at the genotype difference that is closest to following mark, and described mark has maximum associated.Two kinds of such methods that are used to detect the purpose proterties locus are: 1) association analysis (i.e. association mapping) based on colony and 2) traditional linkage analysis.

[0084] Association mapping

[0085] Understanding the degree and pattern of linkage disequilibrium (LD) in the genome is a prerequisite for developing effective association methods to identify and map quantitative trait loci (QTL). Linkage disequilibrium (LD) refers to the non-random association of alleles in an individual collection. When LD is observed in the alleles at the linked locus, it is measured as the LD decay in a specific region across the chromosome. The degree of LD reflects the recombination history of the region. The average rate of LD decay in the genome can help predict the quantity and density of the markers required for whole genome association studies and provide an estimate of the resolution that can be expected.

[0086] Association or LD mapping aims to identify important genotype-phenotype associations. It has been used for the identification of significant genotype-phenotype associations in outcrossing species, such as humans (Corder et al. (1994) "Protective effect of apolipoprotein-Etype-2 allele for late-onset Alzheimer-disease," Nat Genet 7:180-184; Hastbacka et al. (1992) “Linkage disequilibrium mapping in isolated founder populations: diastrophic dysplasia in Finland,” Nat Genet 2:204-211; Kerem et al. (1989) “Identification of the cystic fibrosis gene: genetic analysis,” Science 245:1073-1080) and maize (Remington et al., (2001) “Structure of linkage disequilibrium and phenotype associations in the maize genome,” Proc Natl Acad Sci USA 98:11479-11484; Thornsberry et al. (2001) Dwarf8 polymorphismsassociate with variation in flowering time,” Nat Genet 28:286-289; by Flint-Garcia et al. (2003) “Structure of linkage disequilibrium in plants,” Annu Rev Plant Biol. 54:357-374), where recombination between heterozygotes is frequent and leads to a rapid decay of LD. In inbred species where recombination between homozygous genotypes cannot be detected genetically, the extent of LD is greater (i.e., larger blocks of linked markers are inherited together), and this dramatically enhances the detection power of association mapping (Wall and Pritchard (2003) "Haplotype blocks and linkage disequilibrium in the human genome," Nat Rev Genet 4:587-597).

[0087] The recombination and mutation history of a population depends on mating habits as well as the effective size and age of the population. Large population sizes provide an enhanced probability of detecting recombination, while older populations are generally associated with higher levels of polymorphism, both of which contribute to the observed accelerated LD decay rates. On the other hand, smaller effective population sizes, such as populations that have experienced a recent genetic bottleneck, tend to show slower LD decay rates, leading to more extensive haplotype conservation (Flint-Garcia et al. (2003) "Structure of linkage disequilibrium in plants," Annu Rev Plant Biol. 54:357-374).

[0088] Elite breeding lines provide a valuable starting point for association analyses. Association analyses use quantitative phenotypic scores (e.g., a disease tolerance rating of 1 to 9 for each line) in the analysis (as opposed to looking only at tolerance and resistance allele frequency distributions in interpopulation allele distribution type analyses). The availability of detailed phenotypic performance data for a large number of elite lines collected over many years and environments through breeding programs provides a valuable dataset for genetic marker association mapping analyses. This paves the way for seamless integration between research and application, leveraging historically accumulated datasets. However, understanding the relationship between polymorphism and recombination is useful for developing appropriate strategies to efficiently extract maximum information from these resources.

[0089] This type of association analysis neither generates nor requires any map data, but is independent of map position. The analysis compares the phenotypic score of the plant with the genotype at each locus. Subsequently, any suitable map (e.g., a composite map) can optionally be used to help visualize the distribution of identified QTL markers and / or QTL marker clusters using the map positions of the previously determined markers.

[0090] Traditional linkage analysis

[0091] The basic principles of traditional linkage analysis are the same; however, LD is generated by creating a population from a small number of founders. The founders are selected to maximize the level of polymorphism within the constructed population, and the level of co-segregation of the polymorphic sites with a given phenotype is evaluated. Many statistical methods have been used to identify significant marker-trait associations. One such method is the interval mapping method (Lander and Botstein, Genetics121: 185-199 (1989), wherein for each of a number of sites along a genetic map (e.g., intervals of 1 cM), the probability that the gene controlling the trait of interest is located at that site is tested. Genotype / phenotype data are used to calculate the LOD score (logarithm of the probability ratio) for each test site. When the LOD score is greater than a threshold, there is significant evidence that the gene controlling the trait of interest is located at that site on the genetic map (it will be located between two specific marker loci).

[0092] Provided herein are marker loci that demonstrate statistically significant cosegregation with disease resistance traits, as determined by traditional linkage analysis and by genome-wide association analysis. Detection of these loci or additional linked loci can be used in marker-assisted breeding programs to produce plants with disease resistance.

[0093] Activities in a marker-assisted breeding program may include, but are not limited to: selecting among new breeding populations based on historical genotype and agronomic trait associations to identify which population has the highest frequency of a favorable nucleic acid sequence, selecting among progeny within a breeding population for favorable nucleic acid sequences, selecting among parental lines based on predictions of progeny performance, and advancing lines in germplasm improvement activities based on the presence of a favorable nucleic acid sequence.

[0094] Chromosome interval

[0095] Chromosome intervals associated with disease resistance traits are provided. Various methods can be used to identify chromosomal intervals. The boundaries of such chromosomal intervals are stretched to encompass markers that will be linked to genes controlling the desired trait. In other words, stretching the chromosomal interval enables any marker within the interval (including terminal markers that define the interval boundaries) to be used as a marker for the disease resistance trait.

[0096] Conversely, if, for example, two markers in close proximity show co-segregation with a desired phenotypic trait, it is sometimes unclear whether each of those markers identifies the same gene or two different genes or genes. In any case, knowledge of how many genes are in a particular physical / genomic interval is not necessary to perform or practice what is presented in this disclosure.

[0097] The chromosome 4 interval may encompass any marker identified herein as being associated with the SCR resistance trait, including "T" at PM01-000058W (position 99 of reference sequence SEQ ID NO: 11), "G" at SOURST-83_1284720 (position 51 of reference sequence SEQ ID NO: 21), "C" at SOURST-83_1314662 (position 24 of reference sequence SEQ ID NO: 19), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at PZE-104005694 (position 25 of reference sequence SEQ ID NO: 23), "C" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), "G" at SOURST-83_1936804 (position 33 of reference sequence SEQ ID NO: 25), "C" at SOURST-83_1936805 (position 47 of reference sequence SEQ ID NO: 26), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_1 NO: 32, position 30), "C" at PZE-104001404 (position 26 of reference sequence SEQ ID NO: 25), "C" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "G" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), "A" at SOURST-83_2036602 (position 24 of reference sequence SEQ ID NO: 20), "T" at SOURST-83_2035716 (position 35 of reference sequence SEQ ID NO: 28), "T" at PZE-104001592 (position 46 of reference sequence SEQ ID NO: 29), "C" at SOURST-83_1926276 (position 26 of reference sequence SEQ ID NO: 26), "G" at SOURST-83_1652968 (position 26 of reference sequence SEQ ID NO: 27), NO:30, position 51) and "G" at SOURST-83_2679982 (reference sequence SEQ ID NO:31, position 26). Any marker located within these intervals can be used as a marker for SCR resistance and can be used in the context of the methods presented herein to identify and / or select plants with resistance to SCR, whether newly conferred or enhanced compared to control plants. In certain embodiments, markers located upstream and downstream of the SCR gene position are very closely linked genetically and physically and can therefore be used to select SCR genes for trait introgression and product development.

[0098] Chromosomal intervals can also be defined by markers that are linked (showing linkage disequilibrium) to disease resistance genes, and r 2 is a common measure of linkage disequilibrium (LD) in the context of association studies. If the r of LD between a chromosome 7 marker locus in the interval of interest and another adjacent chromosome 7 marker locus is2 Values ​​greater than 1 / 3 (Ardlie et al., Nature Reviews Genetics 3:299-309 (2002)), the loci are in linkage disequilibrium with each other.

[0099] Tags and chain relationships

[0100] The common measure of linkage is the frequency of proterties cosegregation.This can be expressed as cosegregation percentage (recombination frequency), or centimorgan (cM).cM is the unit of measurement of gene recombination frequency.1cM equals in a single generation, and the probability that the proterties on a locus that causes due to exchange will be separated with the proterties on another locus is 1% (meaning that proterties cosegregate under 99% of the time).Because the frequency of exchange events between chromosome distance and proterties is approximately proportional, there is the approximate physical distance associated with recombination frequency.

[0101] The marker loci themselves are traits and can be assessed according to standard linkage analysis by tracking the marker loci during segregation. Thus, 1 cM is equal to a 1% probability that a marker locus will segregate with another locus due to crossover in a single generation.

[0102] In some embodiments, the marker is from a gene that is controlled by a target gene, and the marker is more effective and advantageous as the index of the desired proterties. Tightly linked loci show about 10% or lower, preferably about 9% or lower, still more preferably about 8% or lower, again more preferably about 7% or lower, still more preferably about 6% or lower, again more preferably about 5% or lower, still more preferably about 4% or lower, again more preferably about 3% or lower, and still more preferably about 2% or lower loci between exchange frequency. In highly preferred embodiments, related loci (such as marker loci and target loci) show about 1% or lower recombination frequency, for example, about 0.75% or lower, more preferably about 0.5% or lower, or again more preferably about 0.25% or lower recombination frequency. Therefore, the loci are separated by about 10cM, 9cM, 8cM, 7cM, 6cM, 5cM, 4cM, 3cM, 2cM, 1cM, 0.75cM, 0.5cM or 0.25cM or lower. In other words, two loci are said to be "proximal" to each other if they are located on the same chromosome and are located at a distance between them such that the frequency of recombination between the two loci is less than 10% (e.g., about 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or less).

[0103] Although specific marker alleles can be co-segregated with disease resistance traits, it is important to note that marker loci are not necessarily responsible for expressing disease resistance phenotypes. For example, it is not required that the marker polynucleotide sequence is part of the gene responsible for the disease resistance phenotype (e.g., part of the gene open reading frame). The association between specific marker alleles and disease resistance traits is due to the initial "coupling" linkage phase between the marker allele and the allele in the ancestral strain from which the allele originates. Finally, through repeated recombination, exchange events between markers and loci can change this orientation. For this reason, favorable marker alleles can change according to the linkage phase present in parents with disease resistance, and the parents with disease resistance are used to produce segregating populations. This does not change the fact that markers can be used to monitor the separation of phenotypes. It only changes which marker allele is considered to be favorable in a given segregating population.

[0104] The method presented herein comprises the existence of one or more marker allelotrope that is associated with disease resistance in the detection plant, and identifies and / or is selected at those marker loci and has the plant of favorable allele then.Mark has been accredited as and is associated with disease resistance trait in this article, and therefore can be used for predicting the disease resistance in plant.Any mark (according to the genetic map based on single meiosis) in 50 cM, 40 cM, 30 cM, 20 cM, 15 cM, 10 cM, 9 cM, 8 cM, 7 cM, 6 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.75 cM, 0.5 cM or 0.25 cM also can be used for predicting the disease resistance in plant.

[0105] Marker-assisted selection

[0106] Molecular markers can be used in a variety of plant breeding applications (see, for example, Staub et al. (1996) Hortscience 31:729-741; Tanksley (1983) Plant Molecular Biology Reporter.1:3-8). One of the main areas of interest is to increase the efficiency of backcrossing and introgressing genes using marker-assisted selection (MAS). Molecular markers that display linkage to loci that affect desired phenotypic traits provide a useful tool for selecting traits in plant populations. This is especially true when the phenotype is difficult to determine. Because DNA marker assays are more labor-saving and take up less physical space than field phenotyping, much larger populations can be assayed, increasing the probability of finding recombinants with the target segment that moved from the donor line to the recipient line. The tighter the linkage, the more useful the marker, because recombination is less likely to occur between the marker and the gene causing the trait, which can lead to false positives. Since double recombination events are required, flanking markers reduce the probability that false positive selection will occur. Ideally, the gene itself has the marker so that recombination between the marker and the gene cannot occur. In some embodiments, the methods disclosed herein generate markers in disease resistance genes, where the gene is identified by inferring the genomic location from clustering or cluster analysis of conserved domains.

[0107] When a gene is introgressed by MAS, not only the gene but also the flanking regions are introduced (Gepts. (2002) Crop Sci ; 42: 1780-1790). This is called "linkage drag". In cases where the donor plant is highly unrelated to the recipient plant, these flanking regions carry additional genes that can encode agronomically undesirable traits. This "linkage drag" can also lead to reduced yield or other negative agronomic characteristics, even after multiple rounds of backcrossing with elite lines. This is sometimes also called "yield drag". The size of the flanking regions can be reduced by additional backcrossing, although this is not always successful because breeders cannot control the size of the region or the recombination breakpoints (Young et al. (1998) Genetics 120:579-585). In classical breeding, recombinations that contribute to a reduction in the size of the donor segment are often selected by chance alone (Tanksley et al. (1989) Biotechnology7: 257-264). Even after 20 backcrosses in this type of backcross, it is expected to find a fairly large fragment of the donor chromosome that is still linked to the selected gene. However, using markers makes it possible to select those rare individuals that undergo recombination near the target gene. In 150 backcross plants, there is a 95% probability that at least one plant will undergo an exchange within 1 cM (based on the map distance of a single meiosis) of the gene. Markers will allow those individuals to be clearly identified. For an additional backcross of 300 plants, there is a 95% probability of exchange within a single meiotic map distance of 1 cM on the other side of the gene, resulting in a fragment of less than 2 cM (based on the map distance of a single meiosis) near the target gene. This can be achieved in two generations using markers, while without markers it would take an average of 100 generations (see Tanksley et al., supra). When the exact location of a gene is known, flanking markers around the gene can be used to select for recombination in different population sizes. For example, in smaller population sizes, recombination can be expected to occur further away from the gene, thus requiring more distal flanking markers to detect recombination.

[0108] The key components of implementing MAS are: (i) defining the population in which marker-trait associations will be determined, which can be a segregating population, or a random or structured population; (ii) monitoring the segregation or association of polymorphic markers with respect to traits and using statistical methods to determine linkage or association; (iii) defining a set of desired markers based on the results of the statistical analysis, and (iv) using and / or extrapolating this information to the current set of breeding germplasm to enable marker-based selection decisions. The markers described in this disclosure, as well as other marker types such as SSRs and FLPs, can be used in marker-assisted selection schemes.

[0109] SSRs can be defined as relatively small numbers of tandemly repeated DNA sequences of 6 bp or less in length (Tautz (1989) Nucleic Acid Research 17: 6463-6471; Wang et al. (1994) Theoretical and Applied Genetics , 88:1-6). Polymorphism arises from variations in the number of repeat units, possibly due to slippage during DNA replication (Levinson and Gutman (1987) Mol Biol Evol 4: 203-221). Repeat length variations can be detected by designing PCR primers to conserved non-repeat flanking regions (Weber and May (1989) Am J Hum Genet.44:388-396). SSRs are well suited for mapping and MAS because they are multiallelic, codominant, reproducible, and amenable to high-throughput automation (Rafalski et al. (1996) Generating and using DNA markers inplants. Non-mammalian genomic analysis: a practical guide. Academicpress. pp 75-135).

[0110] Various types of SSR markers can be generated, and SSR profiles can be obtained by gel electrophoresis of amplified products. The scoring of marker genotypes is based on the size of the amplified fragments.

[0111] Various types of FLP markers can also be generated. Most commonly, amplification primers are used to generate fragment length polymorphisms (FLPs). These FLP markers are similar in many respects to SSR markers, except that the region amplified by the primers is typically not a highly repetitive region. The amplified region, or amplicon, also exhibits sufficient variability between germplasm, often due to insertions or deletions, to allow the fragments generated by the amplification primers to be distinguished between polymorphic individuals. Such insertions and deletions are known to occur frequently in maize (Bhattramakki et al. (2002)). Plant Mol Biol 48, 539–547; Rafalski (2002b), ibid.).

[0112] SNP markers detect single base pair nucleotide substitutions. Of all molecular marker types, SNPs are the most abundant and therefore have the potential to provide the highest genetic map resolution (Bhattramakki et al. 2002). Plant Molecular Biology 48:539-547). SNPs can be detected at an even higher throughput level than SSRs, i.e., in a so-called "ultra-high throughput" format, because SNPs do not require large amounts of DNA and can be directly automated. SNPs also have the potential to be a relatively low-cost system. These three factors together make SNPs highly attractive for use in MAS. Several methods can be used for SNP genotyping, including but not limited to hybridization, primer extension, oligonucleotide ligation, nuclease cleavage, microsequencing, and coding spheres. Such methods have been reviewed in: Gut (2001) Hum Mutat 17 pp. 475-492; Shi (2001) Clin Chem 47, pp. 164-172; Kwok (2000) Pharmacogenomics1, pp. 95-100; and Bhattramakki and Rafalski (2001) Discovery and application of single nucleotide polymorphism markers in plants. In: RJ Henry, Ed, Plant Genotyping: The DNA Fingerprinting of Plants , CABI Publishing, Wallingford. A variety of commercially available technologies utilize these and other methods to detect SNPs, including Masscode.™. (Qiagen), INVADER®. (Third Wave Technologies) and Invader PLUS®, SNAPSHOT®. (Applied Biosystems), TAQMAN®. (Applied Biosystems), and BEADARRAYS®. (Illumina).

[0113] Many SNPs that are together within a sequence, or across linked sequences, can be used to describe the haplotype of any particular genotype (Ching et al. (2002), BMC Genet. 3:19 pp Gupta et al. 2001, Rafalski (2002b), Plant Science 162:329-333). Haplotypes can provide more information than a single SNP and are more descriptive of any particular genotype. For example, a single SNP may be the allele "T" for a particular strain or variety that is disease-resistant, but the allele "T" may also occur in a breeding population used as a recurrent parent. In this case, a haplotype, such as the allele combination at a linked SNP marker, can provide more information. Once a unique haplotype has been assigned to a donor chromosome region, the haplotype can be used in the population or any subpopulation thereof to determine whether an individual has a specific gene. See, for example, WO2003054229. The use of an automated high-throughput marker detection platform makes this method efficient and effective.

[0114] Many marks presented herein can be easily used as single nucleotide polymorphism (SNP) markers, to select SCR resistance genes. Using PCR, primers are used to amplify DNA segments from individuals (preferably inbred lines) representing the diversity in the target population. PCR products are directly sequenced in one or two directions. The resulting sequences are compared and polymorphisms are identified. Polymorphisms are not limited to single nucleotide polymorphisms (SNPs), but also include insertions and deletions, CAPS, SSRs, and VNTRs (variable number of tandem repeats). Specifically, about fine map information as described herein, it is possible to easily use the information provided herein to obtain additional polymorphic SNPs (and other markers) in the region amplified by primers disclosed herein. The markers in the described map region can be hybridized with BAC or other genomic libraries, or electronically compared with genomic sequences, to find new sequences in the approximate position identical to the described markers.

[0115] In addition to the SSRs, FLPs, and SNPs described above, other types of molecular markers are also widely used, including but not limited to expressed sequence tags (ESTs), SSR markers derived from EST sequences, randomly amplified polymorphic DNA (RAPD), and other nucleic acid-based markers.

[0116] In some cases, isozyme profiles and linked morphological features can also be used indirectly as markers. Although they do not directly detect DNA differences, they are often also affected by specific genetic differences. However, markers that detect DNA variation are far more numerous than isozyme or morphological markers and are more polymorphic (Tanksley (1983) Plant Molecular Biology Reporter 1:3-8).

[0117] Sequence alignment or overlapping group also can be used for finding the sequence of the upstream or downstream of the specific mark that this paper lists.These new sequences near mark as herein described are then used to find and develop functionally equivalent mark.For example different physics and / or gene maps are compared with location equivalent mark, and described mark is not described in the present disclosure, but is positioned at similarity zone.These collections of illustrative plates can be in species, or even cross over and have carried out heredity or other species that physically compare.

[0118] Typically, MAS uses polymorphic markers that have been identified as having a significant probability of co-segregation with a trait such as an SCR disease resistance trait. It is speculated that such markers are mapped close to one or more genes that confer the plant its disease resistance phenotype, and it is believed that such markers are indicators, or markers, of the desired trait. The test plant has the desired allele in the marker, and it is expected that plants containing the desired genotype at one or more loci will transfer the desired genotype to their progeny along with the desired phenotype. Therefore, plants with SCR disease resistance can be selected by detecting one or more marker alleles, and in addition, progeny plants derived from those plants can also be selected. Thus, a plant containing the desired genotype (i.e., a genotype associated with disease resistance) in a given chromosome region is obtained and then hybridized with another plant. The progeny of this hybrid will then be evaluated genotypically using one or more markers, and progeny plants with the same genotype in a given chromosome region will then be selected as having disease resistance.

[0119] SNPs can be used individually or in combination (ie, SNP haplotypes) to select for favorable resistance gene alleles that are associated with SCR disease resistance. For example, the following SNP haplotypes: "T" at PM01-000058W (position 99 of reference sequence SEQ ID NO: 11), "G" at SOURST-83_1284720 (position 51 of reference sequence SEQ ID NO: 21), "C" at SOURST-83_1314662 (position 24 of reference sequence SEQ ID NO: 19), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), "C" at SOURST-83_1936804 (position 30 of reference sequence SEQ ID NO: 32), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), "A" at SOURST-83_III (position 32 of reference sequence SEQ ID NO: 24), "C" at SOURST-83_1936804 (position 30 of reference sequence SEQ ID NO: 32), "T" at SOURST-83_1542053 (position 51 of reference sequence SEQ ID NO: 22), NO: 25), “G” at position 26 of reference sequence SEQ ID NO: 25), “C” at position 26 of reference sequence SEQ ID NO: 26), “G” at position 26 of reference sequence SEQ ID NO: 27), “A” at position 24 of reference sequence SEQ ID NO: 20), “T” at position 35 of reference sequence SEQ ID NO: 28, “T” at position 46 of reference sequence SEQ ID NO: 29, “G” at position 51 of reference sequence SEQ ID NO: 30, and “T” at position 26 of reference sequence SEQ ID NO: 31, or a combination thereof.

[0120] The skilled person will anticipate that there may be additional polymorphic sites in and around the marker loci of the chromosome markers identified by the methods described herein, wherein the alleles of one or more polymorphic sites and the one or more polymorphic sites in the haplotype are in linkage disequilibrium (LD), and therefore can be used in marker-assisted selection programs to infiltrate target gene alleles or genomic fragments. If there is an allele at one of the sites that tends to predict the presence of an allele at another site on the same chromosome, then two specific alleles at different polymorphic sites are said to be in LD, (Stevens, et al. Mol. Diag.4:309-17 (1999). The marker locus can be located within 5 cM, 2 cM, or 1 cM (on a genetic map based on a single meiotic division) of the QTL for the disease resistance trait.

[0121] The skilled artisan will appreciate that allele frequencies (and therefore haplotype frequencies) may vary from one germplasm pool to another. Germplasm pools differ due to differences in maturity, heterosis grouping, geographical distribution, etc. Therefore, in some germplasm pools, SNPs and other polymorphisms may not be informative.

[0122] Plant composition

[0123] Plants identified, modified and / or selected by any of the methods described above are also of interest.

[0124] Proteins and their variants and fragments

[0125] The present disclosure encompasses SCR-resistant polypeptides. As used herein, "SCR-resistant polypeptide" and "SCR-resistant protein" refer interchangeably to polypeptides having SCR-resistant activity and having sufficient identity to the SCR-resistant polypeptide of any one of SEQ ID NOs: 5-8. A variety of SCR-resistant polypeptides are contemplated.

[0126] As used herein, "sufficient identity" refers to amino acid sequences having at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity. In some embodiments, the sequence identity is for the full length of the polypeptide. When used in conjunction with percent sequence identity in this article, the term "about" means + / - 1.0%.

[0127] As used herein, "recombinant protein" refers to a protein that is no longer in its native environment, for example, in vitro or in a recombinant bacterial or plant host cell; a protein expressed from a polynucleotide that has been edited from its native version; or a protein expressed from a polynucleotide in a genomic location that is different from the native sequence.

[0128] As used herein, "substantially free of cellular material" refers to preparations of polypeptides, including proteins, having less than about 30%, 20%, 10%, or 5% (by dry weight) of non-target proteins (also referred to herein as "contaminating proteins").

[0129] "Fragments" or "biologically active portions" include polypeptide or polynucleotide fragments that comprise a sequence having sufficient identity to an SCR resistance polypeptide or polynucleotide, respectively, and that exhibit disease resistance when expressed in a plant.

[0130] As used herein, a "variant" refers to a protein or polypeptide having an amino acid sequence that is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the parent amino acid sequence.

[0131] In some embodiments, the SCR resistance polypeptide comprises an amino acid sequence that is at least about 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the full length or a fragment of the amino acid sequence of any one of SEQ ID NOs: 5-8, wherein the SCR resistance polypeptide confers SCR resistance when expressed in a plant.

[0132] Methods for such manipulations are generally known in the art. For example, amino acid sequence variants of SCR-resistant polypeptides can be prepared by mutations in DNA. This can also be accomplished by one of several forms of mutagenesis (e.g., site-specific double-strand break technology) and / or in directed evolution. In some aspects, the changes encoded in the amino acid sequence will not substantially affect the function of the protein. Such variants will have the desired activity. However, it should be understood that the ability of SCR-resistant polypeptides to confer disease resistance can be improved by using such techniques on the compositions of the present disclosure.

[0133] Nucleic acid molecules and variants and fragments thereof

[0134] Provided are isolated or recombinant nucleic acid molecules comprising a nucleic acid sequence encoding an SCR resistance polypeptide or a biologically active portion thereof, as well as nucleic acid molecules sufficient for use as hybridization probes to identify nucleic acid molecules encoding proteins having regions of sequence homology. As used herein, the term "nucleic acid molecule" refers to DNA molecules (e.g., recombinant DNA, cDNA, genomic DNA, plasmid DNA, mitochondrial DNA) and RNA molecules (e.g., mRNA) and analogs of DNA or RNA produced using nucleotide analogs. The nucleic acid molecule can be single-stranded or double-stranded, but is preferably double-stranded DNA.

[0135] As used herein, "isolated" nucleic acid molecules (or DNA) refer to nucleic acid sequences (or DNA) that are no longer in their natural environment, e.g., in vitro. "Recombinant" nucleic acid molecules (or DNA) are used herein to refer to nucleic acid sequences (or DNA) that are in recombinant bacteria or plant host cells; that have been edited from their native sequence; or that are located in a position different from the native sequence. In some embodiments, an "isolated" or "recombinant" nucleic acid is one that does not contain sequences that naturally flank the nucleic acid (i.e., sequences located at the 5' and 3' ends of the nucleic acid) in the genomic DNA of the organism from which the nucleic acid was derived (preferably sequences that encode proteins). For the purposes of this disclosure, "isolated" or "recombinant" excludes isolated chromosomes when used to refer to nucleic acid molecules. For example, in various embodiments, a recombinant nucleic acid molecule encoding an SCR-resistant polypeptide may comprise less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb of nucleic acid sequence that naturally flanks the nucleic acid molecule in the genomic DNA of the cell from which the nucleic acid was derived.

[0136] In some embodiments, the isolated nucleic acid molecule encoding the SCR resistance polypeptide has one or more changes in the nucleic acid sequence compared to the native or genomic nucleic acid sequence. In some embodiments, changes in the native or genomic nucleic acid sequence include, but are not limited to: changes in the nucleic acid sequence due to the degeneracy of the genetic code; changes in the nucleic acid sequence due to amino acid substitutions, insertions, deletions and / or additions compared to the native or genomic sequence; removal of one or more introns; deletion of one or more upstream or downstream regulatory regions; and deletion of the 5' and / or 3' untranslated regions relative to the genomic nucleic acid sequence. In some embodiments, the nucleic acid molecule encoding the SCR resistance polypeptide is a non-genomic sequence.

[0137] A variety of polynucleotides encoding SCR-resistant polypeptides or related proteins are contemplated. When operably linked to appropriate promoter, transcription termination, and / or polyadenylation sequences, such polynucleotides can be used to produce SCR-resistant polypeptides in host cells. Such polynucleotides can also be used as probes for isolating homologous or substantially homologous polynucleotides encoding SCR-resistant polypeptides or related proteins.

[0138] In some embodiments, the nucleic acid molecule encoding the SCR resistance polypeptide is a polynucleotide having a sequence as shown in any one of SEQ ID NOs: 5-8, and variants, fragments, and complementary sequences thereof. As used herein, "complementary sequence" refers to a nucleic acid sequence that is sufficiently complementary to a given nucleic acid sequence such that it can hybridize with the given nucleic acid sequence to form a stable duplex. As used herein, "polynucleotide sequence variant" refers to a nucleic acid sequence that encodes the same polypeptide except for the degeneracy of the genetic code.

[0139] In some embodiments, the nucleic acid molecule encoding the SCR resistance polypeptide is a non-genomic nucleic acid sequence. As used herein, a "non-genomic nucleic acid sequence" or "non-genomic nucleic acid molecule" or "non-genomic polynucleotide" refers to a nucleic acid molecule having one or more changes in the nucleic acid sequence compared to a natural or genomic nucleic acid sequence. In some embodiments, changes in natural or genomic nucleic acid molecules include, but are not limited to: changes in the nucleic acid sequence due to the degeneracy of the genetic code; optimization of the nucleic acid sequence for expression in plants; changes in the nucleic acid sequence that introduce at least one amino acid substitution, insertion, deletion and / or addition compared to the natural or genomic sequence; removal of one or more introns associated with the genomic nucleic acid sequence; insertion of one or more heterologous introns; deletion of one or more upstream or downstream regulatory regions associated with the genomic nucleic acid sequence; insertion of one or more heterologous upstream or downstream regulatory regions; deletion of the 5' and / or 3' untranslated regions associated with the genomic nucleic acid sequence; insertion of heterologous 5' and / or 3' untranslated regions; and modification of polyadenylation sites. In some embodiments, the non-genomic nucleic acid molecule is a synthetic nucleic acid sequence.

[0140] In some embodiments, the nucleic acid molecule encoding an SCR resistance polypeptide disclosed herein is a non-genomic polynucleotide having a nucleotide sequence that is at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the nucleic acid sequence of any one of SEQ ID NOs: 5-8, wherein the SCR resistance polypeptide has SCR resistance activity when expressed in a plant.

[0141] In some embodiments, the nucleic acid molecule encodes an SCR-resistant polypeptide variant comprising one or more amino acid substitutions to the amino acid sequence of any one of SEQ ID NOs: 5-8.

[0142] The embodiments also encompass nucleic acid molecules that are fragments of these nucleic acid sequences encoding SCR resistance polypeptides. As used herein, "fragment" refers to a portion of a nucleic acid sequence encoding an SCR resistance polypeptide. A fragment of a nucleic acid sequence can encode a biologically active portion of an SCR resistance polypeptide, or it can be a fragment that can be used as a hybridization probe or PCR primer using the methods disclosed below. Nucleic acid molecules that are fragments of a nucleic acid sequence encoding an SCR resistance polypeptide comprise at least about 150, 180, 210, 240, 270, 300, 330, 360, 400, 450 or 500 consecutive nucleotides or up to the number of nucleotides present in the full-length nucleic acid sequence encoding an SCR resistance polypeptide identified by the methods disclosed herein, depending on the intended use. "Contiguous nucleotides" is used herein to refer to nucleotide residues that are immediately adjacent to each other. Fragments of the nucleic acid sequences of the embodiments will encode protein fragments that retain the biological activity of the SCR resistance polypeptide and therefore retain disease resistance. As used herein, "retains disease resistance" refers to a polypeptide that has at least about 10%, at least about 30%, at least about 50%, at least about 70%, 80%, 90%, 95% or more of the disease resistance of a full-length SCR resistance polypeptide as shown in any one of SEQ ID NOs: 5-8.

[0143] "Percent (%) sequence identity" relative to a reference sequence (subject) is determined as the percentage of amino acid residues or nucleotides in a candidate sequence (query) that are identical to each amino acid residue or nucleotide in the reference sequence, after aligning the sequences and introducing gaps (if necessary) to achieve maximum percent sequence identity and not considering any conservative amino acid substitutions as part of the sequence identity. Alignments for determining percent sequence identity can be achieved in a variety of ways, for example, using publicly available computer software such as BLAST, BLAST-2. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithm required to achieve maximum alignment over the full length of the compared sequences. The percent identity between two sequences is a function of the number of identical positions shared by the sequences (e.g., percent identity of query sequence = number of identical positions between query and subject sequences / total number of positions of query sequence × 100).

[0144] In some embodiments, the SCR-resistance polynucleotide encodes an SCR-resistance polypeptide comprising an amino acid sequence that is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more identical to the amino acid sequence of any one of SEQ ID NOs: 5-8.

[0145] The embodiments also encompass nucleic acid molecules encoding variants of SCR-resistant polypeptides. "Variants" of SCR-resistant polynucleotides encoding polypeptide sequences include those that encode SCR-resistant polypeptides identified by the methods disclosed herein, but that have conservative differences due to the degeneracy of the genetic code, as well as those sequences with sufficient identity as described above. Naturally occurring allelic variants can be identified using well-known molecular biology techniques, such as polymerase chain reaction (PCR) and hybridization techniques as outlined below. Variant nucleic acid sequences also include synthetically derived nucleic acid sequences that have been generated, for example, by using site-directed mutagenesis, but still encode SCR-resistant polypeptides disclosed herein.

[0146] The skilled artisan will further appreciate that changes can be introduced by mutation of the nucleic acid sequence, thereby resulting in changes in the amino acid sequence of the encoded SCR-resistant polypeptide without altering the biological activity of these proteins. Thus, variant nucleic acid molecules can be generated by introducing one or more nucleotide substitutions, additions, and / or deletions into the corresponding nucleotide sequences disclosed herein, such that one or more amino acid substitutions, additions, or deletions are introduced into the encoded protein. Mutations can be introduced by standard techniques, such as site-directed mutagenesis and PCR-mediated mutagenesis. Such variant nucleic acid sequences are also encompassed by the present disclosure.

[0147] Alternatively, variant nucleic acid sequences can be prepared by randomly introducing mutations along all or part of the coding sequence (e.g., by saturation mutagenesis), and the resulting mutants can be screened for their ability to confer activity to identify mutants that retain activity. Following mutagenesis, the encoded protein can be recombinantly expressed, and standard assay techniques can be used to determine the activity of the protein.

[0148] In addition to the standard cloning methods described in, for example, Ausubel, Berger, and Sambrook, the polynucleotides of the present disclosure and fragments thereof are optionally used as substrates for various recombination and recursive recombination reactions, i.e., to produce other polypeptide homologs and fragments thereof with desired properties. Various such reactions are known. Methods for producing variants of any nucleic acid listed herein (which include recursively recombining such a polynucleotide with a second (or more) polynucleotide to form a variant polynucleotide library) are also embodiments of the present disclosure, as are the libraries produced, the cells comprising the libraries, and any recombinant polynucleotides produced by such methods. Additionally, such methods optionally include selecting variant polynucleotides from these libraries based on activity, as where such recursive recombination is performed in vitro or in vivo.

[0149] Various diversity generation schemes, including nucleic acid recursive recombination schemes, are available. The programs can be used alone and / or in combination to produce one or more variants of a nucleic acid or nucleic acid set, and variants of an encoded protein. Individually or in their entirety, these programs provide robust and widely applicable methods for producing diverse nucleic acids and nucleic acid sets (including, for example, nucleic acid libraries) that can be used, for example, to engineer or rapidly evolve nucleic acids, proteins, pathways, cells, and / or organisms with new and / or improved characteristics.

[0150] Although for the sake of clarity, distinctions and classifications are made in the subsequent discussion, it should be understood that these techniques are generally not mutually exclusive. In fact, various methods can be used alone or in combination, in parallel or in series, to obtain different sequence variants.

[0151] The result of any diversity generation program described herein can be the generation of one or more nucleic acids, which can select or screen for nucleic acids having or conferring desired characteristics or encoding nucleic acids having or conferring proteins of desired characteristics. After being diversified by one or more methods that can be obtained herein or otherwise by those skilled in the art, any nucleic acid produced can be selected for desired activity or characteristic, for example, such activity under desired pH, etc. This can include identifying any activity that can be detected, for example, in an automated or automated form, by any assay in this area. Various related (or even unrelated) characteristics can be evaluated in series or in parallel as appropriate by the practitioner.

[0152] The nucleotide sequence of embodiment can also be used for separating corresponding sequence from different sources.In this way, such sequence (based on the sequence homology of the sequence identified by the method disclosed herein) can be identified using methods such as PCR, hybridization. The embodiment encompasses the sequence selected based on the sequence identity of the full sequence illustrated with it or its fragment herein. Such sequence includes the sequence as the ortholog of this sequence.Term " ortholog " refers to the gene derived from the common ancestor gene and found in different species due to species formation.When its nucleotide sequence and / or its encoded protein sequence have the substantial identity as defined elsewhere herein, the gene found in different species is considered to be ortholog.

[0153] In the PCR method, oligonucleotide primers can be designed for PCR reactions to amplify corresponding DNA sequences from cDNA or genomic DNA extracted from any organism of interest. Methods for designing PCR primers and PCR cloning are disclosed in Sambrook et al., (1989) Molecular Cloning: A Laboratory Manual (2nd Edition, Cold Spring Harbor Laboratory Press), Plainview, New York), hereinafter referred to as "Sambrook". See also, Innis et al., ed., (1990) PCR Protocols: A Guide to Methods and Applications (Academic Press, New York); Innis and Gelfand, ed., (1995) PCR Strategies (Academic Press, New York); and Innis and Gelfand, ed., (1999) PCR Methods Manual (Academic Press, New York). Known PCR methods include, but are not limited to, methods using paired primers, nested primers, single specific primers, degenerate primers, gene-specific primers, vector-specific primers, partially mismatched primers, etc.

[0154] In hybridization methods, all or part of the nucleic acid sequence can be used to screen cDNA or genomic libraries. Methods for constructing such cDNA and genomic libraries are disclosed in Sambrook and Russell, (2001), supra. So-called hybridization probes can be genomic DNA fragments, cDNA fragments, RNA fragments or other oligonucleotides, and can be labeled with a detectable group (such as 32P or any other detectable label, such as other radioisotopes, fluorescent compounds, enzymes or enzyme cofactors). Probes for hybridization can be prepared by labeling synthetic oligonucleotides, which are based on the nucleic acid sequences encoding known polypeptides disclosed herein. Degenerate primers can be used in addition, which are designed based on conserved nucleotides or amino acid residues in the nucleic acid sequence or the encoded amino acid sequence. The probe typically comprises a region of the following nucleic acid sequence that hybridizes under stringent conditions to at least about 12, at least about 25, at least about 50, 75, 100, 125, 150, 175 or 200 consecutive nucleotides of the nucleic acid sequence encoding the polypeptide or its fragment or variant. Methods and stringency conditions for preparing probes for hybridization are disclosed in Sambrook and Russell, (2001), supra.

[0155] Nucleotide constructs, expression cassettes and vectors

[0156] The term "nucleotide construct" is used herein and is not intended to limit the embodiment to the nucleotide construct comprising DNA. It will be appreciated by those of ordinary skill in the art that nucleotide constructs, particularly polynucleotides and oligonucleotides consisting of ribonucleotides and combinations of ribonucleotides and deoxyribonucleotides, can also be used in the method disclosed herein. The nucleotide constructs, nucleic acids and nucleotide sequences of the embodiment additionally encompass all complementary forms of such constructs, molecules and sequences. In addition, the nucleotide constructs, nucleotide molecules and nucleotide sequences of the embodiment encompass all nucleotide constructs, molecules and sequences in the plant transformation method that can be used for the embodiment, including but not limited to those consisting of deoxyribonucleotides, ribonucleotides and combinations thereof. Such deoxyribonucleotides and ribonucleotides comprise both naturally occurring molecules and synthetic analogs. The nucleotide constructs, nucleic acids and nucleotide sequences of the embodiment also encompass all forms of nucleotide constructs, including but not limited to single-stranded forms, double-stranded forms, hairpins, stem-loop structures etc.

[0157] Additional embodiments relate to transformed organisms, such as organisms selected from the group consisting of plant cells, bacteria, yeast, baculovirus, protozoa, nematodes, and algae. The transformed organism comprises a DNA molecule of the embodiments, an expression cassette comprising the DNA molecule, or a vector comprising the expression cassette, which can be stably incorporated into the genome of the transformed organism.

[0158] The sequence of embodiment is provided in DNA construct, for expression in target organism.This construct will comprise the 5' and 3' regulatory sequence of the sequence that is operably connected to embodiment.Term " operably connected " refers to the functional connection between promoter and / or regulatory sequence and the second sequence as used herein, wherein promoter and / or regulatory sequence starts, mediates and / or affects the transcription of the DNA sequence dna corresponding to the second sequence.Usually, being operably connected means that the connected nucleotide sequence is continuous, and connects two protein coding regions in identical reading frame when necessary.This construct can contain at least one other gene to be co-transformed into organism in addition.Or, one or more other genes can be provided on multiple DNA constructs.

[0159] Such DNA constructs are provided with a plurality of restriction sites for inserting the polypeptide gene sequence of the present disclosure, which will be placed under the transcriptional regulation of the regulatory regions.The DNA constructs may additionally comprise a selectable marker gene.

[0160] In the transcription direction from 5' to 3', the DNA construct will generally include: a transcription and translation initiation region (i.e., a promoter), a DNA sequence of the embodiment, and a transcription and translation termination region (i.e., a termination region) that functions in the organism used as the host. For the host organism and / or for the sequence of the embodiment, the transcription initiation region (i.e., a promoter) can be natural, similar, exogenous, or heterologous. In addition, the promoter or regulatory sequence can be a natural sequence or, alternatively, a synthetic sequence. As used herein, the term "exogenous" means that the promoter is not found in the natural organism into which the promoter is introduced. As used herein, the term "heterologous" with respect to a sequence means a sequence that is derived from an alien species, or if derived from the same species, is substantially modified by intentional human intervention in the natural form of its composition and / or genomic locus. As used herein, a chimeric gene comprises a coding sequence that is operably connected to a transcription initiation region that is heterologous to the coding sequence. When the promoter is a native or natural sequence, expression of the operably linked sequence is altered from wild-type expression, which results in a change in phenotype.

[0161] In some embodiments, the DNA construct comprises a polynucleotide encoding an SCR-resistance polypeptide of the embodiments. In some embodiments, the DNA construct comprises a polynucleotide encoding a fusion protein comprising an SCR-resistance polypeptide of the embodiments.

[0162] In some embodiments, the DNA construct may further include a transcriptional enhancer sequence. As used herein, the term "enhancer" refers to a DNA sequence that can stimulate promoter activity and may be an innate element or heterologous element inserted to enhance the level or tissue specificity of a promoter. Various enhancers include, for example, introns with gene expression enhancing properties in plants (U.S. Patent Application Publication No. 2009 / 0144863), ubiquitin introns (i.e., maize ubiquitin intron 1 (see, e.g., NCBI sequence S94464)), ω enhancers or ω major enhancers (Gallie et al., (1989)). Molecular Biology of RNA , eds.: Cech (Liss, New York) 237-256 and Gallie et al., (1987) Gene 60:217-25), the CaMV 35S enhancer (see, e.g., Benfey et al., (1990) EMBO J . 9:1685-96) and the enhancers of U.S. Patent No. 7,803,992 can also be used. The above list of transcription enhancers is not intended to be limiting. Any suitable transcription enhancer can be used in the embodiments.

[0163] The termination region may be native with the transcriptional initiation region, may be native with the operably linked DNA sequence of interest, may be native with the plant host, or may be derived from another source (i.e., foreign or heterologous to the promoter, sequence of interest, plant host, or any combination thereof).

[0164] Convenient termination regions are available from Agrobacterium tumefaciens (A. tumefaciens ) of the Ti plasmid, such as the octopine synthase and nopaline synthase terminator regions. See also, Guerineau, et al., (1991) Mol. Gen. Genet. 262:141-144;Proudfoot, (1991) Cell 64:671-674; Sanfacon, et al. (1991) Genes Dev. 5:141-149; Mogen, et al. (1990) Plant Cell 2:1261-1272; Munroe, et al., (1990) Gene 91:151-158; Ballas, et al. (1989) Nucleic Acids Res. 17:7891-7903 and Joshi, et al., (1987) Nucleic Acid Res. 15:9627-9639.

[0165] Where appropriate, the nucleic acid can be optimized to increase expression in the host organism. Thus, where the host organism is a plant, the synthetic nucleic acid can be synthesized using plant-preferred codons to improve expression. For a discussion of host-preferred codon usage, see, for example, Campbell and Gowri, (1990). Plant Physiol 92:1-11. For example, although the nucleic acid sequences of the embodiments can be expressed in both monocot and dicot species, the sequences can be modified to account for monocot or dicot specific preferences and GC content preferences, as these preferences have been shown to differ (Murray et al. (1989) Nucleic Acids Res . 17:477-498). Thus, plant preferences for specific amino acids can be derived from known gene sequences from plants.

[0166] Other sequence modifications are known to enhance gene expression in cellular hosts. These include eliminating sequences encoding false polyadenylation signals, exon-intron splice site signals, transposon-like repeats, and other sequences that are well characterized and may be detrimental to gene expression. The GC content of the sequence can be adjusted to the average level of a given cellular host, as calculated with reference to known genes expressed in the host cell. As used herein, the term "host cell" refers to a cell that contains a vector and supports replication and / or expression of a desired expression vector. The host cell can be a prokaryotic cell such as Escherichia coli, or a eukaryotic cell such as a yeast, insect, amphibian, or mammalian cell, or a monocotyledonous or dicotyledonous plant cell. An example of a monocotyledonous host cell is a corn host cell. When possible, the sequence is modified to avoid predicted hairpin secondary mRNA structures.

[0167] In preparing the expression cassette, the various DNA fragments may be manipulated to provide a DNA sequence in the proper orientation and, where appropriate, in the proper reading frame. To this end, adapters or linkers may be employed to connect the DNA fragments, or other manipulations may be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. For this purpose, in vitro mutagenesis, primer repair, restriction enzyme digestion, annealing, resubstitution (e.g., conversion and transversion) may be involved.

[0168] Many promoters can be used to implement the embodiments. A promoter can be selected based on the desired results. The nucleic acid can be combined with a constitutive, tissue-preferred, inducible, or other promoter for expression in a host organism.

[0169] Plant transformation

[0170] The methods of the embodiments involve introducing a polypeptide or polynucleotide into a plant. As used herein, "introducing" means presenting a polynucleotide or polypeptide to a plant in a manner such that the sequence is able to enter the interior of a plant cell. The methods of the embodiments do not depend on the specific method used to introduce the polynucleotide or polypeptide into a plant, as long as the polynucleotide or polypeptide is able to enter the interior of at least one cell of the plant. Methods for introducing a polynucleotide or polypeptide into a plant include, but are not limited to, stable transformation methods, transient transformation methods, and viral-mediated methods.

[0171] As used herein, "stable transformation" means that the nucleotide construct introduced into a plant is integrated into the genome of the plant and can be inherited by its progeny. As used herein, "transient transformation" means that a polynucleotide is introduced into the plant and is not integrated into the genome of the plant, or that a polypeptide is introduced into the plant. As used herein, "plant" refers to a complete plant, plant organs (e.g., leaves, stems, roots, etc.), seeds, plant cells, propagules, and embryos and progeny thereof. Plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, and pollen).

[0172] Transformation protocols and protocols for introducing nucleotide sequences into plants may vary depending on the type of plant or plant cell targeted for transformation (i.e., monocot or dicot). Suitable methods for introducing nucleotide sequences into plant cells and subsequent insertion into the plant genome include microinjection (Crossway et al., (1986) Biotechniques 4:320-334), electroporation (Riggs et al., (1986) Proc. Natl. Acad. Sci. USA 83:5602-5606), Agrobacterium-mediated transformation (U.S. Pat. Nos. 5,563,055 and 5,981,840), direct gene transfer (Paszkowski et al., (1984) EMBO J 3:2717-2722) and ballistic particle acceleration (see, e.g., U.S. Patent Nos. 4,945,050; 5,879,918; 5,886,244 and 5,932,782; Tomes et al., (1995) Plant Cell, Tissue, and Organ Culture: Fundamental Methods , Gamborg and Phillips, eds. (Springer-Verlag, Berlin); and McCabe et al., (1988) Biotechnology 6:923-926); and Lec1 transformation method (WO 00 / 28058). For potato transformation method, see Tu et al., (1998) Plant Molecular Biology 37:829-838 and Chong et al., (2000) Transgenic Research9:71-78. Additional transformation procedures can be found in: Weissinger et al., (1988) Ann. Rev. Genet. 22:421-477; Sanford et al., (1987) Particulate Science and Technology 5:27-37 (onion); Christou et al., (1988) Plant Physiol . 87:671-674 (soybean); McCabe et al., (1988) Bio / Technology 6:923-926 (soybean); Finer and McMullen, (1991) In Vitro Cell Dev. Biol . 27P:175-182 (soybean); Singh et al., (1998) Theor. Appl. Genet . 96:319-324 (soybean); Datta et al., (1990) Biotechnology 8:736-740 (Rice); Klein et al., (1988) Proc. Natl. Acad. Sci. USA 85:4305-4309 (maize); Klein et al., (1988) Biotechnology 6:559-563 (maize); U.S. Patent Nos. 5,240,855; 5,322,783 and 5,324,646; Klein et al. (1988) Plant Physiol. 91:440-444 (maize); Fromm et al. (1990) Biotechnology 8:833-839 (maize); Hooykaas-Van Slogteren et al., (1984) Nature (London) 311:763-764; U.S. Patent No. 5,736,369 (cereals); Bytebier et al., (1987) Proc. Natl. Acad. Sci. USA 84:5345-5349 (Liliaceae); De Wet et al., (1985) The Experimental Manipulation of Ovule Tissues , Chapman et al., eds. (Longman, New York), pp. 197–209 (pollen); Kaeppler et al., (1990) Plant Cell Reports 9:415-418 and Kaeppler et al., (1992) Theor. Appl. Genet 84:560-566 (whisker-mediated transformation); D'Halluin et al., (1992) Plant Cell 4:1495-1505 (electroporation); Li et al., (1993) Plant Cell Reports, 12:250-255 and Christou and Ford, (1995) Annals of Botany 75:407-413 (Rice); Osjoda et al., (1996) Nature Biotechnology 14:745-750 (via Agrobacterium tumefaciens ( Agrobacterium tumefaciens ) of corn).

[0173] Methods for introducing genome editing technology into plants

[0174] In some embodiments, genome editing technology can be used to introduce the polynucleotide composition into the genome of the plant, or the polynucleotides previously introduced in the genome of the plant can be edited using genome editing technology. For example, the identified polynucleotides can be introduced into the desired position in the plant genome by using double-strand break technology (such as TALEN, meganucleases, zinc finger nucleases, CRISPR-Cas, etc.). For example, for the purpose of site-specific insertion, the identified polynucleotides can be introduced into the desired position in the genome using the CRISPR-Cas system. The desired position in the plant genome can be any target site required for insertion, such as a genomic region suitable for breeding, or can be a target site located in a genomic window with an existing purpose trait. The existing purpose trait can be an endogenous trait or a previously introduced trait.

[0175] In some embodiments, in the case of having identified SCR resistance gene alleles in the genome, genome editing technology can be used to change or modify the polynucleotide sequence. Site-specific modifications can be introduced into the site-specific modifications of the desired SCR resistance gene allele polynucleotides, including those produced using any method for introducing site-specific modifications, including but not limited to using gene repair oligonucleotides (e.g., U.S. Publication 2013 / 0019349), or by using double-strand break technology, such as TALEN, large-range nucleases, zinc finger nucleases, CRISPR-Cas, etc. Such technology can be used to modify the polynucleotides previously introduced by insertion, deletion or substitution of nucleotides in the introduced polynucleotides. Alternatively, double-strand break technology can be used to add additional nucleotide sequences to the introduced polynucleotides. Additional sequences that can be added include additional expression elements (e.g., enhancer and promoter sequences). In another embodiment, genome editing technology can be used to locate additional disease resistance proteins adjacent to the SCR resistance polynucleotide composition in the plant genome to produce molecular stacks of disease resistance proteins.

[0176] "Altered target site," "altered target sequence," "modified target site," and "modified target sequence" are used interchangeably herein and refer to a target sequence as disclosed herein that includes at least one alteration when compared to a non-altered target sequence. Such "alterations" include, for example: (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, or (iv) any combination of (i)-(iii). Example

[0177] The following examples are provided to illustrate but not limit the claimed subject matter. It should be understood that the examples and embodiments described herein are for illustrative purposes only, and those skilled in the art will recognize that various reagents or parameters can be changed without departing from the spirit of the present disclosure or the scope of the appended claims.

[0178] Example 1. QTL Mapping

[0179] Southern corn rust (SCR) is a maize disease caused by the fungal pathogen Puccinia polysora. CIMBL83, an inbred line resistant to SCR, was crossed with a susceptible line. The F1 progeny from this cross were then backcrossed with the susceptible line to create a BC1F1 population, with the susceptible line as the recurrent parent. A mapping population of 117 recombinant inbred lines (RILs) was developed from individual seeds from the BC1F1 population.

[0180] The RILs were planted at Changge Station in Henan Province, China in 2016 and 2017 and at Ledong Station in Hainan Province, China in 2017. Natural infection with the SCR pathogen was used to assess the resistance level of plants in the field. Disease severity was evaluated 30-35 days after pollination using a scale from 1 (most resistant phenotype) to 9 (most susceptible phenotype). The resistant parent, CIMBL83, was completely resistant (scored 1), while the susceptible parent showed an average score of 7.4. For the RILs, the average score of the graph was used for subsequent analysis. Plants were genotyped with 9,433 SNPs using the MaizeSNP9.4k BeadChip, and composite interval mapping was performed using QTL Cartographer. A QTL was identified in bin 4.01 (qSCR4.01) on chromosome 4. This QTL was reproducible across years and multiple locations and explained 48-65% of the phenotypic variation.

[0181] Example 2. QTL fine mapping and candidate gene identification

[0182] To develop populations for fine mapping, qSCR4.01Heterozygous RIL lines were self-pollinated to develop near-isogenic lines (NIL) F2 populations. A total of 1289 NIL F2 individuals were used for fine mapping. These plants were genotyped using two SNP markers PM01-000058W (SEQ ID NO: 11) and PM01-00002MH (SEQ ID NO: 12) flanking the QTL interval. A total of 449 recombinants were identified. These recombinants were grown, genotyped with newly developed markers, and self-crossed. 70 offspring of these F3 families were grown in Ledong, Hainan and evaluated for SCR resistance. Based on the analysis of the phenotype, combined with the genotype of the F2 NIL, the SCR resistance of the recombinants was evaluated. qSCR4.01 Delimited as an interval flanked by markers SOURST-83_1314662 (SEQ ID NO: 19) and SOURST-83_2036602 (SEQ ID NO: 20).

[0183] Table 2. SNP markers used for QTL mapping

[0184] Tag Name Reference sequence type resistance Susceptible B73v4 SNP location PM01-000058W SEQ ID NO: 11 SNP markers T A 1279994 PM01-00002MH SEQ ID NO: 12 SNP markers T C 3248271

[0185] Table 3. Primers and probes for markers used for QTL mapping

[0186] Tag Name Forward primer (5'-3') Reverse primer (5'-3') probe type resistance Susceptible B73v4 SNP location SOURST-83_1314662 SEQ ID NO: 13 SEQ ID NO: 14 SEQ ID NO: 15 KASP mark G A 1546133 SOURST-83_2036602 SEQ ID NO: 16 SEQ ID NO: 17 SEQ ID NO: 18 KASP mark A T 2537720

[0187] For identification qSCR4.01 The genes in the QTL interval were sequenced in the CIMBL83 genome. RNA-seq data from infected and uninfected leaf tissues were generated to facilitate gene annotation and expression analysis. qSCR4.01 The interval is approximately 668kb. Two R genes have been annotated in this interval. The main candidate gene NLR01 appears to be a fusion of three NLRs, with multiple transcription start sites, which result in a single NLR (NLR01-1, SEQ ID NO:7), a 2-NLR fusion (NLR01-2, SEQ ID NO:6) and a 3-NLR fusion (NLR01-3, SEQ ID NO:5). A second gene, NLR02 (SEQ ID NO:8), has the NB-ARC domain, but lacks a leucine-rich repeat domain. By Iso-Seq (full-length isoform sequencing) and extra RNA-seq, more NLR01 transcript isoforms (SEQ ID NO 33-58) have been identified.

[0188] Table 4. Candidate R gene sequences

[0189] SEQ ID NO: Description of SEQ ID NO: 1 NLR01-3 CDS 2 NLR01-2 CDS 3 NLR01-1 CDS 4 NLR02 CDS 5 NLR01-3 amino acids 6 NLR01-2 amino acids 7 NLR01-1 amino acids 8 NLR02 amino acids 33 NLR01-4_CDS 34 NLR01-5_CDS 35 NLR01-6_CDS 36 NLR01-7_CDS 37 NLR01-8_CDS 38 NLR01-9_CDS 39 NLR01-10_CDS 40 NLR01-11_CDS 41 NLR01-12_CDS 42 NLR01-13_CDS 43 NLR01-14_CDS 44 NLR01-15_CDS 45 NLR01-16_CDS 46 NLR01-4_Protein 47 NLR01-5_Protein 48 NLR01-6_Protein 49 NLR01-7_Protein 50 NLR01-8_Protein 51 NLR01-9_Protein 52 NLR01-10_Protein 53 NLR01-11_Protein 54 NLR01-12_Protein 55 NLR01-13_Protein 56 NLR01-14_Protein 57 NLR01-15_Protein 58 NLR01-16_Protein

[0190] Based on SNPs from various maize lines, primers were designed within our mapping interval. Primer Picker was used to design primers for KASP markers. Several SNPs were able to distinguish between resistant and susceptible genotypes in our mapping population. These markers were used to narrow the chromosomal interval. To identify CIMBL83 haplotypes, after obtaining the CIMBL83 genomic sequence, KASP-marked primer sequences were used to map to the CIMBL83 genomic sequence. SNPs present in CIMBL83 were recorded as resistance alleles. Other SNPs were identified as susceptible alleles using KASP primers. Table 5 shows the KASP markers for CIMBL83 resistance haplotypes.

[0191] Table 5. CIMBL83 QTL and fine mapping markers

[0192] mark chromosome B73 v4 position CIMBL83 Susceptible Reference sequence SEQ ID NO: PM01-000058W Chr4 1279994 T A 11 SOURST-83_1284720 Chr4 1518539 G A 21 SOURST-83_1314662 Chr4 1546133 C T 19 SOURST-83_1542053 Chr4 1751040 T G 22 PZE-104005694 Chr4 1768240 A G 23 SOURST-83_1II Chr4 1785157 C T 24 SOURST-83_1936804 Chr4 2462979 C G 32 PZE-104001404 Chr4 2432052 G A 25 SOURST-83_1926276 Chr4 2430969 C T 26 SOURST-83_1652968 Chr4 1864853 G C 27 SOURST-83_2036602 Chr4 2537720 A T 20 SOURST-83_2035716 Chr4 2538606 T C 28 PZE-104001592 Chr4 2541019 T C 29 SOURST-83_2465654 Chr4 3000813 G A 30 SOURST-83_2679982 Chr4 3227087 T C 31

[0193] Example 5. Transgenic verification of SCR resistance causal gene

[0194] The entire NLR01 fragment (SEQ ID NO 9) and the genomic sequences of the NLR01-1, NLR01-2 and NLR01-3 genes (including the promoter region) were synthesized and cloned into a binary vector for transformation in HC69. Single copy mass events were hybridized with PHR03 to generate segregating T1 seeds. T1 plants were inoculated with multiple piles of Puccinia uredia spores and maintained in the greenhouse for 10 days. Greenhouse disease was visually scored as susceptible (S) or resistant (R) based on the presence or absence of burst uredia (a sporulating postule). Approximately 30 T1 plants with a single copy transgene and 30 T1 plants without the transgene, all from the same 6 events, were tested for SCR resistance in the greenhouse. All 30 transgene-positive plants were expected to produce a resistant phenotype, while all transgene-negative plants were expected to be scored as susceptible.

Claims

1. A method for identifying a corn plant having increased resistance to southern corn rust, the method comprising: a. detecting a haplotype associated with increased resistance to southern corn rust in the corn plant, wherein the haplotype comprises: a "T" at position 99 of SEQ ID NO: 11, a "G" at position 51 of SEQ ID NO: 21, a "C" at position 24 of SEQ ID NO: 19, a "T" at position 51 of SEQ ID NO: 22, an "A" at position 25 of SEQ ID NO: 23, a "C" at position 32 of SEQ ID NO: 24, a "C" at position 30 of SEQ ID NO: 32, a "G" at position 26 of SEQ ID NO: 25, a "C" at position 26 of SEQ ID NO: 26, a "G" at position 26 of SEQ ID NO: 27, an "A" at position 24 of SEQ ID NO: 20, a "T" at position 35 of SEQ ID NO: 28, a "T" at position 46 of SEQ ID NO: 29, a "G" at position 51 of SEQ ID NO: 30, and a "T" at position 26 of SEQ ID NO: 31; and b. Identifying the corn plant as having the resistance allele, wherein the plant has increased resistance to southern corn rust.

2. A method of selecting corn plants having increased resistance to southern corn rust, the method comprising: a. screening a population to determine whether one or more plants from the population contain a haplotype comprising a "T" at position 99 of SEQ ID NO: 11, a "G" at position 51 of SEQ ID NO: 21, a "C" at position 24 of SEQ ID NO: 19, a "T" at position 51 of SEQ ID NO: 22, an "A" at position 25 of SEQ ID NO: 23, a "C" at position 32 of SEQ ID NO: 24, a "C" at position 30 of a reference sequence of SEQ ID NO: 32, a "G" at position 26 of SEQ ID NO: 25, a "C" at position 26 of SEQ ID NO: 26, a "G" at position 26 of SEQ ID NO: 27, an "A" at position 24 of SEQ ID NO: 20, a "T" at position 35 of SEQ ID NO: 28, a "T" at position 46 of SEQ ID NO: 29, a "G" at position 51 of SEQ ID NO: 30, and a "T" at position 26 of SEQ ID NO: 31; and b. Selecting at least one plant from the population that comprises the gene allele.

3. The method of claim 2, further comprising: c. Crossing the corn plant of (b) with a second corn plant; and d. Obtaining progeny corn plants having the gene allele.

Citation Information

Patent Citations

  • Expression Enhancing Intron Sequences

    US20090144863A1

  • Epsps mutants

    US20130019349A1

  • Method for transporting substances into living cells and tissues and apparatus therefor

    US4945050A

  • Particle gun

    US5240855A

  • Soybean transformation by microparticle bombardment

    US5322783A