Compact PRR03 compositions and methods for gray leaf spot resistance
By introducing the compact PRR03 protein and its coding sequence into maize, the problem of resistance to gray leaf spot disease was solved, achieving the effects of enhancing resistance and reducing environmental side effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies make it difficult to quickly and effectively introduce gray leaf spot resistance genes into superior maize varieties, leading to maize yield losses, and fungicides have environmental side effects.
A compact PRR03 protein (PRR03c) and its coding sequence are provided. This protein is introduced into the maize genome through transgenic modification or genome editing to enhance resistance to gray leaf spot disease. The enhanced resistance is ensured by utilizing amino acid and nucleic acid sequences with high identity to SEQ ID NO:3 and SEQ ID NO:1,2.
This method has enabled enhanced resistance to gray leaf spot disease in maize during rapid breeding processes, reducing negative environmental impacts and improving maize yield and quality.
Smart Images

Figure SMS_1 
Figure SMS_3 
Figure SMS_4
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 578,471, filed August 24, 2023. TECHNICAL FIELD
[0002] The present disclosure relates to plants, plant breeding, and methods of identifying, selecting, and / or producing plants having compact PRR03 associated with gray leaf spot resistance. Provided are polynucleotides and constructs encoding proteins capable of providing gray leaf spot resistance and uses thereof. The polynucleotides, constructs, and marker alleles of the present disclosure are useful in the production of disease resistant plants through breeding, transgenic modification, and / or genome editing. Reference to a Sequence Listing Submitted in Text File
[0003] The official copy of the sequence listing is submitted in XML format as an electronic file named 9683-US-PSP_ST26, created on August 21, 2023, and having a size of 69,223 bytes. The sequence listing contained in this XML formatted document is part of the specification and is herein incorporated by reference in its entirety. BACKGROUND
[0004] Gray leaf spot is a major disease of maize (Zea mays). Gray leaf spot is a major concern due to significant reduction in yield, grain weight, and quality. Immature plant death (interrupted grain filling) and stalk breakage and lodging (resulting in loss of ears in the field) result in yield loss. Gray leaf spot occurs in all corn-growing regions and can result in 10-20% loss.
[0005] While farmers can combat fungal infection (e.g., gray leaf spot) by using fungicides, these fungicides have side effects on the environment and require field monitoring and diagnostic techniques to determine which fungus is causing the infection so that the correct fungicide can be used. It is more practical to use a corn line that carries a source of the gene for resistance or a transgenic source if the gene for resistance can be incorporated into elite, high yielding germplasm without reducing yield. Sources of the gene for resistance have been described (White et al. (1979) Annu. Corn Sorghum Res. Conf. Proc. 34:1-15; Carson. 1981. Sources of inheritance of resistance to gray leaf spot of corn. Ph.D. Thesis, University of Illinois, Urbana-Champaign; Badu-Apraku et al. (1987) Phytopathology 77:957-959; Toman et al. 1993. Phytopathology, 83:981-986; Cowen, N et al. (1991) Maize Genetics Conference Abstracts 33; Jung et al. (1994). Theoretical and Applied Genetics, 89:413-418). However, introgression of resistance can be very complex.
[0006] Selection by using a genetic sequence or molecular marker associated with the gray leaf spot resistance trait allows selection based only on the genetic makeup of the progeny. As a result, plant breeding can occur more rapidly, thereby producing commercially acceptable, higher levels of gray leaf spot resistant maize plants. There are multiple sequences and QTLs that control gray leaf spot resistance, each of which has a different effect on the trait. Thus, it is desirable to provide compositions and methods for identifying and selecting maize plants having a sequence that confers or enhances gray leaf spot resistance. There is a continuing need for disease resistant plants and methods for finding disease resistance genes. SUMMARY
[0007] Provided herein are compositions and methods useful in identifying and selecting plant markers or genetic sequences associated with increased Gray Leaf Spot ("GLS") disease resistance. Accordingly, the compositions and methods disclosed herein can be used to select GLS disease resistant plants, breed disease resistant plants, produce transgenic disease resistant plants, and / or produce genome edits of plants for GSL disease resistance. Also provided herein are plants and methods for making plants having a disclosed marker and / or gene associated with enhanced disease resistance as compared to a control plant. In some embodiments, the compositions and methods can be used to introgress disease resistance into a plant. In some embodiments, provided herein are disease resistance markers having increased disease resistance to GLS as compared to a control plant.
[0008] Disclosed herein is a compact PRR03 protein (PRR03c) having a length of less than 1,000 amino acids, less than 900 amino acids, less than 700 or less than 600 amino acids, or less than 500 amino acids, which provides increased or improved disease resistance to GLS. The PRR03c can comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3.Accordingly, PRR03c can (i) have a length of less than 1,000 amino acids and comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (ii) have a length of less than 900 amino acids and comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (iii) have a length of less than 800 amino acids and comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (iv) have a length of less than 700 amino acids and comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (iv) have a length of less than 600 amino acids and comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (v) have a length of less than 500 amino acids and comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3.
[0009] The foregoing PRR03c can be expressed from a genomic sequence ("PRR03c genomic sequence") having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity to SEQ ID NO: 1. As used herein, the term PRR03c genomic sequence refers to a genomic sequence comprising a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity to SEQ ID NO: 1.The PRR03c genomic sequence can encode a PRR03c that (i) has a length of less than 1,000 amino acids and comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (ii) has a length of less than 900 amino acids and comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (iii) has a length of less than 800 amino acids and comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (iv) has a length of less than 700 amino acids and comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (iv) has a length of less than 600 amino acids and comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3; (v) has a length of less than 500 amino acids and comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity to SEQ ID NO: 3.
[0010] In an embodiment, provided herein is a plant comprising a PRR03c genomic sequence disclosed herein at a non-native genomic locus. As used herein, a "non-native genomic locus" is a genomic locus other than the native locus between 90 and 115 cM on Zea mays chromosome 4. Thus, for example, a disclosed plant comprising a PRR03c genomic sequence at a non-native genomic locus can comprise a PRR03c genomic sequence disclosed herein having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity to SEQ ID NO: 1 and located on Zea mays chromosome 1, 2, 3, 5, 6, 7, 8, 9, or 10, or alternatively, at a location on Zea mays chromosome 4 other than its native locus between 90 and 115 cM on Zea mays chromosome 4.The disclosed plant may contain a PRR03c genome sequence encoding PRR03c, wherein PRR03c (i) has a length of less than 1,000 amino acids and contains an amino acid sequence that is identical to SEQ ID NO:3 in at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the amino acid sequence; (ii) has a length of less than 900 amino acids and contains an amino acid sequence that is identical to SEQ ID NO:3 in at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% of the amino acid sequence; or (iii) has a length of less than 8 ...0%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90%, at least 90 NO:3 has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3; (iv) has a length of less than 700 amino acids and contains an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3; (iv) has a length of less than 600 amino acids and contains an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3; (v) has a length of less than 500 amino acids and contains an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3. NO:3 has an amino acid sequence with at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity.
[0011] The aforementioned PRR03c can be expressed by a coding sequence (“PRR03c coding sequence”) having a length of less than 2,000 nucleotides, less than 1,900 nucleotides, less than 1,800 nucleotides, less than 1,700 nucleotides, less than 1,600 nucleotides, or less than 1,500 nucleotides, and having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2. In some instances, the aforementioned PRR03c coding sequence is located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively, on chromosome 4 of maize, at a location other than its natural locus between cM 90 and cM 115 of chromosome 4.As used herein, the term PRR03c coding sequence refers to a genome sequence that (A) contains a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2, and (B) encodes PRR03c, which (i) has a length of less than 1,000 amino acids and contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3; and (ii) has a length of less than 900 amino acids and contains an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3; NO:3 has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence of SEQ ID NO:3; (iii) has a length of less than 800 amino acids and contains an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence of SEQ ID NO:3; (iv) has a length of less than 700 amino acids and contains an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence of SEQ ID NO:3; (iv) has a length of less than 600 amino acids and contains an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the amino acid sequence of SEQ ID NO:3; IDNO:3 has an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3; (v) has a length of less than 500 amino acids and contains an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3.
[0012] This document also provides a plant comprising any PRR03c coding sequence disclosed herein, wherein such PRR03c coding sequence has a length of less than 2,000 nucleotides, less than 1,900 nucleotides, or less than 1,800 nucleotides, or less than 1,700 nucleotides, or less than 1,600 nucleotides, or less than 1,500 nucleotides, or less than 1,400 nucleotides, and has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2. The plant may be a transgenic plant PRR03c comprising the disclosed coding sequence operatively linked to a heterologous promoter.
[0013] In one aspect, methods are provided for identifying and / or selecting one or more plant materials having the PRR03c polypeptide, PRR03c genomic sequence, and / or PRR03c coding sequence disclosed herein, and associated with increased or improved resistance to gray leaf spot compared to control plants. As used herein, the term "plant material" refers to one or more plants, plant cells, plant tissues, seeds, or germplasm thereof. In some instances, methods for identification and / or selection include detecting or selecting one or more plant materials containing the PRR03c genome sequence disclosed herein, the PRR03c genome sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:1, and located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively located on chromosome 4 of maize at a location other than its natural locus between 90 and 115 cM on chromosome 4. In some instances, the method includes detecting or selecting plant material containing any of the disclosed PRR03c coding sequences that have at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2, and are located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively at a location on chromosome 4 of maize other than its natural locus between cM 90 and cM 115. The identified or selected plants may possess newly conferred or enhanced disease resistance (e.g., GLS resistance) relative to control plants that do not have genomic regions containing PRR03c genome sequences or PRR03c coding sequences.
[0014] In each method for identifying and / or selecting one or more plant materials having the PRR03c polypeptide, PRR03c genomic sequence, and / or PRR03c coding sequence disclosed herein, the method may further include confirming that the identified and / or selected plant contains the PRR03c polypeptide (or the genomic sequence and / or coding sequence encoding the PRR03c polypeptide) having a length of less than 1,000 amino acids, less than 900 amino acids, less than 800 amino acids, less than 700 amino acids, less than 600 amino acids, or less than 500 amino acids (e.g., about 489 amino acids).
[0015] In specific instances, the aforementioned method for identifying plant material having nucleic acids or polypeptides associated with increased or improved resistance to gray leaf spot may include obtaining nucleic acid samples from each of one or more plants, seeds, tissues, or germplasm in a population; screening each sample for the presence of: (i) any compact PRR03 protein (PRR03c) disclosed herein, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3; (ii) the PRR03c genomic sequence disclosed herein, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:1; or (iii) The disclosed PRR03c coding sequence has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2. In specific instances, the screening method further includes detecting the presence of the aforementioned PRR03c genomic sequence or PRR03c coding sequence located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively on chromosome 4 of maize at a location other than its natural locus between 90 and 115 cM on chromosome 4; and selecting plant material having the PRR03c genomic sequence or PRR03c coding sequence located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively on chromosome 4 of maize at a location other than its natural locus between 90 and 115 cM on chromosome 4.
[0016] In another example, the aforementioned method for identifying and / or selecting plant materials resistant to gray leaf spot may include obtaining nucleic acid samples from one or more plants, seeds, tissues, or germplasm, each sample representing multiple (e.g., a population) of plants, seeds, tissues, or germplasm; screening each sample for the presence of one or more of the aforementioned PRR03c polypeptide, PRR03c genomic sequence, or PRR03c coding sequence; and selecting multiple plants, seeds, tissues, or germplasm, wherein representative samples of the selected multiple samples have been screened for the PRR03c polypeptide, PRR03c genomic sequence, or PRR03c coding sequence.
[0017] The aforementioned methods for identifying and / or selecting plants may further include hybridizing at least one of the selected plants with a second plant that does not possess the PRR03c gene to produce offspring plants whose genomes comprise one or more of the following: (i) the PRR03c genome sequence disclosed herein, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:1; or (ii) the disclosed PRR03c coding sequence, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2. The PRR03c genome sequence and / or PRR03c coding sequence are located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively on chromosome 4 of maize at a location other than its natural locus between cM 90 and cM 115. In another example, the second plant is one of a plant line (“recurrent parent line”), and the method further includes crossing the progeny plant with another plant of the recurrent parent line to produce a second generation of progeny whose genome contains one or more of the aforementioned PRR03c genome sequence and / or PRR03c coding sequence. Optionally, the second generation of progeny may be crossed with the recurrent parent line to produce a third generation of progeny whose genome contains one or more of the aforementioned PRR03c genome sequence and / or PRR03c coding sequence. This method can be repeated three, four, five, six, seven, or more times, with each subsequent generation of progeny crossing with the recurrent parent line, thereby introducing the PRR03c genome sequence and / or the PRR03c coding sequence into the recurrent parent line. For example, these repeated backcrosses can produce plants with the PRR03c genome sequence and / or the PRR03c coding sequence, which exhibit improved GLS disease resistance compared to the original recurrent parent line (lacking either the PRR03c genome or the PRR03c coding sequence).
[0018] In an alternative approach, plants with increased resistance to gray leaf spot are crossed with a second plant to produce progeny plants. The progeny plants are then screened for QTLs or marker alleles associated with increased resistance to gray leaf spot according to the methods disclosed herein. Typically, such screening involves obtaining nucleic acid samples from each progeny plant and screening for the presence of nucleic acids in the samples, which comprise (i) the PRR03c genome sequence disclosed herein, having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:1, or (ii) the disclosed PRR03c coding sequence, having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2. The PRR03c genomic sequence and / or PRR03c coding sequence are located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively on chromosome 4 of maize at a location other than its natural locus between cM 90 and cM 115. The method may further include detecting and selecting one or more progeny plants whose nucleic acid samples contain the PRR03c genomic sequence and / or PRR03c coding sequence, thereby identifying novel progeny plants containing the PRR03c sequence associated with increased GLS resistance.
[0019] In each of the methods disclosed herein that include hybridizing one or more plants having a PRR03c polypeptide, a PRR03c genomic sequence, and / or a PRR03c coding sequence, the method may further include confirming that the progeny of such hybridization contains a PRR03c polypeptide (or a genomic sequence and / or coding sequence encoding a PRR03c polypeptide) having a length of less than 1,000 amino acids, less than 900 amino acids, less than 800 amino acids, less than 700 amino acids, less than 600 amino acids, or less than 500 amino acids (e.g., about 489 amino acids).
[0020] On the other hand, methods are provided that include expressing heterologous nucleic acids in plant material capable of increasing resistance to gray leaf spot. These methods may include introducing nucleic acid sequences or marker alleles associated with increased gray leaf spot resistance into the plant material, for example, through transgenic modification or genome editing methods. In some instances, the plant material was susceptible to gray leaf spot before the introduction of the heterologous nucleic acid. For example, the genome of a plant (e.g., a plant susceptible to gray leaf spot) may be altered by transgenic modification or genome editing to include one or more of the following: (i) the PRR03c genome sequence disclosed herein, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleotide sequence identity with SEQ ID NO:1; or (ii) the PRR03c coding sequence disclosed herein, which has a length of less than 2,000 nucleotides, less than 1,900 nucleotides, or less than 1,800 nucleotides, or less than 1,700 nucleotides, or less than 1,600 nucleotides, or less than 1,500 nucleotides, or less than 1,400 nucleotides, and is identical to SEQ ID NO:1. NO:2 has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleotide sequence identity. The PRR03c genome sequence and / or PRR03c coding sequence are located on maize chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10, or alternatively on maize chromosome 4 at a location other than its natural locus between cM 90 and cM 115 on chromosome 4.Plant material whose genome has been transgenic modified or genome-edited to include (i) the PRR03c genome sequence disclosed herein, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleotide sequence identity with SEQ ID NO:1, or (ii) the disclosed PRR03c coding sequence having a length of less than 2,000 nucleotides, less than 1,900 nucleotides, or less than 1,800 nucleotides, or less than 1,700 nucleotides, or less than 1,600 nucleotides, or less than 1,500 nucleotides, or less than 1,400 nucleotides, and is identical to SEQ ID NO:1. NO:2 has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleotide sequence identity. The PRR03c genome sequence and / or the PRR03c coding sequence are located on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively on chromosome 4 of maize, at a location other than its natural locus between cM 90 and cM 115. In some instances, transgenic or genome-edited plant material provides increased resistance to GLS diseases compared to syngeneic plants lacking genome editing.
[0021] In each of the methods disclosed herein for expressing heterologous nucleic acids encoding a PRR03c polypeptide (e.g., a PRR03c genomic sequence and / or a PRR03c coding sequence) in plant material, the method may further include confirming that the plant expresses a PRR03c polypeptide having a length of less than 1,000 amino acids, less than 900 amino acids, less than 800 amino acids, less than 700 amino acids, less than 600 amino acids, or less than 500 amino acids (e.g., about 489 amino acids).
[0022] This document provides a method for introducing a construct into plant material that does not contain a PRR03c sequence associated with increased GLS resistance. The introduced construct contains a nucleic acid heterologous to the plant material, and the heterologous nucleic acid contains (i) the PRR03c genomic sequence disclosed herein, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:1, or (ii) the disclosed PRR03c coding sequence, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:2. In some instances, the introduced constructs are integrated into different (non-natural) genomic loci, for example, at loci on maize chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10, or alternatively at locations on maize chromosome 4 other than its natural locus between 90 and 115 cM on maize chromosome 4.
[0023] Some of the aforementioned methods relate to methods for transforming host cells, which may be plant cells. The method includes transforming host or plant cells with isolated polynucleotide constructs disclosed herein. The method may further include generating a plant by transforming plant cells with constructs disclosed herein and regenerating the plant from the transformed plant cells, thereby producing a plant having the PRR03c genome sequence or PRR03c coding sequence disclosed herein. In some instances, such as compared to syngeneic plants lacking the PRR03c genome sequence or coding sequence, the regenerated plant exhibits improved GLS disease resistance.
[0024] In each of the methods disclosed herein for introducing a construct comprising a nucleic acid encoding a PRR03c polypeptide (e.g., a PRR03c genomic sequence and / or a PRR03c coding sequence) into plant material, the method may further include confirming that, upon introduction of the construct, the plant material expresses the PRR03c polypeptide having a length of less than 1,000 amino acids, less than 900 amino acids, less than 800 amino acids, less than 700 amino acids, less than 600 amino acids, or less than 500 amino acids (e.g., about 489 amino acids).
[0025] In some embodiments, the compositions and methods relate to modified plant material having increased disease resistance, wherein prior to modification, the plant material lacks the PRR03c genome sequence or PRR03c coding sequence disclosed herein. The plant material is modified, for example, by genome editing to include the nucleotide sequence (i) the PRR03c genome sequence disclosed herein, which has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity with SEQ ID NO:1, or (ii) the PRR03c coding sequence disclosed herein, which has a length of less than 2,000 nucleotides, less than 1,900 nucleotides, or less than 1,800 nucleotides, or less than 1,700 nucleotides, or less than 1,600 nucleotides, or less than 1,500 nucleotides, or less than 1,400 nucleotides, and is identical to SEQ ID NO:1. NO:2 has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% nucleic acid sequence identity.
[0026] In each of the methods disclosed herein for modifying plant material that previously lacks a PRR03c genomic sequence or coding sequence to produce modified plant material comprising nucleic acid encoding a PRR03c polypeptide (e.g., a PRR03c genomic sequence and / or a PRR03c coding sequence), the method may further include confirming that the modified plant material expresses a PRR03c polypeptide having a length of less than 1,000 amino acids, less than 900 amino acids, less than 800 amino acids, less than 700 amino acids, less than 600 amino acids, or less than 500 amino acids (e.g., about 489 amino acids).
[0027] Methods for confirming the presence or absence of PRR03c peptides with lengths of less than 1,000 amino acids, less than 900 amino acids, less than 800 amino acids, less than 700 amino acids, less than 600 amino acids, or less than 500 amino acids (e.g., approximately 489 amino acids) may include, for example, confirming the nucleotide sequence encoding the PRR03c peptide by nucleic acid amplification and / or nucleotide sequencing. In some instances, confirmation methods (including amplification-based and sequencing-based methods) can be readily adapted to high-throughput analysis, for example, by using available high-throughput sequencing methods (e.g., hybridization sequencing). Other methods include, but are not limited to, label-assisted analysis, protein isolation and visualization hybridization, primer extension, oligonucleotide ligation, nuclease cleavage, microsequencing, and coded spheres. Such methods are reviewed in the following publications, including Gut, 2001, Hum. Mutat. [Human Mutation] 17:475; Shi, 2001, Clin. Chem. [Clinical Chemistry] 47:164; Kwok, 2000, Pharmacogenomics [Pharmacogenomics] 1:95; Bhattramakki and Rafalski, “Discovery and application of single nucleotide polymorphism markers in plants”, in PLAANT GENOTYPING: THE DNA FINGERPRINTING OFPLANTS (CABI Publishing, Wallingford, 2001). In other instances, PRR03c peptide confirmation may include protein analysis performed by methods such as electrophoresis (e.g., SDS-PAGE, IEF, or 2-D gel electrophoresis), Western blotting, mass spectrometry (e.g., liquid chromatography-coupled mass spectrometry (LC-MS), LC-MS and tandem mass spectrometry, MS / MS or tandem mass spectrometry, electrospray ionization (ESI) or ESI-LC, matrix-assisted laser desorption / ionization (MALDI) or MALDI ESI), etc.
[0028] This disclosure also provides for identifying, selecting, or producing plants using any of the methods presented herein. Detailed Implementation
[0029] As used herein, the singular forms “a / an” and “the” include a plural indicator unless the context clearly indicates otherwise. Thus, for example, reference to “cell” includes multiple such cells, and reference to “protein” includes reference to one or more proteins and their equivalents, and so on. All technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains, unless otherwise expressly stated.
[0030] The NBS-LRR (“NLR”) group of R genes is the largest class of R genes discovered to date. In Arabidopsis thaliana, more than 150 NLR genes are expected to exist in the genome (Meyers et al. (2003), Plant Cell, 15:809-834; Monosi et al. (2004), Theoretical and Applied Genetics, 109:1434-1447), while in rice, approximately 500 NLR genes have been predicted (Monosi (2004) ibid.). The NBS-LRR class of R genes consists of two subclasses. One class of NLR genes contains a TIR-Toll / interleukin-1-like domain at its N' terminus; these have so far been found only in dicotyledons (Meyers (2003) ibid.; Monosi (2004) ibid.). The second type of NBS-LRR contains a coiled-coil domain or (nt) domain at its N-terminus (Baiet et al. (2002) Genome Research, 12:1871-1884; Monosi (2004) ibid.; Pan et al. (2000) Journal of Molecular Evolution, 50:203-213). Both types of NBS-LRR have been found in dicotyledonous and monocotyledonous species (Bai (2002) ibid.; Meyers (2003) ibid.; Monosi (2004) ibid.; Pan (2000) ibid.).
[0031] The NBS domain of genes appears to play a role in signal transduction in plant defense mechanisms (van der Biezen et al. (1998), Current Biology: CB, 8:R226-R227). The LRR region appears to be the region that interacts with pathogen AVR products (Michelmore et al. (1998), Genome Research, 8:1113-1130; Meyers (2003) ibid.). Compared with the NB-ARC (NBS) domain, the LRR region is subjected to greater selective pressure to diversify (Michelmore (1998) ibid.; Meyers (2003) ibid.; Palomino et al. (2002), Genome Research, 12:1305-1315). LRR domains can also be found in other contexts; these 20-29 residue motifs are tandemly arranged in many proteins with multiple functions, such as hormone-receptor interactions, enzyme inhibition, cell adhesion, and cell transport. Recent studies have shown that LRR proteins are involved in early development, neural development, cell polarization, regulation of gene expression, and apoptosis signaling in mammals.
[0032] An allele or gene sequence is "associated" with a trait when it is part of or linked to a DNA sequence that influences trait expression. The presence of an allele is an indicator of how the trait will be expressed.
[0033] As used herein, “disease resistance” or “resistance to disease” refers to a plant that exhibits increased disease resistance compared to a control plant (e.g., a control plant may be a plant lacking a QTL or PRR03c gene that provides disease resistance but is otherwise identical to a disease-resistant plant). Disease resistance can manifest as fewer and / or smaller lesions, increased plant health, increased yield, increased root weight, increased plant vigor, less or no fading, increased growth, reduced necrotic area, or reduced wilting. In some embodiments, alleles may exhibit resistance to one or more diseases.
[0034] Diseases affecting maize plants include, but are not limited to, bacterial leafblight and stalk rot; bacterial leaf spot; bacterial stripe; chocolate spot; Goss's bacterial wilt and blight; holcus spot; purple leafsheath; seed rot-seedling blight; bacterial wilt; corn dwarf virus; anthracnose leaf blight; gray leaf spot; aspergillus ear and kernel rot; banded leaf and sheath spot; black bundle disease; and black kernel rot. Rot); white edge disease; brown spot; black spot; stalk rot; cephalosporium kernel rot; charcoal rot; corticium ear rot; curvularia leafspot; didymella leaf spot; diplomadia ear rot and stalk rot; seed rot; corn seedling blight; diplomadia leaf spot or leaf streak; downy mildews; brown stripe downy mildew; crazy top downy mildew; green ear downy mildew; graminicola downy mildew. Java downy mildew; Java downy mildew;Philippine downy mildew; sorghum downy mildew; spontaneum downy mildew; sugarcane downy mildew; dry ear rot; ergot; horse's tooth; corn eyespot; fusarium ear and stalk rot; fusarium blight; seedling root rot; gibberella ear and stalk rot; gray ear rot; gray leafspot; cercospora leaf spot; helminthosporium root rot; hormodendrum ear rot rot); cladosporium rot; hyalothyridium leaf spot; late wilt; northern leaf blight; white blast; crown stalk rot; corn stripe; northern leaf spot; helminthosporium ear rot; penicillium ear rot; corn blue eye; downy mildew; phaeocytostroma stalk rot and root rot; phaeosphaeria leaf spot; physalospora ear rot; botryosphaeria ear rot; pyrenochaeta stalk rot and root rot root rot); Pythium root rot; Pythium stalk rot; Red kernel disease;Rhizoctonia ear rot; sclerotial rot; Rhizoctonia root rot and stalk rot; rostratum leaf spot; common corn rust; southern corn rust; tropical corn rust; sclerotium ear rot; southern leaf blight; selenophoma leaf spot; sheath rot; shuck rot; silage mold; common smut; false smut; head smut; southern corn leafblight and stalk rot; southern leaf spot; black spore (tar) Trichoderma ear rot and root rot; white ear rot, root and stalk rot; yellow leaf blight; zonate leafspot; wheat striate mosaic; barley stripe mosaic; barley yellow dwarf; bromemosaic; cereal chlorotic mottle; maize lethal necrosis disease; cucumber mosaic; Johnson's grass mosaic; maize bushy stunt; maize chlorotic dwarf; maize chlorotic mottle. Maize dwarf mosaic; Maize leaf fleck;Maize clear ring spot; Maize rayado fino; Red leaf and red stripe; Maize red stripe; Maize ring mottle; Maize rough dwarf; Maize sterile stunt; Maize streak; Maize stripe; Maize tassel abortion; Maize vein enation; Maize wallaby ear; Maize white leaf; Maize white line mosaic. Millet red leaf disease; and northern cereal mosaic.
[0035] Diseases affecting plants include, but are not limited to, bacterial blight; bacterial leaf streak; foot rot; grain rot; sheath brown rot; blast; brown spot; crown sheath rot; downy mildew; eyespot; false smut; kernelsmut; leaf smut; leaf scald; narrow brown leafspot; root rot; seedling blight; sheath blight; sheath rot; sheath spot; Alternaria leaf spot; and stem rot.
[0036] Disease-resistant plants may exhibit an increase in resistance of 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% compared to control plants. In some embodiments, plants may exhibit an increase in plant health of 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% compared to control plants in the presence of disease.
[0037] In this article, "control plant" refers to a plant that does not express the PRR03c protein or lacks one or more copies of the PRR03c gene sequence disclosed herein, but is otherwise genetically or nearly genetically identical to the disease-resistant plant. Therefore, a control plant can be a plant that is more susceptible to gray leaf spot disease compared to a plant containing the PRR03c gene sequence disclosed herein.
[0038] As used herein, the term "chromosomal interval" refers to a continuous linear span of genomic DNA located on a single chromosome of a plant. Genetic elements or genes located on a single chromosomal interval are physically linked. There is no particular limitation on the size of a chromosomal interval. In some respects, genetic elements located within a single chromosomal interval are genetically linked, typically having a genetic recombination distance of, for example, less than or equal to 20 cM, or alternatively, less than or equal to 10 cM. That is, two genetic elements within a single chromosomal interval recombine at a frequency of less than or equal to 20% or less than or equal to 10%.
[0039] The term “hybrid” or “hybrid” refers to sexual hybridization and involves the fusion of two haploid gametes through pollination to produce diploid offspring (e.g., cells, seeds, or plants). The term encompasses both pollination of one plant by another and self-pollination (or self-pollination, such as when pollen and ovules come from the same plant).
[0040] A “superior strain” is any strain produced through breeding that focuses on superior agronomic traits.
[0041] "Exotic varieties," "tropical lines," or "exotic germplasm" are varieties derived from plants that do not belong to any available superior lines or germplasm. In the case of hybridization between two plant or germplasm varieties, the offspring of the exotic germplasm are not closely related to the superior germplasm from which it is hybridized. Most commonly, exotic germplasm is not derived from any known superior line, but is selected to introduce new genetic elements (usually new alleles) into the breeding program.
[0042] A "favorable allele" is an allele (marker, QTL, gene, etc.) at a specific locus that confers or contributes to an agronomically desired phenotype (e.g., disease resistance) and allows for the identification of plants possessing that agronomically desired phenotype. A marked favorable allele is a marker allele that separates from the favorable phenotype.
[0043] "Genetic marker" refers to a polymorphic nucleic acid within a population, and the alleles of this genetic marker can be detected and distinguished by one or more analytical methods (e.g., RFLP, AFLP, isoenzyme, SNP, SSR, etc.). The term also refers to a nucleic acid sequence complementary to a genomic sequence (e.g., nucleic acid) used as a probe. Markers corresponding to genetic polymorphisms among population members can be detected using methods recognized in the art. These methods include, for example, PCR-based sequence-specific amplification methods, restriction fragment length polymorphism detection (RFLP), isoenzyme marker detection, polynucleotide polymorphism detection via allele-specific hybridization (ASH), amplified variable sequence detection of plant genomes, autonomous sequence replication detection, simple repeat sequence detection (SSR), single nucleotide polymorphism detection (SNP), or amplified fragment length polymorphism detection (AFLP). Recognized methods are also known for detecting expressed sequence tags (ESTs) and SSR markers derived from EST sequences, as well as randomly amplified polymorphic DNA (RAPD).
[0044] "Germium" refers to genetic material that belongs to or originates from an individual (e.g., a plant), a group of individuals (e.g., a plant strain, variety, or family), or a clone derived from a strain, variety, species, or culture; or more generally, all individuals of one or more species (e.g., a maize germplasm collection or an Andean germplasm collection). Germplasm can be a part of an organism or a cell, or can be isolated from an organism or cell. Generally, germplasm provides genetic material with a specific molecular structure that provides the physical basis for some or all of the genetic qualities of an organism or cell culture. As used herein, germplasm includes cells, seeds, or tissues from which new plants can grow, or plant parts, such as leaves, stems, pollen, or cells, that can be cultured into a whole plant.
[0045] A "haplotype" is an individual's genotype at multiple genetic loci, that is, a combination of alleles. Typically, the genetic loci described by haplotypes are physically and genetically linked, that is, located on the same segment of chromosome.
[0046] The term "heterogeneity" is used to indicate that individuals within a group have different genotypes at one or more specific loci.
[0047] The heterosis response, or "heterogeneity," of a material can be defined by its performance exceeding the average of its parents (or high parent) when crossed with other dissimilar or unrelated groups.
[0048] A “hybrid group” comprises a set of genotypes that perform well when crossed with genotypes from different hybrid groups (Hallauer et al. (1998) Corn breeding, pp. 463-564. In GFSprague and JWDudley (ed.) Corn and corn improvement). Inbred lines are classified into hybrid groups and further subdivided into families within hybrid groups based on several criteria, such as pedigree, marker-based associations, and performance in hybrid combinations (Smith et al. (1990) Theor. Appl. Gen. 80:833-840). In the United States, the two most widely used hybrid groups are called the “Iowa Stiff Stalk Synthetic” (also referred to as “rigid stem” in this text) and “Lancaster” or “Lancaster Sure Crop” (sometimes referred to as NSS or non-rigid stem).
[0049] Some heterotic groups possess the traits required to be the female parent, while others possess the traits required to be the male parent. For example, in maize, yield results from publicly released inbred lines from a population called BSSS (Stiff Stalk Synthetics, Iowa) have led to these inbred lines and their derivatives becoming the female pool in the central maize belt. BSSS inbred lines have been hybridized with other inbred lines (e.g., SD 105 and Maiz Amargo), and the general group of this material is known as Stiff Stalk Synthetics (SSS), even though not all inbred lines are derived from the original BSSS population (Mikel and Dudley (2006) Crop Sci [Crop Science]: 46:1193-1205). By default, all other inbred lines that inbreed well with SSS inbred lines are assigned to the male pool and are named NSS (Non-Stiff Stalk) due to a lack of a better name. This group includes several major heterosis groups, such as Lancaster Surecrop, Iodent, and Leaming Corn.
[0050] The term "homogeneity" means that members of a group have the same genotype at one or more specific loci.
[0051] The term "hybrid" refers to offspring obtained from a cross between at least two genetically distinct parents.
[0052] The term "inbred line" refers to a line that has been bred to achieve genetic homogeneity.
[0053] The term "insertion or deletion" refers to an insertion or deletion in which one strain may be referred to as having an inserted nucleotide or DNA fragment relative to a second strain, or the second strain may be referred to as having a deleted nucleotide or DNA fragment relative to the first strain.
[0054] The term "introgression" refers to the phenomenon of a desired allele at a genetic locus being transferred from one genetic background to another. For example, the introgression of a desired allele at a designated locus can be transferred to at least one offspring via sexual hybridization between two parents of the same species, where at least one of these parents carries the desired allele in its genome. Alternatively, allele transfer can occur, for example, through recombination between two donor genomes, such as in fusion protoplasts, where at least one donor protoplast carries the desired allele in its genome. The desired allele can be detected, for example, by markers associated with the phenotype, at QTLs, transgenes, etc. Offspring containing the desired allele can be repeatedly backcrossed with lines having the desired genetic background and the desired allele selected to produce alleles fixed in the selected genetic background.
[0055] When "infiltration" is repeated two or more times, the method is often referred to as "backcrossing".
[0056] A "strain" or "variety" is a group of individuals that share the same parents, are usually inbred to some extent, and are typically homozygous and homogeneous (homogeneous or nearly homogeneous) at most loci. A "substrain" refers to a subgroup of inbred lines that are genetically distinct from other similar inbred subgroups that originated from the same ancestor.
[0057] As used herein, the term "linked" or "linkage" describes the degree to which one marker locus is associated with another marker locus or several other loci. The linkage between a molecular marker and a locus influencing a phenotype is expressed as a "probability" or "adjusted probability." Linkage can be expressed as a desired limitation or range. For example, in some embodiments, markers are linked (genetically or physically) when any marker is separated from any other marker by less than 50, 40, 30, 25, 20, or 15 map distance units (or cM) on a single meiotic map (based on a genetic map of a population that has undergone one round of meiosis (e.g., F2; the IBM2 map consists of multiple meiotic divisions). In some aspects, it is advantageous to define a bracketed linkage range, for example, between 10 cM and 20 cM, between 10 cM and 30 cM, or between 10 cM and 40 cM. The stronger the linkage of a marker to a second locus, the better the marker indicates the second locus. In this application, the phrase "closely linked" refers to recombination between two linked loci occurring at a frequency equal to or less than about 10% (i.e., separated by no more than 10 cM on a genetic map). In other words, the closely linked loci have at least a 90% chance of co-segregation. Marker loci are particularly useful to the subject matter of this disclosure when they show a significant probability of co-segregation (linkage) with a desired trait (e.g., resistance to GLS). Thus, closely linked loci (e.g., marker loci and second loci) can exhibit recombination frequencies of 10% or less, preferably about 9% or less, even more preferably about 8% or less, even more preferably about 7% or less, even more preferably about 6% or less, even more preferably about 5% or less, even more preferably about 4% or less, even more preferably about 3% or less, and even more preferably about 2% or less. In a highly preferred embodiment, the relevant loci show a recombination frequency of about 1% or less, such as about 0.75% or less, more preferably about 0.5% or less, or even more preferably about 0.25% or less. Two loci located on the same chromosome and having a distance such that recombination between the two loci occurs at a frequency of less than 10% (e.g., about 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25% or less) are also considered to be “proximate” to each other. Because one cM is the distance between two markers showing a 1% recombination frequency, any marker is closely linked (genetically and physically) to any other closely adjacent marker (e.g., at a distance equal to or less than 10 cM). Two closely linked markers on the same chromosome may be located to each other at a distance of 9, 8, 7, 6, 5, 4, 3, 2, 1, 0.75, 0.5, or 0.25 cM or closer. In some cases, two different markers may have the same genetic map coordinates.In this situation, the two markers are so close to each other that recombination between them occurs at such a low frequency that it is undetectable.
[0058] The term “linkage disequilibrium” refers to the non-random segregation of a genetic locus or trait (or both). In either case, linkage disequilibrium means that the associated loci are physically close enough along a segment of chromosome that they segregate together at a higher-than-random (i.e., non-random) frequency. Markers exhibiting linkage disequilibrium are considered linked. Linked loci have a greater than 50% chance (e.g., about 51% to about 100% chance) of co-segregating. In other words, two co-segregating markers have a recombination frequency of less than 50% (and, by definition, less than 50 cM separated on the same linkage group). As used herein, linkage can exist between two markers, or alternatively, between a marker and a locus influencing a phenotype. A marker locus can be “associated” (linked) with a trait. The degree of linkage between a marker locus and a locus influencing a phenotypic trait is measured, for example, by the statistical probability of co-segregation of the molecular marker with the phenotype (e.g., F-statistic or LOD score).
[0059] The most common measure of chain imbalance is r. 2 The assessment, the measure r 2 Calculate using the formula from the following literature: Hill, WG and Robertson, A, Theor. Appl. Genet. [Theoretical and Applied Genetics] 38:226-231 (1968). When r 2 When r = 1, there is a complete LD between the two marker loci, meaning that these markers have not yet undergone recombination segregation and have the same allele frequency. 2 The value depends on the group used. 2 A value greater than 1 / 3 indicates a sufficiently strong LD for localization (Ardlie et al. (2002) Nature Reviews Genetics 3:299-309). Therefore, when r between paired marker loci... 2 When the value is greater than or equal to 0.33, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0, the allele is in linkage disequilibrium.
[0060] As used in this article, “linkage equilibrium” describes a situation where two markers are independently separated, i.e., randomly assigned in offspring. Markers showing linkage equilibrium are considered unlinked (regardless of whether they are located on the same chromosome).
[0061] A "locus" is a position on a chromosome, such as the location of a nucleotide, gene, sequence, or marker.
[0062] The "log dominance (LOD) value" or "LOD score" (Risch, Science 255(5046):803-804(1992)) is used in genetic locus mapping to describe the degree of linkage between two marker loci. A LOD score of three between two markers indicates a 1000-fold higher probability of linkage than no linkage, while a LOD score of two indicates a 100-fold higher probability of linkage than no linkage. LOD scores greater than or equal to two can be used to detect linkage. The LOD score can also be used to show the strength of the association between marker loci and quantitative traits in quantitative trait locus mapping. In this case, the size of the LOD score depends on the tightness of the association between the marker locus and the locus influencing the quantitative trait, as well as the magnitude of the quantitative trait effect.
[0063] The term "plant" includes the whole plant, plant cells, plant protoplasts, plant cells or tissue cultures from which new plants can be regenerated, plant callus, plant clumps, and complete plant cells or parts, such as seeds, flowers, cotyledons, leaves, stems, buds, roots, and root tips. As used herein, "modified plant" means any plant that has undergone genetic changes due to human intervention. Modified plants may have genetic changes introduced through plant transformation, genome editing, mutagenesis, or conventional plant breeding.
[0064] A "marker" is a way of discovering a location on a genetic or physical map or a linkage between a marker and a trait locus (a locus that influences a trait). The location detected by a marker can be determined by detecting polymorphic alleles and their genetic localization, or by hybridization, sequence matching, or amplification of a physically mapped sequence. Markers can be DNA markers (detecting DNA polymorphisms), proteins (detecting variations in encoded polypeptides), or simple inherited phenotypes (such as the "waxy" phenotype). DNA markers can be developed from genomic nucleotide sequences or from expressed nucleotide sequences (e.g., from spliced RNA or cDNA). According to DNA marker technology, the marker can consist of a primer complementary to a sequence flanked at the locus and / or a probe that hybridizes to a polymorphic allele at the locus. DNA markers or genetic markers can also be used to describe genes, DNA sequences, or nucleotides on the chromosome itself (rather than components for detecting the gene or DNA sequence), and they are typically used when the DNA marker is associated with a specific trait in human genetics (e.g., breast cancer markers). The term marker locus is the locus (gene, sequence, or nucleotide) that the marker detects.
[0065] "Marker" can refer to the type of polymorphism detected and the marking technique used to detect that polymorphism. Marker types include, but are not limited to: restriction fragment length polymorphism detection (RFLP), isoenzyme labeling detection, randomly amplified polymorphic DNA (RAPD), amplified fragment length polymorphism detection (AFLP), simple repeat sequence detection (SSR), amplified variable sequence detection in plant genomes, autonomous sequence replication detection, or single nucleotide polymorphism detection (SNP). SNPs can be detected, for example, by DNA sequencing, PCR-based sequence-specific amplification methods, polynucleotide polymorphism detection via allele-specific hybridization (ASH), dynamic allele-specific hybridization (DASH), molecular beacons, microarray hybridization, oligonucleotide ligase analysis, Flap endonuclease, 5' endonuclease, primer extension, single-strand conformation polymorphism (SSCP), or temperature gradient gel electrophoresis (TGGE). DNA sequencing (such as pyrosequencing) has the advantage of being able to detect a series of linked SNP alleles that make up a haplotype. Haplotypes tend to be more informative than SNPs (detecting higher levels of polymorphism).
[0066] A “marker allele” can refer to one of several polymorphic nucleotide sequences found at a marker locus in a population.
[0067] Marker-assisted selection (MAS) is a method for selecting individual plants based on marker genotypes.
[0068] "Marker-assisted reverse selection" is a method that uses marker genotypes to identify plants that will not be selected, thereby removing these plants from the breeding process or planting.
[0069] "Marker haplotype" refers to the combination of alleles at a marker locus.
[0070] A "marker locus" is a specific chromosomal location in a species' genome where a specific marker can be found. Marker loci can be used to track the presence of second-linked loci (e.g., linked loci that influence the expression of phenotypic traits). For example, marker loci can be used to monitor the segregation of alleles at genetically or physically linked loci.
[0071] As described above, when identifying linked loci, the term "molecular marker" can be used to refer to a genetic marker or its encoded product (e.g., a protein) used as a reference point. Molecular markers can be derived from genomic nucleotide sequences or expressed nucleotide sequences (e.g., from spliced RNA, cDNA, etc.), or from encoded polypeptides. The term also refers to nucleic acid sequences complementary to or flanked by the marker sequence, such as nucleic acids used as probes or primer pairs capable of amplifying the marker sequence. A "molecular marker probe" is a nucleic acid sequence or molecule that can be used to identify the presence or absence of a marker locus, such as a nucleic acid probe complementary to the marker locus sequence. Alternatively, in some respects, a marker probe refers to a probe capable of distinguishing any type (i.e., genotype) of a specific allele present at a marker locus. They are "complementary" when nucleic acids specifically hybridize in solution. Some of the markers described herein are also called hybridization markers when located in insertion / deletion regions, such as the non-collinear regions described herein. This is because, by definition, the insertion region is a polymorphism concerning the absence of the insertion in a plant. Therefore, the marker only needs to indicate the presence or absence of the insertion / deletion region. Any suitable marker detection technique can be used to identify such hybridization markers, such as the SNP technique used in the examples provided in this article.
[0072] An allele is "negatively" correlated with a trait when it is linked to the trait, and when the presence of an allele is an indication that the desired trait or the form of the trait will not appear in plants containing that allele.
[0073] The terms "phenotype," "phenotypic trait," or "trait" can refer to the observable expression of a gene or gene series. A phenotype can be observable by the naked eye or by any other assessment method, such as weighing, counting, measurement (length, width, angle, etc.), microscopy, biochemical analysis, or electromechanical determination. In some cases, a phenotype is directly controlled by a single gene or genetic locus; this is called a "monogenous trait" or "simple inherited trait." In the absence of large-scale environmental variation, a monogenic trait can segregate in a population to give a "qualitative" or "discrete" distribution, meaning the phenotype is classified into discrete categories. In other cases, a phenotype is the result of multiple genes and can be considered a "polygenic trait" or "complex trait." A polygenic trait segregates in a population to give a "quantitative" or "continuous" distribution; that is, the phenotype cannot be separated into discrete categories. Both monogenic and polygenic traits can be influenced by the environment in which their expression occurs, but polygenic traits tend to have a larger environmental component.
[0074] A genome's "physical map" is a map showing the linear order of identifiable markers (including genes, markers, etc.) on chromosomal DNA. However, compared to a genetic map, the distances between markers are absolute (e.g., measured in base pairs or as separate and overlapping contiguous gene segments) and are not based on gene recombination (which can vary across different populations).
[0075] "Polymorphism" is a variation in the DNA between two or more individuals within a population. Polymorphism preferably has a frequency of at least 1% in the population. Useful polymorphisms can include single nucleotide polymorphisms (SNPs), simple repeat sequences (SSRs), or insertion / deletion polymorphisms (also referred to herein as "insertions and deletions").
[0076] "Production markers" or "production SNP markers" are markers that have been developed for high-throughput purposes. Production SNP markers have been developed for the detection of specific polymorphisms and are designed to be used with a variety of chemical reactions and platforms.
[0077] The term “quantitative trait locus” or “QTL” refers to a region of DNA associated with differential expression of a quantitative phenotypic trait in at least one genetic context (e.g., in at least one breeding population). The region of a QTL encompasses or is closely linked to one or more genes that influence the trait under consideration.
[0078] A "reference sequence" or "shared sequence" is a defined sequence used as the basis for sequence alignment. A labeled reference sequence is obtained by sequencing multiple lines at a given locus, comparing these nucleotide sequences in a sequence alignment program (such as Sequencher), and then obtaining the most universal nucleotide sequence for that alignment. Polymorphisms found in these individual sequences are labeled in this shared sequence. The reference sequence is typically not an exact copy of any individual DNA sequence, but rather represents a mixture of available sequences and is used to design primers and probes targeting polymorphisms within that sequence.
[0079] The marked “unfavorable allele” is a marker allele that is separated from the unfavorable plant phenotype, thus providing the benefit of identifying plants that can be removed from breeding programs or cultivation.
[0080] The term "yield" refers to the productivity of a specific plant product per unit area with commercial value. Yield is influenced by both genetic and environmental factors. "Agronomy," "agronomic traits," and "agronomic trait performance" refer to the traits (and potential genetic elements) of a given plant variety that contribute to yield during its growth period. Individual agronomic traits include emergence vigor, nutrient potential, stress tolerance, disease resistance or tolerance, herbicide resistance, branching, flowering, seed formation, seed size, seed density, lodging resistance, and threshing ability. Therefore, yield is the ultimate culmination of all agronomic traits.
[0081] This article provides statistically significant cosegregating marker loci that exhibit disease resistance traits that confer broad resistance to one or more specific diseases. The detection of these loci, or other linked loci, along with resistance genes, can be used as part of a breeding program in marker-assisted selection to produce plants resistant to one or more diseases.
[0082] Genetic mapping
[0083] It has been recognized that, in many cases, specific genetic loci associated with particular phenotypes (e.g., disease resistance) can be located within the genome of an organism. Plant breeders can advantageously use molecular markers to identify desired individuals by detecting marker alleles that show a statistically significant probability of co-segregation with the desired phenotype, exhibiting linkage disequilibrium. By identifying molecular markers or clusters of molecular markers that co-segregate with the target trait, plant breeders can rapidly select for the desired phenotype by choosing appropriate molecular marker alleles (a method known as marker-assisted selection or MAS).
[0084] Several methods can be used to detect molecular markers or clusters of molecular markers that co-segregate with a target trait (e.g., disease resistance). The basic idea behind these methods is to detect markers of alternative genotypes (or alleles) with significantly different mean phenotypes. Therefore, the magnitude or significance level of the difference between alternative genotypes (or alleles) at marker loci is compared. The location of the trait gene closest to one or more markers with the greatest correlation in genotype difference is inferred. Two methods for detecting loci of a target trait are: 1) population-based association analysis (i.e., association mapping) and 2) traditional linkage analysis.
[0085] Association mapping
[0086] Understanding the extent and patterns of linkage disequilibrium (LD) in the genome is a prerequisite for developing effective association methods for identifying and mapping quantitative trait loci (QTLs). Linkage disequilibrium (LD) refers to the non-random association of alleles in a set of individuals. When LD is observed in alleles at linked loci, it is measured as LD decay across a specific region of the chromosome. The extent of LD reflects the recombination history of that region. The average rate of LD decay in the genome can help predict the number and density of markers needed for genome-wide association studies and provide estimates with predictable resolution.
[0087] Association or LD mapping aims to identify significant genotype-phenotype associations. It has been developed and utilized as a powerful tool for fine mapping in crossbred species such as humans (Corder et al. (1994) "Protective effect of apolipoprotein-E type-2 allele for late-onset Alzheimer-disease", Nat Genet [Nature Genetics] 7:180-184; Hastbacka et al. (1992) "Linkage disequilibrium mapping in isolated founder populations: diastrophic dysplasia in Finland", Nat Genet [Nature Genetics] 2:204-211; Kerem et al. (1989) "Identification of the cystic fibrosis gene: genetic analysis", Science [Science] 245:1073-1080) and maize (Remington et al. (2001) “Structure of linkage disequilibrium and phenotype associations in themaize genome,” Proc Natl Acad SciUSA [Proceedings of the National Academy of Sciences] 98:11479-11484; Thornsberry et al. (2001) “Dwarf8 polymorphisms associate with variation in flowering time,” Nat Genet [Nature Genetics] 28:286-289; Reviewed by Flint-Garcia et al. (2003) “Structure of linkage disequilibrium in plants [Structure of linkage disequilibrium in plants],” Annu Rev Plant Biol. [Annual Review of Plant Biology] 54:357-374), where recombination between heterozygotes is frequent and leads to rapid decay of LD.In inbred species, recombination between homozygous genotypes is not genetically detectable, and the degree of LD is greater (i.e., larger linkage marker blocks are inherited together), which greatly enhances the ability to detect association localization (Wall and Pritchard, (2003) "Haplotype blocks and linkage disequilibrium in the human genome", Nat Rev Genet 4:587-597).
[0088] The recombination and mutation history of a population is a function of mating habits as well as the effective size and age of the population. Larger population sizes provide an enhanced likelihood of recombination detection, while older populations are generally associated with higher levels of polymorphism, both of which contribute to a significant increase in the observable rate of LD decay. On the other hand, smaller effective population sizes, such as those that have recently experienced genetic bottlenecks, tend to exhibit a slower rate of LD decay, resulting in broader haplotype conservation (Flint-Garcia et al., (2003) “Structure of linkage disequilibrium inplants”, Annu Rev Plant Biol. 54:357-374).
[0089] Superior breeding lines provide a valuable starting point for association analysis. This analysis uses quantitative phenotypic scores (e.g., disease tolerance grades from one to nine for each line) (rather than considering only the frequency distribution of tolerance versus resistance alleles in the inter-group allele distribution types analyzed). The availability of detailed phenotypic performance data collected over many years through breeding programs and the environment of numerous superior lines provide valuable datasets for genetic marker association mapping analysis. This paves the way for seamless integration between research and application and leverages historically accumulated datasets. However, understanding the relationship between polymorphism and recombination is useful for developing appropriate strategies to effectively extract the maximum information from these resources.
[0090] This type of association analysis neither generates nor requires any atlas data, and is independent of atlas location. This analysis compares the plant's phenotypic score to the genotype at different loci. Subsequently, using the previously identified atlas locations of these markers, any suitable atlas (e.g., composite atlases) can optionally be used to aid in observing the distribution of identified QTL markers and / or QTL marker clusters.
[0091] Traditional linkage analysis is based on the same principle; however, LD is generated by creating a population from a small number of founders. Founders are selected to maximize the level of polymorphism within the structured population, and the level of co-segregation of polymorphic loci with a given phenotype is assessed. Numerous statistical methods have been used to identify significant marker-trait associations. One such method is the interval localization method (Lander and Botstein, Genetics 121:185-199 (1989), where each of many locations along a genetic map (say, in 1 cM intervals) is tested for the probability that a gene controlling the desired trait is located at that location. Genotype / phenotype data are used to calculate the LOD score (logarithm of the probability ratio) for each tested location. When the LOD score is greater than a threshold, there is significant evidence that a gene controlling the desired trait is located at that location on the genetic map (between two specific marker loci).
[0092] This article presents marker loci that show statistically significant cosegregation with disease resistance traits, as identified through conventional linkage analysis and genome-wide association analysis. Detection of these loci, or other linked loci, can be used in marker-assisted breeding programs to produce disease-resistant plants.
[0093] Activities in marker-assisted breeding programs may include, but are not limited to: selecting from new breeding populations based on historical genotype and agronomic trait associations to identify which population has the highest frequency of favorable nucleic acid sequences; selecting favorable nucleic acid sequences in the progeny of the breeding population; selecting from parental lines based on predictions of progeny performance; and advancing lines in germplasm improvement activities based on the presence of favorable nucleic acid sequences.
[0094] Chromosome interval
[0095] Chromosomal regions associated with disease resistance traits are provided. Various methods can be used to identify these chromosomal regions. The boundaries of such chromosomal regions are extended to encompass markers that will be linked to one or more genes controlling the desired trait. In other words, chromosomal regions are extended such that any marker located within the region (including terminal markers defining the boundaries of the region) can be used as a marker for the disease resistance trait.
[0096] Conversely, for example, if two very close markers show co-segregation with the desired phenotypic trait, it is sometimes difficult to distinguish whether each of those markers identifies the same gene or two different genes or multiple genes. In any case, knowledge about how many genes are within a particular physical / genomic region is unnecessary for formulating or practicing which ones are presented in this disclosure.
[0097] Therefore, this document discloses intervals on chromosome 2 and intervals on chromosome 4. The intervals on chromosome 2 disclosed herein may cover 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13 of the GLS resistance markers disclosed in Table 10 of this disclosure. The intervals on chromosome 4 disclosed herein may cover 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, or 39 of the markers disclosed in Table 11 of this disclosure. Any marker located within these intervals can be used as a marker of GLS resistance and can be used in the context of the methods proposed herein to identify and / or select plants with GLS resistance, regardless of whether the resistance is newly conferred or enhanced compared to control plants. In some embodiments, markers located upstream and downstream of the PRR gene location are genetically and physically closely linked and can therefore be used to select the PRR gene for trait introgression and product development.
[0098] Chromosomal regions can also be defined by markers linked to disease resistance genes (which exhibit linkage disequilibrium), and r 2 Linkage disequilibrium (LD) is a common measure in the context of association studies. If the LD between a marker locus on chromosome 7 in the target region and another closely adjacent marker locus on chromosome 7 is r... 2 If the value is greater than 1 / 3 (Ardlie et al. (2002) ibid.), then the two loci are linked out of balance.
[0099] Markers and Linkages
[0100] A common measure of linkage is the frequency of cosegregation of traits. This can be expressed as a percentage of cosegregation (recombination frequency) or in centimoles (cM). cM is a unit of measurement for the frequency of genetic recombination. One cM equals a 1% chance that a trait at one locus will segregate from a trait at another locus due to hybridization in a single generation (meaning there is a total 99% chance of these traits segregating). Since chromosomal distance is roughly proportional to the frequency of hybridization events between traits, there exists an approximate physical distance associated with recombination frequency.
[0101] A marker locus is a trait in itself and can be evaluated during segregation by tracking the marker locus and performing standard linkage analysis. Therefore, a cM equals a 1% chance that a marker locus will segregate from another locus due to hybridization in a single generation.
[0102] The closer a marker is to the gene controlling the desired trait, the more effective and advantageous the marker is as an indicator of that trait. Closely linked loci exhibit a hybridization frequency of about 10% or less, preferably about 9% or less, even more preferably about 8% or less, even more preferably about 7% or less, even more preferably about 6% or less, even more preferably about 5% or less, even more preferably about 4% or less, even more preferably about 3% or less, and even more preferably about 2% or less. In a highly preferred embodiment, the associated loci (e.g., marker loci and target loci) exhibit a recombination frequency of about 1% or less, for example about 0.75% or less, more preferably about 0.5% or less, or even more preferably about 0.25% or less. Therefore, the loci are separated by a distance of about 10 cM, 9 cM, 8 cM, 7 cM, 6 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.75 cM, 0.5 cM, or 0.25 cM or less. In other words, two loci located on the same chromosome and having a distance that allows recombination between the two loci to occur at a frequency of less than 10% (e.g., about 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25% or less) are considered to be “neighbors” to each other.
[0103] Although specific marker alleles can co-segregate with the disease resistance trait, it is important to note that the marker locus does not necessarily induce the expression of that disease resistance phenotype. For example, it is not necessary for the marker polynucleotide sequence to be part of the gene that produces the disease resistance phenotype (e.g., part of the gene's reading frame). The association between a specific marker allele and the disease resistance trait arises from the initial "coupling" between the marker allele and the allele in the ancestral line from which the allele originated. This orientation can be altered by repeated recombination events between the marker and the genetic locus. For this reason, favorable marker alleles can change based on the linkage present in disease-resistant parents used to create segregating populations. This does not change the fact that markers can be used to monitor phenotypic segregation. It merely changes which marker allele is considered favorable in a given segregating population.
[0104] The method presented in this paper involves detecting the presence of one or more marker alleles associated with disease resistance in plants, and then identifying and / or selecting plants with favorable alleles at those marker loci. Markers have been identified in this paper as being associated with the disease resistance trait and can therefore be used to predict disease resistance in plants. Any marker within 50 cM, 40 cM, 30 cM, 20 cM, 15 cM, 10 cM, 9 cM, 8 cM, 7 cM, 6 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.75 cM, 0.5 cM, or 0.25 cM (based on a genetic map from a single meiotic division) can also be used to predict disease resistance in plants.
[0105] Marker assisted selection
[0106] Molecular markers have applications in a variety of plant breeding fields (e.g., see Staub et al. (1996) Hortscience [Horticultural Science] 31: 729-741; Tanksley (1983) Plant Molecular Biology Reporter. [Journal of Plant Molecular Biology] 1: 3-8). One of the main areas of interest is the use of marker-assisted selection (MAS) to increase the efficiency of backcrossing and gene introgression. Molecular markers demonstrating linkage to loci influencing desired phenotypic traits provide a useful tool for selecting traits in plant populations. This is especially true where phenotype determination is difficult. Because DNA marker determination is less labor-intensive and requires less physical space than field phenotypic analysis, it can be applied to larger populations, increasing the probability of finding recombinants with target segments that have moved from donor lines to recipient lines. The tighter the linkage, the more useful the marker, as recombination is less likely to occur between the marker and the gene causing the trait, which can lead to false positives. Lateralized markers reduce the probability of false positive selection because they require a double recombination event. In the most preferred embodiment, the marker is located within the gene itself, preventing recombination between the marker and the gene. In some embodiments, the methods disclosed herein generate markers in disease resistance genes, wherein the gene is identified by inferring its genomic location from clustering or cluster analysis of conserved domains.
[0107] When genes are introduced via MAS, not only the gene itself but also the lateral graft region is introduced (Gepts. (2002). Crop Sci; 42: 1780-1790). This is called "linkage cumbersomes." In cases where the donor and recipient plants are highly unrelated, these lateral graft regions carry additional genes that can encode traits that are agronomically unwanted. Linkage cumbersomes can lead to reduced yield or other negative agronomic traits even after multiple backcrosses with superior lines. This is sometimes also called "yield cumbersomes." The size of the lateral graft region can be reduced through further backcrosses, although this is not always successful because breeders cannot control the size of the region or the recombination breakpoint (Young et al., (1998) Genetics 120:579-585). In classical breeding, recombinations that help reduce the size of the donor segment are often chosen only by chance (Tanksley et al., (1989). Biotechnology 7: 257-264). Even after 20 backcrosses of this type, a considerable number of donor chromosome fragments still linked to the gene can be expected to be selected. However, with markers, it is possible to select rare individuals that have undergone recombination near the target gene. In 150 backcrossed plants, there is a 95% chance that at least one plant will undergo hybridization within 1 cM (based on the single meiotic map distance) of the gene. Marking allows for the definitive identification of these individuals. A subsequent backcross using 300 plants yields a 95% probability of hybridization within 1 cM of the single meiotic map distance on the other side of the gene, resulting in a segment near the target gene less than 2 cM based on the single meiotic map distance. This can be achieved in two generations with markers, compared to an average of 100 generations without markers (see Tanksley et al., ibid.). When the exact location of the gene is known, flanking markers around the gene can be used to select for recombination at different population sizes. For example, in smaller populations, recombination is expected to occur further away from the gene, thus requiring more distant side markers to detect the recombination.
[0108] The main components of implementing MAS are: (i) defining the population in which marker-trait associations will be measured, which may be a segregating population or a randomized or structured population; (ii) monitoring the segregation or association of polymorphic markers relative to the trait and using statistical methods to determine linkage or association; (iii) defining a set of desired markers based on the results of statistical analysis; and (iv) using and / or extrapolating this information to the current germplasm to enable marker-based selection decisions. The markers described in this disclosure, as well as other marker types such as SSR and FLP, can be used in marker-assisted selection schemes.
[0109] SSRs can be defined as relatively short sequences of tandemly repeated DNA of 6 bp or less in length (Tautz (1989) Nucleic Acid Research 17: 6463-6471; Wang et al. (1994) Theoretical and Applied Genetics 88: 1-6). Polymorphism arises from variations in the number of repeat units, which may be due to slippage during DNA replication (Levinson and Gutman (1987) Mol Biol Evol 4: 203-221). Variations in repeat length can be detected by designing PCR primers to conserved non-repetitive side-joined regions (Weber and May (1989) Am J Hum Genet. 44: 388-396). Because SSRs are multi-allelic, codominant, regenerable, and suitable for high-throughput automation, they are well-suited for mapping and MAS (Rafalski et al. (1996) Generating and using DNA markers in plants. In: Non-mammalian genomic analysis: a practical guide. Academic Press. pp. 75-135).
[0110] Various types of SSR markers can be generated, and SSR profiles can be obtained by gel electrophoresis of the amplification products. The marker genotype score is based on the size of the amplified fragment.
[0111] Various types of FLP markers can also be generated. Most commonly, amplification primers are used to generate fragment length polymorphisms. Except that the regions amplified by the primers are usually not highly repetitive, such FLP markers are similar to SSR markers in many ways. Usually, due to insertions or deletions, the amplified regions or amplicones still have enough variability across germplasm to allow the fragments generated by the amplification primers to be distinguished in polymorphic individuals, and such insertions and deletions are known to occur frequently in maize (Bhattramakki et al. (2002). Plant Mol Biol [Plant Molecular Biology] 48, 539-547; Rafalski (2002b), ibid.).
[0112] SNP markers detect single-base pair nucleotide substitutions. Among all molecular marker types, SNPs are the most abundant, thus possessing the potential to provide the highest genetic map resolution (Bhattramakki et al., 2002 Plant Molecular Biology 48:539-547). Because SNPs do not require large amounts of DNA and the automation of assays can be straightforward, they can be determined in a so-called "ultra-high throughput" manner, at throughput levels even higher than SSRs. SNPs also have the potential to be a relatively low-cost system. These three factors together make the use of SNPs in MAS highly attractive. Several methods can be used for SNP genotyping, including but not limited to: hybridization, primer extension, oligonucleotide ligation, nuclease digestion, microsequencing, and coding spheres. These methods have been reviewed in the following literature: Gut (2001) Hum Mutat [Human Mutation] 17 pp. 475-492; Shi (2001) Clin Chem [Clinical Chemistry] 47, pp. 164-172; Kwok (2000) Pharmacogenomics [Pharmacogenomics] 1, pp. 95-100; and Bhattramakki and Rafalski (2001) Discovery and application of single nucleotide polymorphism markers in plants. [Single Nucleotide Polymorphism Markers in Plants] in: RJ Henry, ed., Plant Genotyping: The DNA Fingerprinting of Plants, CABI Publishing, Wallingford. [Plant Genotyping: The DNA Fingerprinting of Plants, CABI Publishing, Wallingford] A wide range of commercially available technologies utilize these and other methods to detect SNPs. These commercially available technologies include: Masscode™ (Qiagen), INVADER® (Third Wave Technologies) and Invader PLUS®, SNAPSHOT® (Applied Biosystems), TAQMAN® (Applied Biosystems), and BEADARRAYS® (Illumina).
[0113] Haplotypes of any particular genotype can be described using numerous SNPs, either within a sequence or across linked sequences (Ching et al. (2002), BMC Genet. 3:19 pp; Gupta et al. 2001, Rafalski (2002b), Plant Science 162:329-333). Haplotypes can be more informative than a single SNP and can describe any particular genotype in more detail. For example, a single SNP might be the allele "T" of a specific line or variety with disease resistance, but the allele "T" could also be present in a breeding population used for recurrent parents. In such cases, haplotypes (e.g., combinations of alleles at linked SNP markers) may be more informative. Once a unique haplotype is assigned to a donor chromosome region, that haplotype can be used in that population or any subpopulation to determine whether an individual possesses the specific gene. The use of automated, high-throughput marker detection platforms makes this method efficient and effective.
[0114] Many of the markers presented herein can be readily used as single nucleotide polymorphism (SNP) markers for selecting the PRR03c gene. Using PCR, primers are used to amplify DNA segments representing the diversity of the target population (preferably inbred lines). The PCR products are sequenced directly in one or both directions. The resulting sequences are aligned and polymorphisms are identified. Polymorphisms are not limited to single nucleotide polymorphisms (SNPs) but also include insertions / deletions, CAPS, SSRs, and VNTRs (variable number of tandem repeats). In particular, based on the fine mapping information described herein, additional polymorphic SNPs (and other markers) within the regions amplified by the primers disclosed herein can be readily obtained. Markers within the described mapping regions can be hybridized with BAC or other genomic libraries, or electronically aligned with genomic sequences, to find new sequences in the same approximate locations as the markers.
[0115] In addition to SSR, FLP, and SNP mentioned above, other types of molecular markers are also widely used, including but not limited to: expressed sequence tags (EST), SSR tags derived from EST sequences, randomly amplified polymorphic DNA (RAPD), and other nucleic acid-based markers.
[0116] Isozyme profiles and linkage morphological features can also be used indirectly as markers in certain situations. Although they do not directly detect DNA differences, they are often influenced by specific genetic differences. However, there are far more and more diverse markers for detecting DNA variation than isozymes or morphological markers (Tanksley (1983) Plant Molecular Biology Reporter: 1:3-8).
[0117] Sequence alignment or contigs can also be used to discover upstream or downstream sequences of the specific markers listed herein. These new sequences, which are close to the markers described herein, are then used to discover and develop functionally equivalent markers. For example, alignments of different physical and / or genetic maps can be performed to locate equivalent markers not described in this disclosure but located in similar regions. These maps may be intra-species or even across other species for which genetic or physical alignments are performed.
[0118] Generally, MAS uses polymorphic markers that have been identified as having a significant probability of co-segregating with traits such as GLS disease resistance. Such markers are presumed to be located on the map near one or more genes that confer a disease resistance phenotype in the plant and are considered indicators of the desired trait or marker. The presence of the desired allele in the marker is tested in the plant, and plants containing the desired genotype at one or more loci are expected to transfer the desired genotype along with the desired phenotype to their progeny. Therefore, plants with GLS disease resistance can be selected by detecting one or more marker alleles, and further, progeny plants derived from these plants can also be selected. Alternatively, plants containing the desired genotype (i.e., the genotype associated with disease resistance) in a given chromosomal region are obtained and then crossed with another plant. The progeny of such a cross are then genotyped using one or more markers, and progeny plants with the same genotype in a given chromosomal region are then selected as disease-resistant.
[0119] SNPs (i.e., SNP haplotypes) can be used alone or in combination to select favorable resistance gene alleles associated with GLS disease resistance. For example, an SNP haplotype may include a combination of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13 of the GLS resistance markers in Table 10 of this disclosure. An SNP haplotype may also include a combination of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, or 39 of the markers in Table 11 of this disclosure.
[0120] Additional polymorphic loci may exist at or near marker loci identified by the methods disclosed herein, where one or more polymorphic loci are in linkage disequilibrium (LD) with alleles at one or more polymorphic loci in the haplotype, and thus can be used in marker-assisted selection procedures to infiltrate target gene alleles or target genomic fragments. Two specific alleles at different polymorphic loci are considered to be in LD if the presence of an allele at one of these loci tends to predict the presence of an allele at other loci on the same chromosome (Stevens, Mol. Diag. [Molecular Diagnostics] 4:309-17 (1999)). This marker locus may be located within 5 cM, 2 cM, or 1 cM of the QTL for the disease resistance trait (on a genetic map based on a single meiotic division).
[0121] Allele frequencies (and consequently haplotype frequencies) can vary between germplasm banks. Germplasm banks vary due to differences in maturity, heterosis grouping, geographical distribution, and other factors. Therefore, SNPs and other polymorphisms may not be available in some germplasm banks.
[0122] Proteins and their variants and fragments
[0123] This disclosure covers the PRR03c polypeptide. As used herein, “PRR03c polypeptide” and “PRR03c protein” can be used interchangeably to refer to one or more polypeptides having a length of less than 1,000 amino acids, less than 900 amino acids, less than 700, or less than 600 amino acids, or less than 500 amino acids, providing disease resistance to GLS. The PRR03c may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3. In some embodiments, sequence identity refers to the full-length sequence of the polypeptide. When used in the context of this document with a percentage of sequence identity, the term “about” means + / - 1.0 percentage point relative to the enumerated percentage.
[0124] As used in this article, “recombinant protein” refers to a protein that is no longer in its natural environment (e.g., in vitro or in recombinant bacterial or plant host cells); a protein expressed from a polynucleotide that has been edited from its natural version; or a protein expressed from a polynucleotide at a different genomic location relative to its natural sequence.
[0125] As used herein, “substantially free of cellular material” refers to a polypeptide comprising a protein formulation having less than about 30%, 20%, 10%, or 5% (by dry weight) of non-target proteins (also referred to herein as “contaminating proteins”).
[0126] The “fragment” or “bioactive part” includes a polypeptide fragment or polynucleotide fragment containing a sequence that is fully identical to the PRR03c polypeptide or polynucleotide, and exhibits disease resistance when expressed in a plant.
[0127] As used herein, “variant” refers to a protein or polypeptide having an amino acid sequence that is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identical to the parent amino acid sequence (SEQ ID NO:3).
[0128] In some instances, the PRR03c peptide has an amino acid sequence with at least about 40%, 45%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher identity with the full length or fragment of the amino acid sequence of SEQ ID NO:3, wherein the PRR03c peptide can provide GLS resistance when expressed in plants.
[0129] Methods for such operations are generally known in the art. For example, amino acid sequence variants of the PRR03c peptide can be prepared by mutation in the DNA. This can also be accomplished by one of several forms of mutagenesis, such as site-specific double-strand break techniques, and / or directed evolution. In some respects, the changes encoded in the amino acid sequence will have substantially no effect on the function of the protein. Such variants will possess the desired activity. However, it should be understood that in some cases, the ability of the PRR03c peptide to confer disease resistance can be improved by using such techniques on the compositions disclosed herein.
[0130] Nucleic acid molecules, variants and fragments thereof
[0131] This document provides isolated or recombinant nucleic acid molecules containing a nucleic acid sequence encoding a PRR03c polypeptide or its biologically active portion, as well as nucleic acid molecules sufficient to be used as hybridization probes to identify nucleic acid molecules encoding proteins having sequence homology regions. As used herein, the term "nucleic acid molecule" refers to DNA molecules (e.g., recombinant DNA, cDNA, genomic DNA, plasmid DNA, mitochondrial DNA) and RNA molecules (e.g., mRNA), as well as DNA or RNA analogs produced using nucleotide analogs. In some instances, nucleic acid molecules may be single-stranded. In some instances, nucleic acid molecules may be double-stranded.
[0132] As used herein, “isolated” nucleic acid molecules (e.g., RNA or DNA) refer to nucleic acid sequences (e.g., RNA or DNA) that are no longer in their natural environment (e.g., in vitro). As used herein, “recombinant” nucleic acid molecules (e.g., RNA or DNA) refer to nucleic acid sequences (e.g., RNA or DNA) in recombinant bacterial or plant host cells; these sequences have been edited from their natural sequences; or the sequences are located at a different position than the natural sequences. In some embodiments, “isolated” or “recombinant” nucleic acids do not contain sequences naturally located flanking the nucleic acid in the genomic DNA of the organism from which it is derived (i.e., sequences located at the 5' and 3' ends of the nucleic acid) (preferably sequences encoding proteins). For the purposes of this disclosure, “isolated” or “recombinant” when used to refer to nucleic acid molecules excludes isolated chromosomes. For example, in different embodiments, the recombinant nucleic acid molecule encoding the PRR03c polypeptide may contain nucleic acid sequences of less than about 5 kb, 4 kb, 3 kb, 2 kb, 1 kb, 0.5 kb, or 0.1 kb, which are naturally located on the flanking side of the nucleic acid molecule in the genomic DNA of the cell from which the nucleic acid is derived.
[0133] In some embodiments, the isolated nucleic acid molecule encoding the PRR03c polypeptide has one or more variations in its nucleic acid sequence compared to the natural or genomic nucleic acid sequence. In some embodiments, the alterations to the natural or genomic nucleic acid sequence include, but are not limited to: nucleic acid sequence changes due to the degeneracy of the genetic code; nucleic acid sequence changes due to amino acid substitutions, insertions, deletions, and / or additions compared to the natural or genomic sequence; removal of one or more introns; deletion of one or more upstream or downstream regulatory regions; and deletion of 5' and / or 3' untranslated regions associated with the genomic nucleic acid sequence. In some embodiments, the nucleic acid molecule encoding the PRR03c polypeptide is a non-genomic sequence.
[0134] Various polynucleotides encoding the PRR03c polypeptide or related proteins were considered. When operatively linked to suitable promoters, transcription terminators, and / or polyadenylated sequences, such polynucleotides can be used to generate the PRR03c polypeptide in host cells. These polynucleotides can also be used as probes to isolate homologous or substantially homologous polynucleotides encoding the PRR03c polypeptide or related proteins.
[0135] This document provides a nucleic acid molecule encoding the PRR03c polypeptide. Such a polynucleotide may have the sequence shown in SEQ ID NO:1 or SEQ ID NO:2, its variants, fragments, and complementary sequences. As used herein, a “complementary sequence” refers to a nucleic acid sequence that is sufficiently complementary to a given nucleic acid sequence, such that it can hybridize with that given nucleic acid sequence to form a stable double strand. An inverse complementary sequence is a complementary sequence formed by exchanging each A with T, T with A, C with G, and G with C in the sequence, and then reversing the 5' through 3' sequence of the exchanged sequences, such that the inverse complementary sequence of 5'-ACCTGAG-3' is 5'-CTCAGGT-3'. As used herein, a “polynucleotide sequence variant” refers to a nucleic acid sequence that encodes the same polypeptide except for the degeneracy of the genetic code.
[0136] In some instances, the nucleic acid molecule encoding the PRR03c polypeptide is a non-genomic nucleic acid sequence. As used herein, “non-genomic nucleic acid sequence” or “non-genomic nucleic acid molecule” or “non-genomic polynucleotide” refers to a nucleic acid molecule that has one or more alterations in its nucleic acid sequence compared to a natural or genomic nucleic acid sequence. In some instances, alterations to a natural or genomic nucleic acid molecule include, but are not limited to: changes in nucleic acid sequence due to degeneracy of the genetic code; optimization of the nucleic acid sequence for expression in plants; changes in the nucleic acid sequence by introducing at least one amino acid substitution, insertion, deletion, and / or addition compared to a natural or genomic sequence; removal of one or more introns associated with the genomic nucleic acid sequence; insertion of one or more heterologous introns; deletion of one or more upstream or downstream regulatory regions associated with the genomic nucleic acid sequence; insertion of one or more heterologous upstream or downstream regulatory regions; deletion of 5' and / or 3' untranslated regions associated with the genomic nucleic acid sequence; insertion of heterologous 5' and / or 3' untranslated regions; and modification of polyadenylation sites. In some instances, the non-genomic nucleic acid molecule is a synthetic nucleic acid sequence.
[0137] This article discloses examples of nucleic acid molecules encoding the PRR03c polypeptide, see, for example, the discussion of the PRR03c genome sequence and the discussion of the PRR03c coding sequence above. When expressed in plants, nucleic acid molecules encoding the PRR03c polypeptide can confer improved GLS resistance, for example, compared to syngeneic or near-syngeneic plants lacking the PRR03c genome sequence or the PRR03c coding sequence.
[0138] In some instances, the nucleic acid molecule encodes a PRR03c polypeptide variant that contains one or more amino acid substitutions relative to the amino acid sequence of SEQ ID NO:3.
[0139] Nucleic acid molecules that are fragments of these nucleic acid sequences encoding the PRR03c polypeptide are also covered in this disclosure. As used herein, "fragment" refers to a portion of the nucleic acid sequence encoding the PRR03c polypeptide. The fragment of the nucleic acid sequence may encode the biologically active portion of the PRR03c polypeptide, or it may be a fragment that can be used as a hybridization probe or PCR primer using the methods disclosed below. The nucleic acid molecule that is a fragment of the nucleic acid sequence encoding the PRR03c polypeptide contains at least about 150, 180, 210, 240, 270, 300, 330, 360, 400, 450, or 500 consecutive nucleotides, or up to the number of nucleotides present in the full-length nucleic acid sequence encoding the PRR03c polypeptide identified by the methods disclosed herein, depending on the intended use. As used herein, "consecutive nucleotides" refers to nucleotide residues that are adjacent to each other. The fragment of the nucleic acid sequence will encode a protein fragment that retains the biological activity of the PRR03c polypeptide and thus retains disease resistance. As used herein, “retained disease resistance” refers to a polypeptide having at least about 10%, at least about 30%, at least about 50%, at least about 70%, 80%, 90%, 95% or higher of the full-length PRR03c polypeptide shown in SEQ ID NO:3.
[0140] The "sequence identity percentage (%)" relative to the reference sequence (topic sequence) is determined as the percentage of amino acid residues or nucleotides in the candidate sequence (query sequence) that are identical to the corresponding amino acid residues or nucleotides in the reference sequence, after alignment and the introduction of gaps (if necessary) to achieve maximum sequence identity, and without considering any conserved substitutions of amino acids as part of sequence identity. Alignments performed for determining the sequence identity percentage can be performed in various ways, for example, using publicly available computer software such as BLAST, BLAST-2. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithms required to achieve maximum alignment across the full length of the sequences being compared. The identity percentage between two sequences is a function of the number of common positions shared by the sequences (e.g., identity percentage of the query sequence = number of common positions between the query and topic sequences / total number of positions in the query sequence × 100).
[0141] In some instances, the PRR03c polynucleotide encodes a PRR03c polypeptide comprising the following amino acid sequences: (i) having a length of less than 1,000, less than 900, less than 700, less than 600, or less than 500 amino acids; (ii) having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or higher identity with the entire length of the amino acid sequence throughout SEQ ID NO:3c; and (iii) providing improved GLS disease resistance relative to control plants. In some instances, the PRR03c polynucleotide comprises a genomic sequence including introns, regulatory elements, and untranslated regions.
[0142] This disclosure also provides nucleic acid molecules encoding variants of the PRR03c polypeptide. "Variants" of nucleic acid sequences encoding the PRR03c polypeptide include those sequences encoding the PRR03c polypeptide identified by the methods disclosed herein but exhibiting conserved differences due to the degeneracy of the genetic code, as well as those sequences that are sufficiently identical as discussed above. Naturally occurring allelic variants can be identified using well-known molecular biology techniques, such as polymerase chain reaction (PCR) and hybridization techniques outlined below. Variant nucleic acid sequences also include synthetically derived nucleic acid sequences that have been generated, for example, by site-directed mutagenesis, but still encode the PRR03c polypeptide disclosed herein.
[0143] Those skilled in the art will further understand that changes can be introduced through mutations in the nucleic acid sequence, resulting in alterations to the amino acid sequence of the encoded PRR03c polypeptide without changing the protein's biological activity. Therefore, variant nucleic acid molecules can be generated by introducing one or more nucleotide substitutions, additions, and / or deletions into the corresponding nucleic acid sequences disclosed herein, thereby introducing one or more amino acid substitutions, additions, or deletions into the encoded protein. Mutations can be introduced using standard techniques, such as site-directed mutagenesis and PCR-mediated mutagenesis. Such variant nucleic acid sequences are also covered by this disclosure.
[0144] Alternatively, variant nucleic acid sequences can be prepared by randomly introducing mutations along all or part of the coding sequence (e.g., through saturation mutagenesis), and the resulting mutants can be screened for activity to identify the ability of mutants to retain activity. After mutagenesis, the encoded protein can be recombinantly expressed, and the activity of the protein can be determined using standard assay techniques.
[0145] The polynucleotides and fragments thereof disclosed herein may be optionally used as substrates for various recombination and recursive recombination reactions, in addition to standard cloning methods described, for example, by Ausubel, Berger, and Sambrook, i.e., to produce additional polypeptide homologs and fragments thereof with desired properties. Various such reactions are known. Methods for generating variants of any of the nucleic acids listed herein (methods involving recursively recombinating such polynucleotides with a second (or more) polynucleotide to form a variant polynucleotide library) are also examples of this disclosure, as are the resulting libraries, the cells containing such libraries, and any recombinant polynucleotides produced by such methods. Furthermore, such methods may optionally include selecting variant polynucleotides from such libraries based on activity, as in which such recursive recombination is performed in vitro or in vivo.
[0146] Various diversity generation protocols, including nucleic acid recursive recombination protocols, are available. These procedures can be used alone and / or in combination to generate one or more variants of nucleic acids or nucleic acid assemblies, as well as variants of the proteins they encode. Individually or collectively, these procedures provide robust and widely applicable methods for generating diverse nucleic acids and nucleic acid assemblies, including, for example, nucleic acid libraries, which can be used for, for example, the engineering or rapid evolution of nucleic acids, proteins, pathways, cells, and / or organisms with novel and / or improved characteristics.
[0147] Although distinctions and classifications have been made in the subsequent discussion for clarity, it should be understood that these techniques are generally not mutually exclusive. In fact, various methods can be used individually or in combination, in parallel or in series, to obtain different sequence variants.
[0148] The result of any diversity-generating procedure described herein may be the production of one or more nucleic acids, which can select or screen nucleic acids having or conferring desired properties or nucleic acids encoding proteins having or conferring desired properties. Following diversification by one or more methods described herein or otherwise available to those skilled in the art, any of the produced nucleic acids can be selected for desired activity or properties, such as such activity at a desired pH. This may include the identification by any assay in the art that can, for example, be detected in an automated or automatable manner. Various related (or even unrelated) properties may be evaluated in series or in parallel by the practitioner at their discretion.
[0149] The nucleotide sequences disclosed herein can also be used to isolate corresponding sequences from different sources. In this way, such sequences can be identified using methods such as PCR, hybridization, etc. (based on their sequence homology with sequences identified by the methods disclosed herein). This disclosure covers sequences selected based on sequence identity with all sequences or fragments thereof shown herein. Such sequences include sequences that are orthologs of these sequences. The term "ortholog" refers to a gene that is derived from a common ancestral gene and is found in different species due to speciation. Genes found in different species are considered orthologs when their nucleotide sequences and / or the protein sequences they encode share a basic identity as defined elsewhere herein.
[0150] In PCR methods, oligonucleotide primers can be designed for PCR reactions to amplify corresponding DNA sequences from cDNA or genomic DNA extracted from any target organism. Methods for designing PCR primers and PCR cloning are disclosed in the following literature: Sambrook et al. (1989) *Molecular Cloning: A Laboratory Manual* (2nd ed., Cold Spring Harbor Laboratory Press, Plainview, New York), hereinafter referred to as "Sambrook". See also, Innis et al., ed. (1990) *PCR Protocols: A Guide to Methods and Applications* (Academic Press, New York); Innis and Gelfand, ed. (1995) *PCR Strategies* (Academic Press, New York); and Innis and Gelfand, ed. (1999) *PCR Methods Manual* (Academic Press, New York). Known PCR methods include, but are not limited to, methods using paired primers, nested primers, single-specific primers, degenerate primers, gene-specific primers, vector-specific primers, and partially mismatched primers.
[0151] In hybridization methods, all or part of a nucleic acid sequence can be used to screen cDNA or genomic libraries. Methods for constructing such cDNA and genomic libraries are disclosed in Sambrook and Russell, (2001), ibid. Hybridization probes can be genomic DNA fragments, cDNA fragments, RNA fragments, or other oligonucleotides, and can be coupled with a detectable group (such as…). 32The probe is labeled with P or any other detectable marker, such as other radioisotopes, fluorescent compounds, enzymes, or enzyme cofactors. Probes for hybridization can be prepared by labeling synthetic oligonucleotides based on nucleic acid sequences encoding known polypeptides disclosed herein. Degenerate primers, designed based on conserved nucleotides or amino acid residues in the nucleic acid sequence or the encoded amino acid sequence, may also be used. The probe typically contains a region of a nucleic acid sequence that, under stringent conditions, hybridizes to at least about 12, at least about 25, at least about 50, 75, 100, 125, 150, 175, or 200 consecutive nucleotides of a nucleic acid sequence encoding a polypeptide or a fragment or variant thereof. Methods and stringent conditions for preparing probes for hybridization are disclosed in Sambrook and Russell (2001), ibid.
[0152] Nucleotide constructs, expression cassettes, and vectors
[0153] The use of the term "construction" in connection with isolated and / or heteropolynucleotides herein is not intended to limit this disclosure to constructs containing DNA. Polynucleotide constructs, particularly polynucleotides and oligonucleotides composed of ribonucleotides, as well as combinations of ribonucleotides and deoxyribonucleotides, may also be used in the methods disclosed herein. The isolated polynucleotide constructs, nucleic acids, and nucleotide sequences disclosed herein further cover all complementary forms (e.g., reverse complementary sequences) for each sequence disclosed for such constructs. Additionally, the polynucleotide constructs and nucleotide sequences disclosed herein may cover any such constructs, molecules, and sequences suitable for use in methods for transforming the plant material disclosed herein. Such constructs may include naturally occurring molecules and / or synthetic analogs. The nucleotide constructs, nucleic acids, and nucleotide sequences disclosed herein also cover all forms of nucleotide constructs, including but not limited to single-stranded forms, double-stranded forms, hairpins, stem-loop structures, etc.
[0154] The transformed organisms disclosed herein include plant cells, bacteria, yeast, baculoviruses, protozoa, nematodes, and algae. The transformed organisms contain the disclosed sequences (e.g., as part of a construct, expression cassette, or vector containing the disclosed nucleotide sequences associated with GLS disease resistance).
[0155] The sequences disclosed herein can be used in constructs expressed in a target organism. Constructs may include a 5' and 3' regulatory sequence operatively linked to the coding sequence of the disclosed PRR03c polypeptide. As used herein, the term "operatively linked" refers to a functional linkage between a promoter and / or regulatory sequence and a second sequence, wherein the promoter and / or regulatory sequence initiates, mediates, and / or influences transcription of a DNA sequence corresponding to the second sequence. Typically, operatively linked means that the linked nucleic acid sequences are contiguous and, if necessary, link two protein-coding regions in the same reading frame. Constructs may additionally contain at least one additional gene to be co-transformed into an organism. Alternatively, one or more additional genes may be provided on multiple DNA constructs.
[0156] The provided DNA construct has multiple restriction sites for inserting the polypeptide gene sequence disclosed herein, which will be under transcriptional regulation in the regulatory region. The DNA construct may additionally contain a selective marker gene.
[0157] Generally, the DNA construct will include, in the transcriptional direction from 5' to 3', a transcription and translation initiation region (e.g., a promoter), the DNA sequence of the embodiment, and a transcription and translation termination region (e.g., a termination region) that functions in the host organism. The transcription initiation region (e.g., the promoter) may be natural, similar, exogenous, or heterologous to the host organism and / or sequence of the embodiment. Furthermore, the promoter or regulatory sequence may be a natural sequence, or alternatively, a synthetic sequence. As used herein, the term "exogenous" means that the promoter is not found in the natural organism to which it is introduced. As used herein, the term "heterogeneous" with respect to a sequence means that the sequence originates from a foreign species, or, if it originates from the same species, is a sequence substantially modified from its natural form in the composition and / or genomic locus through deliberate human intervention. As used herein, a chimeric gene contains a coding sequence operatively linked to a transcription initiation region that is heterologous to that coding sequence. When the promoter is a native sequence, the expression of the operablely linked sequence changes from wild-type expression, which leads to a change in phenotype.
[0158] In some embodiments, the DNA construct comprises a polynucleotide encoding the PRR03c polypeptide of the embodiment. In some embodiments, the DNA construct comprises a polynucleotide encoding a fusion protein comprising the PRR03c polypeptide of the embodiment.
[0159] In some embodiments, the DNA construct may also include a transcriptional enhancer sequence. As used herein, the term "enhancer" refers to a DNA sequence that can stimulate promoter activity and may be an innate or heterologous element of a promoter inserted to enhance promoter level or tissue specificity. Various enhancers include, for example, introns that have gene expression-enhancing properties in plants (US Patent Application Publication No. 2009 / 0144863), ubiquitin introns (i.e., corn ubiquitin intron 1 (see, for example, NCBI sequence S94464)), ω enhancers or ω major enhancers (Gallie et al. (1989) Molecular Biology of RNA, ed.: Cech (Liss, NY) 237-256 and Gallie et al. (1987) Gene 60:217-25), CaMV 35S enhancers (see, for example, Benfey et al. (1990) EMBOJ. 9:1685-96), and enhancers of US Patent No. 7,803,992 may also be used. The above list of transcriptional enhancers is not intended to be limiting. Any suitable transcriptional enhancer may be used in the examples.
[0160] The termination region may be natural for the transcription start region, natural for the target DNA sequence to which it is operatively linked, natural for the plant host, or derived from another source (i.e., exogenous or heterologous for the promoter, target sequence, plant host, or any combination thereof).
[0161] Convenient termination regions can be obtained from Ti plasmids of Agrobacterium tumefaciens, such as the termination regions of octopaline synthase and carmine synthase. See also Guerineau et al. (1991) Mol. Gen. Genet. [Molecular Genetics and General Genetics] 262:141-144; Proudfoot (1991) Cell [Cell] 64:671-674; Sanfacon et al. (1991) Genes Dev. [Genes and Development] 5:141-149; Mogen et al. (1990) Plant Cell [Plant Cell] 2:1261-1272; Munroe et al. (1990) Gene [Genes] 91:151-158; Ballas et al. (1989) Nucleic Acids Res. [Nucleic Acids Research] 17:7891-7903; and Joshi et al. (1987) Nucleic Acids Res. [Nucleic Acids Research] 15:9627-9639.
[0162] Nucleic acids can be optimized to increase expression in the host organism when appropriate. Therefore, in the case of a plant as the host organism, synthetic nucleic acids can be synthesized using plant-preferred codons to improve expression. For a discussion of the use of host preferences, see, for example, Campbell and Gowri (1990) Plant Physiol. [Plant Physiology] 92:1-11. For example, while the nucleic acid sequences of the examples can be expressed in both monocotyledonous and dicotyledonous plant species, the sequences can be modified to account for the specific preferences and GC content preferences of monocotyledonous or dicotyledonous plants, as these preferences have shown differences (Murray et al. (1989) Nucleic Acids Res. [Nucleic Acid Research] 17:477-498). Thus, plant preferences for specific amino acids can be derived from known gene sequences of plants.
[0163] Other sequence modifications are known to enhance gene expression in the host cell. These include the removal of sequences encoding pseudopolyadenylation signals, exon-intron splicing site signals, transposon-like repeat sequences, and other well-characterized sequences that may be detrimental to gene expression. The GC content of a sequence can be adjusted to the average level for a given host cell, as calculated by referencing known genes expressed in that host cell. As used herein, the term “host cell” refers to a cell that contains a vector and supports the replication and / or expression of the expression vector. Host cells can be prokaryotic cells such as *E. coli*, or eukaryotic cells such as yeast, insect, amphibian, or mammalian cells, or monocotyledonous or dicotyledonous plant cells. An example of a monocotyledonous host cell is a corn host cell. When possible, sequences are modified to avoid the occurrence of predicted hairpin secondary mRNA structures.
[0164] In preparing expression cassettes, various DNA fragments can be manipulated to provide DNA sequences that are oriented correctly and, at the appropriate time, within the correct reading frame. This can be achieved by using adapters or linkers to ligate the DNA fragments, or by employing other techniques to provide convenient restriction sites, remove redundant DNA, and eliminate restriction sites. For this purpose, in vitro mutagenesis, primer repair, restriction enzyme digestion, annealing, and substitution (e.g., conversion and transversion) may be involved.
[0165] Many promoters are available for implementing these embodiments. A promoter can be selected based on the desired results. Nucleic acids can be combined with constitutive, tissue-biased, inducible, or other promoters for expression in a host organism.
[0166] Plant transformation
[0167] The methods of these embodiments involve introducing a polypeptide or polynucleotide into a plant. As used herein, “introduction” means presenting the polynucleotide or polypeptide to the plant in such a way that the sequence enters the interior of the plant cell. The methods of these embodiments are not dependent on a specific method for introducing one or more polynucleotides or one or more polypeptides into a plant, as long as the polynucleotide or polypeptide enters the interior of at least one cell of the plant. Methods for introducing one or more polynucleotides or one or more polypeptides into a plant include, but are not limited to, stable transformation, transient transformation, and virus-mediated transformation.
[0168] As used herein, "stable transformation" means that a nucleotide construct introduced into a plant integrates into the plant's genome and is inherited by its offspring. As used herein, "transient transformation" means the introduction of a polynucleotide into a plant that does not integrate into the plant's genome, or the introduction of a polypeptide into a plant. As used herein, "plant" refers to the whole plant, plant organs (e.g., leaves, stems, roots), seeds, plant cells, propagules, and their embryos and offspring. Plant cells can be differentiated or undifferentiated (e.g., callus, suspension culture cells, protoplasts, leaf cells, root cells, phloem cells, and pollen).
[0169] The transformation scheme and the scheme for introducing nucleotide sequences into plants can vary depending on the type of plant or plant cell to be targeted for transformation (i.e., monocots or dicots). Suitable methods for introducing nucleotide sequences into plant cells and subsequently inserting them into the plant genome include microinjection (Crossway et al. (1986) Biotechniques [Biotechnology] 4:320-334), electroporation (Riggs et al. (1986) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 83:5602-5606), Agrobacterium-mediated transformation (US Patent Nos. 5,563,055 and 5,981,840), direct gene transfer (Paszkowski et al. (1984) EMBO J [Journal of the European Society for Molecular Biology] 3:2717-2722), and ballistic particle acceleration (see, for example, US Patent Nos. 4,945,050; 5,879,918; 5,886,244 and 5,932,782; Tomes et al. (1995) in Plant Cell, Tissue, and Organ Culture: Fundamental In Methods [Plant Cell, Tissue and Organ Culture: Basic Methods], edited by Gamborg and Phillips (Springer-Verlag, Berlin [Springer Berlin, Germany]); and McCabe et al. (1988) Biotechnology 6:923-926; and the Lecl transformation method (WO 00 / 28058). For potato transformation, see Tu et al. (1998) Plant Molecular Biology 37:829-838 and Chong et al. (2000) Transgenic Research 9:71-78. Other transformation methods can be found in the following literature: Weissinger et al. (1988) Ann. Rev. Genet. [Annals of Genetics] 22:421-477; Sanford et al. (1987) Particulate Science and Technology [Particle Science and Technology] 5:27-37 (Onion); Christou et al. (1988) Plant Physiol.[Plant Physiology] 87:671-674 (Soybean); McCabe et al. (1988) Bio / Technology 6:923-926 (Soybean); Finer and McMullen (1991) In Vitro Cell Dev. Biol. 27P:175-182 (Soybean); Singh et al. (1998) Theor. Appl. Genet. 96:319-324 (Soybean); Datta et al. (1990) Biotechnology 8:736-740 (Rice); Klein et al. (1988) Proc. Natl. Acad. Sci. USA 85:4305-4309 (Corn); Klein et al. (1988) Biotechnology 6:559-563 (Corn); US Patent Nos. 5,240,855; 5,322,783 and 5,324,646; Klein et al. (1988) Plant Physiol. 91:440-444 (Corn); Fromm et al. (1990) Biotechnology 8:833-839 (Corn); Hooykaas-Van Slogteren et al. (1984) Nature (London) 311:763-764; US Patent No. 5,736,369 (Cereals); Bytebier et al. (1987) Proc. Natl. Acad. Sci. USA 84:5345-5349 (Liliaceae); De Wet et al. (1985) The Experimental Manipulation of Ovule Tissues, edited by Chapman et al. (Longman, New York), pp. 197-209 (pollen); Kaeppler et al. (1990) Plant Cell Reports 9:415-418 and Kaeppler et al. (1992) Theor. Appl. Genet.[Theoretical and Applied Genetics] 84:560-566 (whisker-mediated transformation); D'Halluin et al. (1992) Plant Cell 4:1495-1505 (electroporation); Li et al. (1993) Plant Cell Reports 12:250-255; and Christou and Ford (1995) Annals of Botany 75:407-413 (rice); Osjoda et al. (1996) Nature Biotechnology 14:745-750 (maize via Agrobacterium tumefaciens).
[0170] Methods of introducing genome editing technology into plants
[0171] In some embodiments, genome editing techniques can be used to introduce a polynucleotide encoding the PRR03c polypeptide into the plant genome. For example, the identified polynucleotide can be introduced into the desired location in the plant genome using endonucleases or double-strand breaking techniques (e.g., TALEN, large-scale nucleases, zinc finger nucleases, CRISPR-Cas, etc.). For example, for site-specific insertion purposes, the CRISPR-Cas system can be used to introduce the PRR03c genomic sequence or the PRR03c coding sequence into the desired location in the genome. The desired location in the plant genome can be any target site required for insertion, such as a genomic region suitable for breeding, or it can be a target site located within a genomic window with an existing target trait. The existing target trait may be an endogenous trait or a previously introduced trait. Thus, for example, the PRR gene can be altered by genome editing at its native site to encode the PRR03c polypeptide having the amino acid sequence shown in SEQ ID NO:3. Alternatively or additionally, the PRR03c gene can be introduced into different genomic locations by genome editing. For example, a nucleotide construct encoding the PRR03c polypeptide can be inserted into chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10, or alternatively, into a location on chromosome 4 other than its natural locus between cM 90 and cM 115 on chromosome 4, wherein the PRR03c polypeptide (i) has a length of less than 1,000 amino acids, less than 900 amino acids, less than 700 amino acids, less than 600 amino acids, or less than 500 amino acids, and (ii) has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% amino acid sequence identity with SEQ ID NO:3.
[0172] For example, the use of double-strand break techniques (e.g., Cas endonuclease-gRNA complexes) has been described in U.S. Patent Application Publications 2015 / 0082478 and 2015 / 0059010, International Application Publications WO 2015 / 026886, WO 2016 / 007347 and WO 2016 / 25131, and U.S. Patent No. 10,934,536. As used herein, a Cas endonuclease refers to an enzyme derived from Cas ( C Cas proteins are polypeptides encoded by RISPR-associated genes. Cas proteins include, but are not limited to: Cas9, Cpf1 (Cas12), C2c1, C2c2, C2c3, Cas3, Cas3-HD, Cas5, Cas7, Cas8, Cas10, or combinations or complexes thereof. When complexed with a guiding polynucleotide, the "guiding polynucleotide / Cas endonuclease complex" (or "guiding polynucleotide / Cas endonuclease system," "guiding polynucleotide / Cas complex," "guiding polynucleotide / Cas system," and "guided Cas system," or "polynucleotide-guided endonuclease," "PGEN") guides the Cas endonuclease to the DNA target site, enabling the Cas endonuclease to recognize, bind to, and create a nick or cut (introducing single-strand or double-strand breaks) at the DNA target site. The guiding Cas system mentioned in this article may contain one or more Cas proteins and one or more suitable polynucleotide components of any known CRISPR system (Horvath and Barrangou, 2010, Science 327:167-170; Makarova et al. 2015, Nature Reviews Microbiology Vol. 13:1-15; Zetsche et al. 2015, Cell 163, 1-13; Shmakov et al. 2015, Molecular Cell 60, 1-13).
[0173] In some embodiments, where the GLS resistance PRR03c gene allele has been identified in the genome, genome editing techniques can be used to alter or modify polynucleotide sequences. Site-specific modifications that can be introduced into the desired PRR gene allele polynucleotide include modifications produced using any method for introducing site-specific modifications, including but not limited to the use of gene repair oligonucleotides (e.g., U.S. Publication 2013 / 0019349) or by using double-strand break techniques such as TALEN, large-scale nucleases, zinc finger nucleases, CRISPR-Cas, etc. Such techniques can be used to modify previously introduced polynucleotides by inserting, deleting, or substituting nucleotides within the introduced polynucleotide. Alternatively, double-strand break techniques can be used to add additional nucleotide sequences to the introduced polynucleotide. Additional sequences that can be added include additional expression elements (such as enhancer sequences and promoter sequences). In another embodiment, genome editing techniques can be used to localize additional disease resistance proteins within the plant genome near the PRR polynucleotide composition to produce a molecular stack of disease resistance proteins.
[0174] The terms “altered target site,” “altered target sequence,” “modified target site,” and “modified target sequence” are used interchangeably herein and mean a target sequence as disclosed herein that contains at least one alteration when compared to an unaltered target sequence. Such an “alteration” includes, for example: (i) substitution of at least one nucleotide, (ii) deletion of at least one nucleotide, (iii) insertion of at least one nucleotide, or (iv) any combination of (i)-(iii). Example
[0175] The following examples are provided to illustrate, but do not limit, the claimed subject matter. It should be understood that the examples and embodiments described herein are for illustrative purposes only, and those skilled in the art will recognize that various reagents or parameters may be changed without departing from the spirit of this disclosure or the scope of the appended claims.
[0176] Example 1. QTL mapping
[0177] Gray leaf spot (GLS), caused by the fungal pathogen *Cercospora zeae-maydis*, is a devastating foliar disease of corn that leads to persistent and significant yield losses (Wiebold et al. 2020). To identify natural corn genes that can confer GLS resistance, a localization population was generated by crossing GLS-resistant lines with susceptible inbred lines, as described in WO 2018013323. Plant materials were backcrossed with susceptible inbred lines to generate localization populations. First-generation backcrosses (BC1) were screened to determine phenotypic segregation of indicators of major dominant QTLs. Between growth stages V10 and V15, plants were inoculated 3–4 times with *Cercospora zeae-maydis* spores on a sorghum seed carrier. Four to six weeks after flowering, GLS severity was assessed using a visual score based on leaf area affected by *Cercospora zeae-maydis* lesions (GLFSPT). Severity was determined using a scale of 1-9, where 9 represents the highest resistance and 1 represents the highest susceptibility. Scores of 1-3 were considered susceptible, scores between 4-6 were considered median, and scores of 7-9 were classified as resistant. Individual plants were genotyped using SNP markers for marker-trait association analysis. The Kruskal-Wallis method was used with TIBCO... ® Spotfire ® The software package (version 10.3.3) analyzed the location data to compare numerical and categorical variables. Initial QTL mapping placed the quantitative trait locus (QTL) for GLS resistance between 94.78 and 113.94 cM on Chr04, with a correlation peak at 111.72 cM (PHM586-10). Table 1 provides p-values representing the correlation between genotype and gray leaf spot phenotype at a given genetic and physical location. Genetic locations were based on an internally proprietary B73 corn cultivar genome map, and physical genomic locations were based on publicly available B73 corn cultivar genome sequences (version 2). See Schnable et al. (2009) Science [Science] 326(5956): 1112-5.
[0178] Table 1
[0179]
[0180] Example 2. QTL fine mapping and candidate gene identification
[0181] To finely map resistance genes, the segregating materials were further backcrossed with susceptible parents to generate BC2, BC3, or BC4 segregating populations. Recombination events in the region were screened, and the plants were characterized using previously described methods. Additional SNP markers were generated using combinations of 56K SNPs and exon capture data for fine mapping.
[0182] The results of phenotypic analysis of the selected recombination events are shown in Tables 2-4. p-values represent the correlation between genotype and GLS phenotype at a given genetic location. Genetic locations are based on the Corteva B73 genome (version 2). Peak correlations between genotype and phenotype occur at 111.49–111.74 cM on Chr04, and associated markers are shown in bold. Lateral markers listed in Tables 2-4 are indicated by asterisks (…). Table 3 shows the marker trait analysis of GLS QTLs on chromosome 4 based on the 2016 experimental results. Tables 3 and 4 show the marker trait analysis of GLS QTLs on chromosome 4 based on two experiments in 2017.
[0183] Table 2
[0184]
[0185] Table 3
[0186]
[0187] Table 4
[0188]
[0189] To obtain candidate genes within the mapping region, complete whole-genome sequences were generated. RNA-seq reads from NIL were used to refine the gene model of the Chr04 region. Based on the FGENESH gene model, the inbred line A Chr04 region is predicted to contain seven pattern recognition receptor (PRR) genes. These are among the few genes within these regions known to play a role in host defense against pathogens. Due to the large physical size but small genetic size of the Chr04 fine mapping region, further fine mapping would require screening thousands of plants with no guarantee of success. Therefore, genes were tested using proprietary bioinformatics identification methods, rather than attempting to identify specific genes causing the effect via recombination. Table 5 shows the SNPs and locations of GLS resistance markers.
[0190] Table 5
[0191]
[0192] Example 3. Transgenic validation of PRR03 candidate genes
[0193] The efficacy of all seven PRR genes from chromosome 4 regions in a susceptible background was tested. A transgenic construct of PRR03 was constructed, containing a natural promoter (1500 bp) (SEQ ID NO:4), a deduced natural coding sequence (3903 bp) (SEQ ID NO:5), and a natural corn terminator (473 bp) (SEQ ID NO:6). The PRR03 coding sequence (SEQ ID NO:5 encoding SEQ ID NO:7) was shown to have field efficacy in seeds isolated with T1, with an effect size of 2.2–2.4 in gray leaf spot grades. Table 6 shows the enhanced gray leaf spot resistance provided by this PRR03 construct (higher values indicate better resistance).
[0194] Table 6
[0195]
[0196] Example 4. Refinement of PRR03 c sequences
[0197] The transgenic construct used in Example 3 above includes a 3' frameshift sequence encoding a kinase domain. Such a kinase domain is expected to be present in the receptor-like kinase (RLK) PRR gene. However, further sequence analysis showed that the kinase domain is not part of the PRR03c protein sequence disclosed herein as associated with increased GLS resistance. RNA was extracted from a near-isogenic line (NIL) developed to contain a QTL region on chromosome 4. NIL RNA was used to construct Illumina. ® Paired-terminal RNA-seq libraries were used. After sequencing, the reads were aligned to the original PRR03c donor strain reference genome using HISAT2 (with default parameters) (Kim et al. 2019 Nat Biotechnol [Nature Biotechnology] 37:907-15). Transcripts of all candidate genes were then predicted using StringTie (Pertea et al. 2015, Nat Biotechnol [Nature Biotechnology] 33:290-5). The results showed that the protein conferring GLS resistance was the compact PRR03c protein SEQ ID NO:3, which can be expressed from the genomic sequence SEQ ID NO:1 and encoded by the coding sequence SEQ ID NO:2.
Claims
1. A method for increasing resistance to gray leaf spot disease in plant material, the method comprising introducing a heterologous nucleic acid sequence into the genome of the plant material or expressing the heterologous nucleic acid sequence in the plant material, wherein the heterologous nucleic acid (i) encodes a polypeptide having at least 90% amino acid sequence identity with SEQ ID NO:3, or (ii) comprises a sequence having at least 90% nucleotide sequence identity with SEQ ID NO:1 or at least 90% nucleotide sequence identity with SEQ ID NO:2; wherein the plant material expressing the heterologous polynucleotide has increased resistance to gray leaf spot disease (GLS) compared with a control plant that does not express the heterologous polynucleotide.
2. The method of claim 1, wherein the heterologous nucleic acid is inserted at a location on chromosomes 1, 2, 3, 5, 6, 7, 8, 9, or 10 of maize, or alternatively at a location on chromosome 4 of maize other than its natural locus between 90 and 115 cM on chromosome 4.
3. The method of claim 1, wherein the heterologous nucleic acid encodes a polypeptide having a length of less than 500 amino acids.
4. The method of claim 1, wherein the heterologous nucleic acid further comprises a heterologous promoter.
5. The method of claim 1, wherein the method includes introducing the heteropolynucleotide using a double-strand break.
6. The method of claim 5, wherein the method comprises introducing the heteropolynucleotide using TALEN, a wide range of nucleases, zinc finger nucleases, or CRISPR-Cas technology.
7. The method of claim 5, wherein the method comprises introducing the heteropolynucleotide using a Cas endonuclease.
8. The method of any one of claims 1-7, wherein the method further comprises confirming that the heterologous nucleic acid encodes a polypeptide having a length of less than 500 amino acids.
9. A method for introducing the PRR03c allele associated with increased resistance to gray leaf spot (GLS) into a plant, the method comprising: a. Crossing GLS-resistant plants with plants from a second plant line ("second plant") to obtain progeny plants; b. Obtain a sample containing nucleic acid from each of one or more of the said progeny plants; c. Screening the nucleic acids in the sample, wherein the nucleic acids (i) encode a polypeptide having at least 90% amino acid sequence identity with SEQ ID NO:3, or (ii) contain a sequence having at least 90% nucleotide sequence identity with SEQ ID NO:1 or at least 90% nucleotide sequence identity with SEQ ID NO:2; and d. Select one or more progeny plants that have the nucleic acid sequence of c).
10. The method of claim 7, wherein the plant with GLS disease resistance comprises a heterologous nucleic acid, the heterologous nucleic acid (i) encoding a polypeptide having at least 90% amino acid sequence identity with SEQ ID NO:3, or (ii) comprising a sequence having at least 90% nucleotide sequence identity with SEQ ID NO:1 or at least 90% nucleotide sequence identity with SEQ ID NO:2; step c. comprising screening the heterologous nucleic acid in the sample, and step d. comprising selecting one or more progeny plants having the heterologous nucleic acid.
11. The method of claim 8, wherein the heterologous nucleic acid is introduced into the plant with GLS resistance by genome editing or by transgenic modification, or into an ancestor of the plant with GLS resistance.
12. The method of claim 8, further comprising crossing one or more selected progeny plants with the second plant to produce backcross progeny plants.
13. The method of claim 10, wherein the method further comprises Samples containing nucleic acids were obtained from one or more of the backcross progeny plants; The nucleic acids in each sample are screened to (i) encode a polypeptide having at least 90% amino acid sequence identity with SEQ ID NO:3, or (ii) contain a sequence having at least 90% nucleotide sequence identity with SEQ ID NO:1 or at least 90% nucleotide sequence identity with SEQ ID NO:2; and Select one or more backcross progeny plants that have been screened for nucleic acids.
14. The method of claim 11, further comprising: One or more selected backcross progeny plants are crossed with the second plant to produce additional backcross progeny plants; Samples containing nucleic acids were obtained from one or more of the other backcross progeny plants; The nucleic acids in each sample are screened to (i) encode a polypeptide having at least 90% amino acid sequence identity with SEQ ID NO:3, or (ii) contain a sequence having at least 90% nucleotide sequence identity with SEQ ID NO:1 or at least 90% nucleotide sequence identity with SEQ ID NO:2; and Select one or more additional backcross progeny plants that have the screened nucleic acids.
15. The method of claim 12, further comprising one or more of the following repeated steps: The most recent backcross progeny plant is crossed with the second plant to produce further backcross progeny plants; Obtain samples containing nucleic acids from one or more further backcross progeny plants; The nucleic acids in each sample are screened to (i) encode a polypeptide having at least 90% amino acid sequence identity with SEQ ID NO:3, or (ii) contain a sequence having at least 90% nucleotide sequence identity with SEQ ID NO:1 or at least 90% nucleotide sequence identity with SEQ ID NO:2; and Select one or more further backcross progeny plants that have the screened nucleic acids.
16. The method of any one of claims 8-13, wherein the method further comprises identifying one or more of the progeny plants, backcross progeny plants, or other backcross progeny plants that encode a polypeptide having a length of less than 500 amino acids.
Citation Information
Patent Citations
CRISPR-CAS systems for genome editing
US10934536B2
Expression Enhancing Intron Sequences
US20090144863A1
Epsps mutants
US20130019349A1
Genome modification using guide polynucleotide / CAS endonuclease systems and methods of use
US20150059010A1
Plant genome modification using guide RNA / CAS endonuclease systems and methods of use
US20150082478A1