InDel sites, molecular markers, primers and their applications related to drought resistance in maize

By developing the InDel site at 178050bp on maize chromosome 2 and the corresponding molecular marker, the problem of lacking drought markers in existing technologies for maize has been solved, enabling efficient identification and breeding of new drought-resistant maize varieties.

CN117144047BActive Publication Date: 2025-12-02HENAN ACAD OF AGRI SCI INST OF GRAIN CROPS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311216969.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-20
Publication Date
2025-12-02
Estimated Expiration
2043-09-20

AI Technical Summary

Technical Problem

The lack of InDel markers associated with drought in existing technologies makes it difficult to effectively identify and breed new drought-resistant maize varieties.

Method used

The InDel locus located at 178050 bp on maize chromosome 2 was developed, and corresponding molecular markers and primers were designed. The drought resistance of maize plants was identified by PCR amplification and sequencing analysis.

Benefits of technology

It provides a high-throughput, high-efficiency molecular marker method that can accurately identify drought-resistant and non-drought-resistant segregating populations of maize and locate drought-responsive genes, providing a basis for the screening and breeding of drought-resistant maize materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117144047B_ABST
    Figure CN117144047B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of molecular marker technology, specifically relating to the InDel locus, molecular markers, primers, and their applications related to drought resistance in maize. The locus is located at 178050 bp on maize chromosome 2; maize plants without a 24 bp deletion in the allele at this locus are considered drought-intolerant. This invention also discloses molecular markers containing this InDel locus, amplification primers for the molecular markers, a kit containing the amplification primers, and their application in identifying drought sensitivity in maize. The molecular markers described in this invention can identify drought-resistant and drought-intolerant maize lines, and can be used to locate drought-responsive genes and screen drought-resistant maize materials, playing an important role in the study of maize drought resistance and molecularly assisted selection breeding for drought resistance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular marker technology, specifically relating to InDel sites, molecular markers, primers and their applications related to drought resistance in maize. Background Technology

[0002] Maize is a major food crop, feed crop, and important energy crop, especially as an energy crop, where it has attracted much attention. However, many regions currently face problems such as severe water shortages, high evaporation rates, rapid water loss, large precipitation variability, and interannual instability. Analysis shows that drought is the main factor causing fluctuations in maize yields in my country; at the same time, the impact of extreme weather in recent years has exacerbated the problem of reduced maize yields due to water scarcity. Therefore, research on maize drought resistance is of significant practical importance, and how to improve maize drought resistance has gradually become a research hotspot both domestically and internationally. To avoid further damage to maize yields caused by drought, the efficient discovery and utilization of superior genes in germplasm resources and the breeding of maize varieties with high yield potential and environmental stability have become the main goals of breeders.

[0003] In recent years, with the rapid development of molecular biology, molecular marker technology has provided a new approach for drought resistance research. Scholars both domestically and internationally have located numerous QTL sites related to drought resistance using various materials and applied them to the breeding of drought-resistant maize varieties using marker-assisted selection (MAG). Simultaneously, the molecular mechanisms of drought resistance in maize have been further elucidated. The identification of drought-resistant QTLs in maize mainly focuses on several aspects, including the average day interval between male and female flowering (ASI), greenness retention, ear setting rate, ABA, proline, free radicals, and grain yield. Therefore, research on maize drought resistance has significant practical implications, and improving maize drought resistance has gradually become a research hotspot both domestically and internationally.

[0004] Indel markers primarily refer to base insertions or deletions between two parental individuals. Primers are designed using these differentially expressed sites for PCR amplification to identify these molecular markers. Indels are widely distributed in the genome, with high density, numerous individuals, good polymorphism, strong stability, and easy differentiation, making them suitable for multiple species. Indel markers are currently being applied in various fields such as population genetic analysis, germplasm analysis, molecular-assisted breeding, and genetics. The continuous development of Indel markers will facilitate the further utilization of superior genes. Based on the magnitude of the difference, these markers can rapidly detect large numbers of samples and accurately determine their genotypes, making them highly suitable for marker-assisted selection breeding.

[0005] However, in actual production, there is a lack of drought-related SNPs or InDel markers that can be truly applied to drought-resistant molecular breeding of maize. Therefore, in-depth research on maize genes that respond to drought stress and the development of closely linked molecular markers are of great significance for elucidating the mechanism of maize response to drought stress and for breeding new drought-resistant, high-yielding and stable maize varieties. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the deficiency of the lack of InDel markers related to drought in maize in the prior art, and to provide InDel sites, molecular markers, primers and their applications related to drought resistance in maize.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution.

[0008] The first aspect of the present invention provides an InDel locus associated with drought resistance in maize, the locus being located at 178050 bp on maize chromosome 2, wherein maize plants without the 24 bp base deletion shown in SEQ ID NO:1 at this locus are drought-resistant maize plants.

[0009] A second aspect of the invention provides molecular markers developed incorporating the InDel site.

[0010] In some embodiments of the present invention, when the nucleotide sequence of the molecular marker is as shown in SEQ ID NO:2, corn without the 24bp base deletion shown in SEQ ID NO:1 is non-drought-resistant corn; when the nucleotide sequence of the molecular marker is as shown in SEQ ID NO:3, corn with the 24bp base deletion shown in SEQ ID NO:1 is drought-resistant corn.

[0011] A third aspect of the invention provides amplification primers for the aforementioned molecular markers.

[0012] In some embodiments of the present invention, the nucleotide sequences of the amplification primers are shown in SEQ ID NO:4 and SEQ ID NO:5.

[0013] A fourth aspect of the present invention provides a kit comprising the amplification primers.

[0014] A fifth aspect of the invention provides the use of the InDel site, the molecular marker, the amplification primers, or the kit in identifying drought resistance in maize.

[0015] The sixth aspect of this invention provides a method for identifying drought resistance in maize, comprising the following steps:

[0016] S1, Maize genomic DNA to be identified was extracted using the CTAB method;

[0017] S2, using the maize genomic DNA as a template, perform PCR amplification using the amplification primers to obtain the amplification product;

[0018] S3. The amplified product is sequenced and analyzed, and the drought resistance of the maize plant is determined based on the results of the sequencing analysis.

[0019] In some embodiments of the present invention, in S3, the criterion for judgment is: when the 113th position of the 5' end of the amplification product has a 24bp base deletion as shown in SEQ ID NO:1, it is judged to be a drought-resistant maize plant; when the 113th position of the 5' end of the amplification product has a 24bp base insertion as shown in SEQ ID NO:1, it is judged to be a non-drought-resistant maize plant.

[0020] The seventh aspect of the present invention provides the application of the InDel site, the molecular marker, the amplification primer, or the kit in the cultivation of drought-resistant maize plants.

[0021] Compared with existing technologies, the present invention has the following beneficial effects: The molecular markers described in this invention can identify drought-resistant and non-drought-resistant segregating populations of maize, as well as different drought-resistant lines in natural populations. They can also be used to locate drought-responsive genes, laying the foundation for cloning candidate genes at this locus. Furthermore, they can provide reference and basis for screening drought-resistant maize materials, studying maize drought resistance, and molecularly assisted selection breeding for drought resistance. In addition, the InDel marker Indel15 described in this invention is a co-dominant marker, possessing advantages such as high throughput, high efficiency, stable amplification, and convenient detection. Attached Figure Description

[0022] Figure 1 This is a comparison diagram of wilted plants and normal plants in the BC1 segregating population.

[0023] Figure 2 The image shows the screening results of SSR primers for whole genome polymorphism; lanes 1-4 represent umc1227 (104bp); lanes 5-8 represent umc1552 (157bp); lanes 21-24 represent SSR28 (326bp); and lanes 25-28 represent SSR29 (282bp).

[0024] Figure 3 This is a diagram showing the preliminary localization results of the wi6 gene.

[0025] Figure 4 This is a flowchart illustrating the specific process of BSA bioinformatics analysis for resequencing.

[0026] Figure 5 This is a base distribution diagram of Clean Reads from each sequencing pool; among them, Figure 5 A represents the base distribution diagram of Clean Reads for wilted individual plants. Figure 5B represents the base distribution diagram of a normal single Clean Reads plant.

[0027] Figure 6 This is a statistical graph showing the distribution of inserted segments; where, Figure 6 A represents the statistical distribution of inserted fragments in wilted individual plants. Figure 6 B represents the statistical distribution of normal single-plant insertion fragments.

[0028] Figure 7 This is a statistical graph showing the distribution of sequencing depth; among which, Figure 7 A represents the statistical distribution of sequencing depth in wilted single plants. Figure 7 B represents the statistical graph of normal single-strain sequencing depth distribution.

[0029] Figure 8 This is a map showing the chromosome coverage depth distribution of the sample; where, Figure 8 A represents the chromosome coverage depth distribution map of wilted individual plants. Figure 8 B represents the chromosome coverage depth distribution map of a normal single plant.

[0030] Figure 9 This is a quality distribution diagram of SNPs; where, Figure 9 A represents the cumulative number of SNP reads supported by the graph; Figure 9 B is the cumulative distance graph between adjacent SNPs.

[0031] Figure 10 Venn plot for inter-sample SNP statistics.

[0032] Figure 11 This is a Small InDel statistical Venn plot for the samples.

[0033] Figure 12 This is a distribution map of the ED association values ​​of SNP sites on chromosomes.

[0034] Figure 13 This is a distribution map of ED association values ​​at the InDel site on the chromosome.

[0035] Figure 14 Gene GO annotation cluster diagrams for candidate regions.

[0036] Figure 15 This is a pathway distribution map of genes within the candidate region.

[0037] Figure 16 A COG annotation classification map for genes in SNPs within the candidate region.

[0038] Figure 17 This is a fine-map of the wi6 gene.

[0039] Figure 18 To develop a marker-localized result map.

[0040] Figure 19 Image showing the sequence alignment results of the Indel15 marker in the parent. Detailed Implementation

[0041] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments, but this should not be construed as limiting the invention. Unless otherwise specified, the technical means used in the following embodiments are conventional means well known to those skilled in the art, and the materials, reagents, etc. used in the following embodiments are commercially available unless otherwise specified.

[0042] 1. Experimental Materials

[0043] The mutant used in this invention is a mutant derived from the MU9 transposon insertion synodyne 31, tentatively named wilted6 (wi6). Using the wi6 mutant as the male parent and the commonly used maize backbone lines B73 and Mo17 as the female parents, the mutant was planted at the Modern Agricultural Science and Technology Experimental Demonstration Base of the Henan Academy of Agricultural Sciences in the summer of 2016, yielding F1 seeds. In the winter of 2016, the F1 plants were planted in Sanya, Hainan, for self-pollination and backcrossing to obtain F2 and BC1 genetic analysis populations, respectively. In 2016, the F2 and BC1 segregating populations were planted at the Modern Agricultural Science and Technology Experimental Demonstration Base of the Henan Academy of Agricultural Sciences for phenotypic identification and subsequent experiments.

[0044] Example 1: Phenotypic Investigation of F2 and BC1 Recombinant Progeny

[0045] (1) Inheritance patterns of F2 and BC1

[0046] During the corn seedling stage (around the 7-leaf stage), the number of normal and wilted plants in the parent (wi6 mutant), F1, F2 and BC1 populations was counted, photographed and the segregation ratio was recorded.

[0047] The results showed that, except for the wi6 mutant which exhibited wilting, all other plants were normal. Small populations of approximately 300 plants each were planted in the BC1 and F2 segregating populations. When phenotypic expression was more pronounced, the normal and wilted plants in the segregating populations were investigated. The results are shown in Table 1. Figure 1 As shown, the ratio of normal to wilted plants in F2 was approximately 3:1, while in BC1 it was approximately 1:1. The segregation ratio in the offspring conformed to Mendelian segregation, indicating that this trait is controlled by a single recessive nuclear gene. A large-scale mapping population was constructed based on B73 and the wi6 mutant, providing a material basis for map-based cloning of this mutant gene.

[0048] Table 1. Analysis of the genetic patterns of the wi6 mutant.

[0049]

[0050] Note: "-" indicates that the item is not available.

[0051] (2) Preliminary localization of the wi6 gene

[0052] Using the near-isogenic pooling method, a BC1 mapping population of approximately 400 plants was constructed using the drought-sensitive mutants wi6 and B73. After field phenotypic surveys, DNA was extracted from 30 normal plants and 30 wilted plants, and then mixed in equal amounts to construct normal and mutant pools. Polymorphism analysis was performed on wi6, B73, the normal pool, and the mutant pool using over 800 pairs of SSR primers covering the entire maize genome. Primers showing polymorphism among parents and between extreme pools were linked to the target gene. The results of polymorphism screening of over 800 pairs of SSR primers are as follows: Figure 2 As shown, four pairs of polymorphic markers linked to the target gene were found on chromosome 2: umc1552 (region 2.02), umc1227 (region 2.02), SSR29 (region 2.00), and SSR28 (region 2.00). The physical locations of the four markers on the chromosome are umc1552 (4832136-5585579bp), umc1227 (4407802-4732458), and SSR29 (1258875-1257158). SSR28 (1256643-1256988) were used as polymorphic markers. These four markers were used to amplify over 1000 BC1 populations. The amplification results showed a decreasing trend in the number of exchanged single plants from umc1552, umc1227, SSR29, and SSR28. SSR29 and SSR28 were physically close, with the same number of exchanged single plants, but all were contained within umc1552 and umc1227. This indicates that these four markers are located on one side of the target gene and extend towards the telomere, approximately 1.2 Mb from the telomere (e.g., SSR1552, SSR28, SSR29, SSR28). Figure 3 (As shown).

[0053] Example 2: Mixed-pool resequencing and data analysis

[0054] 1. Field sampling

[0055] During the corn seedling stage (around the 7-leaf stage), the parent plants and the F2 and BC1 populations were numbered. When sampling, each individual plant sample was placed in a corresponding numbered 5mL centrifuge tube, placed in liquid nitrogen, and brought back to the laboratory and placed in a -80℃ freezer for later DNA extraction.

[0056] 2. DNA extraction from maize leaves and detection of DNA concentration in maize leaves

[0057] Based on field survey data of maize drought resistance experiments, DNA was extracted from leaves of 30 extremely normal and wilted individual plants from the BC1 population and mixed to form two DNA extreme pools. Whole-genome DNA was extracted using the CTAB method; the CTAB formulation is shown in Table 2.

[0058] Table 21.67% CTAB solution formulation

[0059]

[0060]

[0061] Note: "-" indicates that the item is not available.

[0062] The concentration and quality of DNA from individual samples of the BC1 isolate population were determined using a Nanodrop concentration meter. Table 3 shows that the DNA concentration of individual samples ranged from 122.56 to 454.34 ng / μL, with an OD of... 260 / 280 The range is between 1.80 and 1.99, OD 260 / 230 The range is between 1.79 and 2.13. This indicates that the extracted DNA is of good quality and can guarantee the success of subsequent library construction experiments.

[0063] Table 3. DNA concentration and quality detection of extreme individuals in the BC1 segregating population.

[0064]

[0065]

[0066]

[0067] 3. DNA Mixed Pool Resequencing Experimental Procedure

[0068] After the sample DNA passed the initial testing, the DNA was randomly fragmented into 350bp fragments using ultrasonic disruption. These fragments underwent end repair, 3' end A addition, sequencing adapter addition, purification, and PCR amplification to construct the sequencing library. After passing quality control, the library was sequenced using Illumina HiSeq, following the specific procedure described below. Figure 4 As shown.

[0069] 4. Raw sequencing data

[0070] The raw image data files obtained from the Illunima HiSeq high-throughput sequencing platform are converted into raw sequencing reads (Raw Data or Raw Reads) through base calling analysis. These reads contain sequence information and corresponding sequencing quality information. The sequencing error rate for each base is obtained through a sequencing Phred score, which is calculated during base calling using a model that predicts the probability of base discrimination errors. The correspondence is shown in Table 4.

[0071] Table 4. Concise Correspondence between Illumina Casava Base Recognition and Phred Score

[0072] Phred score Base sequencing error rate (e) Base correct identification rate Q-Score 10 1 / 10 90% Q10 20 1 / 100 99% Q20 30 1 / 1000 99.9% Q30 40 1 / 10000 99.99% Q40

[0073] (1) Filtering of raw low-quality data

[0074] To ensure the quality of information analysis, the raw sequencing sequences are filtered to obtain Clean Reads, which are then used for subsequent analysis. The specific filtering steps are as follows:

[0075] 1) Remove reads with connectors;

[0076] 2) If the proportion of N (the specific base type could not be determined) in a read is greater than 10%, then filter out that pair end read;

[0077] 3) Remove low-quality reads (the number of bases with a quality value of Q≤10 accounts for more than 50% of the entire read).

[0078] The data filtering statistics are shown in Table 5. Wilted individual plants in the BC1 segregating population were named R01, and normal individual plants were named R02. The two pooled sequencing results generated a total of 92 Gbp of clean data, with the proportion of clean reads after filtering approaching 90%. This indicates that the library quality is good and meets sequencing quality requirements, making it suitable for subsequent experiments.

[0079] Table 5 Statistics on Resequencing Data Filtering

[0080]

[0081] (2) Base type distribution

[0082] like Figure 5 A and Figure 5As shown in Figure B, a base distribution map is obtained by plotting the base positions of the filtered Reads on the x-axis and the proportion of ACGTN bases at each position on the y-axis. Figure 5 It can be seen that the distribution of A and T, C and G bases is the same in each sequencing pool, and AT and CG bases are basically not separated, which is consistent with the expected base distribution. Moreover, the curve is relatively flat, indicating that the sequencing results are normal.

[0083] 5. Comparison and quality control of raw data

[0084] (1) Comparison information with the reference genome

[0085] The resequencing reads obtained need to be repositioned onto the reference genome before subsequent variant analysis can be performed. Using Bwa software, the positions of clean reads on the reference genome were located by alignment, and information such as sequencing depth and genome coverage for each sample was statistically analyzed, along with variant detection. The alignment results for the samples are shown in Table 6. Furthermore, the average alignment efficiency for all samples was above 80%, indicating that the sample sequencing was normal.

[0086] Table 6. Comparison Results Statistics

[0087]

[0088]

[0089] Note: "-" indicates that the item is not available.

[0090] (2) Distribution of inserted fragments

[0091] The distribution of insert fragments can assess the library construction process and determine whether the fragment sizes match expectations. By detecting the start and end positions of the paired-end sequences on the reference genome, the actual size of the sequencing fragments obtained after DNA fragmentation in the sample can be obtained, i.e., the insert fragment size. The distribution of insert fragment sizes generally follows a normal distribution with a single peak. Insert fragment size distribution plots can show the length distribution of insert fragments in each sample.

[0092] Because the size of the inserted fragments was designed to be around 350 when building the library. For example... Figure 6 A and Figure 6 As shown in Figure B, the inserted fragments are mainly concentrated between 250bp and 400bp, consistent with the size of the fragments recovered during library construction. Furthermore, the length distribution of the inserted fragments follows a normal distribution, indicating that there were no abnormalities in the sequencing data library construction.

[0093] (3) Sequencing depth distribution

[0094] After readings locate the reference genome, the base coverage on the reference genome can be analyzed. The base coverage depth distribution curve and coverage distribution curve of the sample are shown below. Figure 7 A and Figure 7 As shown in B. The average coverage depth of each sample and the corresponding genome coverage ratio at each depth are shown in Table 7. The average genome coverage depth is approximately 21.50X, and the genome coverage is approximately 98.91% (at least 1X coverage).

[0095] Table 7. Statistics on sample coverage depth and coverage ratio

[0096]

[0097] Plot the data based on the coverage depth of each point on the sample chromosome, such as... Figure 8 A and Figure 8 As shown in Figure B, the genome is relatively evenly covered, indicating good sequencing randomness. Areas with uneven depth in the figure may be due to repetitive sequences or PCR bias.

[0098] 6. SNP Detection and Annotation

[0099] SNP detection primarily utilizes the GATK software toolkit. Based on the Clean Reads' localization results in the reference genome, preprocessing is performed using Picard (mark duplicates), GATK (local realignment), and base recalibration to ensure the accuracy of the detected SNPs. Then, GATK is used to detect and filter single nucleotide polymorphisms (SNPs) to obtain the final SNP site set.

[0100] (1) Detection of SNPs between the sample and the reference genome

[0101] SNP quality distribution as follows Figure 9 A and Figure 9 As shown in B. For maize, a SNP site on a homologous chromosome containing only the same base is called a homozygous SNP site; a SNP site on a homologous chromosome containing different types of bases is called a heterozygous SNP site. The more homozygous SNPs there are, the greater the difference between the sample and the reference genome; the more heterozygous SNPs there are, the higher the degree of heterozygosity of the sample.

[0102] (2) Detection of SNPs between samples

[0103] Based on the alignment results between the sample and the reference genome, all differentially identified variant sites between the samples are summarized. For example... Figure 10 As shown, 7,857,636 common SNP differential sites were found among the samples.

[0104] (3) SNP results annotation

[0105] Using SnpEff software, based on the location of the variant site on the reference genome and the gene location information on the reference genome, the region in which the variant site occurred (intergenic region, gene region, or CDS region, etc.) and the impact of the variant (synonymous and non-synonymous mutations, etc.) can be obtained. Specific statistical results of SNP annotation between the two samples are shown in Table 8.

[0106] Table 8. Statistics of SNP annotation results

[0107]

[0108]

[0109] 7. Small InDel Detection and Annotation

[0110] (1) Small InDel detection

[0111] Based on the location of the clean reads of the sample on the reference genome, the presence of small in-Del (1-5 bp) insertions and deletions between the sample and the reference genome was detected. Insertions and deletions in the samples were detected using GATK. Based on the Small InDel detection results of the sample and the reference genome, differentially expressed variant sites were extracted between the samples; these were identified as Small InDel variant sites between the samples. The statistical results of Small InDel between samples are shown below. Figure 11 As shown, there are a total of 1,312,429 Small InDel differential sites.

[0112] (2) Small InDel's annotation

[0113] Based on the location information of Small InDel sites on the reference genome obtained from sample detection, and by comparing them with gene and CDS locations in the reference genome, annotations can be made to determine whether the InDel sites occur in intergenic regions, gene regions, or CDS regions, and whether they are frameshift mutations. Small InDel annotation is performed using SnpEff software. Specific annotation results are shown in Table 9.

[0114] Table 9 Statistics of InDel annotation results

[0115]

[0116]

[0117] 8. Association Analysis (SNP)

[0118] (1) High-quality SNP screening

[0119] Before association analysis, SNPs were first filtered using the following criteria: first, SNPs with multiple genotypes were filtered out; second, SNPs with read support less than 4 were filtered out; and third, SNPs with identical genotypes across pools were filtered out. This resulted in 8,239,626 high-quality, reliable SNPs (as shown in Table 10). During the analysis, the depth of each base in different pools was calculated using SNPs with genotype differences between pools, and the ED value for each site was calculated. The fifth power of the original ED was used as the association value to eliminate background noise. The DISTANCE method was then used to fit the ED values. The association value distribution is shown below. Figure 12 As shown in Table 11, the original ED values ​​were exponentially multiplied. The median + 3SD of the fitted values ​​for all sites was taken as the association threshold for analysis, which was calculated to be 0.06. Based on the association threshold, a total of one region was obtained, with a total length of 11.46 Mb, containing 2,250 genes, of which 242 were non-synonymous mutation SNP sites (as shown in Table 11). The calculation formula of the ED method is shown below. The larger the ED value, the greater the difference of the marker between the two pools.

[0120]

[0121] Where: Amut is the frequency of A base in the mutant pool, Awt is the frequency of A base in the wild-type pool; Cmut is the frequency of C base in the mutant pool, Cwt is the frequency of C base in the wild-type pool; Gmut is the frequency of G base in the mutant pool, Gwt is the frequency of G base in the wild-type pool; Tmut is the frequency of T base in the mutant pool, Twt is the frequency of T base in the wild-type pool.

[0122] Table 10 SNP Filtering Statistics

[0123]

[0124] Table 11 Statistical Table of Related Areas

[0125]

[0126] Note: "-" indicates that the item is not available.

[0127] (2) Association Analysis (InDel)

[0128] Before performing association analysis using InDel, the InDel sites were first filtered using the same criteria as for SNP analysis, resulting in 1,228,371 high-quality, reliable InDel sites (as shown in Table 12). The analysis was performed using the same analytical method (ED method) as for SNP association analysis, and the association value distribution is as follows: Figure 13 As shown in Table 13, the median+3SD of the fitted values ​​of all sites was taken as the association threshold for analysis, which was calculated to be 0.07. According to the association threshold, a total of 1 region was obtained, with a total length of 12.43Mb, containing 2,434 genes, of which 45 genes were at the InDel frameshift mutation site.

[0129] Table 12 InDel Filtering Statistics

[0130]

[0131] Table 13 Statistical Table of Related Areas

[0132]

[0133] Note: "-" indicates that the item is not available.

[0134] (3) Candidate region screening

[0135] The intersection of the results for the associated regions corresponding to SNPs and InDels was taken. The intersection is shown in Table 14. A total of 1 region was obtained, located on chromosome 2, with a total length of 11.46 Mb and containing 2,250 genes.

[0136] Table 14 Statistical Table of Related Areas

[0137]

[0138] Note: "-" indicates that the item is not available.

[0139] 9. Functional annotations for candidate regions

[0140] (1) Gene annotation within candidate regions

[0141] BLAST software was used to perform deep annotation of coding genes within candidate regions using multiple databases (NR, Swiss-Prot, GO, KEGG, and COG). Detailed annotation enabled rapid screening of candidate genes. A total of 2,223 genes were annotated within the candidate regions, including 242 genes with non-synonymous mutations in the pooled regions and 46 frameshift mutations. The annotation results are shown in Table 15.

[0142] Table 15. Statistical analysis of gene function annotation results in SNPs and InDels within candidate regions.

[0143] database Number of genes Number of nonsynonymous mutant genes Number of frameshift mutation genes NR 2221 242 46 NT 2223 242 46 trEMBL 2221 242 46 SwissProt 1378 164 30 GO 1812 204 36 KEGG 994 103 22 COG 874 95 15 Total 2223 242 46

[0144] (2) GO enrichment analysis of genes within candidate regions

[0145] The GO classification statistics of the corresponding genes in the candidate regions are as follows: Figure 14 As shown, in cellular component types, it is mainly enriched in processes such as cells, membranes, and organelles; in molecular function types, it is mainly enriched in processes such as catalytic activity and binding; and in biological process types, it is mainly enriched in metabolic processes, cellular processes, and single organism processes.

[0146] (3) KEGG enrichment analysis of genes in candidate regions

[0147] The KEGG annotation results of the corresponding genes in the selected region are categorized according to pathway type. For example... Figure 15 As shown, pathway analysis of differentially expressed genes revealed that they were classified into five major categories: cellular processes, environmental information processing, genetic information processing, metabolism, and organic systems. Cellular processes were mainly enriched in pathways such as endocytosis and peroxisomes; environmental information processing was mainly enriched in pathways such as plant hormone single transport proteins; genetic information processing was mainly enriched in pathways such as homologous recombination, basic transcription factors, and RNA transport; metabolism was mainly enriched in pathways such as arginine and proline metabolism, oxidative phosphorylation, carbon metabolism, and metabolic processes; and organic systems were mainly enriched in pathways such as plant-pathogen interactions.

[0148] (4) Statistical analysis of COG classification of genes in candidate regions

[0149] Statistical results of COG classification of genes within the associated region are as follows: Figure 16 As shown, these genes are mainly involved in transcription, replication, recombination and repair, general function prediction, and signal transduction mechanisms.

[0150] Example 3: Validation and Marker Development of Candidate Sites

[0151] Using the maize backbone line B73 and the drought-sensitive mutant wi6, polymorphic SSR and Indel markers distributed on chromosome 2 were screened. Using the screened polymorphic markers, normal and wilted segregating plants in the F2 population were selected for localization, and exchange plants at candidate loci were screened.

[0152] Four pairs of markers with good polymorphism selected from chromosome 2 were amplified using over 1000 individuals from the BC1 population. The simplified sequence results were validated, and the candidate region was located within 1.2 Mb between SSR28 and telomeres, consistent with QTL-Seq and QTG-Seq results. SSR28 was used to amplify over 6000 BC1 individuals, and exchanged individuals were screened, resulting in 124 individuals with right-side exchange. Based on the B73 genome sequence between SSR28 and telomeres, new SSR and Indel markers were sought. Simultaneously, new SSR markers were developed using SSRHunter 1.3 software. The obtained SSR markers were compared on NCBI, and primers were designed using Primer 5.0 for single-copy markers. Polymorphism screening was then performed between parents. Of the 55 pairs of SSR and Indel primers developed, four pairs showed polymorphism between parents and between parental lines: SSR20, SSR18, umc2246, and Indel15. Their physical locations on the second chromosome were SSR20 (958884-959148), SSR18 (957605-957812), umc2246 (952153-952966), and Indel15 (177937-178229), respectively. The number of exchanged individuals decreased sequentially from SSR28 (124 individuals) to Indel15 (24 individuals), and the Indel15 exchanged individuals were included among the exchanged individuals of other markers. Therefore, as... Figure 17 As shown, the wi6 gene was finely mapped between Indel15 and telomeres, with a physical distance of 180 kb.

[0153] Using drought-resistant and drought-sensitive segregating plants from the F2 and BC1 populations, such as Figure 18 As shown in Group A, 10 drought-resistant exchange plants were selected respectively, and... Figure 18 The 10 drought-sensitive exchanged single plants shown in Group B were analyzed. The Indel15 marker was found to be a drought-sensitive genotype in the drought-sensitive plants and a drought-resistant genotype in the drought-resistant plants, and was tightly linked to drought in the F2 population. Subsequently, using Indel15-marked primers (as shown in Table 16, the nucleotide sequence of Indel15F is shown in SEQ ID NO:4, and the nucleotide sequence of Indel15R is shown in SEQ ID NO:5), a 293 bp fragment was amplified. DNA fragments from the drought-sensitive mutant wi6 (amplified sequence shown in SEQ ID NO:2) and the amplified inbred line B73 (amplified sequence shown in SEQ ID NO:3) were sequenced, and the results are as follows: Figure 19 As shown, it was found that the mutant wi6 had 24 bp more bases than the B73 sequence from position 113 at the 5′ end of the amplified product. The 24 bp base sequence is shown in SEQ ID NO:1.

[0154] Table 16 Indel15 labeled primers

[0155]

[0156]

[0157] Example 4: Application of Closely Linked Markers

[0158] Using the selected SSR and Indel markers with good polymorphism, the initially identified exchange plants were further screened. The linkage degree of the selected closely linked markers was verified in different normal and wilting natural populations. The study found that the Indel15 marker was closely linked to drought in the BC1 segregating population. To verify the correlation between this marker and drought resistance in different maize materials, as shown in Table 17, analysis revealed that the genotype of this marker was consistent with the drought-resistant phenotype in different materials, which can be used for marker-assisted selection breeding to improve the drought resistance of different materials.

[0159] Table 17 Genotypic analysis of the Indel15 marker in different inbred lines Note: Drought-resistant and drought-sensitive inbred lines are derived from phenotypic identification and relative expression levels that directly reflect drought resistance. Genotype 1 represents non-drought-resistant plants; genotype 2 represents drought-resistant plants. All materials in the table are inbred lines whose drought resistance has been identified.

[0160] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0161] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. The application of amplification primers containing a molecular marker of the InDel site associated with drought resistance in maize in identifying drought resistance, characterized in that, When the nucleotide sequence of the molecular marker is as shown in SEQ ID NO: 2, corn without the 24bp base deletion shown in SEQ ID NO: 1 is non-drought-resistant corn; when the nucleotide sequence of the molecular marker is as shown in SEQ ID NO: 3, corn with the 24bp base deletion shown in SEQ ID NO: 1 is drought-resistant corn. The nucleotide sequences of the amplification primers are shown in SEQ ID NO: 4 and SEQ ID NO:

5.

2. The application of a kit containing the amplification primers of claim 1 in identifying drought resistance in maize, characterized in that, Includes the following steps: S1, Maize genomic DNA to be identified was extracted using the CTAB method; S2, using the maize genomic DNA as a template, perform PCR amplification using the amplification primers described in claim 1 to obtain the amplification product; S3, perform sequencing analysis on the amplified products, and determine the drought resistance of the maize plants based on the sequencing analysis results; In S3, the criteria for judgment are as follows: when the 113th position of the 5' end of the amplification product has a 24bp base deletion as shown in SEQ ID NO: 1, it is judged to be a drought-resistant maize plant; when the 113th position of the 5' end of the amplification product has a 24bp base insertion as shown in SEQ ID NO: 1, it is judged to be a non-drought-resistant maize plant.

3. The application of the amplification primers as described in claim 1 or the kit as described in claim 2 in cultivating drought-resistant maize plants, characterized in that... When the amplification product of the amplification primer has a 24bp base deletion at position 113 of the 5′ end as shown in SEQ ID NO: 1, it is identified as a drought-resistant maize plant; when the amplification product of the amplification primer has a 24bp base insertion at position 113 of the 5′ end as shown in SEQ ID NO: 1, it is identified as a non-drought-resistant maize plant.