Application of Zm00001d028188 gene in maize breeding
By digging and utilizing the corn functional gene Zm00001d028188, its expression is regulated to increase the weight of corn grains by 100 grains, the problem of difficulty in effectively increasing the weight of 100 grains in the existing technology is solved, and technical support for corn breeding is achieved.
Patent Information
- Application Number
- CN202411883935.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The prior art is difficult to effectively increase the weight of corn grains by 100 grains, affecting yield.
By digging out the functional gene Zm0001d028188, which is associated with corn grains, and by regulating the expression of this gene, it is possible to assist in breeding of corn grains, molecular markers of corn grains, by regulating the amount of expression of this gene.
This gene can explain 6.3% of the phenotype variants of 100 grains, and the GG haplotype significantly increases the weight of 100 grains, providing theoretical basis and technical support for corn breeding.
Smart Images

Figure CN119662709B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural biotechnology, and particularly relates to the application of the Zm00001d028188 gene in maize breeding. Background Art
[0002] Maize (Zea mays L.) is one of the most important crops globally and plays an important role in food, feed, and industrial product sources. The 100-kernel weight (HKW) of maize is an important component affecting yield, and increasing yield is an important goal in maize breeding (Z. Zhang, Z. Liu, Y. Hu, W. Li, Z. Fu, D. Ding, H. Li, M. Qiao, J. Tang, QTL analysis of Kernel-related traits in maize using an immortalized F2 population, PLoS One, 9 (2014) e89645).
[0003] As a typical quantitative trait, HKW is regulated by multiple genes and environmental factors (Y. Xiao, H. Liu, L. Wu, M. Warburton, J. Yan, Genome-wide Association Studies in Maize: Praise and Stargaze, Mol Plant, 10 (2017) 359-374; Z. Zhang, Z. Liu, Y. Hu, W. Li, Z. Fu, D. Ding, H. Li, M. Qiao, J. Tang, QTL analysis of Kernel-related traits in maize using an immortalized F2 population, PLoS One, 9 (2014) e89645). In-depth understanding of the natural genetic variation regulating HKW and its molecular mechanism is of great significance for increasing maize yield. In summary, increasing the 100-kernel weight is a key goal in plant breeding and biotechnology-assisted improvement. Therefore, mining functional genes closely associated with the 100-kernel weight of maize can provide technical support for molecular marker-assisted selection of high-quality maize. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides the application of the Zm00001d028188 gene in maize breeding, mines the functional gene Zm00001d028188 associated with the 100-kernel weight of maize kernels, and provides a theoretical basis for molecular marker-assisted selection to regulate the 100-kernel weight of maize kernels.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0006] Application of Zm00001d028188 gene in molecular marker-assisted breeding for regulating 100-kernel weight of maize kernels. The nucleotide sequence of the Zm00001d028188 gene is shown in SEQ ID NO.1, and the amino acid sequence of the protein encoded by this gene is shown in SEQ ID NO.2.
[0007] Preferably, regulating the 100-kernel weight of maize kernels is to increase the 100-kernel weight of maize kernels by adjusting the overexpression of the Zm00001d028188 gene in maize.
[0008] Preferably, regulating the 100-kernel weight of maize kernels is to reduce the 100-kernel weight of maize kernels by reducing the expression level of the Zm00001d028188 gene in maize.
[0009] Preferably, the method for reducing the expression level of the Zm00001d028188 gene in maize is any one or more of silencing, knocking out, or knocking down.
[0010] The present invention provides an application of the Zm00001d028188 gene in maize breeding. Compared with the prior art, the advantages are as follows:
[0011] In the present invention, a multi-parent population (MPP) of 813 F2:7 RILs containing 5 recombinant inbred line (RIL) subgroups was constructed by crossing the common parent Ye107 (male parent) with 5 donor parents (female parents) that were significantly different in kernel size and weight. By using GWAS analysis and genetic linkage analysis, SNP_25750352 significantly related to the 100-kernel weight of maize located on chromosome 1 was co-localized, and then the functional gene Zm00001d028188 regulating the 100-kernel weight was mined. This gene can explain 6.3% of the phenotypic variation in the 100-kernel weight. Haplotype analysis showed that in 813 RILs, the gene Zm00001d028188 had 2 haplotypes (CC and GG), and the number of haplotype B (allele GG) in pop3 was significantly higher than that of haplotype A (allele CC) ( Figure 7 d-g). Haplotype analysis showed that the GG gene played an important role in the HKW of maize. Therefore, the GG of the Zm00001d028188 gene is the haplotype type that significantly increases the 100-kernel weight. The results of the present invention contribute to further studying the regulation mechanism of the 100-kernel weight of maize and also provide technical support for breeding high-quality maize varieties. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1For the comparison of population development and 100-kernel weight, where a is the number of each RIL in the multi-parent population, b is the grain phenotype of each parent, c is the 100-kernel weight of the parent, d is the 100-kernel weight of the RIL in DH, and e is the 100-kernel weight of the RIL in BS;
[0013] Figure 2 For the correlation results of pop1, pop2, pop3, pop4, and pop5 with HKW under two environments;
[0014] Figure 3 Among them, a is the phylogenetic tree of pop1, pop2, pop3, pop4, and pop5; b is the kinship analysis of 813 maize RILs; c is the Bayesian clustering map of 813 maize RILs when k = 5;
[0015] Figure 4 For the genotype diversity and LD decay map, where a is the chromosomal-specific SNP density within a 1-Mb interval, and the range of the number of SNPs is represented by a green-to-red scale; b is the frequency distribution of SNP missing data points; c is the allele frequency (MAF) distribution; d is the change of the genome-wide LD decay (r2) of 813 maize RILs with physical distance (Kb);
[0016] Figure 5 Among them, a is Dehong, b is Baoshan, and c is the phenotypic distribution, GWAS Manhattan plot, and q-q plot (from left to right) of the 100-kernel weight (HKW) trait of BLUP;
[0017] Figure 6 For the QTL mapping of 5 subgroups, where a is the QTL mapping of HKW in Pop1; b is the QTL mapping of HKW in Pop2; c is the QTL mapping of HKW in Pop3; d is the QTL mapping of HKW in Pop4; e is the QTL mapping of HKW in Pop5; and blue represents bin markers, orange represents DH, and purple represents the BS environment;
[0018] Figure 7 For the identification of candidate genes related to maize HKW, where a is the QTL identification of different chromosomes of maize pop3 in the DH environment; b is the QTL identification of different chromosomes of maize pop3 in the BS environment; c is the haplotype analysis of Zm00001d028188; d is the comparison of the 100-kernel weight of two haplotypes (CC and GG) in DH and e in BS; f is the comparison of HKW of two haplotypes of each RIL in DH and g in BS; h is the proportion of the two haplotypes in RILs; i is the position of Zm00001d028188 and its related SNPs. Detailed implementation methods
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] The present invention provides the Zm00001d028188 gene for regulating the hundred-kernel weight of maize. The nucleotide sequence of the Zm00001d028188 gene is preferably as shown in SEQ ID NO.1, and the amino acid sequence of the protein encoded by the Zm00001d028188 gene is as shown in SEQ ID NO.2:
[0021] SEQ ID NO.1:
[0022] 5’
[0023]
[0024] AGACTACTGTTAAATGTAGTAGCA
[0025]
[0026] SEQ ID NO.2:
[0027] MLIRQRVICSPTSFRFKLSDCLGRRIGPSFLGRQGGDSTRLVQDLYRTFDQVNNEESPSDEKLPESFRDFLLEMKDNHYDARTFAVRLKATMENMDKEVKRSRLAEQLYKHYAATAIPKGIHCLSLRLTDEYSSNAHARKQLPPPELLPLLSDNSFQHYILASDNILAASVVVSSTVRSSSVPEKVVFHVITDKKTYPGMHSWFALNSVSPAIVEVKGVHQFDWLTRENVPVLEAIESHRGVRNHYHGDHGTVSSASDNPRMLASKLQARSPKYISLLNHLRIYLPELFPNLNKVVFLDDDIVVQRDLSPLWAINLEGKVNGAVETCRGEDSWVMSKRFRTYFNFSHPVIARSLDPDECAWAYGMNIFDLAVWRKTNIRDTYHFWLKENLKSGLTLWKFGTLPPALIAFRGHVHGIDPSWHLLGLGYQDKTDIESVRRAAVIHYNGQCKPWLDIAFKNLQPFWTNHVNYSNDFVRNCHILEPERVKE。
[0028] Tropical and subtropical maize germplasm contains abundant genetic variations that are lacking in temperate maize, and is an important germplasm resource for maize breeding. In this invention, a multi-parent population with rich variation in grain oil content was constructed using 6 maize inbred lines mainly from tropical and subtropical regions with obvious differences in grain size and weight. A functional gene Zm00001d028188 closely associated with 100-kernel weight in maize was mined from the multi-parent population derived from these 6 parents. This gene can explain 6.3% of the phenotypic variation in 100-kernel weight. Haplotype analysis showed that in 813 RILs, the gene Zm00001d028188 had 2 haplotypes (CC and GG), and the number of haplotype B (allele GG) in pop3 was significantly higher than that of haplotype A (allele CC) ( Figure 7 d-g). Haplotype analysis showed that the GG gene played an important role in maize HKW. The results of this invention contribute to further research on the regulation mechanism of maize 100-kernel weight, and also provide an important reference for molecular-assisted selection in maize high-yield breeding.
[0029] Example 1:
[0030] 1.1 Plant materials
[0031] The experiment was conducted in 2019 in experimental fields planted in two different ecological environments, DH (longitude: 98.6°E, latitude: 24.4°N) and BS (longitude: 99.2°E, latitude: 25.1°N). Five multi-parent populations were obtained using the single-seed descent method: pop1 (Ye107 × CML312), pop2 (Ye107 × YML32), pop3 (Ye107 × CML373), pop4 (Ye107 × CML395), and pop5 (Ye107 × Q11); the six parents (Ye107, CML312, YML32, CML373, CML395, and Q11) were published in the literature
Jiang F, Liu L, Li Z, et al. Identification of Candidate QTLs and Genes for Ear Diameter by Multi-Parent Population in Maize. Genes (Basel). 2023;14(6):1305. Published 2023 Jun 20. doi:10.3390 / genes14061305
Jiang F, Liu L, Li Z, et al. Identification of Candidate QTLs and Genes for Ear Diameter by Multi-Parent Population in Maize. Genes (Basel). 2023;14(6):1305. Published 2023 Jun 20. doi:10.3390 / genes14061305
Zhen, S., Gao, G., Wang, X., Ning, H., and Duan, X. (2004). Appraisal of drought-enduring quality of several maize inbred lines. J Maize Sci. (in Chinese) 12, 18-19.
[0032] The pedigrees, ecotypes, and 100-kernel weights of the six parental lines are shown in Table 1. Among them, pop1 has 125 RILs, pop2 has 156 RILs, pop3 has 156 RILs, pop4 has 196 RILs, and pop5 has 180 RILs. Finally, an MPP population of 813 maize RILs was constructed.
[0033] Table 1 Parent Information
[0034]
[0035]
[0036] 1.2. Experimental Design
[0037] In 2019, at Dehong (denoted as DH) and Baoshan (denoted as BS), a randomized complete block design (RCBD) was adopted. There were three replicates at each location. Each experimental plot was 4.0 m long, with a row spacing of 0.7 m and an inter-plant spacing of 0.25 m. There were 14 plants in each row, and 10 plants were sampled from the middle of each row. The maize in the experimental field was cultivated according to the local standard agronomic practices. The measurement method was to randomly select 100 grains, measure their average weight after three replicates, and weigh them with an electronic balance.
[0038] 1.3. BLUP Analysis
[0039] After preliminary processing of the phenotypic data collected at two time points and locations, the Ime4 package in R software (V4.0.5) was used to analyze the correlation of the 100-kernel weight (HKW) of different populations in different environments, calculate the mean, standard deviation, skewness, kurtosis, and coefficient of variation of the oil content, and perform a normal distribution test. The lme4 software package in R (v3.2.2) was used to perform the best linear unbiased prediction (BLUP) analysis on the collected data. The formula is shown in Equation Ⅰ:
[0040] Y = μ + Line + Loc + (Line × Loc) + Rep(Loc) + ε Equation Ⅰ;
[0041] Among them, Y, μ, Line, and Loc represent phenotype, intercept, variety effect, and environment effect respectively. Rep represents different replicates, and ε represents the random effect. Line × Loc represents the interaction between variety and environment, and Rep(Loc) represents the associated effect of the within-environment replicate effect.
[0042] 1.4. DNA Extraction and Genome Sequencing
[0043] First, the cetyltrimethylammonium bromide (CTAB) method was used to extract genomic DNA from maize seedling leaves. Subsequently, the genomic DNA isolated from each F7RIL was digested with the restriction enzymes PstI and MspI, and then ligated with barcode adapters using T4 ligase (New England BioLabs). A GBS DNA library was constructed according to the GBS protocol and sequenced.
[0044] All ligated samples were pooled and purified using the QIAquick PCR Purification Kit (QIAGEN, Valencia, CA, USA). Polymerase chain reaction (PCR) amplification was performed using primers matching the adapters. Finally, the PCR products were purified and quantified using the Qubit dsDNA HS Assay Kit (Life Technologies, Grand Island, NY, USA). After selecting 200 - 300 bp PCR products using the Egel system (Life Technologies), the library concentration was estimated using a Qubit 2.0 fluorometer and the Qubit dsDNA HS Assay Kit (Life Technologies). Subsequently, sequencing reads were generated using TASSEL v5.0 (Li C, Guan H, Jing X, et al. Genomic insights into historical improvement of heterotic groups during modern hybrid maize breeding. Nat Plants. 2022;8(7):750 - 763.). Before performing TASSEL analysis, 80 poly(A) bases were appended to the 3' ends of all sequencing reads. For comparative analysis, the B73_V4 reference genome sequence was used, and analysis was performed using Sentieon software (parameters "bwa mem - k 32 - M - R") (Pei S, Liu T, Ren X, Li W, Chen C, Xie Z. Benchmarking variant callers in next - generation and third - generation sequencing analysis. Brief Bioinform. 2021;22(3):bbaa148). The alignment results were sorted and duplicate - removed using Samtools (using the parameter rmdup). A total of 591,483 high - quality SNPs were finally generated and annotated using the ANNOVAR (Wang K, Li M, Hakonarson H. ANNOVAR: functional annotation of genetic variants from high - throughput sequencing data. Nucleic Acids Res. 2010;38(16):e164.) software tool.
[0045] 1.5, SNP characterization, phylogenetic tree, PCA, and linkage disequilibrium analysis
[0046] Phylogenetic tree analysis was performed using Tassel v5.0 software, and 591,483 high-quality SNPs were used to evaluate the genetic relationships among 813 RILs. Principal component analysis (PCA) was performed using the R package 4.3.2, and the results were visualized using the scatterplot3d package. LD decay was evaluated using the original SNP data with Pop LD decay v3.42 (Zhang C, Dong SS, Xu JY, He WM, Yang TL. PopLDdecay: a fast and effective tool for linkage disequilibrium decay analysis based on variant call format files. Bioinformatics. 2019;35(10):1786-1788.). The parameter settings for calculating the r2 (correlation coefficient) value were set to the default values. The LD decay plot was drawn using the default parameters.
[0047] 1.6, Genome-wide association study
[0048] The efficient mixed-model association (EMME) analysis method in the GEMMA (Genome-wide efficient mixed-model analysis for association studies) (Zhou X, Stephens M. Genome-wide efficient mixed-model analysis for association studies. Nat Genet. 2012;44(7):821-824. Published 2012 Jun 17.) software package was used for GWAS. The following mixed-model method was used for the analysis:
[0049] y = Xa + Sb + Km + e Equation II;
[0050] where y represents the phenotype, a and b are fixed effects, representing the marker and non-marker effects respectively, and m represents the unknown random effect. The incidence matrices of a, b, and m are represented by X, S, and K respectively, and e is the vector of random residual effects. To correct for population structure, the present invention uses the first three principal components (PCs) to construct the S matrix, while the kinship (K) matrix is constructed using the simple matching coefficient matrix. The genetic relationships between individuals are modeled as random effects using the K matrix. In the association analysis, a significant P-value threshold of p < 1×10-6 was set to control type I error.
[0051] The present invention uses PLINK (Purcell S, Neale B, Todd-Brown K, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet. 2007;81(3):559-575.) to calculate independent markers with the parameters -indeppairwise 50 50 2. The significance threshold -log10(p)>4.5 is calculated using the formula -log10(1 / number of SNPs) to identify significant SNPs associated with maize HKW. SNPs that meet or exceed the threshold are extracted using bedtools v1.7 (Strable J, Wallace JG, Unger-Wallace E, et al. Maize YABBY Genes drooping leaf1 and drooping leaf2 Regulate Plant Architecture. Plant Cell. 2017;29(7):1622-1641.), and candidate genes associated with maize HKW are identified in the 50 kb regions upstream and downstream of the significantly associated SNPs based on the B73 (RefGen_v4) reference genome and annotation information. Based on the observation that the r2 value shows a plateau at 50 kb in the LD decay plot, the present invention decides to screen for candidate genes associated with maize HKW within the 50 kb regions upstream and downstream of the significantly associated SNPs.
[0052] 1.7, Construction of genetic map and QTL mapping
[0053] By setting the parameter MAF≥0.05, the SNP data is filtered by aligning it with the maize reference genome B73 (RefGen_v4). Then, the genetic linkage map is constructed using JoinMap4.0 through allelic SNPs (Ooijen, J.W., Ooijen, J.W., Verlaat, J.V., Ooijen, J.W., Tol, J., Dalén, J., Buren, J.B., Meer, J.V., Krieken, J.H., Ooijen, J.W., Kessel, J.S., Van, O., Voorrips, R.E., & Heuvel, L.P. (2006). 4. Software for the calculation of genetic linkage maps in experimental populations. Linkage groups were formed using a LOD threshold ≥ 2.5. Quantitative trait locus (QTL) mapping of HKW was performed using the composite interval mapping (CIM) method with Windows QTL Cartographer v2.0 (Zeng ZB. Precision mapping of quantitative trait loci. Genetics. 1994;136(4):1457 - 1468). The LOD threshold was set based on 1000 permutation tests at a significance level of p ≤ 0.05. QTLs with a LOD threshold ≥ 2.5 were considered significant. The proportion of phenotypic variance explained (PVE) by each QTL was determined using the squared correlation coefficient (R2). The QTL names were constructed by starting with the letter "q" to indicate a QTL, followed by the abbreviation of the trait name, the corresponding chromosome number, and the marker position (Ribaut JM, Hoisington DA, Deutsch JA, Jiang C, Gonzalez - de - Leon D. Identification of quantitative trait loci under drought conditions in tropical maize. 1. Flowering parameters and the anthesis - silking interval. Theor Appl Genet. 1996;92(7):905 - 914.).
[0054] 1.8. Identification and functional annotation of candidate genes
[0055] The significant SNPs identified in GWAS were compared with the QTL mapping results to identify concordant sites, and SNPs overlapping within the QTL intervals were selected to screen for candidate genes. Candidate genes were searched within a 50 - kb range upstream and downstream of the significant SNPs. Candidate genes were predicted in MaizeGDB (https: / / www.maizegdb.org / ) using the maize B73 v4 reference genome. Functional annotations of candidate genes were obtained using the InterPro database.
[0056] 1.9. Haplotype analysis
[0057] Haploview v4.2 software was used to perform haplotype analysis on SNPs related to HKW in two environments. First, a haplotype map was constructed using high-density whole-genome SNPs, and the haplotypes of SNPs significantly associated with maize 100-kernel weight were determined based on the positions of these loci and the results of LD analysis. Finally, the genes within the haplotypes were annotated to identify functionally related gene loci.
[0058] 2. Results
[0059] 2.1. 100-kernel weight phenotypic analysis
[0060] An MPP population consisting of 813 RILs from 5 parents was constructed ( Figure 1 a). The differences in kernels of the MPP population parents are shown in Figure 1 b. As shown in Figure 1 c, the phenotypic value of the male parent Ye107 was significantly lower than that of all female parents in the HKW trait. Analysis of variance showed significant differences in the HKW trait among the 5 populations ( Figure 1 d - e). Table 2 presents the descriptive statistical analysis of HKW for the 5 RIL subgroups and calculates the coefficient of variation for each subgroup in the two environments. The CV indicates that the variation in HKW among the 5 subgroups ranges from 14.3% (YML32) to 26.8% (CML395) in DH and from 15.6% (YML32) to 28.9% (CML395) in BS. To further evaluate the reliability of the phenotypic identification, we calculated the correlation coefficient between the two environments, and the results showed a highly significant positive correlation for the HKW phenotype. The correlation coefficients for the five subgroups in the two environments were 0.767, 0.734, 0.987, 0.979, and 0.975 (p < 0.001), respectively. The absolute value of skewness for all five subgroups was less than 1, indicating a small degree of bias. Overall, the HKW of the MPP population showed wide differences; however, the variation in HKW was consistent between the two environments, indicating that the phenotypic data had high reliability for further analysis.
[0061] Table 2 Statistical analysis results of 100-kernel weight phenotype
[0062]
[0063]
[0064] Note: DH represents the experiment conducted in Dehong in 2019, and BS represents the experiment conducted in Baoshan in 2019.
[0065] Pearson correlation analysis (significance level p < 0.001) showed that there was a significant and strong correlation between the overall performances of RILs in the same population under different environmental conditions, and the correlation coefficients ranged from 0.73 to 0.99.Figure 2 )。This strong correlation indicates that in multiple environments, the response patterns of RILs to 100-seed weight are highly consistent. Even under different environmental variations, their performances still maintain a high degree of consistency. This phenomenon not only proves the stability and reliability of the experimental data but also further indicates that the genetic control of the 100-seed weight trait may be relatively stable in different environments. In addition, the strong correlation results also suggest that environmental factors may not have produced significant variations in the performance of 100-seed weight, or even if there are, the genetic backgrounds among RILs enable them to have a strong internal consistency in response to environmental changes.
[0066] 2.2, Population Structure of the RIL Population
[0067] In this study, the Admixture software was used to analyze the population structure of 813 materials. The distributions of pop1, pop2, and pop5 populations are relatively concentrated, indicating less genetic variation within these populations, while the significant separation between populations reflects greater genetic differences. Due to the influence of the common parent Ye107, there is a certain overlap between pop3 and pop4, but the remaining 3 RIL populations are relatively independent ( Figure 3 a, 3c). The significant overlap between CML373 (pop3) and CML395 (pop4) indicates that they have similar genetic backgrounds and both belong to the non-Reid group, showing relatively high genetic diversity and frequent gene exchanges. While CML312 (pop1), YML32 (pop2), and Q11 (pop5) form their own independent branches, demonstrating different evolutionary trajectories and relatively homogeneous genetic backgrounds. As Figure 3 shown, the results of population structure, principal component analysis (PCA), and genetic distance or correlation ( Figure 3 a-c) are consistent. In the PCA plot, the scattered points may come from within-population heterogeneity or outliers. Based on pedigree or genetic background, the MPP population can be divided into five major clusters. When K = 5 ( Figure 3 c), the population structure of the MPP population becomes clear. Phylogenetic tree analysis also shows five genetic clusters, which are consistent with the population structure based on genetic relationships.
[0068] 2.3, SNP Characterization and LD Decay
[0069] Through genotyping-by-sequencing (GBS), a total of 591,483 high-quality genome-wide SNPs were identified, which are distributed on all 10 chromosomes of maize. Figure 4The heatmap of a shows the marker density of SNPs on maize chromosomes. The number of SNPs found on chromosomes 1-10 is as follows: 82,889, 67,249, 67,934, 75,637, 57,899, 47,683, 53,494, 50,217, 44,840, and 43,631. Chromosome 1 has the largest number of SNPs, while chromosome 10 has the smallest. The SNP density per million base pairs on chromosomes 1-10 is 269.99, 265.05, 275.11, 288.26, 306.30, 258.60, 273.99, 293.31, 277.25, and 280.65, respectively. SNPs are evenly distributed on the chromosomes. In the filtered SNP dataset, the average missing rate is 0.19, and the average minor allele frequency (MAF) is 0.20, indicating that this dataset is suitable for subsequent GWAS( Figure 4 b, c). 591,483 SNPs were used to evaluate the linkage disequilibrium (LD) decay in the association mapping population. LD in the association mapping panel was estimated using r2 of paired combinations of SNPs across chromosomes. The minimum threshold was set at 0.38. The physical distance range for estimating LD decay was approximately 50 kb( Figure 4 d). LD decayed rapidly, indicating that the higher the degree of domestication, the greater the selection intensity, forming more favorable genetic relationships among populations.
[0070] 2.4, Genome-wide association analysis of HKW
[0071] In this study, there were significant differences in the HKW traits of the MPP population under the DH and BS environments( Figure 5 ). GWAS analysis of the HKW traits was performed using the MLM model in both environments. This model considered population structure and kinship. In the association study, population structure and kinship matrix were used as covariates to reduce false positives. The q-q plot showed that false positives for the HKW traits were effectively controlled( Figure 5 a-c). A threshold of -log10(P) > 4.5 was set, and 21 SNPs significantly associated with HKW were identified. Among them, 5 were identified in DH, 6 in BS, and 10 in BLUP( Figure 5, (Table 3). These SNPs are distributed on 7 chromosomes (c1, c2, c4, c5, c6, c8, and c9). SNP-25750352 located on chromosome 1 was detected in BS and BLUP; SNP-73091888 on chromosome 2 was detected in DH, BS, and BLUP, and SNP-176207154 was detected in DH and BLUP; SNP-233438083 on chromosome 4 was detected in DH, BS, and BLUP, and SNP-182679880 was detected in BS and BLUP; SNP-51610984 on chromosome 5 was detected in DH and BLUP; SNP-36178213 on chromosome 9 was detected in DH and BLUP. Using MaizeGDB, InterPro, UniProt, and NCBI, combined with relevant published studies, candidate genes were comprehensively screened within the 50Kb flanking regions of significant SNPs. Four potential candidate genes (Table 4) (Zm00001d028185, Zm00001d028186, Zm00001d028187, Zm00001d028188) were identified within the 50Kb regions flanking these loci.
[0072] Details of SNPs significantly associated with the HKW trait in Table 3
[0073]
[0074]
[0075] Table 4 Potential candidate genes and their functions identified within the 50Kb regions flanking important SNP loci
[0076]
[0077] 2.5, QTL mapping of the MPP population
[0078] QTL mapping and effect analysis of HKW in 5 subgroups under two different environments, screening out SNP markers with a missing rate above 10% and loci with a minor allele frequency less than 5%, and setting the LOD threshold to ≥2.5. pop1( Figure 6 ) A total of 3 HKW QTLs were detected, namely qHKW4-1, qHKW1-1, and qHKW6-1, which explained 7.3%, 7.8%, and 7.4% of the phenotypic variation, respectively, under two different environments (Table 5). Among them, qHKW4-1 identified on chromosome 4 had the highest LOD of 3.26 and the largest additive effect of 0.97 under two different environments. The remaining QTLs were only detected in the BS environment, and all additive effect values were negative. In pop2 under two environments ( Figure 6) Two HKW QTLs, qHKW3-1 and qHKW4-2 (Table 5), were detected in both, explaining 6.9% and 6.5% of the phenotypic variation, respectively. Among them, the LOD value of qHKW4-2 was the largest, at 4.3. pop3( Figure 6 ) A total of four HKW QTLs were detected, namely qHKW1-2, qHKW3-2, qHKW5-1, and qHKW1-3, which explained 6.3%, 5.6%, 5.9%, and 6.2% of the phenotypic variation, respectively, in two different environments (Table 5). The qHKW1-3 identified on chromosome 1 had the highest LOD in the two different environments, at 4.34. The additive effect values of qHKW1-2 and qHKW1-3 were both positive, indicating that these two QTLs had a positive effect on HKW. pop4( Figure 6 ) A total of eleven HKW QTLs were detected, with the phenotypic variation ranging from 4.7% to 8.0%. Among them, qHKW2-1 had the highest LOD in the two different environments, at 4.82. The additive effect values of qHKW8-1, qHKW10-1, qHKW5-2, qHKW7-2, qHKW10-3, and qHKW10-4 were all positive, indicating that these six QTLs had a positive effect on HKW. Pop5( Figure 6 ) A total of ten HKW QTLs were detected, with the phenotypic variation ranging from 4.6% to 15.2%. qHKW5-8 had the highest LOD in the two different environments, at 7.95. The additive effect values of qHKW5-4 and qHKW5-7 were both positive, indicating that these two QTLs had a positive effect on HKW.
[0079] Table 5 Positions and effects of HKW QTLs detected in the MPP population
[0080]
[0081]
[0082]
[0083] 2.6, Identification of genes related to HKW
[0084] In this study, QTL mapping and GWAS analysis were used to identify genes related to HKW. A comparison of the results of the two analyses showed that the SNP (SNP-25750352) on chromosome 1 was located within the QTL intervals of qHKW1-2( Figure 7 a) and qHKW1-3( Figure 7 b). Therefore, we consider these QTL intervals to be reliable and consistent (Table 6). Based on functional annotation, one candidate gene related to HKW, Zm00001d028188, was identified within these QTL intervals Figure 7c). SNP-25750352 is located 9,968 bp downstream of the candidate gene Zm0000d028188 ( Figure 7 i). Further analysis based on 813 maize RILs showed that in 19DH and 19BS, the number of haplotype B (allele GG) of pop3 was significantly higher than that of haplotype A (allele CC) ( Figure 7 d-g). Haplotype analysis showed that the GG gene played an important role in maize HKW. As Figure 7 shown in h, the CC haplotype was present in pop1, pop2, pop3, pop4, and pop5. The GG haplotype was present in pop1, pop2, pop3, and pop4.
[0085] Table 6 Co-localized genes of GWAS and QTL
[0086]
[0087] 2.7, SNP additive and dominance analysis
[0088] For the candidate genes related to HKW, a significant SNP locus: SNP-26750352 was deeply analyzed. By evaluating the additive and dominance effects of this locus, we found that SNP-26750352 showed strong positive additive effects and dominance effects under different environmental conditions. This result indicates that this SNP locus has a major positive role in regulating maize HKW, and shows stable synergistic characteristics in all environments, suggesting its value as a potentially important regulatory gene. In addition, the above results further indicate that additive effects and dominance effects play important control roles in the development of maize HKW. It is particularly noteworthy that the parent Ye107 showed a significant increasing effect in HKW, indicating that Ye107 has good donor potential in improving maize grain-related traits.
[0089] Table 7 Additive and dominance effect values of SNP
[0090]
[0091] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. Application of a reagent for detecting SNP molecular markers related to 100-grain weight of corn in screening high 100-grain weight corn assisted breeding, characterized in that: The molecular marker is located at the 25750352bp position of chromosome 1 of the corn genome, the gene reference version of the position is the genome version B73_v4, the molecular marker has two haplotypes, the alleles of the two haplotypes are CC and GG respectively, and the number of alleles in the haplotype with GG is significantly higher than that of the allele CC in the haplotype, which indicates a higher 100-grain weight.
Citation Information
Patent Citations
Application of Gene Zm00001d040827 in Increasing Maize Yield
AU2020104114A4
Molecular marker relevant to yield and quality of zea mays grains
CN110541046A