Application of Zm00001eb378820 gene in molecular marker-assisted breeding for regulating 100-kernel weight of maize kernels

By combining QTL localization and GWAS method, SNP_27712538 and functional gene Zm0001eb378820 on chromosome 9 solved the problem of difficult to analyze the genetic basis of corn grains with 100 grains, and provided the gene resources and regulatory mechanism for high-yield and high-quality corn breeding.

CN119662708BActive Publication Date: 2025-06-20FOOD CROPS RES INST YUNNAN ACAD OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411883934.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-06-20
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively analyze the genetic basis and regulatory mechanism of corn grains weighing hundreds of grains, affecting high-yield and high-quality breeding of corn.

Method used

By combining QTL localization and GWAS methods, the genetic basis of corn 100000 weight was studied, and SNP_27712538 located on chromosome 9 was located as a significant correlation site, and the functional gene Zm00001eb378820, which is closely related to corn 1000 weight was excavated.

Benefits of technology

The functional mechanism of the Zm00001eb378820 gene in the regulation of corn 100 grains in size and size is revealed, providing potential genetic resources for the breeding of high-yield corn varieties, helping to understand the genetic regulation mechanism of corn 100 grains and providing key targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119662708B_ABST
    Figure CN119662708B_ABST
Patent Text Reader

Abstract

The present invention provides an application of the Zm00001eb378820 gene in molecular marker-assisted breeding for regulating the 100-kernel weight of maize, which relates to the field of agricultural biotechnology. Using the backbone maize inbred line Ye107 as the male parent, it was crossed with 3 excellent inbred lines (D39, R-2-1-1, and YML1218) to construct a multi-parent population with significantly different 100-kernel weights. By measuring the 100-kernel weight phenotypic data in multiple environments and combining high-density SNP markers, GWAS analysis, and QTL mapping, the SNP_27712538 located on chromosome 9 was identified as a significantly associated locus, and the functional gene Zm00001eb378820 closely related to the regulation of maize 100-kernel weight was mined. The present invention provides new insights into the genetic regulation mechanism of maize 100-kernel weight and provides key targets for the breeding of high-yield maize varieties, thus significantly improving the breeding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of agricultural biotechnology, and particularly relates to the application of the Zm00001eb378820 gene in molecular marker-assisted breeding for regulating the 100-kernel weight of maize kernels. Background Art

[0002] With the continuous growth of the global population, the shortage of feed grains has become the main challenge facing future food security. Maize, known as the "king of feeds", plays an important role in making up for the insufficient supply of feed grains. Maize kernel-related traits directly affect the yield level and are one of the core target traits for maize variety improvement, showing relatively high heritability. In particular, the 100-kernel weight, as a key trait with strong heritability, is an important component determining the maize yield level and is often used as an indicator to measure kernel quality. Although there is usually a positive correlation between the 100-kernel weight and the kernel weight, the two are not strictly linearly correlated, and the 100-kernel weight is more representative and stable in explaining the variation of maize yield. Therefore, in-depth analysis of the genetic basis and its regulatory mechanism of the 100-kernel weight has important theoretical value for revealing the genetic laws of maize yield formation. QTL mapping and genome-wide association study (GWAS) are two main research methods for analyzing maize kernel size and weight-related genetic loci. These two methods provide important means for accurately mapping the key genes or genomic regions regulating the 100-kernel weight of maize. In this study, by combining QTL and GWAS methods, the genetic basis and regulatory mechanism of the 100-kernel weight of maize were studied, and its potential application value in high-yield and high-quality maize breeding was analyzed, aiming to provide a theoretical basis and technical support for cultivating new high-yield and high-quality maize varieties. Summary of the Invention

[0003] Aiming at the deficiencies of the prior art, the present invention provides the application of the Zm00001eb378820 gene in molecular marker-assisted breeding for regulating the 100-kernel weight of maize kernels, which not only reveals the functional mechanism of the Zm00001eb378820 gene in regulating the 100-kernel weight of maize, but also provides potential gene resources for the breeding of high-yield maize varieties.

[0004] To achieve the above object, the present invention is realized by the following technical solutions:

[0005] The application of a Zm00001eb378820 gene in molecular marker-assisted breeding for regulating the 100-kernel weight of maize kernels, wherein the nucleotide sequence of the Zm00001eb378820 gene is as shown in SEQ ID NO.1, and the amino acid sequence of the encoded protein of the Zm00001eb378820 gene is as shown in SEQ ID NO.2.

[0006] Preferably, the regulation of the 100-kernel weight of maize grains is achieved by overexpressing the Zm00001eb378820 gene in maize to increase the 100-kernel weight of maize grains.

[0007] Preferably, the breeding includes cultivating maize with a large 100-kernel weight, increasing the maize yield, and selecting new maize varieties.

[0008] The present invention provides an application of the Zm00001eb378820 gene in molecular marker-assisted breeding for regulating the 100-kernel weight of maize grains. Compared with the prior art, the advantages are as follows:

[0009] The present invention provides the Zm00001eb378820 gene for regulating the 100-kernel weight of maize and its application. The present invention uses the temperate maize inbred line Ye107 with a large 100-kernel weight as a common parent and crosses it with 3 tropical and temperate maize inbred lines with a small 100-kernel weight to construct a maize multi-parent population with a significant difference in 100-kernel weight. By measuring the 100-kernel weight phenotypic data in multiple environments and combining high-density SNP markers, GWAS analysis, and QTL mapping, the SNP_27712538 located on chromosome 9 was identified as a significantly associated locus. This locus shows large additive and dominant effects, and through further analysis, the functional gene Zm00001eb378820 closely related to the regulation of maize 100-kernel weight was mined. Haplotype analysis showed that in 485 RILs, the gene Zm00001eb378820 has 4 haplotypes (Hap1, Hap2, Hap3, and Hap4), and there are significant differences in the 100-kernel weight phenotypes corresponding to Hap1 and Hap4 (p<0.05), and the 100-kernel weight corresponding to Hap4 is significantly lower than that of Hap1. Therefore, Hap4 of the Zm00001eb378820 gene is a haplotype type that significantly increases the 100-kernel weight of maize. The results of the present invention help to further provide new insights into the genetic regulation mechanism of maize 100-kernel weight and provide key targets for the breeding of maize varieties with a high 100-kernel weight. Description of the Drawings

[0010] Figure 1 Shows the 100-kernel weight phenotypic distribution results of pop1, pop2, and pop3 in three environments;

[0011] Figure 2 Shows the correlation analysis results of the 100-kernel weight phenotypic data of pop1, pop2, and pop in three environments;

[0012] Figure 3 Shows the distribution of SNPs detected by whole-genome resequencing on 10 maize chromosomes and their related characteristics;

[0013] Figure 4 Shows the genetic structure analysis of the multi-parent population;

[0014] Figure 5 LD decay results for 3 recombinant inbred line populations;

[0015] Figure 6 Manhattan plots and Q-Q plots for BLUP, 21YS, 22YS, and 23YS; among them, the left figures in (a) to (d) are Manhattan plots, and the right figures are Q-Q plots. Each dot in the left figure represents an SNP, the black line represents the significance threshold, different colors represent different chromosomes, and the red line in the right figure is the trend line corresponding to the ideal Q-Q plot in each case; (a) is the GWAS result of BLUP, (b) is the GWAS result of YS21, (c) is the GWAS result of YS22, and (d) is the GWAS result of YS23;

[0016] Figure 7 Detection results of significant QTLs co-localized with significant SNPs in GWAS;

[0017] Figure 8 Haplotype analysis results of the Zm00001eb378820 gene, * indicates p≤0.05; (a) is the distribution frequency of 4 haplotypes in 3 populations, and (b) is the box plot of the change in 100-kernel weight corresponding to 4 haplotypes;

[0018] Figure 9 Genomic estimated breeding value (GEBV) and additive and dominant effects of SNPs significantly associated with 100-kernel weight. Specific implementation manners

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0020] The present invention provides the Zm00001eb378820 gene for regulating the 100-kernel weight of maize, and the amino acid sequence of the protein encoded by the Zm00001eb378820 gene is shown in SEQ ID NO.2. In the present invention, the nucleotide sequence of the Zm00001eb378820 gene is preferably shown in SEQ ID NO.1, and the specific sequence is as follows:

[0021] SEQ ID NO.1

[0022]

[0023]

[0024]

[0025] SEQ ID NO.2

[0026]

[0027]

[0028] Tropical and temperate maize germplasms are important germplasm resources for maize breeding. In this invention, a multi-parent population with rich variation in 100-kernel weight was constructed using 4 tropical and temperate maize inbred lines with significantly different 100-kernel weights. A functional gene Zm00001eb378820 closely associated with maize 100-kernel weight was mined from the multi-parent population derived from these 4 parents. Haplotype analysis showed that in 485 RILs, the gene Zm00001eb378820 had 4 haplotypes (Hap1, Hap2, Hap3, and Hap4), and there were significant differences in the 100-kernel weight phenotypes corresponding to Hap1 and Hap4 (p<0.05), and the 100-kernel weight corresponding to Hap4 was significantly lower than that of Hap1. Therefore, Hap4 of the Zm00001eb378820 gene is a haplotype type that can significantly increase the 100-kernel weight of maize. The results of this invention help to further provide new insights into the genetic regulation mechanism of maize 100-kernel weight and provide key targets for the breeding of maize varieties with high 100-kernel weight.

[0029] Example 1:

[0030] 1.1 Plant materials

[0031] The experiments were conducted in three different ecological environments in Yanshan County, Yunnan Province (denoted as YS, altitude 1540 m, longitude 104.5°E, latitude 23.6°N) in 2021, 2022, and 2023, respectively, with experimental fields planted. Three multi-parent populations were obtained using the single-seed descent method: Pop1 (D39 × Ye107), Pop2 (R-2-1-1 × Ye107), and Pop3 (YML1218 × Ye107). The four parents (Ye107, D39, R-2-1-1, and YML1218) are disclosed in the literature

Jiang F, Liu L, Li Z, et al. Identification of Candidate QTLs and Genes for Ear Diameter by Multi-Parent Population in Maize. Genes (Basel). 2023;14(6):1305. Published 2023 Jun 20. doi:10.3390 / genes14061305

Bi, Y., Jiang, F., Yin, X. et al. Identification of candidate gene associated with maize northern leaf blight resistance in a multi-parent population. Plant Cell Rep 43, 189 (2024). https: / / doi.org / 10.1007 / s00299-024-03269-w

[0032] Table 1

[0033]

[0034] 1.2, Field experiments and collection of phenotypic data

[0035] In this study, a completely randomized block design (RCBD) was adopted, and the research objects included a multi-parent population composed of 485 F7 RILs and their 4 parents. The plants were grown in Yanshan County, Yunnan Province, China (longitude 104.27°E, latitude 23.61°N) in 2021, 2022, and 2023 respectively. The experiment was set with 3 replicates to ensure the reliability and statistical significance of the results. In this experiment, each F7 RIL was planted in two rows of holes, with 14 plants in each row, the row length was 4 m, the row spacing was 0.70 m, and the plant spacing was 25 cm. Standard agronomic practices were followed during the experiment.

[0036] 1.3, Phenotypic data analysis

[0037] After the preliminary processing of the phenotypic data collected in 3 years with 3 replicates, the Anderson-Darling test in Minitab software was used for the normality test of the data, and the mean, standard deviation, variance, skewness, kurtosis, coefficient of variation, and variation range of the data were analyzed. The Pearson test in R software (V4.1.3) was used to analyze the correlation of the 100-seed weight data of different populations and years. Subsequently, the "aov" function was used for one-way analysis of variance, and the heritability was calculated through the "lme4" package. The calculation of heritability referred to the method of Knapp et al.

Knapp SJ. Confidence intervals for heritability for two-factor mating design single environment linear models. Theor Appl Genet. 1986;72(5):587-591.

[0038]

[0039] where, h 2 represents the broad-sense heritability, σg 2 is the genotype variance, σge 2 is the variance of the interaction between the environment and the genotype, σε 2 is the error variance, e is the environment, and r is the replicate.

[0040] 1.4, DNA extraction and whole-genome sequencing (WGS)

[0041] The genomic DNA of the seedling leaves of the maize multi-parent population was extracted using the cetyltrimethylammonium bromide (CTAB) method (Maroof, 1994), and the quality and purity of the extracted DNA were detected. After ensuring that the data met the experimental requirements, an Illumina TruSeq DNA Sample Preparation Kit was used to construct a genomic DNA library. After fragmenting the genomic DNA to approximately 350 bp, 150-bp paired-end reads sequencing was performed using the Illumina HiSeq X Ten platform. The raw reads generated by the sequencing were processed using the Trimmomatic software to remove adapter sequences and low-quality reads, obtaining high-quality clean reads. Subsequently, the clean reads were aligned to the maize reference genome B73 (RefGen_v5) using the BWA software to generate an alignment file BAM. The GATK software was used to detect single nucleotide polymorphisms (SNPs) and insertion-deletion variations (InDels) to generate a variant file VCF. To ensure the reliability of the data, the PLINK tool was used to filter the variations in the VCF file, setting the filtering criteria as the missing rate > 0.2 and the minor allele frequency (MAF) < 0.05, retaining high-quality variant data. ANNOVAR was used to annotate the high-quality SNPs to determine the distribution of the variant sites on the genome and their mutation types [Wang, K.; Li, M.; Hakonarson, H. ANNOVAR: Functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 2010, 38, e164.]. The annotated variant data was subjected to association analysis with the phenotypic data, and subsequent analyses such as population structure analysis, GWAS, and QTL mapping were carried out.

[0042] 1.5, Population Structure Analysis and LD Decay

[0043] To understand the genetic structure of the F7 RILs, this study used the STRUCTURE software for population structure analysis. We set different numbers of populations (K values) and determined the optimal K value through cross-validation error. The R4.1.3 software was used to perform PCA (principal component analysis) on the genotype data to further verify the population stratification results, and the scatterplot3d package was used to visualize the PCA results. Analyze the decay of linkage disequilibrium (LD) across the genome, use the PopLDdecay software under the Linux system to calculate the LD values between SNPs, and determine the degree of linkage disequilibrium (r 2), thus generating a set of corresponding data points of LD values and physical distances. By fitting curve to analyze the relationship between the physical distance of SNPs and r 2 to evaluate the decay trend of LD with the increase of distance. Subsequently, use the built-in script Plot-OnePop.pl of the software to plot the LD decay curve, and at the same time use the "LD decay" package in R language for visualization. To determine the critical r 2 value, set the minimum threshold to 0.2, which indicates the maximum physical distance at which LD significantly decays

Zhang C, Dong S, Xu J, et al. PopLDdecay: a fast and effective tool for linkage disequilibrium decay analysis based on variant call format files[J]. Bioinformatics, 2019, 35(10): 1786 - 1788.

[0044] 1.6, Genome-wide association study (GWAS)

[0045] GWAS was performed using genotype and phenotype data from three different environments. First, we performed quality control on the initially obtained genotype data, filtering out low-quality SNP sites and retaining high-quality SNPs. The FarmCPU model in GEMMA (http: / / www.xzlab.org / software.html) was used to perform GWAS analysis on these high-quality SNPs. The PVE (phenotypic variance explained rate) was calculated using the formula proposed by Shim et al. [Shim H, Chasman DI, Smith JD, et al. A multivariate genome-wide association analysis of 10 LDL subfractions, and their response to statin treatment, in 1868 Caucasians [J]. PLoS One, 2015, 10(4): e0120758.]. In the model, we incorporated the first few principal components of population structure (such as PC1, PC2, and PC3) and other known significantly associated genetic markers (SNPs) as covariates into the model to better control for possible spurious associations. The significance threshold was determined based on the total number of SNPs and calculated using the formula -log10(1 / total number of SNPs) at the p < 0.05 level. Candidate genes that were closely related to the 100-seed weight variation and had specific functions were further identified among the significant SNPs with reference to the maize genome B73 RefGen_v5 (Zhou, 2012). Manhattan plots and QQ plots were generated using the "ggplot2" package in R software (V4.1.3). According to the LD decay results and the relative positions of the SNPs, candidate genes were screened within a 20-kb range upstream and downstream of the significant SNPs [Zhang, X.; Ren, Z.Y.; Luo, B.W.; Zhong, H.X.; Ma, P.; Zhang, H.K.; Hu, H.M.; Wang, Y.K.; Zhang, H.Y.; Liu, D.; et al. Genetic architecture of maize yield traits dissected by QTL mapping and GWAS in maize. Crop J. 2022, 10, 436 - 446.].

[0046] 1.7, Construction of Genetic Linkage Maps and QTL Mapping

[0047] Based on high-density genotyping technology, high-quality SNP markers were screened out, and markers with high missing rates and significant segregation distortion were excluded. The JoinMap 4.0 software was used to perform linkage analysis on the screened high-quality SNP markers. According to the recombination frequencies between SNP markers, the markers were assigned to different linkage groups, and the maximum likelihood method was used to sort the markers within each linkage group. The Kosambi mapping function was used to calculate the genetic distance (cM) between markers, and finally, a genetic linkage map of three recombinant inbred line populations was constructed. Subsequently, the appropriate LOD threshold was determined through 1000 random permutation tests, and the significance level was set at p ≤ 0.05

Delannoy, E.; Stanley, W. A.; Bond, C. S.; Small, I. D. Pentatricopeptide repeat (PPR) proteins as sequence-specificity factors in post-transcriptional processes in organelles. Biochem. Soc. Trans. 2007, 35, 1643-1647.

[0048] 1.8, Joint Analysis of GWAS and QTL

[0049] Based on B73 RefGen_v5, the physical positions of the significant SNPs identified in the GWAS analysis were compared with the significant intervals obtained from the QTL mapping analysis to observe whether there was co-localization. Through GWAS and linkage analysis, candidate genes related to the target trait were obtained. Subsequently, the candidate genes were annotated and functionally predicted through databases such as Maize GDB, UniProt, NCBI, and InterPro to find candidate genes closely related to the hundred-kernel weight.

[0050] 1.9, Gene Haplotype Analysis

[0051] Haplotype analysis was performed on the genes associated with the significant association SNP loci that were consistently detected in multiple environments and co-localized with QTLs. Haplotype blocks in the target regions were constructed based on the genotype data, the linkage disequilibrium (LD) structure of these regions was analyzed, and its effect on the 100-seed weight trait was evaluated. The calculation of LD values was completed using the Haploview software. By comparing the associations between different haplotypes and the 100-seed weight phenotype, the dominant haplotypes that might significantly affect this trait were screened out, further revealing potential functional variations and their genetic effects.

[0052] 1.10, SNP effect analysis and calculation of genomic estimated breeding value (GEBV)

[0053] The GCTA software was used to calculate the additive, dominant effects of SNPs and the epistatic effects between significant SNPs, and the GEBVs of the three subgroups were calculated simultaneously. First, the genomic relationship matrix (GRM) was constructed using the genotype data in PLINK format, and the additive effects were estimated through the generalized linear mixed model (GREML) of GCTA. The dominant GRM was constructed to estimate the dominant effects of SNPs, and the epistatic effects of significant SNP pairs were further analyzed. Subsequently, the GRM and phenotypic data were used for the calculation of GEBV, thereby obtaining the GEBV of each population.

[0054] 2. Results

[0055] 2.1. Phenotypic analysis of 100-seed weight in the multi-parent population

[0056] The 100-seed weight data in this study were obtained by planting the multi-parent population in Yanshan, Yunnan for three years. The statistical analysis results of the 100-seed weight data of the three subgroups are shown in Table 2. The average 100-seed weight of pop2 was the lowest, while that of pop3 was the highest. The coefficient of variation (CV) of the three subgroups in the three environments ranged from 0.19 to 0.24, indicating the differences in 100-seed weight among samples. Further analysis showed that the frequency distribution histograms of 100-seed weight of the three populations in different environments (such as Figure 1 ) were consistent with the normal distribution, and the absolute values of skewness and kurtosis were both less than 1, indicating that the degree of data bias was small and in line with the characteristics of the normal distribution. The broad-sense heritabilities of 100-seed weight of the three populations were 85.96%, 84.87%, and 91.59% respectively. This relatively high heritability indicates that the variation of this trait is mainly driven by genotype variation and is suitable for GWAS and QTL mapping studies.

[0057] Table 2 Descriptive statistical analysis of 100-seed weight in the multi-parent population

[0058]

[0059]

[0060] Correlation analysis was performed on the 100-seed weight data of three populations under different environments. As Figure 2 shown, in the YS21, YS22, and YS23 environments, the correlation coefficients of the 100-seed weight of pop1 were 0.68, 0.63, and 0.73 respectively, those of pop2 were 0.65, 0.61, and 0.72 respectively, and those of pop3 were 0.79, 0.75, and 0.81 respectively. The correlation coefficients were all significant (p < 0.05). The consistent high correlations among the three environments indicated that the responses of pop1, pop2, and pop3 to the 100-seed weight were stable in different environments, once again indicating that the 100-seed weight was mainly affected by genetic factors, ensuring the phenotypic reliability of subsequent GWAS analysis and QTL mapping. Analysis of variance (ANOVA) was performed on the 100-seed weight data of the F7RIL population under different environments (Table 3), and it was found that there were extremely significant differences in the 100-seed weight of different populations under different environments (p < 0.001). This indicated that the influence of different environments on the 100-seed weight was extremely significant, meaning that environmental factors such as climate, soil type, water supply, etc. had an undeniable impact on the 100-seed weight.

[0061] Table 3 Analysis of variance for multi-parent populations, *** represents p < 0.001

[0062] Population Degree of freedom Sum of squares Mean square F value P value F7RILs 2 4282 2141.2 67.3 <2e-16*** Pop1 2 2579 1289.7 43.27 <2e-16*** Pop2 2 883 441.7 12.83 3.82e-06*** Pop3 2 1151 575.7 19.63 6.87e-09***

[0063] 2.2. Population structure analysis and LD decay

[0064] Through whole-genome resequencing (WGS), a total of 194,963,454 variant sites were identified. To ensure data quality, the original data were strictly filtered, including removing sites with quality score (QUAL) ≤ 20, sequencing depth (DP) ≤ 3, missing rate > 20%, and minor allele frequency (MAF) < 0.05. Finally, 6,395,521 high-quality SNP sites were obtained, and these sites were evenly distributed on 10 chromosomes of maize ( Figure 3 ). Among them, chromosome 1 had the largest number of SNPs, reaching 2,467,471, while chromosome 10 had the smallest number of SNPs, only 1,200,136. The average marker density of the whole genome was 2975.46 SNPs / Mb, and the marker density of chromosome 4 was the highest (3453.33 SNPs / Mb), while that of chromosome 6 was the lowest (2776.07 SNPs / Mb). In the filtered SNP dataset, the average value of MAF was about 0.125, and the average value of the missing rate was about 0.175, indicating that the data met the high-quality standards for subsequent GWAS.

[0065] Further principal component analysis divided the F7RILs into three main subpopulations ( Figure 4a), this result is highly consistent with the subpopulation structure revealed by the kinship analysis. The phylogenetic tree analysis showed three genetic clusters, with the vast majority of families from the same subpopulation clustering together, while a few families clustered with individuals from other subpopulations.( Figure 4 b, Figure 4 c), which is consistent with the population structure distribution based on kinship, and this admixture between subpopulations may be caused by gene introgression or hybridization of the common male parent Ye107 during the breeding process.

[0066] To evaluate the LD decay in the association population, we analyzed the distribution of 6,395,521 high-quality SNPs on the chromosomes and used the r 2 value (LD intensity between two SNPs) to estimate LD. Figure 5 showed a trend of gradual decrease in the r 2 value with the increase of physical distance. In the short-distance range, the r 2 value decreased rapidly, showing a strong LD effect. When the physical distance increased to about 20 Kb, the r 2 value decreased to 0.1 and leveled off at greater distances. This indicates that within the range of 20 Kb, LD decays significantly, and the rate of LD decay slows down and gradually levels off after exceeding this distance. This result provides an important reference for determining the physical distance for screening candidate genes upstream and downstream of significant SNPs across the genome.

[0067] 2.3 GWAS Analysis and Candidate Gene Mining under Different Environments

[0068] Across the genome, a significance threshold of 5.5 was determined for SNPs using the formula -log10(1 / number of SNPs), and 42 significantly associated SNP loci were identified at the p < 0.05 level, with the PVE ranging from 0.37% to 8.91% and an average PVE of 5.54%. These were distributed on chromosomes 1, 2, 3, 4, 5, 6, 7, 9, and 10, among which the concentrated distribution of SNPs on chromosomes 5 and 9 was the most significant( Figure 6 ). Most of these SNPs were detected in a single environment, but 4 SNPs showed consistent associations in multiple environments. Among them, 5-51394339 and 9-27712538 were consistently detected in the YS22, YS23 environments and in BLUP, and 9-27712538 had the largest effect value (Table 4). These findings laid an important foundation for the screening of candidate genes.

[0069] Combined with the LD analysis results of the associated population, by comparing with B73 RefGen_v5 on the maize GDB website, we mined 50 candidate genes within the 20 kb regions upstream and downstream of 42 SNPs significantly associated with 100-kernel weight. Based on the PVE values of the SNPs, the functional annotations of the candidate genes, and their consistency in different environments, a total of 5 candidate genes that may affect maize 100-kernel weight were screened out, namely Zm00001eb225640, Zm00001eb225650, Zm00001eb228310, Zm00001eb228320, and Zm00001eb378820, which are located on chromosomes 5 and 9 (Table 5).

[0070] Table 4 SNPs significantly associated with 100-kernel weight consistently identified in multiple environments

[0071]

[0072] Table 5 Candidate genes revealed by GWAS

[0073]

[0074]

[0075] 2.4, Construction of linkage map and QTL mapping

[0076] Genetic linkage maps were constructed for three subgroups (RIL-D39, RIL-R-2-1-1, and RIL-YML1218). The genetic maps of pop1, pop2, and pop3 contained 910, 2,837, and 2,400 SNP markers, respectively, covering all chromosomes. The total genetic distances were 588.07 cM, 2,487.83 cM, and 1,501.77 cM, respectively, and the average genetic distances between markers were 0.65 cM, 0.88 cM, and 0.63 cM, respectively. Under three different environments, QTL mapping analysis was performed on the 100-seed weight and BLUP values of the three populations. Through 1,000 permutation tests, the LOD threshold was determined to be 2.5 at a significance level of p ≤ 0.05. The results showed that a total of 17 significant QTLs were detected in the three populations, distributed on chromosomes 1, 2, 4, 6, 7, 8, and 9. The range of LOD was 2.52-6.34, and the range of phenotypic variation explained rate (R2) was 5.73%-13.93%, with an average R2 of 8.66% (Table 6). Notably, in Pop2, three significant QTLs (qHGW1-1, qHGW1-2, and qHGW1-3) located on chromosome 1 had an overlapping interval of 119,325,456-129,429,998. In Pop3, two significant QTLs qHGW9-3 and qHGW9-4 located on chromosome 9 had an overlapping interval of 23,960,895-47,646,228 (Table 6).

[0077] Table 6 Distribution of significant QTLs in different populations

[0078]

[0079]

[0080] 2.5, Joint analysis of GWAS and QTL

[0081] We compared the QTL mapping results of the three subgroups with the GWAS results of MPP. In GWAS, SNP 9-27712538 was identified as a significantly associated locus in YS22, YS23, and BLUP, and this SNP overlapped with two co-localized QTLs (qHGW9-3 and qHGW9-4) in pop3, with LOD values of 4.96 and 4.24 for the two QTLs, and R2 values of 13.09% and 12.32%, respectively. In addition, SNP 9-108046020 co-localized with qHGW9-1 in pop1, and the LOD value of qHGW9-1 was 3.46, and R 2 was 7.89% (Table 6,[[]]END]] Figure 7) Therefore, we speculate that the gene Zm00001eb378820 mined within 20 kb upstream and downstream of locus 9-27712538, as well as the genes Zm00001eb388530 and Zm00001eb388540 mined within 20 kb upstream and downstream of locus 9-108046020, may play a regulatory role in 100-seed weight. We queried Zm00001eb388530 and Zm00001eb388540 and found that they encode pentatricopeptide repeat protein (PPR) and transcription factor protein (GLK73) respectively, and both proteins are closely related to the growth and development of maize plants. Therefore, it is preliminarily speculated that a total of 7 candidate genes were screened out by the combined analysis of GWAS and QTL, namely Zm00001eb225640, Zm00001eb225650, Zm00001eb228310, Zm00001eb228320, Zm00001eb378820, Zm00001eb388530 and Zm00001eb388540.

[0082] 2.7, Haplotype analysis of genes

[0083] The gene Zm00001eb378820 was consistently identified in GWAS and QTL analyses, appeared in multiple environments, and had a high expression level in grain-related tissues. Therefore, it was used as the key candidate gene in this study. Haplotype analysis was performed on it, and a total of 4 haplotypes were identified (Table 6). There were significant differences in the 100-seed weight phenotypes corresponding to Hap1 and Hap4 (p<0.05), and the 100-seed weight corresponding to Hap4 was significantly lower than that of Hap1 ( Figure 8 b). In the haplotype frequency distribution, the Pop2 population contained two haplotypes, Hap1 and Hap4, and the frequency of Hap4 exceeded 60%; in contrast, the Pop3 population only contained the Hap1 haplotype ( Figure 8 a). This haplotype distribution pattern was consistent with the 100-seed weight performance among populations. The mean 100-seed weight of the Pop2 population was the smallest, while the mean 100-seed weight of the Pop3 population was the largest. This indicates that Hap4 is closely related to a lower 100-seed weight, while Hap1 is closely related to a higher 100-seed weight.

[0084] Table 7 Summary of haplotypes corresponding to Zm00001eb378820

[0085]

[0086] 2.8, SNP effect analysis and genomic estimated breeding value (GEBV)

[0087] In this study, the additive, dominant, and epistatic effects among SNPs (5-51,394,339, 5-64,069,455, 6-73,588,308, and 9-27,712,538) significantly associated with 100-seed weight were evaluated. The results showed that the additive and dominant effects of 9-27712538 were the largest and both were positive values (Table 4, Figure 9 ), indicating that this SNP had a positive effect on 100-seed weight and could increase the 100-seed weight of maize. Moreover, it not only exerted its effect through the additive effect of individual alleles but was also significantly affected by the dominant effect of different allele combinations. The epistatic effects among the four selected significant SNPs were analyzed, and the results showed that there were significant interactions between 5-51394339 and 6-73588308 and between 5-51394339 and 5-64069445 (p<0.05). Among them, the epistatic effect between 5-51394339 and 6-73588308 was 2.24, and that between 5-51394339 and 5-64069445 was -2.05. This indicated that there were important gene-gene interactions among these three SNPs, and such interactions might lead to non-additive genetic effects, thereby affecting the size of 100-seed weight. In addition, by calculating the GEBV of the three subgroups, it was found that there were significant differences in GEBV among the three populations. Among them, pop3 showed the highest mean GEBV among the three subgroups, and the GEBV differences among individuals were significant, indicating that this population had great potential for genetic improvement. In contrast, the mean GEBV of pop2 was the lowest. These results provided a reference for further population selection and genotype optimization.

[0088] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An application of Zm00001eb378820 gene in regulating 100-grain weight of maize by molecular marker-assisted breeding, characterized in that: The nucleotide sequence of the Zm00001eb378820 gene is shown in SEQ ID NO.1; the regulation of corn grain 100-grain weight is to increase the corn grain 100-grain weight by adjusting the overexpression of the Zm00001eb378820 gene in corn; the breeding is to cultivate corn with high 100-grain weight, increase corn yield, and breed new corn varieties.

Citation Information

Patent Citations

  • Molecular marker of maize drought resistance related gene and application thereof

    CN115029475A

  • Methods and Compositions for Haploid Mapping

    US20090064361A1