Use of a Zm00001eb257890 gene in regulating the kernel percentage of corn
By locating and regulating the expression level of the Zm00001eb257890 gene, the problem of corn seed yield regulation was solved, achieving efficient regulation of corn seed yield, increasing the yield per corn plant, and providing support for the breeding of high-yielding corn varieties.
Patent Information
- Application Number
- CN202510654299.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Existing technologies are insufficient to effectively regulate corn seed yield, thus affecting the yield per corn plant, and there is a lack of efficient functional gene regulation methods.
Gene editing of the Zm00001eb257890 gene was used to regulate maize seed yield by adjusting its expression level. CRISPR/Cas9, TALEN, zinc finger nuclease and other technologies were used, combined with genome-wide association analysis and genetic linkage analysis, to locate and mine functional genes related to seed yield, providing technical support for molecular marker-assisted selection of high-yield maize varieties.
This study explained 4.68% of the phenotypic variation in seed yield, improved the precision of maize seed yield regulation, provided technical support for breeding high-yielding and stable maize varieties, and revealed the regulatory mechanism of maize seed yield.
Smart Images

Figure CN120193125B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of agricultural biotechnology, specifically to the application of the Zm00001eb257890 gene in regulating maize seed yield. Background Technology
[0002] As one of the world's most important food crops, maize is not only a vital source of human nutrition and animal feed, but also plays a crucial role in the field of bioenergy. Increasing its yield is of paramount strategic importance for ensuring global food security. Among the many yield traits of maize, shelling percentage, calculated as the ratio of dry ear grain weight to dry ear weight, is an important indicator of the efficiency of photosynthetic product distribution in maize ears. As a key trait measuring the efficiency of photosynthetic product distribution to kernels, its level directly affects the yield per maize plant.
[0003] Existing research has confirmed that shelling percentage is closely related to the number of kernels per ear, ear weight, and kernel yield, and has become an important indicator for breeding high-yielding maize varieties. In summary, improving shelling percentage is a key objective of plant breeding and biotechnology-assisted improvement. Therefore, identifying functional genes closely related to maize shelling percentage can provide technical support for marker-assisted selection in breeding high-yielding and stable maize varieties. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides an application of the Zm00001eb257890 gene in regulating maize seed yield. It was found that the Zm00001eb257890 gene is a functional gene that regulates seed yield, and this gene can explain 4.68% of the phenotypic variation in seed yield.
[0005] To achieve the above objectives, the technical solution of the present invention is implemented through the following technical solution:
[0006] An application of the Zm00001eb257890 gene in regulating maize seed yield, wherein the CDS sequence of the Zm00001eb257890 gene is shown in SEQ ID NO.1, and the physical location of the QTL site is chr5: 215.05-224.10Mb; the amino acid sequence of the protein encoded by the Zm00001eb257890 gene is shown in SEQ ID NO.2.
[0007] A recombinant vector comprising a regulatory element containing the Zm00001eb257890 gene.
[0008] Preferably, the regulatory element comprises a tissue-specific promoter or an inducible promoter.
[0009] The method to regulate plant seed yield is achieved by altering the expression level of the Zm00001eb257890 gene, specifically as follows:
[0010] (1) Reduce gene expression to increase seed yield;
[0011] (2) Enhance gene expression to reduce seed yield.
[0012] Preferably, the alteration of gene expression level is achieved through gene editing technology, and the gene editing technology is selected from any one of CRISPR / Cas9, TALEN, and zinc finger nucleases.
[0013] Preferably, the gene editing technology is used to modify the 5'UTR region or exon region of the sequence shown in SEQ ID NO.1.
[0014] Preferably, the alteration of gene expression level is achieved through RNA interference, antisense RNA technology, or promoter substitution.
[0015] Preferably, the plant is a grass (Poaceae).
[0016] The kit contains a probe set for detecting the Zm00001eb257890 gene, the sequence of which is shown in SEQ ID NO: 1. This kit can be used to detect the corn seed yield.
[0017] This invention provides an application of the Zm00001eb257890 gene in regulating maize seed yield, which has the following advantages compared with existing technologies:
[0018] This invention utilizes the temperate maize inbred line Ye107 as a common parent, and crosses it with five tropical and subtropical maize inbred lines to construct a multi-parental maize population with significant differences in seed yield. Genome-wide association analysis and genetic linkage analysis were used to locate a single nucleotide polymorphism (SNP) on chromosome 5 that is significantly associated with seed yield. This SNP, located at 222146671 bases on chromosome 5, was named SNP5_222146671. Furthermore, the functional gene Zm00001eb257890, which regulates seed yield, was identified. This gene explains 4.68% of the phenotypic variation in seed yield. qRT-PCR expression level detection and genetic structure analysis of Zm00001eb257890 suggest that this gene may negatively regulate ear development, leading to a significant increase in its expression in parents with low seed yield. Haplotype analysis further confirmed the association between the G / A mutation and the seed yield phenotype, providing a basis for the development of molecular markers. This result also helps to further study the regulatory mechanism of maize seed yield and provides technical support for the breeding of high-yield maize varieties. Attached Figure Description
[0019] Figure 1 This is a schematic diagram illustrating a multi-parent population with significant differences in seed yield constructed in an embodiment of the present invention;
[0020] Figure 2 This is a frequency distribution diagram of the seed production rate of the multi-parent population in three environments: Jinghong (21), Yanshan (22), and Yanshan (23) in this embodiment of the invention.
[0021] Figure 3 This document presents the genotype diversity, linkage disequilibrium decay analysis, and population structure analysis in this embodiment of the invention; where a is the chromosome-specific SNP density of a 1 million base pair (Mb) genomic region; b is the linkage disequilibrium decay plot of 678 recombinant inbred lines; c is the principal component analysis of 678 recombinant inbred lines; and d is the Bayesian clustering plot of 678 recombinant inbred lines when K = 5.
[0022] Figure 4 The figures shown are Manhattan plots and QQ plots for Jinghong (21), Yanshan (22), Yanshan (23), and the optimal linear unbiased prediction value in this embodiment of the invention. In the figures a-d, the left figure is the Manhattan plot, and the right figure is the QQ plot. In the left figure, each dot represents a SNP, the black line represents a threshold <1×10−4, and different colors represent different chromosomes. The red line in the right figure is the trend line corresponding to the ideal QQ plot for each case. Figure a represents the result for Jinghong (21), figure b represents the result for Yanshan (22), figure c represents the result for Yanshan (23), and figure d represents the result of the optimal linear unbiased prediction value.
[0023] Figure 5 This is a schematic diagram of the co-location sites of genome-wide association analysis and QTL in an embodiment of the present invention; where a is the joint location map of SNP5_222146671 in genome-wide association analysis and QTL; b is the linkage analysis of SNP5_222146671; c is the positional relationship between SNP5_222146671 and Zm00001eb257890; d is the ratio of the two haplotypes of non-synonymous SNPs; e is the ratio of the two haplotypes in the multi-parent population;
[0024] Figure 6 This invention presents the development of spike length and diameter of Zm00001eb257890 in six parents at four different time periods and their correlation with relative expression levels in this embodiment of the invention; where a represents the development of spike length and diameter of each parent at different time periods; and b represents the correlation between spike length and diameter of each parent and the relative expression level of Zm00001eb257890 at different time periods.
[0025] Figure 7 This document describes the relative expression levels of Zm00001eb257890 in the middle and tip segments of the spikelets of six parents at four different time periods in this invention embodiment.
[0026] Figure 8 This is an example of the relative expression levels of Zm00001eb257890 in the spikelets of six parents at four different time periods in this invention.
[0027] Figure 9 This embodiment of the invention illustrates the effect of non-synonymous SNPs in the coding region of Zm00001eb257890 on its domains and motifs. Where a represents a conserved domain of Zm00001eb257890; b represents amino acid changes caused by non-synonymous SNPs; and c represents mutant amino acids (V...). 1322 I) Location in the protein tertiary structure: The red box indicates the location of the site, and the blue area indicates the serine / threonine phosphorylation site. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] The Zm00001eb257890 gene, which regulates maize seed yield, is provided. The amino acid sequence of the protein encoded by the Zm00001eb257890 gene is shown in SEQ ID NO.2. The preferred CDS sequence of the Zm00001eb257890 gene is shown in SEQ ID NO.1, and the specific sequence is as follows:
[0030] SEQ ID NO.1:
[0031]
[0032] SEQ ID NO.2:
[0033] Example 1:
[0034] 1. Materials and Methods
[0035] 1.1 Parental lines and population structure
[0036] Using the temperate maize backbone inbred line Ye107 as the male parent, it was crossed with five tropical and temperate maize inbred lines (YML226, Q11, Shen137, Chang 7-2, and YML1218). After nine generations of self-pollination, a multi-parent population composed of five different recombinant inbred lines was constructed. Significant differences were observed in shelling percentage-related traits such as ear length and ear diameter in this multi-parent population. Figure 1 It contains a total of 678 recombinant inbred lines, with 132, 140, 115, 146, and 135 recombinant inbred lines in subgroups 1, 2, 3, 4, and 5, respectively. The ecotypes and heterosis groups of the six parent lines are shown in Table 1.
[0037] Table 1 Parental Information
[0038]
[0039] 1.2 Field Trials and Phenotypic Data Analysis
[0040] The experimental materials were planted in Jinghong City, Yunnan Province (100°52′E, 21°41′N) in 2021, and in Yanshan County, Yunnan Province (103°35′E, 23°18′N) in 2022 and 2023, respectively, and named 21 Jinghong, 22 Yanshan, and 23 Yanshan. A completely randomized block design was used, with two replicates per group. Each experimental field had rows 4 meters long, row spacing of 0.70 meters, and plant spacing of 25 centimeters, with 14 plants per row. Field management for all experimental materials was consistent with local production management. After maize maturity, individual ears were harvested, bagged, and numbered for each population. The maize ears were completely dried, and five ears of uniform size were selected from each recombinant inbred line for phenotypic analysis, including ear weight and grain weight. The shelling percentage (shelling percentage % = grain weight / ear weight × 100%) was further calculated. The average of the two replicates for each recombinant inbred line was used as the phenotypic data for statistical analysis. Statistical analysis of seed yield phenotypic data from multi-parent populations was performed using SPSS 27.0, including assessment of normality and analysis of variance. Broadly sensed heritability (HbA1c) was calculated using the lme4 package (v1.1-31.1) in R. 2The variance refers to the proportion of genetic variance to total variance, calculated with reference to [Knapp SJ. Confidence intervals for heritability for two-factor mating design single environment linear models. Theor Appl Genet. 1986;72(5):587-591.].
[0041] Formula I;
[0042] σ 2 g Indicates genotype; σ 2 ge σ represents the variance of the genotype-environment interaction; 2 e denoted as the error term; r is the number of repetitions, with 2 repetitions set, r=2; n is the number of environments, which is 3, as the experiment was conducted in three environments: 21 Jinghong, 22 Yanshan, and 23 Yanshan.
[0043] Based on the phenotypic values of all three environments, the optimal linear unbiased prediction of seed yield in each recombinant inbred line was estimated using a mixed linear model, following the method described in [Zhang Z, Ersoz E, Lai CQ, Todhunter RJ, TiwariHK, Gore MA, Bradbury PJ, Yu J, Arnett DK, Ordovas JM et al: Mixed linearmodel approach adapted for genome-wide association studies. Nature Genetics 2010, 42(4):355-360.].
[0044] Formula II;
[0045] In the formula, Y ijlk Represents the phenotypic value, μ represents the intercept, and Line represents the phenotypic value. i Loc represents the genetic effect of the i-th genotype. j Rep represents the influence of the j-th environment. j This indicates the j-th repetition, and (Line × Loc) ij This represents the interaction effect between the i-th genotype and the j-th environment. Additionally, Rep(Loc) jl ε represents the effect of the j-th repetition in the l-th environment. ijk This indicates a random effect.
[0046] 1.3 Genomic DNA Extraction and Sequencing Genotyping
[0047] Genomic DNA was extracted from maize seedling leaves using a modified CTAB method [Allen GC, Flores-Vergara MA, Krasnyanski S, Kumar S, Thompson WF: A modified protocol for rapid DNA isolation from plant tissues using cetyltrimethylammonium bromide. Nature Protocols 2006, 1(5):2320-2325.]. After quantification using Qubit (v4.0), the DNA was diluted to 20 ng / μL for subsequent library construction. Whole-genome resequencing was performed on 678 recombinant inbred lines using the Illumina HiSeq™ platform (Illumina, San Diego, CA, USA). Clean reads were obtained after Trimmomatic filtering of low-quality sequences. These clean reads were then aligned with the maize reference genome B73 RefGen_v5 using BWA software (parameters: mem -t 4 -k 32 -M -R) [Li H, Durbin R: Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 2009, 25(14):1754-1760.]. The alignment results were then converted into SAM / BAM files using SAMtools [Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, GenomeProject Data P: The Sequence Alignment / Map format and SAMtools. Bioinformatics]. 2009, 25(16):2078-2079.]; Perl scripts were used to calculate the alignment rate and coverage; SAMtools was used to sort and remove duplicates from the alignment results (parameters: sort & rmdup) for mutation detection.GATK (v4.2.6.1) [McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A, Garimella K, Altshuler D, Gabriel S, Daly M et al: The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Research 2010, 20(9):1297-1303.] was used to detect population SNPs in the filtered BAM files, and SNPs with MAF < 0.1 or deletion rate > 0.6 were filtered out; finally, a high-quality variant dataset was recorded in vcf format. ANNOVAR v2020-06-08 software and Phytozome v13.0 maize genome annotation database were used to perform functional analysis on variant sites, identify variants located in coding, regulatory, or non-coding regions of genes, and screen for potential harmful mutations based on SIFT scores (< 0.05).
[0048] 1.4 LD Decay Estimation and Population Structure Analysis
[0049] The chain imbalance coefficient (r) was calculated using popLDdecay v3.42 software. 2 [Zhang C, Dong SS, Xu J-Y, He WM, Yang TL: popLDdecay: a fast and effective tool for linkage disequilibrium decay analysis based on variant call format files. Bioinformatics 2018, 35(10):1786-1788.], and then use the script Plot_Onepop.pl in the package to plot the paired r 2 A graph showing the relationship between linkage disequilibrium values and physical distance was generated to produce a linkage disequilibrium decay map. By analyzing the linkage disequilibrium decay curves, the decay of linkage disequilibrium with physical distance in the population was determined, providing key parameters for subsequent genome-wide association studies.
[0050] In R software, principal component analysis (PCA) was performed using the scatterplot3d 0.3-41 package. The first three principal components were extracted using the -pca3 option for PCA analysis [Bradbury PJ, Zhang Z, Kroon DE, Casstevens TM, Ramdoss Y, Buckler ES: TASSEL: software for association mapping of complex traits in diverse samples. Bioinformatics 2007, 23(19):2633-2635.]. Furthermore, population structure analysis was performed using Admixture (v1.3.0) [Tamaki I, Mizuno M, Ohtsuki T, Shutoh K, Tabata R, Tsunamoto Y, Suyama Y, Nakajima Y, Kubo N, Ito T et al: Phylogenetic, population structure, and population demographic analyses reveal that Vicia sepium in Japan is native and not introduced. Scientific Reports 2023, 13(1):20746.], and the grouping results under specific K values or single K values were visualized using the R package. Population structure analysis assesses the degree of genetic differentiation within the population, providing population structure information for subsequent genome-wide association studies to avoid the potential impact of population structure on the association analysis results.
[0051] 1.5 Genome-wide association analysis and haplotype analysis
[0052] This study followed Henderson's optimal linear unbiased prediction method based on the Mixed Linear Model (MLM) [Henderson CR: Estimation of genetic parameters. Annals of Mathematical Statistics 1950, 21:309-310.], and used the EMMAX software package v intel64-20,120,205 for genome-wide association analysis.
[0053] Formula III;
[0054] Where y represents phenotype, a and b are fixed effects, representing labeled and unlabeled effects respectively, and m represents unknown random effects. The occurrence matrices of a, b, and m are denoted by X, S, and K respectively, and e is a vector of random residual effects. To correct for population structure, the S matrix is constructed using the first three principal components (PCAs), while the kinship (K) matrix is constructed using a simple matching coefficient matrix. Genetic relationships between individuals are modeled as random effects using the K matrix. In association analysis, a significance threshold of P-value was set to P < 1 × 10⁻⁶. -6 Independent markers were calculated using PLINK with parameters -indep-painwise50 5 0.2. Significant SNPs associated with maize shelling percentage were identified using the formula -log10(1 / number of SNPs) and adjusted by a significance threshold of -log10(P)>5 from Bonferroni [Kaler AS, Gillman JD, Beissinger T, Purcell LC: Comparing Different Statistical Models and Multiple Testing Corrections for Association Mapping in Soybean and Maize. Front Plant Sci2019, 10:1794.]. SNPs meeting or exceeding the threshold were extracted using bedtools software v1.7 [Jiang F, Liu L, Li Z, Bi Y, Yin X, Guo R, Wang J, Zhang Y, Shaw RK, Fan X: Identification of Candidate QTLs and Genes for Ear Diameter by Multi-Parent population in Maize. Genes2023, 14(6):1305.]. Based on the B73 (RefGen_v5) reference genome and annotation information, candidate genes associated with maize seed yield were screened in the 20kb upstream and downstream regions of significantly relevant SNPs. Genome-wide association analysis was performed using a multi-parent population to identify candidate functional genes regulating seed yield.
[0055] SNP haplotype analysis was performed using a haplotype-based calling procedure [Utsunomiya YT,Milanesi M, Utsunomiya ATH, Ajmone-Marsan P, Garcia JF: GHap: an R package for genome-wide haplotyping. Bioinformatics 2016, 32(18):2861-2862.].
[0056] 1.6 Genetic Mapping and QTL Localization
[0057] Genotyping of progeny was filtered based on a 0.8 integrity threshold and a 0.001 partial segregation threshold to select high-quality population markers. Binary division was performed based on these population markers to obtain the final population markers used for analysis. For each population marker, Joinmap v4.0 was used to sort the binary markers of each population, and SMOOTH software was used to construct genetic maps [HuX-S, Goodwillie C, Ritland KM: Joining genetic linkage maps using a jointlikelihood function. Theoretical and Applied Genetics 2004, 109(5):996-1004.]. The QTL detection of phenotypic values and BULP values of maize seed yield under three conditions was performed using the Inclusive Composite Interval Mapping (ICIM) method in QTL IciMapping v4.2 software [Meng L, Li HH, Zhang LY, Wang JK: QTL IciMapping: Integrated software for genetic linkage map construction and quantitative trait locus mapping inbiparental populations. Crop Journal 2015, 3(3):269-283.]. The LOD threshold was set based on a 1000-times random permutation test with a significance level of P≤0.05. QTLs with an LOD threshold ≥2.5 were considered significant [Churchill GA, Doerge RW: Empirical threshold values for quantitative traitmapping. Genetics 1994, 138(3):963-971.]. The phenotypic variation explained (PVE) for each QTL was determined using the squared correlation coefficient (R²). The QTLs were named according to the nomenclature of McCouch et al. [McCouch S, Cho YG, Yano M, Paul E, Blinstrub M, Morishima H, Kinoshita T: Report on QTL nomenclature. 1997, 14:11-13.].
[0058] In QTL additive effect analysis, "a" represents the marker type of the common paternal parent, and "b" represents the marker type of the maternal parent. Positive additive effects indicate that the paternal allele plays a dominant role in increasing seed yield, while negative additive effects indicate that the maternal allele plays a dominant role in increasing seed yield.
[0059] 1.7 Discovery and Preliminary Functional Analysis of Key Candidate Genes
[0060] Based on colocalized SNPs obtained from genome-wide association analysis and QTL linkage mapping, key candidate genes were screened by combining the positional relationship and functional effects of SNPs with genes. TBtools-II v2.142 software was used to predict and visualize the functional domains of the coding regions of candidate genes, further analyzing their functional characteristics. In addition, the MEME online software was used to predict the motifs of functional genes to reveal their potential regulatory elements. The tertiary structure of candidate genes was predicted on the SWISS-MODEL platform, and disordered regions of the amino acid sequences of candidate genes were predicted using SMART to assess their structural characteristics [Letunic I, Bork P: 20 years of the SMART protein domain annotation resource. NucleicAcids Research 2018, 46(D1):D493-D496.].
[0061] 1.8 Comparative expression analysis of key candidate genes in parental lines
[0062] To investigate the expression patterns of key candidate genes in parental lines, samples were taken from the middle and top tissues at four stages of maize kernel development: V9, 0-DAP, 6-DAP, and 12-DAP, respectively. The top (10% of ear length) and middle (50% of ear length) of the ear were designated as middle and top, respectively [Bi YQ, Jiang FY, Zhang YD, Li ZW, Kuang TH, Shaw RK, Adnan M, Li KZ, Fan XM: Identification of a novel marker and its associated laccasegene for regulating ear length in tropical and subtropical maize lines. Theoretical and Applied Genetics 2024, 137(4).], to assess the relative expression levels of key candidate genes in different parental lines. RE levels were detected by quantitative real-time RT-PCR (qRT-PCR) using the Tiangen SuperReal Premix Plus (SYBR Green) kit (Tiangen, Beijing), repeated 3 times. The total reaction volume for each sample was 20 μL. The primer sequences used for qRT-PCR are shown in Table 1. The maize GAPDH gene was used as an internal reference gene for expression normalization [Li Y, Liu XQ, Chen RM, Tian J, Fan YL, Zhou XJ: Genome-scale mining of root-preferential genes from maize and characterization of their promoter activity. Bmc Plant Biology 2019, 19(1).]. The qRT-PCR program was as follows: 95℃ pre-denaturation for 3 min, 95℃ denaturation for 20 s, and 55℃ annealing extension for 30 s. This cycle was repeated 40 times. Fluorescence signals were captured during the annealing and extension steps. Finally, -2 −△△CTThe relative expression level of the gene was calculated using a method [Zhu Y, Liu Y, Zhou K, Tian C, Aslam M, Zhang B, Liu W, Zou H: Overexpression of ZmEREBP60 enhances drought tolerance in maize. Journal of Plant Physiology 2022, 275:153763.].
[0063] Table 2 Primer sequences
[0064]
[0065] 2. Results
[0066] 2.1 Phenotypic analysis of seed yield
[0067] Phenotypic statistical analysis was conducted on the seed production rates of multiparental populations under three different environments: Jinghong (21), Yanshan (22), and Yanshan (23) (see Table 3).
[0068] Table 3 Statistical analysis results of seed yield phenotype
[0069]
[0070] The results showed that subpopulation 3 had the highest average seed yield under all conditions, followed by subpopulation 4, while subpopulation 1 had the lowest average seed yield, followed by subpopulations 2 and 5. Under different conditions, the coefficient of variation of seed yield among the five recombinant inbred line subpopulations ranged from 4.41% to 6.10%, indicating that seed yield is a genetically stable trait and exhibits good stability during variety breeding. Furthermore, the absolute values of skewness and kurtosis of seed yield under different conditions were all less than 1, showing an approximately normal distribution. Figure 2 The results showed that the population exhibited the genetic characteristics of a quantitative trait, indicating that the multi-parent population was suitable for genome-wide association analysis. Further analysis revealed that the broad-sense heritability of the multi-parent population under the three environments ranged from 0.76 to 0.89, indicating a relatively high broad-sense heritability. This suggests that the seed yield trait is mainly regulated by genetic factors, with relatively little influence from environmental factors.
[0071] Analysis of variance showed that the differences in seed yield among the multi-parent populations were highly significant (P<0.001), while the interactions among the three environments and between the environment and the multi-parent populations were not significant (see Table 4).
[0072] Table 4. Analysis of variance of seed yield of multi-parent populations in three environments
[0073]
[0074] Furthermore, Spearman correlation analysis revealed a high correlation between seed yield phenotypic values among recombinant inbred line subpopulations under the three environments, with correlation coefficients ranging from 0.578 to 0.988 (P < 0.001), further confirming the consistency of the seed yield phenotype across different environments. These results indicate that seed yield exhibits significant genetic differences in multiparental populations and possesses high genetic stability across different environments, providing reliable phenotypic data to support subsequent molecular mapping and genetic analysis.
[0075] 2.2 SNP identification, linkage disequilibrium decay, and population structure analysis of multi-parent populations
[0076] Based on the phenotype of the multi-parent population, after rigorous quality control filtering, a total of 6,386,662 high-quality SNPs were identified. These SNPs are evenly distributed across the 10 chromosomes of maize. Figure 3 a). To assess linkage disequilibrium decay within the population, linkage disequilibrium decay analysis was performed using the original SNP dataset of the multiparent population (a). Figure 3 b). The results showed that the linkage disequilibrium decay rate was faster with the increase of the physical distance between SNP sites, and the linkage disequilibrium decay coefficient (R) was higher than that of the SNP sites. 2 When the distance between SNP sites tends to stabilize, the physical distance is approximately 20 kb. Based on this result, a 20 kb interval upstream and downstream of the SNP site was selected as the range for candidate gene screening to ensure the association between candidate genes and target traits.
[0077] The genetic structure of a multi-parent population was analyzed using principal component analysis. Figure 3 c), the results showed that the multi-parent population generally exhibited significant independence, although some families showed clustering in principal component analysis due to the use of the common parent Ye107. This result is consistent with the experimental design and reflects the genetic diversity of the multi-parent population. Further analysis of population structure ( Figure 3 d) At K = 5, the 678 recombinant inbred lines were clearly divided into 5 subpopulations, and the mixing phenomenon among the populations may be caused by genetic drift. Overall, the population structure analysis and principal component analysis results are highly consistent, verifying that the genetic characteristics of the multi-parent population in this study meet the requirements of the experimental design.
[0078] 2.3 Genetic Map Construction and QTL Localization
[0079] Using ICIMapping v4.2 software, this study constructed a genetic map of a multi-parent population covering all 10 maize chromosomes. The results showed that, through complete interval mapping, multi-environment QTL analysis of the multi-parent population detected 35 QTLs significantly associated with seed yield, distributed on chromosomes 1, 2, 3, 4, 5, 7, 8, 9, and 10. Chromosome 5 had the most significant QTLs (10) (Table 5). A single QTL could explain 6.37%–15.16% of the phenotypic variation, and the phenotypic explanation rate for all significant QTLs was higher than 5%. In terms of population distribution, 5, 9, 6, 6, and 9 QTLs were detected in subpopulations 1, 2, 3, 4, and 5, respectively. On chromosome 5, qKR5-3 (physical location 14.22–34.60 Mb) in subpopulation 2 showed significant overlap with qKR5-1, qKR5-2, and qKR5-4 (16.23–18.49 Mb); and the three QTLs (qKR5-6, qKR5-8, and qKR5-9) (19.48–20.46 Mb) and qKR5-7 (15.36–21.27 Mb) in subpopulation 4, which had completely overlapping physical regions, also showed significant overlap, with phenotypic contributions ranging from 7.03% to 15.16%. Notably, qKR5-10 (215.06–224.10 Mb) in subpopulation 5 is a chromosome 5-specific QTL, showing no overlap with other QTLs, and has a high phenotypic explanatory power of 8.50%.
[0080] Table 5 Significant QTLs for Maize Shelling Rate
[0081]
[0082] 2.4 Combined genome-wide association analysis and QTL analysis to identify candidate genes and analyze their haplotypes
[0083] Genome-wide association analysis showed that, under different environmental conditions, SNP loci significantly associated with seed yield were identified on all 10 maize chromosomes. Figure 4 To screen for reliable QTLs suitable for gene function analysis, a comprehensive analysis of genome-wide association studies (GWAS) and linkage mapping results was performed. Through combined GWAS and linkage mapping analysis, a non-synonymous mutation site, SNP5_222146671, was identified on chromosome 5. This site is located within the physical region of qKR5-10. Figure 5(a, Table 6) The phenotypic explanation rate of this locus was 4.68%, indicating that it has a certain regulatory effect on the seed yield trait. Furthermore, this locus forms a strong linkage block with adjacent SNP loci such as SNP5_222146687 (r²=0.90). Figure 5 (b) This suggests that it may occupy a key position in the gene network. Further gene localization analysis showed that SNP5_222146671 is located on an exon of the gene Zm00001eb257890 (b). Figure 5 c). Haplotype analysis of SNP5_222146671 revealed the existence of two haplotypes, G and A (Figure 5d). Compared to haplotype A, haplotype G showed a higher phenotypic value and a wider frequency across all populations. Figure 5 e). Therefore, haplotype G is considered to be the dominant haplotype regulating maize shelling rate.
[0084] Table 6. Candidate genes co-localized by the two combined methods
[0085]
[0086] 2.5 Preliminary functional analysis of key candidate genes
[0087] Based on SNP effects and gene function annotation, Zm00001eb257890 was proposed as a key candidate gene. To preliminarily explore the effect of Zm00001eb257890 expression in ears of parents with high and low shelling rates, qRT-PCR was used to detect the expression level of this gene in the mid-section and tip tissues of ears during development in six parents. Figure 6 a). qPCR results showed that, at different stages, taking the parent with the lowest seed yield, YML226, as the control, the expression level of Zm00001eb257890 in each parental line was negatively correlated with both ear length and ear diameter, while ear length and ear diameter showed a stable positive correlation, but neither was significant (Fig. 6b). Furthermore, the relative expression level of Zm00001eb257890 in the middle ear tissue of each parental line was higher than that in the tip tissue (Fig. 7). Notably, after stage V9, the expression level of Zm00001eb257890 showed a decreasing trend. Regarding the expression differences among parents, except for stages V9 and 0_DAP, the expression level of Zm00001eb257890 in the parents with lower seed yields, Chang7-2 and Shen137, was significantly higher than that in the parents with higher seed yields, YML1218 and Q11, and the differences reached a highly significant level or above. Figure 8 ).
[0088] 2.6 Genetic variation analysis of Zm00001eb257890
[0089] To further verify whether Zm00001eb257890 has undergone mutations in its parents Chang7-2 and Shen137, a genetic structure analysis was performed. The protein encoded by Zm00001eb257890 is a serine / threonine protein kinase VPS15 isoform. This protein consists of multiple functional domains, including a protein kinase domain, a HEAT repeat sequence, and multiple WD40 domains. Figure 9 a). By comparing the base sequences and analyzing the amino acid variations of Zm00001eb257890 in the six parents, a G / A nonsynonymous mutation was found in the CDS region of Zm00001eb257890 in Shen137 and Chang7-2, resulting in the substitution of valine at position 1322 of the encoded amino acid sequence with isoleucine (V... 1322 I)( Figure 9 b). Further analysis showed that the nonsynonymous mutation at position 1322 of SNP5_222146671 did not occur directly within the WD40 domain (positions 1329-1373), but rather in a nearby location. Furthermore, in the VPS15 protein tertiary structure prediction model, this site is downstream of the serine / threonine phosphorylation site. Figure 9 c). SMART prediction revealed this amino acid site (V 1322 I) Located in an inherently disordered region, it is speculated that it may be a key amino acid affecting the formation of phase transition.
[0090] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. The application of a reagent for detecting SNP molecular markers related to maize seed yield in marker-assisted breeding of maize seed yield, characterized in that, The SNP molecular marker is located at base 222146671 on chromosome 5 of the maize genome. The SNP site has two haplotypes, G and A, with haplotype G being the dominant haplotype for maize seed yield. The reference version of the maize genome is B73 RefGen_v5.
2. An application of a kit for detecting SNP molecular markers related to maize seed yield in screening for high maize seed yield, characterized in that, The SNP molecular marker is located at base 222146671 on chromosome 5 of the maize genome. The SNP site has two haplotypes, G and A, with haplotype G being the dominant haplotype for maize seed yield. The reference version of the maize genome is B73 RefGen_v5.