Use of the gene zm00001eb090970 in high-protein corn breeding

CN122521882APending Publication Date: 2026-08-07FOOD CROPS RES INST YUNNAN ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FOOD CROPS RES INST YUNNAN ACAD OF AGRI SCI
Filing Date
2026-04-22
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

由于SNP 与 SV 之间的连锁不平衡(LD)关系往往较弱,现有的 SNP 图谱难以有效标记这些大尺度变异,导致大量具有潜在育种价值的 SV 仍未被充分解析

Benefits of technology

本发明利用温带优良玉米骨干自交系 Ye107 作为共同父本,以及14个在遗传距离、生态类型及籽粒蛋白含量上具有显著差异的代表性自交系作为供体母本进行杂交,构建了一个具有广泛遗传差异的大型玉米多亲群体。利用SV-GWAS分析挖掘到位于2号染色体的调控籽粒蛋白质含量的功能基因Zm00001eb090970。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122521882A_ABST
    Figure CN122521882A_ABST
Patent Text Reader

Abstract

The application relates to the field of corn molecular marker assisted breeding, and specifically discloses application of a Zm00001eb090970 gene in high-protein corn breeding, wherein the Zm00001eb090970 gene sequence is shown as SEQ ID NO:1, the application mines a functional gene Zm00001eb090970 located on a chromosome 2 by using SV-GWAS analysis, and determines an excellent haplotype Hap2 for high protein, the above natural excellent variation significantly improves the transport efficiency of nitrogen substrates of plants by remodeling protein pore gate kinetics and transcription regulation double mechanisms, and further positively regulates the accumulation of grain protein. Based on the specific molecular marker combination, the application provides a set of efficient detection products and screening methods, can accurately identify high-protein corn germplasm in early breeding, effectively breaks the genetic negative correlation between yield and quality, and greatly improves the genetic improvement efficiency of high-quality corn.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular marker-assisted breeding of maize, specifically to the application of the Zm00001eb090970 gene in the breeding of maize with high protein content in kernels. Background Technology

[0002] As a globally important food and feed crop, the protein content of maize kernels directly affects its nutritional value and processing quality. In the history of modern high-yield maize breeding, a persistent negative correlation has existed between yield and quality; that is, increased kernel yield is often accompanied by a significant decrease in protein content. The root of this yield-quality trade-off lies in the resource competition of carbon and nitrogen metabolic flows at the crop physiological level, and the internal physiological bottlenecks that limit the efficient transport of nitrogen from source organs to the kernel. With the increasing demand for high-protein maize in modern agriculture, how to increase protein content while maintaining high yield has become a major challenge that urgently needs to be overcome in the field of plant biotechnology. While some progress has been made in elucidating the genetic architecture of kernel protein content, many genetic factors driving phenotypic evolution (especially large-scale structural variations) have not yet been effectively discovered. Existing molecular breeding techniques still lack efficient functional molecular markers that can directly reflect and regulate high-protein traits, greatly limiting the targeted genetic improvement of high-protein maize.

[0003] Deciphering the genetic architecture of grain protein content is a prerequisite for achieving this biofortification goal. Previous genome-wide association studies (GWAS) based on single nucleotide polymorphisms (SNPs) have made significant progress. For example, the classic bZIP transcription factor Opaque2 has been confirmed as a key hub regulating zein synthesis; the recently discovered high-protein gene THP9 from rue grass significantly increases protein content by enhancing nitrogen assimilation; Chen et al. (2023) revealed how the endosperm-specific transcription factor ZmNAC128 coordinates carbon and nitrogen balance at the sink end by synergistically activating sugar transporter and storage protein genes. However, despite our in-depth understanding of the regulatory mechanisms at the sink end, the genetic basis of how nitrogen efficiently "flows" from source organs to these sinks remains incomplete. Furthermore, although the discovery of these key sites is a milestone, the phenomenon of "lost heritability" remains significant; relying solely on SNPs often only explains limited phenotypic variation, and a large number of genetic factors driving phenotypic evolution remain hidden in the dark matter of the genome.

[0004] This limitation stems primarily from the high complexity of the maize genome. As an ancient tetraploid species, the maize genome is replete with long terminal repeat (LTR) retrotransposons and complex structural variations (SVs, including large-scale insertions / deletions, inversions, and translocations). Recent pan-genome studies reveal that SVs cover a much larger total number of genomic bases than SNPs and often act as "causal variants," driving phenotypic evolution by reshaping the cis-regulatory landscape or altering gene copy number dosage. Because the linkage disequilibrium (LD) relationship between SNPs and SVs is often weak, existing SNP maps are insufficient to effectively label these large-scale variations, resulting in a large number of potentially valuable SVs remaining unresolved. Therefore, constructing a graph pangenome-based analytical framework that integrates SVs into GWAS is a necessary strategy for capturing "hidden variations" and comprehensively analyzing the protein content genetic architecture.

[0005] Maize kernel protein content is a quantitative trait, and its expression depends not only on the role of individual genes but also on the interactions between genes (epistesis). The contribution of genes to the trait may vary under different environmental conditions. By screening multiple genes, gene combinations that exhibit stable performance under various environmental conditions can be identified, thereby improving crop adaptability and stability. This invention aims to use whole-genome sequencing (WGS) technology to perform genome-wide association analysis (GWAS) on a multi-parent population of tropical maize to identify novel genes related to kernel protein content. Haplotype analysis will be used to screen for tag SVs, reducing redundant testing costs during GS or MAS breeding processes while preserving the integrity of genetic information. The goal is to provide innovative genetic resources and precise improvement strategies for significantly increasing maize kernel protein content, thereby promoting the optimization and efficiency of maize production. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention aims to provide molecular markers located on chromosome 2 that regulate the protein content of maize kernels and their applications. It aims to identify the functional gene Zm00001eb090970, which is closely associated with high protein content in maize kernels, and to clearly determine the core mutation combination of superior high-protein haplotypes. This provides precise marker targets for molecular marker-assisted selection of maize with high kernel protein content, thereby enabling more efficient screening of germplasm with the target trait.

[0007] To achieve the above objectives, the present invention provides the following technical solution: This invention provides the application of the Zm00001eb090970 gene as a molecular marker in molecular marker-assisted selection breeding of high protein content maize kernels. The Zm00001eb090970 gene sequence is shown in SEQ ID NO:1.

[0008] This invention also provides an application of a product that detects the Zm00001eb090970 gene in molecular marker-assisted selection breeding of high-protein maize kernels. The Zm00001eb090970 gene sequence is shown in SEQ ID NO:1.

[0009] Furthermore, the product detects the expression level of the gene, and the expression level of the gene is positively correlated with the protein content of corn kernels.

[0010] Furthermore, the product detects the haplotype of the Zm00001eb090970 gene and its regulatory regions. The specific haplotype locus is based on the Zm-B73-REFERENCE-NAM-5.0 reference genome. The haplotype locus is determined when the haplotype is located on chromosome 2 of the reference genome, starting from the 5' end. 132256790(REF:C / ALT:G)、132256806(REF:C / ALT:T)、132256846(REF:A / ALT:G)、132256886(REF:T / ALT:G)、132256988(REF:C / ALT:T)、132257037(REF:C / ALT:T)、132257154(REF:G / ALT:A)、132257364(REF:C / ALT:G)、132257391(REF:T / ALT:C)、132257394(REF:A / ALT:C)、132257565(REF:A / ALT:G)、132257685(REF:C / ALT:T)、132257712(REF:G / ALT:A)、132257726(REF:T / ALT:A)、132257743(REF:C / ALT:T)、132257912(REF:G / ALT:T)、132258007(REF:C / ALT:T)、132258235(REF:A / ALT:C)、132258258(REF:C / ALT:T)、132258402(REF:A / ALT:G)、132258420(REF:C / ALT:G)、132258688(REF:T / ALT:G)、132258702(REF:C / ALT:T)、132258770(REF:C / ALT:T)、132258877(REF:G / ALT:C)、132259697(REF:T / ALT:A)、132259757(REF:C / ALT:T)、132259763(REF:A / ALT:C)、132259780(REF:T / ALT:G)、132259781(REF:G / ALT:C)、132259795(REF:G / ALT:A)、132259845(REF:C / ALT:T)、132259846(REF:C / ALT:G)、132259857(REF:G / ALT:T)、132259884(REF:T / ALT:C)、132260234(REF:C / ALT:G)、132260332(REF:G / ALT:A)、132260347(REF:G / ALT:T)、132260395(REF:C / ALT:A)、132260418(REF:G / ALT:A)、132260487(REF:G / ALT:A)、132260503(REF:C / ALT:A)、132260560(REF:A / ALT:T)、132260987(REF:C / ALT:T)、132261068(REF:G / ALT:A)、132261179(REF:G / ALT:C)、132261184(REF:A / ALT:G)、132261211(REF:C / ALT:A)、132261266(REF:G / ALT:A)、132261401(REF:G / ALT:C)、132261419(REF:C / ALT:T)、132261422(REF:G / ALT:T)、132261441(REF:G / ALT:A)、132261458(REF:A / ALT:G)、132261879(REF:A / ALT:G)、132262159(REF:C / ALT:T)、132262162(REF:G / ALT:A)、132262215(REF:G / ALT:A)、132262522(REF:T / ALT:C)、132262550(REF:A / ALT:G)、132262680(REF:G / ALT:T)、132263022(REF:C / ALT:A)、132263049(REF:T / ALT:G)、132263062(REF:A / ALT:C)、132263432(REF:C / ALT:T)、132263618(REF:C / ALT:A)、132263714(REF:T / ALT:C)、132264048(REF:C / ALT:T)、132264098(REF:G / ALT:T)、132264216(REF:C / ALT:T)、132264222(REF:T / ALT:G)、132264239(REF:C / ALT:T)、132264393(REF:T / ALT:C)、132264412(REF:G / ALT:A)、132264418(REF:C / ALT:T)、132264427(REF:G / ALT:A)、132264438(REF:G / ALT:A)、132264441(REF:G / ALT:A)、132264446(REF:A / ALT:G)、132264447(REF:T / ALT:C)、132264474(REF:G / ALT:A)、132264486(REF:G / ALT:A)、132264512(REF:A / ALT:G)、132264521(REF:G / ALT:A)、132264533(REF:A / ALT:G)、132264556(REF:T / ALT:C)、132264849(REF:T / ALT:G)、132264915(REF:C / ALT:T)、132264922(REF:C / ALT:T)、132264932(REF:C / ALT:T)、132264967(REF:G / ALT:A)、132265004(REF:C / ALT:T)、132265015(REF:G / ALT:A)、132265070(REF:T / ALT:G)、132265178(REF:A / ALT:SV)、132265250(REF:G / ALT:A)、132265466(REF:G / ALT:A)、132265479(REF:C / ALT:T)、132265499(REF:T / ALT:C)、132265544(REF:C / ALT:T)、132265545(REF:T / ALT:G)、132265549(REF:T / ALT:C)、132265566(REF:C / ALT:T)、132265586(REF:A / ALT:C)、132265589(REF:C / ALT:T)、132265591(REF:T / ALT:A)、132265597(REF:C / ALT:T)、132265639(REF:T / ALT:A)、132265708(REF:A / ALT:T)、132265741(REF:A / ALT:G)、132265825(REF:C / ALT:T)、132265853(REF:T / ALT:C)、132265859(REF:A / ALT:G)、132265942(REF:C / ALT:T)、132265954(REF:T / ALT:G)、132266018(REF:G / ALT:T)、132266051(REF:T / ALT:G)、132266062(REF:T / ALT:G)、132266065(REF:G / ALT:C)、132266066(REF:G / ALT:A)、132266084(REF:G / ALT:A)、132266095(REF:T / ALT:C)、132266145(REF:T / ALT:C)、132266177(REF:C / ALT:T)、132266227(REF:A / ALT:C)、132266238(REF:G / ALT:A)、132266247(REF:A / ALT:G)、132266268(REF:T / ALT:C)、132266273(REF:C / ALT:A)、132266278(REF:G / ALT:T)、132266282(REF:C / ALT:T)、132266359(REF:C / ALT:A)、132266388(REF:G / ALT:A)、132266390(REF:G / ALT:A)、132266405(REF:C / ALT:T)、132266445(REF:C / ALT:G)、132266522(REF:T / ALT:C)、132266559(REF:T / ALT:A)、132266606(REF:A / ALT:C)、132266635(REF:T / ALT:A)、132266641(REF:G / ALT:A)、132266645(REF:G / ALT:A)、132266655(REF:G / ALT:A)、132266663(REF:A / ALT:C)、132266676(REF:G / ALT:A)、132266695(REF:T / ALT:C)、132266708(REF:A / ALT:G)、132266718(REF:C / ALT:G)、132266721(REF:A / ALT:C)、132266733(REF:T / ALT:C)、132266754(REF:C / ALT:T)、132266766(REF:G / ALT:A)、132266795(REF:C / ALT:A)、132266867(REF:T / ALT:C)、132266916(REF:G / ALT:A)、132266920(REF:A / ALT:G)、132266928(REF:T / ALT:C)、132266948(REF:T / ALT:A)、132266959(REF:T / ALT:C)、132267009(REF:A / ALT:C)、132267026(REF:A / ALT:C)、132267036(REF:T / ALT:C)、132267071(REF:C / ALT:T)、132267076(REF:C / ALT:T)、132267114(REF:C / ALT:A)、132267118(REF:G / ALT:T)、132267126(REF:T / ALT:A)、132267144(REF:G / ALT:A)、132267162(REF:G / ALT:A)、132267168(REF:G / ALT:T)、132267249(REF:G / ALT:C)、132267260(REF:T / ALT:C)、At positions 132267306 (REF:A / ALT:G), 132267316 (REF:T / ALT:C), 132267361 (REF:A / ALT:G), 132267382 (REF:T / ALT:C), 132267386 (REF:G / ALT:A), and 132267417 (REF:G / ALT:T), when the base sequence is CTATTTGGCCACGATGCCTAGGCCCTCCGGGTGTCGATAAAATTACGAACTTAGGTAACGGAGCTACTTTTTCATAAGGCAAGAGCGTTTATAGAAATCTGCTCCATATGTCGTGTGGCAACCTCAGTATCAAATGCACAAAACACGGCCCGCTAGCACCCCTTCTAAATCCGCGCAT, maize exhibits the dominant trait of high grain protein content.

[0011] Furthermore, the products include kits, reagents, or gene chips.

[0012] Furthermore, the product is prepared using qPCR or dPCR.

[0013] Furthermore, the product is prepared using Sanger sequencing, high-throughput sequencing, fluorescence in situ hybridization, TaqMan probe method, ARMS-PCR method, or KASP method.

[0014] This invention also provides a method for screening maize germplasm with high protein content in the kernels. The method involves taking a maize sample to be tested, detecting the gene expression level or haplotype in the application, and screening for germplasm that meets the requirements, which is maize germplasm with high protein content in the kernels.

[0015] The term "molecular marker-assisted selection (MAS)" is a breeding technique that uses molecular markers of a target trait to select offspring lines in order to obtain superior individual plants containing the target gene.

[0016] The term "Genomic Selection (GS)" is a modern breeding technique that uses whole-genome marker information for genetic evaluation and selection. It aims to accelerate the breeding process and improve selection efficiency by predicting the breeding value or phenotypic performance of individuals through high-density molecular markers. The term "GS-MAS" refers to a breeding strategy that combines genomic selection (GS) with marker-assisted selection (MAS). GS-MAS uses high-density marker information and phenotypes covering the entire genome to estimate the breeding value of an individual, while also associating major and minor genes. By using breeding values, it can predict and select for complex traits (low heritability, difficult to determine, etc.) at an early stage, thereby shortening the generation interval, accelerating the breeding process, improving selection accuracy, and saving costs.

[0017] The technical effects achieved by this invention are as follows: This invention utilizes the superior temperate maize inbred line Ye107 as a common paternal parent and 14 representative inbred lines with significant differences in genetic distance, ecological type, and grain protein content as donor maternal parents for hybridization, constructing a large maize multi-parent population with extensive genetic diversity. SV-GWAS analysis revealed the functional gene Zm00001eb090970, located on chromosome 2, which regulates grain protein content.

[0018] Haplotype analysis of the Zm00001eb090970 gene and its flanking regulatory regions (where the gene sequence itself is shown in SEQ ID NO:1, corresponding to positions 132256747-132258853 on chromosome 2; the target region used for haplotype analysis covers this sequence and its flanking amplified regions, from position 132256747 to position 132267417) showed that in a population containing 2742 RILs, this region mainly formed three haplotypes (Hap1~Hap3). Samples with complete whole-sequence sequencing and phenotypic data were grouped by haplotype, and the best linear unbiased predictive value (BLUP) of maize kernel protein content was calculated. The analysis results showed that the superior haplotype Hap2 exhibited a significant positive phenotypic trend in the population, with a mean BLUP of 1.26. In contrast, the most prevalent common haplotype, Hap1, had a low mean phenotype (Mean BLUP = -0.24). Furthermore, a very small number of families carrying Hap3 were detected in the population, but they were not included in the box plot analysis of phenotypic differences due to phenotypic absence in their samples. This data, supported by a large sample size, confirms that Hap2 is the core dominant haplotype positively regulating high protein content in maize kernels (Figure 5a).

[0019] In-depth sequence variation analysis confirmed that Zm00001eb090970 exhibits a clear coding region-promoter dual differentiation regulatory pattern. This invention identified a tightly linked variant cluster in the coding region containing three non-synonymous mutations (chr2:132257364, chr2:132257391, and chr2:132257394). Figure 5 c), these mutations led to significant amino acid substitutions (p.Q105E, p.S114P, p.K115Q), confirming a fundamentally superior differentiation in structure and transport activity between the proteins encoded by Hap1 and Hap2. Furthermore, key variations in the promoter region were completely co-separated from the aforementioned coding region variations and SV-GWAS signaling, confirming that the transcriptional abundance of this gene is simultaneously and precisely positively regulated. Figure 5 b).

[0020] The results of this invention contribute to further research on the regulatory mechanism of maize kernel protein content, and provide more stable and accurate markers for genetic resource identification, MAS, genetic map construction, linkage mapping and GS-MAS in the breeding application of maize kernel protein content trait. Attached image description: Figure 1 Phenotypic variation and correlation analysis diagram. (a) Variation in grain protein content of 14 populations under 3 environments. (b) Pearson correlation coefficient among the 3 environments.

[0021] Figure 2 Genetic variation density and genome-wide linkage disequilibrium (LD) decay plot. (a) Structural variation (SV) density (window size 0.1 Mb); (b) Genome-wide LD decay based on SV.

[0022] Figure 3 Population structure of 2742 maize recombinant inbred lines and genetic differentiation and diversity of 14 NAM subpopulations. (a) Principal component analysis (PCA) shows the clustering of recombinant inbred lines by parental ecotype; (b) Neighbor-joining (NJ) phylogenetic tree shows the genetic relationships among recombinant inbred lines; (c) ADMIXTURE analysis at K=14 shows the proportion of ancestral components in each recombinant inbred line; (d) Heatmap of pair fixation index (Fst) among subpopulations; (e) Violin plot shows the distribution of nucleotide diversity (π) within each subpopulation.

[0023] Figure 4SV-GWAS plot of maize kernel protein content. (a) Manhattan plot (left) and QQ plot (right) of kernel protein content under 21JH, (b) 21YS, (c) 22YS and (d) BLUP values.

[0024] Figure 5 Functional analysis diagram of Zm00001eb090970. (a) Box plot of grain protein content of haplotype; (b) Schematic diagram of chromosome linkage related to ZmPAD4; (c) Dual variant module showing differences in coding region and promoter region; (d) Prediction of protein three-dimensional structure.

[0025] Figure 6. Spatiotemporal expression patterns of Zm00001eb090970 in different maize parent lines. Relative expression levels of Zm00001eb090970 in (a) ear leaves and (b) kernels at 15 days (blue), 20 days (orange), and 25 days (green) post-pollination. Data are expressed as mean ± standard error (SEM) (n=3). An asterisk indicates a statistically significant difference compared to the control material Ye107: *P<0.05; **P<0.01; ***P<0.001.

[0026] Figure 7 A schematic diagram of mutation sites that form different haplotypes in the Zm00001eb090970 gene. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention.

[0028] The gene sequence of Zm00001eb090970, as labeled in this invention, is shown in SEQ ID NO: 1, corresponding to positions 132256747-132258853 on chromosome 2 of the reference genome version Zm-B73-REFERENCE-NAM-5.0, starting from the 5' end. The actual target region for haplotype detection covers this sequence (SEQ ID NO: 1) and its flanking amplification regions (including promoters and regulatory sequences), from position 132256747 to position 132267417 on chromosome 2.

[0029] Example 1 1. Experiment 1.1 Plant Materials and Experimental Design This invention selected Ye107, a superior temperate maize backbone inbred line, as the common paternal parent, and 14 representative inbred lines with significant differences in genetic distance, ecological type, and grain protein content as donor maternal parents. These 14 donor parents mainly originated from tropical and subtropical germplasm resources, covering a wide range of phenotypic variations from high protein (e.g., YML226, 13.0%; CML312, 12.31%) to low protein (TRL418, 9.49%) (Table 1), ensuring allele richness in the offspring population. F1 generations were obtained by crossing Ye107 with each donor parent, followed by continuous self-pollination using the Single-Seed Descent (SSD) method up to F9, ultimately constructing a large NAM population containing 14 subpopulations and a total of 2,742 recombinant inbred lines (RILs) (Table 1).

[0030] Table 1 Parental Information To comprehensively assess the impact of environmental effects on phenotypes and obtain stable measurement data, the experimental materials of this invention were subjected to planting trials under three different natural environments (consisting of different planting locations and years), specifically including: 1) Planted in Jinghong City, Yunnan Province in 2021 (tropical climate, abbreviated as 21JH); 2) Planted in Yanshan County, Yunnan Province in 2021 (subtropical climate, abbreviated as 21YS); 3) Planted in Yanshan County, Yunnan Province in 2022 (abbreviated as 22YS).

[0031] The experiment followed a completely randomized block design (RCBD), with three replicates at each location to effectively control for spatial heterogeneity in the field microenvironment. After grain harvest, the grain protein content of all 2742 families was determined using a high-throughput cereal near-infrared spectrometer. To ensure data accuracy, each family underwent nine independent scans, and the average value was used as the phenotypic observation for that family.

[0032] 1.2 Phenotypic Data Analysis All statistical analyses were performed in the R statistical environment (v4.3.2). To accurately assess genetic effects and eliminate environmental noise, this invention constructed a mixed linear model (MLM) using the lme4 package in R software (v4.3.2). In this model, genotype, environment (defined as a combination of location and year), and their interaction (G×E) were set as random effects. Subsequently, the ranef() function was used to extract the BLUP values ​​corresponding to each genotype from the model. These BLUP values ​​represent stable genetic evaluation values ​​across tropical and subtropical environments after removing environmental variance, and were used as single phenotypic inputs for subsequent GWAS. Generalized heritability (H2) was estimated based on variance components, using the following formula: H 2 in, , and represents the genotype variance, G×E interaction variance, and residual variance, respectively; e and r are the environmental number and the number of replicates, respectively.

[0033] 1.3 Whole genome sequencing 1.3.1 Sample Collection and DNA Extraction Healthy leaves were sampled during the seedling to jointing stage of maize. The collected leaf samples were immediately frozen in liquid nitrogen to prevent tissue and DNA degradation. Leaf DNA was extracted using the Tiangen Plant Genomic DNA Extraction Kit (Tiangen Biotech Co., Ltd., Beijing, China). The extracted DNA samples were quality controlled using agarose gel electrophoresis to confirm concentration (above 100 ng / μl) and purity (A260 / A280 ratio should generally be between 1.8 and 2.0) to ensure the extracted DNA was suitable for subsequent whole-genome sequencing.

[0034] 1.3.2 Library Construction and Sequencing DNA fragmentation was performed using a Vibra-Cell series ultrasonic fragmentation instrument, and the fragmented DNA samples were subjected to gel electrophoresis to assess quality (ideally, the DNA fragment size should be between 200-500 bp). End repair was performed using T4 DNA polymerase, and short DNA adapters were ligated to both ends of the DNA fragments. To ensure library quality, DNA fragments without ligated adapters were removed using Exonuclease I. DNA fragments of 200-500 bp length (including portions of the adapter sequence) were selected using the Agencourt AMPure XP magnetic bead method. PCR amplification was performed on the library to increase the number of target DNA fragments. Finally, the library concentration was quantified using Qubit, typically targeting a concentration between 10-20 nM. The size distribution of the library was analyzed using Bioanalyzer, with fragment sizes typically between 200-500 bp.

[0035] All samples were sequenced using the Illumina NovaSeq 6000 platform in paired-end sequencing mode at a depth of 5×. The sequencing results were stored in FASTQ format. The quality of the sequencing data was evaluated using FastQC, with the following criteria: (1) Q-score ≥ 20, (2) GC content of the genome is 40% to 60%, (3) repetitive sequence ratio is less than 30%, (4) N content is less than 1-2%, and (5) read length matches the expected length.

[0036] 1.4 Group Stratification Analysis 1.4.1 Principal Component Analysis To assess population genetic structure and control for the impact of population stratification on GWAS results, we performed principal component analysis (PCA) using GCTA software. First, genotypic data were standardized, and a genetic similarity matrix between samples was calculated. Then, the first few principal components were extracted based on this matrix; these principal components characterize the population structure.

[0037] 1.4.2 Construction of Rootless Tree To further analyze the genetic relationships among subpopulations, we constructed an unrooted tree using MEGA (Molecular Evolutionary Genetics Analysis) software to assess the genetic distance and phylogenetic relationships between subpopulations. During the construction process, we used the Neighbor-Joining (NJ) method, which efficiently and accurately reflects the genetic differentiation of a population by minimizing the sum of genetic distances between samples.

[0038] 1.4.3 Population genetic structure analysis We used ADMIXTURE software for population stratification analysis. This method infers the contribution proportion of different subpopulations to each individual's genome by assuming that each individual is a mixture of components from multiple subpopulations. We first selected and set the expected subpopulation size (K value) based on our research objectives. By trying different K values ​​and selecting the optimal K value through cross-validation, we avoided overfitting or underfitting. We analyzed the population structure under different K values ​​and determined the genetic background of the samples based on the population stratification pattern.

[0039] 1.4.4 Analysis of genetic diversity and differentiation To quantify genetic diversity within populations and genetic differentiation between populations, this invention calculates nucleotide diversity (Pi) and the fixation index (Fst) based on high-quality SNP VCF files. Using VCFtools software, the genome-wide average Pi value within 14 populations was calculated using the method of Nei and Li (1979), and the genome-wide average Fst value between all pairs of the 14 populations was estimated using the method of Weir and Cockerham (1984).

[0040] 1.4.5 Analysis of the density distribution of variant sites and linkage disequilibrium analysis The number of variants within a fixed window on each chromosome was counted using VCFtools software (v0.1.16), with the statistical window for SVs set to 0.1 Mb. Density distribution maps were then plotted using the pheatmap package (v1.0.12) in R language (v4.3.2) https: / / CRAN.R-project.org / package=pheatmap.

[0041] The linkage disequilibrium correlation coefficient (r) of SV was calculated using the PopLDdecay (v3.41) software. 2 The calculation parameters were set to a maximum physical distance of 500 kb. Subsequently, the average r of the entire genome was plotted using R. 2 The decay curve of the value as physical distance increases was plotted, and r was statistically analyzed. 2 The physical distances corresponding to the decay to different thresholds (such as 0.1 and 0.2) are used as the basis for evaluating the degree of LD decay.

[0042] 1.5 Genome-wide association analysis To identify significantly associated loci, GWAS analysis was performed on high-quality SVs using a mixed linear model (MLM) in GEMMA (v0.98.5) software. The model was adjusted for population structure (Q matrix, first 3 principal components) and phylogenetic relationships (K matrix). y=Xβ+Sα+Qυ+Zu+e Where y is the BLUP phenotype vector, X and S are the fixed effects and marker effects matrices, respectively, and Q and Z are the population structure and polygenic background effects matrices, respectively. The significance threshold was set at the Bonferroni correction level (p = 0.05 / Neff, where Neff is the number of effective markers).

[0043] 1.6 Candidate gene haplotype analysis Based on resequencing genotype data, key structural variants (SVs) identified by GWAS were integrated as independent genetic markers for inclusion in the analysis. Strict quality control was performed using VCFtools (missing rate < 0.1, MAF > 0.05). Haplotypes were constructed using the GeneHapR package in R. Subsequently, each haplotype group was correlated with the best linear unbiased predictor (BLUP) of maize kernel protein content, and the significance of differences among major haplotypes was assessed using Student's t-test (P < 0.05). Data visualization was performed using ggplot2.

[0044] 1.7 Sequence alignment and amino acid variation analysis SnpEff software was used to perform functional annotation of variant sites, and non-synonymous variants located in the coding region (CDS) and high-quality variants in the promoter region were screened. The CDS sequences of the B73 reference genome and parental genome were extracted, and multiple sequence alignment was performed using MAFFT to identify key sites leading to amino acid changes and assess their potential impact on protein structure.

[0045] 1.8 RNA Extraction and Real-Time Quantitative PCR (qRT-PCR) Analysis Total RNA was extracted from ear leaves and kernels of 15 maize inbred lines at 15, 20, and 25 days post-pollination (DAP) using TRIzol reagent. First-strand cDNA was synthesized using the HiScript III first-strand cDNA synthesis kit (Vazyme Biotech, China). Real-time quantitative PCR (qRT-PCR) amplification was performed on a Bio-Rad CFX96 real-time quantitative PCR instrument using the SuperReal PreMix Plus kit (Tiangen Biotech, China). Gene-specific primers were designed for the target gene and the internal reference gene LUG (Table 2). Stable-expressed LUG was used as the internal reference gene for data normalization. The qRT-PCR amplification program was: 95℃ pre-denaturation for 15 min; followed by 95℃ denaturation for 10 s, 54–58℃ annealing for 20 s, and 72℃ extension for 30 s, for a total of 40 cycles. Three technical replicates were set for each sample. Two... -ΔΔCt The relative expression level of the target gene was calculated using a method, and statistical analysis was performed using a t-test.

[0046] Table 2 Primer sequences for qRT-PCR of candidate genes 2. Results 2.1 Phenotypic Data Analysis Phenotypic analysis of 14 RIL subpopulations under three environments revealed extensive phenotypic variation in grain protein content among different subpopulations (Figure 1a, Table 3). Although the overall distribution was similar among the subpopulations, certain genetic differences remained. The average grain protein content was relatively high in RIL-NK40-1 (9.87%), while it was relatively low in RIL-Q11 (8.85%). Analysis of variance showed that genotype, environment, and interaction effects were all highly significant (P < 0.0001) (Table 4). Despite significant G × E interaction, a very strong positive correlation was observed between environments (r = 0.956 - 0.980) (Figure 1b), and broad-sense heritability (H... 2 The percentage is as high as 98% or more.

[0047] Table 3 Phenotypic Data Analysis of Grain Protein Content Table 4. Analysis of variation in grain protein content (ANOVA) 2.2 Genomic Variation and Linkage Disequilibrium Based on whole-genome resequencing data, this invention identified a total of 180,767 structural variants (SVs). The distribution of SVs across chromosomes showed a significant unevenness. Chromosome 1 contained the most SVs (25,487), while chromosome 9 contained the fewest (11,885). SV density distribution maps showed that most structural variations were enriched in the arms of chromosomes, while the density decreased significantly in the centromere region, consistent with the distribution characteristics of functional regions in the genome. Figure 2 a). The LD mode of SVs exhibits significant complexity and long-distance effects, without showing obvious regular decay (Fig. 2b).

[0048] 2.3 Population genetic structure analysis PCA results showed that the 2742 RILs exhibited a clear clustering trend based on parental ecotype, but there was extensive overlap between populations. Figure 3 a). PC1 (1.39%) primarily reflects the genetic differentiation between temperate and tropical / subtropical germplasm. Although all RILs share the genetic background of the temperate paternal parent Ye107, PC1 clearly reveals the differentiation of ecotypes: populations derived from temperate parents (such as RIL-Q11, RIL-Chang7-2, etc.) cluster on the left, while populations derived from tropical parents (such as the CML series) are mainly distributed in the upper right. PC2 (1.17%) further distinguishes different populations within the same ecotype, for example, clearly separating RIL-Q11 and RIL-Chang7-2 on the vertical axis.

[0049] Phylogenetic tree and population structure analysis further validated the above results. (Evolutionary tree) Figure 3 b) shows that all RILs clustered into 14 major branches, and the deep clustering relationships between branches also reflect the genetic differences between temperate and tropical zones. Population structure analysis ( Figure 3 c) indicates that when K=14, the RILs in each population form independent genetic clusters. However, significant genetic background sharing also exists between the populations. This mixing phenomenon is due, on the one hand, to the fact that all RILs share a common paternal parent, Ye107, and on the other hand, to the fact that some maternal parents themselves have certain kinship relationships.

[0050] In summary, the results of the three analyses corroborate each other, revealing a significant dual-structure characteristic of this population: that is, there is clear genetic differentiation among the 14 populations, but at the same time, they maintain a close genetic connection due to sharing the same paternal line.

[0051] 2.4 Population genetic diversity and differentiation The genome-wide average nucleotide diversity (pi) for all 14 populations was 1.51 × 10⁻⁶.-3 Among them, RIL-YML46 exhibited the highest genetic diversity, pi = 1.83 × 10⁻⁶. -3 RIL-Q11 had the lowest (pi = 1.23 × 10⁻⁶), while RIL-Q11 had the lowest (pi = 1.23 × 10⁻⁶). -3 Considering that these populations originate from tropical and subtropical germplasm, even after multiple generations of self-pollination and purification, they still maintain a relatively rich genetic variation base. Figure 3 e).

[0052] Regarding population differentiation, the paired genetic differentiation index (Fst) ranged from 0.013 to 0.181. Figure 3 d), with an overall average of 0.126. The highest degree of differentiation was observed between RIL-Chang7-2 and RIL-TRL418 (Fst = 0.181), suggesting limited gene flow or strong divergent selection between them. Conversely, the lowest degree of differentiation was observed between RIL-Shen137 and RIL-YML1218 (Fst = 0.013), indicating a closer kinship. These results confirm that the population structure of this invention exhibits significant genetic heterogeneity, reflecting the broad genetic diversity potential of tropical germplasm.

[0053] 2.5 Genome-wide association analysis of seed protein content based on SV As an important supplement to SNP analysis, this invention used SV-GWAS to identify one environmentally stable locus (chr2:132265178) on chromosome 2, independent of the SNP signal, in four environments (including BLUP) (Table 5). Its stability across tropical and subtropical environments confirms the specific role of structural variation in regulating maize kernel protein content. Unlike SNP-GWAS, which detected only weak background noise in this region, SV-GWAS detected a highly significant associated signal at this locus (Avg -log10P = 5.92), and this signal was stably detected in all test environments. Figure 4The presence of an insertion / deletion variant (ad) indicates that the functional gene marked by this structural variation possesses strong environmental stability. Sequence analysis revealed a 1.2 kb insertion / deletion variant in this region, physically located adjacent to the amino acid permease gene ZmAAP3 (Zm00001eb090970). Given the rate-limiting role of AAP family proteins in nitrogen transport from source organs to sink organs (grains), this core structural variation confirms that it directly affects and enhances the transport function and expression activity of Zm00001eb090970 by reshaping genomic structural features, thereby positively regulating the accumulation of protein synthesis substrates. This functional site was not detected in SNP-GWAS, strongly demonstrating that the SV analysis used in this invention can effectively capture key genetic variation information.

[0054] Table 5 Candidate genes co-localized by SV-GWAS 2.6 Candidate gene haplotype analysis, sequence variation analysis, and three-dimensional protein structure analysis Haplotype analysis of the Zm00001eb090970 gene and its flanking regulatory regions (where the gene sequence itself is shown in SEQ ID NO:1, corresponding to positions 132256747-132258853 on chromosome 2; the target region used for haplotype analysis covers this sequence and its flanking amplified regions, from position 132256747 to position 132267417) showed that in a population containing 2742 RILs, this region mainly formed three haplotypes (Hap1~Hap3) ( Figure 5 a).

[0055] Samples with complete whole-sequence sequencing and phenotypic data were grouped by haplotype, and the best linear unbiased predictor value (BLUP) for maize kernel protein content was calculated. Analysis showed that the superior haplotype Hap2 exhibited a significant positive phenotypic trend in the population, with a mean BLUP of 1.26. In contrast, the most prevalent common haplotype, Hap1, had a lower mean phenotypic value, with a mean BLUP of -0.24. Furthermore, a very small number of families carrying Hap3 were detected in the population; however, due to phenotypic deficiencies in their samples, they were not included in the boxplot analysis of phenotypic differences. This data, supported by a large sample size, confirms that Hap2 is the core dominant haplotype positively regulating high protein content in maize kernels (Figure 5a).

[0056] In-depth sequence variation analysis confirmed that Zm00001eb090970 exhibits a clear coding region-promoter dual differentiation regulatory pattern. This invention identified a tightly linked variant cluster in the coding region containing three non-synonymous mutations (chr2:132257364, chr2:132257391, and chr2:132257394). Figure 5 c), these mutations led to significant amino acid substitutions (p.Q105E, p.S114P, p.K115Q), confirming a fundamentally superior differentiation in structure and transport activity between the proteins encoded by Hap1 and Hap2. Furthermore, key variations in the promoter region were completely co-separated from the aforementioned coding region variations and SV-GWAS signaling, confirming that the transcriptional abundance of this gene is simultaneously and precisely positively regulated. Figure 5 b).

[0057] To further elucidate the molecular effects of coding variations, this invention utilizes three-dimensional protein structure analysis to confirm that three non-synonymous mutations specific to the superior haplotype (p.Q105E, p.S114P, p.K115Q) aggregate on the extracellular hydrophilic loop connecting TM2 and TM3. Figure 5 d). In particular, the substitution of serine for proline at residue 114 (p.S114P) confirms that this mutation introduces a rigid conformational distortion in this flexible linker, thereby successfully reshaping the channel-gated dynamics and achieving high-affinity capture and efficient transport of amino acid substrates.

[0058] The results of this invention contribute to further research on the regulatory mechanism of maize kernel protein content, and provide more stable and accurate markers for genetic resource identification, MAS, genetic map construction, linkage mapping and GS-MAS in the breeding application of maize kernel protein content trait.

[0059] The specific haplotype loci, using Zm-B73-REFERENCE-NAM-5.0 as the reference genome, are as follows on chromosome 2 of the genome version, viewed from left to right starting at the 5' end: 132256790 (REF:C / ALT:G), 132256806 (REF:C / ALT:T), 132256846 (REF:A / ALT:G), 132256886 (REF:T / ALT:G), 132256988 (REF:C / ALT:T), 132257037 (REF:C / ALT:T), 132257154 (REF:G / ALT:A), 132257364 (REF:C / A). LT:G), 132257391(REF:T / ALT:C), 132257394(REF:A / ALT:C), 132257565(REF:A / ALT:G), 132257685(REF:C / ALT:T), 132257712(REF:G / ALT:A), 13 2257726(REF:T / ALT:A), 132257743(REF:C / ALT:T), 132257912(REF:G / ALT:T), 132258007(REF:C / ALT:T), 132258235(REF:A / ALT:C), 132258258(R EF:C / ALT:T), 132258402(REF:A / ALT:G), 132258420(REF:C / ALT:G), 132258688(REF:T / ALT:G), 132258702(REF:C / ALT:T), 132258770(REF:C / ALT ) 9781(REF:G / ALT:C), 132259795(REF:G / ALT:A), 132259845(REF:C / ALT:T), 132259846(REF:C / ALT:G), 132259857(REF:G / ALT:T), 132259884(REF: T / ALT:C), 132260234(REF:C / ALT:G), 132260332(REF:G / ALT:A), 132260347(REF:G / ALT:T), 132260395(REF:C / ALT:A), 132260418(REF:G / ALT:A),132260487(REF:G / ALT:A)、132260503(REF:C / ALT:A)、132260560(REF:A / ALT:T)、132260987(REF:C / ALT:T)、132261068(REF:G / ALT:A)、132261179(REF:G / ALT:C)、132261184(REF:A / ALT:G)、132261211(REF:C / ALT:A)、132261266(REF:G / ALT:A)、132261401(REF:G / ALT:C)、132261419(REF:C / ALT:T)、132261422(REF:G / ALT:T)、132261441(REF:G / ALT:A)、132261458(REF:A / ALT:G)、132261879(REF:A / ALT:G)、132262159(REF:C / ALT:T)、132262162(REF:G / ALT:A)、132262215(REF:G / ALT:A)、132262522(REF:T / ALT:C)、132262550(REF:A / ALT:G)、132262680(REF:G / ALT:T)、132263022(REF:C / ALT:A)、132263049(REF:T / ALT:G)、132263062(REF:A / ALT:C)、132263432(REF:C / ALT:T)、132263618(REF:C / ALT:A)、132263714(REF:T / ALT:C)、132264048(REF:C / ALT:T)、132264098(REF:G / ALT:T)、132264216(REF:C / ALT:T)、132264222(REF:T / ALT:G)、132264239(REF:C / ALT:T)、132264393(REF:T / ALT:C)、132264412(REF:G / ALT:A)、132264418(REF:C / ALT:T)、132264427(REF:G / ALT:A)、132264438(REF:G / ALT:A)、132264441(REF:G / ALT:A)、132264446(REF:A / ALT:G)、132264447(REF:T / ALT:C)、132264474(REF:G / ALT:A)、132264486(REF:G / ALT:A)、132264512(REF:A / ALT:G)、132264521(REF:G / ALT:A)、132264533(REF:A / ALT:G)、132264556(REF:T / ALT:C)、132264849(REF:T / ALT:G)、132264915(REF:C / ALT:T)、132264922(REF:C / ALT:T)、132264932(REF:C / ALT:T)、132264967(REF:G / ALT:A)、132265004(REF:C / ALT:T)、132265015(REF:G / ALT:A)、132265070(REF:T / ALT:G)、132265178(REF:A / ALT:SV)、132265250(REF:G / ALT:A)、132265466(REF:G / ALT:A)、132265479(REF:C / ALT:T)、132265499(REF:T / ALT:C)、132265544(REF:C / ALT:T)、132265545(REF:T / ALT:G)、132265549(REF:T / ALT:C)、132265566(REF:C / ALT:T)、132265586(REF:A / ALT:C)、132265589(REF:C / ALT:T)、132265591(REF:T / ALT:A)、132265597(REF:C / ALT:T)、132265639(REF:T / ALT:A)、132265708(REF:A / ALT:T)、132265741(REF:A / ALT:G)、132265825(REF:C / ALT:T)、132265853(REF:T / ALT:C)、132265859(REF:A / ALT:G)、132265942(REF:C / ALT:T)、132265954(REF:T / ALT:G)、132266018(REF:G / ALT:T)、132266051(REF:T / ALT:G)、132266062(REF:T / ALT:G)、132266065(REF:G / ALT:C)、132266066(REF:G / ALT:A)、132266084(REF:G / ALT:A)、132266095(REF:T / ALT:C)、132266145(REF:T / ALT:C)、132266177(REF:C / ALT:T)、132266227(REF:A / ALT:C)、132266238(REF:G / ALT:A)、132266247(REF:A / ALT:G)、132266268(REF:T / ALT:C)、132266273(REF:C / ALT:A)、132266278(REF:G / ALT:T)、132266282(REF:C / ALT:T)、132266359(REF:C / ALT:A)、132266388(REF:G / ALT:A)、132266390(REF:G / ALT:A)、132266405(REF:C / ALT:T)、132266445(REF:C / ALT:G)、132266522(REF:T / ALT:C)、132266559(REF:T / ALT:A)、132266606(REF:A / ALT:C)、132266635(REF:T / ALT:A)、132266641(REF:G / ALT:A)、132266645(REF:G / ALT:A)、132266655(REF:G / ALT:A)、132266663(REF:A / ALT:C)、132266676(REF:G / ALT:A)、132266695(REF:T / ALT:C)、132266708(REF:A / ALT:G)、132266718(REF:C / ALT:G)、132266721(REF:A / ALT:C)、132266733(REF:T / ALT:C)、132266754(REF:C / ALT:T)、132266766(REF:G / ALT:A)、132266795(REF:C / ALT:A)、132266867(REF:T / ALT:C)、132266916(REF:G / ALT:A)、132266920(REF:A / ALT:G)、132266928(REF:T / ALT:C)、132266948(REF:T / ALT:A)、132266959(REF:T / ALT:C)、132267009(REF:A / ALT:C)、132267026(REF:A / ALT:C)、132267036(REF:T / ALT:C)、132267071(REF:C / ALT:T)、132267076(REF:C / ALT:T)、132267114(REF:C / ALT:A)、132267118(REF:G / ALT:T)、132267126(REF:T / ALT:A)、132267144(REF:G / ALT:A)、132267162(REF:G / ALT:A)、132267168(REF:G / ALT:T), 132267249(REF:G / ALT:C), 132267260(REF:T / ALT:C), 132267306(REF:A / ALT:G), 132267316(R EF:T / ALT:C), 132267361(REF:A / ALT:G), 132267382(REF:T / ALT:C), 132267386(REF:G / ALT:A), 132267417(REF:G / ALT:T), Hap1:CCATCCGCTAACGTTTCCTAGGCTCATATCATGGCGATAAAATTACGAACTTAGGTAACGGAGCTACTTTGTCATAAGGCAAGAGCGTTTATAGAAATCTGCTCCATATGTCGTGTGGCAACCTCAGCATCAAATGCACAAAACACGGCCCACTAGCACCCCTTCTAAATCCGCGCAT Hap2:CTATTTGGCCACGATGCCTAGGCCCTCCGGGTGTCGATAAAATTACGAACTTAGGTAACGGAGCTACTTTTTCATAAGGCAAGAGCGTTTATAGAAATCTGCTCCATATGTCGTGTGGCAACCTCAGTATCAAATGCACAAAACACGGCCCGCTAGCACCCCTTCTAAATCCGCGCAT Hap3:CTATTTGGCCACGATGCCTAGGCTCATATCATGGCGATAAAATTACGAACTTAGGTAACGGAGCTACTTTTTCATAAGGCAAGAGCGTTTATAGAAATCTGCTCTATAAACTATGTGGCAACCTCAGCATCAAATGCACAAAACACGGCCCGCTAGCACCCCTTCTAAATCCGCGCAT 2.7 Spatiotemporal expression validation of candidate genes To deeply elucidate the molecular mechanism by which the candidate gene Zm00001eb090970 regulates grain protein content, this invention employs real-time quantitative PCR (qRT-PCR) to analyze the expression patterns of this gene in 15 maize inbred lines with different genetic backgrounds at different tissue (ear leaf, grain) and developmental time (15 DAP, 20 DAP, 25 DAP).

[0060] (1) Dynamic time-series switching of "source-sink" that matches the characteristics of nitrogen transport According to gene function annotation, Zm00001eb090970 is involved in the transmembrane transport of nitrogen substrates in plants. qRT-PCR results revealed a highly precise spatiotemporal relay expression pattern between the "source" (leaves) and "sink" of this gene. During the early grain-filling stage (15 DAP), the gene was first activated in the panicle leaves, which serve as the nitrogen supply source (e.g., the relative expression levels of the high-potential parents CML395 and YML46 reached 6.10 times and 5.34 times that of the control line Ye107, respectively), while it was generally under transcriptional repression in the grains at the same time (relative quantification was generally below 0.6). As grain filling progressed to the core stage (20 DAP), the expression in the panicle leaves gradually stabilized, while the transcriptional activity at the grain end showed an explosive upregulation. This dynamic switching of "source-end activation followed by sink-end activation" physiologically confirms the gene's key transporter function in driving the efflux of organic nitrogen from leaves and the nitrogen reception in grains. Figure 6 ab).

[0061] (2) The grain filling core stage exhibits strong and abundant transcriptional driving force. During the critical window of 20 DAP, which determines the massive influx of nitrogen into grains, this gene exhibits extremely rich transcriptional variation in the grains of the parental population, providing a strong potential for nutrient uptake in the offspring population. Experimental results show that materials with specific genetic backgrounds exhibit remarkable transcriptional activity during this stage; for example, inbred lines NK40-1 (11.13%), YML46 (10.93%), and R-2-1-1 (11.69%) showed their relative expression levels in 20 DAP grains soaring to 37.29 times, 16.41 times, and 10.72 times that of the control line Ye107, respectively. This extremely high expression provides a sufficient reserve of channel proteins for transmembrane nitrogen uptake substrates in grains. Figure 6 ab).

[0062] (3) Breaking through the limitations of single expression level, confirming the dual mechanism of transcription-gating. Most importantly, the expression profile data objectively show that the transcriptional abundance of this gene at 20 DAP is not a simple linear equivalence with the final static protein content of the parent (for example, the expression level of the extremely high-protein parent YML226 remains steady, while the low-protein parent TRL418 also has a high expression peak). This seemingly non-linear phenotype-expression correlation is the direct molecular evidence for the core innovation of this invention—"the dual mechanism of reshaping protein pore gating dynamics and transcriptional regulation".

[0063] These results indicate that achieving extremely high protein phenotypes depends not only on a transcriptional burst at a specific time (providing a large number of channels), but more importantly, on the optimized pore-gating dynamics conferred by the naturally superior variant (Hap2) (enhancing the permeability efficiency of individual channels). In breeding practice, recombination and aggregation of alleles containing high transcriptional activity and superior pore structure (Hap2) are necessary to break through the rate-limiting bottleneck of a single mechanism and endow offspring with superior nitrogen transport efficiency. This objective finding effectively eliminates the limitations of assessing the value of transport proteins based solely on single expression levels or single phenotypes, establishing the core application status of this gene and its specific haplotypes in high-quality molecular breeding.

[0064] In summary, the Zm00001eb090970 gene and its upstream regulatory regions can be used as molecular markers for identifying maize kernel protein content. Allele-specific markers can be developed, allowing for direct detection of crop genotypes and accurate identification of the presence or absence of target genes or traits. This avoids the blindness and uncertainty inherent in traditional selection, ensuring the accuracy of offspring selection. In breeding, these markers can be applied to genetic resource identification, MAS (Magnetic Image Processing), genetic mapping, linkage mapping, and GS-MAS (Graduate-Growth Modeling). MAS enables rapid and accurate identification of crop genotypes, avoiding the tediousness, time-consuming nature, and screening difficulties of traditional selection, significantly improving selection efficiency and accelerating the breeding process.

[0065] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. The application of the Zm00001eb090970 gene as a molecular marker in marker-assisted selection breeding of high-protein maize kernels, characterized by, The gene sequence Zm00001eb090970 is shown in SEQ ID NO:

1.

2. The application of products containing the Zm00001eb090970 gene in marker-assisted selection breeding of high-protein maize kernels, characterized by: The gene sequence Zm00001eb090970 is shown in SEQ ID NO:

1.

3. The application according to claim 2, characterized in that, The product detects the expression level of the gene, and the expression level of the gene is positively correlated with the protein content of corn kernels.

4. The application according to claim 2, characterized in that, The product detects the haplotype of the Zm00001eb090970 gene and its regulatory regions. Specific haplotype loci are determined using Zm-B73-REFERENCE-NAM-5.0 as a reference genome. The haplotype loci are defined on chromosome 2 from the 5' end as follows: 132256790 (REF:C / ALT:G), 132256806 (REF:C / ALT:T), 132256846 (REF:A / ALT:G), 132256886 (REF:T / ALT:G), 132256988 (REF:C / ALT:T), 132257037 (REF:C / ALT:T), and 132257154 (REF:G / A). LT:A), 132257364(REF:C / ALT:G), 132257391(REF:T / ALT:C), 132257394(REF:A / ALT:C), 132257565(REF:A / ALT:G), 132257685(REF:C / ALT:T), 13 2257712(REF:G / ALT:A), 132257726(REF:T / ALT:A), 132257743(REF:C / ALT:T), 132257912(REF:G / ALT:T), 132258007(REF:C / ALT:T), 132258235(R EF:A / ALT:C), 132258258(REF:C / ALT:T), 132258402(REF:A / ALT:G), 132258420(REF:C / ALT:G), 132258688(REF:T / ALT:G), 132258702(REF:C / ALT :T)、132258770(REF:C / ALT:T)、132258877(REF:G / ALT:C)、132259697(REF:T / ALT:A)、132259757(REF:C / ALT:T)、132259763(REF:A / ALT:C)、13225 9780(REF:T / ALT:G), 132259781(REF:G / ALT:C), 132259795(REF:G / ALT:A), 132259845(REF:C / ALT:T), 132259846(REF:C / ALT:G), 132259857(REF: G / ALT:T), 132259884(REF:T / ALT:C), 132260234(REF:C / ALT:G), 132260332(REF:G / ALT:A), 132260347(REF:G / ALT:T), 132260395(REF:C / ALT:A),132260418(REF:G / ALT:A)、132260487(REF:G / ALT:A)、132260503(REF:C / ALT:A)、132260560(REF:A / ALT:T)、132260987(REF:C / ALT:T)、132261068(REF:G / ALT:A)、132261179(REF:G / ALT:C)、132261184(REF:A / ALT:G)、132261211(REF:C / ALT:A)、132261266(REF:G / ALT:A)、132261401(REF:G / ALT:C)、132261419(REF:C / ALT:T)、132261422(REF:G / ALT:T)、132261441(REF:G / ALT:A)、132261458(REF:A / ALT:G)、132261879(REF:A / ALT:G)、132262159(REF:C / ALT:T)、132262162(REF:G / ALT:A)、132262215(REF:G / ALT:A)、132262522(REF:T / ALT:C)、132262550(REF:A / ALT:G)、132262680(REF:G / ALT:T)、132263022(REF:C / ALT:A)、132263049(REF:T / ALT:G)、132263062(REF:A / ALT:C)、132263432(REF:C / ALT:T)、132263618(REF:C / ALT:A)、132263714(REF:T / ALT:C)、132264048(REF:C / ALT:T)、132264098(REF:G / ALT:T)、132264216(REF:C / ALT:T)、132264222(REF:T / ALT:G)、132264239(REF:C / ALT:T)、132264393(REF:T / ALT:C)、132264412(REF:G / ALT:A)、132264418(REF:C / ALT:T)、132264427(REF:G / ALT:A)、132264438(REF:G / ALT:A)、132264441(REF:G / ALT:A)、132264446(REF:A / ALT:G)、132264447(REF:T / ALT:C)、132264474(REF:G / ALT:A)、132264486(REF:G / ALT:A)、132264512(REF:A / ALT:G)、132264521(REF:G / ALT:A)、132264533(REF:A / ALT:G)、132264556(REF:T / ALT:C)、132264849(REF:T / ALT:G)、132264915(REF:C / ALT:T)、132264922(REF:C / ALT:T)、132264932(REF:C / ALT:T)、132264967(REF:G / ALT:A)、132265004(REF:C / ALT:T)、132265015(REF:G / ALT:A)、132265070(REF:T / ALT:G)、132265178(REF:A / ALT:SV)、132265250(REF:G / ALT:A)、132265466(REF:G / ALT:A)、132265479(REF:C / ALT:T)、132265499(REF:T / ALT:C)、132265544(REF:C / ALT:T)、132265545(REF:T / ALT:G)、132265549(REF:T / ALT:C)、132265566(REF:C / ALT:T)、132265586(REF:A / ALT:C)、132265589(REF:C / ALT:T)、132265591(REF:T / ALT:A)、132265597(REF:C / ALT:T)、132265639(REF:T / ALT:A)、132265708(REF:A / ALT:T)、132265741(REF:A / ALT:G)、132265825(REF:C / ALT:T)、132265853(REF:T / ALT:C)、132265859(REF:A / ALT:G)、132265942(REF:C / ALT:T)、132265954(REF:T / ALT:G)、132266018(REF:G / ALT:T)、132266051(REF:T / ALT:G)、132266062(REF:T / ALT:G)、132266065(REF:G / ALT:C)、132266066(REF:G / ALT:A)、132266084(REF:G / ALT:A)、132266095(REF:T / ALT:C)、132266145(REF:T / ALT:C)、132266177(REF:C / ALT:T)、132266227(REF:A / ALT:C)、132266238(REF:G / ALT:A)、132266247(REF:A / ALT:G)、132266268(REF:T / ALT:C)、132266273(REF:C / ALT:A)、132266278(REF:G / ALT:T)、132266282(REF:C / ALT:T)、132266359(REF:C / ALT:A)、132266388(REF:G / ALT:A)、132266390(REF:G / ALT:A)、132266405(REF:C / ALT:T)、132266445(REF:C / ALT:G)、132266522(REF:T / ALT:C)、132266559(REF:T / ALT:A)、132266606(REF:A / ALT:C)、132266635(REF:T / ALT:A)、132266641(REF:G / ALT:A)、132266645(REF:G / ALT:A)、132266655(REF:G / ALT:A)、132266663(REF:A / ALT:C)、132266676(REF:G / ALT:A)、132266695(REF:T / ALT:C)、132266708(REF:A / ALT:G)、132266718(REF:C / ALT:G)、132266721(REF:A / ALT:C)、132266733(REF:T / ALT:C)、132266754(REF:C / ALT:T)、132266766(REF:G / ALT:A)、132266795(REF:C / ALT:A)、132266867(REF:T / ALT:C)、132266916(REF:G / ALT:A)、132266920(REF:A / ALT:G)、132266928(REF:T / ALT:C)、132266948(REF:T / ALT:A)、132266959(REF:T / ALT:C)、132267009(REF:A / ALT:C)、132267026(REF:A / ALT:C)、132267036(REF:T / ALT:C)、132267071(REF:C / ALT:T)、132267076(REF:C / ALT:T)、132267114(REF:C / ALT:A)、132267118(REF:G / ALT:T)、132267126(REF:T / ALT:A)、132267144(REF:G / ALT:A)、132267162(REF:G / ALT:A), 132267168(REF:G / ALT:T), 132267249(REF:G / ALT:C), 132267260(REF:T / ALT:C), 132267306(REF:A / A LT:G), 132267316(REF:T / ALT:C), 132267361(REF:A / ALT:G), 132267382(REF:T / ALT:C), 132267386(REF:G / ALT:A), 132267417(RE At the F:G / ALT:T site, when the bases sequentially present as CTATTTGGCCACGATGCCTAGGCCCTCCGGGTGTCGATAAAATTACGAACTTAGGTAACGGAGCTACTTTTTCATAAGGCAAGAGCGTTTATAGAAATCTGCTCCATATGTCGTGTGGCAACCTCAGTATCAAATGCACAAAACACGGCCCGCTAGCACCCCTTCTAAATCCGCGCAT, maize exhibits the dominant trait of high kernel protein content.

5. The application according to any one of claims 2-4, characterized in that, The products include kits, reagents, or gene chips.

6. The application according to claim 3, characterized in that, The product is prepared using qPCR or dPCR.

7. The application according to claim 4, characterized in that, The product is prepared using Sanger sequencing, high-throughput sequencing, fluorescence in situ hybridization, TaqMan probe method, ARMS-PCR method, or KASP method.

8. A method for screening maize germplasm with high protein content in the kernels, characterized in that, Take a corn sample to be tested, and test the gene expression level in the application described in claim 2, or the haplotype in the application described in claim 3. Select germplasm that meets the requirements, which is corn germplasm with high protein content in the kernels.