SNP (Single Nucleotide Polymorphism) molecular marker combination for predicting tryptophan content of fresh corn kernels and application of SNP molecular marker combination
By employing a genome-wide selection technique combining 20,000 SNP molecular markers and molecular probes, the problem of low tryptophan content in fresh corn kernels was solved, enabling early, efficient, and low-cost breeding prediction with high accuracy.
Patent Information
- Application Number
- CN202511753079.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-11-26
AI Technical Summary
In existing fresh maize breeding, the low levels of lysine and tryptophan lead to hidden starvation. Traditional phenotypic selection methods are costly, inefficient, and time-consuming, and there is a lack of genome-wide selection breeding research using molecular marker technology.
Using a combination of 20,000 SNP molecular markers and molecular probes, a predictive model was constructed through genome-wide selection technology. Genotypes were then detected using methods such as resequencing to predict the tryptophan content in fresh corn kernels, thereby reducing breeding costs and shortening the breeding cycle.
It enables the detection of high-quality protein content in grains at the seedling or early seed stage, improving breeding efficiency, reducing costs, and achieving a prediction accuracy of 95.16%.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of fresh corn selection and breeding, and particularly relates to a SNP molecular marker combination for predicting the high and low content of tryptophan in fresh corn kernels and application thereof. BACKGROUND
[0002] Fresh corn is a type of corn that is eaten at the milk stage, mainly including sweet corn, waxy corn and sweet waxy corn. As a high-nutrient economic crop, the planting area of fresh corn in China has exceeded 28 million mu, and plays an important role in promoting farmers' income. However, the content of essential amino acids lysine and tryptophan in corn kernels is too low, which can cause hidden hunger, and it is of great significance to improve the content of lysine and tryptophan in fresh corn kernels through breeding. High lysine and tryptophan corn is also known as high-quality protein corn.
[0003] At present, the breeding of fresh corn in China still mainly relies on traditional phenotypic selection method, but this method has inherent defects such as high cost, low efficiency and long breeding cycle. For complex traits such as lysine and tryptophan content in corn kernels, high-performance liquid chromatography needs to be used for detection, which is very costly. In contrast, whole genome selection has significant advantages, including low cost, high efficiency, improved selection efficiency and shortened breeding cycle. This technology has become a routine technical means in corn breeding for international seed companies. It should be particularly pointed out that molecular marker technology constitutes an important technical basis for modern breeding system.
[0004] Whole genome selection is a breeding technique that evaluates the total genetic value of individuals through markers distributed in the whole genome of corn and selects them according to their genetic value. It has been widely used in a variety of animals and plants. However, there is no report on the use of molecular marker technology for whole genome selection breeding of high-quality protein fresh corn. SUMMARY
[0005] In view of this, one of the purposes of the present application is to provide a SNP molecular marker combination for predicting the high and low content of tryptophan in fresh corn kernels, a molecular probe combination and application thereof.
[0006] The second purpose of the present application is to provide a method for predicting the high and low content of tryptophan in fresh corn kernels, which only needs to detect the genotype of breeding materials to predict the content of high-quality protein in fresh corn kernels, and has the advantages of clear selection target and being not affected by the environment.
[0007] In order to achieve the above-mentioned purposes of the application, the present application provides the following technical solutions: The application provides a SNP molecular marker combination for predicting the content of tryptophan in fresh corn kernels, which is composed of 20000 SNP loci located on a B73 reference genome version v5, and information of the 20000 SNP loci is shown in Table 1. Table 1 information of 20000 SNP loci
[0008] The application further provides a molecular probe combination for specifically recognizing the SNP molecular marker combination.
[0009] The application further provides application of the SNP molecular marker combination or the molecular probe combination in any one of the following aspects: (1) constructing a fresh corn kernel tryptophan content high-low prediction model; (2) high-quality protein fresh corn genetic selection breeding; (3) fresh sweet corn multi-traits aggregation breeding.
[0010] The application further provides a method for constructing a fresh corn kernel tryptophan content high-low prediction model, comprising the following steps: detecting the genotypes of the 20,000 SNP molecular markers and the tryptophan content in multiple corn materials respectively, performing missing data filling analysis of the genotype data by using beagle software, inputting the detected genotypes and tryptophan content into an rrBLUP software package, performing whole genome selection research, and obtaining a prediction model.
[0011] The specific method for detecting the genotypes of the 20,000 SNP molecular markers in the sample to be detected is not particularly limited in the present application, and methods such as resequencing, liquid chip, KASP, etc. can be used. In the embodiments of the present application, a low-depth resequencing method is used. The specific method for detecting the tryptophan content in the corn material is not particularly limited in the present application, and a conventional method for detecting the tryptophan content in the art can be used. In the present application, the plurality of corn materials preferably include waxy corn inbred lines, sweet corn inbred lines, and sweet waxy double recessive inbred lines. In the present application, the number of corn materials required for constructing the prediction model is preferably 200 or more. The present application does not have a special limitation on the number of waxy corn inbred lines, sweet corn inbred lines, and sweet waxy double recessive inbred lines, as long as the total number of corn materials is 200 or more. In the present application, when performing genotype data missing data filling analysis using beagle software and performing whole genome selection research using rrBLUP software package, genotype and tryptophan content are used for model construction, and in addition to inputting genotype and tryptophan content, the remaining parameters are set to the default parameters of the software.
[0012] The present application also provides a prediction model for the high and low tryptophan content of fresh corn kernels, which is obtained by the above method.
[0013] The present application also provides the use of the above prediction model in predicting the high and low tryptophan content of fresh corn kernels.
[0014] The present application also provides a method for predicting the high and low tryptophan content of fresh corn kernels, comprising the following steps: identifying the genotype of the above SNP molecular marker combination in the sample to be detected, performing genotype data missing data filling analysis using beagle software, inputting the genotype into the above prediction model, performing whole genome selection research using rrBLUP software package, obtaining breeding value, and the sample to be detected with high breeding value indicating high tryptophan content.
[0015] The application uses the genotype information of the above 20000 SNP sites to predict the samples without determining the phenotype by using the prediction model (or referred to as the whole genome selection model) constructed by the training population of the application, and obtain the predicted value of the phenotype for breeding selection. In breeding, the determination of the phenotype value is high in cost, and when the amino acid determination is performed, 300 yuan / sample is needed, and the sample is taken for determination after the whole growth period of corn is completed. The cost and time cost of material planting and management are also high. The method provided by the application only needs to extract DNA in the seedling stage or seed for genotype detection, and the breeding value (phenotype value) can be predicted by the method constructed by the application. The selection according to the predicted phenotype can greatly reduce the breeding cost, shorten the breeding cycle and improve the breeding efficiency. In the application, after the breeding value is obtained, the high and low of the tryptophan content in the corn sample to be measured is predicted according to the size of the breeding value, and the high breeding value indicates that the tryptophan content is high. The high and low of the breeding value mentioned in the application is relative, that is, the corn samples to be measured are compared, and the high indicates that the tryptophan content is relatively high. In the application, after the breeding values of a plurality of corn samples to be measured are obtained, the breeding values are sorted from top to bottom, and the corn samples corresponding to the top 20% of the breeding values are preferably regarded as high tryptophan content and have subsequent breeding prospects.
[0016] The application uses the five-fold cross-validation method to generate the training population and the verification population for evaluating the prediction accuracy, that is, 205 inbred lines are selected from the population as the training population to construct the prediction model, then the constructed prediction model is used to predict the phenotype values of the remaining 51 corn inbred lines, the predicted phenotype values of the 51 corn inbred lines are compared with the actually determined phenotype values, the sampling is repeated 100 times, and the prediction accuracy of the constructed model is evaluated. In the experiment, the correlation coefficient between the predicted value and the observed value of the grain protein content of the verification population generated by the five-fold cross-validation is used as the prediction accuracy, and the "cor()" function in the R language is used for calculation. The results show that the average value of the prediction accuracy of the prediction method of the fresh corn kernel tryptophan content constructed by using the 20000 SNPs shown in Table 1 is 95.16%.
[0017] The application also provides a high-quality protein fresh corn gene selection breeding method, which comprises the following steps: identifying the genotype of the above SNP molecular marker combination in the sample to be measured, performing missing data filling analysis of the genotype data by using the beagle software, inputting the genotype into the above prediction model, and then performing whole genome selection research by using the rrBLUP software package to obtain the breeding value.
[0018] The method for selecting and breeding high-quality protein fresh corn provided by the application is constructed according to the genotype data and the phenotype data (tryptophan content in the kernel) of the training population, and is constructed by using the rrBLUP software; only the genotype of the individual needs to be determined, the genotype is input into the rrBLUP software, the rest are all the default parameters of the software, and then the breeding value can be calculated based on the prediction model; the tryptophan content in the corn sample to be measured is predicted according to the high and low of the breeding value, and since the corn is lack of lysine and tryptophan, which are two kinds of essential amino acids for human body, the corn containing lysine or tryptophan or both of them is called high-quality protein corn.
[0019] The specific method for identifying the genotype of the SNP molecular marker in the sample to be measured is not particularly limited in the application, and methods such as resequencing, liquid chip and KASP can be used, and in the embodiments of the application, a low-depth resequencing method is used. In the application, the sample to be measured with a high breeding value is preferably selected for breeding. In the application, the high breeding value refers to a relative value, that is, the breeding values of the samples to be measured are compared, and the high value indicates that it has breeding prospects.
[0020] The application has the following beneficial effects: The SNP molecular marker combination for predicting the high and low of the tryptophan content in the kernel of fresh corn provided by the application can be used for molecular marker assisted breeding of high-quality protein (high tryptophan content) fresh corn, and can also be used for multi-traits of fresh sweet corn.
[0021] The method for predicting the high and low of the tryptophan content in the kernel of fresh corn provided by the application only needs to detect the genotype of the breeding material, and can predict the content of high-quality protein (tryptophan) in the kernel of fresh corn, has the advantages of clear selection target and not affected by the environment, and meanwhile, the kernel high-quality protein (tryptophan) content trait prediction provided by the application is mainly in the harvesting period of fresh corn, and early detection can be realized. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 The box plot of the tryptophan content in the kernel in the breeding population; Figure 2 The distribution diagram of the SNP on the ten chromosomes of corn; Figure 3 The Manhattan plot and the QQ plot of the whole genome association analysis result of the tryptophan content, wherein the left graph is the Manhattan plot and the right graph is the QQ plot; Figure 4 The prediction accuracy of the whole genome prediction model for tryptophan. DETAILED DESCRIPTION
[0023] The technical solutions provided by the application will be described in detail below in combination with the embodiments, but they should not be understood as limiting the protection scope of the application.
[0024] In the following examples, unless otherwise stated, conventional methods were used.
[0025] In the following examples, unless otherwise stated, the materials, reagents, etc. used can be obtained commercially.
[0026] Example 1 1. Planting of Fresh Corn Breeding Population In this experiment, 256 representative inbred lines were selected from the core breeding germplasm of fresh corn, including 81 waxy corn inbred lines, 157 sweet corn inbred lines, and 18 sweet-waxy double recessive inbred lines. The population materials were planted in the spring of 2024 at the Zhuanghang Comprehensive Test Station of Shanghai Agricultural Academy. The experiment used a completely randomized design, with 20 plants per material, 2 rows, row length of 2.5 meters, and routine field management. Self-pollination was performed by bagging, and after the ears matured, 3 evenly growing ears were selected from each material for subsequent phenotype determination.
[0027] 2. Determination of Grain Tryptophan Content In this experiment, the corn kernels were first ground into powder using a sample crusher, then mixed with hydrochloric acid solution, and subjected to 24 hours of acid hydrolysis at 110°C. After hydrolysis, the supernatant was taken and neutralized to neutral with sodium hydroxide solution. Then, AccQ•Tag Ultra Borate buffer and AccQ•Tag reagent were added to the neutralized sample, heated at 55°C for 10 minutes for derivatization reaction. After the reaction, the sample solution cooled to room temperature was analyzed using ultra-high performance liquid chromatography-tandem high-resolution Orbitrap mass spectrometry (UHPLC-QE, Thermo, USA) to detect the content of tryptophan in the sample.
[0028] Using Meta-R software, the Best Linear Unbiased Prediction (BLUP) of tryptophan content in each material was calculated. After calculation, among the 256 fresh corn breeding materials, the tryptophan content was between 269~1085µg / g, with an average of 686µg / g, as shown in Figure 1 .
[0029] 3. Genotype Determination of Association Population Genomic DNA was extracted from the young leaves of 256 fresh corn inbred lines using the CTAB method, and whole genome resequencing was performed using the Illumina sequencing platform. After sequencing, the raw sequencing data of the 256 fresh corn breeding materials was between 12.25~35.69 Gb, and the average sequencing data was 16.04 Gb. Subsequently, the sequencing data quality control, alignment to the reference genome, removal of PCR repetitive sequences in the sequencing data and SNP identification were carried out, and the data processing process was as follows: (1) using FASTP software based on Q20 standard to control and filter the raw sequencing data; (2) using BWA software to align the filtered data to the corn reference genome B73v5; (3) using Picard software package to mark the repetitive sequences in the sequencing data; (4) using GATK software for variation site detection. After obtaining the SNP sites, the genotype data was filtered using bcftools software, and the screening criteria included: ① only double allele SNP sites were retained; ② minor allele frequency (MAF)>0.05; ③ deletion rate<5%; ④ adjacent SNP linkage disequilibrium strength>0.5. After filtering, a total of 886066 SNPs were retained, as shown in Table 2 and Figure 2 , for subsequent whole genome association analysis.
[0030] Table 2. Statistical results of filtered SNPs
[0031] 4. Whole genome association analysis of tryptophan content The Gapit software was used to perform whole genome association analysis (GWAS) of the grain tryptophan content trait of the 256 fresh corn population based on the mixed linear model (with the kinship matrix and population structure matrix as covariates). The significance threshold of GWAS analysis was set to 1×10 -5 . The markers were sorted according to the effect of GWAS analysis from large to small, and the top 20000 SNP markers (as shown in Table 1) were selected for model construction and subsequent prediction. Among the 20000 SNP markers, 10 SNP sites significantly associated with tryptophan content were identified (as shown in Figure 3 and Table 3).
[0032] Table 3. Significant SNP information of whole genome association analysis
[0033] Example 2 Construction of prediction model for high and low tryptophan content of fresh corn kernels The 20000 SNPs screened in Example 1 were used for model construction. Specifically, the genotypes of the 20000 SNP markers as described in Table 1 in the 205 corn samples in Example 1 were identified by resequencing, and the tryptophan content of each of the 205 corn samples was detected. The genotype data was analyzed for missing data filling by beagle software, and the genotypes and tryptophan content were input into the rrBLUP software. Subsequently, the whole genome selection study was performed by the rrBLUP software (https: / / CRAN.R-project.org / package=rrBLUP) package to obtain a prediction model for the high and low tryptophan content of fresh corn kernels.
[0034] Example 3 The prediction model constructed in Example 2 was used to predict the tryptophan content in the kernels of the remaining 51 corn materials in Example 1. Specifically, The genotypes of the 20000 SNP markers as described in Table 1 in the 51 corn samples were identified by resequencing. The genotype data was analyzed for missing data filling by beagle software, and the genotypes were input into the prediction model of Example 2. The whole genome selection study was performed by the rrBLUP software package to obtain breeding values, and the test samples with high breeding values indicated high tryptophan content.
[0035] The predicted tryptophan content of the 51 samples was compared with the actually determined tryptophan content, and the prediction accuracy of the above constructed model was evaluated by repeating 100 times of sampling. The correlation coefficient between the predicted and observed grain protein content of the validation population generated by five-fold cross-validation was used as the prediction accuracy, which was calculated by the "cor()" function in R language.
[0036] The results showed that the prediction method of fresh corn kernel tryptophan content constructed by the 20000 SNPs shown in Table 1 had an average prediction accuracy of 95.16% (as shown in Figure 4 ).
[0037] Example 4 The prediction model constructed in Example 2 was used for high-quality protein fresh corn gene selection breeding. Specifically, the genotypes of the 20000 SNP molecular markers as described in Table 1 in another 30 corn samples were identified by resequencing. The genotype data was analyzed for missing data filling by beagle software, and the genotypes were input into the prediction model of Example 2. Subsequently, the whole genome selection study was performed by the rrBLUP software (https: / / CRAN.R-project.org / package=rrBLUP) package to obtain breeding values.
[0038] The breeding values are arranged from large to small, and the corn samples with the top 20% breeding values (i.e. the top 6) are determined as high-quality protein fresh corn. The grain tryptophan content in the 30 corn samples is determined, and the result is completely consistent with the prediction by the method of the application. The grain tryptophan content in the 6 corn samples is higher than that in the remaining 24 corn samples, and the order of the grain tryptophan content in the 6 corn samples is consistent with the order of the breeding values determined by the application.
[0039] The above only describes the preferred embodiments of the present application, and it should be noted that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements should also be considered within the protection scope of the present application.
Claims
1. A combination of SNP molecular markers for predicting the tryptophan content in fresh corn kernels, characterized in that, It consists of 20,000 SNP sites located on the maize B73 reference genome version v5, as shown in Table 1 of the specification.
2. A molecular probe assembly, characterized in that, The molecular probe combination is used to specifically identify the SNP molecular marker combination of claim 1.
3. The application of the SNP molecular marker combination of claim 1 or the molecular probe combination of claim 2 in any one of the following, characterized in that, (1) Construct a predictive model for high and low tryptophan content in fresh corn kernels; (2) Gene selection breeding of high-quality protein fresh corn; (3) Multi-trait aggregation breeding of fresh sweet corn.
4. The application according to claim 3, characterized in that, The high-quality protein fresh corn mentioned is corn with a high tryptophan content.
5. A method for constructing a predictive model for the tryptophan content in fresh corn kernels, characterized in that, The process includes the following steps: detecting the genotype and tryptophan content of the 20,000 SNP molecular markers described in claim 1 in multiple maize materials, performing missing data imputation analysis of genotype data using Beagle software, inputting the detected genotype and tryptophan content into the rrBLUP software package, performing a genome-wide selection study, and obtaining a prediction model.
6. The method according to claim 5, characterized in that, The plurality of corn materials include waxy corn inbred lines, sweet corn inbred lines, and sweet-waxy double recessive inbred lines; the plurality of materials is more than 200.
7. A predictive model for the tryptophan content in fresh corn kernels, characterized in that, It is obtained by the method described in claim 5 or 6.
8. The application of the prediction model of claim 7 in predicting the tryptophan content of fresh corn kernels.
9. A method for predicting the tryptophan content in fresh corn kernels, characterized in that, The process includes the following steps: identifying the genotype of the SNP molecular marker combination described in claim 1 in the sample to be tested; performing missing data imputation analysis of the genotype data using the Beagle software; inputting the genotype into the prediction model described in claim 7; and then performing a genome-wide selection study using the rrBLUP software package to obtain the breeding value.
10. The method according to claim 9, characterized in that, The high breeding value of the test sample indicates a high tryptophan content.
Citation Information
Patent Citations
Molecular marker for high lysine maize Opaque7 gene and application thereof
CN102586241A
Method for breeding high-lysine sweet and waxy fresh corn variety
CN115281072A
Molecular marker, primer pair and kit for corn high-lysine site, application of molecular marker, primer pair and kit, and genotype detection method
CN118064620A
Obtaining of corn kernel protein major QTL qHP5 and development and application of molecular marker primer of corn kernel protein major QTL qHP5
CN120648837A
KASP molecular marker related to corn kernel protein content and application of KASP molecular marker
CN120843725A