Application of SNPs and / or InDels in Evaluating Heterosis in Maize
By detecting the number of hybrid SNPs and InDels in specific regions of the entire genome of corn, the problems of low efficiency and poor prediction performance in the prior art are solved, and efficient screening and prediction of corn hybrid populations and parent populations are achieved.
Patent Information
- Application Number
- CN202411425933.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-10-12
AI Technical Summary
The prior art has problems of inefficiency and poor predictive performance in evaluating corn hybrid advantages, especially in screening corn hybrid advantage groups and predicting single-plant yields.
A new idea is provided to evaluate and predict corn hybrid advantages by detecting the number of heterozygous SNPs and/or heterozygous InDels in specific regions of the whole genome (2kb region and exon region upstream of the gene).
It is achieved to simply and effectively evaluate and predict single hybrid populations with high yields in corn hybrid populations, and to screen corn parent combinations with relatively high yields in corn parent populations, providing a new idea for the screening of corn hybrid dominant groups.
Smart Images

Figure CN119418763B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of corn hybrid breeding, and specifically relates to the application of the number of SNPs and / or the number of InDels in evaluating the heterosis of corn. Background Art
[0002] Corn is an important food crop. The utilization of heterosis is a key factor in breeding high-yield corn hybrids. Studies have shown that germplasm resources are of great significance to corn heterosis. By grouping these germplasm resources into heterosis groups, breeders can better predict and utilize heterosis. The division of heterosis groups and the study of heterosis patterns are the basis of corn hybrid breeding.
[0003] Methods for dividing heterotic groups include combining ability, pedigree method, molecular marker method, etc. Combining ability is an important tool widely used to select excellent parental lines and determine the type of gene action involved in the target traits, and is often used in the formulation of breeding plans. Combining ability is also an important criterion for dividing heterotic groups. Dividing heterotic groups by SCA of grain yield was a common research method used by breeders in the early days. Since the SCA effect is greatly affected by the interaction between the two inbred lines and the interaction between the hybrid and the environment, different studies may divide the same inbred line into different heterotic groups. There have always been some differences in breeding efficiency between the combining ability division method and the molecular marker method. Genetic distance (GD) is a standard based on molecular markers to describe the genomic similarity between parents. With the development of molecular marker technology, restriction fragment length polymorphism (RFLP), amplified length polymorphism (AFLP), randomly amplified polymorphic DNA (RAPD), simple sequence repeats (SSR), and single nucleotide polymorphism (SNP) are widely used to calculate GD and use the obtained GD to predict maize heterosis. Some studies have obtained the results that GD is positively correlated with maize heterosis. However, the performance of GD and heterosis effects is affected by many factors, and there are certain differences in their prediction performance. Previous studies on maize heterosis prediction mainly focused on the use of GD methods, and other molecular marker methods were less, and the various factors affecting GD prediction of heterosis have not yet been clarified. Therefore, how to screen maize heterosis groups is still one of the key issues that need to be studied in maize breeding. Summary of the invention
[0004] In order to solve the problems existing in the prior art, the purpose of the present invention is to provide an application of the number of SNPs and / or InDels in a specific region of the whole genome of corn in evaluating corn hybrid vigor (single plant yield), providing a new idea for the screening of corn hybrid vigor groups.
[0005] The present invention also aims to provide a method for screening a hybrid population of corn with a higher yield than its parents, which can effectively predict and screen hybrid offspring with a higher yield per plant than its parents.
[0006] The present invention also aims to provide a method for screening a corn parent group with relatively high-yield offspring, which can effectively predict and screen parent combinations whose offspring after hybridization exhibit relatively high-yield traits per plant.
[0007] In order to achieve the above-mentioned object of the invention, the present invention provides the following technical solutions:
[0008] The present invention provides an application of the number of SNPs and / or the number of InDels in evaluating the heterosis of corn, wherein the SNPs are located in the 2kb region upstream of the gene of the corn full gene and the exon region, the SNPs are heterozygous SNPs, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the gene of the corn full gene and the exon region, the InDels are heterozygous InDels; the heterosis of corn is the heterosis of single plant yield.
[0009] Preferably, in corn hybrids, the number of SNPs and / or the number of InDels is positively correlated with corn heterosis.
[0010] The present invention also provides a method for screening a hybrid population of corn with a higher yield than its parents, comprising the following steps: counting the number of SNPs and the number of InDels of the corn hybrid population to be screened, and selecting a hybrid population with the highest number of SNPs and / or InDels in the corn hybrid population; the SNPs are located in the 2kb region upstream of the gene and the exon region of the corn full gene, the SNPs are heterozygous SNPs, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the gene and the exon region of the corn full gene, and the InDels are heterozygous InDels.
[0011] Preferably, the parents of the corn hybrid population to be screened are homozygous; the parents are any one or two of a Tropical population, a Reid population and a Non-Reid population.
[0012] Preferably, the number of SNPs of the corn hybrid population to be screened is obtained according to the combination of the number of paternal SNPs and the number of maternal SNPs, and the number of InDels of the corn hybrid population to be screened is obtained according to the combination of the number of paternal InDels and the number of maternal InDels.
[0013] The present invention also provides a method for screening a corn parent group with relatively high yield in offspring, comprising the following steps: counting the number of SNPs and the number of InDels in the corn parent group to be screened, combining the corn parents to be screened in pairs to obtain the number of heterozygous SNPs and the number of InDels in the offspring, and selecting the corn parent group combination with the highest number of heterozygous SNPs and / or InDels in the offspring; the SNPs are located in the 2kb region upstream of the gene and the exon region of the corn full gene, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the gene and the exon region of the corn full gene.
[0014] Preferably, the corn parent population is homozygous.
[0015] Preferably, the number of SNPs and the number of InDels of the corn parent population to be screened are obtained based on the whole genome sequence of the corn parent population.
[0016] Preferably, the corn parent population to be screened is any one or two of a Tropical population, a Reid population and a Non-Reid population.
[0017] The invention also provides application of the method in corn hybrid breeding.
[0018] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0019] The present invention discovers for the first time that by only detecting the number of heterozygous SNPs and / or heterozygous InDels in specific regions of the corn genome (2kb regions upstream of the gene and exon regions), it is possible to simply and effectively evaluate / predict / screen single-plant hybrid populations with higher yields than their parents in corn hybrid populations, and to screen corn parent combinations whose offspring have higher yields than their parents in corn parent populations, which provides a new idea for screening corn hybrid advantage groups and has good prospects in corn breeding. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 : Principal component analysis of 37 parents;
[0021] Figure 2 : HSGCA of GYPP between 34 tested lines and 3 tested species;
[0022] Figure 3 : The performance of UE-SNPs, UE-InDels of three heterotic groups and HUE-SNPs, HUE-InDels, and GD of three heterotic patterns;
[0023] Figure 4 : MEAN of GYPP, prediction method, and heterosis correlation diagram. DETAILED DESCRIPTION
[0024] The present invention provides an application of the number of SNPs and / or the number of InDels in evaluating corn heterosis, wherein the corn heterosis is the heterosis of single plant yield, wherein the SNPs are located in the 2kb region upstream of the gene and the exon region, and the SNPs are heterozygous SNPs, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the gene and the exon region, and the InDels are heterozygous InDels. The 2kb region upstream of the gene described in the present invention is the region of the first 2000 bases in the 5' end direction of the transcribed gene in the whole genome, and the exon region described in the present invention is the expression sequence retained after shearing in the whole genome. The present invention first discovered that in corn hybrids, the number of heterozygous SNPs and / or the number of heterozygous InDels is positively correlated with corn heterosis (single plant yield). The number of SNPs and the number of InDels in corn show a positive correlation. The present invention can simply and effectively evaluate / predict / screen relatively high-yield single-plant hybrid populations in corn hybrid populations by detecting the number of heterozygous SNPs or the number of heterozygous InDels in specific regions of the whole genome of corn hybrid populations, or can simultaneously detect the number of heterozygous SNPs and the number of heterozygous InDels in specific regions of the whole genome of corn hybrid populations, and screen corn parents with relatively high-yield offspring in corn parent populations.
[0025] The present invention provides a method for screening a hybrid population with high yield relative to a parent of corn, comprising the following steps: counting the number of SNPs and the number of InDels of the corn hybrid population to be screened, and selecting the hybrid population with the highest number of SNPs and / or InDels in the corn hybrid population; the SNPs are located in the 2kb region upstream of the gene of the corn full gene and the exon region, and the SNPs are heterozygous SNPs, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the gene of the corn full gene and the exon region, and the InDels are heterozygous InDels. In the present invention, the higher the number of heterozygous SNPs and / or the number of heterozygous InDels of the corn hybrid population, the higher the yield per plant of the corn hybrid population.
[0026] The parents of the hybrid population to be screened in the present invention are homozygous, and the parents are any one or two of the tropical population, the Reid population and the non-reid population, that is, the tropical population, the Reid population and the non-reid population are hybridized between populations or within populations to obtain the hybrid population to be screened. The tropical population, the Reid population and the non-reid population in the present invention are the population divisions of corn by the HSGCA method, and all corn varieties are divided into three populations, namely the tropical population, the Reid population and the non-reid population.
[0027] The present invention does not limit the method for obtaining the number of heterozygous SNPs and the number of heterozygous InDels of the hybrid population to be screened, and the statistical method of the number of heterozygous SNPs and the number of heterozygous InDels adopts the method commonly used in the art. As an optional embodiment, the number of heterozygous SNPs of the corn hybrid population to be screened according to the present invention is obtained according to the combination of the number of paternal SNPs and the number of maternal SNPs, and the number of heterozygous InDels of the corn hybrid population to be screened is obtained according to the combination of the number of paternal InDels and the number of maternal InDels. The number of paternal SNPs and the number of InDels of the present invention are obtained according to the paternal whole genome sequence, and the paternal whole genome sequence is obtained by extracting the paternal leaves at the seedling stage; the number of maternal SNPs and the number of InDels of the present invention are obtained according to the maternal whole genome sequence, and the maternal whole genome sequence is obtained by extracting the maternal leaves at the seedling stage. In the present invention, the number of parental SNPs and InDels includes heterozygous and homozygous, which is the total number of heterozygous and homozygous. The number of SNPs and InDels in the hybrid population (offspring) to be screened are all heterozygous SNPs and heterozygous InDels.
[0028] In the present invention, the number of parental SNPs located in the 2 kb region upstream of the gene and exons (except synonymous mutations) is referred to as UE-SNPs, the number of parental InDels located in the 2 kb region upstream of the gene and exons is referred to as UE-InDels, the number of hybrid (offspring) heterozygous SNPs located in the 2 kb region upstream of the gene and exons (except synonymous mutations) is referred to as HUE-SNPs, and the number of hybrid (offspring) heterozygous InDels located in the 2 kb region upstream of the gene and exons is referred to as HUE-InDels.
[0029] The method for obtaining UE-SNPs (the number of parental SNPs) and UE-InDels (the number of parental InDels) of the present invention is preferably as follows: extracting paternal / maternal DNA from seedling leaves, sequencing the paternal / maternal whole genome and quality control, and then using BWAv0.7.17 software to align the Clean reads with the maize reference genome B73_RefGen_v4, with the parameter set to mem-t4-k 32-M. GATK4 and Vcftools are used to identify and filter SNPs and InDels, with the parameter set to varFilter-w5
[0030] -w10,clusterSize 2clusterWindowSize 5,QUAL<30,QD<2.0,MQ<40,FS>60.0. ANNOVAR software tool and maize reference genome B73_RefGen_v4 were used to generate high-quality SNPs and InDels and annotate them, and the number of different types of SNPs and InDels was obtained. SNPs located in the 2kb region upstream of the gene and exons (except synonymous mutations) and InDels located in the 2kb region upstream of the gene and exon regions were selected and counted, and the number of SNPs and InDels (total number of heterozygous and homozygous) of the paternal / maternal parent were obtained respectively.
[0031] The method for obtaining HUE-SNPs (the number of heterozygous SNPs in the offspring) and HUE-InDels (the number of heterozygous InDels in the offspring) of the present invention is preferably: according to the combination relationship of the hybrid (the number of SNPs and InDels of the father, the number of SNPs and InDels of the mother), the pure and different UE-SNPs and UE-InDels between the parents are selected to obtain the HUE-SNPs and HUE-InDels of the hybrid.
[0032] The present invention also provides a method for screening a relatively high-yield corn parent population of offspring, comprising the following steps: counting the number of SNPs and the number of InDels of the corn parent population to be screened, combining the corn parents to be screened in pairs, obtaining the number of heterozygous SNPs and the number of heterozygous InDels of the offspring, and selecting the corn parent population combination with the highest number of heterozygous SNPs and / or InDels of the offspring; in the present invention, the higher the number of heterozygous SNPs and / or InDels of the offspring, the better the hybrid vigor (yield per plant) of the corn parent combination obtained after hybridization of the offspring; the SNPs are located in the 2kb region upstream of the gene and the exon region, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the gene and the exon region. In the present invention, the number of the parent SNPs and InDels includes heterozygous and pure, which is the total number of heterozygous and homozygous.
[0033] The parents of the hybrid population to be screened in the present invention are homozygous, and the parents are any one or two of the tropical population, the Reid population and the non-Reid population, that is, the hybrid population to be screened is obtained by inter-group hybridization or intra-group hybridization of the tropical population, the Reid population and the non-Reid population.
[0034] According to the combination relationship of the hybrid (the number of SNPs and InDels of the father, the number of SNPs and InDels of the mother), the present invention selects the pure and different UE-SNPs and UE-InDels between the parents to obtain the HUE-SNPs and HUE-InDels of the hybrid. The parent combination that can obtain the offspring with high heterosis (single plant yield) is screened by the level of the HUE-SNPs and / or HUE-InDels of the obtained offspring hybrid. The method for obtaining UE-SNPs (the number of parental heterozygous SNPs) and UE-InDels (the number of parental heterozygous InDels) of the present invention is preferably: extracting parental DNA from seedling leaves, and after whole genome sequencing and quality control, using BWA v0.7.17 software to compare Clean reads with the corn reference genome B73_RefGen_v4, and the parameters are set to mem-t 4-k 32-M. GATK4 and Vcftools were used to identify and filter SNPs and InDels, with parameters set to varFilter-w5-w10, clusterSize 2clusterWindowSize 5, QUAL<30, QD<2.0, MQ<40, FS>60.0. ANNOVAR software tools and maize reference genome B73_RefGen_v4 were used to generate high-quality SNPs and InDels and annotate them, and the number of different types of SNPs and InDels was obtained. SNPs located in the 2kb region upstream of the gene, exons (except synonymous mutations) and InDels located in the 2kb region upstream of the gene and exon regions were selected and counted to obtain the UE-SNPs and UE-InDels of the parents.
[0035] The present invention also provides an application of the above method in corn hybrid breeding. Preferably, the present invention obtains three hybrid single plant yield advantage patterns when the parents are tropical group, Reid group and Non-Reid group through the above method: Tropical×Reid>Tropical×Non-Reid>Reid×Non-Reid.
[0036] The technical solutions in the present invention will be described clearly and completely below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] In the embodiments of the present invention, the plant materials Ye107, NK40, Shen137, CML171, D39, TRL02, YML226, YML1218, Q11, Chang7-2, TML418, R-2, YML335, and S08 are all from the Institute of Grain Crops, Yunnan Academy of Agricultural Sciences. Among them, Ye107 and NK40 (NK40-1) are disclosed in the document “Li S, Jiang F, Bi Y, Yin X, Li L, Zhang X, Li J, Liu M, Shaw RK, Fan X: Utilizing Two Populations Derived from Tropical Maize for Genome-Wide Association Analysis of Banded Leaf and Sheath Blight Resistance. Plants (Basel) 2024, 13 (3)”; Shen137 and CML171 are disclosed in the document “Ran F, Wang Y, Jiang F, Yin X, Bi Y, Shaw RK, Fan X: Studies on Candidate Genes Related to Flowering Time in a Multiparent Population of Maize Derived from Tropical and Temperate Germplasm. Plants (Basel) 2024, 13 (7)”; D39 and TRL02 are disclosed in the document “Jiang F, Liu L, Li Z, Bi Y, Yin X, Guo R, Wang J, ZhangY, Shaw RK, Fan 7-2) In the literature "Bi Y, Jiang F, Yin X, Shaw RK, Guo R, Wang J, Fan X: Identification of candidategene associated with maize northern leaf blight resistance in a multi-parent population.Plant Cell Rep2024,43(7):189."; TML418 was disclosed in the document "WangY, Bi Y, JiangF, Shaw RK, Sun J, Hu C, Guo R, Fan 2023,45(5):4416-4430." published in
[0038] In the following embodiments, unless otherwise specified, all of them are conventional methods.
[0039] Unless otherwise specified, the materials and reagents used in the following examples can be obtained from commercial sources.
[0040] Example 1
[0041] 1. Plant materials
[0042] Ten F8 RILs (F8 generation recombinant inbred lines) were constructed with Ye 107 as the common male parent and 10 different inbred lines (NK40, Shen 137, CML171, D39, YML226, YML1218, Q11, Chang 7-2, TML418, R-2) as the female parents. 34 families were selected from the 10 F8 RILs as the tested lines (male parents) in this experiment, and three inbred lines S08, YML335, and TRL02 from the Tropical, Reid, and Non-Reid heterotic groups were selected as the test varieties (female parents). The pedigrees and ecotypes of the test varieties and tested lines are shown in Table 1.
[0043] Table 1 Pedigree of parents of the tested lines
[0044]
[0045]
[0046] 2. Experimental methods
[0047] The experiment adopted a test line × test variety design, with 34 F8 RILs as the test line and 3 inbred lines as the test variety, producing 102 hybrid combinations. The test line, test variety, and test line × test variety hybrid combination were planted together in Jinghong City, Yunnan Province (22°01'N, 100°49'E, 552.7m above sea level) in the winter of 2021, numbered 21JH, and planted in Yanshan County, Yunnan Province (1540m above sea level, 23°36'N, 104°18'E) in the summer of 2022, numbered 22YS. Each location adopted a completely randomized block design with 2 replications. Each plot had a row length of 4m, a row spacing of 0.7m, and 14 plants per row. Field management was implemented according to standard production management. When the seed moisture content was adjusted to 140 g / kg, 5 plants were randomly selected from each plot, and eight traits, including ear length (EL, cm), ear diameter (ED, cm), number of ear rows (KRN, No.), number of grains per row (KNR, No.), plant height (PH, cm), ear height (EH, cm), 100-grain weight (HKW, g), and yield per plant (GYPP, g), were measured.
[0048] 3. Data Statistical Analysis
[0049] (1) The R genetic design analysis software (Analysis of Genetic Designs with R for Windows) was used to input the measured data of ear length, ear diameter, number of ear rows, number of grains per row, plant height, ear height, 100-grain weight, and yield per plant according to its instructions, and the variance analysis results (using the general linear mixed model) were directly displayed, with genotype (test line effect, test variety effect, test line × test variety effect) and environmental effect as fixed effects, repeated effects, and environment × genotype interaction as random effects. The calculation formula used in the software is as follows:
[0050] y ijk =μ+E d +REP k (E d )+g ij +E d ×g ij +e ij
[0051] g ij = l i +t j +l i ×t j
[0052] y ijk is the observed value, μ is the general mean, E d is the environmental effect, REP k (E d ) is the repetition effect within the environment, g ijis the genotype effect, e ij is the residual, l i is the measured system effect, t j The data results are shown in Table 2.
[0053] Table 2 Mean squares of eight traits in the cross-environmental test line × test species experiment
[0054]
[0055] * indicates P < 0.05, ** indicates P < 0.01, and *** indicates P < 0.001, which are statistically significant at the probability level.
[0056] Based on the mean square of variance analysis, the environmental effect represents the difference of the corresponding traits of all hybrids between environments (locations), and P < 0.05 indicates that there are significant differences between environments (locations); the repeated effect represents the difference of the corresponding traits of hybrids between different planting replicates, and P < 0.05 indicates that there are significant differences between planting replicates; the genotype effect represents the difference between the corresponding traits of each hybrid, and P < 0.05 indicates that there are significant differences between the corresponding traits of each hybrid; GCA 被测系 The effect represents the difference in GCA (general combining ability) of the corresponding traits of each tested line. P < 0.05 indicates that there is a significant difference in GCA (general combining ability) of the corresponding traits of each tested line. 测验种 Effect represents the difference between the GCA (general combining ability) of the corresponding traits of each test species. P < 0.05 indicates that there is a significant difference between the GCA (general combining ability) of the corresponding traits of each test species. SCA 被测系×测验种 The effect indicates the difference between the corresponding trait SCA (special combining ability) of each hybrid, and P < 0.05 indicates that there is a significant difference between the corresponding trait SCA (special combining ability) of each hybrid; the residual indicates the error size.
[0057] The cross-environment variance analysis of 21JH and 22YS showed that the environmental effects of all traits reached an extremely significant level (P < 0.01), and all traits except the number of ear rows and the number of grains per row also reached a significant level of difference between replicates (P < 0.05), indicating that the setting of block groups effectively reduced the experimental error. 被测系 , GCA 测验种 、SCA 被测系×测验种 All of them reached a significant level (P<0.05), indicating the feasibility of eight traits including single plant yield in field experiments.
[0058] (2) Using R genetic design analysis software (Analysis of Genetic Designs with R for Windows), according to its instructions, input the measurement data of ear length, ear diameter, number of ear rows, number of grains per row, plant height, ear height, 100-grain weight, and single plant yield, and directly display the GCA of the eight traits and the SCA of each hybrid combination. And calculate GCA_SCAratio based on the values of GCA and SCA. The relative importance of GCA and SCA is represented by GCA_SCAratio. The closer this value is to 1, the stronger the predictive power of GCA alone for the performance of a specific hybrid combination. The calculation formula of GCA_SCAratio is as follows:
[0059]
[0060] where σ 2 GCA is the variance of GCA, σ 2 SCA is the variance of SCA.
[0061] The calculation of narrow-sense heritability and broad-sense heritability was completed by AGD-R software. According to the AGD-R software manual, the table that the software can recognize was set up, including the trait value, replication, environment, paternal and maternal information and input into the software. The Line×Tester model was selected, the data type was balanced, the variance estimation method was Harderson, the experimental design was RCBD, and the trait selection was all traits. After that, the output data table was obtained after analysis in the software, and the result was obtained directly. is the narrow sense heritability, is the broad-sense heritability, and the calculation formula is as follows:
[0062]
[0063] in is the narrow sense heritability, is the broad sense heritability, is the additive variance, is the explicit variance, is the phenotypic variance.
[0064] The data results are shown in Tables 3 and 4.
[0065] Table 3 GCA effects of eight traits of Line×Tester across environments
[0066]
[0067]
[0068] * indicates P < 0.05, which is statistically significant at the probability level. SE: standard error.
[0069] The results showed that the GCA of 7 tested lines (L8, L10, L20, L23, L24, L26, L31) reached a significant level, and the GCA effect of no tested varieties reached a significant level, but S08 showed a positive GCA effect in 8 agronomic traits. The tested lines had 1 and 2 significant positive GCA effects in ear diameter and ear row number, respectively, and 2, 1, 1, 1, and 1 negative significant GCA effects in EL, ear diameter, ear row number, 100-grain weight, and single plant yield, respectively, while there was no significant GCA effect in ear row number, plant height, and ear position. L26 showed a significant negative GCA effect in ear length and single plant yield, which indicated that the GCA significance levels of different traits were different, and some traits showed a certain correlation in the same tested line. It showed that the experimental sample with 8 trait data was experimentally feasible.
[0070] Table 4 Variance components of eight traits in the cross-environment Line×Tester experiment
[0071] Variance components Spike length Thick spike Number of ear rows Number of rows Plant height Ear height Hundred Grain Weight Yield per plant Additive variance 1.62 0.11 1.75 10.9 608.28 113.69 6.54 842.13 Dominant Variance 3.88 0.08 0.56 15.79 235.07 233.46 8.55 1207.06 Environmental variance 0.47 0.02 0.19 3.35 83.97 41.3 4.28 351.82 Broad sense heritability 0.92 0.93 0.92 0.89 0.91 0.89 0.78 0.85 Narrow sense heritability 0.27 0.54 0.7 0.36 0.66 0.29 0.34 0.35
[0072] Additive effects, dominant effects, and environment all have important influences on different phenotypic traits. The genetic characteristics of different traits are revealed by calculating the variance components of each trait. The results show that the additive variance (σ 2 Additive) is greater than the explicit variance (σ 2 Dominance and GCA_SCA ratio were high, indicating that ear diameter, number of rows per ear, and plant height were mainly controlled by additive genes, while ear length, number of grains per row, ear height, 100-grain weight, and yield per plant were mainly controlled by additive genes. 2 Dominance is greater than σ 2 Additive, mainly controlled by dominant genes. Ear diameter and 100-grain weight have the highest and lowest broad-sense heritability, indicating that ear diameter is less affected by the environment, while 100-grain weight is more affected by the environment. Number of ear rows and ear length have the highest and lowest narrow-sense heritability, indicating that number of ear rows has the largest additive effect in the phenotype, while ear length has the smallest additive effect in the phenotype. The eight agronomic traits in this study all showed different genetic characteristics. This shows that the experimental sample with eight trait data is experimentally feasible.
[0073] (3) Calculate the GCA of the yield per plant (GYPP) of each hybrid combination sum and HSGCA for subsequent correlation analysis with HUE-SNPs and HUE-InDels.
[0074] GCA sum The calculation formula is as follows:
[0075] GCA sum =GCA Line +GCA Tester
[0076] The HSGCA calculation formula is as follows:
[0077] HSGCA=GCA Line +SCA
[0078] GCA sum is the sum of the GCA of the hybrid combination, GCA Line GCA is the general compatibility of the combination corresponding to the tested system. Tester GCA is the general combining ability of the corresponding test species. Line and GCA Tester The data used are shown in Table 3. The data used for SCA are shown in Table 5.
[0079] Table 5 SCA and GCA of single plant yield across environments sum ,HSGCA
[0080]
[0081]
[0082] * indicates P < 0.05, which is statistically significant at the probability level. SE: standard error.
[0083] The results show that in SCA, there are 2 combinations that reach a positive significant level, namely L28×YML335 and L6×TRL02, and 4 combinations that reach a negative significant level, namely L6×YML335, L3×S08, L17×S08, and L23×TRL02, among which L28×YML335 is the highest and L23×TRL02 is the lowest. sum Among them, L6×S08 was the highest and L26×YML335 was the lowest. In HSGCA, L6×TRL02 was the highest and L3×S08 was the lowest. The emergence of significant combinations in SCA indicated that there were large differences between SCAs. sum The differences in the highest and lowest combinations in HSGCA indicate the differences among the three combining ability methods, which is beneficial to the subsequent correlation analysis between the three combining ability methods and heterosis.
[0084] 4. Genetic data acquisition
[0085] DNA from the seedling leaves of 37 parents (34 F8 RILs tested lines and 3 inbred lines tested) was extracted by the cetyl trimethylammonium bromide (CTAB) method, and the total DNA of each sample was 1.5 μg, which was used as the input material for DNA sample preparation. The sequencing library was generated using the Truseq Nano DNA HT Sample Preparation Kit (Illumina USA). The constructed library was sequenced using the IlluminaNovaSeq platform to obtain 150 bp end-paired reads with an insert length of approximately 350 bp. The obtained Raw reads were filtered to obtain Clean reads, which were aligned to the maize reference genome B73_RefGen_v4 using the BWA software, with the parameters set to mem-t 4-k 32-M. The software GATK4 and Vcftools were used for SNP detection and filtering, and the parameters were set to varFilter-w5-w10, clusterSize 2clusterWindowSize 5, QUAL<30, QD<2.0, MQ<40, FS>60.0. The ANNOVAR software tool and the maize reference genome B73_RefGen_v4 were used to generate high-quality SNPs and InDels and annotate them.
[0086] 5. Principal Component Analysis
[0087] Using MAF < 0.05, INF > 0.8 as the filtering criteria, 37 highly consistent SNPs of the parents were obtained. The smartPCA program in the EIGENSOFT software package was used for principal component analysis, and the three largest principal components PC1, PC2 and PC3 were selected to obtain a three-dimensional plane cluster diagram. The largest principal component PC1 explained 8.41% of the total variation, the second principal component PC2 explained 7.8%, and PC3 explained 7.07% of the total variation. The results are shown in Figure 1 As shown (the line group in the figure is 34 tested systems, and the Tester group is 3 test types).
[0088] The 34 male parents (RILs) selected from the RIL population showed greater genetic variation and differentiation due to recombination events that occurred over a period of time during selfing, and were roughly divided into three categories. Among them, YML335 and TRL02 were closer to the tested lines, while tropical germplasm S08 was farther away from the tested lines. The tested species were divided into three directions. In order to better show the heterotic groups to which the tested species and the tested lines belonged, the combining ability method was further used to classify the heterotic groups of the tested lines.
[0089] 6. Division of heterosis groups (general combining ability and special combining ability of single plant yield)
[0090] Using the HSGCA method to divide heterotic groups: When a tested line has the smallest (most negative) HSGCA value among the tested varieties of a certain heterotic group, the tested line should belong to that group. The heterotic groups to which all 34 tested lines belong were divided according to the heterotic groups of the tested varieties and the HSGCA of the tested lines. The results are as follows: Figure 2 And as shown in Table 6.
[0091] Table 6 Division of heterotic groups of 3 tested varieties and 34 tested lines
[0092]
[0093] The results showed that 10 tested lines were classified into the Non-Reid (TRL02) heterotic group, 13 tested lines were classified into the Tropical (S08) heterotic group, and 11 tested lines were classified into the Reid (YML335) heterotic group.
[0094] 7. Differences in SNPs and InDels between heterotic groups
[0095] The numbers of parental UE-SNPs and UE-InDels and the numbers of HUE-SNPs and HUE-InDels of offspring combinations were counted for the three heterotic populations (Tropical×Reid, Tropical×Non-Reid, and Reid×Non-Reid) divided by the HSGCA method:
[0096] The parental SNPs filtered in step 4 were analyzed by IBS kinship matrix using TASSEL 5.0 software to obtain the corresponding IBS kinship, and further calculate the GD between the parents of each hybrid combination. The GD calculation formula is as follows:
[0097] GD=1-IBS
[0098] Through step 4, the number of SNPs and InDels of different annotation types of parents is obtained, and the SNPs located in the 2kb region upstream of the gene and the exon region (except synonymous mutations) and the InDels located in the 2kb region upstream of the gene and the exon region are extracted, which are referred to as UE-SNPs and UE-InDels for short. Then, after the sites are extracted, the UE-SNPs and UE-InDels of the parents that are homozygous and different are found according to the combination relationship to obtain the heterozygous UE-SNPs and UE-InDels of the theoretical F1 and count the number, and the molecular markers are referred to as HUE-SNPs and HUE-InDels for short.
[0099] The results are as follows Figure 3As shown in the figure, a: the number in the circle represents the average number of parental UE-SNPs in the three heterotic groups, and the number at the intersection of the circles represents the average number of HUE-SNPs in the hybrid combination between the groups; b: the number in the circle represents the average number of parental UE-InDels in the three heterotic groups, and the number at the intersection of the circles represents the average number of HUE-InDels in the hybrid combination between the groups; c: t-test of GD between the three groups of heterotic patterns. Including Tropical×Reid (T×R), Tropical×Non-Reid (T×NR), Reid×Non-Reid (R×NR). ns, *, *, respectively indicate no statistical significance, statistical significance at the 0.05 probability level, and statistical significance at the 0.01 probability level.
[0100] The results showed that the number of UE-SNPs and UE-InDels in the three heterotic groups was as follows: Tropical>Reid>Non-Reid, indicating that the tropical materials in this experiment had more molecular marker variations than other heterotic groups. In the pairwise hybrid combinations between heterotic groups, the number of HUE-SNPs and HUE-InDels was as follows: Tropical×Reid>Tropical×Non-Reid>Reid×Non-Reid( Figure 3 a, b).
[0101] In GD, no significant difference was found between Tropical×Reid and Tropical×Non-Reid, while Tropical×Reid and Tropical×Non-Reid were significantly higher than Reid×Non-Reid ( Figure 3 c) in.
[0102] The conclusions drawn based on the number of UE-SNPs and UE-InDels indicate that the number of UE-SNPs and UE-InDels is a method for predicting tropical germplasm. The conclusions drawn based on the number of HUE-SNPs and HUE-InDels are consistent with those drawn based on GD, indicating that HUE-SNPs and HUE-InDels can be used to predict the maize heterosis pattern (Tropical×Reid) containing tropical populations.
[0103] 8. Analysis of heterosis
[0104] SPSS23.0 was used to analyze the mean (MEAN), MPH, BPH, SCA, and GCA of single plant yield sum , HSGCA, GD, HUE-SNPs, and HUE-InDels.
[0105] The calculation formulas for hybrid vigor (MPH) and super hybrid vigor (BPH) are as follows:
[0106] MPH = [(F1-MP) / MP]
[0107] BPH=[(F1-BP) / BP]
[0108] Among them, MPH is the hybrid vigor of the middle parent, BPH is the hybrid vigor of the super parent, F1 is the mean of the hybrid, MP is the mean of the two parents of the hybrid, and BP is the value of the high parent of the hybrid.
[0109] Heterosis correlation Figure 4 The value on the lower left of the figure indicates R 2 (Correlation coefficient); The upper right corner of the figure indicates the significant level, * indicates P < 0.05, ** indicates P < 0.01, *** indicates P < 0.001, and is statistically significant at the probability level. The results show that HUE-SNPs, HUE-InDels and other prediction methods (SCA, GCA sum The results of the correlation between HUE-SNPs, HUE-InDels and GD for the heterosis of maize plant yield (MPH, BPH) were consistent with those for HUE-SNPs, HUE-InDels and GD, and showed a significant positive correlation (P≤0.05). Within the molecular marker indicators, HUE-SNPs, HUE-InDels and GD all showed a significant positive correlation, indicating that HUE-SNPs, HUE-InDels and GD are similar to each other and can predict the heterosis of maize plant yield.
[0110] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. Application of the number of SNPs and / or the number of InDels in evaluating heterosis in maize, characterized in that: The SNPs are located in the 2kb region upstream of the whole corn gene and the exon region, and the SNPs are heterozygous SNPs, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the whole corn gene and the exon region, and the InDels are heterozygous InDels; The corn heterosis is the heterosis of single plant yield; The method for obtaining the number of SNPs and / or InDels comprises: extracting parental DNA, generating high-quality SNPs and InDels and annotating them, which are called UE-SNPs and UE-InDels; After the sites of UE-SNPs and UE-InDels are proposed, the UE-SNPs and UE-InDels of the parents that are homozygous and different are found according to the combination relationship, and the heterozygous UE-SNPs and UE-InDels of the theoretical F1 are obtained and the number is counted, which is called HUE-SNPs and HUE-InDels; The number of UE-SNPs and UE-InDels predicts tropical germplasm; The number of HUE-SNPs and HUE-InDels predicted the heterosis pattern of maize containing tropical populations.
2. The use according to claim 1, characterized in that: In corn hybrids, the number of SNPs and / or InDels is positively correlated with corn heterosis.
3. A method for screening a hybrid population of corn with a higher yield than its parents, characterized in that: The following steps are involved: Counting the number of SNPs and InDels of the corn hybrid population to be screened, and selecting the hybrid population with the highest number of SNPs and / or InDels in the corn hybrid population; the SNPs are located in the 2kb region upstream of the gene of the corn full gene and the exon region, and the SNPs are heterozygous SNPs, and the SNPs do not include synonymous mutations; the InDels are located in the 2kb region upstream of the gene of the corn full gene and the exon region, and the InDels are heterozygous InDels; The method for obtaining the number of SNPs and / or InDels comprises: extracting parental DNA, generating high-quality SNPs and InDels and annotating them, which are called UE-SNPs and UE-InDels; After the loci of UE-SNPs and UE-InDels are proposed, the UE-SNPs and UE-InDels that are homozygous and different in the parents are found according to the combination relationship, and the heterozygous UE-SNPs and UE-InDels of the theoretical F1 are obtained and the number is counted, which is the number of SNPs and InDels in the corn hybrid population.
4. The method according to claim 3, characterized in that The parents of the corn hybrid population to be screened are homozygous; the parents are any one or two of a tropical population, a Reid population and a non-Reid population.
5. The method according to claim 3, characterized in that: The number of SNPs of the corn hybrid population to be screened is obtained according to the combination of the number of paternal SNPs and the number of maternal SNPs, and the number of InDels of the corn hybrid population to be screened is obtained according to the combination of the number of paternal InDels and the number of maternal InDels.
6. A method for screening a corn parent population with relatively high yield in offspring, characterized in that: The following steps are involved: Counting the number of SNPs and InDels of the corn parent population to be screened, combining the corn parents to be screened in pairs, obtaining the number of heterozygous SNPs and InDels of the offspring, and selecting the corn parent population combination with the highest number of heterozygous SNPs and / or InDels of the offspring; the SNPs are located in the 2kb region upstream of the gene and the exon region of the corn full gene, and the SNPs do not include synonymous mutations; The InDels are located in the 2kb region upstream of the gene and the exon region of the corn full gene; The method for obtaining the number of SNPs and / or InDels comprises: extracting parental DNA, generating high-quality SNPs and InDels and annotating them, which are called UE-SNPs and UE-InDels; After the sites of UE-SNPs and UE-InDels are proposed, the UE-SNPs and UE-InDels that are homozygous and different in the parents are found according to the pairing relationship, and the heterozygous UE-SNPs and UE-InDels of the theoretical F1 are obtained and the number is counted, which is the number of heterozygous SNPs and / or InDels in the offspring.
7. The method according to claim 6, characterized in that The corn parent population is homozygous.
8. The method according to claim 6, characterized in that The number of SNPs and the number of InDels of the corn parent population to be screened are obtained based on the whole genome sequence of the corn parent population.
9. The method according to claim 6, characterized in that The corn parent population to be screened is any one or two of a Tropical population, a Reid population and a Non-Reid population.
10. Use of the method according to any one of claims 3 to 9 in corn hybrid breeding.
Citation Information
Patent Citations
Methods of creating drought tolerant corn plants and compositions thereof
US20140130211A1
Rapmap method for rapid and high-throughput positioning and cloning of plant QTL gene
WO2021196255A1