Training group efficient construction method for waxy corn whole genome selective breeding
By constructing an efficient whole-genome selection breeding training population for waxy maize, the problem of insufficient representativeness of training populations in existing technologies has been solved, enabling efficient and accurate breeding prediction and material screening under limited resources.
Patent Information
- Application Number
- CN202511110324.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-12-30
AI Technical Summary
The existing training populations used in whole-genome selection breeding of waxy maize only contain hybrid combinations of a few backbone inbred lines, which lacks representativeness and leads to a decline in the predictability of newly introduced germplasm or germplasm with distant genetic relationships.
We extensively collected commonly used or genetically improved inbred lines or DH lines in breeding, performed simplified genome sequencing and phylogenetic tree analysis, divided multiple heterosis groups, selected representative inbred lines, used sparse partial diallel hybridization design to breed hybrids, conducted multi-environment planting identification, used mixed linear models to predict the phenotype of hybrids, and screened parental inbred lines with high combining ability.
While ensuring that each inbred line participates in hybridization, the number of hybrids required for fusion is reduced, saving resources and costs, while improving the predictive ability for actual breeding materials and the representativeness of genetic variation.
Smart Images

Figure CN121237211A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of genetic breeding technology, specifically to a method for efficiently constructing training populations for whole-genome selection breeding of waxy maize. Background Technology
[0002] Genomic selection (GS) is a method that establishes associations between marker genotypes and phenotypes based on known molecular marker genotype and phenotype information on the genome of a population. It simultaneously estimates the effects of all markers across the entire genome and makes reasonable predictions for populations with unknown phenotypes. This provides a new method for plant and animal breeding and has been successfully applied.
[0003] Currently, almost all waxy corn production utilizes single-cross hybrids derived from two inbred lines. In conventional breeding models, large-scale hybridization experiments to screen for strong hybrid combinations are a common approach. However, due to constraints such as labor costs, funding, and land area, the breeding of a waxy corn hybrid often takes 6-8 years. Applying the GS method to predict waxy corn hybrids only requires whole-genome sequencing analysis of the inbred lines or DH lines (double haploid systems) currently used by breeders to obtain genotypic information for each inbred line or DH line. This allows for the deduction of genotypic information for all possible hybrids obtained from pairwise hybridization of these inbred lines or DH lines. Furthermore, a genetic model constructed using a training population can then be used to predict the possible phenotypes of these hybrids.
[0004] In genome-wide selection breeding, the training population serves as the bridge between genotype and phenotype. The quality of training population construction is a core factor determining the accuracy of predictive models and directly impacts breeding efficiency. Constructing an efficient training population requires comprehensive consideration of factors such as trait heritability, inter-population correlation, marker density, breeding objectives, and resources. For traits with high heritability, such as plant height, gene effects are more pronounced, and key genetic information can be captured with fewer individuals, allowing for a relatively small training population. If the training population and the prediction population are genetically closely related, such as closely related populations with similar gene compositions, a smaller training population can achieve high predictive accuracy. Conversely, for distantly related populations, due to greater genetic differences, a larger training population is needed to encompass sufficient genetic variation. Higher marker density covers a wider genomic range, enabling more precise capture of gene-trait associations. When marker density is sufficiently high, the training population size can be appropriately reduced. If the breeding objective is to improve a single trait, the training population size can be relatively small; if the breeding objective involves multiple complex quantitative traits, a larger training population is required. Different species have different genome sizes, genetic diversity, and reproductive characteristics, resulting in varying training population sizes. Generally, species with complex genomes and high genetic diversity require larger training populations. For example, crops like corn and wheat, due to their large planting areas and abundant genetic resources, typically require large-scale training populations. At the same time, human, material, and financial resources must be considered. Phenotyping and genotyping of large-scale populations are costly; therefore, a reasonable population size must be determined based on available resources, balancing costs and benefits, while ensuring a certain level of predictive accuracy.
[0005] The performance of waxy maize hybrids depends on the genetic complementarity between parents, and the genetic diversity of the training population needs to cover the core breeding germplasm and focus on the target breeding traits. Existing training populations used in whole-genome selection breeding of waxy maize only contain hybrid combinations of a few backbone inbred lines, which lacks representativeness and leads to a decline in predictability for newly introduced germplasm or germplasm with distant genetic relationships. Summary of the Invention
[0006] To overcome the shortcomings of existing technologies, a method for efficiently constructing training populations for whole-genome selection breeding of waxy maize is provided. This method addresses the problem that existing training populations used in whole-genome selection breeding of waxy maize only contain hybrid combinations of a few backbone inbred lines and lack representativeness, leading to a decline in predictability for newly introduced germplasm or germplasm with distant genetic relationships.
[0007] To achieve the above objectives, a method for efficiently constructing training populations for genome-wide selection breeding of waxy maize is provided, comprising the following steps:
[0008] Collect widely used or genetically improved inbred lines or DH lines as breeding materials;
[0009] Simplified genome sequencing or resequencing of the breeding materials was performed to obtain high-quality SNP markers, which were then used to perform phylogenetic tree and population structure analysis on the multiple maize inbred lines.
[0010] Based on the phylogenetic tree and combined with pedigree information and population structure analysis, multiple heterotic groups were obtained;
[0011] Several representative inbred lines were selected from the aforementioned heterotic groups;
[0012] Multiple hybrids were prepared by using a sparse partial diallel hybridization design method among different heterotic groups on the aforementioned representative inbred lines;
[0013] Multiple hybrids were subjected to multi-environment planting identification to obtain multiple trait data;
[0014] Predict the phenotypes of all possible hybrids formed by combining the aforementioned representative inbred lines;
[0015] By estimating the general combining ability effect of each representative inbred line based on the predicted phenotype of the hybrid, parental inbred lines with high general combining ability are screened for breeding applications.
[0016] Furthermore, using the BLUP values of the multiple trait data as phenotypes, a mixed linear model is employed to predict the phenotypes of all possible hybrids formed by combining the multiple representative inbred lines.
[0017] Furthermore, the multiple trait data include fatty acid content, protein content, husk residue rate, lysine content, and amylopectin content under various environmental conditions.
[0018] The beneficial effect of this invention is that the efficient construction method for training populations in whole-genome selection breeding of waxy maize can fully consider representative parental materials of different heterotic groups, resulting in greater genetic variation and more representativeness of the training population.
[0019] The efficient construction method for training populations in whole-genome selection breeding of waxy maize of the present invention requires the least number of hybrids to be combined under the condition that the scale of parental materials is fixed and that each inbred line participates in hybridization. This method can effectively reduce the scale of the experiment and save resources and costs to the maximum extent, while maintaining sufficient information.
[0020] In actual breeding, target varieties are almost all hybrid combinations between different heterotic groups. The efficient construction method of training population for whole-genome selection breeding of waxy maize in this invention is based directly on the core heterotic groups in breeding, which significantly improves the model's predictive ability for actual breeding materials. Attached Figure Description
[0021] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0022] Figure 1 In this embodiment of the invention, phylogenetic trees were constructed for 260 breeding materials.
[0023] Figure 2 This is a schematic diagram of a sparse partial diallel cross design method between different heterotic groups according to an embodiment of the present invention.
[0024] Figure 3 This is a schematic diagram showing the number of hybridizations involving 260 inbred lines in an embodiment of the present invention.
[0025] Figure 4 This invention provides field validation selection gain analysis of top80 and bottom80 predictions using three GS methods (BayesB, GBLUP, and LSSO) for embodiments of the present invention. Detailed Implementation
[0026] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0027] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0028] This invention provides a method for efficiently constructing training populations for genome-wide selection breeding of waxy maize, comprising the following steps:
[0029] S1. Widely collect inbred lines or DH lines (Double Haploid lines are formed by doubling haploid cells through natural or artificial chromosomes) commonly used in waxy corn breeding or genetically improved.
[0030] S2. Based on the simplified genome sequencing results of the breeding inbred lines, multiple SNP markers were obtained, and phylogenetic tree and population structure analysis were performed on multiple waxy maize inbred lines.
[0031] S3. Based on phylogenetic trees and combined with pedigree information and population structure analysis, multiple heterotic groups were divided.
[0032] S4. Select several representative inbred lines from multiple heterotic groups.
[0033] S5. Multiple hybrids were prepared by using a sparse partial diallel cross design method between different heterotic groups on multiple representative inbred lines.
[0034] S6. Conduct multi-environment planting identification of multiple hybrids to obtain multiple trait data.
[0035] Multiple trait data include fatty acid content, protein content, husk residue rate, lysine content, and amylopectin content under various environmental conditions.
[0036] S7. Predict the phenotypes of all possible hybrids formed by combining multiple representative inbred lines.
[0037] Specifically, using the BLUP values of multiple trait data as phenotypes, a mixed linear model was used to predict the phenotypes of all possible hybrids from combinations of multiple representative inbred lines.
[0038] S8. Estimate the general combining ability effect of each representative inbred line by predicting the phenotype of the hybrid, and screen out parental inbred lines with high general combining ability for breeding applications.
[0039] This invention provides an efficient method for constructing training populations for whole-genome selection breeding of waxy maize, establishing a highly efficient genetic design framework for training populations in whole-genome selection of waxy maize hybrids. This method fully considers representative parental materials from different heterotic groups, resulting in greater genetic variation and representativeness in the training population. With a given number of parental materials, this method minimizes the number of hybrids required while ensuring that each inbred line participates in hybridization, thus maximizing resource and cost savings.
[0040] To further illustrate the implementation of the method for efficiently constructing training populations for whole-genome selection breeding of waxy maize according to the present invention, the following examples are provided.
[0041] Example 1
[0042] This embodiment provides a method for efficiently constructing a training population for whole-genome selection breeding of waxy maize, including the following steps:
[0043] S10. Collect widely used or genetically improved inbred lines in waxy corn breeding.
[0044] First, a breeding population of 400 waxy maize inbred lines was constructed.
[0045] S20. Based on the simplified genome sequencing results of the breeding inbred lines, multiple SNP markers were obtained to construct phylogenetic trees for multiple waxy maize inbred lines.
[0046] From the breeding inbred lines constructed in step S10, 176,000 SNP markers were obtained based on simplified genome sequencing results, which were used to construct phylogenetic trees and analyze population structure of 400 breeding materials.
[0047] S30. Based on phylogenetic trees and combined with pedigree information and population structure analysis, multiple heterotic groups were obtained.
[0048] Based on the phylogenetic tree and combined with pedigree information and population structure analysis, the results show that it is divided into five heterotic groups: Tongxiu 5 groups, Hengbai 522 group, T2 group, PW group and mixed group.
[0049] S40. Select several representative inbred lines from multiple heterotic groups.
[0050] Representative inbred lines were selected from each heterosis group, numbering 77, 68, 36, 12, and 67 respectively, for a total of 260 inbred lines (e.g., ...). Figure 1 (As shown).
[0051] S50. Multiple hybrids are prepared by using a sparse partial diallel cross design method between different heterotic groups on multiple representative inbred lines.
[0052] The design method of sparse partial diallel cross between different heterotic groups (e.g.) Figure 2 As shown in the diagram, solid circles represent proposed hybrids, and hollow circles represent hybrids to be predicted. A total of 770 hybrids were prepared. Due to the influence of field environmental conditions and flowering time, these 770 hybrids are not entirely in accordance with... Figure 2 The sparse portion of the double-column hybridization table shown is used to construct the hybrids. As shown in Table 1, the actual number of hybrids obtained is relatively small compared to the theoretically possible number of hybrids. Each inbred line participates in at least one hybridization, with an average of 5.9 hybridizations per inbred line (e.g., ...). Figure 3 (As shown).
[0053] Table 1. Actual composition of different heterotic groups in 770 hybrid populations
[0054]
[0055] Note: The number outside the parentheses represents the actual number of hybrids obtained, and the number inside the parentheses represents the theoretical number of hybrids obtained. S60. Multiple hybrids were planted and identified to obtain multiple trait data.
[0056] As a preferred implementation method, multiple trait data include fatty acid content, protein content, peel residue rate, lysine content, and amylopectin content under various environmental conditions.
[0057] Specifically, the aforementioned 945 maize hybrids were planted and identified in Rugao and Haimen, Jiangsu Province, and data on five quality traits, including fatty acid content, protein content, bran content, lysine content, and amylopectin content, were collected from the two environments.
[0058] S70, predict the phenotypes of all possible hybrids from combinations of multiple representative inbred lines.
[0059] Specifically, using the BLUP values of multiple trait data as phenotypes, a mixed linear model was used to predict the phenotypes of all possible hybrids from combinations of multiple representative inbred lines.
[0060] In this embodiment, to better facilitate subsequent analysis and minimize the impact of environmental effects on genetic effects, the BLUP values of five quality-related traits from 770 hybrids were used as phenotypic data for further analysis. The genotypes of all possible hybrids (n = 230 × 229 × 0.5 = 26335) derived from the corresponding parental genotypes were extrapolated. Cross-validation was performed on the five quality-related traits of waxy maize using three commonly used GS methods: GBLUP, LASSO, and BayesB, and the predictive accuracy of the three methods was compared. The results showed that the predictive power of the five quality traits ranged from 0.454 to 0.726 (Table 2), with the husk percentage exhibiting the highest predictive power among all methods, while lysine content showed the lowest predictive power. Using 770 hybrids as the training population, the protein content (EW) values of all 26,335 possible hybrids were predicted using three genetic sequencing (GS) methods: GBLUP, LASSO, and BayesB. The top 80 hybrids with the highest EW and the bottom 80 hybrids with the lowest EW were selected from each of the three GS methods. Comparison revealed that the top 80 and bottom 80 hybrids selected by the three GS methods included 108 and 94 hybrids, respectively. Figure 4(a) and (c) are mentioned, where 55 and 65 hybrids in the top 80 and bottom 80 respectively were selected using three GS methods, showing a high degree of consistency. The predicted top 80 hybrids mainly involve 35 parental inbred lines, including 18 from the Tongxi 5 group, 10 from the Hengbai 522 group, 2 from the mixed group, 3 from the T2 group, and 2 from the PW group. Each inbred line participated in at least one hybridization, with W10W361 participating in the most hybridizations (9 times). The predicted top 80 hybrids were crossbred in Hainan Nanfan in 2021, successfully producing 41 hybrids for field verification. The predicted bottom 100 hybrids mainly involve 30 inbred lines, including 7 from the Tongxi 5 group, 6 from the Hengbai 522 group, 6 from the mixed group, 10 from the T2 group, and 1 from the PW group. Each inbred line participated in at least one hybridization, with inbred line T21 participating in the most hybridizations (6 times). The predicted bottom80 hybrid was crossbred in Hainan's southern breeding program in 2021, and 52 hybrids were successfully crossbred and field-planted for verification. Figure 4 b, d).
[0061] Table 2. Prediction accuracy of five quality-related traits using three statistical methods in 770 hybrid populations.
[0062]
[0063] Note: Letters indicate the significance level of multiple comparisons.
[0064] The mean protein content estimates for all potential hybrids validated in the field using the BayesB, GBLUP, and LASSO methods were 9.1, 9.4, and 9.4, respectively. The top hybrids validated in the field using the three GS methods showed protein content increases of 101.4%, 101.3%, and 102.5% higher than the bottom hybrids, respectively. Figure 4 e) The protein content was 34.3%, 35.6%, and 34.4% higher than the average of all potential hybrids, respectively. The average protein content of the control variety, Suyunuo 2, was 10.1%. Among the top hybrid combinations verified in the field, 16 waxy corn hybrids had protein content higher than the control variety, and 7 hybrids had protein content higher than the control variety by more than 5%.
[0065] Furthermore, among the 41 top hybrids and 52 bottom hybrids validated in the field, 26 top hybrids and 32 bottom hybrids were detected by all three methods. The average protein content of the 41 field-validated top hybrids was 13.6, which was 109.2% higher than that of the 52 field-validated bottom hybrids (6.5). Comparative analysis revealed that the average protein content of the 26 top hybrids detected by all three methods (13.8) was 106.0% higher than that of the 32 bottom hybrids (6.7), which is similar to the results obtained by the three prediction methods. This result also indicates that the selection gain cannot be improved by using the three prediction methods to screen the intersection of the hybrid sets.
[0066] S80. Estimate the general combining ability effect of each representative inbred line by predicting the phenotype of the hybrid, and screen out parental inbred lines with high general combining ability for breeding applications.
[0067] The general combining ability (GBA) effect of each inbred line was estimated by predicting the phenotype of the hybrids. Then, parental inbred lines with higher GBA were selected.
[0068] Based on the training population, we used a mixed linear model to predict the phenotypes of all possible hybrids. Then, based on genomic information, we used a mixed linear model to accurately predict the GCA values (general combining ability) of five quality-related traits in 230 inbred lines. The mean GCA values of these five quality traits were all 0, and all showed a normal distribution with a large range of variation (Table 3). Analysis of variance of GCA for each trait in each heterotic group showed that for fatty acid content and protein content, there were no significant differences at the 0.05 level between the five Tong-line groups, the Hengbai 522 group, the T2 group, and the PW group. There were also no significant differences between the PW group and the mixed group (P<0.05). However, there were significant differences between the five Tong-line groups, the Hengbai 522 group, the T2 group, and the mixed group (Table 3). For the GCA of other traits, there were also significant differences among different heterotic groups. For example, among the five groups of Tongxi, the four traits of fatty acid content, protein content, husk residue rate, and lysine content showed the highest GCA values, while the GCA value of amylopectin content was the lowest. This indicates that the five groups of Tongxi have great breeding potential in terms of fatty acid content, protein content, and husk residue rate, but they have the potential to reduce amylopectin content. In breeding practice, the superior inbred lines of this group can be used for backcrossing and breeding to improve them.
[0069] Table 3. Analysis of variance of GCA for 9 traits in different heterotic groups
[0070]
[0071] Compared with existing NCII or partial diallel cross designs (almost all waxy maize production uses single crosses derived from two parent-cross lines), the efficient construction method of training populations for whole-genome selection breeding of waxy maize in this invention can fully consider representative parental materials of different heterotic groups, resulting in greater genetic variation and more representativeness of the training population.
[0072] The efficient construction method for training populations in whole-genome selection breeding of waxy maize of the present invention requires the least number of hybrids to be combined under the condition that the scale of parental materials is fixed and that each inbred line participates in hybridization. This method can effectively reduce the scale of the experiment and save resources and costs to the maximum extent, while maintaining sufficient information.
[0073] In actual breeding, target varieties are almost all hybrid combinations between different heterotic groups. The efficient construction method of training population for whole-genome selection breeding of waxy maize in this invention is based directly on the core heterotic groups in breeding, which significantly improves the model's predictive ability for actual breeding materials.
[0074] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A method for efficient construction of training population for whole genome selection breeding of waxy maize, characterized in that, The method comprises the following steps: a. Collecting a large number of inbred lines or DH lines commonly used or genetically improved in waxy maize breeding; b. Simplified genome sequencing or re-sequencing of the inbred lines or DH lines to obtain high-quality SNP markers, and phylogenetic tree and population structure analysis of the inbred lines or DH lines using the obtained high-quality SNPs; c. Dividing each inbred line or DH line into different heterotic groups based on the phylogenetic tree and combining pedigree information and population structure analysis results; d. Selecting a plurality of representative inbred lines from each of the different heterotic groups; e. Preparing a plurality of hybrid seeds by using a sparse partial diallel cross design method among different heterotic groups; f. Obtaining a plurality of trait data by planting and identifying the plurality of hybrid seeds in multiple environments; g. Predicting phenotypes of all possible hybrid seeds formed by the plurality of representative inbred lines; h. Estimating general combining ability effects of each of the representative inbred lines by the predicted phenotypes of the hybrid seeds, and screening parent inbred line materials with high general combining ability for breeding application.
2. The method for constructing a training population for whole genome selection breeding of waxy maize according to claim 1, characterized in that, The BLUP value of the plurality of trait data is used as the phenotype, and a mixed linear model is used to predict the phenotypes of all possible hybrid seeds formed by the plurality of representative inbred lines. 3.The method for constructing a training population for maize whole genome selection breeding according to claim 1, wherein, The plurality of trait data includes fatty acid content, protein content, skin and core rate, lysine content, and amylopectin content under multiple environmental conditions.