Crop breeding optimization combination design method based on multi-character whole genome selection index model
By constructing a multi-trait selection index model based on the whole-genome QTL-allele matrix, the problem of low efficiency in multi-trait selection in existing technologies is solved, and the optimal combination design of multiple traits in crop breeding and accurate prediction of breeding objectives are realized, thereby improving breeding efficiency and genetic diversity.
Patent Information
- Application Number
- CN202510949330.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2026-02-06
AI Technical Summary
Existing multi-trait selection methods in crop breeding suffer from low efficiency, reduced genetic diversity, and difficulty in achieving breeding goals. In particular, under complex genetic structures and multi-trait conditions, it is difficult to obtain ideal genotype individuals using existing methods.
Based on the whole-genome QTL-allele matrix, computer simulation technology is used to generate homozygous population genotypes of offspring from any parental combination, construct a multi-trait selection index model, and optimize the design of parental combinations by integrating multi-trait breeding information to achieve comprehensive evaluation of multiple traits.
It improves breeding efficiency, enables more accurate prediction of the performance of multiple traits in individuals, achieves optimized design of parental combinations, meets the goals of multi-trait breeding, and improves the targeting and efficiency of breeding.
Smart Images

Figure CN121483377A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular quantitative genetics and molecular breeding technology, specifically relating to a method for constructing a multi-trait index model based on the optimized combination design of crop hybridization breeding. Background Technology
[0002] Crop breeding, as a primary means of increasing grain yield and improving grain quality, has long been considered an important pathway for crop genetic improvement and new variety development due to its provision of abundant genetic variation. Pureline breeding involves two key steps: first, selecting parental individuals and generating a population of offspring with variations; second, selecting superior materials from the offspring population and upgrading them into new varieties. The generation of new varieties is inseparable from selection, which not only expands the degree of variation in a particular trait within a population but also achieves comprehensive improvement of multiple traits. In practical breeding work, researchers usually need to focus on both yield and quality simultaneously; therefore, comprehensive improvement of multiple traits is a crucial task in breeding.
[0003] Compared to selecting multiple traits individually, multi-trait comprehensive selection is more efficient and effective. Methods for improving multiple traits include sequential selection, independent elimination, and index selection. Sequential selection involves selecting only one trait within a certain timeframe until it reaches the breeding objective, after which other traits are improved. Independent elimination sets minimum standard values for all traits to be improved, retaining only individuals that meet the standards in all traits. Although these two methods are widely used in multi-trait selection, they each have significant limitations. While sequential selection is simple and easy to implement, it may lead to some traits being overlooked during the improvement process, thus affecting the overall breeding effect; independent elimination may result in the elimination of superior individuals due to overly strict selection criteria, reducing genetic diversity. Furthermore, since different traits contribute differently to the breeding objective, the rational allocation of weights among traits is particularly important.
[0004] To address these issues, the index selection method was proposed. This method integrates multiple traits requiring improvement into a quantitative index reflecting the breeding value of multiple traits through linear combination, and then selects individuals based on the selection index. The index selection method has gradually become an important method for breeding superior varieties in plant and animal breeding and has been widely used in crop breeding. With the development of genetics and high-throughput sequencing technologies, combining a large number of molecular markers with phenotypic data for QTL mapping has become a common method in quantitative trait genetic research, which has also brought new opportunities for the application of selection indices. In recent years, the combination of genomic selection (GS) and selection indices has gradually attracted attention. However, current multi-trait selection methods mainly focus on the selection of offspring populations, with relatively limited research on parental pairing. When the genetic structure of breeding traits is complex and the number of traits is large, obtaining individuals with ideal genotypes may face certain difficulties. In this case, larger-scale populations and more individuals are usually required for hybridization to aggregate multiple superior allelic variations and achieve the expected breeding goals.
[0005] To improve breeding efficiency, based on genetic analysis of breeding traits, a design breeding approach was adopted, and genome-wide QTL-allele-based parental pairing prediction was conducted. Computer simulation technology was used to fully explore existing genetic information and reasonably predict the phenotypic expression of homozygous individuals in hybrid offspring, providing a theoretical basis and breeding guidance for combinatorial design and hybridization breeding. Field phenotypic identification of local soybean varieties in China shows that these populations still have significant variation space in multiple traits, possessing significant improvement potential. Previous studies have used genome-wide association analysis to genetically analyze multiple breeding traits in resource populations and established QTL-allele matrices, screening out a batch of superior combinations that may surpass their parents in genotype based on offspring simulation results. However, existing combination selection methods mainly focus on single-trait improvement, while actual breeding processes usually require comprehensive consideration of the performance of multiple traits. Therefore, research on parental pairing prediction based on multiple traits still has important practical significance. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a method for constructing a multi-trait index model based on optimized combination design in crop hybrid breeding. This method utilizes the QTL-allele matrix of multiple traits in the parental population and employs computer simulation technology to generate homozygous genotypes of offspring from any parental combination. This allows for accurate estimation of predicted values for multiple traits in individuals, thereby assessing the breeding potential of the parental combination. Furthermore, a multi-trait selection index model is constructed, integrating multi-trait breeding information to comprehensively evaluate the breeding potential of the parental combination, achieving optimized design of the parental combination and providing more targeted and forward-looking guidance for breeding work.
[0007] The technical solution provided by this invention is as follows:
[0008] A method for constructing a multi-trait index model based on crop hybrid breeding optimization combination design includes the following steps:
[0009] Genome-wide SNP genotype data and breeding-related phenotypic data were obtained, and QTL-allele matrices for multiple traits were obtained using genome-wide association analysis.
[0010] Based on the QTL-allele matrix of multiple traits in the germplasm resource population, simulated offspring of all possible hybridization combinations within the crop population are obtained;
[0011] Calculate the genotype and phenotypic predicted values of the simulated offspring of the hybrid combinations; evaluate the breeding potential of different hybrid combinations by specifying the percentile of the phenotypic predicted values after the hybrid combination simulation, and screen out high-quality hybrid combination simulated offspring.
[0012] Based on the high-quality combination of hybrid combinations to simulate offspring, a multi-trait selection index is constructed for offspring individuals to evaluate the breeding performance of multiple traits.
[0013] Furthermore, obtaining all possible hybrid combinations of simulated offspring within the crop population includes: generating multiple homozygous offspring individuals through continuous self-pollination of F1 seed to form the offspring population of the parental combination.
[0014] Furthermore, the method for calculating the genotype value includes:
[0015] (1)
[0016] in, Let be the genotype estimate of the i-th individual; K is the number of QTLs. This is a genotype indicator variable for the i-th individual at the k-th locus. The value is 0 if the individual does not possess the allele, and 1 if the individual possesses the allele. This represents the effect of the l-th allele at the k-th locus.
[0017] Furthermore, the method for calculating the phenotypic predicted value includes:
[0018] (2)
[0019] in, The phenotypic predicted value for the r-th offspring produced by the cross between parent i and parent j. The group mean Let be the genotype estimate of the r-th offspring individual. and These are the phenotypic observations of the two parents. and These are the estimated genotype values for the two parents.
[0020] Furthermore, methods for calculating the multi-trait selection index of offspring individuals include:
[0021] (3)
[0022] in, The selection index is used for comprehensive selection, where T represents the number of traits. Let t be the weight of the t-th trait, and let the weight satisfy t = t / t. ; t represents the standardized phenotypic predicted value of the t-th trait.
[0023] Furthermore, methods for calculating phenotypic predicted values of traits include:
[0024] (4)
[0025] in, For the t-th trait, the standardized phenotypic predicted value is... Let t be the phenotypic value of the t-th trait. Let t be the standardized location parameters of the t-th trait. Let be the scaling parameter for the t-th trait, which is standardized.
[0026] Furthermore, the breeding performance of multiple traits can be evaluated based on the quantiles of the multi-trait selection index of the offspring population. The larger the index value, the better the overall performance of multiple traits of the parent combination.
[0027] Furthermore, the assessment of the breeding potential of different hybrid combinations by specifying the percentile of the phenotypic predicted values after simulation of the hybrid combination includes: using the percentile of the phenotypic predicted values of the offspring population of the parental combination as a comparison index, selecting the percentile with the larger percentage for high-value directional traits, and selecting the percentile with the smaller percentage for low-value directional traits.
[0028] This invention also provides a multi-trait index model based on the optimized combination design of crop hybrid breeding, which is obtained by the above method.
[0029] The present invention also provides the application of the above-mentioned multi-trait index model based on crop hybridization breeding optimization combination design in crop breeding.
[0030] Beneficial effects
[0031] To verify the effectiveness of the multi-trait optimal combination design method, this study used two traits—100-seed weight (100SW) and soy oil content (SOC)—of a local soybean variety population in China as examples to construct a multi-trait comprehensive selection index. Based on the breeding objective, the weights were divided into four types: both traits were selected at high values; one trait was selected at a high value and the other at a low value; and both traits were selected at low values. Nine different weight combinations were listed for each type. A positive weight indicates that the breeding objective is to select a high value for that trait, while a negative weight indicates the selection of a low value. The absolute value of the weight measures the relative importance of the trait; the larger the absolute value, the more important the trait is in the selection process.
[0032] When selecting high-value 100-grain weight and oil content, as the weight of 100-grain weight increases, the mean 100-grain weight of the selected superior offspring continuously increases; conversely, as the weight of oil content decreases, the mean of the offspring decreases. Furthermore, to demonstrate the effectiveness of the index selection, comparing the mean values of the 100 individuals with the highest and lowest indices for both traits reveals that, under any weighting condition, the 100 individuals with the highest indices outperform the 100 individuals with the lowest indices in both 100-grain weight and oil content. This pattern also holds true under the other three weighting conditions.
[0033] When selecting high-value 100-seed weight and low-value oil content, as the weight of 100-seed weight increases, the average 100-seed weight of superior offspring individuals continuously increases; as the weight of oil content decreases, the average oil content of superior offspring individuals gradually increases. When the weight of oil content is the smallest, the average oil content of the selected individuals is the largest, and when the weight is the largest, the average oil content of the individuals is the lowest. This trend is consistent with the breeding goal of low oil content.
[0034] When selecting individuals with low 100-grain weight and those with high oil content, the average 100-grain weight of individuals with the highest index was lower than that of individuals with the lowest index, while the average oil content was higher. As the weight of 100-grain weight increased, the average 100-grain weight of superior offspring individuals continued to decrease; conversely, as the weight of oil content decreased, the average oil content of superior offspring individuals also continued to decrease.
[0035] Finally, when selecting individuals with low 100-grain weight and low oil content, the 100 individuals with the highest indices all had lower mean values for both traits than the individuals with the lowest indices, which aligns with the breeding objectives of low 100-grain weight and low oil content. This demonstrates that the offspring performance of the index selection method, under different weight settings, meets the expectations of each breeding objective. Attached Figure Description
[0036] Figure 1 This is a flowchart of a multi-trait parental optimization combination design method based on QTL-allele. Detailed Implementation
[0037] The invention will be further described below with reference to the accompanying drawings.
[0038] Example 1
[0039] like Figure 1 The diagram shows a flowchart of a multi-trait parental optimization combination design method based on QTL-allele, outlining the implementation process of this method, which mainly includes the following four core steps:
[0040] (1) Constructing a genetic information matrix: Obtain whole-genome SNP genotype data and breeding-related phenotypic data, and construct a QTL-allele genetic information matrix for multiple traits using genome-wide association analysis;
[0041] (2) Simulated hybrid offspring: Simulates offspring from all possible hybrid combinations within the population. Users can specify a particular parental combination and set the number of simulated offspring individuals according to actual needs;
[0042] (3) Calculate predicted values and assess breeding potential: Calculate the genotype and phenotypic predicted values of the simulated offspring. By specifying the percentile of the predicted phenotypic values of the offspring population, assess the breeding potential of different hybrid combinations to provide a basis for screening high-quality combinations;
[0043] (4) Constructing a multi-trait selection index: Based on the calculation results in step 3, construct a multi-trait selection index for offspring individuals. Users can choose whether to standardize the selection. If standardization is selected, a specific standardization method (such as STD or RANGE) must be specified to evaluate the breeding performance of multiple traits.
[0044] Example 2
[0045] (1) QTL-allele analysis of breeding traits
[0046] Based on the germplasm resource population, a restricted two-stage genome-wide association analysis (GWAS) method was used to perform GWAS on multiple breeding traits, obtaining QTLs (SNPLDB markers) significantly associated with each trait and the allelic effects at each locus. This genetic information within the population was summarized into a QTL-allele matrix: each column represents an individual, and each row represents the effect of a locus on all individuals. A complete QTL-allele matrix encompasses all the basic genetic information of the breeding trait. For multiple traits, the same method needs to be used to construct QTL-allele matrices for multiple breeding-related traits.
[0047] (2) Simulation of offspring populations of parental combinations
[0048] Based on the QTL-allele matrix of multiple traits in the germplasm resource population, computer simulations were used to generate the offspring population of all possible hybrid combinations within the population. Assuming the sample size of the resource population is n, there are n(n-1) / 2 possible parental combinations (single crosses). For any combination, 2000 homozygous offspring individuals (determined based on actual breeding scale) were generated through continuous self-pollination of F1 individuals, forming the offspring population of that parental combination. Without considering linkage between loci (independent model), all loci on the same chromosome are independent, and alleles assort independently. Considering linkage between loci (linkage model), it is assumed that the number of crossing over during meiosis follows a Poisson distribution, with the distribution parameter λ being the chromosome length (in Morgan units), and the crossing over sites are uniformly distributed. Genotypes of offspring individuals were generated through simulation using parental gametes.
[0049] (3) Individual genotype value estimation
[0050] Based on a single QTL-allele matrix, the genotype estimate of the i-th individual is:
[0051] (1)
[0052] in, Let be the genotype estimate of the i-th individual; K is the number of QTLs. This is a genotype indicator variable for the i-th individual at the k-th locus. The value is 0 if the individual does not possess the allele, and 1 if the individual possesses the allele. This represents the effect of the l-th allele at the k-th locus.
[0053] The predicted phenotypic value of the r-th offspring produced by the cross between parent i and parent j is:
[0054] (2)
[0055] in, The phenotypic predicted value for the r-th offspring produced by the cross between parent i and parent j. The group mean Let be the genotype estimate of the r-th offspring individual. and These are the phenotypic observations of the two parents. and These are the estimated genotype values for the two parents.
[0056] (4) Construction of the multi-trait selection index model
[0057] To integrate breeding information across multiple traits, a multi-trait selection index is constructed to represent the overall performance of offspring individuals, thereby selecting parental combinations that excel in multiple traits or meet breeding objectives. The weighted multi-trait selection index for offspring individuals is as follows:
[0058] (3)
[0059] in, The selection index is used for comprehensive selection, where T represents the number of traits. Let t be the weight of the t-th trait, and let the weight satisfy t = t / t. The absolute value of the weight reflects the relative importance of the trait, and the sign of the weight reflects the method of trait selection (positive weight indicates the direction of high value, and negative weight indicates the direction of low value). Let t be the standardized phenotypic predicted value of the t-th trait. The standardized phenotypic predicted value is:
[0060] (4)
[0061] in, For the t-th trait, the standardized phenotypic predicted value is... Let t be the phenotypic value of the t-th trait. Let t be the standardized location parameters of the t-th trait. Let be the scaling parameter for the t-th trait, where location and scaling parameter are the same for all combinations. There are various standardization methods, with standardization based on arithmetic mean and standard deviation (STD) and standardization based on minimum and range (RANGE) being the most commonly used.
[0062] (5) Evaluation of the recombination potential of parental combinations
[0063] For any parental combination, the predicted phenotypic values of its offspring population have a specific distribution, making direct comparison of the breeding potential of different parental combinations impossible. This study uses quantiles of the predicted phenotypic values of the offspring population as a comparison indicator. Quantiles with higher phenotypic values represent larger percentages of selected traits, while those with lower values represent smaller percentages. For example, the 90th percentile reflects the top 10% of individuals in the offspring population with the highest phenotypic values; therefore, different quantiles also reflect the intensity of selection. For multiple traits, this study uses quantiles of the multi-trait selection index of the offspring population; the larger the index value, the better the overall performance of the parental combination across multiple traits.
[0064] Example 3
[0065] A multi-trait optimization design was conducted on local Chinese soybean varieties to optimize 100-seed weight and oil content. With high oil content as the breeding objective, a weighted ratio of 100-seed weight to oil content of 0.1:0.9 was set. The selection index was used to screen superior combinations, ultimately selecting the 20 combinations with the highest and lowest indices. Results showed that combinations with higher indices consistently outperformed those with lower indices in both traits. The average oil content of the top 20 combinations was 23.06%, with a 95th percentile of 26.09%. Compared to the maximum oil content of the parent population (24.14%), many individuals in the offspring of the top 20 combinations exceeded their parents in oil content. Conversely, the 20 combinations with the lowest indices had an average oil content of only 17.20%, with a 95th percentile of 18.45%. For 100-seed weight, despite a weight of only 0.1, the breeding objective still favored selecting higher values. The average 100-seed weight of the top 20 combinations was approximately 14.38 g, with a 95th percentile of 19.86 g, which was 8.39 g and 5.78 g higher than the 20 combinations with the lowest indices, respectively (Tables 1 and 2). Furthermore, with the weights of the two traits set at 100-seed weight:oil content = -0.5: -0.5, selection was made for both traits towards lower values. Combinations with larger indices had average values of 5.93 g and 17.12% for the two traits, with 95th percentiles of 14.08 g and 20.62%, respectively. They also had smaller genotype values, all achieving the breeding goals of low 100-seed weight and low oil content, further demonstrating the rationality of index selection in optimizing combination design.
[0066] Table 1. Predicted Genotype Values of the 20 Hybrid Combinations with the Highest Selection Index and Their Offspring
[0067]
[0068] Note: The weight of 100 grains and oil content is set to 0.1:0.9.
[0069] Table 2. Predicted genotype values of the 20 hybrid combinations with the smallest selection index and their offspring.
[0070]
[0071] Note: The weight of 100 grains and oil content is set to 0.1:0.9.
Claims
1. A method for constructing a multi-trait index model based on optimal combination design in crop hybrid breeding, characterized in that, Includes the following steps: Genome-wide SNP genotype data and breeding-related phenotypic data were obtained, and QTL-allele matrices for multiple traits were obtained using genome-wide association analysis. Based on the QTL-allele matrix of multiple traits in the germplasm resource population, simulated offspring of all possible hybridization combinations within the crop population are obtained; Calculate the genotype and phenotypic predicted values of the simulated offspring of the hybrid combinations; evaluate the breeding potential of different hybrid combinations by specifying the percentile of the phenotypic predicted values after the hybrid combination simulation, and screen out high-quality hybrid combination simulated offspring. Based on the high-quality combination of hybrid combinations to simulate offspring, a multi-trait selection index is constructed for offspring individuals to evaluate the breeding performance of multiple traits.
2. The method for constructing a multi-trait index model based on optimized combination design in crop hybrid breeding according to claim 1, characterized in that, The simulated offspring of all possible hybrid combinations within the crop population include: generating multiple homozygous offspring individuals through continuous self-pollination of F1 single seeds, which constitute the offspring population of the parental combination.
3. The method for constructing a multi-trait index model based on optimized combination design in crop hybrid breeding according to claim 1, characterized in that, The method for calculating the genotype value includes: (1) in, Let be the genotype estimate of the i-th individual; K is the number of QTLs. This is a genotype indicator variable for the i-th individual at the k-th locus. The value is 0 if the individual does not possess the allele, and 1 if the individual possesses the allele. This represents the effect of the l-th allele at the k-th locus.
4. The method for constructing a multi-trait index model based on optimized combination design in crop hybrid breeding according to claim 1, characterized in that, The method for calculating the phenotypic predicted value includes: (2) in, The phenotypic predicted value for the r-th offspring produced by the cross between parent i and parent j. The group mean Let be the genotype estimate of the r-th offspring individual. and These are the phenotypic observations of the two parents. and These are the estimated genotype values for the two parents.
5. The method for constructing a multi-trait index model based on optimized combination design in crop hybrid breeding according to claim 1, characterized in that, Methods for calculating the multi-trait selection index of offspring individuals include: (3) in, The selection index is used for comprehensive selection, where T represents the number of traits. Let t be the weight of the t-th trait, and let the weight satisfy t = t / t. ; t represents the standardized phenotypic predicted value of the t-th trait.
6. The method for constructing a multi-trait index model based on optimized combination design in crop hybrid breeding according to claim 5, characterized in that, Methods for calculating standardized phenotypic predictors of traits include: (4) in, For the t-th trait, the standardized phenotypic predicted value is... Let t be the phenotypic value of the t-th trait. Let t be the standardized location parameters of the t-th trait. Let be the scaling parameter for the t-th trait, which is standardized.
7. The method for constructing a multi-trait index model based on optimal combination design in crop hybrid breeding according to claim 1, characterized in that, The breeding performance of multiple traits is evaluated based on the quantiles of the multi-trait selection index of the offspring population. The larger the index value, the better the overall performance of multiple traits of the parent combination.
8. The method for constructing a multi-trait index model based on optimal combination design in crop hybrid breeding according to claim 1, characterized in that, The method of evaluating the breeding potential of different hybrid combinations by using the percentiles of the phenotypic predicted values after simulation of the specified hybrid combinations includes: using the percentiles of the phenotypic predicted values of the offspring population of the parental combinations as a comparison index, selecting the percentiles with larger percentages for high-value directional traits, and selecting the percentiles with smaller percentages for low-value directional traits.
9. A multi-trait index model based on optimal combination design in crop hybrid breeding, characterized in that, Obtained by the method described in any one of claims 1 to 8.
10. The application of the multi-trait index model based on crop hybridization breeding optimization combination design as described in claim 9 in crop breeding.