A genome-wide selection breeding method based on pooled simplified genome sequencing
By simplifying genome sequencing and using multi-trait grouping algorithms through pool mixing, the problems of high sequencing costs and poor prediction accuracy of low heritability traits in whole-genome selection breeding have been solved. This has achieved the goal of improving the accuracy of breeding value prediction while reducing costs, with particularly significant results in the breeding of Litopenaeus vannamei and dairy cows.
Patent Information
- Application Number
- CN202511509707.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Existing whole-genome selection breeding technologies suffer from high sequencing costs and poor accuracy in predicting traits with low heritability, especially with large datasets where computational efficiency is low, making it difficult to effectively reduce sequencing costs.
A pooled simplified genome sequencing method was adopted, which grouped individuals by multiple traits and mixed sequencing, used β distribution to fit genotype data, and combined RRBLUP, BayesA, BayesB or RKHS models to estimate marker effect values, thereby reducing the cost of establishing a reference population. At the same time, an individual multi-trait grouping algorithm was used to ensure the consistency of the distribution after grouping.
At a cost of one-fifth or even one-tenth of the reference population, it significantly improves the accuracy of breeding value prediction, especially showing a clear advantage in Litopenaeus vannamei and dairy cow breeding, reducing sequencing costs while improving prediction accuracy.
Smart Images

Figure CN120989258B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of breeding, and in particular to a whole-genome selection breeding method based on mixed-pool simplified genome sequencing. Background Technology
[0002] Genome-wide selection breeding combines phenotypic and genomic marker information to estimate the breeding value of an individual. Existing genome-wide selection techniques include linear models and nonlinear models such as neural networks. While nonlinear models like neural networks offer high accuracy, their computational efficiency is low when dealing with large datasets. Linear models are mainly divided into indirect and direct methods. Indirect methods treat each molecular marker effect value as a random effect, while direct methods treat the individual breeding value as a random effect.
[0003] Indirect methods, by making different assumptions about the prior distribution of molecular marker effect values, have led to various models, including the rrBLUP and a series of Bayesian models. However, most existing models have poor predictive accuracy for traits with low heritability. To improve the predictive accuracy of traits with low heritability, the compressed blup model, which groups individuals together, has been proposed. Because this model groups closely related individuals—that is, individuals with similar genetic backgrounds—it reduces the impact of random errors on the accuracy of estimated breeding values, thereby improving prediction accuracy. The results show that, compared to other models, cBLUP exhibits a significant advantage in traits with low heritability.
[0004] cBLUP merges individuals into groups when constructing the kinship matrix, but each individual still needs to be sequenced separately during the sequencing stage. This does not solve the problem of current whole-genome selection and still faces the issue of high sequencing costs. Summary of the Invention
[0005] Therefore, the purpose of this invention is to propose a whole-genome selection breeding method based on pooled simplified genome sequencing, which can obtain breeding value prediction results that are superior to those of traditional whole-genome selection (GS) with a reference population establishment cost of one-fifth or even one-tenth.
[0006] The technical solution of this invention is implemented as follows:
[0007] A genome-wide selection breeding method based on pooled simplified genome sequencing includes the following steps:
[0008] Step 1: Obtain the minor allele ratio and genotype data for SNP loci;
[0009] Step 2: Mix the suballele ratio and genotype data according to the target group size, and statistically analyze the probability density distribution of allele frequencies corresponding to each possible outcome after genotype mixing to obtain the parameters α and β of the β distribution.
[0010] Step 3: Mix the genotype and phenotypic data of the target species according to the target group size, and then sample the proportion of minor alleles to simulate real pool sequencing using the β distribution fitted in Step 2.
[0011] Step 4: Use the proportion of minor alleles obtained in Step 3 as the genotype, and select a regression model to estimate the marker effect value.
[0012] Step 5: Using the obtained marker effect value and the genotype value of the predicted population, estimate the breeding value of each individual in the target population.
[0013] Furthermore, in step 1, the data comes from raw reads of 2b-rad sequencing.
[0014] Furthermore, in steps 2 and 3, the target group size is set at 5 to 10 individuals per group.
[0015] Furthermore, regression models include RRBLUP, BayesA, BayesB, or RKHS models.
[0016] Furthermore, it also includes using an individual multi-trait grouping algorithm to group the reference population.
[0017] Individual multi-trait grouping algorithms can minimize the difference in the number of individuals in each group while ensuring that the original distribution of individual traits is not altered after grouping. To achieve these two points, two constraints are added to the traditional K-means algorithm: one is the constraint on the number of members in each group, and the second is the constraint on the consistency of probability density before and after grouping. Ultimately, this ensures that the grouped population maintains the original phenotypic distribution, and that the number of members within each group is uniform.
[0018] The algorithm for grouping individuals with multiple traits is as follows:
[0019]
[0020] M It is the maximum number of iterations;
[0021] c tIndicates the current iteration number;
[0022] c t+1 Indicates the next iteration;
[0023] It is the h-th center in the t-th generation;
[0024] This indicates the h-th center in the next iteration;
[0025] It is the i-th point to be clustered;
[0026] N Indicates the number of cluster points;
[0027] K· Indicates the number of cluster centers;
[0028] h This represents the nth cluster, and is the cluster index, with values ranging from 1 to... K ; ; ;
[0029] m Indicates the number of points that need to be clustered;
[0030] t Indicates the number of iterations index ;
[0031] i Representing cluster points index ;
[0032] The subject to keyword indicates a constraint declaration;
[0033] min_size Indicates the minimum number of samples;
[0034] max_size Indicates the maximum number of samples;
[0035] The loop `for h=1 to K` represents a grouped traversal loop.
[0036] `return` indicates that the final result is returned.
[0037] Furthermore, multiple traits include 3 to 5 traits.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0039] (1) This invention provides a whole-genome selection breeding method based on mixed pool simplified genome sequencing. It adopts the numerical representation of population SNP locus genotypes, which avoids the problem of single sample data. At the same time, this continuous numerical representation is more suitable for existing regression models. It can improve the accuracy of breeding value prediction with one-fifth or even one-tenth of the reference population establishment cost, and better carry out breeding of Litopenaeus vannamei, dairy cows and other species.
[0040] (2) The individual multi-trait grouping algorithm of the present invention can make the difference in the number of individuals in each group small, while ensuring that the distribution of the original single shape is not changed after grouping. Attached Figure Description
[0041] Figure 1 Fitting the probability density histogram and β distribution of the minor allele frequencies;
[0042] Figure 2 This is a probability density distribution map of the frequency of the second allele.
[0043] Figure 3 This is a diagram showing the results of simulated pooled sequencing and random selection of the same number of individuals for genome selection in Example 2.
[0044] Figure 4 This is a histogram showing the frequency distribution of cluster sizes after K-means clustering.
[0045] Figure 5 This is a frequency distribution histogram of cluster sizes after clustering using the clustering algorithm in Example 3;
[0046] Figure 6 This is a probability density distribution diagram of the trait distribution before and after mixing.
[0047] Figure 7 To improve the accuracy of estimating breeding values for multiple traits;
[0048] Figure 8 A comparison chart showing the accuracy of various models in estimating breeding values. Detailed Implementation
[0049] To better understand the technical content of this invention, specific embodiments are provided below to further illustrate the invention.
[0050] Unless otherwise specified, the experimental methods used in the embodiments of this invention are all conventional methods.
[0051] Unless otherwise specified, all materials and reagents used in the embodiments of this invention are commercially available.
[0052] For diploid organisms, each SNP site has two base types, called alleles. The alleles in this invention are diallelic alleles. Suballelic alleles refer to the base types that occur less frequently at a certain SNP site in a population.
[0053] Genotype (0, 1, 2) data: For a diploid organism, if at a certain locus, the SNP is two major allele base types, the genotype at that locus is 0; if the SNP is two minor allele base types, the genotype at that locus is 2; otherwise, it is 1.
[0054] Example 1
[0055] The method for simulating real pooled sequencing involves the following steps:
[0056] 1. Obtain the proportion of minor allele reads and genotype (0, 1, 2) data for SNP sites from the raw reads of 2b-rad sequencing.
[0057] 2. Mix the proportion of minor allele reads and genotype data according to the target group size, and statistically analyze the probability density distribution of allele frequencies corresponding to each possible outcome after genotype mixing. Obtain the α and β parameters of the β distribution and fit the β distribution.
[0058] 3. Simulated pooling: Obtain genotype (unpublished raw reads data) and phenotype data, mix the genotypes according to the target group size, and then sample the proportion of minor alleles to simulate real pooling sequencing using the β distribution fitted in step 2.
[0059] 4. Estimate the labeled effect size: Select one of the four regression models, RRBLUP, BayesA, BayesB, and RKHS, for estimating the labeled effect size.
[0060] 5. Estimate the breeding value GEBV: Use the obtained marker effect value and the genotype value of the predicted population to estimate the GEBV of each individual in the selected population or other populations.
[0061] Example 2
[0062] 1. The proportion of minor alleles was obtained from the 2b-rad raw sequencing data of the ten parents (Miller-Crews et al. 2021), and genotyping was performed on them.
[0063] 2. A pooling method was used, with five individuals per group. The sample (Miller-Crews et al. 2021) was randomly pooled into groups of five. The minor allele proportions and genotypes (0, 1, 2) were pooled separately, resulting in pools of five individuals each. The resulting genotype pools were (0, 0.2, 0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6, 1.8, 2). The mean and variance of the pooled minor allele proportions for each outcome were calculated, and a probability density histogram was plotted. Figure 1 The sequencing capture of bistate sites in mixed-pool sequencing shows that it conforms to a bata distribution.
[0064] The form of the Bata distribution is:
[0065]
[0066] The expected value and variance are:
[0067]
[0068] α and β are both shape parameters. The larger α is, the more the distribution is biased towards 1, and the larger β is, the more the distribution is biased towards 0.
[0069] mu: the mean of X;
[0070] σ: Standard deviation of X.
[0071] Then, the β distribution is used for fitting, and the required parameters α and β of the β distribution are obtained from... Figure 2 The probability density distribution of the five individuals shown is a graph.
[0072] Table 1. Parameters α and β values of the probability density β distribution plot after mixing five individuals.
[0073]
[0074] Note: mixed_genotype: mixed genotype
[0075] As shown in Table 1, sequencing multiple individuals together will generate a theoretical mixed genotype. For the actual sequencing results, we statistically analyzed that the observed genotype (proportion of reads) at each locus conforms to a beta distribution centered on the theoretical mixed genotype. This lays a statistical foundation for using the observed mixed genotype as the actual genotype for full genotype breeding.
[0076] 3. Application to breeding value prediction in dairy cows (data used: genotypes and phenotypes of 5024 dairy cows at 50,000 loci; milk yield; heritability: 0.65) (Zhang et al. 2015). 10% of the individuals were randomly assigned to the test set to evaluate the accuracy of the breeding value estimation. The remaining 90% of individuals were used as the training set, with 5 individuals per group, for a total of 900 groups. After mixing the genotypes, the proportion of minor alleles simulating real pool sequencing was obtained using the beta distribution fitted in step 2 and the acceptance / rejection method. This proportion was used as the genotype, and a regression model was selected to predict the breeding value. Simultaneously, 900 individuals were randomly selected from the training set for comparison with the pooled individuals. Figure 2 The results shown are those of simulated pooled sequencing and genomic selection performed on randomly selected individuals of the same number (using the rrBLUP model). Figure 3 This indicates that, at the same cost, pooled sequencing has a significant advantage.
[0077] Example 3
[0078] The method for real pool sequencing and genome-wide selection is as follows:
[0079] 1. Perform phenotypic analysis on the selected species and establish a reference population.
[0080] 2. DNA extraction from the reference and test populations: Genomic DNA was extracted from the samples, and the quality and concentration of the DNA samples were tested. Muscle tissue was excised from each individual, frozen in liquid nitrogen, and stored. 50 mg of the sample was taken, and DNA was extracted from each individual sample using the phenol-chloroform extraction method.
[0081] 3. Group and mix the reference population using an individual multi-trait grouping algorithm: Determine the amount of DNA to be added based on the concentration of qualified DNA samples to ensure that the amount of DNA is the same among the individuals in the mixed sample.
[0082] The algorithm for grouping individuals with multiple traits is as follows:
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091] M It is the maximum number of iterations;
[0092] c t Indicates the current iteration number;
[0093] c t+1 Indicates the next iteration;
[0094] It is the h-th center in the t-th generation;
[0095] This indicates the h-th center in the next iteration;
[0096] It is the i-th point to be clustered;
[0097] N Indicates the number of cluster points;
[0098] K· Indicates the number of cluster centers;
[0099] h This represents the nth cluster, and is the cluster index, with values ranging from 1 to... K ; ; ;
[0100] m Indicates the number of points that need to be clustered;
[0101] t Indicates the number of iterations index ;
[0102] i Representing cluster points index ;
[0103] The subject to keyword indicates a constraint declaration;
[0104] min_size Indicates the minimum number of samples;
[0105] max_size Indicates the maximum number of samples;
[0106] The loop `for h=1 to K` represents a grouped traversal loop.
[0107] `return` indicates that the final result is returned.
[0108] This multi-trait grouping algorithm for individuals adds constraints to make each cluster size more uniform, and adds a condition to ensure that the trait distribution after pooling is consistent with the original distribution. Figure 4 and 5 The results show that this multi-trait grouping algorithm can achieve relatively uniform cluster sizes. Figure 6 The results show that the proposed multi-trait grouping algorithm can maintain the probability density distribution of the distributed traits after mixing.
[0109] According to Example 2, 900 individuals are randomly selected from the training set, and each individual is trained on one of the three traits. The results are then validated using a test set as a control group. The training set is then grouped using a multi-trait grouping algorithm and mixed to obtain 900 groups. Each of these groups is trained on one of the three traits and validated using a test set. Figure 7 The results showed that using a multi-trait grouping algorithm for individual grouping and pooling had a significant advantage in estimating breeding values compared to single-individual library construction and sequencing at the same cost.
[0110] 4. Perform 2b library construction and sequencing: Construct a 2b-rad library of mixed samples of Litopenaeus vannamei using a 2b-rad library construction kit, followed by sequencing analysis. Obtain SNP locus information of the genome, and obtain the minor allele ratio of each SNP locus as the genotype value.
[0111] 5. Estimate the labeled effect size: Select one of the four regression models, RRBLUP, BayesA, BayesB, and RKHS, for estimating the labeled effect size.
[0112] 6. Estimate the breeding value GEBV: Use the obtained marker effect value and the genotype value of the predicted population to estimate the GEBV of each individual in the selected population or other populations.
[0113] Depend on Figure 8 It can be seen that, under the same cost, when mixed-pool sequencing and single-sample sequencing are applied to four different models, mixed-pool sequencing has a significant advantage over single-sample sequencing.
[0114] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A whole genome selection breeding method based on metapooling reduced genome sequencing, characterized by, Comprising the following steps: Step 1, obtaining the allele frequency and genotype data of SNP sites; the alleles are di-allelic; for the genotype data, if at a certain site, the SNP is two major allele base types, the site genotype is 0, the SNP is two minor allele base types, the site genotype is 2, otherwise it is 1; Step 2, mixing the minor allele frequency and genotype data according to the target group size, respectively, to obtain the probability density distribution of the allele frequency corresponding to each possible result after mixing the genotypes, and obtain the parameters a and β of the beta distribution; Step 3, mixing the genotype and phenotype data of the breeding target species according to the target group size, respectively, and then sampling the simulated real pool sequencing minor allele frequency through the beta distribution fitted in step 2; Step 4, using the regression model to estimate the marker effect value by taking the minor allele frequency obtained in step 3 as the genotype; Step 5, using the obtained marker effect value and the genotype value of the predicted population to estimate the breeding value of each individual in the breeding target population; Further comprising grouping the reference population by using an individual multi-trait grouping algorithm, and the individual multi-trait grouping algorithm is as follows: M is the maximum number of iterations; c t denotes the current iteration number; c t+1 denotes the next iteration; is the hth center of the tth generation; represents the hth center of the next iteration; is the ith point to be clustered; N represents the number of cluster points; K represents the number of cluster center points; h represents the nth cluster, is an index of the cluster, and the value thereof ranges from 1 K ; ; m represents the number of points that need to be clustered; t represents the number of iterations index ; i representing the cluster points index ; subject to indicates a constraint condition declarator; min_size represents the minimum sample number; max_size denotes the maximum number of samples; for h=1 to K indicates a grouping traversal loop; return indicates returning the final result.
2. The whole genome selection breeding method based on pooled simplified genomic sequencing according to claim 1, characterized in that, In step 1, the data source is the original reads of 2b-rad sequencing.
3. The whole genome selection breeding method based on pooled simplified genome sequencing according to claim 1, characterized in that, In steps 2 and 3, the target group size is 5 individuals per group.
4. The whole genome selection breeding method based on pooled simplified genome sequencing according to claim 1, characterized in that, The regression model includes RRBLUP, BayesA, BayesB or RKHS model.
5. The whole genome selection breeding method based on pooled simplified genome sequencing according to claim 1, characterized in that, The multi-trait includes 3-5 traits.
Citation Information
Patent Citations
Non-hybrid offspring identification method based on simplified genome sequencing and SNP minor allele frequency
CN111826429A
Method for deducing abundance of each family in population based on mixed pool simplified genome sequencing
CN120591388A