Whole genome selective breeding method based on mixed pool simplified genome sequencing

By simplifying the genome sequencing method through pooling and combining β-distribution fitting and regression models, the problems of high sequencing costs and insufficient accuracy in predicting low heritability traits in whole-genome selection breeding were solved, achieving improved accuracy in breeding value prediction while reducing costs.

CN120989258AActive Publication Date: 2025-11-21SANYA INST OF OCEANOGRAPHY OCEAN UNIV OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511509707.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2025-11-21
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing whole-genome selection breeding technologies are computationally inefficient when dealing with large amounts of data and have high sequencing costs, especially in terms of accuracy in predicting traits with low heritability.

Method used

A pooled simplified genome sequencing method was adopted, which combines multiple individuals into a group for sequencing. By using the minor allele ratio and genotype data of SNP loci, combined with β distribution fitting and regression models, the marker effect value was estimated to predict the breeding value.

Benefits of technology

While reducing sequencing costs by one-fifth or even one-tenth, it improves the accuracy of breeding value prediction, showing a significant advantage, especially in the breeding of Litopenaeus vannamei and dairy cows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120989258A_ABST
    Figure CN120989258A_ABST
Patent Text Reader

Abstract

The invention provides a whole genome selective breeding method based on mixed pool simplified genome sequencing, and relates to the field of breeding. Comprising the following steps: acquiring a sub-allele proportion and genotype data of an SNP site, respectively mixing according to the size of a target group, counting a probability density distribution diagram of allele frequency corresponding to each possibility result after genotype mixing, and acquiring parameters alpha and beta of beta distribution; the method comprises the following steps: respectively mixing genotype and phenotype data of a breeding target species according to the size of a target group, sampling through fitted beta distribution to obtain a sub-allele proportion for simulating real mixed pool sequencing, taking the sub-allele proportion as a genotype, and selecting a regression model for estimating a marker effect value and estimating a breeding value of each individual in a breeding target group. By adopting the method disclosed by the invention, the problem of singleness of single sample data is avoided, and meanwhile, the breeding value prediction accuracy can be improved and breeding of penaeus vannamei boone, dairy cows and the like can be better carried out under one fifth or even one tenth reference group establishment cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of breeding, in particular to a whole genome selection breeding method based on pool-based reduced genome sequencing. BACKGROUND

[0002] Whole genome selection is a method of combining phenotype information and genomic marker information to estimate the breeding value of individuals. Existing whole genome selection techniques include linear models and non-linear models such as neural networks. Although non-linear models such as neural networks have high accuracy, their computational efficiency is low when facing large amounts of data. Linear models are mainly divided into indirect methods and direct methods. The indirect method regards each molecular marker effect value as a random effect, and the direct method regards the individual breeding value as a random effect.

[0003] In the indirect method, different assumptions are made about the prior distribution of molecular marker effect values, which leads to a variety of models such as rrBLUP and a series of Bayesian models. However, most existing models have poor prediction accuracy for low heritability traits. In order to improve the prediction accuracy of low heritability traits, the compressed blup model, which combines individuals into groups, is proposed. Since this model combines individuals with close genetic relationships, that is, individuals with the same genetic background, it can reduce the impact of random errors on the accuracy of estimated breeding values, thereby improving prediction accuracy. As can be seen from the results, cBLUP has a clear advantage over other models in low heritability traits.

[0004] cBLUP combines individuals into groups when constructing the kinship matrix, but each individual still needs to be sequenced separately at the sequencing stage, which does not solve the problem of high sequencing costs faced by current whole genome selection. SUMMARY

[0005] In view of this, the present application aims to provide a whole genome selection breeding method based on pool-based reduced genome sequencing, which can obtain better breeding value prediction results than traditional genomic selection (GS) at one-fifth or even one-tenth of the reference population establishment cost.

[0006] The technical solution of the present application is as follows: A whole genome selection breeding method based on pool-based reduced genome sequencing, comprising the following steps: Step 1: Obtain the proportion of the second allele and the genotype data of the SNP site; Step 2, mix the allelic ratio and genotype data according to the target group size, respectively, and calculate the probability density distribution of the allelic frequency corresponding to each possible result after mixing the genotypes to obtain the parameters a and β of the β distribution; Step 3, mix the genotypes and phenotypes of the breeding target species according to the target group size, respectively, and then sample the simulated true pool sequencing allelic ratio by the β distribution fitted in step 2; Step 4, use the allelic ratio obtained in step 3 as the genotype, and select a regression model to estimate the marker effect value; Step 5, use the obtained marker effect value and the genotype value of the predicted population to estimate the breeding value of each individual in the breeding target population.

[0007] Further, in step 1, the data source is 2b-rad sequencing raw reads.

[0008] Further, in steps 2 and 3, the target group size is 5-10 individuals per group.

[0009] Further, the regression model includes RRBLUP, BayesA, BayesB or RKHS model.

[0010] Further, it also includes using individual multi-trait grouping algorithm to group the reference population.

[0011] The individual multi-trait grouping algorithm can make the number of individuals in each group differ less, while ensuring that the original single shape distribution is not changed after grouping. In order to achieve the above two points, two conditional constraints are added to the traditional K-mean algorithm, one is the constraint of the number of members in each group, and the second is the consistency constraint of the probability density before and after grouping. Finally, the grouped population can guarantee the original distribution in phenotype, and the number of members in each group is uniform.

[0012] The individual multi-trait grouping algorithm is as follows:

[0013] M is the maximum number of iterations; c t represents the current number of iterations; c t+1 represents the next round of iterations; is the hth center of the tth iteration; is the hth center of the next iteration; is the ith point to be clustered; N is the number of cluster points; K· is the number of cluster center points; h is the nth cluster, which is the index of the cluster, and the value range is 1 K ;; ; m is the number of points to be clustered; t is the number of iterations index ; i is the number of cluster points index ; subject to is the constraint declarator; min_size is the minimum sample number; max_size is the maximum sample number; for h=1 to K is the group traversal loop; return is the return of the final result.

[0014] Further, the multiple traits include 3-5 traits.

[0015] Compared with the prior art, the present application has the following beneficial effects: (1) The present application provides a whole genome selection breeding method based on pool-based simplified genome sequencing, which adopts a numerical representation method of population SNP site genotype, avoids the problem of single sample data singleness, and at the same time, this continuous numerical representation method is more suitable for existing regression models, can improve the breeding value prediction accuracy under one-fifth or even one-tenth of the reference population establishment cost, and better performs Penaeus vannamei, dairy cattle and other breeding.

[0016] (2) The individual multi-trait grouping algorithm of the present application can make the number of individuals in each group small, while ensuring that the original single shape distribution is not changed after grouping. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is the probability density histogram of the minor allele frequency and the fitting of the beta distribution; Figure 2Figure 2 is a probability density distribution diagram of the minor allele frequency; Figure 3 Figure 5 is a diagram of the results of the simulated pool sequencing of Example 2 and randomly sampling the same number of individuals to perform genome selection respectively; Figure 4 Figure 6 is a histogram of the frequency distribution of the cluster size after Kmeans clustering; Figure 5 Figure 7 is a histogram of the frequency distribution of the cluster size after clustering by the clustering algorithm of Example 3; Figure 6 Figure 8 is a probability density distribution diagram of the trait distribution before and after mixing; Figure 7 Figure 9 is the accuracy of the estimated breeding value of multiple traits; Figure 8 Figure 10 is a comparison diagram of the accuracy of the estimated breeding value of multiple models. DETAILED DESCRIPTION

[0018] In order to better understand the technical content of the present application, specific examples are provided below to further illustrate the present application.

[0019] The experimental methods used in the embodiments of the present application are conventional methods unless otherwise specified.

[0020] The materials, reagents, etc. used in the embodiments of the present application can be obtained from commercial channels unless otherwise specified.

[0021] For a diploid organism, each SNP site has two base types, called alleles, and the alleles of the present application are di-alleles; the minor allele refers to the base type with a lower frequency at a SNP site in a population.

[0022] Genotype (0, 1, 2) data: for a diploid organism, if at a certain site, the SNP is two major allele base types, the site genotype is 0, the SNP is two minor allele base types, the site genotype is 2, otherwise it is 1.

[0023] Example 1 The method for simulating real pool sequencing is as follows: 1. Obtain the minor allele read proportion and genotype (0, 1, 2) data of the SNP site from the original reads of 2b-rad sequencing.

[0024] 2. Mix the minor allele read proportion and genotype data according to the target group size, respectively, and count the probability density distribution diagram of the allele frequency corresponding to each possible result after genotype mixing to obtain the a, β parameters of the β distribution and fit the β distribution.

[0025] 3. Simulated pool: Obtain the genotype (not release the original reads data) and phenotype data, mix the genotype according to the target group size, and sample the simulated true pool sequencing allele ratio by the beta distribution fitted in step 2.

[0026] 4. Estimate marker effect value: select one of the four regression models of RRBLUP, BayesA, BayesB and RKHS for the estimation of marker effect value.

[0027] 5. Estimate breeding value GEBV: use the obtained marker effect value and the genotype value of the predicted population to estimate the GEBV of each individual in the breeding population or other population.

[0028] Example 2 1. Obtain the ratio of secondary alleles from the 2b-rad original sequencing data of ten parents (Miller-Crews et al. 2021), and genotype them.

[0029] 2. Mix five individuals into a group, randomly mix the samples (Miller-Crews et al. 2021) into five groups, mix the secondary allele ratio and genotype (0, 1, 2) data of each group, and the genotype mixed results are (0, 0.2, 0.4, 0.6, 0.8, 1.0, 1.2, 1.4, 1.6, 1.8, 2), and the mean and variance of the mixed results corresponding to each result are calculated, and the probability density histogram is drawn. Figure 1 It can be seen that the sequencing capture of biallelic sites in pool sequencing is consistent with bata distribution; The form of bata distribution is:

[0030] The expectation and variance are:

[0031] Both α and β are shape parameters, the larger α is, the more the distribution is biased towards 1, and the larger β is, the more the distribution is biased towards 0; mu: mean of X; σ: standard deviation of X.

[0032] And fit the beta distribution, and then obtain the parameters a, β required by the beta distribution, as shown in the following formula: Figure 2 The probability density distribution diagram of the five individuals after mixing.

[0033] Table 1 Parameters a, β of the probability density beta distribution diagram of the five individuals after mixing

[0034] Note: mixed_genotype: mixed genotype From Table 1, we can see that due to the sequencing of multiple individuals mixed together, a theoretical mixed genotype is generated. For the actual sequencing results, we fit the observed genotype (proportion of reads) at each locus to a beta distribution centered on the theoretical mixed genotype, which lays a statistical foundation for using the observed mixed genotype as the actual genotype for whole-genome selection in the future.

[0035] 3. Application to the prediction of breeding values for dairy cows (data used: genotypes of 5024 dairy cows at 50,000 loci, phenotype: milk yield, heritability: 0.65) (Zhang et al. 2015). Randomly divide 10% of the individuals into a test set to evaluate the accuracy of breeding value estimation. The remaining 90% of individuals are used as the training set, with 5 individuals in a group, a total of 900 groups. After mixing the genotypes, use the beta distribution fitted in step 2 to sample the simulated true pool sequencing allele proportions using the acceptance-rejection method, which are used as genotypes to select a regression model to estimate the prediction of breeding values. At the same time, randomly select 900 individuals in the training set and compare them with the pool. Figure 2 The results of simulated pool sequencing and randomly selected the same number of individuals for genome selection are shown (the model used is rrBLUP), which shows that Figure 3 It can be seen that mixed pool sequencing has a clear advantage at the same cost.

[0036] Example 3 The method of real pool sequencing and whole-genome selection is as follows: 1. Phenotype determination and reference population establishment for selected species.

[0037] 2. DNA extraction for reference population and test population: extract genomic DNA from samples and detect the quality and concentration of DNA samples. Cut the muscle of each individual, freeze in liquid nitrogen and store, take 50mg sample, extract DNA of each individual sample by phenol chloroform extraction method.

[0038] 3. Grouping and mixing of reference population by individual multi-trait grouping algorithm: determine the amount of DNA to be added according to the concentration of the DNA sample to ensure the same amount of DNA between mixed individuals.

[0039] The individual multi-trait grouping algorithm is as follows:

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046]

[0047] M is the maximum number of iterations; c t is the current iteration number; c t+1 is the next iteration; is the hth center of the tth generation; is the hth center of the next iteration; is the ith point to be clustered; N is the number of cluster points; K· is the number of cluster center points; h is the nth cluster, which is the index of the cluster, and its value ranges from 1 K to K; ; m is the number of points to be clustered; t is the iteration number; index ; i is the number of cluster points; index ; subject to is the constraint declarator; min_size is the minimum sample number; max_size is the maximum sample number; for h = 1 to K is the grouping traversal loop; return is the return of the final result.

[0048] The individual multi-trait grouping algorithm adds a constraint condition, so that each cluster size is relatively uniform, and adds a condition to make the trait distribution after mixing consistent with the original distribution, thereby Figure 4 and5 The results show that the individual multi-trait grouping algorithm can make the cluster size more uniform. Figure 6 The results show that the individual multi-trait grouping algorithm can maintain the distribution of trait probability density distribution after mixing.

[0049] According to embodiment 2, 900 individuals are randomly selected in the training set, and three traits are trained respectively, and the test set is verified. As a comparison group, the multi-trait grouping algorithm is used for grouping after mixing in the training set, 900 groups are obtained, and three traits are trained respectively, and the test set is verified. By Figure 7 The results show that after using the multi-trait grouping algorithm for individual grouping and mixing, compared with single individual library sequencing under the same cost, the accuracy of estimated breeding value has obvious advantages.

[0050] 4, 2b library sequencing: 2b-rad library kit is used to establish 2b-rad library of Penaeus vannamei mixed sample, and then sequencing analysis is carried out. SNP site information of the genome is obtained, and the proportion of the secondary allele of each SNP site is obtained as the genotype value.

[0051] 5, Estimate marker effect value: select one from RRBLUP, BayesA, BayesB and RKHS four regression models for estimation of marker effect value.

[0052] 6, Estimate breeding value GEBV: the obtained marker effect value and the genotype value of the prediction population are used to estimate the GEBV of each individual in the breeding population or other population.

[0053] It can be seen that under the same cost, compared with single measurement, pool sequencing has obvious advantages after being applied to four different models respectively. Figure 8

[0054] The above only describes the preferred embodiments of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.​

Claims

1. A whole genome selection breeding method based on metapooling reduced genome sequencing, characterized by, Comprising the following steps: Step 1, obtaining the allele frequency and genotype data of SNP sites; Step 2, mixing the allele frequency and genotype data according to the target group size, respectively, and calculating the probability density distribution of the allele frequency corresponding to each possible result after genotype mixing to obtain the parameters a and β of the beta distribution; Step 3, mixing the genotype and phenotype data of the breeding target species according to the target group size, respectively, and sampling the simulated true pool sequencing allele frequency by the beta distribution fitted in step 2; Step 4, using the allele frequency obtained in step 3 as the genotype, and selecting a regression model to estimate the marker effect value; Step 5, using the obtained marker effect value and the genotype value of the predicted population to estimate the breeding value of each individual in the breeding target population.

2. The whole genome selection breeding method based on pooled simplified genomic sequencing according to claim 1, characterized in that, In step 1, the data source is the original reads of 2b-rad sequencing.

3. The whole genome selection breeding method based on pooled simplified genome sequencing according to claim 1, characterized in that, In steps 2 and 3, the target group size is 5 individuals per group.

4. The whole genome selection breeding method based on pooled simplified genome sequencing according to claim 1, characterized in that, The regression model includes RRBLUP, BayesA, BayesB or RKHS model.

5. The whole genome selected breeding method based on pooled simplified genomic sequencing according to any one of claims 1-4, characterized in that, It also includes grouping the reference population using an individual multi-trait grouping algorithm, which is as follows: M is the maximum number of iterations; c t denotes the current iteration number; c t+1 denotes the next iteration; It is the h-th center in the t-th generation; represents the hth center of the next iteration; is the ith point to be clustered; N represents the number of cluster points; K· represents the number of cluster center points; h represents the nth cluster, is an index of the cluster, and has a value ranging from 1 K ;; ; m represents the number of points that need to be clustered; t represents the number of iterations index ; i representing the cluster points index ; subject to indicates a constraint declarator; min_size represents the minimum sample number; max_size represents the maximum number of samples; for h=1 to K indicates a grouping traversal loop; return indicates returning the final result.

6. The whole genome selection breeding method based on pooled simplified genome sequencing according to claim 5, characterized in that, The multi-trait includes 3-5 traits.

Citation Information

Patent Citations

  • Non-hybrid offspring identification method based on simplified genome sequencing and SNP minor allele frequency

    CN111826429A

  • Method for deducing abundance of each family in population based on mixed pool simplified genome sequencing

    CN120591388A

  • Breeding cross-generation phenotype prediction method and system based on ensemble learning, and electronic device

    WO2024212036A1