A parameter-free Bayesian genomic selection method for Huaxi cattle
Through the non-parametric Bayesian genomic selection method, the model complexity is adaptively determined, which solves the problem of insufficient prediction accuracy of traditional methods in complex breeding environments and achieves more efficient breeding results.
Patent Information
- Application Number
- CN202510579601.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Traditional genomic selection methods have difficulty capturing the complexity of gene-environment interactions, nonlinear genetic effects, and high-dimensional features in complex animal breeding environments, resulting in insufficient prediction accuracy.
A non-parametric Bayesian genomic selection method was used to construct a reference population, genotyping, association analysis and a non-parametric Bayesian model, adaptively determine the model complexity, calculate the effect value of the SNP locus, and use the Dirichlet process and Markov chain Monte Carlo method to obtain the genomic estimated breeding value.
The prediction accuracy of genomic selection has been improved, which can better identify genetic variants that have a significant impact on target traits and optimize breeding results.
Smart Images

Figure CN120108495B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of beef cattle breeding, and in particular to a parameter-free Bayesian genomic selection method for Huaxi cattle. Background Art
[0002] Genomic selection (GS) uses genome-wide markers to predict the genetic merit of individuals for breeding decisions. Traditional genomic selection methods are often based on linear mixed models, which typically assume that effects follow a sparse or normal distribution. However, in the complex environment of animal breeding, these assumptions may not be sufficient to capture the complexity of gene-environment interactions, nonlinear genetic effects, and high-dimensional traits. Summary of the Invention
[0003] The object of the present invention is to provide a parameter-free Bayesian genomic selection method for Huaxi cattle to solve the problems raised in the above background technology.
[0004] The present invention provides a reference-free Bayesian genomic selection method for Huaxi cattle, the method comprising:
[0005] Step 1: construct a reference population and measure the important taxonomic traits of each Huaxi cattle;
[0006] Step 2: Blood was collected from each Huaxi cattle in the reference group, DNA was extracted, and genotyping was performed using a biochip to obtain the genotype of each SNP site, and quality control was performed on the genotyping data;
[0007] Step 3: For each Huaxi cattle individual, the genotype file obtained in step 2 and the phenotype file recorded in step 1 are subjected to association analysis to obtain summary statistics of each phenotype, and each SNP site is divided into blocks according to the position of the chromosome to obtain an LD reference file; wherein the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level and initial site effect value, and the LD reference file includes the segment divided by each chromosome, the SNP site corresponding to each chromosome and the allele frequency;
[0008] Step 4: For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual are used as the input files of the pre-built non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector to obtain the genomic estimated breeding value of the candidate population for the phenotype.
[0009] As a preferred embodiment, the important classification traits include calving difficulty, pH and meat color.
[0010] As a preferred embodiment, the quality control of the genotyping data includes: eliminating alleles with a frequency less than 0.05 and a p less than 10 in the Hardy-Weinberg equilibrium test. -6 and SNP sites with a missing genotype ratio of less than 0.05.
[0011] As a preferred embodiment, the quality control of the genotyping data is performed using plink1.9 software.
[0012] As a preferred embodiment, the biochip in step 2 is a Cattle110K high-density chip.
[0013] As a preferred embodiment, the genotype file obtained in step 2 and the phenotype file recorded in step 1 are subjected to association analysis to obtain summary statistics of each phenotype using plink2 software.
[0014] As a preferred embodiment, the constructed non-parametric Bayesian model is expressed as follows:
[0015] ;
[0016] in, represents the phenotype vector, X express N × p The genotype matrix; N Indicates the number of individuals ,p Indicates the number of sites ; express p ×1 vector of site effect values, where Following the mean of 0 and variance σ 2 The normal distribution of σ 2 , σ 2 ~DP( H , ), H represents the basic distribution, Display Control σ 2 The concentration parameter at which the upper distribution shrinks toward H; represents the residual;
[0017] The calculation process of the effect value of the SNP site is as follows:
[0018] ;
[0019] in, represents the initial site effect value; represents the hyperparameters of the non-parametric Bayesian model; represents random weights and is independent of , Satisfy 0≤ ≤1 and , Indicates from Random variables generated independently from the distribution, 1 and is the distribution parameter, by controlling The distribution parameters 1 and Ability to flexibly adjust the generated distribution characteristics, Determines the weight generated after each break ; represents a random distribution; represents the initial prior distribution of the hypothesis; represents the LD matrix; represents a normal distribution; Represents the identity matrix.
[0020] As a preferred embodiment, in step 4, the site effect value vector is obtained using the Markov Chain Monte Carlo method based on the calculated effect value of the SNP site.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] The present invention analyzes the genotype data of Huaxi cattle through a non-parametric Bayesian model, and can adaptively determine the complexity of the model without having to set the number of clusters or other parameters in advance; and the non-parametric Bayesian model can more accurately evaluate the relationship between genotype and traits, thereby improving prediction accuracy to optimize the accuracy of genomic selection. Compared with traditional methods, the method of the present invention can better utilize the information in the data, identify genetic variations that have an important impact on the target traits, and improve the effect of selection. The purpose of the present invention is to establish an efficient and high-quality Huaxi cattle breeding method based on non-parametric Bayesian genomic selection technology. When this method is applied to traits with a high proportion of non-additive effects, the genomic prediction accuracy of this method is significantly improved compared with other methods, providing a new technical path for improving the accuracy of genomic selection prediction in Huaxi cattle. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a flow chart of the reference-free Bayesian genomic selection method for Chinese cattle in Example 1.
[0024] Figure 2 Flowchart of the processing of summary statistics and LD reference files by the non-parametric Bayesian model in Example 1.
[0025] Figure 3This is a schematic diagram comparing the genomic prediction accuracy of three traits of Chinese Western Cattle under three models in Example 2. DETAILED DESCRIPTION
[0026] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0027] Example 1
[0028] Combine Figure 1 This embodiment provides a parameter-free Bayesian genomic selection method for Huaxi cattle, which includes:
[0029] Step 1: construct a reference population and measure the important classification traits of the reference population; in a specific embodiment, the important classification traits include calving ease (CE), meat color (MC) and pH.
[0030] Step 2: Blood is collected from each cow in the reference population, DNA is extracted, and genotyping is performed using a biochip to obtain the genotype of each SNP (single nucleotide polymorphism) site; quality control is performed on the genotyping data. In some specific embodiments, the quality control of the genotyping data includes: eliminating unqualified individuals and eliminating individuals with an allele frequency less than 0.05 and a Hardy-Weinberg equilibrium test p less than 10 -6 and SNP sites with a ratio of missing genotypes less than 0.05, the biochip in step 2 is a Cattle110K high-density chip. In a specific embodiment, Figure 1 The number of SNP sites before quality control was 770K and after quality control was 550K.
[0031] Step 3: Use plink2 software to perform association analysis on the genotype file obtained in step 2 and the phenotype file recorded in step 1 to obtain summary statistics for each phenotype; the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level, and initial site effect value; and each SNP site is divided into blocks according to the location of the chromosome to obtain an LD reference file; the LD reference file includes the segments divided by each chromosome, the SNP site corresponding to each chromosome, and the allele frequency of the SNP site;
[0032] Step 4: For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual in the reference population are used as the input files of the pre-built non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector to obtain the genomic estimated breeding value (GEBV) of the candidate population for the phenotype. The genomic estimated breeding value (GEBV) is calculated as follows: ;in, represents the genomic breeding value of individual i, represents the genotype of individual i at site j, Represents the effect value of site j, and the site effect value vector is obtained based on the calculated effect value of the SNP site using mcmc sampling (Markov Chain Monte Carlo method).
[0033] Combine Figure 2 In this embodiment, the parameter distribution of the non-parametric Bayesian model is determined by the Dirichlet Process based on the existing data (summary statistics and LD reference files). The prior distribution is obtained from the data by the model without manually entering any parameters. Finally, the Markov chain Monte Carlo method is used for posterior sampling to obtain the genomic estimated breeding value.
[0034] Among them, the formula of the non-parametric Bayesian model is:
[0035] ;
[0036] in, represents the phenotype vector, X express N × p The genotype matrix; N Indicates the number of individuals ,p Indicates the number of sites ; express p ×1 vector of site effect values, where Following the mean of 0 and variance σ 2 The normal distribution of σ 2 , σ 2 ~DP( H , ), H represents the basic distribution, Display Control σ 2The concentration parameter at which the upper distribution shrinks toward H; Denotes the residual. Specifically, there are several equivalent probability representations for the Dirichlet process, among which the truncated process is widely used because it is convenient for model fitting. The truncated process representation in the Dirichlet process regards the Dirichlet process as an infinite Gaussian mixture model.
[0037] Furthermore, the calculation process of the effect value of the SNP site in step 4 is expressed as follows:
[0038] ;
[0039] in, represents the initial site effect value; Represents the hyperparameters of the non-parametric Bayesian model; the summary statistics include a column representing the initial site effect value obtained by processing each SNP site through the previous GWAS model (Genome-Wide Association Studies whole genome association analysis model) ; represents a random variable and is independent of , Satisfy 0≤ ≤1 and , Indicates from Random variables generated independently from the distribution, 1 and is the distribution parameter, by controlling The distribution parameters 1 and Ability to flexibly adjust the generated distribution characteristics, Determines the weight generated after each break in the Dirichlet stick-breaking process , It reflects that the non-parametric Bayesian model automatically identifies the probability mass of each potential cluster (or class) in the generated random distribution G according to the characteristics of the data; represents the LD matrix constructed based on the LD reference file; represents the initial prior distribution of the hypothesis; represents the hyperparameters of the non-parametric Bayesian model; represents a normal distribution; Represents the identity matrix.
[0040] It is understandable that the present embodiment analyzes the Huaxi cattle genome data by a parameter-free Bayesian model. Compared with traditional genome selection methods, first, the method in the present embodiment can automatically determine the model complexity. The parameter-free Bayesian model can adaptively determine the complexity of the model without setting the number of clusters or other parameters in advance. Through the Dirichlet process, the parameter-free Bayesian model can automatically identify potential groups or clusters according to the characteristics of the data, thereby avoiding the overfitting or underfitting problems that may exist in traditional methods. Second, the method in the present embodiment can flexibly process high-dimensional data. The parameter-free Bayesian method has strong flexibility and robustness when processing high-dimensional data, so that information can be effectively extracted from complex genome data, and the structure and association in the data can be automatically captured. It adapts to the potential complex patterns in the data, which is difficult to achieve in traditional linear models or fixed parameter models. Third, the method in the present embodiment can optimize the accuracy of genome selection. The parameter-free Bayesian model can more accurately evaluate the relationship between genotype and trait, thereby improving prediction accuracy. Compared with traditional methods, it can better utilize the information in the data, identify genetic variants that have important effects on target traits, and improve the effectiveness of selection.
[0041] Example 2
[0042] This example provides a specific embodiment of the method for selecting the genome of Chinese cattle in Example 1, and the specific steps are as follows:
[0043] Step 1: Construct a reference population and measure key taxonomic traits. The Cattle Genetics and Breeding Innovation Team at the Beijing Institute of Animal Husbandry and Veterinary Medicine began establishing a foundation herd of 2,320 Huaxi cattle cows in the Ulagai Management District of Xilin Gol League, Inner Mongolia, in 2008. With annual expansion, the herd numbered over 4,000 cows by 2019. The offspring of this foundation herd were used to construct a reference population for Huaxi cattle. Traits measured included calving ease (CE), pH, and meat color (MC).
[0044] In the second step, blood was collected from each cow in the reference population, DNA was extracted, and genotyping was performed using the Cattle110K (110k) high-density chip to obtain the genotype of each SNP locus. The data after genotyping were quality controlled using Plink 1.9 software. The quality control included the following: excluding unqualified individuals and excluding those with an allele frequency less than 0.05 and a Hardy-Weinberg equilibrium test with a p less than 10 -6The Cattle110K high-density microarray was used in step 2. The number of SNPs after quality control was 100,000. Data on CE traits were collected from 1,165 Huaxi cattle, pH traits from 1,173 Huaxi cattle, and MC traits from 1,226 Huaxi cattle. Specific data are shown in Table 1.
[0045]
[0046] Step 3: For each Huaxi cattle individual, use plink2 software to perform association analysis on the genotype file obtained in step 2 and the phenotype file recorded in step 1 to obtain summary statistics of each phenotype; and block each SNP site according to the position of the chromosome to obtain an LD reference file, wherein the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level and initial site effect value, and the LD reference file includes the segment divided by each chromosome, the SNP site corresponding to each chromosome and the allele frequency;
[0047] Step 4: Construct a non-parametric Bayesian model. For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual are used as the input files of the non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector obtained using the Markov chain Monte Carlo method to obtain the genomic estimated breeding value (GEBV) of the candidate population for the phenotype. ,in is the genomic breeding value of individual i, is the genotype of individual i at site j, is the effect value of site j), and the Pearson correlation coefficient between the genomic estimated breeding value (GEBV) and the true phenotype of the test set was calculated.
[0048] Step 5: Combine Figure 3 In this example, genomic and phenotype files are used as input for the GBLUP (Genomic Best Linear Unbiased Prediction) model and the SDPR (Summary-Data-Based Penalized Regression Model) model to predict genomic-estimated breeding values. Summary statistics and LD reference files are used as input for the SBayesRC (Non-parametric Bayesian) model and the SDPR model to predict genomic-estimated breeding values. Pearson correlation coefficients are calculated between the predicted genomic-estimated breeding values (GEBVs) using the three methods and the true phenotypes of the pre-obtained test set. Figure 3This is a comparison of the Pearson correlation coefficients of the above three methods under five-fold cross validation. It can be seen from the box plot that for the genomic breeding values of each phenotype, the non-parametric Bayesian model has the best prediction performance.
[0049] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A reference-free Bayesian genomic selection method for Huaxi cattle, characterized in that: include: Step 1: construct a reference population and measure the important taxonomic traits of each Huaxi cattle; Step 2: Blood was collected from each Huaxi cattle in the reference group, DNA was extracted, and genotyping was performed using a biochip to obtain the genotype of each SNP site, and quality control was performed on the genotyping data; Step 3: For each Huaxi cattle individual, the genotype file obtained in step 2 and the phenotype file recorded in step 1 are subjected to association analysis to obtain summary statistics of each phenotype, and each SNP site is divided into blocks according to the position of the chromosome to obtain an LD reference file; wherein, the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level and initial site effect value; the LD reference file includes the segment divided by each chromosome, the SNP site and allele frequency corresponding to each chromosome; Step 4: For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual are used as input files of the pre-built non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector to obtain the genomic estimated breeding value of the candidate individual for the phenotype; The constructed non-parametric Bayesian model is expressed as follows: ; in, represents the phenotype vector, X express N × p The genotype matrix; N Indicates the number of individuals ,p Indicates the number of sites ; express p ×1 vector of site effect values, where Following the mean of 0 and variance σ 2 The normal distribution of σ 2 , σ 2 ~DP( H , ), H represents the basic distribution, Display Control σ 2 The concentration parameter at which the upper distribution shrinks toward H; represents the residual; The calculation process of the effect value of the SNP site is as follows: ; in, represents the initial site effect value; represents the hyperparameters of the non-parametric Bayesian model; represents random weights and is independent of , Satisfy 0≤ ≤1 and , Indicates from Random variables generated independently from the distribution, 1 and are all distribution parameters; represents a random distribution; represents the initial prior distribution of the hypothesis represents the LD matrix; represents a normal distribution; Represents the identity matrix.
2. The method for Bayesian genome selection of Huaxi cattle according to claim 1, wherein The important classification traits include calving ease, pH and meat color.
3. The method for Bayesian genome selection of Huaxi cattle without reference according to claim 1, wherein The quality control of the genotyping data includes: eliminating alleles with a frequency less than 0.05 and p less than 10 in the Hardy-Weinberg equilibrium test. -6 and SNP sites with a missing genotype ratio of less than 0.
05.
4. The method for Bayesian genome selection of Huaxi cattle without reference according to claim 3, wherein The quality control of the genotyping data was performed using plink 1.9 software.
5. The method for Bayesian genome selection of Huaxi cattle without reference according to claim 1, wherein The biochip in step 2 is a Cattle110K high-density chip.
6. The method for Bayesian genome selection of Huaxi cattle without reference according to claim 1, wherein The genotype file obtained in step 2 and the phenotype file recorded in step 1 were subjected to association analysis to obtain summary statistics of each phenotype using plink2 software.
7. The method for Bayesian genomic selection of Huaxi cattle without reference according to claim 1, characterized in that In step 4, the site effect value vector is obtained based on the calculated effect value of the SNP site using the Markov chain Monte Carlo method.
Citation Information
Patent Citations
New fast Bayes method estimating genomic estimated breeding value
CN107590364A
Western China cattle genome selection method
CN111243667A