West China cattle parameter-free Bayesian genome selection method
By analyzing the genotype data using the Bayesian model in West China cattle breeding, the problem of difficulty in capturing complex gene-environment interactions was solved, and higher genome selection accuracy and selection effect were achieved.
Patent Information
- Application Number
- CN202510579601.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Traditional genome selection methods are difficult to capture the complexity of gene-environmental interactions, nonlinear genetic effects, and high-dimensional features in complex animal breeding environments.
The genotype data of West China cattle were analyzed using the Bayesian model without the mark. The potential population or clusters were automatically identified through the Dilikre process, the complexity of the model was adaptively determined, and the effect value of the SNP site was calculated by the Markov chain Monte Carlo method.
It improves the accuracy of genomic selection, can more accurately evaluate the relationship between genotype and trait, identify genetic variants that have important effects on the target trait, and improves the selection effect.
Smart Images

Figure CN120108495A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of beef cattle breeding, and in particular to a parameter-free Bayesian genome selection method for Huaxi cattle. Background Art
[0002] Genomic selection (GS) is the use of whole genome markers to predict the genetic value of individuals for breeding decisions. Traditional genomic selection methods are mostly based on linear mixed models, which usually assume that the effects follow a sparse distribution or a normal distribution. However, in complex animal breeding environments, these assumptions may not be sufficient to capture the complexity of gene-environment interactions, nonlinear genetic effects, and high-dimensional traits. Summary of the invention
[0003] The object of the present invention is to provide a parameter-free Bayesian genomic selection method for Huaxi cattle to solve the problems raised in the above background technology.
[0004] The present invention provides a parameter-free Bayesian genome selection method for Huaxi cattle, the method comprising: Step 1: construct a reference population and measure the important classification traits of each Huaxi cattle; Step 2: Blood is collected and preserved from each Huaxi cattle in the reference group, DNA is extracted and genotyped using a biochip to obtain the genotype of each SNP site, and quality control is performed on the genotyping data; Step 3: For each Huaxi cattle individual, the genotype file obtained in step 2 and the phenotype file recorded in step 1 are subjected to association analysis to obtain summary statistics of each phenotype, and each SNP site is divided into blocks according to the position of the chromosome to obtain an LD reference file; wherein the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level and initial site effect value, and the LD reference file includes the segments divided by each chromosome, the SNP site corresponding to each chromosome and the allele frequency; Step 4: For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual are used as the input files of the pre-built non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector to obtain the genomic estimated breeding value of the candidate population for the phenotype.
[0005] As a preferred embodiment, the important classification traits include calving difficulty, pH and meat color.
[0006] As a preferred embodiment, the quality control of the genotyping data includes: eliminating alleles with a frequency less than 0.05 and a p less than 10 in the Hardy-Weinberg equilibrium test. -6And SNP sites with missing genotype ratio less than 0.05.
[0007] As a preferred embodiment, the quality control of the data after genotyping is performed using plink1.9 software.
[0008] As a preferred embodiment, the biochip in step 2 is a Cattle110K high-density chip.
[0009] As a preferred embodiment, association analysis is performed on the genotype file obtained in step 2 and the phenotype file recorded in step 1 to obtain summary statistical data of each phenotype using plink2 software.
[0010] As a preferred embodiment, the constructed non-parametric Bayesian model is expressed as follows: ; in, represents the phenotype vector, X express N × p The genotype matrix; N Indicates the number of individuals ,p Indicates the number of sites ; express p ×1 vector of site effect values, where Following the mean of 0 and variance σ 2 The normal distribution of σ 2 , σ 2 ~DP( H , ), H represents the basic distribution, Indicates control σ 2 The concentration parameter at which the upper distribution shrinks toward H; represents the residual; The calculation process of the effect value of the SNP locus is as follows: ; in, represents the initial site effect value; represents the hyperparameters of the non-parametric Bayesian model; represents random weights and is independent of , Satisfy 0≤ ≤1 and , Indicates from Random variables generated independently from the distribution, 1 and is a distribution parameter, by controlling The distribution parameters 1 and Ability to flexibly adjust the generated distribution characteristics, Determines the weight generated after each break ; represents a random distribution; represents the initial prior distribution of the hypothesis; represents the LD matrix; represents normal distribution; Represents the identity matrix.
[0011] As a preferred embodiment, in step 4, the site effect value vector is obtained based on the calculated effect value of the SNP site using the Markov Chain Monte Carlo method.
[0012] Compared with the prior art, the present invention has the following beneficial effects: The present invention analyzes the genotype data of Huaxi cattle through a non-parametric Bayesian model, and can adaptively determine the complexity of the model without setting the number of clusters or other parameters in advance; and the non-parametric Bayesian model can more accurately evaluate the relationship between genotype and trait, thereby improving the prediction accuracy to optimize the accuracy of genome selection. Compared with traditional methods, the method of the present invention can better utilize the information in the data, identify genetic variations that have an important impact on the target trait, and improve the effect of selection. The purpose of the present invention is to establish an efficient and high-quality Huaxi cattle breeding method based on non-parametric Bayesian genome selection technology. When this method is applied to traits with a high proportion of non-additive effects, the genome prediction accuracy of this method is significantly improved compared with other methods, providing a new technical path for improving the accuracy of Huaxi cattle genome selection prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a flow chart of the reference-free Bayesian genome selection method for Chinese western cattle in Example 1.
[0014] Figure 2 Flow chart of the processing of summary statistics and LD reference files by the non-parametric Bayesian model in Example 1.
[0015] Figure 3 This is a schematic diagram comparing the genomic prediction accuracy of three traits of Chinese Western Cattle under three models in Example 2. DETAILED DESCRIPTION
[0016] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0017] Example 1
[0018] Combination Figure 1 This embodiment provides a parameter-free Bayesian genome selection method for Huaxi cattle, which includes: Step 1: construct a reference population and determine the important classification traits of the reference population; in a specific embodiment, the important classification traits include calving ease (CE), meat color (MC) and pH.
[0019] Step 2: collect blood from each cow in the reference population, extract DNA and perform genotyping using a biochip to obtain the genotype of each SNP (single nucleotide polymorphism) site; perform quality control on the data after genotyping. In some specific embodiments, the quality control of the data after genotyping includes: eliminating unqualified individuals and eliminating individuals with an allele frequency of less than 0.05 and a p value of less than 10 in the Hardy-Weinberg equilibrium test. -6 and SNP sites with a missing genotype ratio of less than 0.05, the biochip in step 2 is a Cattle110K high-density chip. In a specific embodiment, Figure 1 The number of SNP sites before quality control was 770K, and after quality control was 550K.
[0020] Step 3: Use plink2 software to perform association analysis on the genotype file obtained in step 2 and the phenotype file recorded in step 1 to obtain summary statistics of each phenotype; the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level and initial site effect value; and divide each SNP site into blocks according to the position of the chromosome to obtain an LD reference file; the LD reference file includes the segments divided by each chromosome, the SNP site corresponding to each chromosome, and the allele frequency of the SNP site; Step 4: For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual in the reference population are used as the input files of the pre-built non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector to obtain the genomic estimated breeding value (GEBV) of the candidate population for the phenotype. The genomic estimated breeding value (GEBV) is calculated as follows: ;in, represents the genomic breeding value of individual i, represents the genotype of individual i at site j, Represents the effect value of site j, and the site effect value vector is obtained based on the calculated effect value of the SNP site using mcmc sampling (Markov Chain Monte Carlo method).
[0021] Combination Figure 2 In this embodiment, the parameter distribution of the non-parametric Bayesian model is determined by the Dirichlet Process based on the existing data (summary statistics and LD reference files). The prior distribution is obtained from the data by the model without manually entering any parameters. Finally, the Markov chain Monte Carlo method is used for posterior sampling to obtain the genomic estimated breeding value.
[0022] Among them, the formula of the non-parametric Bayesian model is: ; in, represents the phenotype vector, X express N × p The genotype matrix; N Indicates the number of individuals ,p Indicates the number of sites ; express p ×1 vector of site effect values, where Following the mean of 0 and variance σ 2 The normal distribution of σ 2 , σ 2 ~DP( H , ), H represents the basic distribution, Indicates control σ 2 The concentration parameter at which the upper distribution shrinks toward H; Represents the residual. Specifically, there are several equivalent probability representations of the Dirichlet process, among which the truncated process is widely used because of its convenience in model fitting. The truncated process representation in the Dirichlet process regards the Dirichlet process as an infinite Gaussian mixture model.
[0023] Furthermore, the calculation process of the effect value of the SNP locus in step 4 is expressed as follows: ; in, represents the initial site effect value; Represents the hyperparameters of the non-parametric Bayesian model; one column in the summary statistics represents the initial site effect value obtained by processing each SNP site through the previous GWAS model (Genome-Wide Association Studies whole genome association analysis model) ; represents a random variable and is independent of , Satisfy 0≤ ≤1 and , Indicates from Random variables generated independently from the distribution, 1 and is a distribution parameter, by controlling The distribution parameters 1 and Ability to flexibly adjust the generated distribution characteristics, Determines the weight generated after each break in the Dirichlet stick-breaking process , It reflects that the non-parametric Bayesian model automatically identifies the probability mass of each potential cluster (or class) in the generated random distribution G according to the characteristics of the data; represents the LD matrix constructed according to the LD reference file; represents the initial prior distribution of the hypothesis; represents the hyperparameters of the non-parametric Bayesian model; represents normal distribution; Represents the identity matrix.
[0024] It can be understood that the present embodiment analyzes the Huaxi cattle genome data by a non-parametric Bayesian model. Compared with the traditional genome selection method, first, the method in the present embodiment can automatically determine the model complexity. The non-parametric Bayesian model can adaptively determine the complexity of the model without setting the number of clusters or other parameters in advance. Through the Dirichlet process, the non-parametric Bayesian model can automatically identify potential groups or clusters according to the characteristics of the data, thereby avoiding the overfitting or underfitting problems that may exist in the traditional method. Second, the method in the present embodiment can flexibly process high-dimensional data. The non-parametric Bayesian method has strong flexibility and robustness when processing high-dimensional data, so that it can effectively extract information from complex genomic data, automatically capture the structure and association in the data, and adapt to the potential complex patterns in the data, which is difficult to achieve in traditional linear models or fixed parameter models. Third, the method in the present embodiment can optimize the accuracy of genome selection, and the non-parametric Bayesian model can more accurately evaluate the relationship between genotype and trait, thereby improving prediction accuracy. Compared with traditional methods, it can better utilize the information in the data, identify genetic variants that have important effects on target traits, and improve the effectiveness of selection.
[0025] Example 2
[0026] This embodiment provides a specific embodiment of the method for selecting the genome of Western Chinese cattle in embodiment 1, and the specific steps are as follows: Step 1: Construct a reference population and measure important classification traits; the cattle genetic breeding innovation team of Beijing Animal Husbandry and Veterinary Research Institute has established a basic herd of 2,320 Huaxi cattle cows in the Ulagai Management Area of Xilin Gol League, Inner Mongolia since 2008. After expanding the herd year by year, the number of Huaxi cattle basic cows has exceeded 4,000 as of 2019. The offspring of this basic herd of cows are used to construct a reference population of Huaxi cattle, and the traits measured include calving difficulty (CE), pH and meat color (MC).
[0027] Step 2: Blood was collected from each cow in the reference population, DNA was extracted, and genotyping was performed using the Cattle110K (110k) high-density chip to obtain the genotype of each SNP locus. The data after genotyping were quality controlled using plink 1.9 software. The quality control included: eliminating unqualified individuals and eliminating individuals with an allele frequency less than 0.05 and a p less than 10 in the Hardy-Weinberg equilibrium test. -6 The SNP sites with a missing genotype ratio of less than 0.05 were selected. The biochip in step 2 was the Cattle110K high-density chip. The number of SNP sites after quality control was 100K, with 1165 Huaxi cattle data for CE traits, 1173 Huaxi cattle data for pH traits, and 1226 Huaxi cattle data for MC traits; the specific data are shown in Table 1.
[0028]
[0029] Step 3: For each Huaxi cattle individual, use plink2 software to perform association analysis on the genotype file obtained in step 2 and the phenotype file recorded in step 1 to obtain summary statistics of each phenotype; and divide each SNP site into blocks according to the position of the chromosome to obtain an LD reference file, wherein the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level and initial site effect value, and the LD reference file includes the segments divided by each chromosome, the SNP site and allele frequency corresponding to each chromosome; Step 4: Construct a non-parametric Bayesian model. For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual are used as the input files of the non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector obtained using the Markov chain Monte Carlo method to obtain the genomic estimated breeding value (GEBV) of the candidate population for the phenotype. ,in is the genomic breeding value of individual i, is the genotype of individual i at site j, is the effect value of site j), and the Pearson correlation coefficient between the genomic estimated breeding value (GEBV) and the true phenotype of the test set was calculated.
[0030] Step 5: Combine Figure 3 In this embodiment, the genome file and phenotype file are used as input data for the GBLUP model (Genomic best linear unbiased prediction) and the SDPR model (Summary-Data-based Penalized Regression Model) to predict the genome-estimated breeding value, and the summary statistics and LD reference file are used as input files for the SBayesRC (parameter-free Bayesian) model and the SDPR model to predict the genome-estimated breeding value. The Pearson correlation coefficients of the predicted genomic estimated breeding values (GEBV) under the three methods and the true phenotypes of the pre-acquired test set are calculated respectively. Figure 3 This is a comparison of the Pearson correlation coefficients of the above three methods under five-fold cross validation. From the box plot, it can be seen that for the genomic breeding values of each phenotype, the non-parametric Bayesian model has the best prediction performance.
[0031] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A parameter-free Bayesian genomic selection method for Huaxi cattle, characterized in that: include: Step 1: construct a reference population and measure the important classification traits of each Huaxi cattle; Step 2: Blood is collected and preserved from each Huaxi cattle in the reference group, DNA is extracted and genotyped using a biochip to obtain the genotype of each SNP site, and quality control is performed on the genotyping data; Step 3: For each Huaxi cattle individual, the genotype file obtained in step 2 and the phenotype file recorded in step 1 are subjected to association analysis to obtain summary statistics of each phenotype, and each SNP site is divided into blocks according to the position of the chromosome to obtain an LD reference file; wherein the summary statistics include the name of the SNP site, allele frequency, regression coefficient, significance level and initial point effect value, and the LD reference file includes the segments divided by each chromosome, the SNP site and allele frequency corresponding to each chromosome; Step 4: For each phenotype, the summary statistics and LD reference files of each Huaxi cattle individual are used as the input files of the pre-built non-parametric Bayesian model to calculate the effect value of the SNP site; the genotype vector of the candidate individual is multiplied by the site effect value vector to obtain the genomic estimated breeding value of the candidate population for the phenotype.
2. The method for selecting Huaxi cattle without reference Bayesian genome according to claim 1, characterized in that: The important classification traits include calving ease, pH and meat color.
3. The method for selecting Huaxi cattle without reference Bayesian genome according to claim 1, characterized in that: The quality control of the genotyping data includes: eliminating alleles with a frequency less than 0.05 and p less than 10 in the Hardy-Weinberg equilibrium test. -6 And SNP sites with missing genotype ratio less than 0.
05.
4. The method for selecting Huaxi cattle without reference Bayesian genome according to claim 3, characterized in that: The quality control of the genotyping data was performed using plink1.9 software.
5. The method for selecting Huaxi cattle without reference Bayesian genome according to claim 1, characterized in that: The biochip in step 2 is a Cattle110K high-density chip.
6. The method for selecting Huaxi cattle without reference Bayesian genome according to claim 1, characterized in that: Association analysis was performed between the genotype file obtained in step 2 and the phenotype file recorded in step 1 to obtain summary statistics of each phenotype using plink2 software.
7. The method for selecting Huaxi cattle without reference Bayesian genome according to claim 1, characterized in that: The constructed non-parametric Bayesian model is expressed as follows: ; in, represents the phenotype vector, X express N × p The genotype matrix; N Indicates the number of individuals ,p Indicates the number of sites ; express p ×1 vector of site effect values, where Following the mean of 0 and variance σ 2 The normal distribution of σ 2 , σ 2 ~DP( H , ), H represents the basic distribution, Indicates control σ 2 The concentration parameter at which the upper distribution shrinks toward H; represents the residual; The calculation process of the effect value of the SNP locus is as follows: ; in, represents the initial site effect value; represents the hyperparameters of the non-parametric Bayesian model; represents random weights and is independent of , Satisfy 0≤ ≤1 and , Indicates from Random variables generated independently from the distribution, 1 and are all distribution parameters; represents a random distribution; represents the initial prior distribution of the hypothesis; represents the LD matrix; represents normal distribution; Represents the identity matrix.
8. The method for selecting Huaxi cattle without reference Bayesian genome according to claim 7, characterized in that: In step 4, the site effect value vector is obtained based on the calculated effect value of the SNP site using the Markov Chain Monte Carlo method.
Citation Information
Patent Citations
New fast Bayes method estimating genomic estimated breeding value
CN107590364A
Western China cattle genome selection method
CN111243667A
Whole genome SNP (Single Nucleotide Polymorphism) information-based western China cattle genome matching method and application
CN118888008A
METHOD OF GENOMIC SELECTION OF CATTLE
RU2014110321A
Genomic selection (GS) breeding chip of huaxi cattle and use thereof
US20240043912A1
Cited By
Beef cattle genome matching method integrating genetic advantages and defects and application
CN121171334A