Multi-population genome selection method based on population differentiation index FST
By screening SNP sites and constructing a mixed model based on the population differentiation index FST method, the accuracy problem caused by LD inconsistency between populations in multi-population genome selection was solved, achieving higher genome breeding value prediction accuracy and lower prediction error.
Patent Information
- Application Number
- CN202510524683.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-10-03
AI Technical Summary
In multi-population genomic selection, inconsistent linkage disequilibrium caused by genetic differences between populations affects the accuracy of genomic selection, and existing methods such as GBLUP cannot effectively deal with this problem.
A method based on the population differentiation index FST was used to screen SNP sites in different populations, calculate the FST values and divide them into five intervals, and construct a mixed model to consider the LD inconsistency between QTLs and SNPs among populations, which was divided into the additive effects explained by the SNPs screened by the population differentiation index FST and the remaining SNPs. The conjugate gradient iteration method was used to estimate the genomic breeding value.
The accuracy of multi-population genomic selection has been improved, and the predictive power of genomic breeding values has been enhanced, which is manifested in higher prediction accuracy and lower prediction bias and error.
Smart Images

Figure BDA0005374932590000021 
Figure BDA0005374932590000022 
Figure BDA0005374932590000051
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological breeding, and in particular to a genome selection method based on a population differentiation index FST. Background Art
[0002] Genomic selection (GS) is an animal genetic assessment method proposed by Meuwissen et al. in 2001. This method uses genome-wide single nucleotide polymorphism (SNP) markers to estimate genomic estimated breeding values (GEBVs) and selects superior individuals based on the size of their GEBVs. A fundamental assumption of GS is that every quantitative trait locus (QTL) affecting a quantitative trait is in linkage disequilibrium (LD) with at least one SNP marker across the genome. Therefore, genomic selection can trace all QTLs affecting a trait, thereby capturing a large amount of additive genetic variance and improving prediction accuracy. Currently, the genomic best linear unbiased prediction (GBLUP) method is a common genomic selection method for animals, plants, and aquatic animals. Theoretical and practical breeding practices have demonstrated that the GBLUP method is more accurate than traditional breeding methods, accelerating breeding progress and improving breeding efficiency.
[0003] Population size is a key factor influencing the accuracy of genomic selection; larger populations lead to higher accuracy. Multi-population genomic selection, which increases the size of the reference population, has become a common practice in breeding due to its cost-effectiveness. However, genetic differences often exist between populations due to drift, environmental adaptation, or artificial selection, leading to inconsistent LD between QTLs and SNPs across populations. Because the GBLUP method assumes consistent LD between populations, directly using GBLUP can affect the accuracy of genomic selection.
[0004] The fixation index (FST) is an important statistic for evaluating allele frequency differences between populations. It is widely used in population structure analysis, especially for detecting gene regions that may be affected by environmental or human selection, the so-called "genomic selection signals."
[0005] Based on the above-mentioned problem of genome selection accuracy, the present invention establishes a multi-population genome selection method based on the population differentiation index FST, aiming to improve the accuracy of multi-population genome selection. Summary of the Invention
[0006] In response to the problems existing in existing genomic selection methods, the present invention proposes a multi-population genomic selection method based on the population differentiation index FST. The method of the present invention can take into account the LD inconsistency of QTLs and SNPs between populations, accurately predict genomic breeding values, and improve the accuracy of multi-population genomic selection.
[0007] In order to achieve the above object, the present invention adopts the following technical measures:
[0008] A multi-population genome selection method based on population differentiation index FST comprises the following steps:
[0009] Step 1: Screen all SNP sites in different populations of the species to be tested and verify the inconsistency of LD among different populations;
[0010] Step 2: If LD is inconsistent between populations, use PLINK software to calculate the FST value of each locus. According to the FST value of each SNP, the SNP is divided into five intervals, that is, FST value <0.05, 0.05-0.15, 0.15-0.25, 0.05-0.25 and greater than 0.05; any one of the groups of 0.05-0.15, 0.05-0.25 and greater than 0.05 is selected for the estimation of the following genomic breeding value;
[0011] Step 3: Divide the genomic breeding value GEBV into two parts: the additive effect explained by the SNPs selected by the population differentiation index FST and the additive effect explained by the remaining SNPs. The matrix form of the genomic selection model can be expressed as:
[0012] y=Xb+Zf+Zr+e,
[0013] Where y is the phenotypic value vector, b is the fixed effect, and f is the additive genetic effect vector, which obeys the normal distribution N(0,fstGσ afst 2 ), where fstG is the additive genomic kinship matrix constructed by SNPs screened by the population differentiation index FST, σ afst 2 is the additive genetic variance; r is the additive genetic effect vector explained by the remaining SNP markers, which follows the normal distribution N(0,Gσ r 2 ), where G is the genomic kinship matrix constructed by the remaining SNP markers, σ r 2is the additive genetic variance; e is the random residual, which obeys the normal distribution e~N(0,R)=N(0,Iσ e 2 ), where I is the identity matrix, σ e 2 is the random residual variance; X and Z are the corresponding association matrices. Among them, fstG and G are constructed by the SNPs screened by the population differentiation index FST and the remaining SNP markers, respectively, and the construction formula is the same, both:
[0014]
[0015] where p i is the frequency of the second allele at the i-th site, W is an n×m matrix (n is the number of individuals, m is the number of markers), where the value for genotype AA is 0-2p i , for genotype Aa the value is 1-2p i , for genotype aa the value is 2-2p i .
[0016] Step 4: Estimation of genomic breeding values
[0017] A mixed model equation set was established and the conjugate gradient iteration method was used to estimate the individual genomic breeding values. The mixed model equation set can be expressed as:
[0018]
[0019] Based on the estimated f and r values, the genomic breeding value GEBV = f + r.
[0020] In the above method, preferably, the interval in step 2 is selected as a group of FST values between 0.05 and 0.25 or greater than 0.05;
[0021] The application of the above-described method in biological breeding.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] Compared with the GBLUP model, the model fstGBLUP proposed in the present invention can take into account the LD inconsistency between QTLs and SNPs among populations, and establish an additive genomic kinship matrix constructed by SNPs screened by the population differentiation index FST and an additive genomic kinship matrix constructed by the remaining SNPs, respectively, which has higher predictive power, that is, higher genomic selection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 Principal component analysis and linkage disequilibrium decay plots calculated for the whole genome sequencing data of three selected carp populations.
[0025] Figure 2 Schematic diagram of the population differentiation index FST calculated for the whole-genome sequencing data of three common carp populations.
[0026] Figure 3 Schematic diagram of the genomic prediction accuracy (Accuracy), bias (Bias), mean square error (Mse), and mean absolute error (Mae) of SNP markers and whole genome sequencing (WGS) markers for sex reversal traits based on the fstGBLUP method to compare different population differentiation indices FST.
[0027] Figure 4 Schematic diagram of the genomic prediction accuracy (Accuracy), bias (Bias), mean square error (Mse), and mean absolute error (Mae) of SNP markers and whole genome sequencing (WGS) markers for sex reversal traits based on the GBLUP method. DETAILED DESCRIPTION
[0028] Unless otherwise specified, the technical solutions described in the present invention are all conventional solutions in the field; the reagents or materials described are all from commercial channels unless otherwise specified.
[0029] Example 1:
[0030] SNP screening, population genetic structure analysis and FST calculation of carp population
[0031] 1) Screening of SNPs in three carp populations
[0032] Three common carp populations were selected, with 75, 70, and 70 individuals from each population subjected to high-temperature-induced sex reversal phenotypes (CN118435911A). Fin ray samples were collected, and DNA was extracted from the fin ray tissue of each broodstock. DNA from these 215 fish was sent to a sequencing company for next-generation resequencing at a depth of 10× to establish a reference population. After obtaining the whole-genome resequencing data, the raw data were quality-controlled using FastQC software, removing adapter sequences, reads less than 50 bp in length, and reads of low base sequencing quality. Based on the high conservation of common carp genomes, the quality-controlled reads were aligned to the common carp reference genome (GCA_018340385.1) using BWA software with default parameters. The resulting sam files were converted to bam files using SAMtools. Duplicate sequences were marked using the MarkDuplicates module in GATK software, and the marked duplicates were removed using the rmdup module in Samtools software to generate deduplicated bam files. The HaplotypeCaller in GATK software was used to call SNPs in the deduplicated bam files, and the VariantFiltration module in GATK software was used to filter low-quality SNPs. The sample genotype data were then filled using Beagle / 5.1 software to ensure that the number of SNP markers for all individuals was consistent. All individuals and SNPs were then quality controlled according to the following criteria: individuals with a call rate < 95% were deleted; individuals with a call rate < 95%, a minimum allele frequency (MAF) < 0.01, and a Hardy-Weinberg equilibrium < 10 were deleted. -6 Finally, 19,111,782 SNP sites were obtained for subsequent analysis.
[0033] 2) Genetic structure analysis confirmed inconsistent LD among the three populations
[0034] The genetic structure of the three carp populations was then analyzed, and principal component analysis (PCA) was performed on the quality-controlled data (including 19,111,782 SNP sites). PCA was performed using PLINK software, and the results were visualized using the ggplot package in the R language. At the same time, linkage disequilibrium (LD) was calculated using PopLDDecay software, and the LD decay plot of the population was obtained using the ggplot package in the R language. Figure 1 The two figures shown in the figure correspond to the principal component analysis and LD decay plot, respectively. The principal component analysis plot has the first principal component on the horizontal axis and the second principal component on the vertical axis. The results show that the three populations have significant population genetic differences and are three different populations. The LD decay plot has physical position on the horizontal axis and LD (r²) value on the vertical axis, indicating that the LD among the three populations is inconsistent.
[0035] 3) Calculation and grouping of differentiation index FST
[0036] Then, the FST values of the 19,111,782 loci screened above were calculated using PLINK software to assess the degree of genetic differentiation among these populations. The results are as follows: Figure 2 As shown, it shows that the FST values of different SNP markers are different.
[0037] Based on the applicant's long-term experience, larger values indicate greater differentiation. Therefore, based on the FST value, the applicant divided all the SNPs screened above into five FST intervals to represent different levels of genetic differentiation: FST values <0.05, 0.05-0.15, 0.15-0.25, 0.05-0.25, and greater than 0.05. The number of SNPs corresponding to the five FST intervals was 2,926,693, 5,572,917, 5,010,089, 10,583,008, and 10,583,008, respectively.
[0038] Example 2:
[0039] Construction of fstGBLUP selection model:
[0040] In this example, the genomic breeding value (GEBV) based on the fstGBLUP model is composed of the additive effect explained by the SNPs selected by the population differentiation index (FST) and the additive effect explained by the remaining SNPs. The SNP selected by the population differentiation index (FST) is one of the SNPs in the five intervals described above, and the remaining SNPs is the difference between the 19,111,782 SNPs and the number of SNPs in one of the five intervals selected.
[0041] In this example, the genomic breeding values obtained by substituting whole genome sequencing (WGS) markers (ie, 19,111,782 SNPs) into the conventional GBLUP model were used as a control.
[0042] The fstGBLUP constructed by the present invention is as follows:
[0043] The genomic breeding value GEBV is divided into two parts: the additive effect explained by the SNPs screened by the population differentiation index FST and the additive effect explained by the remaining SNPs. The matrix form of the genomic selection model fstGBLUP can be expressed as:
[0044] y=Xb+Zf+Zr+e,
[0045] Where y is the phenotypic value vector, b is the fixed effect, and f is the additive genetic effect vector, which obeys the normal distribution N(0,fstGσ afst 2), where fstG is the additive genomic kinship matrix constructed by SNPs screened by the population differentiation index FST, σ afst 2 is the additive genetic variance; r is the additive genetic effect vector explained by the remaining SNP markers, which follows the normal distribution N(0,Gσ r 2 ), where G is the genomic kinship matrix constructed by the remaining SNP markers, σ r 2 is the additive genetic variance; e is the random residual, which obeys the normal distribution e~N(0,R)=N(0,Iσ e 2 ), where I is the identity matrix, σ e 2 is the random residual variance; X and Z are the corresponding association matrices. Among them, fstG and G are constructed by the SNPs screened by the population differentiation index FST and the remaining SNP markers, respectively, and the construction formula is the same, both:
[0046]
[0047] where p i is the frequency of the second allele at the i-th site, W is an n×m matrix (n is the number of individuals, m is the number of markers), where the value for genotype AA is 0-2p i , for genotype Aa the value is 1-2p i , for genotype aa the value is 2-2p i .
[0048] A mixed model equation group was established and the conjugate gradient iteration method was used to estimate the individual genomic breeding values. The mixed model equation group can be expressed as:
[0049]
[0050] Based on the estimated f and r values, the genomic breeding value GEBV = f + r.
[0051] Based on the above model, the applicant constructed the genomic breeding value of SNP using 5 intervals as follows: Figure 3 shown and compared with the control group.
[0052] The results of the two models described above were evaluated using a cross-validation strategy. Genotyped individuals were randomly divided into five subpopulations of similar size. In each iteration, one of the subpopulations was selected as the validation set, and the remaining four subpopulations formed the reference set. Each subpopulation was used as the validation set in turn until all folds were validated. To ensure the robustness of the results, the entire 5-fold cross-validation process was repeated 20 times across all scenarios.
[0053] Prediction accuracy was assessed by calculating the Pearson correlation coefficient between the observed phenotypic values and the genomic estimated breeding values (GEBVs) in the validation set.
[0054] The evaluation of prediction bias (Bias) is based on the regression coefficient of phenotypic values to genomic estimated breeding values. The degree of bias is quantified by the absolute value of the regression coefficient deviating from 1, which is used to reflect the systematic error of the prediction.
[0055] In addition, the performance of the models was compared using two metrics: Mean Squared Error (MSE) and Mean Absolute Error (MAE). MSE measures the squared error between the predicted value and the true value, while MAE measures the mean absolute difference between the predicted value and the true value.
[0056] The results are as follows Figure 3 As shown, the results show that there is no significant difference in prediction bias (Bias), mean square error (MSE) and mean absolute error (MAE) under different SNP data sets. The accuracy of SNP markers screened by the population differentiation index FST for the fstGBLUP method is higher than that of whole genome sequencing (WGS) markers for GBLUP. The three regions of FST values 0.05-0.15, 0.05-0.25 and greater than 0.05 are suitable for the fstGBLUP model of the present invention.
[0057] Example 3:
[0058] Comparison of SNPs screened based on the population differentiation index FST with SNPs from whole genome sequencing:
[0059] The GBLUP method was used to compare the genomic prediction effects of SNP markers screened by different population differentiation indices FST (i.e., 5 groups of SNP markers were substituted into GBLUP for estimation of genomic breeding values) and whole genome sequencing (WGS) markers on sex reversal traits. The cross-validation strategy of the results is shown in Example 2.
[0060] The results are as follows Figure 4 As shown, the four figures are prediction accuracy (Accuracy), prediction bias (Bias), mean square error (MSE) and mean absolute error (MAE).
[0061] Results showed no significant differences in prediction bias (Bias), mean squared error (MSE), and mean absolute error (MAE) across different SNP datasets. However, SNP markers selected using the population differentiation index (FST) showed higher accuracy in the GBLUP model than those using whole-genome sequencing (WGS). The GBLUP model of the present invention was also applicable to SNP markers with a range of 0.05-0.25 and greater than 0.05.
Claims
1. A multi-population genome selection method based on population differentiation index FST, comprising the following steps: Step 1: Screen all SNP sites in different populations of the species to be tested and verify the inconsistency of LD among different populations; Step 2: If LD is inconsistent between populations, use PLINK software to calculate the FST value of each locus. According to the FST value of each SNP, the SNP is divided into five intervals, that is, FST value <0.05, 0.05-0.15, 0.15-0.25, 0.05-0.25 and greater than 0.05; any one of the groups of 0.05-0.15, 0.05-0.25 and greater than 0.05 is selected for the estimation of the following genomic breeding value; Step 3: Divide the genomic breeding value GEBV into two parts: the additive effect explained by the SNPs selected by the population differentiation index FST and the additive effect explained by the remaining SNPs. The matrix form of the genomic selection model can be expressed as: y=Xb+Zf+Zr+e, in, y is the phenotypic value vector, b is the fixed effect; f is the additive genetic effect vector, which obeys the normal distribution N(0,fstGσ afst 2 ), where fstG is the additive genomic kinship matrix constructed by SNPs screened by the population differentiation index FST, σ afst 2 is the additive genetic variance; r is the additive genetic effect vector explained by the remaining SNP markers, which follows the normal distribution N(0,Gσ r 2 ), where G is the genomic kinship matrix constructed by the remaining SNP markers, σr 2 is the additive genetic variance; e is the random residual, which obeys the normal distribution e~N(0,R)=N(0,Iσ e 2 ), where I is the identity matrix, σ e 2 is the random residual variance; X and Z are the corresponding association matrices; fstG and G are constructed by the SNPs screened by the population differentiation index FST and the remaining SNP markers, respectively, and the construction formulas are the same, both: where p i is the frequency of the second allele at the i-th site, W is an n×m matrix (n is the number of individuals, m is the number of markers), where the value for genotype AA is 0-2p i , for genotype Aa the value is 1-2p i , for genotype aa the value is 2-2p i ; Step 4: Estimation of genomic breeding values A mixed model equation group was established and the conjugate gradient iteration method was used to estimate the individual genomic breeding values. The mixed model equation group can be expressed as: Based on the estimated f and r values, the genomic breeding value GEBV = f + r.
2. The method according to claim 1, characterized in that The interval in step 2 is selected as a group of FST values between 0.05 and 0.25 or greater than 0.
05.
3. Application of the method according to claim 1 in biological breeding.
Citation Information
Patent Citations
Method for inducing cyprinus carpio from secondary female to male at high temperature and application
CN118435911A