Marker screening and model building method for precise evaluation of parent genome breeding value
By introducing MAF and LD score to screen core feature subsets in the genome prediction model, the problem of insufficient cross-generation prediction stability in existing technologies is solved, and the accurate evaluation of parental breeding values and the stability of cross-generation application are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- OCEAN UNIV OF CHINA
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-05
AI Technical Summary
Existing genomic selection technologies suffer from insufficient predictive stability in cross-generational applications and difficulties in balancing predictive accuracy and cross-sample stability with marker utilization strategies. This is particularly evident in aquatic breeding, leading to inconsistent parental breeding value predictions and affecting the reliability of breeding decisions.
By using minor allele frequency (MAF) and linkage disequilibrium score (LD score) as joint evaluation indicators, a subset of core features was screened out and given high weight in the genome prediction model, thus constructing a model suitable for predicting parental breeding values and its cross-generational application.
It improves the prediction accuracy of parental genome breeding values and the consistency of cross-generational predictions, reduces noise interference, enhances the ability to transmit genetic effects in complex genetic backgrounds, and improves the prediction performance for cross-generational applications.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular breeding and genomics technology, specifically relating to a marker screening and model construction method for accurate evaluation of parental genome breeding values. Background Technology
[0002] With the continuous development of life sciences and information technology, molecular breeding technologies, represented by genomic selection (GS), have gradually become important breeding techniques following traditional phenotypic selection and molecular marker-assisted selection. Genomic selection deploys high-density single nucleotide polymorphism (SNP) markers across the entire genome and combines them with individual phenotypic data to construct predictive models between genotypes and target traits, thereby enabling early assessment and screening of the genomic estimated breeding value (GEBV) of candidate individuals. This technology allows for breeding decisions without relying on complete phenotypic determination, helping to shorten breeding cycles and reduce breeding costs. It has significant advantages in handling traits with low heritability or difficult phenotypic determination (such as disease resistance and growth efficiency), and has been widely applied in the breeding practices of livestock, poultry, crops, and aquaculture species. In existing applications, genomic selection technology typically employs the following workflow: First, a reference population containing known genotype and phenotypic data is constructed as a training sample set. Statistical modeling methods are then used to establish a mapping relationship between genotype characteristics and target traits. Subsequently, the constructed predictive model is applied to a candidate population with only genotype information to achieve early selection and evaluation of individuals. Common modeling methods include prediction methods based on linear mixture models, Bayesian regression models, and other machine learning methods, which have been applied in breeding practices for multiple species.
[0003] However, in practical industrial applications, existing genomic selection technologies still face a series of pressing technical problems in predicting parental breeding values and their cross-generational applications. First, the predictive models lack stability in cross-generational scenarios. Because the number of whole-genome SNP markers is typically much higher than the number of reference population samples available for modeling, parental breeding value prediction models are prone to problems such as unstable parameter estimation, feature redundancy, and noise accumulation during the establishment process, leading to increased dependence of parental individuals' GEBV on training data. Under cross-validation conditions within the reference population, the test set GEBV often shows a high correlation with its own phenotype. However, in practical breeding applications, the main purpose of candidate parental GEBV is to predict the phenotypic performance of their offspring populations. When the candidate parental GEBV is used for cross-generational breeding decisions, the correlation between its GEBV and its offspring representative type, or with family phenotypic indicators obtained based on offspring representative type statistics, often decreases significantly. This phenomenon indicates that existing methods generally suffer from insufficient cross-generational predictive stability in predicting parental breeding values: some markers or effects that perform well in the reference population may fail to maintain consistent predictive contributions in offspring populations, leading to rearrangement of parental breeding value order across generations and affecting the reliability of parental mating schemes and genetic gain assessments. Secondly, existing marker utilization strategies struggle to simultaneously ensure predictive accuracy and cross-sample stability. Current techniques often screen for marker effects based on magnitude or statistical significance to improve the model's explanatory power for trait variation. However, during generational succession, factors such as genetic recombination, changes in allele frequencies, and population restructuring can alter the linkage relationships and statistical characteristics between different SNP markers, causing some markers with high predictive contributions in training samples to exhibit unstable predictive performance in subsequent samples or offspring populations. This instability is particularly pronounced in cross-generational applications. Furthermore, these problems are even more significant in applications such as aquatic breeding. Many aquatic species are characterized by high genetic diversity, rapid linkage disequilibrium decay, and large family sizes. In actual breeding processes, it is usually necessary to measure and evaluate offspring traits at the family level. If modeling is directly based on whole-genome markers, the prediction model is easily affected by interference from close relatives or background genetic signals, making it difficult for the genetic information captured by the model to remain consistent across different generations. This affects the reliability of the parental breeding value prediction results in actual breeding decisions.
[0004] Therefore, how to effectively reduce redundant noise and interference from unstable cross-generational features in massive, high-dimensional genomic feature data, screen out a set of key markers with high predictive stability across different generations, and construct a genomic prediction model based on the marker set that is suitable for predicting parental breeding values and its cross-generational applications has become an important technical problem that urgently needs to be solved in the field of molecular breeding technology. Summary of the Invention
[0005] To address the shortcomings of existing genomic selection technologies in cross-generational applications, such as insufficient predictive stability due to marker redundancy, genetic recombination, and allele frequency variations, as well as declining predictive performance in highly heterozygous, highly polymorphic populations with rapid linkage disequilibrium (LD) decay, this invention provides a marker screening and model construction method for accurate evaluation of parental genomic breeding values. The core of this invention lies in using minor allele frequency (MAF) and linkage disequilibrium score as joint evaluation indicators of locus robustness and local linkage structure representativeness. A core feature subset is determined according to a joint threshold or joint quantile screening rule. In the genomic prediction model, core feature markers are assigned higher weights than non-core feature markers, or modeling is performed solely based on core feature markers, thereby improving the consistency and stability of parental genomic breeding values in cross-generational validation.
[0006] To achieve the above objectives, the present invention provides the following technical solution, comprising the following steps:
[0007] Step 1: Obtain whole-genome SNP genotyping data and target trait phenotypic data of the reference population, and perform quality control and numerical coding on the SNP genotyping data to obtain genotype matrix X and phenotype vector y;
[0008] Step 2: For each SNP locus in the genotype matrix X, calculate the minor allele frequency (MAF) and linkage disequilibrium score (LD score) of that SNP locus.
[0009] Step 3: Based on the MAF and the LD score, construct a joint screening rule, and screen the SNP sites that simultaneously meet the conditions according to the preset threshold or quantile criterion, as the core feature subset Score;
[0010] Step 4: Construct a genome prediction model based on the core feature subset Score, assign higher weights to SNP sites within the core feature subset, or construct the prediction model based solely on the core feature subset, in order to establish a mapping relationship between genotype data and predicted breeding values.
[0011] Step 5: Type the candidate parent population, use the genome prediction model to obtain the predicted breeding value of the candidate parents, and select parents to construct families accordingly. Perform phenotypic determination on the offspring and calculate the phenotypic mean of the offspring in the family.
[0012] Step Six: Using families as units, calculate the mean or weighted combination of the predicted breeding values of the parental individuals in the same family as the family breeding value, and calculate the correlation coefficient between the family breeding value and the mean of the corresponding family subtype, as an evaluation index of the accuracy of the parental genome breeding value prediction.
[0013] Furthermore, in step one, the quality control includes removing individuals and loci with a genotype detection rate of less than 80%; the numerical coding uses an additive model to encode homozygous reference genotype as 0, heterozygous genotype as 1, and homozygous mutant genotype as 2.
[0014] Furthermore, in step two, the calculation method for the linkage disequilibrium fraction parameter includes: dividing the SNP sites according to chromosomes; for each chromosome, sequentially traversing each SNP site located on that chromosome; constructing a sliding window with the SNP site as the center, within a physical distance range of 100 kb upstream and downstream; and calculating the Pearson correlation coefficient r between the target SNP site and other SNP sites within the sliding window. 2 and for the r 2 The values are summed to obtain the linkage disequilibrium fraction parameter corresponding to the SNP site.
[0015] Furthermore, in step three, the minor allele frequency threshold and linkage disequilibrium fraction threshold are set based on the distribution of SNP sites across the whole genome, and the core feature subset is the set of SNP sites that are simultaneously located in the high quantile interval of the minor allele frequency distribution and the high quantile interval of the linkage disequilibrium fraction distribution.
[0016] Furthermore, the high quantile interval of the minor allele frequency distribution is the top 10% quantile interval of the minor allele frequency distribution of SNP sites in the whole genome, and the high quantile interval of the linkage disequilibrium fraction distribution is the top 10% quantile interval of the linkage disequilibrium fraction distribution of SNP sites in the whole genome.
[0017] Furthermore, in step four, the genome prediction model is a linear model or a nonlinear model, including one or more of linear mixed models, random effects models, and machine learning models, or equivalent variations thereof.
[0018] Furthermore, in step five, the target trait for the sub-representative type is the same as that of the reference population, and the measurement method, evaluation criteria, and experimental conditions are consistent with those of the reference population phenotypic vector y to obtain the sub-representative type data.
[0019] Furthermore, in step six, the correlation coefficient is calculated as follows:
[0020] ;
[0021] in, This represents the breeding value of the i-th family. This represents the mean of the breeding values of all families. Let represent the sub-representative mean of the i-th family lineage. denoted as the mean of the subtypes of all families, and N represents the number of families.
[0022] Compared with existing technologies, the advantages of this invention are as follows:
[0023] This invention introduces minor allele frequency (MAF) and linkage disequilibrium score (LD score), which characterizes the abundance of variation information represented by a locus in a local linkage structure, as core evaluation indicators to quantitatively characterize the population genetic stability and representativeness of local genome structure. Based on the joint high-quantile screening rule of MAF and LD score, this invention can screen out a subset of core features with high genetic stability and representativeness from high-dimensional marker loci across the entire genome. Through the above screening method, an effective dimensionality reduction is achieved from the massive number of genetic loci across the entire genome to key genetic information loci. This reduces the number of genotyping markers, lowers gene detection and computational costs, and helps improve the signal consistency and effectiveness of features involved in modeling.
[0024] Furthermore, this invention assigns differentiated weights to markers within the core feature subset and non-core feature subset in the genome prediction model. This results in markers within the core feature subset having a higher contribution to the genetic effect in the model, thereby enhancing the ability to characterize additive genetic effects that can be transmitted across generations in complex genetic contexts and suppressing noise interference introduced by low-stability sites. Using the method of this invention, the prediction accuracy of parental individual genomic breeding value (GEBV) can be improved, and the consistency between parental prediction results and offspring measured phenotypes can be enhanced. This alleviates the problem of declining prediction performance of existing genome prediction models in cross-generational applications, providing an effective technical means for achieving low-cost, robust cross-generational molecular breeding. Attached Figure Description
[0025] The above content and the following detailed embodiments of the present invention will be better understood when read in conjunction with the accompanying drawings. It should be noted that the drawings are merely examples of the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention.
[0026] Figure 1 A flowchart illustrating the construction of the marker screening and genome prediction model provided by this invention.
[0027] Figure 2 is a logic diagram for joint screening of genome-wide loci using minor allele frequency (MAF) and linkage disequilibrium score (LD Score); where: Figure 2AThis is a map showing the distribution of SNP sites in the reference population of Pacific oyster. Figure 2B The figure shows the distribution of SNP loci in the reference population of large yellow croaker; the horizontal axis represents minor allele frequency (MAF), and the vertical axis represents linkage disequilibrium score (LD Score); the red vertical and horizontal dashed lines in the figure represent the top 10% quantile thresholds set based on the whole genome distribution; the loci clustered in the shaded area in the upper right corner of the figure (labeled "Score") simultaneously satisfy high polymorphism and high genetic representativeness, which is the core feature subset constructed by this invention.
[0028] Figure 3 compares the prediction accuracy of parental genomic breeding values under different SNP screening strategies. The prediction accuracy of different SNP sets in predicting offspring phenotypic traits across generations of parental breeding values was evaluated using progeny validation to assess the model's robustness. Figure 3A Predicted parental genomic breeding values for oyster resistance to Vibrio alginolyticus infection. Figure 3B The parental genome breeding values for the trait of resistance to Cryptocaryon irritans infection in large yellow croaker are predicted. The SNP sets represented by each box plot are defined as follows: All SNPs: The set of all SNP sites in the whole genome after quality control (as the basic background); Prioritized SNPs: The core feature subset obtained by the joint screening of MAF and LD Score in this invention. The accuracy evaluation results show that, with a significantly reduced number of sites, the cross-generation prediction accuracy of this set is significantly higher than that of the all-genome site set; Uniformly spaced SNPs: As a negative control, sites with the same number of sites as the core feature subset are uniformly extracted throughout the whole genome. Detailed Implementation
[0029] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0030] Example 1:
[0031] like Figure 1 As shown, a marker screening and model construction method for accurate evaluation of parental genome breeding values is described, and the steps of the method are as follows:
[0032] S1: Obtain whole-genome SNP genotyping data and target trait-related phenotypic data of the reference population, and generate a digital genotype matrix;
[0033] S2, traverse the matrix to calculate the minor allele frequency (MAF) and linkage disequilibrium score (LD score) at each point in the whole genome.
[0034] S3, set dual quantile thresholds and use joint screening rules to construct a highly robust subset of core features;
[0035] S4. Construct a weighted relation matrix based on the core subset to estimate the genomic breeding value of candidate individuals;
[0036] S5, construct pedigrees for the candidate population based on the breeding value screening criteria;
[0037] S6 uses correlation analysis between parental predicted values and offspring pedigree phenotypes to evaluate the accuracy of intergenerational prediction.
[0038] In one implementation method, step S1 specifically involves selecting a representative breeding population with a known pedigree or genetic background as a reference population; using high-throughput SNP chips or genome resequencing technology to perform whole-genome DNA sampling and genotyping on individuals in the reference population to obtain raw SNP locus information; measuring the target phenotypic data of the individuals in the reference population under specific environmental or experimental conditions to form a phenotypic vector y; performing quality control filtering on the raw SNP data to remove individuals and loci with a detection rate of less than 80%; and converting different genotype states into numerical representations according to an additive effect model, i.e., homozygous reference genotype is encoded as 0, heterozygous as 1, and homozygous mutant as 2, generating a genotype value matrix X.
[0039] As one implementation method, step S2 specifically involves: traversing the genotype value matrix and calculating the following statistical characteristic parameters for each SNP locus: (a) minor allele frequency parameter, used to describe the allele distribution of the SNP locus in the reference population; (b) linkage disequilibrium score parameter, used to describe the characteristic quantity of the surrounding genetic variation captured by the SNP locus.
[0040] As one implementation method, step S3 specifically involves: constructing a joint feature evaluation rule based on the minor allele frequency parameter and linkage disequilibrium score parameter, setting a minor allele frequency threshold and a linkage disequilibrium score threshold, mapping SNP sites that simultaneously meet the threshold conditions to a core feature subset, and removing low-frequency or weakly correlated sites that do not meet the threshold conditions.
[0041] As one implementation method, step S4 specifically involves: constructing a genome prediction model based on the core feature subset, setting differential weights for the markers in the core feature subset and non-core feature subset, performing regression estimation or random effect estimation of SNP genetic effects, and establishing a mapping relationship between genotype data and predicted breeding values.
[0042] As one implementation method, step S5 specifically involves: genotyping the candidate population to be predicted and extracting SNP loci information corresponding to the core feature subset; evaluating the breeding value of candidate parents based on the constructed model, and selecting individuals with high breeding values as parents for family construction based on the breeding objectives; after the offspring grow to the phenotypic determination standard, obtaining offspring phenotypic data using the same determination method, evaluation standard, and experimental conditions as the reference population phenotypic vector y; statistically analyzing the phenotypic data of offspring individuals in each family, and calculating the mean of the offspring phenotypic values of each family as the family phenotype.
[0043] As one implementation method, step S6 specifically involves comparing the predicted breeding value of the family with the statistical value of the sub-representative type of the corresponding family, and using the Pearson correlation coefficient between the two as an evaluation index for the accuracy of the parental breeding value prediction.
[0044] Example 2:
[0045] Accuracy assessment of prediction of anti-Vibrio trait in Crassula longifolia based on screening with high MAF and high LD markers
[0046] S1. Reference Population Construction and Data Acquisition
[0047] (1) Reference group construction
[0048] In May 2022, a G4 generation resistant population was bred using the G3 generation of anti-Vibrio oryzae families as parents, and Haida No. 1 and Haida No. 4 were used as parental control populations, resulting in a total of 45 families. After the juveniles reached 2-3 mm in size, they were transferred to the natural sea area of Rongcheng for culture. After reaching one year of age, 2124 oysters were selected as a reference population for phenotypic evaluation.
[0049] (2) Determination of resistance phenotype
[0050] Artificial Vibrio infection and challenge experiments were conducted on the reference population. For each individual, the following were recorded: ① Survival status (0 for death, 1 for survival); ② Survival time (calculated in hours from the start of the challenge experiment to death). Adductor muscle DNA was extracted simultaneously, and the Vibrio viral load in each individual was assessed using Vibrio genome alignment rate as an auxiliary resistance indicator.
[0051] (3) Genotyping and quality control
[0052] Raw SNP data were obtained using high-throughput sequencing genotyping. Quality control was performed to obtain high-quality genomes: samples with an individual deletion rate >10% were removed; SNPs with a locus detection rate <90% were removed; and the remaining data were filled to obtain a genotype dataset that could be used for matrix operations.
[0053] (4) Digital coding
[0054] The genotype at each locus is converted into a numerical value according to the additive coding rule: homozygous reference allele is recorded as 0, heterozygous as 1, and homozygous variant allele as 2, forming a genotype matrix. , where n is the number of individuals and m is the number of SNP loci.
[0055] S2. Rules for Extracting and Calculating Marker Feature Parameters
[0056] For each SNP site i, the following characteristic parameters are calculated.
[0057] (1) Minor allele frequency (MAF)
[0058] Calculate the minor allele frequency at locus i: ,in denoted as the minor allele frequency at the i-th locus; n is the number of samples in the reference population. Let be the genotype value of the j-th individual at the i-th locus, ( ) 2n represents the total number of alleles with mutations at this locus; 2n represents the total number of alleles at this locus in all individuals; the smaller of the two allele frequencies is taken as the final MAF parameter by the min function.
[0059] (2) Chain Disequilibrium Score (LD score)
[0060] Centered on site i, a sliding window with physical distances of 100 kb upstream and downstream is set, and the set of sites within the window is denoted as W(i). For any Calculate the Pearson correlation coefficient between standardized genotype vectors. And take the square. Then the LD score of site i is defined as:
[0061]
[0062] The l i is the LD score of site i, used to characterize the abundance of local linkage structure information represented by site i.
[0063] S3, Core Feature Subset S core Filtering
[0064] For all sites of the whole genome and The statistical distribution is used to calculate quantiles, and a dual-threshold filtering rule is set.
[0065] Set the filter percentile threshold to the top 10 percentile, defined as:
[0066]
[0067]
[0068] The joint selection and subset mapping will simultaneously satisfy: The set of sites is mapped to a subset of core features: ( Figure 2A The remaining sites are considered as a set of non-core sites. This screening strategy enriches a set of markers that are more robust to cross-sample prediction by combining the constraints of "high frequency + strong linkage representativeness" and reduces the instability of the model caused by low frequency or weak linkage sites.
[0069] S4. Construction of Genomic Selection Model
[0070] (1) Model framework
[0071] Using the Genomic Best Linear Unbiased Prediction (GBLUP) model as an example, its mathematical expression is as follows:
[0072]
[0073] Where y is the phenotypic vector, μ is the population mean, Z is the correlation matrix between individuals and random genetic effects, u is the additive genetic effect, and e is the residual term. Assume... , .
[0074] (2) Construction of the genome relationship matrix
[0075] Let M be the centered genotype matrix (subtracting from each column) Construct a diagonal weight matrix Weights are assigned to core subset sites, and non-core sites are assigned zero weights, as follows: The genome relationship matrix can then be obtained according to the following calculation rules: By setting the weights of non-core sites to 0, the model focuses on the stable genetic effects carried by core markers.
[0076] (3) Calculation of breeding value
[0077] Using the weighted relation matrix G obtained in step S4(2), a system of mixed model equations is established in combination with phenotypic data:
[0078]
[0079] In the formula, For solutions involving fixed effects such as the environment, This represents the individual breeding value (GEBV) vector to be determined. An iterative method is used to solve the above system of equations, directly extracting the estimated breeding value for each individual. The model estimates the genomic breeding value (GEBV) of individuals in the reference population, which is then used for subsequent parent selection and offspring prediction.
[0080] S5. Assessment of the accuracy of parental breeding value prediction
[0081] (1) Genotyping and breeding value acquisition of parental candidate populations
[0082] Before constructing families and evaluating the predicted effects, this embodiment also includes steps for genotyping and calculating the genomic breeding values of the parent individuals in the breeding candidate population to ensure the completeness and feasibility of the prediction process.
[0083] Specifically, parental individuals for breeding pairing are selected from the breeding candidate population, and tissue samples are collected for genomic DNA extraction. Using a genotyping platform consistent with the reference population, whole-genome single nucleotide polymorphism (SNP) typing is performed on the parental individuals to obtain corresponding genotype data. The SNP typing data of the parental individuals are preprocessed according to the same quality control and digitization coding rules as the reference population, including deletion site filtering, allele consistency correction, and additive coding conversion, to ensure consistency in genotype data structure between the candidate population and the reference population. Subsequently, the preprocessed parental genotype data is input into a genome prediction model trained based on the reference population to calculate the estimated genomic breeding value (GEBV) for each parental individual. The GEBV serves as a predictive index characterizing the additive genetic effects of the parents and is used for subsequent parental selection, family construction, and evaluation of the accuracy of parental genome breeding value predictions.
[0084] (2) Family lineage construction and virus challenge verification
[0085] To verify the technical effectiveness of the genome prediction model constructed based on MAF and LD information screening markers in cross-generational application scenarios, parental individuals in the breeding candidate population were selectively mated and bred to construct full-sib families for offspring verification.
[0086] Specifically, through pairing and breeding of parent individuals, and selective breeding of the breeding candidate population, a total of 44 full-sib families were constructed. Vibrio challenge experiments were then conducted on the offspring of these families at one age. This challenge experiment included 1245 offspring individuals. After the challenge experiment, resistance phenotypic indicators of each family's offspring were statistically analyzed, including family survival rate and average family survival time. The experimental results showed significant differences in resistance phenotypes among different families, with the average family survival rate ranging from 0% to 64.29%. These results indicate sufficient phenotypic variation in the offspring population, providing a reliable data basis for assessing the predictive accuracy of parental breeding values across generations.
[0087] (3) Calculation of the accuracy of parent breeding value prediction
[0088] To evaluate the accuracy of the prediction model built on the core marker subset in characterizing the ability of parental additive genetic effects to be transmitted to offspring, a family-level cross-generational correlation analysis method was used to evaluate the prediction accuracy.
[0089] The specific calculation process is as follows: First, based on the genotype data and resistance-related phenotypic data of the reference population, the estimated genomic breeding value (GEBV) of the parent individuals is calculated using the aforementioned genome prediction model. Then, the breeding values of the parent individuals within the same family line are averaged to obtain the predicted breeding value at the family level. The calculation method is as follows:
[0090]
[0091] Among them, GEBV dam and GEBV sire These represent the predicted breeding values for the maternal and paternal parents in the corresponding family, respectively.
[0092] Then, the family-predicted breeding value GEBV is... mid Correlation analysis was performed between the predicted breeding values and the measured resistance phenotypic indices of the corresponding offspring families. These offspring family phenotypic indices included family survival rate and average family survival time. The Pearson correlation coefficient *r* between the predicted breeding values and the offspring family phenotypic indices was calculated, and the average correlation coefficient obtained for different resistance phenotypic indices was used as an evaluation index of the predictive accuracy of the predictive model. The magnitude of the correlation coefficient *r* reflects the stability of the predictive model in capturing transferable additive genetic effects. A higher *r* value indicates a stronger predictive ability for the target trait in cross-generational applications, higher cross-generational stability of the model, and more accurate estimation of parental breeding values.
[0093] (4) Effect evaluation
[0094] Taking the anti-Vibrio trait as an example, after screening with the core markers of this invention, which feature "high frequency + strong linkage structure representativeness", the prediction accuracy of the parental genome breeding value was improved by 23.66% compared with the model constructed with all markers. Figure 3A However, uniform sampling of the same number of genomes did not improve prediction accuracy. This result indicates that the core marker subset obtained through combined MAF and LD screening can more robustly capture cross-generational genetic effects, thereby improving the cross-generational prediction stability of resistance traits.
[0095] Example 3:
[0096] Accuracy assessment of predictive value of Cryptocaryon resistance trait in large yellow croaker based on high MAF and high LD marker screening
[0097] S1, Reference Group Construction
[0098] Using the large yellow croaker (Larimichthys crocea) breeding population as the research object, reference populations and candidate breeding populations were constructed using existing breeding experimental populations to verify the applicability of the method of this invention for cross-generational prediction in different aquatic species.
[0099] The reference population consisted of two generations of experimental populations, derived from artificial infection experiments in 2018 and 2020, respectively. In each generation, approximately 600 large yellow croakers were randomly selected from the experimental populations for artificial infection with Cryptocaryon irritans. Individual survival time and status were recorded, and tissue samples were collected after infection for genotyping. After screening, 480 individuals (2018) and 276 individuals (2020) were obtained as reference populations, respectively.
[0100] The candidate breeding population was derived from the next generation (G2 generation) population, from which 381 individuals were randomly selected as breeding candidate parents. Individual markers and body size measurements were performed on these candidate individuals, and fin samples were collected for subsequent genotyping.
[0101] S2, Label Filtering and Model Building
[0102] According to the method described in Embodiment 2 of the present invention, the minor allele frequency (MAF) and linkage disequilibrium score (LD score) are calculated for each SNP site in the whole genome of large yellow croaker.
[0103] The linkage disequilibrium score is calculated using a fixed physical window method. Taking each SNP site as the center, the linkage disequilibrium determination coefficient between the target site and other sites within the window is calculated within a 100 kb genomic region upstream and downstream of it. The determination coefficients are then summed to obtain the corresponding LD score.
[0104] Subsequently, the distribution of MAF and LD scores across the entire genome was statistically analyzed using computing equipment. A joint high-quantile screening threshold was set, and SNP loci that simultaneously met the criteria of having MAF and LD scores within the top 10% quantile range of the entire genome were selected to construct a core feature subset. Figure 2B The remaining unselected sites are treated as a set of non-core markers.
[0105] S3. Core tag subset filtering based on MAF and LD
[0106] A weighted genome prediction model is constructed based on the core feature subset. During the model construction process, SNP sites belonging to the core feature subset are assigned non-zero weights, while non-core sites are assigned zero weights, so that non-core sites do not participate in the calculation during the construction of the genome relationship matrix and the estimation of genetic effects.
[0107] S4. Construction of Genome Prediction Model and Estimation of Parental Breeding Value
[0108] Using disease resistance-related phenotypic data (survival time, survival status) and genotype data of the reference population, the model parameters are estimated, and the trained prediction model is further applied to the candidate breeding population to calculate the estimated genomic breeding value (GEBV) of the candidate parent individuals.
[0109] S5, Offspring Reproduction and Virus Challenge
[0110] Based on the predicted breeding values of the candidate parents, male and female individuals were ranked and screened, and selective breeding was performed based on reproductive maturity to construct the next generation (G3 generation) population. The obtained fertilized eggs were incubated and managed under uniform conditions. After the offspring reached the preset stage, a certain number of individuals were randomly selected from each breeding population to undergo an artificial infection challenge experiment with Cryptocaryon stimuli. The survival time, survival status, and other disease resistance phenotypic indicators of the offspring individuals were recorded and statistically analyzed at the family or population level. A total of 19 families were obtained through breeding, and 7903 offspring individuals were challenged with the virus, obtaining the mean resistance phenotype of each family.
[0111] S6. Cross-generational prediction accuracy assessment
[0112] Using families as units, the combined results of parental predicted breeding values were compared and analyzed with the corresponding disease resistance phenotypic indicators of their offspring. The prediction accuracy of the prediction model based on a subset of core markers in cross-generational applications was evaluated by calculating the correlation between predicted breeding values and offspring phenotypic indicators. Results showed that the prediction model based on the high MAF and high LD core marker subsets selected by the method of this invention exhibited higher stability and predictive ability in predicting offspring phenotypic values across generations for parental breeding values of the Cryptocaryon irritans resistance trait in large yellow croaker, achieving a relative improvement of 81.35% compared to models constructed using all markers. However, uniform sampling of the same number of genomes did not improve prediction accuracy. Figure 3B This indicates that even in fish species with rapidly changing linkage disequilibrium structures and complex population structures, this method can still effectively capture the transmission of parental additive genetic effects to offspring, further demonstrating the applicability of the method in different species breeding scenarios.
[0113] In summary, the marker screening and model construction method for accurately evaluating breeding values based on parental genome estimation provided by this invention can reduce the interference of low-frequency or weakly linked sites on the estimation of genetic effects, improve the model's ability to capture transferable additive genetic effects, and thus enhance the prediction stability and accuracy in cross-generational application scenarios. Furthermore, the method of this invention can be widely applied in the prediction of complex traits and molecular breeding practices in different populations and species, demonstrating good versatility and application value.
Claims
1. A marker screening and model construction method for accurate evaluation of parental genome breeding values, characterized in that, Includes the following steps: Step 1: Obtain whole-genome SNP genotyping data and target trait phenotypic data of the reference population, and perform quality control and numerical coding on the SNP genotyping data to obtain genotype matrix X and phenotype vector y; Step 2: For each SNP locus in the genotype matrix X, calculate the minor allele frequency (MAF) and linkage disequilibrium score (LD score) of that SNP locus. Step 3: Based on the MAF and the LD score, construct a joint screening rule, and select the SNP sites that simultaneously meet the conditions according to a preset threshold or quantile criterion, which will be used as the core feature subset S. core ; Step 4: Based on the core feature subset S core Construct a genome prediction model, assign higher weights to SNP sites within the core feature subset, or construct the prediction model based solely on the core feature subset, in order to establish a mapping relationship between genotype data and predicted breeding values; Step 5: Type the candidate parent population, use the genome prediction model to obtain the predicted breeding value of the candidate parents, and select parents to construct families accordingly. Perform phenotypic determination on the offspring and calculate the phenotypic mean of the offspring in the family. Step Six: Using families as units, calculate the mean or weighted combination of the predicted breeding values of the parental individuals in the same family as the family breeding value, and calculate the correlation coefficient between the family breeding value and the mean of the corresponding family subtype, as an evaluation index of the accuracy of the parental genome breeding value prediction.
2. The method according to claim 1, characterized in that: In step one, the quality control includes removing individuals and loci with a genotype detection rate of less than 80%; the numerical coding uses an additive model to encode homozygous reference genotype as 0, heterozygous genotype as 1, and homozygous mutant genotype as 2.
3. The method according to claim 1, characterized in that, In step two, the calculation method for the linkage disequilibrium fraction parameter includes: dividing the SNP sites according to chromosomes; for each chromosome, sequentially traversing each SNP site located on that chromosome; constructing a sliding window with the SNP site as the center, within a physical distance range of 100 kb upstream and downstream; and calculating the Pearson correlation coefficient r between the target SNP site and other SNP sites within the sliding window. 2 and for the r 2 The values are summed to obtain the linkage disequilibrium fraction parameter corresponding to the SNP site.
4. The method according to claim 1, characterized in that, In step three, the minor allele frequency threshold and linkage disequilibrium fraction threshold are set based on the distribution of SNP sites across the whole genome, and the core feature subset is the set of SNP sites that are simultaneously located in the high quantile interval of the minor allele frequency distribution and the high quantile interval of the linkage disequilibrium fraction distribution.
5. The method according to claim 1, characterized in that: In step three, the high quantile interval of the minor allele frequency distribution is the top 10% quantile interval of the minor allele frequency distribution of SNP sites in the whole genome, and the high quantile interval of the linkage disequilibrium fraction distribution is the top 10% quantile interval of the linkage disequilibrium fraction distribution of SNP sites in the whole genome.
6. The method according to claim 1, characterized in that, In step four, the genome prediction model is a linear model or a nonlinear model, including one or more of linear mixed models, random effects models, and machine learning models, or equivalent variations thereof.
7. The method according to claim 1, characterized in that: In step five, the target trait for sub-representative phenotype measurement is the same as that of the reference population, and the measurement method, evaluation criteria, and experimental conditions are consistent with those of the reference population phenotypic vector y to obtain sub-representative phenotype data.
8. The method according to claim 1, characterized in that: In step six, the correlation coefficient is calculated as follows: ; in, This represents the breeding value of the i-th family. This represents the mean of the breeding values of all families. Let represent the sub-representative mean of the i-th family lineage. denoted as the mean of the subtypes of all families, and N represents the number of families.
Citation Information
Cited By
A method for screening hybrid parent combinations based on genomic functional site information
CN122337325A