A method for integrated hybrid information for purebred genome prediction

By constructing additive and dominant effect partial kinship correlation matrices and integrating the genetic information of hybrid populations, the problem of inaccurate prediction of purebred breeding values ​​in existing technologies has been solved. This has enabled precise quantification of genetic associations between purebred and hybrid populations, improving breeding efficiency and the production performance of end-product groups.

CN120564826BActive Publication Date: 2026-04-14INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSTITUTE OF ANIMAL SCIENCES OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES
Filing Date
2025-07-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate the genetic information of hybrid populations and cannot accurately distinguish the origin of alleles in breeds, resulting in insufficient accuracy in predicting purebred breeding values, particularly in traits that are difficult to measure directly, such as meat quality and the lifetime fertility of female animals.

Method used

By constructing additive effect partial kinship correlation matrix (GBOA model) and dominant effect partial kinship correlation matrix (GDBOA model), the genetic information of hybrid populations is integrated, the variety origin of alleles is clarified, the additive and dominant effect associations between purebreds and hybrid populations are quantified, and a three-trait mixed linear model is constructed to achieve accurate prediction of purebred genome breeding values.

Benefits of technology

It significantly improves the predictive accuracy of purebred genome breeding values, especially for traits that are difficult to measure directly, expands the applicability of genetic assessment, and maintains high predictive reliability in different breeding practices, thus promoting the improvement of purebred genetic improvement and the performance of terminal hybrid commercial populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564826B_ABST
    Figure CN120564826B_ABST
Patent Text Reader

Abstract

The application discloses a kind of integrated hybrid information purebred genome prediction method, belong to animal genetic breeding and propagation technical field.The population genetic dataset comprising purebred population A, B and hybrid population CB is constructed, and the population genetic differentiation is verified by principal component analysis and genetic differentiation index;After the genotypes of purebred and hybrid population CB are phased, the ancestral information inference method is used to determine the origin information of each allele and process missing sites;Based on the origin information of allele, additive effect partial kinship correlation matrix and dominant effect partial kinship correlation matrix are constructed;The above matrix is integrated into three-trait mixed linear model, forming the prediction model that simultaneously incorporates purebred and hybrid population genetic information;The application distinguishes the breed origin of allele and quantifies its additive and dominant effect, significantly improves the prediction accuracy of purebred genome breeding value, is suitable for the traits difficult to directly determine in purebred population, and provides reliable technical support for animal breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of animal genetics, breeding and reproduction technology, and specifically relates to a purebred genome prediction method that integrates hybridization information. Background Technology

[0002] In animal breeding, genetic improvement of purebred populations is a core element in enhancing the production performance of terminal hybrid commercial herds. Currently, traditional genetic assessment methods mostly focus on performance selection of purebred core populations, primarily relying on phenotypic and pedigree information of purebred populations to estimate genetic parameters and predict breeding values. However, the performance of terminal hybrid commercial herds directly reflects their breeding value, and relying solely on purebred information is insufficient to accurately reflect their actual performance in a hybrid context. This is especially true for traits in purebred populations that are difficult to measure (such as meat quality and lifetime reproductive capacity of female animals), where the accuracy of traditional methods is limited.

[0003] While the application of genomic selection technology has significantly improved the efficiency of breeding value prediction, existing models (such as MT-BLUP) often fail to fully consider the integrated value of hybrid population information. The genome of a hybrid population consists of alleles from different purebred parents. The same allele can produce differential effects due to different varietal origins, and conventional models do not distinguish the varietal origin of alleles, making it difficult to accurately quantify the genetic associations between purebred and hybrid populations. Furthermore, existing methods still have shortcomings in handling the synergistic effects of additive and dominant effects, and in how to utilize hybrid phenotypic information to optimize purebred genetic assessment, resulting in room for improvement in the accuracy of purebred breeding value prediction.

[0004] Therefore, how to effectively integrate the genetic information of hybrid populations and construct a predictive model that can distinguish the origin of allele varieties and quantify their effects has become a key issue in improving the accuracy of purebred genome prediction and promoting breeding efficiency. Summary of the Invention

[0005] In view of this, the purpose of this invention is to overcome the shortcomings of existing related technologies and provide a purebred genome prediction method that integrates hybridization information. This method aims to effectively integrate the genetic information of hybrid populations, clearly distinguish the varietal origin of alleles in hybrid individuals, and construct a genetic model that can quantify the additive and dominant effects between purebred and hybrid populations. This improves the accuracy of purebred population genome breeding value prediction, especially for traits in purebred populations that are difficult to measure directly (such as female animal lifespan, meat quality, lifetime fertility, etc.), enabling more precise genetic assessment. This provides more reliable technical support for the genetic improvement of purebred populations in animal breeding, ultimately improving the production performance and breeding efficiency of the final hybrid commercial population.

[0006] To achieve the above objectives, this invention provides a method for predicting purebred genomes by integrating hybridization information, comprising the following steps:

[0007] Step S1: Constructing a population genetic dataset

[0008] We simulated purebred populations A and B, and a hybrid population CB with significant genetic differences, and obtained genomic markers, quantitative trait loci, and phenotypic data for the three populations. The phenotypic data covered traits with different combinations of narrow-sense heritability and dominance variance ratio plus variance. We verified the genetic differentiation between purebred populations A and B through principal component analysis and genetic differentiation index.

[0009] Step S2: Tracing the origin of hybrid alleles

[0010] After haplotype analysis and phasing of the genotypes of the purebred and hybrid populations CB, the ancestry information inference method was used to determine the origin information of each allele from the purebred population A or B. Unassigned and conflicting sites were deleted and the allele origin frequencies were recalculated.

[0011] Step S3: Construct a partial kinship correlation matrix

[0012] Based on allele origin information, an additive effect partial kinship correlation matrix was constructed to quantify additive genetic associations between individuals in purebred and hybrid populations; and a dominant effect partial kinship correlation matrix was constructed to quantify dominant effect associations between individuals in purebred and hybrid populations.

[0013] Step S4: Establish an integrated prediction model

[0014] The partial kinship correlation matrix was integrated into the three-trait mixed linear model to form a predictive model that simultaneously incorporates the genetic information of purebred populations A and B and hybrid population CB, including the GBOA model with only additive effects and the GDBOA model with synergistic additive-dominant effects;

[0015] Step S5: Predict purebred breeding value based on different scenarios

[0016] Scenario 1: Parental information of hybrid individuals: Input phenotypic data, genomic data and partial kinship correlation matrix of purebred population A, B and hybrid population CB, and predict the estimated breeding value of the purebred population through the model;

[0017] Scenario 2: Parental information without hybrid individuals: Input phenotypic and genomic data of purebred populations A and B, and hybrid population CB, and use the model to predict the genomic estimated breeding value of the purebred population.

[0018] Furthermore, in step S1, the phenotypic data consists of 9 traits, corresponding to narrow-sense heritability (h²) of 0.1, 0.3, and 0.5, and dominant variance to additive variance of 0, 0.25, and 0.5, respectively; the genomic data includes 18 chromosomes, each containing 10,000 SNPs and 500 QTLs, with a mutation rate of 2.5 × 10⁻⁶. -5Secondary allele frequency > 0.05; principal component analysis (PCA) and a fixed coefficient Fst were used to verify population genetic differentiation. An Fst value of 0.1302 was used to confirm moderate genetic differentiation. The formula for calculating Fst is as follows:

[0019] Where HS represents the average heterozygosity in the subpopulation and HT represents the average heterozygosity in the total population.

[0020] Furthermore, the deletion processing in step S2 includes: setting unassigned origin sites as deletions, setting minor allele origin sites in haplotypes as deletions; filling deletion values ​​in additive effect models with allele frequencies at the site, and setting deletion values ​​to 0 in dominant effect models.

[0021] Furthermore, in step S3, the additive effect partial kinship correlation matrix is ​​used to construct the corresponding GBOA model, which is implemented as follows:

[0022] Relationship between phenotype and additive inheritance effects:

[0023]

[0024]

[0025]

[0026] in, , and A vector of phenotypic observations; , and This is a fixed effects vector that includes variety effects. , and Design matrix associated with fixed effects; and It is the vector of additive genetic effects in a purebred; and It is the vector of additive genetic effects of purebred gametes in hybrid individuals; , and It is a design matrix associated with additive effects; , and This is the vector of random residual effects, and the variance-covariance matrix of the random residual effects:

[0027]

[0028] The variance-covariance of the additive genetic effects originating from variety A is:

[0029]

[0030] in, This is an additive genetic effect of variety A. This is an additive genetic effect of the parents in hybrids. It is the vector of additive genetic effects of purebred gametes in hybrid individuals. It is an artificially created random vector, which was added to enable the use of the Kronecker product.

[0031] Furthermore, the additive effect partial kinship correlation matrix The matrix representing the relationships between purebreds and hybrids, constructed from alleles of variety A, comprises four parts: For variety A, and Used for Variety A and hybrids Used for hybrids;

[0032] The additive effect partial kinship correlation matrix is ​​constructed by dividing the additive partial relation matrix based on the label, using the following formula:

[0033]

[0034] Each block matrix is ​​calculated using the following formula:

[0035] Purebred A additive genetic association matrix:

[0036]

[0037]

[0038] Additive genetic association matrix between purebred A and hybrid population CB:

[0039] ,

[0040]

[0041] Additive genetic association matrix of hybrid populations (CB):

[0042] ,

[0043]

[0044] In the formula, This is based on the marker allele matrix derived from A in purebred and hybrid animals; It is the allele frequency from breed A, which is based on the marker genotype of purebred animals and the specific marker alleles from breed A in hybrid animals;

[0045] It is the identity matrix. The number of marked sites; "sym" indicates the symmetric part of the matrix.

[0046] The molecular scaling coefficients were transformed. This represents the complete chromosome variance of variety A. This represents the covariance of breed A and hybrid animal CB. Represents the haplotype variance of hybrid animal CB; covariance between A and CB, and haplotype variance of CB.

[0047] The matrix is ​​constructed as follows:

[0048] For purebred A:

[0049]

[0050] For hybrid animals, CB is:

[0051] .

[0052] Furthermore, in step S3, the dominant effect partial kinship correlation matrix is ​​used to construct the corresponding GDBOA model, which is implemented as follows:

[0053] The relationship between phenotype and additive-dominant inheritance effects:

[0054]

[0055]

[0056]

[0057] in, , and This is a vector of phenotypic observations. , and This is a fixed effects vector that includes variety effects. , and It is a design matrix associated with fixed effects; and It is the vector of additive genetic effects in a purebred. and It is the vector of additive genetic effects of purebred gametes in hybrid individuals. , and It is the corresponding correlation matrix; and It is the vector of dominant genetic effects in a purebred. and It is the dominant genetic effect vector of purebred gametes in hybrid individuals. , and It is the corresponding correlation matrix; , and It is the vector of random residual effects; the variance-covariance of the dominant genetic effect originating from variety A is:

[0058]

[0059] This is a partial kinship correlation matrix for dominant effects;

[0060] The construction method of the variance-covariance of the dominant genetic effect originating from variety B is similar to that of the variance-covariance of the dominant genetic effect originating from variety A.

[0061] Furthermore, the dominant partial relation matrix originating from purebred A Build as:

[0062]

[0063] The block matrix and scaling coefficients are calculated using the following formulas:

[0064] Purebred A endogenous dominant inheritance association matrix:

[0065]

[0066]

[0067] Dominant genetic association matrix between purebred A and hybrid population CB:

[0068]

[0069]

[0070] Dominant genetic association matrix within hybrid populations (CB):

[0071]

[0072]

[0073] in, It is the frequency of a specific allele from breed A, which is based on the marker genotype of purebred animals and the specific marker allele from breed A in hybrid animals; It is the frequency of specific alleles from breed B, based on the marker genotype of purebred animals and the specific marker gene from breed B in hybrid animals.

[0074] The matrix is ​​constructed as follows:

[0075] For purebred A, the genotype corresponds to the dominant deviation:

[0076]

[0077] For hybrid animals (CB), the genotype corresponds to the dominant deviation:

[0078]

[0079] In the formula: For specific allele frequencies of variety A, i.e., markers based on purebred and hybrid origin A;

[0080] For specific allele frequencies of variety B; ; The number of marker sites; "sym" indicates the symmetric part of the matrix.

[0081] Furthermore, the dominant-subject relationship matrix originating from variety B. Build as:

[0082]

[0083] Each block matrix and scaling coefficient is symmetric and is calculated using the following formula:

[0084] Intra-dominant genetic association matrix for variety B:

[0085]

[0086]

[0087] Dominant genetic association matrix between variety B and hybrid population CB:

[0088]

[0089]

[0090] Dominant genetic association matrix within hybrid populations (CB):

[0091]

[0092]

[0093] in, It is the frequency of a specific allele from breed A, which is based on the marker genotype of purebred animals and the specific marker allele from breed A in hybrid animals; It is the frequency of specific alleles from breed B, which is based on the marker genotype of purebred animals and the specific marker genes from breed B in hybrid animals.

[0094] The matrix is ​​constructed as follows:

[0095] For variety B, the genotype corresponds to the dominant deviation:

[0096]

[0097] For hybrid animals (CB), the genotype corresponds to the dominant deviation:

[0098]

[0099] In the formula: For specific allele frequencies of variety B, based on markers of purebred and hybrid origin B;

[0100] For the specific allele frequencies of variety A; "sym" indicates the symmetric part of a matrix, and... The rules are consistent.

[0101] The present invention, by adopting the above technical solution, has at least the following beneficial effects:

[0102] In this invention, by tracing the varietal origin of alleles in hybrid populations (from purebred A or B), additive effect partial kinship correlation matrices (GBOA model) and dominant effect partial kinship correlation matrices (GDBOA model) are constructed and integrated into a three-trait mixed linear model, achieving for the first time the varietal-specific quantification of genetic associations between purebred and hybrid populations. Compared with traditional MT-BLUP models that only utilize purebred information or do not distinguish allele origins, this invention can accurately capture the differences in additive and dominant effects of alleles from different varietal origins, significantly improving the prediction accuracy of purebred genomic breeding values.

[0103] In this invention, by performing breed-specific transformations on the scaling coefficients of the additive effect matrix (distinguishing between purebred complete chromosome variance, purebred and hybrid covariance, and hybrid haplotype variance), and by classifying dominant variances according to allele origin (e.g., calculating dominant variances originating from A separately in purebred A and hybrid population CB), a refined analysis of genetic effects is achieved. This design allows the model to be adapted to traits with different narrow-sense heritability (0.1, 0.3, 0.5) and dominant variance ratios (0, 0.25, 0.5), especially for traits that are difficult to measure directly in purebred populations (such as female animal lifespan, meat quality, lifetime fertility, etc.), which can be indirectly assessed through hybrid information, thus expanding the applicability of genetic assessment.

[0104] This invention enhances the applicability of the method in different breeding practices by designing prediction models for different scenarios (including / excluding information on hybrid parent individuals). In high-relatedness scenarios with parental information, the model can fully utilize kinship to improve accuracy; in low-relatedness scenarios without parental information, it can still maintain high prediction reliability through allele origin information. Simultaneously, the synergistic design of the GBOA model (additive only) and the GDBOA model (additive and dominant) allows for flexible selection based on trait inheritance mechanisms, ensuring both computational efficiency for simple traits and prediction accuracy for complex traits. This provides efficient and accurate technical support for breeding across the entire industry chain, promoting the synergistic optimization of purebred genetic improvement and performance enhancement of terminal hybrid commercial populations. Attached Figure Description

[0105] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0106] Figure 1 This is a flowchart of the purebred genome prediction method based on hybridization information of the present invention;

[0107] Figure 2 This is a schematic diagram of LD analysis between groups in this invention;

[0108] Figure 3 This is a schematic diagram of PCA analysis between populations in this invention;

[0109] Figure 4 This is a schematic diagram of the average invalid sites for each haplotype of each chromosome in this invention;

[0110] Figure 5 This is a schematic diagram of the average allocation of error sites for each haplotype of each chromosome in this invention;

[0111] Figure 6 This is a schematic diagram illustrating the prediction accuracy of each model in this invention. Figure 1 ;

[0112] Figure 7 This is a schematic diagram of the prediction reliability of each model in this invention. Figure 1 ;

[0113] Figure 8 This is a schematic diagram illustrating the prediction accuracy of each model in this invention. Figure 2 ;

[0114] Figure 9 This is a schematic diagram of the prediction reliability of each model in this invention. Figure 2 ;

[0115] Figure 10 This is a schematic diagram illustrating the prediction accuracy of each model in this invention. Figure 3 ;

[0116] Figure 11 This is a schematic diagram of the prediction reliability of each model in this invention. Figure 3 ;

[0117] Figure 12 This is a schematic diagram illustrating the prediction accuracy of each model in this invention. Figure 4 ;

[0118] Figure 13 This is a schematic diagram of the prediction reliability of each model in this invention. Figure 4 . Detailed Implementation

[0119] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0120] like Figures 1 to 13 As shown, this embodiment provides a method for predicting purebred genomes by integrating hybridization information, including the following steps:

[0121] Step S1: Constructing a population genetic dataset

[0122] We simulated purebred populations A and B, and a hybrid population CB with significant genetic differences, and obtained genomic markers, quantitative trait loci, and phenotypic data for the three populations. The phenotypic data covered traits with different combinations of narrow-sense heritability and dominance variance ratio plus variance. We verified the genetic differentiation between purebred populations A and B through principal component analysis and genetic differentiation index.

[0123] Step S2: Tracing the origin of hybrid alleles

[0124] After haplotype analysis and phasing of the genotypes of the purebred and hybrid populations CB, the ancestry information inference method was used to determine the origin information of each allele from the purebred population A or B. Unassigned and conflicting sites were deleted and the allele origin frequencies were recalculated.

[0125] Step S3: Construct a partial kinship correlation matrix

[0126] Based on allele origin information, an additive effect partial kinship correlation matrix was constructed to quantify additive genetic associations between individuals in purebred and hybrid populations; and a dominant effect partial kinship correlation matrix was constructed to quantify dominant effect associations between individuals in purebred and hybrid populations.

[0127] Step S4: Establish an integrated prediction model

[0128] The partial kinship correlation matrix was integrated into the three-trait mixed linear model to form a predictive model that simultaneously incorporates the genetic information of purebred populations A and B and hybrid population CB, including the GBOA model with only additive effects and the GDBOA model with synergistic additive-dominant effects;

[0129] Step S5: Predict purebred breeding value based on different scenarios

[0130] Scenario 1: Parental information of hybrid individuals: Input phenotypic data, genomic data and partial kinship correlation matrix of purebred population A, B and hybrid population CB, and predict the estimated breeding value of the purebred population through the model;

[0131] Scenario 2: Parental information without hybrid individuals: Input phenotypic and genomic data of purebred populations A and B, and hybrid population CB, and use the model to predict the genomic estimated breeding value of the purebred population.

[0132] In one implementation method, in step S1 of this embodiment, the phenotypic data consists of 9 traits, corresponding to narrow-sense heritability (h²) of 0.1, 0.3, and 0.5, and dominant variance to additive variance of 0, 0.25, and 0.5, respectively; the genomic data includes 18 chromosomes, each chromosome containing 10,000 SNPs and 500 QTLs, with a mutation rate of 2.5 × 10⁻⁶. -5 Secondary allele frequency > 0.05; principal component analysis (PCA) and a fixed coefficient Fst were used to verify population genetic differentiation. An Fst value of 0.1302 was used to confirm moderate genetic differentiation. The formula for calculating Fst is as follows:

[0133] Where HS represents the average heterozygosity in the subpopulation and HT represents the average heterozygosity in the total population.

[0134] As one implementation method, the deletion processing in step S2 of this embodiment includes: setting unassigned origin sites as deletions, setting minor allele origin sites in haplotypes as deletions; filling the deletion value in the additive effect model with the allele frequency of that site, and setting the deletion value in the dominant effect model to 0.

[0135] Furthermore, in step S3, the additive effect partial kinship correlation matrix is ​​used to construct the corresponding GBOA model, which is implemented as follows:

[0136] Relationship between phenotype and additive inheritance effects:

[0137]

[0138]

[0139]

[0140] in, , and A vector of phenotypic observations; , and This is a fixed effects vector that includes variety effects. , and Design matrix associated with fixed effects; and It is the vector of additive genetic effects in a purebred; and It is the vector of additive genetic effects of purebred gametes in hybrid individuals; , and It is a design matrix associated with additive effects; , and This is the vector of random residual effects, and the variance-covariance matrix of the random residual effects:

[0141]

[0142] The variance-covariance of the additive genetic effects originating from variety A is:

[0143]

[0144] in, This is an additive genetic effect of variety A. This is an additive genetic effect of the parents in hybrids. It is the vector of additive genetic effects of purebred gametes in hybrid individuals. It is an artificially created random vector, which was added to enable the use of the Kronecker product.

[0145] As one implementation method, this embodiment uses an additive effect partial kinship correlation matrix. The matrix representing the relationships between purebreds and hybrids, constructed from alleles of variety A, comprises four parts: For variety A, and Used for Variety A and hybrids Used for hybrids;

[0146] The additive effect partial kinship correlation matrix is ​​constructed by dividing the additive partial relation matrix based on the label, using the following formula:

[0147]

[0148] Each block matrix is ​​calculated using the following formula:

[0149] Purebred A additive genetic association matrix:

[0150]

[0151]

[0152] Additive genetic association matrix between purebred A and hybrid population CB:

[0153] ,

[0154]

[0155] Additive genetic association matrix of hybrid populations (CB):

[0156] ,

[0157]

[0158] In the formula, This is based on the marker allele matrix derived from A in purebred and hybrid animals; It is the allele frequency from breed A, which is based on the marker genotype of purebred animals and the specific marker alleles from breed A in hybrid animals;

[0159] It is the identity matrix. The number of marked sites; "sym" indicates the symmetric part of the matrix.

[0160] The molecular scaling coefficients were transformed. This represents the complete chromosome variance of variety A. This represents the covariance of breed A and hybrid animal CB. Represents the haplotype variance of hybrid animal CB; covariance between A and CB, and haplotype variance of CB.

[0161] The matrix is ​​constructed as follows:

[0162] For purebred A:

[0163]

[0164] For hybrid animals, CB is:

[0165] .

[0166] As one implementation method, in step S3 of this embodiment, the dominant effect partial kinship correlation matrix is ​​used to construct the corresponding GDBOA model, which is implemented as follows:

[0167] The relationship between phenotype and additive-dominant inheritance effects:

[0168]

[0169]

[0170]

[0171] in, , and This is a vector of phenotypic observations. , and This is a fixed effects vector that includes variety effects. , and It is a design matrix associated with fixed effects; and It is the vector of additive genetic effects in a purebred. and It is the vector of additive genetic effects of purebred gametes in hybrid individuals. , and It is the corresponding correlation matrix; and It is the vector of dominant genetic effects in a purebred. and It is the dominant genetic effect vector of purebred gametes in hybrid individuals. , and It is the corresponding correlation matrix; , and It is the vector of random residual effects; the variance-covariance of the dominant genetic effect originating from variety A is:

[0172]

[0173] This is a partial kinship correlation matrix for dominant effects; the construction method of the variance-covariance of dominant genetic effects originating from variety B is similar to that of the variance-covariance of dominant genetic effects originating from variety A.

[0174] As one implementation method, this embodiment originates from the dominant partial relation matrix of purebred A. Build as:

[0175]

[0176] The block matrix and scaling coefficients are calculated using the following formulas:

[0177] Purebred A endogenous dominant inheritance association matrix:

[0178]

[0179]

[0180] Dominant genetic association matrix between purebred A and hybrid population CB:

[0181]

[0182]

[0183] Dominant genetic association matrix within hybrid populations (CB):

[0184]

[0185]

[0186] in, It is the frequency of a specific allele from breed A, which is based on the marker genotype of purebred animals and the specific marker allele from breed A in hybrid animals; It is the frequency of specific alleles from breed B, based on the marker genotype of purebred animals and the specific marker gene from breed B in hybrid animals.

[0187] The matrix is ​​constructed as follows:

[0188] For purebred A, the genotype corresponds to the dominant deviation:

[0189]

[0190] For hybrid animals (CB), the genotype corresponds to the dominant deviation:

[0191]

[0192] In the formula: For specific allele frequencies of variety A, i.e., markers based on purebred and hybrid origin A;

[0193] For specific allele frequencies of variety B; ; The number of marker sites; "sym" indicates the symmetric part of the matrix.

[0194] As one implementation method, this embodiment originates from the dominant partial relation matrix of variety B. Build as:

[0195]

[0196] Each block matrix and scaling coefficient is symmetric and is calculated using the following formula:

[0197] Intra-dominant genetic association matrix for variety B:

[0198]

[0199]

[0200] Dominant genetic association matrix between variety B and hybrid population CB:

[0201]

[0202]

[0203] Dominant genetic association matrix within hybrid populations (CB):

[0204]

[0205]

[0206] in, It is the frequency of a specific allele from breed A, which is based on the marker genotype of purebred animals and the specific marker allele from breed A in hybrid animals; It is the frequency of specific alleles from breed B, which is based on the marker genotype of purebred animals and the specific marker genes from breed B in hybrid animals.

[0207] The matrix is ​​constructed as follows:

[0208] For variety B, the genotype corresponds to the dominant deviation:

[0209]

[0210] For hybrid animals (CB), the genotype corresponds to the dominant deviation:

[0211]

[0212] In the formula: For specific allele frequencies of variety B, based on markers of purebred and hybrid origin B;

[0213] For the specific allele frequencies of variety A; "sym" indicates the symmetric part of a matrix, and... The rules are consistent.

[0214] This embodiment details the complete process of a purebred genome prediction method based on hybridization information. The core of this method lies in tracing the varietal origin of alleles in hybrid populations, constructing a varietal-specific additive and dominant partial kinship correlation matrix, and integrating it into a three-trait mixed linear model to achieve precise quantification of the genetic association between purebred and hybrid populations.

[0215] Specifically, by simulating purebred populations A and B with significant differences in genetic background, and a hybrid population CB, ancestry inference technology was used to clarify the origin information of hybrid alleles. Then, GBOA (additive effect) and GDBOA (additive-dominant effect) models were constructed, and the effectiveness of the method was verified in different scenarios (with and without hybrid parent information). Experimental results show that compared with the traditional MT-BLUP model, this method can significantly improve the prediction accuracy of purebred genome breeding values, and is particularly suitable for traits with significant dominant effects or those difficult to measure directly in purebred populations (such as meat quality and female fertility).

[0216] This embodiment verifies the adaptability of the method under different heritability (0.1, 0.3, 0.5) and dominance variance ratios (0, 0.25, 0.5). Its core innovation lies in the refined analysis of genetic effects through the construction of a variety-specific matrix and effect classification. It provides an operable technical solution for integrating hybrid information to optimize purebred genetic evaluation and has practical application value for promoting the synergistic improvement of purebred improvement and terminal hybrid commercial population performance in animal breeding.

[0217] This embodiment analyzes nine traits from three populations (purebred A, purebred B, and hybrid CB). Population A contains 1,698 individuals, population B contains 1,989 individuals, and population CB contains 2,000 individuals. Descriptive statistical results of the phenotypic data are shown in Table 1. Overall, all phenotypic data are relatively concentrated and have low dispersion, providing scientifically reliable phenotypic values ​​for further analysis.

[0218] Table 1. Descriptive statistics of population phenotypes

[0219]

[0220] Table 1 (continued)

[0221]

[0222] The average kinship coefficient between groups in Scenario 1 is 0.0021, and the average kinship coefficient between groups in Scenario 2 is 0.0005.

[0223] In this embodiment, the genetic correlations between purebred and hybrid traits were estimated using a two-trait animal model with DMU software, as shown in Table 2. Among all traits in population A, the highest genetic correlation with all traits in the hybrid population was 0.84, the lowest was 0.50, and the average was 0.60. Among all traits in population B, the highest genetic correlation with all traits in the hybrid population was 0.99, the lowest was 0.49, and the average was 0.68. For trait 3, both populations A and B showed strong correlations with the hybrid population traits.

[0224] Table 2 Genetic correlation coefficients between purebred and hybrid populations

[0225]

[0226] Linkage disequilibrium analysis was performed on 1,698 purebred individuals A and 1,989 purebred individuals B as follows: Figure 2 As shown, the LD decay rates are similar between the two groups.

[0227] PCA analysis was performed on 1,698 purebred individuals A and 1,989 purebred individuals B. The PCA analysis is as follows: Figure 3 As shown, PC1 explained 81.03% of the variation in the data, and PC2 explained 10.66% of the variation. The samples from populations A and B showed significant differences, indicating distinctly different genetic backgrounds.

[0228] Fst analysis was performed on 1,698 purebred individuals A and 1,989 purebred individuals B. In practical studies, Fst values ​​between 0 and 0.05 indicate very little genetic differentiation between populations; values ​​between 0.05 and 0.15 indicate moderate genetic differentiation; values ​​between 0.15 and 0.25 indicate significant genetic differentiation; and values ​​above 0.25 indicate substantial genetic differentiation. In this study, the Fst value obtained by genepop analysis using the R software package was 0.1302, indicating moderate genetic differentiation between populations A and B.

[0229] In Scenario 1 and Scenario 2, the allele origins of 2,000 hybrid pigs were traced using AllOr and RFMix software, respectively. This described the average number of invalid sites per haplotype on each chromosome, the proportion of invalid sites per haplotype, and the specific number of average misallocation sites per haplotype. Invalid sites here include both unallocated and misallocated sites. The average number of invalid sites per haplotype on each chromosome is shown below. Figure 4 As shown, the AllOr software has the fewest invalid sites in scenario one. The average number of incorrect sites assigned to each haplotype on each chromosome is as follows: Figure 5 As shown, AllOr software had the fewest misallocation sites in Scenario 1. At the individual level, in Scenario 1, using AllOr software, each individual had an average of 1,065 invalid sites and an average of 0.81 misallocation sites. In Scenario 2, each individual had an average of 1,384 invalid sites and an average of 7.89 misallocation sites. In Scenario 1, using RFMix software, each individual had an average of 1,494 invalid sites and an average of 15.97 misallocation sites. In Scenario 2, each individual had 1,514 invalid sites and an average of 15.42 misallocation sites. At the individual level, AllOr software was more accurate than RFMix software in tracking alleles, and the closeness of kinship between populations had a greater impact on the accuracy of AllOr software in tracking alleles than on RFMix software.

[0230] In this embodiment, the models are first divided into: additive models containing additive effects and residual effects; and additive-dominant models containing additive effects, dominant effects, and residual effects. Conventional MT-GBLUP and BOA models are used in both the additive and additive-dominant models. The BOA model is further divided into three types based on allele origin: BOA model (true allele origin), BOA-Ⅰ model (AllOr traces allele origin), and BOA-Ⅱ model (RFMix traces allele origin). Purebred prediction is performed in scenarios one and two, respectively.

[0231] Prediction results under different additive models in scenario 1

[0232] Purebredity predictions for nine traits in the simulated population of Scenario 1 were performed using the MT-GBLUP, GBOA, GBOA-Ⅰ, and GBOA-Ⅱ methods. The accuracy and reliability of the estimated genomic breeding values ​​are evaluated in Table 3. Figure 6 and Figure 7 As shown in the figure. Comparing the GBOA, GBOA-I, and GBOA-II methods with the conventional MT-GBLUP, in trait 1, the prediction accuracy improved by 0.1766, decreased by 0.1479, and improved by 0.0502, respectively; the prediction reliability improved by 0.1471, decreased by 0.0773, and improved by 0.0368, respectively. In trait 2, the prediction accuracy improved by 0.1674, decreased by 0.0648, and improved by 0.0291, respectively; the prediction reliability improved by 0.1237, decreased by 0.0382, and improved by 0.0144, respectively. In trait 3, the prediction accuracy improved by 0.1044, decreased by 0.1150, and improved by 0.0130, respectively; the prediction reliability improved by 0.0852, decreased by 0.0631, and improved by 0.0077, respectively. In trait 4, prediction accuracy improved by 0.2929, 0.1745, and 0.1991, respectively, and prediction reliability improved by 0.2590, 0.1300, and 0.1543, respectively. In trait 5, prediction accuracy improved by 0.0757, 0.0563, and 0.0462, respectively, and prediction reliability improved by 0.0828, 0.0595, and 0.0471, respectively. In trait 6, prediction accuracy improved by 0.1219, 0.0508, and 0.0968, respectively, and prediction reliability improved by 0.1343, 0.0546, and 0.1043, respectively. In trait 7, prediction accuracy improved by 0.1386, 0.0531, and 0.1022, respectively, and prediction reliability improved by 0.1644, 0.0505, and 0.1146, respectively. In trait 8, prediction accuracy improved by 0.0459, 0.0199, and decreased by 0.0610, respectively; prediction reliability improved by 0.0559, 0.0205, and decreased by 0.0841, respectively. In trait 9, prediction accuracy improved by 0.0505, decreased by 0.0178, and improved by 0.0223, respectively; prediction reliability increased by 0.0682, decreased by 0.0240, and improved by 0.0294, respectively.

[0233] Table 3. Prediction results for Scenario 1 under different additive models.

[0234]

[0235] Table 3 (continued)

[0236]

[0237] Purebredity predictions for nine traits in the simulated population of Scenario 2 were performed using the MT-GBLUP, GBOA, GBOA-Ⅰ, and GBOA-Ⅱ methods. The accuracy and reliability of the estimated genomic breeding values ​​are evaluated in Table 4. Figure 8 and Figure 9 As shown in the figure. Comparing the GBOA, GBOA-Ⅰ, and GBOA-Ⅱ methods with the conventional MT-GBLUP, in trait 1, the prediction accuracy decreased by 0.1568, increased by 0.0254, and decreased by 0.1373, respectively; the prediction reliability decreased by 0.1313, increased by 0.0264, and decreased by 0.1168, respectively. In trait 2, the prediction accuracy decreased by 0.1791, increased by 0.0094, and decreased by 0.0703, respectively; the prediction reliability decreased by 0.1309, increased by 0.0098, and decreased by 0.0538, respectively. In trait 3, the prediction accuracy decreased by 0.1274, increased by 0.0169, and decreased by 0.1023, respectively; the prediction reliability decreased by 0.1084, increased by 0.0154, and decreased by 0.0920, respectively. In trait 4, prediction accuracy decreased by 0.0561, increased by 0.0211, and decreased by 0.0264, respectively; prediction reliability decreased by 0.0656, increased by 0.0259, and decreased by 0.0320, respectively. In trait 5, prediction accuracy decreased by 0.0114, 0.0068, and 0.0455, respectively; prediction accuracy decreased by 0.0153, 0.0091, and 0.0590, respectively. In trait 6, prediction accuracy decreased by 0.0452, increased by 0.0132, and increased by 0.0060, respectively; prediction reliability decreased by 0.0515, increased by 0.0166, and increased by 0.0076, respectively. In trait 7, prediction accuracy remained unchanged and decreased by 0.0108 and 0.0589, respectively; prediction reliability remained unchanged and decreased by 0.0159 and 0.0835, respectively. In trait 8, prediction accuracy decreased by 0.0338, 0.0066, and 0.0705, respectively, and prediction reliability decreased by 0.0454, 0.0100, and 0.0996, respectively. In trait 9, prediction accuracy decreased by 0.0163, 0.0089, and 0.0497, respectively, and prediction reliability decreased by 0.0235, 0.0131, and 0.0706, respectively.

[0238] Table 4. Prediction results for Scenario 2 under different additive models.

[0239]

[0240] Table 4 (continued)

[0241]

[0242] Prediction results under different additive-explicit models in scenario 1

[0243] Purebredity predictions for nine traits in the simulated population of Scenario 1 were performed using the MT-GDBLUP, GDBOA, GDBOA-Ⅰ, and GDBOA-Ⅱ methods. The accuracy and reliability of the estimated genomic breeding values ​​are evaluated in Table 5. Figure 10 and Figure 11 As shown in the figure. Comparing the GDBOA, GDBOA-Ⅰ, and GDBOA-Ⅱ methods with the conventional MT-GDBLUP, in trait 1, the prediction accuracy improved by 0.1307, decreased by 0.0239, and improved by 0.0994, respectively; the prediction reliability improved by 0.0924, decreased by 0.0179, and improved by 0.0720, respectively. In trait 2, the prediction accuracy improved by 0.1921, 0.0050, and 0.0766, respectively; the prediction reliability improved by 0.1277, decreased by 0.0027, and improved by 0.0362, respectively. In trait 3, the prediction accuracy improved by 0.0988, 0.0174, and 0.0477, respectively; and the prediction reliability improved by 0.0735, 0.0109, and 0.0317, respectively. In trait 4, prediction accuracy improved by 0.2005, decreased by 0.0715, and improved by 0.0040, respectively; prediction reliability improved by 0.2048, decreased by 0.0469, and improved by 0.0070, respectively. In trait 5, prediction accuracy improved by 0.0757, 0.0121, and 0.0827, respectively; prediction reliability improved by 0.0819, 0.0062, and 0.0898, respectively. In trait 6, prediction accuracy improved by 0.3274, 0.1552, and 0.2002, respectively; prediction reliability improved by 0.2842, 0.1146, and 0.1456, respectively. In trait 7, prediction accuracy improved by 0.0229, decreased by 0.0271, and improved by 0.0180, respectively; prediction reliability improved by 0.0327, decreased by 0.0364, and improved by 0.0257, respectively. In trait 8, prediction accuracy improved by 0.0098, decreased by 0.0044, and decreased by 0.0721, respectively; prediction reliability improved by 0.0135, and decreased by 0.0071 and decreased by 0.0942, respectively. In trait 9, prediction accuracy decreased by 0.0091, 0.0765, and decreased by 0.0725, respectively; prediction reliability decreased by 0.0136, 0.1071, and decreased by 0.1014, respectively.

[0244] Table 5. Prediction results for Scenario 1 under different explicit models.

[0245]

[0246] Table 5 (continued)

[0247]

[0248] Prediction results under different additive-explicit models in scenario 2

[0249] Purebredity predictions for nine traits in the simulated population of Scenario 2 were performed using the MT-GDBLUP, GDBOA, GDBOA-Ⅰ, and GDBOA-Ⅱ methods. The accuracy and reliability of the estimated genomic breeding values ​​are evaluated in Table 6. Figure 12 and Figure 13 As shown in the figure. Comparing the GDBOA, GDBOA-Ⅰ, and GDBOA-Ⅱ methods with the conventional MT-GDBLUP, in trait 1, the prediction accuracy decreased by 0.1428, increased by 0.0370, and decreased by 0.1357, respectively; the prediction reliability decreased by 0.1202, increased by 0.0383, and decreased by 0.1152, respectively. In trait 2, the prediction accuracy decreased by 0.1864, increased by 0.0166, and decreased by 0.0762, respectively; the prediction reliability decreased by 0.1332, increased by 0.0155, and decreased by 0.0608, respectively. In trait 3, the prediction accuracy decreased by 0.1247, 0.1004, and 0.1297, respectively; and the prediction reliability decreased by 0.1122, 0.0953, and 0.1171, respectively. In trait 4, prediction accuracy decreased by 0.1318, increased by 0.0226, and decreased by 0.0102, respectively; prediction reliability decreased by 0.1232, increased by 0.0276, and decreased by 0.018, respectively. In trait 5, prediction accuracy decreased by 0.0191, decreased by 0.0102, and decreased by 0.0474, respectively; prediction reliability decreased by 0.0252, decreased by 0.0144, and decreased by 0.0609, respectively. In trait 6, prediction accuracy increased by 0.0050, increased by 0.0304, and decreased by 0.0088, respectively; prediction reliability increased by 0.0054, increased by 0.0374, and decreased by 0.0095, respectively. In trait 7, prediction accuracy decreased by 0.0086, increased by 0.0037, and decreased by 0.0472, respectively; prediction reliability decreased by 0.0120, increased by 0.0058, and decreased by 0.0654, respectively. In trait 8, prediction accuracy decreased by 0.0739, 0.0036, and 0.0649, respectively, and prediction reliability decreased by 0.1016, 0.0057, and 0.0839, respectively. In trait 9, prediction accuracy decreased by 0.0264, 0.0060, and 0.0590, respectively, and prediction reliability decreased by 0.0371, 0.0092, and 0.0785, respectively.

[0250] Table 6. Prediction results for Scenario 2 under different explicit models.

[0251]

[0252] Table 6 (continued)

[0253]

[0254] This embodiment presents descriptive statistics on population phenotypic data, showing that all phenotypic data are relatively concentrated and have low dispersion. Genetic structure analysis was performed on populations A and B, revealing a moderate degree of genetic differentiation. Genetic correlation analysis was conducted on purebred and hybrid populations. The average genetic correlation between all traits in population A and all traits in the hybrid population was 0.60. The genetic correlation between all traits in population B and all traits in the hybrid population was 0.68. In genome selection, under different additive models in Scenario 1, the GBOA model's predictive accuracy was higher than other models, with GBOA-II outperforming GBOA-I. Under different additive-dominant models, the GDBOA model performed best, with GDBOA-II outperforming GDBOA-I. In Scenario 2, under different additive models, the GBOA-I model outperformed the GBOA-II model; under different additive-dominant models, the GDBOA-I model outperformed GDBOA-II.

[0255] Based on the genetic background of the pig population, two purebred populations, A and B, and a hybrid population CB of these two populations were simulated. The distribution of the three populations simulated nine traits including additive, dominant, and residual effects. Based on parental information or average kinship coefficient, the populations were divided into two scenarios: Scenario 1 included parental information of hybrid individuals (average kinship coefficient of 0.0021); Scenario 2 did not include parental information of hybrid individuals (average kinship coefficient of 0.0005). In Scenario 1 and Scenario 2, additive and additive-dominant models were constructed using the MT-BLUP model (true allele origin), the BOA-I model (AllOr tracing allele origin), and the BOA-II model (RFMix tracing allele origin), respectively, for genome purebred prediction. The main conclusions are as follows:

[0256] In Scenario 1, under different additive models: the GBOA model improved prediction accuracy by 0.05–0.29 compared to the MT-GBLUP model; the GBOA-II model improved accuracy by 0.03–0.20 compared to the GBOA-I model, except for traits 5 and 8. Under different additive-dominant models: the GDBOA model improved accuracy by 0.01–0.33 compared to the MT-GDBLUP model, except for trait 9; the GDBOA-II model improved accuracy by 0.004–0.12 compared to the GDBOA-I model, except for trait 8.

[0257] In Scenario 2, under different additive models: the GBOA model did not improve upon the MT-GBLUP model; the prediction accuracy of the GBOA-Ⅰ model was higher than that of the GBOA-Ⅱ model, improving by 0.1~0.16. Under different additive-dominant models: the GDBOA model improved by 0.01 on trait 6 compared to the MT-GDBLUP model; the GDBOA-Ⅰ model improved by 0.01~0.17 on all traits compared to the GDBOA-Ⅱ model.

[0258] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for predicting purebred genomes by integrating hybridization information, characterized in that: Includes the following steps: Step S1: Constructing a population genetic dataset We simulated purebred populations A and B, and a hybrid population CB with significant genetic differences, and obtained genomic markers, quantitative trait loci, and phenotypic data for the three populations. The phenotypic data covered traits with different combinations of narrow-sense heritability and dominance variance ratio plus variance. We verified the genetic differentiation between purebred populations A and B through principal component analysis and genetic differentiation index. Step S2: Tracing the origin of hybrid alleles After haplotype analysis and phasing of the genotypes of the purebred and hybrid populations CB, the ancestry information inference method was used to determine the origin information of each allele from the purebred population A or B. Unassigned and conflicting sites were deleted and the allele origin frequencies were recalculated. Step S3: Construct a partial kinship correlation matrix Based on allele origin information, an additive effect partial kinship correlation matrix was constructed to quantify the additive genetic associations between individuals in purebred and hybrid populations; And construct a dominant effect partial kinship correlation matrix to quantify the dominant effect associations between individuals in purebred and hybrid populations; Step S4: Establish an integrated prediction model The partial kinship correlation matrix is ​​integrated into the three-trait mixed linear model to form a predictive model that simultaneously incorporates the genetic information of purebred populations A and B and hybrid population CB, including the GBOA model with only additive effects and the GDBOA model with synergistic additive and dominant effects. Step S5: Predict purebred breeding value based on different scenarios Scenario 1: Parental information of hybrid individuals: Input phenotypic data, genomic data and partial kinship correlation matrix of purebred population A, B and hybrid population CB, and predict the estimated breeding value of the purebred population through the model; Scenario 2: Parental information without hybrid individuals: Input phenotypic and genomic data of purebred populations A and B, and hybrid population CB, and use the model to predict the genomic estimated breeding value of the purebred population; In step S3, the additive effect partial kinship correlation matrix is ​​used to construct the corresponding GBOA model, which is implemented as follows: Relationship between phenotype and additive inheritance effects: in, , and A vector of phenotypic observations; , and This is a fixed effects vector that includes variety effects. , and Design matrix associated with fixed effects; and It is the vector of additive genetic effects in a purebred; and It is the vector of additive genetic effects of purebred gametes in hybrid individuals; , and It is a design matrix associated with additive effects; , and This is the vector of random residual effects, and the variance-covariance matrix of the random residual effects: The variance-covariance of the additive genetic effects originating from variety A is: in, This is an additive genetic effect of variety A. This is an additive genetic effect of the parents in hybrids. It is the vector of additive genetic effects of purebred gametes in hybrid individuals. It is an artificially created random vector, which was added to enable the use of the Kronecker product.

2. The method according to claim 1, characterized in that: In step S1, the phenotypic data consists of 9 traits, corresponding to narrow-sense heritability ( The variance ratios were 0.1, 0.3, and 0.5, respectively, with dominant variances being 0, 0.25, and 0.5 compared to additive variances. The genomic data consisted of 18 chromosomes, each containing 10,000 SNPs and 500 QTLs, with a mutation rate of [missing value]. Secondary allele frequency > 0.05; principal component analysis (PCA) and a fixed coefficient Fst were used to verify population genetic differentiation. An Fst value of 0.1302 was used to confirm moderate genetic differentiation. The formula for calculating Fst is as follows: Where HS represents the average heterozygosity in the subpopulation and HT represents the average heterozygosity in the total population.

3. The method according to claim 1, characterized in that: The deletion handling in step S2 includes: setting unassigned origin sites as deletions, setting minor allele origin sites in haplotypes as deletions; filling deletion values ​​in additive effect models with allele frequencies at the site, and setting deletion values ​​to 0 in dominant effect models.

4. The method according to claim 1, characterized in that: Additive effects partial kinship correlation matrix The matrix representing the relationships between purebreds and hybrids, constructed from alleles of variety A, comprises four parts: For variety A, and Used for Variety A and hybrids Used for hybrids; The additive effect partial kinship correlation matrix is ​​constructed by dividing the additive partial relation matrix based on the label, using the following formula: Each block matrix is ​​calculated using the following formula: Purebred A additive genetic association matrix: Additive genetic association matrix between purebred A and hybrid population CB: , Additive genetic association matrix of hybrid populations (CB): , In the formula, This is based on the marker allele matrix derived from A in purebred and hybrid animals; It is the allele frequency from breed A, which is based on the marker genotype of purebred animals and the specific marker alleles from breed A in hybrid animals; It is the identity matrix. To indicate the number of marker sites, "sym" indicates the symmetric part of the matrix; The molecular scaling coefficients were transformed. This represents the complete chromosome variance of variety A. This represents the covariance of breed A and hybrid animal CB. Represents the haplotype variance of hybrid animal CB; covariance between A and CB, and haplotype variance of CB; The matrix is ​​constructed as follows: For purebred A: For hybrid animals, CB is: 。 5. The method according to claim 1, characterized in that: In step S3, the dominant effect partial kinship correlation matrix is ​​used to construct the corresponding GDBOA model, which is implemented as follows: The relationship between phenotype and additive-dominant inheritance effects: in, , and This is a vector of phenotypic observations. , and This is a fixed effects vector that includes variety effects. , and It is a design matrix associated with fixed effects; and It is the vector of additive genetic effects in a purebred. and It is the vector of additive genetic effects of purebred gametes in hybrid individuals. , and It is the corresponding correlation matrix; and It is the vector of dominant genetic effects in a purebred. and It is the dominant genetic effect vector of purebred gametes in hybrid individuals. , and It is the corresponding correlation matrix; , and It is the vector of random residual effects; the variance-covariance of the dominant genetic effect originating from variety A is: This is a partial kinship correlation matrix for dominant effects; The construction method of the variance-covariance of the dominant genetic effect originating from variety B is similar to that of the variance-covariance of the dominant genetic effect originating from variety A.

6. The method according to claim 5, characterized in that: Dominant Partial Relationship Matrix Originating from Purebred A Build as: The block matrix and scaling coefficients are calculated using the following formulas: Purebred A endogenous dominant inheritance association matrix: Dominant genetic association matrix between purebred A and hybrid population CB: Dominant genetic association matrix within hybrid populations (CB): in, It is the frequency of a specific allele from breed A, which is based on the marker genotype of purebred animals and the specific marker allele from breed A in hybrid animals; It is the frequency of specific alleles from breed B, based on the marker genotype of purebred animals and the specific marker gene from breed B in hybrid animals.

7. The method according to claim 6, characterized in that: The matrix is ​​constructed as follows: For purebred A, the genotype corresponds to the dominant deviation: For hybrid animals (CB), the genotype corresponds to the dominant deviation: In the formula: For specific allele frequencies of variety A, i.e., markers based on purebred and hybrid origin A; For specific allele frequencies of variety B; ; "sym" indicates the number of marker sites; "sym" indicates the symmetric part of the matrix.

8. The method according to claim 7, characterized in that: The dominant-subject relationship matrix originating from variety B Build as: Each block matrix and scaling coefficient is symmetric and is calculated using the following formula: Intra-dominant genetic association matrix for variety B: Dominant genetic association matrix between variety B and hybrid population CB: Dominant genetic association matrix within hybrid populations (CB): in, It is the frequency of a specific allele from breed A, which is based on the marker genotype of purebred animals and the specific marker allele from breed A in hybrid animals; It is the frequency of specific alleles from breed B, which is based on the marker genotype of purebred animals and the specific marker genes from breed B in hybrid animals.

9. The method according to claim 8, characterized in that: The matrix is ​​constructed as follows: For variety B, the genotype corresponds to the dominant deviation: For hybrid animals (CB), the genotype corresponds to the dominant deviation: In the formula: For specific allele frequencies of variety B, based on markers of purebred and hybrid origin B; For the specific allele frequencies of variety A; "sym" indicates the symmetric part of a matrix, and... The rules are consistent.

Citation Information

Patent Citations

  • Monitoring bacterial contamination of a wound involving assay of adenosine triphosphate

    GB2340235A

  • Nozzle for soldering apparatus

    GB2360237A

  • Breeding information processing method and device

    CN118038965A